Serverless Chats: Recent Episodes

Jeremy Daly & Rebecca Marshburn

Serverless Chats is a podcast that geeks out on everything serverless. Join Jeremy Daly and Rebecca Marshburn as they chat with a special guest each week.

View Details

About Michael Hart

A software engineering leader with 20 years of experience growing teams and building distributed systems, from fullstack development to machine learning and big data analytics. He also contributes to open source tools with hundreds of millions of downloads per month, primarily around API integrations, team productivity, and developer optimization for cloud environments. Technologies include: Go, Node.js, Python, Java and C#; Docker; AWS Lambda, DynamoDB, Kinesis. If you've ever used the AWS SAM CLI, or the request npm module, then you've installed his work. He is currently a Principal Engineer for Cloudflare Workers at Cloudflare.

  • Twitter: @hichaelmart
  • Github: https://github.com/mhart
  • Medium: https://medium.com/@hichaelmart
  • Cloudflare Workers: https://workers.cloudflare.com/

View Details

About Kevin Jernigan

Kevin started his career on the first product management team at Oracle, with responsibilities for utilities, benchmarks, and Oracle Parallel Server. After Oracle, he built a consulting business focused on data warehousing and high end transactional systems, and then built a SaaS business providing booking capabilities to the health club industry. He returned to Oracle to manage a team delivering storage and performance features in Oracle Database, and then joined AWS to launch Aurora PostgreSQL, which he helped build into the fastest-growing service in the history of AWS. In early 2021, Kevin joined the Atlas Serverless product team, and is focusing on bringing the Serverless from preview to general availability, and on working with customers to ensure it exceeds customer expectations in all dimensions, including ease of use, performance, pricing, scalability, functionality, and integration with the broader serverless application landscape.

  • Twitter: @kjerniga
  • LinkedIn: https://www.linkedin.com/in/kevinjernigan/
  • MongoDB Atlas: https://www.mongodb.com/atlas
  • MongoDB Atlas Serverless: https://www.mongodb.com/use-cases/serverless

View Details

Gwyn is currently a Regional Cloud Advocate at Microsoft as well as a YouTube content creator. She started in tech at a help desk role, where she was first introduced to cloud computing and the learning hasn't stopped since then. Her favorite topics are .NET and Azure Functions, and she’s always down to try out new things. Gwyn is passionate about introducing others to the cloud; creating friendly and concise content; and her family. When she’s not doing Advocate things, you can find her playing video games, hanging out with her family, or eating mint chocolate chip ice cream.

  • Twitter: @madebygps
  • LinkedIn: https://www.linkedin.com/in/gwyneth-pena/
  • GitHub: https://github.com/madebygps/
  • Personal website: https://www.gwynethpena.com/

What if everyone in tech started out in helpdesk... tweet

View Details

About Lee James Gilmore

Lee is a mentor, blogger, and cloud architect passionate about resolving complex problems with simple solutions, with a key focus on serverless technologies on AWS. He's currently a Global Serverless Architect at City Electrical Factors. Before that he worked as a Principal Developer / AWS Architect at AO across the five CeX (Customer Experience) teams; Customer Interactions, ChatBots, Order Management, My Account and Agent Experience, and also previously worked as a Technical Cloud Architect / Technical Lead on cloud native projects @ Sage PLC, after transitioning from Principal Software Developer, and with over 17 years professional experience in the industry.

He was a member of the extended leadership team at Sage within product delivery, with a keen interest in innovation, serverless architectures, and technology and has historically held long-term senior technology positions in two separate FTSE 100 companies, as well as running his own start-up, writing articles for ‘The Startup’ which has 680K followers, and mentoring in his free time. He's also 6x AWS Certified.

  • Twitter: @leejamesgilmore
  • LinkedIn: https://www.linkedin.com/in/lee-james-gilmore
  • GitHub: https://github.com/leegilmorecode

View Details

Episodes mentioned:

  • Episode #108: Mulling over Multi-cloud with Corey Quinn
  • Episode #123: APIs and the Evolution of Serverless with Dorian Smiley
  • Episode #124: Self-Provisioning Runtimes with Shawn "swyx" Wang
  • Episode #127: Supporting Women in Tech with Kristi Perreault
  • Episode #125: Configuration over Code with Eric Johnson
  • Episode #118: Deploying on Fridays with Charity Majors

View Details

Episodes mentioned:

  • Episode #132: The Evolution of Serverless at AWS with Dr. Werner Vogels
  • Episode #112: Abstracting Stateful Serverless with Jonas Bonér
  • Episode #110: Mapping the Inevitability of Serverless with Simon Wardley
  • Episode #128: Serverless-First Engineers and the Flywheel Effect with David Anderson
  • Episode #129: What To Do When the Servers Go Away with Tom McLaughlin
  • Episode #135: Serverless for Frontend Engineers with Swizec Teller
  • Episode #131: Security in the Cloud with Merritt Baer and Megan O'Neil

View Details

About Sarah HamiltonSarah Hamilton is a Software Engineer at LEGO Group and an AWS Community Builder. Prior to her current role, she was a Cloud Engineer at aleios.

  • Twitter: @serverlesssarah
  • LinkedIn: https://www.linkedin.com/in/hamilton-sarah/
  • Medium: https://medium.com/@08hamiltons
  • GitHub: https://github.com/hamilton-s

View Details

About Swizec Teller
Swizec Teller has been programming for the web since the early 2000's. From a server in his bedroom to web scale cloud ecosystems making millions of dollars. The sysadmin part always annoyed him. Too fiddly. Serverless caught his eye as the perfect answer for quick to get started, easy for engineers to use, fit for scale, no fiddling. You can ask him anything on twitter @swizec, or join the newsletter at swizec.com. He writes about web engineering lessons from practice.

  • Twitter: @Swizec
  • LinkedIn: https://www.linkedin.com/in/swizec/
  • GitHub: https://github.com/swizec
  • Personal website: https://swizec.com/
  • React for Data Visualization (course): https://reactfordataviz.com/
  • Serverless Handbook for Frontend Engineers (book): https://serverlesshandbook.dev/
  • The Senior Mindset Series: https://seniormindset.com/
  • YouTube Channel: https://www.youtube.com/c/SwizecTeller

View Details

About Farrah Campbell
After 10 years of working in healthcare management, a serendipitous 20-minute car ride with Kara Swisher inspired Farrah to make the jump into technology. She has worked at multiple startups in many different capacities, eventually working her way to being the Sr. Product Marketing Manager, Containers & Serverless.

Farrah previously worked as Ecosystems Director, at Stackery where she managed the relationship with AWS including Stackery as an Advanced Technology Partner, achieving the AWS DevOps Competency, a launch partner for Lambda Layers and is an AWS Serverless Hero. Farrah has cultivated the serverless community as an organizer of Portland Serverless Days, the Portland Serverless Meetup, along with numerous serverless workshops and the Portland tech community events from Techfest to bringing multiple luminaries to Portland.

  • Twitter: @FarrahC32
  • LinkedIn: https://www.linkedin.com/in/farrahcampbell/
  • AWS Community Builders: https://aws.amazon.com/developer/community/community-builders/

View Details

About Jeff Williams

Jeff brings more than 20 years of security leadership experience as Co-Founder and Chief Technology Officer of Contrast. Previously, Jeff was Co-Founder and Chief Executive Officer of Aspect Security, a successful and innovative application security consulting company acquired by Ernst & Young. Jeff is also a founder and major contributor to OWASP, where he served as Global Chairman for eight years and created the OWASP Top 10, OWASP Enterprise Security API, OWASP Application Security Verification Standard, XSS Prevention Cheat Sheet, and many other widely adopted free and open projects. Jeff has a BA from the University of Virginia, an MA from George Mason, and a JD from Georgetown.

  • Twitter: @planetlevel
  • LinkedIn: https://www.linkedin.com/in/planetlevel/
  • Contrast Security website: https://www.contrastsecurity.com/
  • OWASP Foundation: https://owasp.org/

View Details

Dr. Werner Vogels is Chief Technology Officer at Amazon.com where he is responsible for driving the company’s customer-centric technology vision.

As one of the forces behind Amazon’s approach to cloud computing, he is passionate about helping young businesses reach global scale, and transforming enterprises into fast-moving digital organizations.

Vogels joined Amazon in 2004 from Cornell University where he was a distributed systems researcher. He has held technology leadership positions in companies that handle the transition of academic technology into industry. Vogels holds a PhD from the Vrije Universiteit in Amsterdam and has authored many articles on distributed systems technologies for enterprise computing.

Twitter: https://twitter.com/Werner
LinkedIn: https://www.linkedin.com/in/wernervogels/
Blog: https://www.allthingsdistributed.com/
AWS: https://aws.amazon.com

View Details

About Merritt Baer
Merritt Baer is an emerging tech and infosec expert. She builds strategic initiatives for security and emerging technologies. She currently is a Principal Security Architect at Amazon Web Services (AWS), where she provides technical cloud security guidance to complex, regulated organizations like the Fortune 100, and advises the leadership of AWS' largest customers on security as a bottom line proposition.

Recently, Merritt served as the Lead Cyber Advisor to the Federal Communications Commission. She also wrote and implemented civilian cybersecurity strategy at the Department of Homeland Security's Office of Cybersecurity and Communications, the nation's cyber firehouse.

Merritt is a double Harvard graduate with experience in all three branches of government and a strong publication record. She is a leader in computer security, an Internet law and business expert, and a technology entrepreneur.

  • Twitter: @MerrittBaer
  • LinkedIn: https://www.linkedin.com/in/merrittbaer/
  • Personal website: https://www.merrittrachelbaer.com/

About Megan O’Neil
Megan is a Principal Security Solutions Architect at AWS. In her more than four years with AWS, she has had experience in threat detection and incident response, as well as in enabling customers to implement sophisticated, scalable, and secure solutions that solve their business challenges.

Megan’s expertise also includes collaborating with internal teams to design and develop secure solutions across multiple technologies and platforms, as well as providing strategic direction on enterprise security architecture and the implementation of appropriate safeguards and controls. She is also well-versed in assessing current and planned applications and systems, identifying security architecture issues and designing solutions for gaps.

  • LinkedIn: https://www.linkedin.com/in/megan-o-neil-aa147311/

View Details

About Matthieu Napoli
Matthieu is a software engineer passionate about helping developers to create. He’s the founder of Null, and currently a Senior Product Manager for Serverless Framework at Serverless Inc. Fascinated by how serverless unlocks creativity, he works on making serverless accessible to everyone.

Apart from consulting for clients, Matthieu also spends his time maintaining open-source projects. That includes Bref, a framework for creating serverless PHP applications on AWS. Alongside Bref, he sends a monthly newsletter containing serverless news relevant to PHP developers.

After years of talking at conferences and training teams on serverless, Matthieu created the Serverless Visually Explained course. Packed with use cases, visual explanations, and code samples, the course focuses on being practical and accessible.

  • Twitter: https://twitter.com/matthieunapoli
  • LinkedIn: https://www.linkedin.com/in/matthieunapoli
  • GitHub: https://github.com/mnapoli/
  • Personal website: https://mnapoli.fr/
  • Serverless Explained course: https://serverless-visually-explained.com/
  • Null: https://null.tc/
  • Serverless Framework: serverless.com

About Mariusz Nowak
Mariusz has been involved with full-stack development of web applications since 2004 and actively engaged in the open source community. He developed and published many JavaScript tools and modules, which play important part in implementation of modern web applications (client & server side) that he's worked with.

He also implemented a light, highly configurable, in-memory database engine that allows decentralized, network independent and (while in network connection) a real-time distribution/replication of database data: https://github.com/medikoo/dbjs.

  • Twitter: https://twitter.com/medikoo
  • LinkedIn: https://www.linkedin.com/in/mariusznowak
  • GitHub: https://github.com/medikoo

View Details

Tom is a cloud infrastructure and operations engineer with 13+ years of platform operations and IT experience, and over 8 years of AWS cloud infrastructure. He has worked in companies ranging from startups to the enterprise. His areas of focus around serverless started largely on the operational aspects of building and running reliable serverless systems. More recently his efforts involve mentoring teams new to AWS and serverless, and helping them successfully adopt these technologies through training and education.

Tom is also a leading serverless advocate in the DevOps community. He is a regular speaker at DevOpsDays conferences where he works to guide operations engineers in identifying and further the skills they need to be successful with serverless infrastructure.

What drew Tom early to serverless was the prospect of having no hosts or container management platform to build and manage which yielded the question: What would he do if the servers he was responsible for went away? As an early DevOps adopter he felt it was time to take a leap again into a new and emerging technology space. He’s found enjoyment in a community of people that are both pushing the future of technology and trying to understand its effects on the future of people and businesses.

When not working, Tom can be found racing his ‘87 Buick Grand National at the dragstrip, dabbling in photography, or playing with his cat Cinnamon.

  • Twitter: @tmclaughbos
  • LinkedIn: https://www.linkedin.com/in/tmclaugh/
  • ServerlessOps.io: https://www.serverlessops.io/
  • Serverless DevOps Ebook: https://www.serverlessops.io/download-the-serverless-devops-ebook
  • Dev.to: https://dev.to/tmclaughbos

View Details

Dave Anderson is currently a Technical Fellow with Bazaarvoice, where he focuses on product development, and technical and strategic leadership. He also is a contributor at The Serverless Edge, a blog for engineers, architects, and leaders interested in serverless, where he explores the narrative building on the new ways to create business value through software and technology.

Prior to these roles, Dave has led transformation, technical excellence, cloud adoption, fintech/insurtech strategies, technical community activity, and both participated and led several enterprise/organizational transformation efforts. His experience also includes frequent collaboration with senior executives, engineers and business sponsors. With Liberty Mutual, he designed and implemented large scale internet eCommerce systems and distributed web platforms. Operating as a tech startup within a Fortune 100 company, Dave led a period of digital disruption that put the organization ahead of the competition.

  • Twitter: https://twitter.com/davidand393
  • LinkedIn: https://www.linkedin.com/in/david-anderson-belfast/
  • The Serverless Edge: https://www.theserverlessedge.com/
  • The Serverless Craic (podcast): https://theserverlessedge.podbean.com/
  • Dev.to: https://dev.to/davidand39
  • The Flywheel Effect book: https://itrevolution.com/the-flywheel-effect/

View Details

Kristi Perreault is a Senior Software Engineer at Liberty Mutual Insurance, where her focus is serverless development and enablement. She has over 4 years of industry experience, holds an M.S. in Electrical & Computer Engineering, and has learned, followed, and preached the best coding practices she knows through it all. When she isn’t promoting Women in Technology and mentoring her dozens of new hires & interns, Kristi can be found in the mountains of Colorado hiking, mountain biking, skiing, golfing, paddle-boarding, or doing just about anything else outdoors.

  • Twitter: https://twitter.com/kperreault95
  • LinkedIn: https://www.linkedin.com/in/kristi-perreault/
  • Medium: https://kristiperreault.medium.com/
  • Business Insider Post: “I gave the wrong answer when I was asked how people can better support women in tech. Here's what I wish I said instead.”

View Details

Tomasz Łakomy is a Frontend Engineer at Stedi, Co-founder of Cloudash, an egghead.io instructor, and a lifelong learner with a passion for learning in public.

Since 2018, he's been diving into the world of AWS and at the same time sharing what he's learned with others. After passing the AWS Certified Solutions Architect: Associate exam in 2019 he recorded multiple courses on serverless technologies, including Build an App with the AWS Cloud Development Kit, and Learn AWS Lambda from scratch.

In addition, he's active on his Twitter, blog - tlakomy.com, as well as The Practical Dev community, where he posts articles on career advice, testing and - of course - AWS.

  • Twitter: @tlakomy
  • LinkedIn: https://www.linkedin.com/in/%F0%9F%9A%80-tomasz-Lakomy-12b2a258
  • GitHub: https://github.com/tlakomy/
  • Personal website: https://tlakomy.com/
  • Dev.to: https://dev.to/tlakomy
  • AWS Community Hero: https://aws.amazon.com/developer/community/heroes/tomasz-lakomy/
  • Cloudash: https://cloudash.dev/
  • Stedi: https://www.stedi.com/

View Details

Eric Johnson is a Principal Developer Advocate for Serverless Applications at Amazon Web Services and is based in Northern Colorado. Eric is a fanatic about serverless and enjoys helping developers understand how serverless technologies introduces a major paradigm shift in how they approach building and running applications at massive scale with minimal administration overhead. Prior to this, Eric has worked as a developer, solutions architect and AWS Evangelist for an AWS partner company.

  • Twitter: https://twitter.com/edjgeek
  • LinkedIn: https://www.linkedin.com/in/singledigit/
  • GitHub: https://github.com/singledigit
  • Serverless Land: https://serverlessland.com/about/eric-johnson/

View Details

Shawn “Swyx” Wang is currently Head of DX at Temporal.io, based out of Seattle. He is also a frequent writer and speaker best known for the Learn in Public movement and recently published The Coding Career Handbook with more advice for engineers going from Junior to Senior.

  • Twitter: @swyx
  • LinkedIn: https://www.linkedin.com/in/shawnswyxwang/
  • Website: https://www.swyx.io/
  • Github: https://github.com/sw-yx
  • YouTube: https://www.youtube.com/swyxTV
  • The Swyx Mixtape: https://swyx.transistor.fm/
  • The Self-Provision Runtime: https://www.swyx.io/self-provisioning-runtime

This episode is sponsored by Stream.

View Details

Dorian Smiley is a dedicated full-stack engineer with more than 15 years of experience. He is currently the VP of Technology at Brainly, the world's largest peer-to-peer learning community for students, parents and teachers. Prior to joining Brainly, Dorian spent a decade with Silicon Publishing Inc., first as Sr. Software Architect, and later as its Chief Scientific Officer. His extensive professional experience includes work with cloud native applications, microservices, serverless, big data architectures, PWAs, MEAN, MERN, and LAMP stacks.

  • LinkedIn: https://www.linkedin.com/in/dorian-smiley-97a72a14/
  • Medium: https://dorians.medium.com/
  • Github: https://github.com/doriansmiley
  • Brainly: https://brainly.com/
  • Brainly Tech Blog: https://medium.com/brainly

This episode is sponsored by Stream and Dexecure.

View Details

Ajay Nair is the General Manager (AWS Lambda Experience) at AWS. Ajay is one of the founding members of the AWS Lambda team, in his current role, drives the serverless product strategy and leads a talented team driving the product roadmap, feature delivery, and business results. Throughout his career, Ajay has focused on building and helping developers build large scale distributed systems, with deep expertise in cloud native application platforms, big data systems, and streamlining development experiences. He is also a co-author of Serverless Architectures on AWS, which teaches you how to design, secure, and manage serverless backend APIs for web and mobile applications on the AWS platform.

  • Twitter: @ajaynairthinks
  • LinkedIn: https://www.linkedin.com/in/ajnair/
  • Serverless Land: https://serverlessland.com

Talia Nassi is a Senior Developer Advocate at AWS Serverless and an international keynote speaker who delivers content on all things testing and quality. Previously, she worked at Split Software as a developer advocate and at WeWork as an engineer, and implemented Testing in Production from start to finish! She is passionate about feature flagging, canary launches, CI/CD, testing in production, and A/B testing. She has spoken at countless conferences internationally, ranging from audiences of 100 to 4000!

  • Twitter: @talia_nassi
  • LinkedIn: https://www.linkedin.com/in/talianassi/
  • Serverless Land: https://serverlessland.com

View Details

Ivonne Roberts is a recently named AWS Serverless Hero and currently a Software Architect at Bill.com. Prior to joining Bill.com, she was a Senior Software Architect, Principal Engineer at Edelman Financial Engines, where she and her team were critical in the company’s adoption of a serverless-first software development philosophy. She has experience in modernizing applications as part of cloud migration initiatives based on serverless architecture, and her expertise includes researching new technologies and design patterns, building prototypes, establishing reference architectures, and gaining buy-in from members across the organization. On her blog ivonneroberts.com and her YouTube channel Serverless DevWidgets, Ivonne focuses on demystifying and removing the hurdles of adopting serverless architecture and on simplifying the software development lifecycle.

  • Twitter: https://twitter.com/ivlo11
  • Website/personal blog: https://ivonneroberts.com
  • Serverless DevWidgets: https://www.youtube.com/c/ServerlessDevWidgets

View Details

Adam is an independent cloud consultant helping startups build products on AWS. He's also the host of AWS FM, a weekly podcast and live audio show where he shares stories from around the AWS community. Adam holds all twelve AWS certifications and is an AWS Community Builder. He's the creator of ness.sh, a CLI tool for deploying web sites and apps into your own AWS account. He's also the co-founder of StatMuse, a Disney and Google backed startup building search technology for sports and financial information. Adam lives in Nixa, Missouri, with his wife and two young boys.

  • Twitter: https://twitter.com/aeduhm
  • LinkedIn: https://www.linkedin.com/in/adamelmore/
  • Consulting Site: https://adam.dev/
  • AWS.FM Podcast: https://aws.fm/

View Details

Brian Scanlan is the Principal Systems Engineer at Intecom where he leads their developer infrastructure efforts, helping teams make products resilient to failure, scalable to customers' needs and need little to no human intervention to work well. Based out of Dublin, Brian has previously held posts with HEAnet and Amazon, and has experience helping teams build their technical strategies, as well as designing and implementing solutions. Brian is a frequent contributor to Intercom’s engineering blog, and has presented at LeadDev Con in London, Turing Fest, and Dash by Datadog.

  • Twitter: https://twitter.com/brian_scanlan
  • LinkedIn: https://www.linkedin.com/in/scanlanb/
  • Intercom’s Engineering Site: https://intercom.engineering/
  • 10 technical strategies to avoid when scaling your startup (and 5 to embrace)
  • How we fixed our on call process to avoid engineer burnout

View Details

Charity Majors is the co-founder and CTO of Honeycomb. Before that she worked at Facebook, Parse and Linden Lab on infrastructure and developer tools, and she always seems to wind up running databases. She is the co-author of "Database Reliability Engineering" and the upcoming "Observability Engineering: Achieving Production Excellence" book published by O'Reilly.

  • Twitter: https://twitter.com/mipsytipsy
  • LinkedIn: https://www.linkedin.com/in/charity-majors/
  • Blog: https://charity.wtf/
  • Honeycomb: https://www.honeycomb.io/

View Details

Doug Moscrop is the Lead Software Engineer at Serverless, Inc. working on Serverless Cloud. He is a pragmatic programmer with a strong engineering discipline, a thirst for knowledge and a desire for accuracy. He's the author of several Serverless Framework plugins and the creator of serverless-http.

Eslam Hefnawy is a Principal Software Engineer at Serverless Inc. working on Serverless Cloud. He's been writing software since the age of 14 and is passionate about open source, dev tools, and serverless technologies. In 2015, he joined Serverless, Inc as their first hire and co-created the Serverless Framework with founder, Austen Collins. In 2018, he lead the design and development of Serverless Components, a next-generation Serverless Framework.

Ben Miner is a Software Engineer at Serverless, Inc. working on Serverless Cloud and has experience all over the spectrum, including DevOps, backend, and frontend technologies. He's also a pursuer of SAAS architectures, functional programming, and great end user experiences.

  • Twitter:
    • Eslam: @eahefnawy
    • Doug: @dougmoscrop
    • Ben: @devvyben
  • Serverless Cloud: https://serverless.com/cloud

View Details

Sam Dengler is a Principal Solutions Architect and Justin Callison is an Engineering manager of Workflow service (including Step Functions) at Amazon Web Services.

  • Twitter: Justin Callison @justincallison, Sam Dengler @samdengler
  • Step Functions: https://aws.amazon.com/step-functions
  • Share Your Integration Innovations in the Flex Your Skills Contest on AWS
  • More Resources:
    • AWS Step Functions integrates with over 200 AWS SERVICES (Marcia Villalba)
    • Serverless Office Hours: AWS Step Functions - AWS SDK Service Integrations
    • Taco Bell: This is My Architecture video (we’ll link to this)

View Details

Ant is a consultant, community organizer, and co-founder of Homeschool from Senzo. He also founded and currently runs the Serverless User Group in London, is part of the ServerlessDays London organizing team and the global ServerlessDays leadership team. Previously Ant was a co-founder of A Cloud Guru, and was responsible for organizing the first ServerlessConf event in New York in May 2016. Living in London since 2009, Ant's background before Serverless is primarily as a Solution Architect at various organisations, from managed service providers to Tier 1 telecommunications providers. He started his career in 1999 doing Y2K upgrades in his native South Africa, and then spent 5 years being paid to write VB6. His current focus is Serverless, GraphQL and Node.js.

  • Twitter: @IamStan
  • Homeschool from Senzo: https://homeschool.dev
  • ServerlessDays: serverlessdays.io
  • For organizer information: organise@serverlessdays.io

View Details

Kesha Williams is an award-winning software engineer and technology leader teaching others how to transform their lives through technology. Forbes, Amazon Web Services (AWS), and Oracle have applauded her contributions to the technology community, and she has spoken on the TED stage about the transformative power of artificial intelligence (AI). Amazon recognized her pioneering work in AI with both its AWS Machine Learning Hero and Alexa Champion honors — the first person to receive both. Williams was named Mentor of the Year by Women Tech Network and received the Innovator Award from Hospitality Technology. She has launched several successful startups and appears in the 2020 tech documentary, "Hello World: The Film." Additionally, Williams serves on the Board of Directors for Women in Voice and as a mentor to women in tech.

  • Twitter: @KeshaWillz
  • Blog: kesha.tech/
  • LinkedIn: linkedin.com/in/java-rock-star-kesha/
  • Salary Overflow: salaryoverflow.com

View Details

Chris Munns is a Tech Lead/Advisor for Startup Solution Architects at Amazon Web Services based in New York City. Chris spent the last 4.5 years working with AWS's developer customers to understand how serverless technologies can drastically change the way they think about building and running applications at potentially massive scale with minimal administration overhead. Before this, Chris a global Business Development Manager for DevOps at AWS, he spent a few years as a Solutions Architect at AWS, and has held senior operations engineering posts at Etsy, Meetup, and other NYC based startups. Chris has a Bachelor of Science in Applied Networking and System Administration from the Rochester Institute of Technology.

  • Twitter: @chrismunns
  • Email: munns@amazon.com
  • AWS Compute Blog: https://aws.amazon.com/blogs/compute/

View Details

Jonas Bonér is founder and CEO of Lightbend, creator of the Akka project, initiator and co-author of the Reactive Manifesto and the Reactive Principles, and a Java Champion.

Website: http://jonasboner.com

Twitter: https://twitter.com/jboner

LinkedIn: https://www.linkedin.com/in/jonasboner/

Akka Serverless: https://www.lightbend.com/akka-serverless

Akka: https://akka.io/

Reactive Manifesto: https://www.reactivemanifesto.org/

Reactive Principles: https://principles.reactive.foundation/

View Details

Ali loves teaching people to code, and is currently doing so as a Senior Developer Advocate at AWS. She has been employed in the tech industry since 2014, holding multiple software engineering positions at startups, and a Distinguished Faculty and Faculty Lead role at General Assembly's Software Engineering Immersive. She blogs a lot about code and her life as a developer and also has a podcast with three other incredible women: Ladybug Podcast. They talk about the tech industry, their backgrounds, and go in depth on code-topics. When she’s not coding you can find her watching her favorite New England sports teams, taking runs with her dog Blair, or rock climbing.

  • Twitter: https://twitter.com/ASpittel
  • LinkedIn: https://www.linkedin.com/in/aspittel/
  • Portfolio: https://alispit.tel
  • Ladybug Podcast: https://www.ladybug.dev/
  • Blog / WeLearnCode: https://welearncode.com/
  • YouTube Channel: https://www.youtube.com/channel/UCOxxRhCHDqgtKplU_Ecu4BA
  • Twitch: https://www.twitch.tv/aspittel
  • Dev.to: https://dev.to/aspittel

View Details

Simon Wardley is a researcher for the Leading Edge Forum focused on the intersection of IT strategy and new technologies. Simon is a seasoned executive who has spent the last 15 years defining future IT strategies for companies in the FMCG, retail, and IT industries—from Canon’s early leadership in the cloud-computing space in 2005 to Ubuntu’s recent dominance as the top cloud operating system. As a geneticist with a love of mathematics and a fascination for economics, Simon has always found himself dealing with complex systems, whether in behavioral patterns, the environmental risks of chemical pollution, developing novel computer systems, or managing companies. He is a passionate advocate and researcher in the fields of open source, commoditization, innovation, organizational structure, and cybernetics.

Simon’s most recent published research, “Clash of the Titans: Can China Dethrone Silicon Valley?,” assesses the high-tech challenge from China and what this means to the future of global technology industry competition. His previous research covers topics including the nature of technological and business change over the next 20 years, value chain mapping, strategies for an increasingly open economy, Web 2.0, and a lifecycle approach to cloud computing. Simon is a regular presenter at conferences worldwide and has been voted one of the UK’s top 50 most influential people in IT in Computer Weekly’s 2011 and 2012 polls.

Twitter: https://twitter.com/swardley
Medium: https://swardley.medium.com
Blog: https://blog.gardeviance.org
Wardley Maps (free online book): https://medium.com/wardleymaps

Simon's slides discussed during the podcast

View Details

Emily Shea is a Sr. Serverless GTM Specialist at AWS. Emily has been at Amazon for 5 years and currently works with customers adopting serverless in the UK & Ireland. In her free time, Emily has learned to code and build her own serverless applications. Emily’s current personal project is a daily Chinese vocabulary app with over 100 subscribers.

Twitter: https://twitter.com/em__shea
Personal blog: https://emshea.com/
Chinese vocabulary app: https://haohaotiantian.com/
re:Invent talk: Getting started building your first serverless web application

View Details

Corey Quinn is the Cloud Economist at The Duckbill Group. Corey’s unique brand of snark combines with a deep understanding of AWS’s offerings, unlocking a level of insight that’s both penetrating and hilarious. He lives in San Francisco with his spouse and daughter.

  • Twitter: https://twitter.com/QuinnyPig
  • LinkedIn: https://www.linkedin.com/in/coquinn/
  • Last Week in AWS: https://www.lastweekinaws.com/
  • The Morning Brief and Screaming in the Cloud: https://www.lastweekinaws.com/podcast/
  • Duckbill Group: https://www.duckbillgroup.com/

View Details

About Ben Kehoe
Ben Kehoe is a Cloud Robotics Research Scientist at iRobot and an AWS Serverless Hero. As a serverless practitioner, Ben focuses on enabling rapid, secure-by-design development of business value by using managed services and ephemeral compute (like FaaS). Ben also seeks to amplify voices from dev, ops, and security to help the community shape the evolution of serverless and event-driven designs.

Twitter: @ben11kehoe
Medium: ben11kehoe
GitHub: benkehoe
LinkedIn: ben11kehoe
iRobot: www.irobot.com

Watch this episode on YouTube: https://youtu.be/B0QChfAGvB0

This episode is sponsored by CBT Nuggets and Lumigo.

Transcript
Jeremy: Hi, everyone. I'm Jeremy Daly.

Rebecca: And I'm Rebecca Marshburn.

Jeremy: And this is Serverless Chats. And this is a momentous occasion on Serverless Chats because we are welcoming in Rebecca Marshburn as an official co-host of Serverless Chats.

Rebecca: I'm pretty excited to be here. Thanks so much, Jeremy.

Jeremy: So for those of you that have been listening for hopefully a long time, and we've done over 100 episodes. And I don't know, Rebecca, do I look tired? I feel tired.

Rebecca: I've never seen you look tired.

Jeremy: Okay. Well, I feel tired because we've done a lot of these episodes and we've published a new episode every single week for the last 107 weeks, I think at this point. And so what we're going to do is with you coming on as a new co-host, we're going to take a break over the summer. We're going to revamp. We're going to do some work. We're going to put together some great content. And then we're going to come back on, I think it's August 30th with a new episode and a whole new show. Again, it's going to be about serverless, but what we're thinking is ... And, Rebecca, I would love to hear your thoughts on this as I come at things from a very technical angle, because I'm an overly technical person, but there's so much more to serverless. There's so many other sides to it that I think that bringing in more perspectives and really being able to interview these guests and have a different perspective I think is going to be really helpful. I don't know what your thoughts are on that.

Rebecca: Yeah. I love the tech side of things. I am not as deep in the technicalities of tech and I come at it I think from a way of loving the stories behind how people got there and perhaps who they worked with to get there, the ideas of collaboration and community because nothing happens in a vacuum and there's so much stuff happening and sharing knowledge and education and uplifting each other. And so I'm super excited to be here and super excited that one of the first episodes I get to work on with you is with Ben Kehoe because he's all about both the technicalities of tech, and also it's actually on his Twitter, a new compassionate tech values around humility, and inclusion, and cooperation, and learning, and being a mentor. So couldn't have a better guest to join you in the Serverless Chats community and being here for this.

Jeremy: I totally agree. And I am looking forward to this. I'm excited. I do want the listeners to know we are testing in production, right? So we haven't run any unit tests, no integration tests. I mean, this is straight test in production.

Rebecca: That's the best practice, right? Total best practice to test in production.

Jeremy: Best practice. Right. Exactly.

Rebecca: Straight to production, always test in production.

Jeremy: Push code to the cloud. Here we go.

Rebecca: Right away.

Jeremy: Right. So if it's a little bit choppy, we'd love your feedback though. The listeners can be our observability tool and give us some feedback and we can ... And hopefully continue to make the show better. So speaking of Ben Kehoe, for those of you who don't know Ben Kehoe, I'm going to let him introduce himself, but I have always been a big fan of his. He was very, very early in the serverless space. I read all his blogs very early on. He was an early AWS Serverless Hero. So joining us today is Ben Kehoe. He is a cloud robotics research scientist at iRobot, as I said, an AWS Serverless Hero. Ben, welcome to the show.

Ben: Thanks for having me. And I'm excited to be a guinea pig for this new exciting format.

Rebecca: So many observability tools watching you be a guinea pig too. There's lots of layers to this.

Jeremy: Amazing. All right. So Ben, why don't you tell the listeners for those that don't know you a little bit about yourself and what you do with serverless?

Ben: Yeah. So I mean, as with all software, software is people, right? It's like Soylent Green. And so I'm really excited for this format being about the greater things that technology really involves in how we create it and set it up. And serverless is about removing the things that don't matter so that you can focus on the things that do matter.

Jeremy: Right.

Ben: So I've been interested in that since I learned about it. And at the time saw that I could build things without running servers, without needing to deal with the scaling of stuff. I've been working on that at iRobot for over five years now. As you said early on in serverless at the first serverless con organized by A Cloud Guru, now plural sites.

Jeremy: Right.

Ben: And yeah. And it's been really exciting to see it grow into the large-scale community that it is today and all of the ways in which community are built like this podcast.

Jeremy: Right. Yeah. I love everything that you've done. I love the analogies you've used. I mean, you've always gone down this road of how do you explain serverless in a way to show really the adoption of it and how people can take that on. Serverless is a ladder. Some of these other things that you would ... I guess the analogies you use were always great and always helped me. And of course, I don't think we've ever really come to a good definition of serverless, but we're not talking about that today. But ...

Ben: There isn't one.

Jeremy: There isn't one, which is also a really good point. So yeah. So welcome to the show. And again, like I said, testing in production here. So, Rebecca, jump in when you have questions and we'll beat up Ben from both sides on this, but, really ...

Rebecca: We're going to have Ben from both sides.

Jeremy: There you go. We'll embrace him from both sides. There you go.

Rebecca: Yeah. Yeah.

Jeremy: So one of the things though that, Ben, you have also been very outspoken on which I absolutely love, because I'm in very much closely aligned on this topic here. But is about infrastructure as code. And so let's start just quickly. I mean, I think a lot of people know or I think people working in the cloud know what infrastructure as code is, but I also think there's a lot of people who don't. So let's just take a quick second, explain what infrastructure as code is and what we mean by that.

Ben: Sure. To my mind, infrastructure as code is about having a definition of the state of your infrastructure that you want to see in the cloud. So rather than using operations directly to modify that state, you have a unified definition of some kind. I actually think infrastructure is now the wrong word with serverless. It used to be with servers, you could manage your fleet of servers separate from the software that you were deploying onto the servers. And so infrastructure being the structure below made sense. But now as your code is intimately entwined in the rest of your resources, I tend to think of resource graph definitions rather than infrastructure as code. It's a less convenient term, but I think it's worth understanding the distinction or the difference in perspective.

Jeremy: Yeah. No, and I totally get that. I mean, I remember even early days of cloud when we were using the Chefs and the Puppets and things like that, that we were just deploying the actual infrastructure itself. And sometimes you deploy software as part of that, but it was supporting software. It was the stuff that ran in the runtime and some of those and some configurations, but yeah, but the application code that was a whole separate process, and now with serverless, it seems like you're deploying all those things at the same time.

Ben: Yeah. There's no way to pick it apart.

Jeremy: Right. Right.

Rebecca: Ben, there's something that I've always really admired about you and that is how strongly you hold your opinions. You're fervent about them, but it's also because they're based on this thorough nature of investigation and debate and challenging different people and yourself to think about things in different ways. And I know that the rest of this episode is going to be full with a lot of opinions. And so before we even get there, I'm curious if you can share a little bit about how you end up arriving at these, right? And holding them so steady.

Ben: It's a good question. Well, I hope that I'm not inflexible in these strong opinions that I hold. I mean, it's one of those strong opinions loosely held kind of things that new information can change how you think about things. But I do try and do as much thinking as possible so that there's less new information that I have to encounter to change an opinion.

Rebecca: Yeah. Yeah.

Ben: Yeah. I think I tend to try and think about how people ... But again, because it's always people. How people interact with the technology, how people behave, how organizations behave, and then how technology fits into that. Because sometimes we talk about technology in a vacuum and it's really not. Technology that works for one context doesn't work for another. I mean, a lot of my strong opinions are that there is no one right answer kind of a thing, or here's a framework for understanding how to think about this stuff. And then how that fits into a given person is just finding where they are in that more general space. Does that make sense? So it's less about finding out here's the one way to do things and more about finding what are the different options, how do you think about the different options that are out there.

Rebecca: Yeah, totally makes sense. And I do want to compliment you. I do feel like you are very good at inviting new information in if people have it and then you're like, "Aha, I've already thought of that."

Ben: I hope so. Yeah. I was going to say, there's always a balance between trying to think ahead so that when you discover something you're like, "Oh, that fits into what I thought." And the danger of that being that you're twisting the information to fit into your preexisting structures. I hope that I find a good balance there, but I don't have a principle way of determining that balance or knowing where you are in that it's good versus it's dangerous kind of spectrum.

Jeremy: Right. So one of the opinions that you hold that I tend to agree with, I have some thoughts about some of the benefits, but I also really agree with the other piece of it. And this really has to do with the CDK and this idea of using CloudFormation or any sort of DSL, maybe Terraform, things like that, something that is more domain-specific, right? Or I guess declarative, right? As opposed to something that is imperative like the CDK. So just to get everybody on the same page here, what is the top reasons why you believe, or you think that DSL approach is better than that iterative approach or interpretive approach, I guess?

Ben: Yeah. So I think we get caught up in the imperative versus declarative part of it. I do think that declarative has benefits that can be there, but the way that I think about it is with the CDK and infrastructure as code in general, I'm like mildly against imperative definitions of resources. And we can get into that part, but that's not my smallest objection to the CDK. I'm moderately against not being able to enforce deterministic builds. And the CDK program can do anything. Can use a random number generator and go out to the internet to go ask a question, right? It can do anything in that program and that means that you have no guarantees that what's coming out of it you're going to be able to repeat.

So even if you check the source code in, you may not be able to go back to the same infrastructure that you had before. And you can if you're disciplined about it, but I like tools that help give you guardrails so that you don't have to be as disciplined. So that's my moderately against. My strongly against piece is I'm strongly against developer intent remaining client side. And this is not an inherent flaw in the CDK, is a choice that the CDK team has made to turn organizational dysfunction in AWS into ownership for their customers. And I don't think that's a good approach to take, but that's also fixable.

So I think if we want to start with the imperative versus declarative thing, right? When I think about the developers expressing an intent, and I want that intent to flow entirely into the cloud so that developers can understand what's deployed in the cloud in terms of the things that they've written. The CDK takes this approach of flattening it down, flattening the richness of the program the developer has written into ... They think of it as assembly language. I think that is a misinterpretation of what's happening. The assembly language in the process is the imperative plan generated inside the CloudFormation engine that says, "Here's how I'm going to take this definition and turn it into an actual change in the cloud.

Jeremy: Right.

Ben: They're just translating between two definition formats in CDK scene. But it's a flattening process, it's a lossy process. So then when the developer goes to the Console or the API has to go say, "What's deployed here? What's going wrong? What do I need to fix?" None of it is framed in terms of the things that they wrote in their original language.

Jeremy: Right.

Ben: And I think that's the biggest problem, right? So drift detection is an important thing, right? What happened when someone went in through the Console? Went and tweaked some stuff to fix something, and now it's different from the definition that's in your source repository. And in CloudFormation, it can tell you that. But what I would want if I was running CDK is that it should produce another CDK program that represents the current state of the cloud with a meaningful file-level diff with my original program.

Jeremy: Right. I'm just thinking this through, if I deploy something to CDK and I've got all these loops and they're generating functions and they're using some naming and all this kind of stuff, whatever, now it produces this output. And again, my naming of my functions might be some function that gets called to generate the names of the function. And so now I've got all of these functions named and I have to go in. There's no one-to-one map like you said, and I can imagine somebody who's not familiar with CloudFormation which is ultimately what CDK synthesizes and produces, if you're not familiar with what that output is and how that maps back to the constructs that you created, I can see that as being really difficult, especially for younger developers or developers who are just getting started in that.

Ben: And the CDK really takes the attitude that it's going to hide those things from those developers rather than help them learn it. And so when they do have to dive into that, the CDK refers to it as an escape hatch.

Jeremy: Yeah.

Ben: And I think of escape hatches on submarines, where you go from being warm and dry and having air to breathe to being hundreds of feet below the sea, right? It's not the sort of thing you want to go through. Whereas some tools like Amplify talk about graduation. In Amplify they aim to help you understand the things that Amplify is doing for you, such that when you grow beyond what Amplify can provide you, you have the tools to do that, to take the thing that you built and then say, "Okay, I know enough now that I understand this and can add onto it in ways that Amplify can't help with."

Jeremy: Right.

Ben: Now, how successful they are in doing that is a separate question I think, but the attitude is there to say, "We're looking to help developers understand these things." Now the CDK could also if the CDK was a managed service, right? Would not need developers to understand those things. If you could take your program directly to the cloud and say, "Here's my program, go make this real." And when it made it real, you could interact with the cloud in an understanding where you could list your deployed constructs, right? That you can understand the program that you wrote when you're looking at the resources that are deployed all together in the cloud everywhere. That would be a thing where you don't need to learn CloudFormation.

Jeremy: Right.

Ben: Right? That's where you then end up in the imperative versus declarative part where, okay, there's some reasons that I think declarative is better. But the major thing is that disconnect that's currently built into the way that CDK works. And the reason that they're doing that is because CloudFormation is not moving fast enough, which is not always on the CloudFormation team. It's often on the service teams that aren't building the resources fast enough. And that's AWS's problem, AWS as an entire company, as an organization. And this one team is saying, "Well, we can fix that by doing all this client side."

What that means is that the customers are then responsible for all the things that are happening on the client side. The reason that they can go fast is because the CDK team doesn't have ownership of it, which just means the ownership is being pushed on customers, right? The CDK deploys Lambda functions into your account that they don't tell you about that you're now responsible for. Right? Both the security and operations of. If there are security updates that the CDK team has to push out, you have to take action to update those things, right? That's ownership that's being pushed onto the customer to fix a lack of ACM certificate management, right?

Jeremy: Right. Right.

Ben: That is ACM not building the thing that's needed. And so AWS says, "Okay, great. We'll just make that the customer's problem."

Jeremy: Right.

Ben: And I don't agree with that approach.

Rebecca: So I'm sure as an AWS Hero you certainly have pretty good, strong, open communication channels with a lot of different team members across teams. And I certainly know that they're listening to you and are at least hearing you, I should say, and watching you and they know how you feel about this. And so I'm curious how some of those conversations have gone. And some teams as compared to others at AWS are really, really good about opening their roadmap or at least saying, "Hey, we hear this, and here's our path to a solution or a success." And I'm curious if there's any light you can shed on whether or not those conversations have been fruitful in terms of actually being able to get somewhere in terms of customer and AWS terms, right? Customer obsession first.

Ben: Yeah. Well, customer obsession can mean two things, right? Customer obsession can mean giving the customer what they want or it can mean giving the customer what they need and different AWS teams' approach fall differently on that scale. The reason that many of those things are not available in CloudFormation is that those teams are ... It could be under-resourced. They could have a larger majority of customer that want new features rather than infrastructure as code support. Because as much as we all like infrastructure as code, there are many, many organizations out there that are not there yet. And with the CDK in particular, I'm a relatively lone voice out there saying, "I don't think this ownership that's being pushed onto the customer is a good thing." And there are lots of developers who are eating up CDK saying, "I don't care."

That's not something that's in their worry. And because the CDK has been enormously successful, right? It's fixing these problems that exists. And I don't begrudge them trying to fix those problems. I think it's a question of do those developers who are grabbing onto those things and taking them understand the full total cost of ownership that the CDK is bringing with it. And if they don't understand it, I think AWS has a responsibility to understand it and work with it to help those customers either understand it and deal with it, right? Which is where the CDK takes this approach, "Well, if you do get Ops, it's all fine." And that's somewhat true, but also many developers who can use the CDK do not control their CI/CD process. So there's all sorts of ways in which ... Yeah, so I think every team is trying to do the best that they can, right?

They're all working hard and they all have ... Are pulled in many different directions by customers. And most of them are making, I think, the right choices given their incentives, right? Given what their customers are asking for. I think not all of them balance where customers ... meeting customers where they are versus leading them where they should, like where they need to go as well as I would like. But I think ... I had a conclusion to that. Oh, but I think that's always a debate as to where that balance is. And then the other thing when I talk about the CDK, that my ideal audience there is less AWS itself and more AWS customers ...

Rebecca: Sure.

Ben: ... to understand what they're getting into and therefore to demand better of AWS. Which is in general, I think, the approach that I take with AWS, is complaining about AWS in public, because I do have the ability to go to teams and say, "Hey, I want this thing," right? There are plenty of teams where I could just email them and say, "Hey, this feature could be nice", but I put it on Twitter because other people can see that and say, "Oh, that's something that I want or I don't think that's helpful," right? "I don't care about that," or, "I think it's the wrong thing to ask for," right? All of those things are better when it's not just me saying I think this is a good thing for AWS, but it being a conversation among the community differently.

Rebecca: Yeah. I think in the spirit too of trying to publicize types of what might be best next for customers, you said total cost of ownership. Even though it might seem silly to ask this, I think oftentimes we say the words total cost of ownership, but there's actually many dimensions to total cost of ownership or TCO, right? And so I think it would be great if you could enumerate what you think of as total cost of ownership, because there might be dimensions along that matrices, matrix, that people haven't considered when they're actually thinking about total cost of ownership. They're like, "Yeah, yeah, I got it. Some Ops and some security stuff I have to do and some patches," but they might only be thinking of five dimensions when you're like, "Actually the framework is probably 10 to 12 to 14." And so if you could outline that a bit, what you mean when you think of a holistic total cost of ownership, I think that could be super helpful.

Ben: I'm bad at enumeration. So I would miss out on dimensions that are obvious if I was attempting to do that. But I think a way that I can, I think effectively answer that question is to talk about some of the ways in which we misunderstand TCO. So I think it's important when working in an organization to think about the organization as a whole, not just your perspective and that your team's perspective in it. And so when you're working for the lowest TCO it's not what's the lowest cost of ownership for my team if that's pushing a larger burden onto another team. Now if it's reducing the burden on your team and only increasing the burden on another team a little bit, that can be a lower total cost of ownership overall. But it's also something that then feeds into things like political capital, right?

Is that increased ownership that you're handing to that team something that they're going to be happy with, something that's not going to cause other problems down the line, right? Those are the sorts of things that fit into that calculus because it's not just about what ... Moving away from that topic for a second. I think about when we talk about how does this increase our velocity, right? There's the piece of, "Okay, well, if I can deploy to production faster, right? My feedback loop is faster and I can move faster." Right? But the other part of that equation is how many different threads can you be operating on and how long are those threads in time? So when you're trying to ship a feature, if you can ship it and then never look at it again, that means you have increased bandwidth in the future to take on other features to develop other new features.

And so even if you think about, "It's going to take me longer to finish this particular feature," but then there's no maintenance for that feature, that can be a lower cost of ownership in time than, "I can ship it 50% faster, but then I'm going to periodically have to revisit it and that's going to disrupt my ability to ship other things," right? So this is where I had conversations recently about increasing use of Step Functions, right? And being able to replace Lambda functions with Step Functions express workflows because you never have to go back to those Lambdas and update dependencies in them because dependent bot has told you that you need to or a version of Python is getting deprecated, right? All of those things, just if you have your Amazon States Language however it's been defined, right?

Once it's in there, you never have to touch it again if nothing else changes and that means, okay, great, that piece is now out of your work stream forever unless it needs to change. And that means that you have more bandwidth for future things, which serverless is about in general, right? Of say, "Okay, I don't have to deal with this scaling problems here. So those scaling things. Once I have an auto-scaling group, I don't have to go back and tweak it later." And so the same thing happens at the feature level if you build it in ways that allow you to do that. And so I think that's one of the places where when we focus on, okay, how fast is this getting me into production, it's okay, but how often do you have to revisit it ...

Jeremy: Right. And so ... So you mentioned a couple of things in there, and not only in that question, but in the previous questions as you were talking about the CDK in general, and I am 100% behind you on this idea of deterministic builds because I want to know exactly what's being deployed. I want to be able to audit that and map that back. And you can audit, I mean, you could run CDK synth and then audit the CloudFormation and test against certain things. But if you are changing stuff, right? Then you have to understand not only the CDK but also the CloudFormation that it actually generates. But in terms of solving problems, some of the things that the CDK does really, really well, and this is something where I've always had this issue with just trying to use raw CloudFormation or Serverless Framework or SAM or any of these things is the fact that there's a lot of boilerplate that you often have to do.

There's ways that companies want to do something specifically. I basically probably always need 1,400 lines of CloudFormation. And for every project I do, it's probably close to the same, and then add a little bit more to actually make it adaptive for my product. And so one thing that I love about the CDK is constructs. And I love this idea of being able to package these best practices for your company or these compliance requirements, excuse me, compliance requirements for your company, whatever it is, be able to package these and just hand them to developers. And so I'm just curious on your thoughts on that because that seems like a really good move in the right direction, but without the deterministic builds, without some of these other problems that you talked about, is there another solution to that that would be more declarative?

Ben: Yeah. In theory, if the CDK was able to produce an artifact that represented all of the non-deterministic dependencies that it had, right? That allowed you to then store that artifacts as you'd come back and put that into the program and say, "I'm going to get out the same thing," but because the CDK doesn't control upstream of it, the code that the developers are writing, there isn't a way to do that. Right? So on the abstraction front, the constructs are super useful, right? CloudFormation now has modules which allow you to say, "Here's a template and I'm going to represent this as a CloudFormation type itself," right? So instead of saying that I need X different things, I'm going to say, "I packaged that all up here. It is as a type."

Now, currently, modules can only be playing CloudFormation templates and there's a lot of constraints in what you can express inside a CloudFormation template. And I think the answer for me is ... What I want to see is more richness in the CloudFormation language, right? One of the things that people do in the CDK that's really helpful is say, "I need a copy of this in every AZ."

Jeremy: Right.

Ben: Right? There's so much boilerplate in server-based things. And CloudFormation can't do that, right? But if you imagine that it had a map function that allowed you to say, "For every AZ, stamp me out a copy of this little bit." And then that the CDK constructs allowed to translate. Instead of it doing all this generation only down to the L one piece, instead being able to say, "I'm going to translate this into more rich CloudFormation templates so that the CloudFormation template was as advanced as possible."

Right? Then it could do things like say, "Oh, I know we need to do this in every AZ, I'm going to use this map function in the CloudFormation template rather than just stamping it out." Right? And so I think that's possible. Now, modules should also be able to be defined as CDK programs. Right? You should be able to register a construct as a CloudFormation tag.

Jeremy: It would be pretty cool.

Ben: There's no reason you shouldn't be able to. Yeah. Because I think the declarative versus imperative thing is, again, not the most important piece, it's how do we move ... It's shifting right in this case, right? That how do you shift what's happening with the developer further into the process of deployment so that more of their context is present? And so one of the things that the CDK does that's hard to replicate is have non-local effects. And this is both convenient and I think of code smell often.

So you can pass a bucket resource from another stack into a piece of code in your CDK program that's creating a different stack and you say, "Oh great, I've got this Lambda function, it needs permissions to that bucket. So add permissions." And it's possible for the CDK programs to either be adding the permissions onto the IAM role of that function, or non-locally adding to that bucket's resource policy, which is weird, right? That you can be creating a stack and the thing that you do to that stack or resource or whatever is not happening there, it's happening elsewhere. I don't think that's a great approach, but it's certainly convenient to be able to do it in a lot of situations.

Now, that's not representable within a module. A module is a contained piece of functionality that can't touch anything else. So things like SAM where you can add events onto a function that can go and create ... You create the API events on different functions and then SAM aggregates them and creates an API gateway for you. Right? If AWS serverless function was a module, it couldn't do that because you'd have these in different places and you couldn't aggregate something between all of them and put them in the top-level thing, right?

This is what CloudFormation macros enable, but they don't have a... There's no proper interface to them, right? They don't define, "This is what I'm doing. This is the kind of resources I can create." There's none of that that would help you understand them. So they're infinitely flexible, but then also maybe less principled for that reason. So I think there are ways to evolve, but it's investment in the CloudFormation language that allows us to shift that burden from being a flattening inside client-side code from the developer and shifting it to be able to be represented in the cloud.

Jeremy: Right. Yeah. And I think from that standpoint too if we go back to the solving people's problems standpoint, that everything you explained there, they're loaded with nuances, it's loaded with gotchas, right? Like, "Oh, you can't do this, you can't do that." So that's just why I think the CDK is so popular because it's like you can do so much with it so quickly and it's very, very fast. And I think that trade-off, people are just willing to make it.

Ben: Yes. And that's where they're willing to make it, do they fully understand the consequences of it? Then does AWS communicate those consequences well? Before I get into that question of, okay, you're a developer that's brand new to AWS and you've been tasked with standing up some Kubernetes cluster and you're like, "Great. I can use a CDK to do this." Something is malfunctioning. You're also tasked with the operations and something is malfunctioning. You go in through the Console and maybe figure out all the things that are out there are new to you because they're hidden inside L3 constructs, right?

You're two levels down from where you were defining what you want, and then you find out what's wrong and you have no idea how to turn that into a change in your CDK program. So instead of going back and doing the thing that infrastructure as code is for, which is tweaking your program to go fix the problem, you go and you tweak it in the Console ...

Jeremy: Right. Which you should never do.

Ben: ... and you fix it that way. Right. Well, and that's the thing that I struggle with, with the CDK is how does the CDK help the developer who's in that situation? And I don't think they have a good story around that. Now, I don't know. I haven't talked with enough junior developers who are using the CDK about how often they get into that situation. Right? But I always say client-side code is not a replacement for a managed service because when it's client-side code, you still own the result.

Jeremy: Right.

Ben: If a particular CDK construct was a managed service in AWS, then all of the resources that would be created underneath AWS's problem to make work. And the interface that the developer has is the only level of ownership that they have. Fargate is this. Because you could do all the things that Fargate does with a CDK construct, right? Set up EC2, do all the things, and represent it as something that looks like Fargate in your CDK program. But every time your EC2 fleet is unhealthy that's your problem. With Fargate, that's AWS's problem. If we didn't have Fargate, that's essentially what CDK would be trying to do for ECS.

And I think we all recognize that Fargate is very necessary and helpful in that case, right? And I just want that for all the things, right? Whenever I have an abstraction, if it's an abstraction that I understand, then I should have a way of zooming into it while not having to switch languages, right? So that's where you shouldn't dump me out the CloudFormation to understand what you're doing. You should help me understand the low-level things in the same language. And if it's not something that I need to understand, it should be a managed service. It shouldn't be a bunch of stuff that I still own that I haven't looked at.

Jeremy: Makes sense. Got a question, Rebecca? Because I was waiting for you to jump in.

Rebecca: No, but I was going to make a joke, but then the joke passed, and then I was like, "But should I still make it?" I was going to be like, "Yeah, but does the CDK let you test in production?" But that was a 32nd ago joke and then I was really wrestling with whether or not I should tell it, but I told it anyway, hopefully, someone gets a laugh.

Ben: Yeah. I mean, there's the thing that Charity Majors says, right? Which is that everybody tests in production. Some people are lucky enough to have a development environment in production. No, sorry. I said that the wrong way. It's everybody has a test environment. Some people are lucky enough that it's not in production.

Rebecca: Yeah. Swap that. Reverse it. Yeah.

Ben: Yeah.

Jeremy: All right. So speaking of talking to developers and getting feedback from them, so I actually put a question out on Twitter a couple of weeks ago and got a lot of really interesting reactions. And essentially I asked, "What do you love or hate about infrastructure as code?" And there were a lot of really interesting things here. I don't know, maybe it might be fun to go through a couple of these and get your thoughts on them. So this is probably not a great one to start with, but I thought it was interesting because this I think represents the frustration that a lot of us feel. And it was basically that they love that automation minimizes future work, right? But they hate that it makes life harder over time. And that pretty much every approach to infrastructure in, sorry, yeah, infrastructure in code at the present is flawed, right? So really there are no good solutions right now.

Ben: Yeah. CloudFormation is still a pain to learn and deal with. If you're operating in certain IDEs, you can get tab completion.

Jeremy: Right.

Ben: If you go to CDK you get tab completion, which is, I think probably most of the value that developers want out of it and then the abstraction, and then all the other fancy things it does like pipelines, which again, should be a managed service. I do think that person is absolutely right to complain about how difficult it is. That there are many ways that it could be better. One of the things that I think about when I'm using tools is it's not inherently bad for a tool to have some friction to use it.

Jeremy: Right.

Ben: And this goes to another infrastructure as code tool that goes even further than the CDK and says, "You can define your Lambda code in line with your infrastructure definition." So this is fine with me. And there's some other ... I think Punchcard also lets you do some of this. Basically extracts out the bits of your code that you say, "This is a custom thing that glues together two things I'm defining in here and I'll make that a Lambda function for you." And for me, that is too little friction to defining a Lambda function.

Because when I define a Lambda function, just going back to that bringing in ownership, every time I add a Lambda function, that's something that I own, that's something that I have to maintain, that I'm responsible for, that can go wrong. So if I'm thinking about, "Well, I could have API Gateway direct into DynamoDB, but it'd be nice if I could change some of these fields. And so I'm just going to drop in a little sprinkle of code, three lines of code in between here to do some transformation that I want." That is all of sudden an entire Lambda function you've brought into your infrastructure.

Jeremy: Right. That's a good point.

Ben: And so I want a little bit of friction to do that, to make me think about it, to make me say, "Oh, yeah, downstream of this decision that I am making, there are consequences that I would not otherwise think about if I'm just trying to accomplish the problem," right? Because I think developers, humans, in general, tend to be a bit shortsighted when you have a goal especially, and you're being pressured to complete that goal and you're like, "Okay, well I can complete it." The consequences for later are always a secondary concern.

And so you can change your incentives in that moment to say, "Okay, well, this is going to guide me to say, "Ah, I don't really need this Lambda function in here. Then I'm better off in the long term while accomplishing that goal in the short term." So I do think that there is a place for tools making things difficult. That's not to say that the amount of difficult that infrastructure as code is today is at all reasonable, but I do think it's worth thinking about, right?

I'd rather take on the pain of creating an ASL definition by hand for express workflow than the easier thing of writing Lambda code. Because I know the long-term consequences of that. Now, if that could be flipped where it was harder to write something that took more ownership, it'd be just easy to do, right? You'd always do the right thing. But I think it's always worth saying, "Can I do the harder thing now to pay off to pay off later?"

Jeremy: And I always call those shortcuts "tomorrow-Jeremy's" problem. That's how I like to look at those.

Ben: Yeah. Yes.

Jeremy: And the funny thing about that too is I remember right when EventBridge came out and there was no CloudFormation support for a long time, which was super frustrating. But Serverless Framework, for example, implemented a custom resource in order to do that. And I remember looking at a clean stack and being like, "Why are there two Lambda functions there that I have no idea?" I'm like, "I didn't publish ..." I honestly thought my account was compromised that somebody had published a Lambda function in there because I'm like, "I didn't do that." And then it took me a while to realize, I'm like, "Oh, this is what this is." But if it is that easy to just create little transform functions here and there, I can imagine there being thousands of those in your account without anybody knowing that they even exist.

Ben: Now, don't get me wrong. I would love to have the ability to drop in little transforms that did not involve Lambda functions. So in other words, I mean, the thing that VTL does for API Gateway, REST APIs but without it being VTL and being ... Because that's hard and then also restricted in what you can do, right? It's not, "Oh, I can drop in arbitrary code in here." But enough to say, "Oh, I want to flip ... These fields should go from a key-value mapping to a list of key-value, right? In the way that it addresses inconsistent with how tags are defined across services, those kinds of things. Right? And you could drop that in any service, but once you've defined it, there's no maintenance for you, right?

You're writing JavaScript. It's not actually a JavaScript engine underneath or something. It's just getting translated into some big multi-tenant fancy thing. And I have a hypothesis that that should be possible. You should be able to do it where you could even do it in the parsing of JSON, being able to do transforms without ever having to have the whole object in memory. And if we could get that then, "Oh, sure. Now I have sprinkled all over the place all of these little transforms." Now there's a little bit of overhead if the transform is defined correctly or not, right? But once it is, then it just works. And having all those little transforms everywhere is then fine, right? And that incentive to make it harder it doesn't need to be there because it's not bringing ownership with it.

Rebecca: Yeah. It's almost like taking the idea of tomorrow-Jeremy's problem and actually switching it to say tomorrow-Jeremy's celebration where tomorrow-Jeremy gets to look back at past-Jeremy and be like, "Nice. Thank you for making that decision past-Jeremy." Because I think we often do look at it in terms of tomorrow-Jeremy will think of this, we'll solve this problem rather than how do we approach it by saying, how do I make tomorrow-Jeremy thankful for it today-Jeremy? And that's a simple language, linguistic switch, but a hard switch to actually make decisions based on.

Ben: Yeah. I don't think tomorrow-Ben is ever thankful for today-Ben. I think it's tomorrow-Ben is thankful for yesterday-Ben setting up the incentives correctly so that today-Ben will do the right thing for tomorrow-Ben. Right? When I think about people, I think it's easier to convince people to accept a change in their incentives than to convince them to fight against their incentives sustainably.

Jeremy: Right. And I think developers and I'm guilty of this too, I mean, we make decisions based off of expediency. We want to get things done fast. And when you get stuck on that problem you're like, "You know what? I'm not going to figure it out. I'm just going to write a loop or I'm going to do whatever I can do just to make it work." Another if statement here, "Isn't going to hurt anybody." All right. So let's move to ... Sorry, go ahead.

Ben: We shouldn't feel bad about that.

Jeremy: You're right.

Ben: I was going to say, we shouldn't feel bad about that. That's where I don't want tomorrow-Ben to have to be thankful for today-Ben, because that's the implication there is that today-Ben is fighting against his incentives to do good things for tomorrow-Ben. And if I don't need to have to get to that point where just the right path is the easiest path, right? Which means putting friction in the right places than today-Ben ... It's never a question of whether today-Ben is doing something that's worth being thankful for. It's just doing the job, right?

Jeremy: Right. No, that makes sense. All right. I got another question here, I think falls under the category of service discovery, which I know is another topic that you love. So this person said, "I love IaC, but hate the fuzzy boundaries where certain software awkwardly fall. So like Istio and Prometheus and cert-manager. That they can be considered part of the infrastructure, but then it's awkward to deploy them when something like Terraform due to circular dependencies relating to K8s and things like that."

So, I mean, I know that we don't have to get into the actual details of that, but I think that is an important aspect of infrastructure as code where best practices sometimes are deploy a stack that has your permanent resources and then deploy a stack that maybe has your more femoral or the ones that are going to be changing, the more mutable ones, maybe your Lambda functions and some of those sort of things. If you're using Terraform or you're using some of these other services as well, you do have that really awkward mix where you're trying to use outputs from one stack into another stack and trying to do all that. And really, I mean, there are some good tools that help with it, but I mean just overall thoughts on that.

Ben: Well, we certainly need to demand better of AWS services when they design new things that they need to be designed so that infrastructure as code will work. So this is the S3 bucket notification problem. A very long time ago, S3 decided that they were going to put bucket notifications as part of the S3 bucket. Well, CloudFormation at that point decided that they were going to put bucket notifications as part of the bucket resource. And S3 decided that they were going to check permissions when the notification configuration is defined so that you have to have the permissions before you create the configuration.

This creates a circular dependency when you're hooking it up to anything in CloudFormation because the dependency depends on the resource policy on an SNS topic, and SQS queue or a Lambda function depends on the bucket name if you're letting CloudFormation name the bucket, which is the best practice. Then bucket name has to exist, which means the resource has to have been created. But the notification depends on the thing that's notifying, which doesn't have the names and the resource policy doesn't exist so it all fails. And this is solved in a couple of different ways. One of which is name your bucket explicitly, again, not a good practice. Another is what SAM does, which says, "The Lambda function will say I will allow all S3 buckets to invoke me."

So it has a star permission in it's resource policy. So then the notification will work. None of which is good or there's custom resources that get created, right? Now, if those resources have been designed with infrastructure as code as part of the process, then it would have been obvious, "Oh, you end up with a circular pendency. We need to split out bucket notifications as a separate resource." And not enough teams are doing this. Often they're constrained by the API that they develop first ...

Jeremy: That's a good point.

Ben: ... they come up with the API, which often makes sense for a Console experience that they desire. So this is where API Gateway has this whole thing where you create all the routes and the resources and the methods and everything, right? And then you say, "Great, deploy." And in the Console you only need one mutable working copy of that at a time, but it means that you can't create two deployments or update two stages in parallel through infrastructure as code and API Gateway because they both talk to this mutable working copy state and would overwrite each other.

And if infrastructure as code had been on their list would have been, "Oh, if you have a definition of your API, you should be able to go straight to the deployment," right? And so trying to push that upstream, which to me is more important than infrastructure as code support at launch, but people are often like, "Oh, I want CloudFormation support at launch." But that often means that they get no feedback from customers on the design and therefore make it bad. KMS asymmetric keys should have been a different resource type so that you can easily tell which key types are in your template.

Jeremy: Good point. Yeah.

Ben: Right? So that you can use things like CloudFormation Guard more easily on those. Sure, you can control the properties or whatever, but you should be able to think in terms of, "I have a symmetric key or an asymmetric key in here." And they're treated completely separately because you use them completely differently, right? They don't get used to the same place.

Jeremy: Yeah. And it's funny that you mentioned the lacking support at launch because that was another complaint. That was quite prevalent in this thread here, was people complaining that they don't get that CloudFormation support right away. But I think you made a very good point where they do build the APIs first. And that's another thing. I don't know which question asked me or which one of these mentioned it, but there was a lot of anger over the fact that you go to the API docs or you go to the docs for AWS and it focuses on the Console and it focuses on the CLI and then it gives you the API stuff and very little mention of CloudFormation at all. And usually, you have to go to a whole separate set of docs to find the CloudFormation. And it really doesn't tie all the concepts together, right? So you get just a block of JSON or of YAML and you're like, "Am I supposed to know what everything does here?"

Ben: Yeah. I assume that's data-driven. Right? And we exist in this bubble where everybody loves infrastructure as code.

Jeremy: True.

Ben: And that AWS has many more customers who set things up using Console, people who learn by doing it first through the Console. I assume that's true, if it's not, then the AWS has somehow gotten on the extremely wrong track. But I imagine that's how they find that they get the right engagement. Now maybe the CDK will change some of this, right? Maybe the amount of interest that is generating, we'll get it to the point where blogs get written with CDK programs being written there. I think that presents different problems about what that CDK program might hide from when you're learning about a service. But yeah, it's definitely not ... I wrote a blog for AWS and my first draft had it as CloudFormation and then we changed it to the Console. Right? And ...

Jeremy: That must have hurt. Did you die a little inside when that happened?

Ben: I mean, no, because they're definitely our users, right? That's the way in which they interact with data, with us and they should be able to learn from that, their company, right? Because again, developers are often not fully in control of this process.

Jeremy: Right. That's a good point.

Ben: And so they may not be able to say, "I want to update this through CloudFormation," right? Either because their organization says it or just because their team doesn't work that way. And I think AWS gets requests to prevent people from using the Console, but also to force people to use the Console. I know that at least one of them is possible in IAM. I don't remember which, because I've never encountered it, but I think it's possible to make people use the Console. I'm not sure, but I know that there are companies who want both, right? There are companies who say, "We don't want to let people use the API. We want to force them to use the Console." There are companies who say, "We don't want people using the Console at all. We want to force them to use the APIs."

Jeremy: Interesting.

Ben: Yeah. There's a lot of AWS customers, right? And there's every possible variety of organization and AWS should be serving all of them, right? They're all customers. And certainly, I want AWS to be leading the ones that are earlier in their cloud journey and on the serverless ladder to getting further but you can't leave them behind, I think it's important.

Jeremy: So that people argument and those different levels and coming in at a different, I guess, level or comfortability with APIs versus infrastructure as code and so forth. There was another question or another comment on this that said, "I love the idea of committing everything that makes my solution to text and resurrect an entire solution out of nothing other than an account key. Loved the ability to compare versions and unit tests, every bit of my solution, and not having to remember that one weird setting if you're using the Console. But hate that it makes some people believe that any coder is now an infrastructure wizard."

And I think this is a good point, right? And I don't 100% agree with it, but I think it's a good point that it basically ... Back to your point about creating these little transformations in Pulumi, you could do a lot of damage, I mean, good or bad, right? When you are using these tools. What are your thoughts on that? I mean, is this something where ... And again, the CDK makes it so easy for people to write these constructs pretty quickly and spin up tons of infrastructure without a lot of guard rails to protect them.

Ben: So I think if we tweak the statement slightly, I think there's truth there, which isn't about the self-perception but about what they need to be. Right? That I think this is more about serverless than about infrastructure as code. Infrastructure as code is just saying that you can define it. Right? I think it's more about the resources that are in a particular definition that require that. My former colleague, Aaron Camera says, "Serverless means every developer is an architect" because you're not in that situation where the code you write goes onto something, you write the whole thing. Right?

And so you do need to have those ... You do need to be an infrastructure wizard whether you're given the tools to do that and the education to do that, right? Not always, like if you're lucky. And the self-perception is again an even different thing, right? Especially if coders think that there's nothing to be learned ... If programmers, software developers, think that there's nothing to be learned from the folks who traditionally define the infrastructure, which is Ops, right? They think, "Those people have nothing to teach me because now I can do all the things that they did." Well, you can create the things that they created and it does not mean that you're as good at it ...

Jeremy: Or responsible for monitoring it too. Right.

Ben: ... and have the ... Right. The monitoring, the experience of saying these are the things that will come back to bite you that are obvious, right? This is how much ownership you're getting into. There's very much a long-standing problem there of devaluing Ops as a function and as a career. And for my money when I look at serverless, I think serverless is also making the software development easier because there's so much less software you need to write. You need to write less software that deals with the hard parts of these architectures, the scaling, the distributed computing problems.

You still have this, your big computing problems, but you're considering them functionally rather than coding things that address them, right? And so I see a lot of operations folks who come into serverless learn or learn a new programming language or just upscale, right? They're writing Python scripts to control stuff and then they learn more about Python to be able to do software development in it. And then they bring all of that Ops experience and expertise into it and look at something and say, "Oh, I'd much rather have step functions here than something where I'm running code for it because I know how much my script break and those kinds of things when an API changes or ... I have to update it or whatever it is."

And I think that's something that Tom McLaughlin talks about having come from an outside ground into serverless. And so I think there's definitely a challenge there in both directions, right? That Ops needs to learn more about software development to be more engaged in that process. Software development does need to learn much more about infrastructure and is also at this risk of approaching it from, "I know the syntax, but not the semantics, sort of thing." Right? We can create ...

Jeremy: Just because I can doesn't mean I should.

Ben: ... an infrastructure. Yeah.

Rebecca: So Ben, as we're looping around this conversation and coming back to this idea that software is people and that really software should enable you to focus on the things that do matter. I'm wondering if you can perhaps think of, as pristine as possible, an example of when you saw this working, maybe it was while you've been at iRobot or a project that you worked on your own outside of that, but this moment where you saw software really working as it should, and that how it enabled you or your team to focus on the things that matter. If there's a concrete example that you can give when you see it working really well and what that looks like.

Ben: Yeah. I mean, iRobot is a great example of this having been the company without need for software that scaled to consumer electronics volumes, right? Roomba volumes. And needing to build a IOT cloud application to run connected Roombas and being able to do that without having to gain that expertise. So without having to build a team that could deal with auto-scaling fleets of servers, all of those things was able to build up completely serverlessly. And so skip an entire level of organizational expertise, because that's just not necessary to accomplish those tasks anymore.

Rebecca: It sounds quite nice.

Ben: It's really great.

Jeremy: Well, I have one more question here that I think could probably end up ... We could talk about for another hour. So I will only throw it out there and maybe you can give me a quick answer on this, but I actually had another Twitter thread on this not too long ago that addressed this very, very problem. And this is the idea of the feedback cycle on these infrastructure as code tools where oftentimes to deploy infrastructure changes, I mean, it just takes time. In many cases things can run in parallel, but as you said, there's race conditions and things like that, that sometimes things have to be ... They just have to be synchronous. So is this something where there are ways where you see in the future these mutations to your infrastructure or things like that potentially happening faster to get a better feedback cycle, or do you think that's just something that we're going to have to deal with for a while?

Ben: Yeah, I think it's definitely a very extensive topic. I think there's a few things. One is that the deployment cycle needs to get shortened. And part of that I think is splitting dev deployments from prod deployments. In prod it's okay for it to take 30 seconds, right? Or a minute or however long because that's at the end of a CI/CD pipeline, right? There's other things that are happening as part of that. Now, you don't want that to be hours or whatever it is. Right? But it's okay for that to be proper and to fully manage exactly what's going on in a principled manner.

When you're doing for development, it would be okay to, for example, change the Lambda code without going through CloudFormation to change the Lambda code, right? And this is what an architect does, is there's a notion of a dirty deploy which just packages up. Now, if your resource graph has changed, you do need to deploy again. Right? But if the only thing that's changing is your code, sure, you can go and say, "Update function code," on that Lambda directly and that's faster.

But calling it a dirty deploy is I think important because that is not something that you want to do in prod, right? You don't want there to be drift between what the infrastructure as code service understands, but then you go further than that and imagine there's no reason that you actually have to do this whole zip file process. You could be R sinking the code directly, or you could be operating over SSH on the code remotely, right? There's many different ways in which the loop from I have a change in my Lambda code to that Lambda having that change could be even shorter than that, right?

And for me, that's what it's really about. I don't think that local mocking is the answer. You and Brian Rue were talking about this recently. I mean, I agree with both of you. So I think about it as I want unit tests of my business logic, but my business logic doesn't deal with AWS services. So I want to unit test something that says, "Okay, I'm performing this change in something and that's entirely within my custom code." Right? It's not touching other services. It doesn't mean that I actually need adapters, right? I could be dealing with the native formats that I'm getting back from a given service, but I'm not actually making calls out of the code. I'm mocking out, "Well, here's what the response would look like."

And so I think that's definitely necessary in the unit testing sense of saying, "Is my business logic correct? I can do that locally. But then is the wiring all correct?" Is something that should only happen in the cloud. There's no reason to mock API gateway into Lambda locally in my mind. You should just be dealing with the Lambda side of it in your local unit tests rather than trying to set up this multiple thing. Another part of the story is, okay, so these deploys have to happen faster, right? And then how do we help set up those end-to-end test and give you observability into it? Right? X-Ray helps, but until X-Ray can sort through all the services that you might use in the serverless architecture, can deal with how does it work in my Lambda function when it's batching from Kinesis or SQS into my function?

So multiple traces are now being handled by one invocation, right? These are problems that aren't solved yet. Until we get that kind of inspection, it's going to be hard for us to feel as good about cloud development. And again, this is where I feel sometimes there's more friction there, but there's bigger payoff. Is one of those things where again, fighting against your incentives which is not the place that you want to be.

Jeremy: I'm going to stop you before you disagree with me anymore. No, just kidding! So, Rebecca, you have any final thoughts or questions for Ben?

Rebecca: No. I just want to say to both of you and to everyone listening that I hope your today self is celebrating your yesterday-self right now.

Jeremy: Perfect. Well, Ben, thank you so much for joining us and being a guinea pig as we said on this new format that we are trying. Excellent guinea pig. Excellent.

Rebecca: An excellent human too but also great guinea pig.

Jeremy: Right. Right. Pretty much so. So if people want to find out more about you, read some of the stuff you're doing and working on, how do they do that?

Ben: I'm on Twitter. That's the primary place. I'm on LinkedIn, I don't post much there. And then I write articles that show up on Medium.

Rebecca: And just so everyone knows your Twitter handle I'll say it out loud too. It's @ben11kehoe, K-E-H-O-E, ben11kehoe.

Jeremy: Right. Perfect. All right. Well, we will put all that in the show notes and hopefully people will like this new format. And again, we'd love your feedback on this, things that you'd like us to do in the future, any ideas you have. And of course, make sure you reach out to Ben. He's an amazing resource for serverless. So again, thank you for everything you do, and thank you for being on the show.

Ben: Yeah. Thanks so much for having me. This was great.

Rebecca: Good to see you. Thank you.

View Details

About Nader Dabit
Nader Dabit is a web and mobile developer, author, and Developer Relations Engineer building the decentralized future at Edge and Node. Previously, he worked as a Developer Advocate at AWS Mobile working with projects like AWS AppSync and AWS Amplify. He is also the author and editor of React Native in Action and OpenGraphQL.

Nader Dabit Twitter: @dabit3
Edge and Node Twitter: @edgeandnode
Graph protocol Twitter: @graphprotocol
Edge and Node: edgeandnode.com
Everest: everest.link
YouTube: YouTube.com/naderdabit
What is Web3? The Decentralized Internet of the Future Explained

Watch this episode on YouTube: https://youtu.be/pSv_cCQyCPQ

This episode is sponsored by CBT Nuggets and Fauna.

Transcript
Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I am joined again by Nader Dabit. Hey Nader, thanks for joining me.

Nader: Hey Jeremy. Thanks for having me.

Jeremy: You are now a developer relations engineer at Edge & Node. I would love it if you could tell the listeners a little bit about yourself. I think a lot of people probably know you already, but a little bit about your background and then what Edge & Node is.

Nader: Yeah, totally. My name is Nader Dabit like you mentioned, and I've been a developer for about, I guess, nine or ten years now. A lot of people might know me from my work with AWS, where I worked with the Amplify team with the front end web and mobile team, doing a lot of full stack stuff there as well as serverless. I've been working as a developer relations person, developer advocate, actually, leading the front end web and mobile team at AWS for a little over three years I was there. I was a manager for the last year and I became really, really interested in serverless while I was there. It led to me writing a book, which is Full Stack Serverless. It also just led me down the rabbit hole of managed services and philosophy and all this stuff.

It's been really, really cool to learn about everything in the space. Edge & Node is my next step, I would say, in doing work and what I consider maybe a serverless area, but it's an area that a lot of people might not associate with the traditional, I would say definition of serverless or the types of companies they often associate with serverless. But Edge & Node is a company that was spun off from a team that created a decentralized API protocol, which is called the Graph protocol. And the Graph protocol started being built in 2017. It was officially launched in a decentralized way at the end of 2020. Now we are currently finalizing that migration from a hosted service to a decentralized service actually this month.

A lot of really exciting things going on. We'll talk a lot about that and what all that means. But Edge & Node itself, we do support the Graph protocol, that's part of what we do, but we also build out decentralized applications ourselves. We have a couple of applications that we're building as engineers. We're also doing a lot of work within the Web3 ecosystem, which is known as the decentralized web ecosystem by investing in different people and companies and supporting different things and spreading awareness around some of the things that are going on here because it does have a lot to do with maybe the work that people are doing in the Web2 space, which would be the traditional webspace, the space that I was in before.

Jeremy: Right, right. Here I am. I follow you on Twitter. Love the videos that you do on your YouTube channel. You're like a shining example of what a really good developer relations dev advocate is. You just produce so much content, things like that, and you're doing all this stuff on serverless and I'm loving it. And then all of a sudden, I see you post this thing saying, hey, I'm leaving AWS Amplify. And you mentioned something about blockchain and I'm like, okay, wait a minute. What is this that Nader is now doing? Explain to me this, or maybe explain to me and hopefully the audience as well. What is the blockchain have to do with this decentralized applications or decentralized, I guess Web3?

Nader: Web3 as defined by definition, what you might see if you do some research, would be what a lot of people are talking about as the next evolution of the web as we know it. In a lot of these articles and stuff that people are trying to formalize ideas and stuff, the original web was the read-only web where we were not creators, the only creators were maybe the developers themselves. Early on, I might've gone and read a website and been able to only interact with the website by reading information. The current version that we're currently experiencing might be considered as Web2 where everyone's a creator. All of the interfaces, all of the applications that we interact with are built specifically for input. I can actually create a comment, I can upload a video, I can share stuff, and I can write to the web. And I can read.

And then the next evolution, a lot of people are categorizing, yes, is Web3. It's like taking a lot of the great things that we have today and maybe improving upon those. A lot of people and everyone kind of, this is just a really, a very old discussion around some of the trade-offs that we currently make in today's web around our data, around advertising, around the way a lot of business models are created for monetization. Essentially, they all come down to the manipulation of user data and different tricks and ways to steal people's data and use that essentially to create targeted advertising. Not only does this lead to a lot of times a negative experience. I just saw a tweet yesterday that resonated a lot with me that said, "YouTube is no longer a video platform, it's now an ad platform with videos in between." And that's the way I feel about YouTube. My kids ...

Jeremy: Totally.

Nader: ... I have kids that use YouTube and it's interesting to watch them because they know exactly what to do when the ads come up and exactly how to time it because they're used to, ads are just part of their experience. That's just what they're used to. And it's not just YouTube, it's every site that's out there, that's a social site, Instagram, LinkedIn. I think that that's not the original vision that people had, right, for the web. I don't think this was part of it. There have been a lot of people proposing solutions, but the core fundamental problem is how these applications are engineered, but also how the applications are paid for. How do these companies pay for developers to build. It's a really complex problem that, the simplest solution is just sell ads or maybe create something like a developer platform where you're charging a weekly or monthly or yearly or something like that.

I would say a lot of the ideas around Web3 are aiming to solve this exact problem. In order to do that you have to rethink how we build applications. You have to rethink how we store data. You have to rethink about how we think about identity as well, because again, how do you build an application that deals with user data without making it public in some way? Right? How do we deal with that? A lot of those problems are the things that people are thinking about and building ways to address those in this decentralized Web3 world. It became really fascinating to me when I started looking into it because I'm very passionate about what I'm doing. I really enjoy being a developer and going out and helping other people, but I always felt there was something missing because I'm sitting here and I love AWS still.

In fact, I would 100% go back and work there or any of these big companies, right? Because you can't really look at a company as, in my opinion, a black or white, good or bad thing, there's companies are doing good things and bad things at the same time. For instance, at AWS, I would meet a developer, teach them something at a workshop, a year later they would contact me and be like, hey, I got my first job or I created a business, or I landed my first client. So you're actually helping improve people's lives, at the same time you're reading these articles about Amazon in the news with some of the negative stuff going on. The way that I look at it is, I can't sit there and say any company is good or bad, but I felt a lot of the applications that people were building were also, at the end goal when you hear some of these VC discussions or people raising money, a lot of the end goal for some of the people I was working with were just selling advertising.

And I'm like, is this really what we're here to do? It doesn't feel fulfilling anymore when you start seeing that over and over and over. I think the really thing that fascinated me was that people are actually building applications that are monetized in a different way. And then I started diving into the infrastructure that enabled this and realized that there was a lot of similarities between serverless and how developers would deploy and build applications in this way. And it was the entry point to my rabbit hole.

Jeremy: I talked to you about this and I've been reading some of the stuff that you've been putting out and trying to educate myself on some of this. It seems very much so that show Silicon Valley on HBO, right? This decentralized web and things like that, but there's kind of, and totally correct me if I'm wrong here, but I feel there's two sides of this. You've got one side that is the blockchain, that I think some people are familiar with in the, I guess in the context of cryptocurrency, right? This is a very popular use of the blockchain because you have that redundancy and you have the agreement amongst multiple places, it's decentralized. And so you have that security there around that. But there's other uses for the blockchain as well.

Especially things like banking and real estate and some of those other use cases that I'd like to talk about. And then there's another side of it that is this decentralized piece. Is the decentralized piece of it like building apps? How is that related to the blockchain or are those two separate things?

Nader: Yeah, absolutely. I'm a big fan of Silicon Valley. Working in tech, it's almost like every single episode resonates with you if you've been in here long enough because you've been in one of those situations. The blockchain is part of the discussion. Crypto is part of the discussion, and those things never really interested me, to be honest. I was a speculator in crypto from 2015 until now. It's been fun, but I never really looked at crypto in any other way other than that. Blockchain had a really negative, I would say, association in my mind for a long time, I just never really saw any good things that people were doing with it. I just didn't do any research, maybe didn't understand what was going on.

When I started diving into it originally what really got me interested is the Graph protocol, which is one of the things that we work on at Edge & Node. I started actually understanding, why does this thing exist? Why is it there? That led me to understanding why it was there and the fact that 90% of dApps, decentralized apps in the Ethereum ecosystem are using it. And billions of queries, companies with billions of dollars in transactions are all using this stuff. I'm like, okay, this whole world exists, but why does it exist? I guess to give you an example, I guess we can talk about the Graph protocol. And there are a lot of other web, I would say Web3 or decentralized infrastructure protocols that are out there that are similar, but they all are doing similar things in the sense of how they're actually built and how they allow participation and stuff like that.

When you think of something like AWS, you think of, AWS has all of these different services. I want to build an app, I need storage. I need some type of authentication layer, maybe with Cognito, and then maybe I need someplace to execute some business logic. So maybe I'll spin up some serverless functions or create an EC2 instance, whatever. You have all these building blocks. Essentially what a lot of these decentralized protocols like the Graph are doing, are building out the same types of web infrastructure, but doing so in a decentralized way. Why does that even matter? Why is that important? Well, for instance, when you live, let's say for example in another country, I don't know, in South America and outside the United States, or even in the United States in the future, you never know. Let's say that you have some application and you've said something rude about maybe the president or something like that.

Let's say that for whatever reason, somebody hacks the server that you're dealing with or whatever, at the end of the day, there is a single point of failure, right? You have your data that's controlled by the cloud provider or the government can come in and they can have control over that. The idea around some of, pretty much all of the decentralized protocols is that they are built and distributed in a way that there is no single point of failure, but there's also no single point of control. That's important when you're living in areas that have to even worry about stuff like that. So maybe we don't have to worry about that as much here, but in other countries, they might.

Building something like a server is not a big deal, right? With AWS, but how would you build a server and make it available for anyone in the world to basically deploy and do so in a decentralized way? I think that's the problem that a lot of these protocols are trying to solve. For the Graph in particular, if you want to build an application using data that's stored on a blockchain. There's a lot of applications out there that are basically using the blockchain for mainly, right now it's for financial, transactional reasons because a lot of the transactions actually cost a lot of money. For instance, Uniswap is one of these applications. If you want to basically query data from a blockchain, it's not as easy as querying data from a traditional server or database.

For us we are used to using something like DynamoDB, or some type of SQL database, that's very optimized for queries. But on the blockchain, you're basically having these blocks that add up every time. You create a transaction, you save it. And then someone comes behind them and they save another transaction. Over time you build up this data that's aggregated over time. But let's say you want to hit that database with the, quote-unquote, database with a query and you want to retrieve data over time, or you want to have some type of filtering mechanism. You can't do that. You can't just query blockchains the way you can from a regular database. Similar to how a database basically indexes data and stores it and makes it efficient for retrieval, the Graph protocol basically does that, but for blockchain data.

Anyone that wants to build an application, one of these decentralized apps on top of blockchain data has a couple of options. They can either build their own indexing server and deploy it to somewhere like AWS. That takes away the whole idea of decentralization because then you have a single point of failure again. You can query data directly from the blockchain, from your client application, which takes a very long time. Both of those are not, I would say the most optimal way to build. But also if you're building your own indexing server, every time you want to come up with a new idea also, you have to think about the resources and time that go into it. Basically, I want to come up with a new idea and test it out, I have to basically build a server index, all this data, create APIs around it. It's time-intensive.

What the Graph protocol allows you to do is, as a developer you can basically define a subgraph using YAML, similar to something like cloud formation or a very condensed version of that maybe more Serverless Framework where you're defining, I want to query data from this data source, and I want to save these entities and you deploy that to the network. And that subgraph will basically then go and look into that blockchain. And will look for all the transactions that have happened, and it will go ahead and save those and make those available for public retrieval. And also, again, one of the things that you might think of is, all of this data is public. All of the data that's on the blockchain is public.

Jeremy: Right. Right. All right. Let me see if I could repeat what you said and you tell me if I'm right about this. Because this was one of those things where blockchain ... you're right. To me, it had a negative connotation. Why would you use the blockchain, unless you were building your own cryptocurrency? Right. That just seemed like that's what it was for. Then when AWS comes out with QLDB or they announced that or whatever it was. I'm like, okay, so this is interesting, but why would you use it, again, unless you're building your own cryptocurrency or something because that's the only thing I could think of you would use the blockchain for.

But as you said, with these blockchains now, you have highly sensitive transactions that can be public, but a real estate transaction, for example, is something really interesting, where like, we still live in a world where if Bank of America or one of these other giant banks, JPMorgan Chase or something like that gets hacked, they could wipe out financial data. Right? And I know that's backed up in multiple regions and so forth, but this is the thing where if you're doing some transaction, that you want to make sure that transaction lives forever and isn't manipulated, then the blockchain is a good place to do that. But like you said, it's expensive to write there. But it's even harder to read off the blockchain because it's that ledger, right? It's just information coming in and coming in.

So event storming or if you were doing event sourcing or something that, it's that idea. The idea with these indexers are these basically separate apps that run, and again, I'm assuming that these protocols, their software, and things that you don't have to build this yourself, essentially you can just deploy these things. Right? But this will read off of the blockchain and do that aggregation for you and then make that. Basically, it caches the blockchain. Right? And makes that available to you. And that you could deploy that to multiple indexers if you wanted to. Right? And then you would have access to that data across multiple providers.

Nader: Right. No single point of failure. That's exactly right. You basically deploy a very concise configuration file that defines how you want your data stored and made available. And then it goes, and it just starts at the very beginning and it queries all those blocks or reads all those blocks, saves the data in a database, and then it keeps up with additional new updates. If someone writes a new transaction after that, it also saves that and makes it available for efficient retrieval. This is just for blockchain data. This is the data layer for, but it's not just a blockchain data in the future. You can also query from IPFS, which is a file storage layer, somewhat S3. You can query from other chains other than Ethereum, which is kind of like the main chamber.

In the future really what we're hoping to have is a complete API on top of all public data. Anybody that wants to have some data set available can basically deploy a subgraph and index it and then anyone can then essentially query for it. It's like when you think of public data, we're not really used to thinking of data in this way. And also I think a good thing to talk about in a moment is the types of apps that you can build because you wouldn't want to store private messages on a blockchain or something like that. Right? The types of apps that people are building right now at least are not 100% in line with everything. You can't do everything I would say right now in Web3 that you can do in Web2.

There are only certain types of applications, but those applications that are successful seem to be wildly successful and have a lot of people interested in them and using them. That's the general idea, is like you have this way to basically deploy APIs and the technology that we use to query is GraphQL. That was one of the reasons that I became interested as well. Right now the main data sources are blockchains like Ethereum, but in the future, we would like to make that available to other data sources as well.

Jeremy: Right. You mentioned earlier too because there are apps obviously being built on this that you said are successful. And the problem though, I think right now, because I remember I speculated a little bit with Bitcoin and I bought a whole bunch of Ripple, so I'm still hanging on to it. Ripple XPR whatever, let's go. Anyways, but it was expensive to make a transaction. Right? Reading off of the blockchain itself, I think just connecting generally doesn't cost money, but if you're, and I know there's some costs with indexers and that's how that works. But in terms of the real cost, it's writing to the blockchain. I remember moving some Bitcoin at one point, I think cost me $30 to make one transaction, to move something like that.

I can see if you're writing a $300,000 real estate transaction, or maybe some really large wire transfer or something that you want to record, something that makes sense where you could charge a fee of $30 or $40 in order to do that. I can't see you doing that for ... certainly not for web streaming or click tracking or something like that. That wouldn't make sense. But even for smaller things there might be writing more to it, $30 or whatever that would be ... seems quite expensive. What's the hope around that?

Nader: That was one of the biggest challenges and that was one of the reasons that when I first, I would say maybe even considered this as a technology back in the day, that I would be considering as something that would possibly be usable for the types of applications I'm used to seeing. It just was like a no-brainer, like, no. I think right now, and that's one of the things that attracted me right now to some of the things that are happening, is a lot of those solutions are finally coming to fruition for fixing those sorts of things. There's two things that are happening right now that solve that problem. One of them is, they are merging in a couple of updates to the base layer, layer one, which would be considered something like Ethereum or Bitcoin. But Ethereum is the main one that a lot of the financial stuff that I see is happening.

Basically, there are two different updates that are happening, I think the main one that will make this fee transactional price go down a little bit is sharding. Sharding is basically going to increase the number of, I believe nodes that are basically able to process the transactions by some number. Basically, that will reduce the cost somewhat, but I don't think it's ever going to get it down to a usable level. Instead what the solutions seem to be right now and one of the solutions that seems to actually be working, people are using it in production really recently, this really just started happening in the last couple of months, is these layer 2 solutions. There are a couple of different layer 2 solutions that are basically layers that run on top of the layer one, which would be something like Ethereum.

And they treat Ethereum as the settlement layer. It's almost like when you interact with the bank and you're running your debit card. You're probably not talking to the bank directly and they are doing that. Instead, you have something like Visa who has this layer 2 on top of the banks that are managing thousands of transactions per second. And then they take all of those transactions and they settle those in an underlying layer. There's a couple different layer 2s that seem to be really working well right now in the Ethereum ecosystem. One of those is Arbitrum and then the other is I think Matic, but I think they have a different name now. Both of those seem to be working and they bring the cost of a transaction down to a fraction of a penny.

You have, instead of paying $20 or $30 for a transaction, you're now paying almost nothing. But now that's still not cheap enough to probably treat a blockchain as a traditional database, a high throughput database, but it does open the door for a lot of other types of applications. The applications that you see building on layer one where the transactions really are $5 to $20 or $30 or typically higher value transactions. Things like governance, things like financial transactions, you've heard of NFTs. And that might make sense because if someone's going to spend a thousand bucks or 500 bucks, whatever ...

Jeremy: NFTs don't make sense to me.

Nader: They're not my thing either, the way they're being, I would say, talked about today especially, but I think in the future, the idea behind NFTs is interesting, but yeah, I'm in the same boat as you. But still to those people, if you're paying a thousand dollars for something then that 5 or 10 or 20 bucks might make sense, but it's not going to make sense if I just want to go to an e-commerce store and pay $5 for something. Right? I think that these layer 2s are starting to unlock those potential opportunities where people can start building these true financial applications that allow these transactions to happen at the same cost or actually a lot cheaper maybe than what you're paying for a credit card transaction, or even what those vendors, right? If you're running a store, you're paying percentages to those companies.

The idea around decentralization comes back to this discussion of getting rid of the middleman, and a lot of times that means getting rid of the inefficiencies. If you can offload this business logic to some type of computer, then you've basically abstracted away a lot of inefficiencies. How many billions of dollars are spent every year by banks flying their people around the world and private jets and these skyscrapers and stuff. Now, where does that money come from? It comes from the consumer and them basically taking fees. They're taking money here and there. Right? That's the idea behind technology in general. They're like whenever something new and groundbreaking comes in, it's often unforeseen, but then you look back five years later and you're like, this is a no-brainer. Right?

For instance Blockbuster and Netflix, there's a million of them. I don't have to go into that. I feel this is what that is for maybe the financial institutions and how we think about finance, especially in a global world. I think this was maybe even accelerated by COVID and stuff. If you want to build an application today, imagine limiting yourself to developers in your city. Unless you're maybe in San Francisco or New York, where that might still work. If I'm here in Mississippi and I want to build an application, I'm not going to just look for developers in a 30-mile radius. That is just insane. And I don't use that word mildly, it's just wild to think about that. You wouldn't do that.

Instead, you want to look in your nation, but really you might want to look around the world because you now have things like Slack and Discord and all these asynchronous ways of doing work. And you might be able to find the best developer in the world for 25% or 50% of what you would typically find locally and an easy way to pay them might just be to just send them some crypto. Right? You don't have to go find out all their banking information and do all the wiring and all this other stuff. You just open your wallet, you send them the money and that's it. It's a done deal. But that's just one thing to think about. To me when I think about building apps in Web2 versus Web3, I don't think you're going to see the Facebook or Instagram use case anytime in the next year or two. I think the killer app for right now, it's going to be financial and e-commerce stuff.

But I do think in maybe five years you will see someone crack that application for, something like a social media app where we're basically building something that we use today, but maybe in a better way. And that will be done using some off-chain storage solution. You're not going to be writing all these transactions again to a blockchain. You're going to have maybe a protocol like Graph that allows you to have a distributed database that is managed by one of these networks that you can write to. I think the ideas that we're talking about now are the things that really excite me anyway.

Jeremy: Let's go back to GraphQL for a second, though. If you were going to build an app on top of this, and again, that's super exciting getting those transaction fees down, because I do feel every time you try to move money between banks or it's the $3 fee, if you go to a foreign ATM and you take money out of an ATM, they charge you. Everybody wants to take a cut somewhere along, and there's probably reasons for it, but also corporate jets cost money. So that makes sense as well. But in terms of the GraphQL protocol here, so if I wanted to build an application on top of it, and maybe my application doesn't write to the blockchain, it just reads from it, with one of these indexers, because maybe I'm summing up some financial transactions or something, or I've got an app we can look things up or whatever, I'm building something.

I'm querying using the GraphQL, this makes sense. I have to use one of these indexers that's aggregating that data for me. But what if I did want to write to the blockchain, can I use GraphQL to do a mutation and actually write something to the blockchain? Or do I have to write to it directly?

Nader: Yeah, that's actually a really, really good question. And that's one of the things that we are currently working on with the Graph. Right now if you want to write a transaction, you typically are going to be using one of these JSON RPC wallets and using some type of client library that interacts with the wallet and signs the transaction with the private key. And then that sends the transaction to the blockchain directly. And you're talking to the blockchain and you're just using something like the Graph to query. But I think what would be ideal and what we think would be ideal, is if someone could use a single technology, a single language, and a single abstraction to do everything, not only with reading and writing but also with subscriptions for real-time updates.

That's where we think the whole idea for this will ultimately be, and that's what we're working on now. Right now you can only query. And if you want to write a transaction, you basically are still going to be using something like ethers.js or Web3 or one of these other libraries that allows you to sign a transaction using your wallet. But in the future and in fact, we're already building this right now as having an end-to-end GraphQL library that allows you to write transactions as well as read. That way someone just learns a single API and it's a lot easier. It would also make it easier for developers that are coming from a traditional web background to come in because there's a little bit of learning curve for understanding how to create one of these signed providers and write the transaction. It's not that much code, but it is a new way of thinking about things.

Jeremy: Well I think both of us coming from the serverless space, we know that new way of thinking about things certainly can throw a wrench in the system when a new developer is trying to pick that stuff up.

Nader: Yeah.

Jeremy: All right. So that's the blockchain side of things with the data piece of it. I think people could wrap their head around that. I think it makes a lot of sense. But I'm still, the decentralized, the other things that you talked about. You mentioned an S3, something that's sort of an S3 type protocol that you can use. And what are some of the other ones? I think I've written some of them down here. Acash was one, Filecoin, Livepeer. These are all different protocols or services that are hosted by the indexers, or is this a different thing than the indexers? How does that work? And then how would you use that to save data, maybe save some blob, a blob storage or something like that?

Nader: Let's talk about the tokenomics idea around how crypto fits into this and how it actually powers a protocol like this. And then we'll talk about some of those other protocols. How do people actually build all this stuff and do it for, are they getting paid for it? Is it free? How does that work and how does this network actually stay up? Because everything costs money, developers' time costs money, and so on and so forth. For something like the Graph, basically during the building phase of this protocol, basically, there was white papers and there was blog posts, and there was people in Discords talking about the ideas that were here. They basically had this idea to build this protocol. And this is a very typical life cycle, I would say.

You have someone that comes up with an idea, they document some of it, they start building it. And the people that start building it are going to be basically part of essentially the founding team you could think of, in the sense of they're going to be having equity. Because at the end of the day, to actually launch one of these decentralized protocols, the way that crypto comes into it, there's typically some type of a token offering. The tokens need to be for a network like this, some type of utility token to keep the network running in the future. You're not just going to create some crypto and that's it like, right? I think that's the whole idea that I thought was going on when in reality, these tokens are typically used for powering the protocol.

But let's say early on you have let's say 20 developers and they all build 5% of the system, whatever percentage that you want to talk about, whatever. Let's say you have these people helping out and then you actually build the thing and you want to go ahead and launch it and you have something that's working. A lot of times what people will do is they'll basically have a token offering, where they'll basically say, okay, let's go ahead and we're going to mint X number of tokens, and we're going to put these on the market and we're going to also pay these people that helped build this system, X number of tokens, and that's going to be their payment. And then they can go and sell those or keep those or trade those or whatever they would like to do.

And then you have the tokens that are then put on the public market essentially. Once you've launched the protocol, you have to have tokens to basically continue to power the protocol and fund it. There are different people that interact with the protocol in different ways. You have the indexers themselves, which are basically software engineers that are deploying whatever infrastructure to something like AWS or GCP. These people are still using these cloud providers or they're maybe doing it at their house, whatever. All you basically need is a server and you want to basically run this indexer node, which is software that is open source, and you run this node. Basically, you can go ahead and say, okay, I want to start being an indexer and I want to be one of the different nodes on the network.

To do that you basically buy some GRT, Graph Token, and in our case you stake it, meaning you are putting this money up to basically affirm that you are an indexer on the protocol and you are going to be accepting subgraph developers to deploy their subgraphs to your indexer. You stake that money and then when people use the API, they're basically paying money just like they might pay money to somewhere like API gateway or AppSync. Instead, they're paying money for their subgraph and that money is paid in GRT and it's distributed to the people in the ecosystem. Like me as a developer, I'm deploying the subgraph, and then if I have a million people using it, then I make some money. That's one way to use tokens in the system.

Another way is basically to, as an outside person looking in, I can say, this indexer is really, really good. They know what they're doing. They're a very strong engineer. I'm going to basically put some money into their indexer and I'm basically backing them as an indexer. And then I will also share the money that comes in from the query fees. And then there are also people that are subgraph developers, which is the stuff that I've been working with mainly, where I can basically come up with a new API. I can be like, it'd be cool if I took data from this blockchain and this file system and merged it together, and I made this really cool API that people can use to build their apps with. I can deploy that. And basically, people can signal to this subgraph using tokens. And when people do that, they can say that they believe that this is a good subgraph to use.

And then when people use that, I can also make money in that way. Basically, people are using tokens to be part of the system itself, but also to use that. If I'm a front end application like Uniswap and I want to basically use the Graph, I can basically say, okay, I'm going to put a thousand dollars in GRT tokens and I'm going to be using this API endpoint, which is a subgraph. And then all of the money that I have put up as someone that's using this, is going to be taken as the people start using it. Let's say I have a million queries and each query is one, 1000th of a cent, then after those million queries are up, I've spent $100 or something like that. Kind of similar to how you might pay AWS, you're now paying, you know, subgraph developers and indexers.

Jeremy: Right. Okay. That makes sense. So then that's the payment method of that. So then these other protocols that get built on top of it, the Acash and Filecoin and Livepeer. So those ...

Nader: They're all operating in a very similar fashion.

Jeremy: Okay. All right. And so it's ...

Nader: They have some type of node software that's run and people can basically run this node on some server somewhere and make it available as part of the network. And then they can use the tokens to participate. There's Filecoin for file storage. There's also IPFS, which is actually more of, it's a completely free service, but it's also not something that's as reliable as something like S3 or Filecoin. And then you have, like you mentioned, I believe Acash, which is a way to execute arbitrary code, business logic, and stuff like that. You have Ceramic Network, which is something that you can use for authentication. You have Livepeer which is something you use for live streaming. So you have all these ideas, these decentralized services fitting in these different niches.

Jeremy: Right, right. Okay. So then now you've got a bunch of people. Now you mentioned this idea of, you could say, this is a good indexer. What about bad indexers? Right?

Nader: That's a really good question.

Jeremy: Yeah. You're relying on people to take data off of a public blockchain, and then you're relying on them to process it correctly and give you back good data. I'm assuming they could manipulate that data if they wanted to. I don't know why, but let's say they did. Is there a way to guarantee that you're getting the correct data?

Nader: Yeah. That's a whole part of how the system works. There's this whole idea and this whole, really, really deep rabbit hole of crypto-economics and how these protocols are structured to incentivize and also disincentivize. In our protocol, basically, you have this idea of slashing and this is also a fairly known and used thing in the ecosystem and in the space. It's this idea of slashing. Basically, you incentivize people to go out and find people that are serving incorrect data. And if that person finds someone that's serving incorrect data, then the person that's serving the incorrect data is, quote-unquote, slashed. And that basically means that they're not only not going to receive the money from the queries that they were serving, but they also might lose the money that they put up to be a part of the network.

I mentioned you have to actually put up money to deploy an indexer to the network, that money could also be at risk. You're very, very, very much so financially disincentivized to do that. And there's actually, again, incentives in the network for people to go and find those people. It's all-around incentives, game theory, and things like that.

Jeremy: Which makes a ton of sense. That's good to know. You mentioned, you threw out the number, five years from now, somebody might build the killer app or whatever, they'll figure out some of these things. Where are we with this though? Because this sounds really early, right? There's still things that need to be figured out. Again, it's public data on the blockchain. How do you see this evolving? When do you think Web3 will be more accessible to the masses?

Nader: Today people are actually building really, really interesting applications that are fitting the current technology stack, what are the things that you can build? People are already building those. But when you think about the current state of the web, where you have something like Twitter, or Facebook or Instagram, where I would say, especially maybe something like Facebook, that's extremely, extremely complex with a lot of UI interaction, a lot of private data, messages and stuff. I think to build something like that, yeah, it's going to be a couple of years. And then you might not even see certain types of applications being built. I don't think there is going to be this thing where there is no longer these types of applications. There are only these new types. I think it's more of a new type of application that people are going to be building, and it's not going to be a winner takes all just like in all tech in my opinion.

I wouldn't say all but in many areas of tech where you're thinking of something as a zero-sum game where I don't think this is. But I do think that the most interesting stuff is around how Web3 essentially enables native payments and how people are going to use these native payments in interesting ways that maybe we haven't thought of yet. One of the ways that you're starting to see people doing, and a lot of venture capitalists are now investing in a lot of these companies, if you look at a lot of the companies coming out of YC and a lot of the new companies that these traditional venture capitalists are investing in, are a lot of TOMS crypto companies.

When you think about the financial incentives, the things that we talked about early on, let's say you want to have the next version of YouTube and you don't want to have ads. How would that even work? Right? You still need to enable payments. But there's a couple of things that could happen there. Well, first of all, if you're building an application in the way that I've talked about, where you basically have these native payments or these native tokens that can be part of the whole process now, instead of waiting 10 years to do an IPO for an application that has been around for those 10 years and then paying back all his investors and all of those people that had been basically pulling money out their pockets to take part in.

What if someone that has a really interesting idea and maybe they have a really good track record, they come out with a new application and they're basically saying, okay, if you want to own a piece of this, we're going to basically create a token and you can have ownership in it. You might see people doing these ICO's, initial coin offerings, or whatever, where basically they're offering portions of the company to anyone that wants to own it and then incentivizing people to basically use those, to govern how the application is built in the future. Let's say I own 1% of this company and a proposal is put up to do something new. I can basically say, I can use that portion of my ownership to vote on things. And then people that are speculating can say, this company is doing interesting things. I'm going to buy into it, therefore driving the price up or down.

Kind of like the same way that you see the traditional stock market there, but without all of the regulation and friction that comes with that. I think that's interesting and you're already seeing companies doing that. You're not seeing the majority of companies doing that or anything like that, but you are starting to see those types of things happening. And that brings around the discussion of regulations. Is ... can you even do something like that in the United States? Well, maybe, maybe not. Does that mean people are going to start building these companies elsewhere? That's an interesting discussion as well. Right now if you want to build an application this way, you need to have some type of utility that these tokens are there for. You can't just do them purely on speculation, at least right now. But I think it's going to be interesting for sure, to watch.

Jeremy: Right. And I think too that, I'm just thinking if you're a bank, right? And you maybe have a bunch of private transactions that you want to keep private. Because again, I don't even know how, I don't know how we get to private transactions on the blockchain. I could see you wanting to have some transactions that were public blockchain and some that were private and maybe a hybrid approach would make sense for some companies.

Nader: I think the idea that we haven't really talked about at all is identity and how identity works compared to how we're used to identity. The way that we're used to identity working is, we basically go to a new website and we're like, this looks awesome. Let me try it out. And they're like, oh wait, we need your name, your email address, your phone number, and possibly your credit card and all this other stuff. We do that over and over and over, and over time we've now given our personal information to 500 people. And then you start getting these emails, your data has been breached, every week you get one of these emails, if you're someone like me, I don't know. Maybe I'm just signing up for too much stuff. Maybe not every week, but maybe every month or two. But you're giving out your personal data.

But we're used to identity as being tied to our own physical name and address and things like that. But what if identity was something that was more abstract? And I think that that's the way that you typically see identity managed in Web3. When you're dealing with authentication mechanisms, one of the most interesting things that I think that is part of this whole discussion is this idea of a single sign-on mechanism, that you own your identity and you can transfer it across all the applications and no one else is in control of it. When you use something like an Ethereum wallet, like MetaMask, for example, it's an extension you can just download and put crypto in and basically make payments on the web with. When you create a wallet, you're given a wallet address. And the wallet address is basically created using public key cryptography, where basically you start with this private key, your public key is derived from the private key, and then your address is dropped from the public key.

And when you send a transaction, you basically sign the transaction with your private key and you send your public key along with the transaction, and the person that receives that can decode the transaction with the public key to verify that that's who signed the transaction. Using this public key cryptography that only you can basically sign with your own address and your own password, it's all stored on the blockchain or in some decentralized manner. Actually in this case stored on the blockchain or it depends on how you use it really, I guess. But anyway, the whole idea here is that you completely own your identity. If you never decide to associate that identity with your name and your phone number, then who knows who's sending these transactions and who knows what's going on, because why would you need to associate your own name and phone number with all of these types of things, in these situations where you're making payments and stuff like that. Right?

What is the idea of a user profile anyway, and why do you actually need it? Well, you might need it on certain applications. You might need it or want it on social network, or maybe not, or you might come up with a pseudonym, because maybe you don't want to associate yourself with whatever. You might want to in other cases, but that's completely up to you and you can have multiple wallet addresses. You might have a public wallet address that you associate your name with that you are using on social media. You might have a private wallet address that you're never associating with your name, that you're using for financial transactions. It's completely up to you, but no one can change that information. One of the applications that I recently built was called Decentralized Identity. I built it and release it a few days ago.

And it's an implementation of this and it's using some of these Web3 technologies. One of them is IDX. One of them is Ceramic, which is a decentralized protocol similar to the Graph but for identity. And then it's using something called DIDs, which are decentralized identifiers, which are a way to have a completely unique ID based off of your address. And then you own the control over that. You can basically go in and make updates to that profile. And then any application across the web that you choose to use can then access that information. You're only dealing with it stored in one place. You have full control over it, at any time you can go in and delete that. You can go in and change it. No one has control over it except for you.

The idea of identity is a mind-bending thing in this space because I think we're so used to just handing everybody our real names and our real phone numbers and all of our personal information and just having our fingers crossed, that we're just not used to anything else.

Jeremy: It's all super interesting. You mentioned earlier about, would it be legal in the United States? I'm thinking of all these recent ransomware attacks and I think they were able to trace back some Bitcoin transaction, they were actually able to trace it back to the individual group that accepted the payment. It opens up a whole can of worms. I love this idea of being anonymous and not being tracked, but then it's also like, what could bad actors do with anonymous financial transactions and things like that? So ...

Nader: There kind of has been anonymous transactional layer for a long time. Cash brought in, you can't really do a lot of illegal stuff these days without cash. So should we get rid of cash? I think with any technology ...

Jeremy: No, but I mean, there's a limit though, right? You can't withdraw more than $10,000 worth of cash without the FBI being flagged and you can't deposit more, you know what I mean?

Nader: You can't take a million dollars worth of Bitcoin that you've gotten from ransomware and turn it into cash either.

Jeremy: That's also true. Right.

Nader: Because it's all tracked on the blockchain, that's probably how they caught those people. Right? They somehow had their personal information tied to a transaction, because if you follow these transactions long enough, you're going to find some origination point. I agree though. There's definitely trade-offs with everything. I don't think I'm ever the type to argue that. There's good things and there's bad things. I think you have to look at the whole picture and decide for yourself, what you think. I'm the type that's like, let's lay out all of the ideas and let the market decide.

Jeremy: Right. Yeah. I totally agree with that. All this stuff is fascinating, there is way too much more for me to learn at this point. I think my brain is filled at this point. Anything else about Edge & Node? Any cool things you're working on there or anything you want people to know?

Nader: We're working on a couple of different projects. I can't really talk about some of them because they're not released yet, but we are working on a new version of something called Everest, and Everest is already out. If you want to check it out, it's at everest.link. It's basically a repository of a bunch of different applications that have already been built in the Web3 ecosystem. It also ties in a lot of the stuff that we talked about, like identity and stuff like that. You can basically sign in with your Ethereum wallet. You can basically interact with different applications and stuff, but you can also just see the types of stuff people are building. It's categorized into games, financial apps. If you've listened to this and you're like, this sounds cool, but are people actually building stuff? This is a place to see hundreds of apps that people have are already built and that are out there and successful.

Jeremy: Awesome. All right. Well, listen, Nader, this was awesome. Thank you so much for sharing this with me. I know I learned a ton. I hope the listeners learned a ton. If people want to learn more about this or just follow you and keep up with what you're doing, what's the best way to do that?

Nader: I would say check out Twitter, we're on Twitter @dabit3 for me, @edgeandnode for Edge & Node, and of course @graphprotocol for Graph protocol.

Jeremy: Okay. And then edgeandnode.com. Your YouTube channel is just youtube.com/naderdabit, N-A-D-E-R D-A-B-I-T. And then you had an article on Web3 and I'll put it in the show notes.

Nader: Yeah. Put it in the show notes. For freeCodeCamp, it's called what is Web3. And it's really a condensed version of a lot of the stuff we talked about. Maybe go into a little bit more depth around native payments and how people might build companies in the way that we've talked about here.

Jeremy: Awesome. All right. Well, I will get all that stuff into the show notes. Thanks again, Nader.

Nader: Thanks for having me. It was good to talk.

View Details

About Patrick Strzelec

Patrick Strzelec is a fullstack developer with a focus on building GraphQL gateways and serverless microservices. He is currently working as a technical lead at NorthOne making banking effortless for small businesses.

LinkedIn: Patrick Strzelec
NorthOne Careers: www.northone.com/about/careers

Watch this episode on YouTube: https://youtu.be/8W6lRc03QNU

This episode sponsored by CBT Nuggets and Lumigo.

Transcript
Jeremy
: Hi everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm joined by Patrick Strzelec. Hey, Patrick, thanks for joining me.

Patrick: Hey, thanks for having me.

Jeremy: You are a lead developer at NorthOne. I'd love it if you could tell the listeners a little bit about yourself, your background, and what NorthOne does.

Patrick: Yeah, totally. I'm a lead developer here at NorthOne, I've been focusing on building out our GraphQL gateway here, as well as some of our serverless microservices. What NorthOne does, we are a banking experience for small businesses. Effectively, we are a deposit account, with many integrations that act almost like an operating system for small businesses. Basically, we choose the best partners we can to do things like check deposits, just your regular transactions you would do, as well as any insights, and the use cases will grow. I'd like to call us a very tailored banking experience for small businesses.

Jeremy: Very nice. The thing that is fascinating, I think about this, is that you have just completely embraced serverless, right?

Patrick: Yeah, totally. We started off early on with this vision of being fully event driven, and we started off with a monolith, like a Python Django big monolith, and we've been experimenting with serverless all the way through, and somewhere along the journey, we decided this is the tool for us, and it just totally made sense on the business side, on the tech side. It's been absolutely great.

Jeremy: Let's talk about that because this is one of those things where I think you get a business and a business that's a banking platform. You're handling some serious transactions here. You've got a lot of transactions that are going through, and you've totally embraced this. I'd love to have you take the listeners through why you thought it was a good idea, what were the business cases for it? Then we can talk a little bit about the adoption process, and then I know there's a whole bunch of stuff that you did with event driven stuff, which is absolutely fascinating.

Then we could probably follow up with maybe a couple of challenges, and some of the issues you face. Why don't we start there. Let's start, like who in your organization, because I am always fascinated to know if somebody in your organization says, “Hey we absolutely need to do serverless," and just starts beating that drum. What was that business and technical case that made your organization swallow that pill?

Patrick: Yeah, totally. I think just at a high level we're a user experience company, we want to make sure we offer small businesses the best banking experience possible. We don't want to spend a lot of time on operations, and trying to, and also reliability is incredibly important. If we can offload that burden and move faster, that's what we need to do. When we're talking about who's beating that drum, I would say our VP, Blake, really early on, seemed to see serverless as this amazing fit. I joined about three years ago today, so I guess this is my anniversary at the company. We were just deciding what to build. At the time there was a lot of architecture diagrams, and Blake hypothesized that serverless was a great fit.

We had a lot of versions of the world, some with Apache Kafka, and a bunch of microservices going through there. There's other versions with serverless in the mix, and some of the tooling around that, and this other hypothesis that maybe we want GraphQL gateway in the middle of there. It was one of those things that we wanted to test our hypothesis as we go. That ties into this innovation velocity that serverless allows for. It’s very cheap to put a new piece of infrastructure up in serverless. Just the other day we wanted to test Kinesis for an event streaming use case, and that was just a half an hour to set up that config, and you could put it live in production and test it out, which is completely awesome.

I think that innovation velocity was the hypothesis. We could just try things out really quickly. They don't cost much at all. You only pay for what you use for the most part. We were able to try that out, and as well as reliability. AWS really does a good job of making sure everything's available all the time. Something that maybe a young startup isn't ready to take on. When I joined the company, Blake proposed, “Okay, let's try out GraphQL as a gateway, as a concept. Build me a prototype." In that prototype, there was a really good opportunity to try serverless. They just ... Apollo server launched the serverless package, that was just super easy to deploy.

It was a complete no-brainer. We tried it out, we built the case. We just started with this GraphQL gateway running on serverless. AWS Lambda. It's funny because at first, it's like, we're just trying to sell them development. Nobody's going to be hitting our services. It was still a year out from when we were going into production. Once we went into prod, this Lambda's hot all the time, which is interesting. I think the cost case breaks down there because if you're running this thing, think forever, but it was this GraphQL server in front of our Python Django monolift, with this vision of event driven microservices, which has fit well for banking. If you just think about the banking world, everything is pretty much eventually consistent.

Just, that's the way the systems are designed. You send out a transaction, it doesn't settle for a while. We were always going to do event driven, but when you're starting out with a team of three developers, you're not going to build this whole microservices environment and everything. We started with that monolith with the GraphQL gateway in front, which scaled pretty nicely, because we were able to sort of, even today we have the same GraphQL gateway. We just changed the services backing it, which was really sweet. The adoption process was like, let's try it out. We tried it out with GraphQL first, and then as we were heading into launch, we had this monolith that we needed to manage. I mean, manually managing AWS resources, it's easier than back in the day when you're managing your own virtual machines and stuff, but it's still not great.

We didn't have a lot of time, and there was a lot of last-minute changes we needed to make. A big refactor to our scheduling transactions functions happened right before launch. That was an amazing serverless use case. And there's our second one, where we're like, “Okay, we need to get this live really quickly." We created this work performance pattern really quickly as a test with serverless, and it worked beautifully. We also had another use case come up, which was just a simple phone scheduling service. We just wrapped an API, and just exposed some endpoints, but it was just a lot easier to do with serverless. Just threw it off to two developers, figure out how you do it, and it was ready to be live. And then ...

Jeremy: I'm sorry to interrupt you, but I want to get to this point, because you're talking about standing up infrastructure, using infrastructure as code, or the tools you're using. How many developers were working on this thing?

Patrick: How many, I think at the time, maybe four developers on backend functionality before launch, when we were just starting out.

Jeremy: But you're building a banking platform here, so this is pretty sophisticated. I can imagine another business case for serverless is just the sense that we don't have to hire an operations team.

Patrick: Yeah, exactly. We were well through launching it. I think it would have been a couple of months where we were live, or where we hired our first dev ops engineer. Which is incredible. Our VP took a lot of that too, I'm sure he had his hands a little more dirty than he did like early on. But it was just amazing. We were able to manage all that infrastructure, and scale was never a concern. In the early stages, maybe it shouldn't be just yet, but it was just really, really easy.

Jeremy: Now you started with four, and I think, what are you now? Somewhere around 25 developers? Somewhere in that space now?

Patrick: About 25 developers now, we're growing really fast. We doubled this year during COVID, which is just crazy to think about, and somehow have been scaling somewhat smoothly at least, in terms of just being able to output as a dev team promote. We'll probably double again this year. This is maybe where I shamelessly plug that we're hiring, and we always are, and you could visit northone.com and just check out the careers page, or just hit me up for a warm intro. It's been crazy, and that’s one of the things that serverless has helped with us too. We haven't had this scaling bottleneck, which is an operations team. We don't need to hire X operations people for a certain number of developers.

Onboarding has been easier. There was one example of during a major project, we hired a developer. He was new to serverless, but just very experienced developer, and he had a production-ready serverless service ready in a month, which was just an insane ramp-up time. I haven't seen that very often. He didn't have to talk to any of our operation staff, and we'd already used serverless long enough that we had all of our presets and boilerplates ready, and permissions locked down, so it was just super easy. It's super empowering just for him to be able to just play around with the different services. Because we hit that point where we've invested enough that every developer when they opened a branch, that branch deploys its own stage, which has all of the services, AWS infrastructure deployed.

You might have a PR open that launches an instance of Kinesis, and five SQS queues, and 10 Lambdas, and a bunch of other things, and then tear down almost immediately, and the cost isn't something we really worry about. The innovation velocity there has been really, really good. Just being able to try things out. If you're thinking about something like Kinesis, where it's like a Kafka, that's my understanding, and if you think about the organizational buy-in you need for something like Kafka, because you need to support it, come up with opinions, and all this other stuff, you'll spend weeks trying it out, but for one of our developers, it's like this seems great.

We're streaming events, we want this to be real-time. Let's just try it out. This was for our analytics use case, and it's live in production now. It seems to be doing the thing, and we’re testing out that use case, and there isn't that roadblock. We could always switch off to a different design if you want. The experimentation piece there has been awesome. We’ve changed, during major projects we've changed the way we've thought about our resources a few times, and in the end it works out, and often it is about resiliency. It's just jamming queues into places we didn't think about in the first place, but that's been awesome.

Jeremy: I'm curious with that, though, with 25 developers ... Kinesis for the most part works pretty well, but you do have to watch those iterator ages, and make sure that they're not backing up, or that you're losing events. If they get flooded or whatever, and also sticking queues everywhere, sounds like a really good idea, and I'm a big fan of that, but it also, that means there's a lot of queues you have to manage, and watch, and set alarms and all that kind of stuff. Then you also talked about a pretty, what sounds like a pretty great CI/CD process to spin up new branches and things like that. There's a lot of dev ops-y ops work that is still there. How are you handling that now? Do you have dedicated ops people, or do you just have your developers looking after that piece of it?

Patrick: I would say we have a very spirited group of developers who are inspired. We do a lot of our code-sharing via internal packages. A few of our developers just figured out some of our patterns that we need, whether it's like CI, or how we structure our events stores, or how we do our Q subscriptions. We manage these internal packages. This won't scale well, by the way. This is just us being inspired and trying to reduce some of this burden. It is interesting, I’ve listened to this podcast and a few others, and this idea of infrastructure as code being part of every developer's toolbox, it’s starting to really resonate with our team.

In our migration, or our swift shift to full, I'd say doing serverless properly, we’ve learned to really think in it. Think in terms of infrastructure in our creating solutions. Not saying we're doing serverless the right way now, but we certainly did it the wrong way in the past, where we would spin up a bunch of API gateways that would talk to each other. A lot of REST calls going around the spider web of communication. Also, I'll call these monster Lambdas, that have a whole procedure list that they need to get through, and a lot of points of failure. When we were thinking about the way we're going to do Lambda now, we try to keep one Lambda doing one thing, and then there's pieces of infrastructure stitching that together. EventBridge between domain boundaries, SQS for commands where we can, instead of using API gateway. I think that transitions pretty well into our big break. I'm talking about this as our migration to serverless. I want to talk more about that.

Jeremy: Before we jump into that, I just want to ask this question about, because again, I call those fat, some people call them fat Lambdas, I call them Lambda lifts. I think there's Lambda lifts, then fat Lambdas, then your single-purpose functions. It's interesting, again, moving towards that direction, and I think it's super important that just admitting that you're like, we were definitely doing this wrong. Because I think so many companies find that adopting serverless is very much so an evolution, and it's a learning thing where the teams have to figure out what works for them, and in some cases discovering best practices on your own. I think that you've gone through that process, I think is great, so definitely kudos to you for that.

Before we get into that adoption and the migration or the evolution process that you went through to get to where you are now, one other business or technical case for serverless, especially with something as complex as banking, I think I still don't understand why I can't transfer personal money or money from my personal TD Bank account to my wife's local checking account, why that's so hard to do. But, it seems like there's a lot of steps. Steps that have to work. You can't get halfway through five steps in some transaction, and then be like, oops we can't go any further. You get to roll that back and things like that. I would imagine orchestration is a huge piece of this as well.

Patrick: Yeah, 100%. The banking lends itself really well to these workflows, I'll call them. If you're thinking about even just the start of any banking process, there's this whole application process where you put in all your personal information, you send off a request to your bank, and then now there's this whole waterfall of things that needs to happen. All kinds of checks and making sure people aren't on any fraud lists, or money laundering lists, or even just getting a second dive from our compliance department. There's a lot of steps there, and even just keeping our own systems in sync, with our off-provider and other places. We definitely lean on using step functions a lot. I think they work really, really well for our use case. Just the visual, being able to see this is where a customer is in their onboarding journey, is very, very powerful.

Being able to restart at any point of their, or even just giving our compliance team a view into that process, or even adding a pause portion. I think that's one of the biggest wins there, is that we could process somebody through any one of our pipelines, and we may need a human eye there at least for this point in time. That's one of the interesting things about the banking industry is. There are still manual processes behind the scenes, and there are, I find this term funny, but there are wire rooms in banks where there are people reviewing things and all that. There are a lot of workflows that just lend themselves well to step functions. That pausing capability and being able to return later with a response, so that allows you to build other internal applications for your compliance teams and other teams, or just behind the scenes calls back, and says, "Okay, resume this waterfall."

I think that was the visualization, especially in an events world when you're talking about like sagas, I guess, we're talking about distributed transactions here in a way, where there's a lot of things happening, and a common pattern now is the saga pattern. You probably don't want to be doing two-phase commits and all this other stuff, but when we're looking at sagas, it's the orchestration you could do or the choreography. Choreography gets very messy because there's a lot of simplistic behavior. I'm a service and I know what I need to do when these events come through, and I know which compensating events I need to dump, and all this other stuff. But now there's a very limited view.

If a developer is trying to gain context in a certain domain, and understand the chain of events, although you are decoupled, there's still this extra coupling now, having to understand what's going on in your system, and being able to share it with external stakeholders. Using step functions, that's the I guess the serverless way of doing orchestration. Just being able to share that view. We had this process where we needed to move a lot of accounts to, or a lot of user data to a different system. We were able to just use an orchestrator there as well, just to keep an eye on everything that's going on.

We might be paused in migrating, but let's say we’re moving over contacts, a transaction list, and one other thing, you could visualize which one of those are in the red, and which one we need to come in and fix, and also share that progress with external stakeholders. Also, it makes for fun launch parties I'd say. It's kind of funny because when developers do their job, you press a button, and everything launches, and there's not really anything to share or show.

Jeremy: There's no balloons or anything like that.

Patrick: Yeah. But it was kind of cool to look at these like, the customer is going through this branch of the logic. I know it's all green. Then I think one of the coolest things was just the retry ability as well. When somebody does fail, or when one of these workflows fails, you could see exactly which step, you can see the logs, and all that. I think one of the challenges we ran into there though, was because we are working in the banking space, we're dealing with sensitive data. Something I almost wish AWS solved out of the box, would be being able to obfuscate some of that data. Maybe you can't, I'm not sure, but we had to think of patterns for tokenization for instance.

Stripe does this a lot where certain parts of their platform, you just get it, you put in personal information, you get back a token, and you use that reference everywhere. We do tokenization, as well as we limit the amount of details flowing through steps in our orchestrators. We'll use an event store with identifiers flowing through, and we'll be doing reads back to that event store in between steps, to do what we need to do. You lose some of that debug-ability, you can't see exactly what information is flowing through, but we need to keep user data safe.

Jeremy: Because it's the use case for it. I think that you mentioned a good point about orchestration versus choreography, and I'm a big fan of choreography when it makes sense. But I think one of the hardest lessons you learn when you start building distributed systems is knowing when to use choreography, and knowing when to use orchestration. Certainly in banking, orchestration is super important. Again, with those saga patterns built-in, that's the kind of thing where you can get to a point in the process and you don't even need to do automated rollbacks. You can get to a failure state, and then from there, that can be a pause, and then you can essentially kick off the unwinding of those things and do some of that.

I love that idea that the token pattern and using just rehydrating certain steps where you need to. I think that makes a ton of sense. All right. Let's move on to the adoption and the migration process, because I know this is something that really excites you and it should because it is cool. I always know, as you're building out applications and you start to add more capabilities and more functionality and start really embracing serverless as a methodology, then it can get really exciting. Let's take a step back. You had a champion in your organization that was beating the drum like, "Let's try this. This is going to make a lot of sense." You build an Apollo Lambda or a Lambda running Apollo server on it, and you are using that as a strangler pattern, routing all your stuff through now to your backend. What happens next?

Patrick: I would say when we needed to build new features, developers just gravitated towards using serverless, it was just easier. We were using TypeScript instead of Python, which we just tend to like as an organization, so it's just easier to hop into TypeScript land, but I think it was just easier to get something live. Now we had all these Lambdas popping up, and doing their job, but I think the problem that happened was we weren't using them properly. Also, there was a lot of difference between each of our serverless setups. We would learn each time and we'd be like, okay, we'll use this parser function here to simplify some of it, because it is very bare-bones if you're just pulling the Serverless Framework, and it took a little ...

Every service looked very different, I would say. Also, we never really took the time to sit back and say, “Okay, how do we think about this? How do we use what serverless gives us to enable us, instead of it just being an easy thing to spin up?" I think that's where it started. It was just easy to start. But we didn't embrace it fully. I remember having a conversation at some point with our VP being like, “Hey, how about we just put Express into one of our Lambdas, and we create this," now I know it's a Lambda lift. I was like, it was just easier. Everybody knows how to use Express, why don't we just do this? Why are we writing our own parsers for all these things? We have 10 versions of a make response helper function that was copy-pasted between repos, and we didn't really have a good pattern for sharing that code yet in private packages.

We realized that we liked serverless, but we realized we needed to do it better. We started with having a serverless chapter reading between some of our team members, and we made some moves there. We created a shared boilerplate at some point, so it reduced some of the differences you'd see between some of the repositories, but we needed a step-change difference in our thinking, when I look back, and we got lucky that opportunity came up. At this point, we probably had another six Lambda services, maybe more actually. I want to say around, we'd probably have around 15 services at this point, without a governing body around patterns.

At this time, we had this interesting opportunity where we found out we're going to be re-platforming. A big announcement we just made last month was that we moved on to a new bank partner called Bancorp. The bank partner that supports Chime, and they're like, I'll call them an engine boost. We put in a much larger, more efficient engine for our small businesses. If you just look at the capabilities they provide, they're just absolutely amazing. It's what we need to build forward. Their events API is amazing as well as just their base banking capabilities, the unit economics they can offer, the times on there, things were just better. We found out we're doing an engine swap. The people on the business side on our company trusted our technical team to do what we needed to do.

Obviously, we need to put together a case, but they trusted us to choose our technology, which was awesome. I think we just had a really good track record of delivering, so we had free reign to decide what do we do. But the timeline was tight, so what we decided to do, and this was COVID times too, was a few of our developers got COVID tested, and we rented a house and we did a bubble situation. How in the NHL or MBA you have a bubble. We had a dev bubble.

Jeremy: The all-star team.

Patrick: The all-star team, yeah. We decided let's sit down, let's figure out what patterns are going to take us forward. How do we make the step-change at the same time as step-change in our technology stack, at the same time as we're swapping out this bank, this engine essentially for the business. In this house, we watched almost every YouTube video you can imagine on event driven and serverless, and I think leading up. I think just knowing that we were going to be doing this, I think all of us independently started prototyping, and watching videos, and reading a lot of your content, and Alex DeBrie and Yan Cui. We all had a lot of ideas already going in.

When we all got to this house, we started off with this exercise, an event storming exercise, just popular in the domain-driven design community, where we just threw down our entire business on a wall with sticky notes, and it would have been better to have every business stakeholder there, but luckily we had two people from our product team there as representatives. That's how invested we were in building this outright, that we have products sitting in the room with us to figure it out.

We slapped down our entire business on a wall, this took days, and then drew circles around it and iterated on that for a while. Then started looking at what the technology looks like. What are our domain boundaries, and what prototypes do we need to make? For a few weeks there, we were just prototyping. We built out what I'd called baby's first balance. That was the running joke where, how do we get an account opened with a balance, with the transactions minimally, with some new patterns. We really embraced some of this domain-driven-design thinking, as well as just event driven thinking. When we were rethinking architecture, three concepts became very important for us, not entirely new, but important. Item potency was a big one, dealing with distributed transactions was another one of those, as well as the eventual consistency. The eventual consistency portion is kind of funny because we were already doing it a lot.

Our transactions wouldn't always settle very quickly. We didn't know about it, but now our whole system becomes eventually consistent typically if you now divide all of your architecture across domains, and decouple everything. We created some early prototypes, we created our own version of an event store, which is, I would just say an opinionated scheme around DynamoDB, where we keep track of revisions, payload, timestamp, all the things you'd want to be able to do event sourcing. That's another thing we decided on. Event sourcing seemed like the right approach for state, for a lot of our use cases. Banking, if you just think about a banking ledger, it is events or an accounting ledger. You're just adding up rows, add, subtract, add, subtract.

We created a lot of prototypes for these things. Our events store pattern became basically just a DynamoDB with opinions around the schema, as well as a package of a shared code package with a simple dispatch function. One dispatch function that really looks at enforcing optimistic concurrency, and one that's a little bit more relaxed. Then we also had some reducer functions built into there. That was one of the packages that we created, as well as another prototype around that was how do we create the actual subscriptions to this event store? We landed on SNS to SQS fan-out, and it seems like fan-out first is the serverless way of doing a lot of things. We learned that along the way, and it makes sense. It was one of those things we read from a lot of these blogs and YouTube videos, and it really made sense in production, when all the data is streaming from one place, and then now you just add subscribers all over the place. Just new queues. Fan-out first, highly recommend. We just landed on there by following best practices.

Jeremy: Great. You mentioned a bunch of different things in there, which is awesome, but so you get together in this house, you come up with all the events, you do this event storming session, which is always a great exercise. You get a pretty good visualization of how the business is going to run from an event standpoint. Then you start building out this event driven architecture, and you mentioned some packages that you built, we talked about step functions and the orchestration piece of this. Just give me a quick overview of the actual system itself. You said it's backed by DynamoDB, but then you have a bunch of packages that run in between there, and then there's a whole bunch of queues, and then you're using some custom packages. I think I already said that but you're using ... are you using EventBridge in there? What's some of the architecture behind all that?

Patrick: Really, really good question. Once we created these domain boundaries, we needed to figure out how do we communicate between domains and within domains. We landed on really differentiating milestone events and domain events. I guess milestone events in other terms might be called integration events, but this idea that these are key business milestones. An account was open, an application was approved or rejected, things that every domain may need to know about. Then within our domains, or domain boundaries, we had these domain events, which might reduce to a milestone event, and we can maintain those contracts in the future and change those up. We needed to think about how do we message all these things across? How do we communicate? We landed on EventBridge for our milestone events. We have one event bus that we talked to all of our, between domain boundaries basically.

EventBridge there, and then each of our services now subscribed to that EventBridge, and maintain their own events store. That's backed by DynamoDB. Each of our services have their own data store. It's usually an event stream or a projection database, but it's almost all Dynamo, which is interesting because our old platform used Postgres, and we did have relational data. It was interesting. I was really scared at first, how are we going to maintain relations and things? It became a non-issue. I don't even know why now that I think about it. Just like every service maintains its nice projection through events, and builds its own view of the world, which brings its own problems. We have DynamoDB in there, and then SNS to SQS fan-out. Then when we're talking about packages ...

Jeremy: That's Office Streams?

Patrick: Exactly, yeah. We're Dynamo streams to SNS, to SQS. Then we use shared code packages to make those subscriptions very easy. If you're looking at doing that SNS to SQS fan-out, or just creating SQS queues, there is a lot of cloud formation boilerplate that we were creating, and we needed to move really quick on this project. We got pretty opinionated quick, and we created our own subscription function that just generates all this cloud formation with naming conventions, which was nice. I think the opinions were good because early on we weren't opinionated enough, I would say. When you look in your AWS dashboard, the read for these aren't prefixed correctly, and there's all this garbage. You're able to have consistent naming throughout, make it really easy to subscribe to an event.

We would publish packages to help with certain things. Our events store package was one of those. We also created a Lambda handlers package, which leverages, there's like a Lambda middlewares compose package out there, which is quite nice, and we basically, all the common functionality we're doing a lot of, like parsing a body from S3, or SQS or API gateway. That's just the middleware that we now publish. Validation in and out. We highly recommend the library Zod, we really embrace the TypeScript first object validation. Really, really cool package. We created all these middlewares now. Then subscription packages. We have a lot of shared code in this internal NPM repository that we install across.

I think one challenge we had there was, eventually you extracted away too much from the cloud formation, and it's hard for new developers to ... It's easy for them to create events subscriptions, it's hard for them to evolve our serverless thinking because they're so far removed from it. I still think it was the right call in the end. I think this is the next step of the journey, is figuring out how do we share code effectively while not hiding away too much of serverless, especially because it's changing so fast.

Jeremy: It's also interesting though that you take that approach to hide some of that complexity, and bake in some of that boilerplate that, someone's mostly didn't have to write themselves anyways. Like you said, they're copying and pasting between services, is not the best way to do it. I tried the whole shared packages thing one time, and it kind of worked. It's just like when you make a small change to that package and you have 14 services, that then you have to update to get the newest version. Sometimes that's a little frustrating. Lambda layers haven't been a huge help with some of that stuff either. But anyways, it's interesting, because again you've mentioned this a number of times about using queues.

You did mention resiliency in there, but I want to touch on that point a little bit because that's one of those things too, where I would assume in a banking platform, you do not want to lose events. You don't want to lose things. and so if something breaks, or something gets throttled or whatever, having to go and retry those events, having the alerts in place to know that a queue is backed up or whatever. Then just, I'm thinking ordering issues and things like that. What kinds of issues did you face, and tell me a little bit more about what you've done for reliability?

Patrick: Totally. Queues are definitely ... like SQS is a workhorse for our company right now. We use a lot of it. Dropping messages is one of the scariest things, so you're dead-on there. When we were moving to event driven, that was what scared me the most. What if we drop an event? A good example of that is if you're using EventBridge and you're subscribing Lambdas to it, I was under the impression early on that EventBridge retries forever. But I'm pretty sure it'll retry until it invokes twice. I think that's what we landed on.

Jeremy: Interesting.

Patrick: I think so, and don't quote me on this. That was an example of where drop message could be a problem. We put a queue in front of there, an SQS queue as the subscription there. That way, if there's any failure to deliver there, it's just going to retry all the time for a number of days. At that point we got to think about DLQs, and that's something we're still thinking about. But yeah, I think the reason we've been using queues everywhere is that now queues are in charge of all your retry abilities. Now that we've decomposed these Lambdas into one Lambda lift, into five Lambdas with queues in between, if anything fails in there, it just pops back into the queue, and it'll retry indefinitely. You can drop messages after a few days, and that's something we learned luckily in the prototyping stage, where there are a few places where we use dead letter queues. But one of the issues there as well was ordering. Ordering didn't play too well with ...

Jeremy: Not with DLQs. No, it does not, no.

Patrick: I think that's one lesson I'd want to share, is that only use ordering when you absolutely need it. We found ways to design some of our architecture where we didn't need ordering. There's places we were using FIFO SQS, which was something that just launched when we were building this thing. When we were thinking about messaging, we're like, "Oh, well we can't use SQS because they don't respect ordering, or it doesn't respect ordering." Then bam, the next day we see this blog article. We got really hyped on that and used FIFO everywhere, and then realized it's unnecessary in most use cases. So when we were going live, we actually changed those FIFO queues into just regular SQS queues in as many places as we can. Then so, in that use case, you could really easily attach a dead letter queue and you don't have to worry about anything, but with FIFO things get really, really gnarly.

Ordering is an interesting one. Another place we got burned I think on dead-letter queues, or a tough thing to do with dead letter queues is when you're using our state machines, we needed to limit the concurrency of our state machines is another wishlist item in AWS. I wish there was just at the top of the file, a limit concurrent executions of your state machine. Maybe it exists. Maybe we just didn't learn to use it properly, but we needed to. There's a few patterns out there. I've seen the [INAUDIBLE] pattern where you can use the actual state machine flow to look back at how many concurrent executions you have, and pause. We landed on setting reserved concurrency in a number of Lambdas, and throwing errors. If we've hit the max concurrency and it'll pause that Lambda, but the problem with DLQs there was, these are all errors. They're coming back as errors.

We're like, we're fine with them. This is a throttle error. That's fine. But it's hard to distinguish that from a poison message in your queue, so when do you dump those into DLQ? If it's just a throttling thing, I don't think it matters to us. That was another challenge we had. We're still figuring out dead letter queues and alerting. I think for now we just relied on CloudWatch alarms a lot for our alerting, and there's a lot you could do. Even just in the state machines, you can get pretty granular there. I know once certain things fail, and announced to your Slack channel. We use that Slack integration, it's pretty easy. You just go on a Slack channel, there's an email in there, you plop it into the console in AWS, and you have your very early alerting mechanism there.

Jeremy: The thing with Elasticsearch ... not Elasticsearch, I'm sorry. I'm totally off-topic here. The thing with EventBridge and Lambda, these are one of those things that, again, they’re nuances, but event bridge, as long as it can deliver to the Lambda service, then the Lambda service kicks off and queues it automatically. Then that will retry at a certain number of times. I think you can control that now. But then eventually if that retries multiple times and eventually fails, then that kicks it over to the DLQ or whatever. There's all different ways that it works like that, but that's why I always liked the idea of putting a queue in between there as well, because I felt you just had a little bit more control over exactly what happens.

As long as it gets to the queue, then you know you haven't lost the message, or you hope you haven't lost a message. That's super interesting. Let's move on a little bit about the adoption issues. You mentioned a few of these things, obviously issues with concurrency and ordering, and some of that other stuff. What about some of the other challenges you had? You mentioned this idea of writing all these packages, and it pulls devs away from the CloudFormation a little bit. I do like that in that it, I think, accelerates a lot of things, but what are some of the other maybe challenges that you've been having just getting this thing up and running?

Patrick: I would say IAM is an interesting one. Because we are in the banking space, we want to be very careful about what access do you give to what machines or developers, I think machines are important too. There've been cases where ... so we do have a separate developer set up with their own permissions, in development's really easy to spin up all your services within reason. But now when we're going into production, there's times where our CI doesn't have the permissions to delete a queue or create a queue, or certain things, and there's a lot of tweaking you have to do there, and you got to do a lot of thinking about your IAM policies as an organization, especially because now every developer's touching infrastructure.

That becomes this shared operational overhead that serverless did introduce. We're still figuring that out. Right now we’re functioning on least privilege, so it's better to just not be able to deploy than deploy something you shouldn't or read the logs that you shouldn't, and that's where we're starting. But that's something that, it will be a challenge for a little while I think. There's all kinds of interesting things out there. I think temporary IAM permissions is a really cool one. There are times we're in production and we need to view certain logs, or be able to access a certain queue, and there's tooling out there where you can, or at least so I've heard, you can give temporary permissions. You have this queue permission for 30 minutes, and it expires and it's audited, and I think there's some CloudTrail tie-in you could do there. I'm speaking about my wishlist for our next evolution here. I hope my team is listening ...

Jeremy: Your team's listening to you.

Patrick: ... will be inspired as well.

Jeremy: What about ... because this is something too that I always found to be a challenge, especially when you start having multiple services, and you've talked about these domain events, but then milestone events. You've got different services that need to communicate across services, or across domains, and realize certain things like that. Service discovery in and of itself, and which queue are we mapping to, or which service am I talking to, and which version of the service am I talking to? Things like that. How have you been dealing with that stuff?

Patrick: Not well, I would say. Very, very ad hoc. I think like right now, at least we have tight communication between the teams, so we roughly know which service we need to talk to, and we output our URLs in the cloud formation output, so at least you could reference the URLs across services, a little easier. Really, a GraphQL is one of the only service that really talks to a lot of our API gateways. At least there's less of that, knowing which endpoint to hit. Most of our services will read into EventBridge, and then within services, a lot of that's abstracted away, like the queue subscription's a little easier. Service discovery is a bit of a nightmare.

Once our services grow, it'll be, I don't know. It'll be a huge challenge to understand. Even which services are using older versions of Node, for instance. I saw that AWS is now deprecating version 10 and we'll have to take a look internally, are we using version 10 anywhere, and how do we make sure that's fine, or even things like just knowing which services now have vulnerabilities in their NPM packages because we're using Node. That's another thing. I don't even know if that falls in service discovery, but it's an overhead of ...

Jeremy: It's a service management too. It's a lot there. That actually made me, it brings me to this idea of observability too. You mentioned doing some CloudWatch alerts and some of that stuff, but what about using some observability tool or tracing like x-ray, and things like that? Have you been implementing any of that, and if you have, have you had any success and or problems with it?

Patrick: I wish we had a better view of some of the observability tools. I think we were just building so quickly that we never really invested the time into trying them out. We did use X-Ray, so we rolled our own tooling internally to at least do what we know. X-Ray was one of those, but the problem with X-Ray is, we do subscribe all of our services, but X-Ray isn't implemented everywhere internally in AWS, so we lose our trail somewhere in that Dynamo stream to SNS, or SQS. It's not a full trace. Also, just digesting that huge graph of information is just very difficult. I don't use it often, I think it's a really cool graphic to show, “Hey, look, how many services are running, and it's going so fast."

It's a really cool thing to look at, but it hasn't been very useful. I think our most useful tool for debugging and observability has been just our logging. We created a JSON logger package, so we get up JSON logs and we can actually filter off of different properties, and we ship those to Elasticsearch. Now you can have a view of all of the functions within a given domain at any point in time. You could really see the story. Because I think early on when we were opening up CloudWatch and you'd have like 10 tabs, and you're trying to understand this flow of information, it was very difficult.

We also implemented our own trace ID pattern, and I think we just followed a Lumigo article where we introduced some properties, and in each of our Lambdas at a higher level, and one of our middlewares, and we were able to trace through. It's not ideal. Observability is something that we'll probably have to work on next. It’s been tolerable for now, but I can't see the scaling that long.

Jeremy: That's the other thing too, is even the shared package issue. It's like when you have an observability tool, they'll just install a layer or something, where you don't necessarily have to worry about updating your own tool. I always find if you are embracing serverless and you want to get rid of all that undifferentiated heavy lifting, observability tools, there's a lot of really good ones out there that are doing some great stuff, and they're specializing in it. It might be worth letting someone else handle that for you than trying to do it yourself internally.

Patrick: Yeah, 100%. Do you have any that you've used that are particularly good? I know you work with serverless so-

Jeremy: I played around with all of them, because I love this stuff, so it's always fun, but I mean, obviously Lumigo and Epsagon, and Thundra, and New Relic. They’re all great. They all do things slightly differently, but they all follow a similar implementation pattern so that it’s very easy to install them. We can talk more about some recommendations. I think it's just one of those things where in a modern application not having that insight is really hard. It can be really hard to debug stuff. If you look at some of the tools that AWS offers, I think they’re there, it's just, they are maybe a little harder to implement, and not quite as refined and targeted as some of the observability tools. But still, you got to get there. Again, that's why I keep saying it's an evolution, it's a process. Maybe one time you get burned, and you're like, we really needed to have observability, then that's when it becomes more of a priority when you're moving fast like you are.

Patrick: Yeah, 100%. I think there's got to be a priority earlier than later. I think I'll do some reading now that you've dropped some of these options. I have seen them floating around, but it's one of those things that when it's too late, it's too late.

Jeremy: It's never too late to add observability though, so it should. Actually, a lot of them now, again, it makes it really, really easy. So I'm not trying to pitch any particular company, but take a look at some of them, because they are really great. Just one other challenge that I also find a lot of people run into, especially with serverless because there's all these artificial account limits in place. Even the number of queues you can create, and the number of concurrent Lambda functions in a particular region, and stuff like that. Have you run into any of those account limit issues?

Patrick: Yeah. I could give you the easiest way to run into an account on that issue, and that is replay your entire EventBridge archive to every subscriber, and you will find a bottleneck somewhere. That's something ...

Jeremy: Somewhere it'll fall over? Nice.

Patrick: 100%. It's a good way to do some quick check and development to see where you might need to buffer something, but we have run into that. I think the solution there, and a lot of places was just really playing with concurrency where we needed to, and being thoughtful about where is their main concurrency in places that we absolutely needed to stay functioning. I think the challenge there is that eats into your total account concurrency, which was an interesting learning there. Definitely playing around there, and just being thoughtful about where you are replaying. A couple of things. We use replays a lot. Because we are using these milestone events between service boundaries, now when you launch a new service, you want to replay that whole history all the way through.

We've done a lot of replaying, and that was one of the really cool things about EventBridge. It just was so easy. You just set up an archive, and it'll record everything coming through, and then you just press a button in the console, and it'll replay all of them. That was really awesome. But just being very mindful of where you're replaying to. If you replay to all of your subscriptions, you'll hit Lambda concurrency limits real quick. Even just like another case, early on we needed to replace ... we have our own domain events store. We want to replace some of those events, and those are coming off the Dynamo stream, so we were using dynamo to kick those to a stream, to SNS, and fan-out to all of our SQS queues. But there would only be one or two queues you actually needed to subtract to those events, so we created an internal utility just to dump those events directly into the SQS queue we needed. I think it's just about not being wasteful with your resources, because they are cheap. Sure.

Jeremy: But if you use them, they start to cost money.

Patrick: Yeah. They start to cost some money as well as they could lock down, they can lock you out of other functionality. If you hit your Lambda limits, now our API gateway is tapped.

Jeremy: That's a good point.

Patrick: You could take down your whole system if you just aren't mindful about those limits, and now you could call up AWS in a panic and be like, “Hey, can you update our limits?" Luckily we haven't had to do that yet, but it's definitely something in your back pocket if you need it, if you can make the case to AWS, that maybe you do need bigger limits than the default. I think just not being wasteful, being mindful of where you're replaying. I think another interesting thing there is dealing with partners too. It's really easy to scale in the Lambda world, but not every partner could handle that volume really quickly. If you're not buffering any event coming through EventBridge to your new service that hits a partner every time, you're going to hit their API rate limit really quickly, because they're just going to just go right through it.

You might be doing thousands of API calls when you're instantiating a new service. That's one of those interesting things that we have to deal with, and particularly in our orchestrators, because they are talking to different partners, that's why we need to really make sure we could limit the concurrent executions of the state machines themselves. In a way, some of our architecture is too fast to scale.

Jeremy: It's too good.

Patrick: You still have to consider downstream. That, and even just, if you are using relational databases or anything else in your system, now you have to worry about connection limits and ...

Jeremy: I have a whole talk I gave on that.

Patrick: ... spikes in traffic.

Jeremy: Yes, absolutely.

Patrick: Really cool.

Jeremy: I know all about it. Any final advice for companies like you that are trying to bite off a piece of the serverless apple, I guess, That's really bad. Anyways, any advice for people looking to get into this?

Patrick: Yeah, totally. I would say start small. I think we were wise to just try it out. It might not land with your development team. If you don't really buy in, it's one of those things that could just end up unnecessarily messy, so start small, see if you like it in-shop, and then reevaluate, once you hit a certain point. That, and I would say shared boilerplate packages sooner than later. I know shared code is a problem, but it is nice to have an un-opinionated starter pack, that you're at least not doing anything really crazy. Even just things like having opinions around logging. In our industry, it's really important that you're not logging sensitive details.

For us doing things like wrapping our HTTP clients to make sure we're not logging sensitive details, or having short Lambda packages that make sure out-of-the-box you're opinionated about not doing something terribly awful. I would say those two things. Start small and a boiler package, and maybe the third thing is just pay attention to the code smell of a growing Lambda. If you are doing three API calls in one Lambda, chances are you could probably break that up, and think about it in a more resilient way. If any one of those pieces fail, now you could have retry ability in each one of those. Those are the three things I would say. I could probably talk forever about the rest of our journey.

Jeremy: I think that was great advice, and I love hearing about how companies are going through this process, what that process looks like, and I hope, I hope, I hope that companies listen to this and can skip a lot of these mistakes. I don't want to call them all mistakes, and I think it's just evolution. The stuff that you've done, we've all made them, we've all gone through that process, and the more we can solidify these practices and stuff like that, I think that more companies will benefit from hearing stories like these. Thank you very much for sharing that. Again, thank you so much for spending the time to do this and sharing all of this knowledge, and this journey that you've been on, and are continuing to be on. It would great to continue to get updates from you. If people want to contact you, I know you're not on Twitter, but what's the best way to reach out to you?

Patrick: I almost wish I had a Twitter. It's the developer thing to have, so maybe in the future. Just on LinkedIn would be great. LinkedIn would be great, as well as if anybody's interested in working with our team, and just figuring out how to take serverless to the next level, just hit me up on LinkedIn or look at our careers page at northone.com, and I could give you a warm intro.

Jeremy: That's great. Just your last name is spelled S-T-R-Z-E-L-E-C. How do you say that again? Say it in Polish, because I know I said it wrong in the beginning.

Patrick: I guess for most people it would just be Strzelec, but if there are any Slavs in the audience, it's "Strzelec." Very intense four consonants last name.

Jeremy: That is a lot of consonants. Anyways again, Patrick, thanks again. This was great.

Patrick: Yeah, thank you so much, Jeremy. This has been awesome.

View Details

About Patrick McFadin

Patrick McFadin is the VP of Developer Relations at DataStax, where he leads a team devoted to making users of Apache Cassandra successful. He has also worked as Chief Evangelist for Apache Cassandra and consultant for DataStax, where he helped build some of the largest and exciting deployments in production. Previous to DataStax, he was Chief Architect at Hobsons and an Oracle DBA/Developer for over 15 years.

Twitter: @PatrickMcFadin
LinkedIn: Patrick McFadin
DataStax website: datastax.com
K8ssandra: k8ssandra.io
Stargate: stargate.io
DataStax Astra: Cassandra-as-a-Service

Watch this episode on YouTube: https://youtu.be/-BcIL3VlrjE

This episode sponsored by CBT Nuggets and Fauna.

Transcript
Jeremy: Hi everyone, I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Patrick McFadin. Hey Patrick, thanks for joining me.

Patrick: Hi Jeremy. How are you doing today?

Jeremy: I am doing really well. So you are the VP of Developer Relations at DataStax, so I'd love it if you could tell the listeners a little bit about yourself and what DataStax is all about.

Patrick: Sure. Well, I mean mostly I'm just a nerd with a cool job. I get to talk about technology a lot and work with technology. So DataStax, we're a company that was founded around Apache Cassandra, just supporting and making it awesome. And that's really where I came to the company. I've been working with Apache Cassandra for about 10 years now. I've been a part of the project as a contributor.

But yeah, I mean mostly data infrastructure has been my life for most of my career. I did this in the dotcom era, back when it was really crazy when we had dozens of users. And when that washed out, I'm like, oh, then real scale started and during that period of time I worked a lot in just trying to scale infrastructure. It seems like that's been what I've been doing for like 30 years it seems like, 20 years, 20 years, I'm not that old. Yeah. But yeah, right now, I spend a lot of my time just working with developers on what's next in Kubernetes and I'm part of CNCF now, so yeah. I just can't to seem to stay in one place.

Jeremy: Well, so I'm super interested in the work that DataStax is doing because I have had the pleasure/misfortune of managing a Cassandra ring for a start-up that I was at. And it was a very painful process, but once it was set up and it was running, it wasn't too, too bad. I mean, we always had some issues here and there, but this idea of taking a really good database, because Cassandra's great, it's an excellent data store, but managing it is a nightmare and finding people who can manage it is sort of a nightmare, and all that kind of stuff. And so this idea of taking these services and DataStax isn't the only one to do this, but to take these open-source services and turn them into these hosted solutions is pretty fantastic. So can you tell me a little bit more, though? What this shift is about? This moving away from hosting your own databases to using databases as a service?

Patrick: Yeah. Well, you touched on something important. You want to take that power, I mean Cassandra was a database that was built in the scale world. It was built to solve a problem, but it was also built by engineers who really loved distributed computing, like myself, and it's funny you say like, "Oh, once I got it running, it was great," well, that's kind of the experience with most distributed databases, is it's hard to reason around having, "Oh, I have 100 mouths to feed now. And if one of them goes nuts, then I have to figure it out."

But it's the power, that power, it's like stealing fire from the gods, right? It's like, "Oh, we could take the technology that Netflix and Apple and Facebook use and use it in our own stuff." But you got to pay the price, the gods demand their payment. And that's something that we've been really trying to tackle at DataStax for a couple of years now, actually three, which is how ... Because the era of running your own database is coming to an end. You should not run your own database. And my philosophy as a technologist is that proper, really important technology like your data layer should just fade into the background and it's just something you use, it's not something you have to reason through very much.

There's lots of technology that's like that today. How many times have you ... When was the last time you managed your own memory in your code?

Jeremy: Right. Right. Good point. I know.

Patrick: Thank god, huh?

Jeremy: Exactly.

Patrick: Whew.

Jeremy: But I think that you make a really good point, because you do have these larger companies like Facebook or whatever that are using these technologies and you mentioned data layers, which I don't think I've worked for a single company, I don't think I actually ... I founded a start-up one time and we built a data layer as well, because it's like, the complexity of understanding the transaction models and the routing, especially if you're doing things like sharding and all kinds of crazy stuff like that, hiding that complexity from your developers so that they can just say, "I need to get this piece of information," or, "I need to set this piece of information," is really powerful.

But then you get stuck with these data layers that are bespoke and they're generally fragile and things like that, so how is that you can take data as a service and maybe get rid of some of that, I don't know, some of that liability I guess?

Patrick: Yeah. It's funny because you were talking about sharding and things like that. These are things that we force on developers to reason through, and it's just cognitive load. I have an app to get out, and I have some business desire to get this application online, the last thing I need to worry about is my sharding algorithm. Jeremy, friends don't let friends shard.

Jeremy: Right. That's right. That's a good point.

Patrick: But yeah, I mean I think we actually have all the parts that we need and it's just about, this is closer than you think. Look at where we've already started going, and that is with APIs, using REST. Now GraphQL, which I think is deserving its hotness, is starting to bring together some things that are really important for this kind of world we want to live in. GraphQL is uni-fettering data and collecting and actual queries, it's a QL, and why they call it Graph, I have no idea. But it gives you this ability to have this more abstract layer.

I think GraphQL will, here's a prediction is that it's going to be like the SQL of working with data services on the internet and for cloud-native applications. And so what does that mean? Well, that means I just have to know, well, I need some data and I don't really care what's underneath it. I don't care if I have this field indexed or anything like that. And that's pretty exciting to me because then we're writing apps at that point.

Jeremy: Right. Yeah. And actually, that's one of the things I really like about GraphQL too is just this idea that it's almost like a universal data access layer in a sense because it does, you still have to know it, you have to know what you're requesting if you're an end developer, but it makes it easier to request the things that you need and have those mutations set and have some of those other things standardized across the company, but in a common format because isn't that another problem? Where it's like, I'm working with company A and I move to company B maybe and now company B is using a different technology and a different bespoke data layer and some of these other things.

So, I think data as a service for one, maybe with GraphQL in front of it is a great way to have this alignment across companies, or I guess, just makes it easier for developers to switch and start developing right away when they move into a new company.

Patrick: Yeah, and this is a concept I've been trying to push pretty hard and it's driven by some conversations I've had with some friends that they're engineering leaders and they have this common desire. We want to have a zero day dev, which is the first day that someone starts, they should be producing production code. And I don't think that's crazy talk, we can do this, but there's a lot of things that are in front of it. And the database is one of them. I think that's one of the first things you do when you show up at company X is like, "Okay, what database are you using? What flavor of SQL or GRPC or CQL, Cassandra query language? What's the data model? Quick, where's that big diagram on the wall with my ERD? I got to go look at that for a while."

Jeremy: How poorly did you structure your Git repositories? Yeah.

Patrick: Yeah, exactly. It's like all these things. And no, I would love to see a world where the most troublesome part of your first day is figuring out where the coffee and the bathroom are, and then the rest of it is just total, "Hey, I can do this. This is what I get paid to do."

Jeremy: Right. Yeah. So that idea of zero day developer, I love that idea and I know other companies are trying to do that, but what enables that? Is it getting the idea of having to understand something bespoke? Is it getting that off of the table? Or not having to deal with the low-level database aspect of things? I mean because APIs, I had this conversation with Rob Sutter, actually, a couple weeks ago. And we were talking about the API economy and how everything is moving towards APIs. And even data, it was around data as well.

So, is that the interface, you think, of the future that just says, "Look, trying to interface directly with a database or trying to work with some other layer of abstraction just doesn't make sense, let's just go straight from code right to the data, with a very simple API interface?"

Patrick: Yeah, I think so. And it's this idea of data services because if you think of if you're doing React, or something like a front-end code, I don't want to have a driver. Drivers are a total impediment. It's like, driver hell can be difficult at large organizations, getting the matching right. Oh, we're using this database so you have to use this driver. And if you don't, you are now rejected at the gate. So it's using HTTP protocols, but it's also things like when you're using React or Angular, View, whatever you're using on the front-end, you have direct access.

But most times what you're needing is just a collection or an object. And so just do a get, "I need this thing right now. I'm doing a pick list. I need your collection." I don't need a complicated setup and spend the first three days figuring out which driver I'm using and make sure my Gradle file is just perfect. Yeah. So, I think that's it.

Jeremy: Yeah. No, I'd be curious how you feel about ORMs, or O-R-Ms, certainly for relational databases, I know a lot of people love them. I can't stand them. I think it adds a layer of abstraction and just more complexity where I just want access to the database. I want to write the query myself, and as soon as you start adding in all this extra stuff on top of it to try to make it easier, I don't know, it just seems to mess it up for me.

Patrick: All right. So yeah, I think we have an accord. I am really not a fan of ORMs at all. And I mean this goes back to Hibernate. Everyone's like, "Oh, Hibernate's going to be the end of databases." No, it's not. Oh yeah, it was the end of the database at the other side because it would create these ridiculous queries. It's like, why is every query a full table scan?

Jeremy: Exactly.

Patrick: Because that's the way Hibernate wanted it. Yeah. I actually banned Hibernate at one company I was working at. I was Chief Architect there and I just said, "Don't ever put Hibernate in our production." Because I had more meetings about what it was doing wrong than what it was doing right.

Jeremy: Right. Right. Yeah. No, that's sounds, yeah.

Patrick: Is that a long answer? Like, no.

Jeremy: No, I've had the same experience where certain ORMs you're just like, no. Certain things, you can't do this because it's going to one, I think it locks you in in a sense, I mean there's all kind of lock-in in the cloud, and if you're using a data service or an API or you're using something native in AWS, or IBM Cloud, you're still going to be locked in in some way, but I do feel like whenever you start going down that path of building custom things, or forcing developers to get really low level, that just builds up all kinds of tech debt, right? That you eventually are going to have to work down.

Patrick: Well, it's organizational inertia. When you start getting into this, when you start using annotations in Hibernate where you're just cutting through all the layers and now you're way down in the weeds, try to move that. There's a couple of companies that I've worked with now that are looking at the true reality of portability in their data stores. Like, "Oh, we want to move from one to a different, from a key value to a document without developers knowing." Well, how do you get to that point?

Jeremy: Right. Yeah.

Patrick: And it's just, that's not giving access to those things, first of all, but this is that tech debt that's going to get in your way. We're really good, technologists, we're really good at just wracking up the charges on our tech debt credit card, especially whenever we're trying to get things out the door quickly. And I think that's actually one of the problems that we all face. I mean, I don't think I've ever talked to a developer who was ahead of schedule and didn't have somebody breathing down their neck.

Jeremy: Very true.

Patrick: You take shortcuts. You're like, "We've got to shift this code this week. Skip the annotations and go straight into the database and get the data you need." Or something. You start making trade-offs real fast.

Jeremy: What can we hard code that will just get us past.

Patrick: Yeah. Is it green? Shift it. Yeah.

Jeremy: Yeah, no, I totally, totally agree. All right. So let's talk a little bit more about, I guess, skillsets and things like that. Because there are so many different databases out there. Cassandra is just one and if you're a developer working just at the driver level, I guess, with something like Cassandra, it's not horrible to work with. It's relatively easy once a lot of these things are set up for you.

Same is true of MongoBD, or I mean, DynamoDB, or any of these other ones where the interface to it isn't overly difficult, but there's always some sort of something you want to build on top of it to make it a little bit easier. But I'm just curious, in terms of learning these different things and switching between organizations and so forth, there is a cognitive load going from saying, "I'm working on Cassandra," to going to saying, "I'm working on DynamoDB," or something like that. There's going to be a shift in understanding of how the data can be brought back, what the limitations are, just a whole bunch of things that you kind of have to think about. And that's not even including managing the actual thing. That's a whole other thing.

So, hiring people, I guess, or hiring developers, how much do we want developers to know? Are you on board with me where it's like, I mean I like understanding how Cassandra works and I like understanding how DynamoDB works, and I like knowing the limits, but I also don't want to think about them when I'm writing code.

Patrick: Yeah. Well, it's interesting because Cassandra, one of the things I really loved about Cassandra initially was just how it works. As a computer scientist, I was like, "This is really neat." I mean, my degree field is in distributed computing, so of course, I'm going to nerd out.

Jeremy: There you go.

Patrick: But that doesn't mean that it doesn't have mass appeal because it's doing the thing that people want. And I think that's going to be the challenge of any properly built service layer. I think I've mentioned to you before we started this, I work on a project called Stargate. And Stargate is a project that is meant to build a data layer on top of databases. And right now it's with Cassandra. And it's abstracting away some of the harder to understand or reason things.

For instance, with distributed computing, we're trying to reduce the reliance on coordination. There is a great article about this by Pat Helland about how coordination is the last really expensive thing that we have in development. Memory, CPU, super cheap. I can rent that all day long. Coordination is really, really hard, and I don't expect a new programmer to understand, to reason through coordination problems. "Oh, yeah, the just in time race conditions," and things like that.

And I think that's where distributed computing, it's super powerful, but then whenever people see what eventual consistency are, they freak out and they're like, "I just want my SQL Lite on my laptop. It's very safe." But that's not going to get you there. That's not a global database, it's not going to be able to take you to a billion users. Come on, don't cut ...

Jeremy: Maybe you don't need to be.

Patrick: ... your apps short Jeremy. You're going to have a billion users.

Jeremy: You should strive for it, at least, is how I feel about it. So that's, I guess, the point I was trying to get to is that if the developers are the ones that you don't want learning some of this stuff, and there's ways to abstract it away again, going like we talked about data as a service and APIs and so forth. And I think that's where I would love to see things shifting. And as you said earlier, that's probably where things are going.

But if you did want to run your own database cluster, and you wanted to do this on your own, I mean you have to hire people that know how to do this stuff. And the more I see the market heating up for this type of person, there is very, very few specialists out there that are probably available. So how would you even hire somebody to run your Cassandra ring? They probably all work at DataStax.

Patrick: No, not all of them. There's a few that work at Target and FedEx, Apple, the biggest Cassandra users in the world. Huawei. We just found out lately that Huawei now has the biggest cluster on the planet. Yeah. They just showed up at ApacheCon and said, "Oh yeah, hold my beer." But I mean, you're right, it's a specialized skillset and one of the things we're doing at DataStax, we feel, yeah, you should just rent that. And so we have Astra, which is our database as a service.

It's fully compatible with open-source Cassandra. If you don't like it, you can just take it over and use open-source. But we agree and we actually can run Cassandra cheaper than you can, and it's just because we can do it at scale. And right now Astra, the way we run it is truly serverless, you only pay for what you need, and that's something that we're bringing to the open-source side of Cassandra as well, but we're getting Cassandra closer to Kubernetes internally.

So if you don't want to think about Kubernetes, if you don't want to think about all that stuff, you can just rent it from us, or you could just go use it in open-source, either way. But you're right. I mean, it should not be a 2020s skillset is, "Get better at running Cassandra." I think those days should be, leave it to, if you want to go work at DataStax and run Cassandra, great, we're hiring right now, you will love it. You don't have to. Yeah.

Jeremy: So the idea of it being open-source, so again, I'm not a huge fan of this idea of vendor lock-in. I think if you want to run on AWS Lambda, yeah, most of what you can do can only run on AWS Lambda, but changing the compute, switching that over to Azure or switching that over to GCP or something like that, the compute itself is probably not that hard to move, right? I think especially depending on what you're doing, setting up an entire Kubernetes cluster just to run a few functions is probably not worth it. I mean, obviously, if you've got a much bigger implementation, that's a little different.

But with data, data is just locked in. No matter where you go, it is very hard to move a lot of data. So even with the open-source flair that you have there, do you still see a worry about lock in from a data side?

Patrick: Yeah. And it's becoming more of a concern with larger companies too, because options, #options. There was a pretty famous story a few years ago where the CEO of Target said, "I am not paying Amazon any more money," and they just picked up shop and moved from AWS to Google Cloud. And the CEO made a technical decision. It was like everybody downstream had to deal with that. And I think that luckily Target's a huge Cassandra shop and they were just like, "Okay, we'll just move it over there."

But the thing is that you're right, I mean, and I love talking about this because back when cloud was first starting and I was talking about it and thinking about it, just what do the clouds promise you? Oh, you get commodity scale of CPU and network and storage. And that's what they want to sell you because that what they're building. Those big buildings in north Virginia, they are full of compute network and storage, but the thing they know they need to hook you in and the way that they're hooking you in, there's some services that are really handy, they're great, but really the hook is the data.

Once you get into the database, the bespoke database for the cloud, one of the features of that database is it will not connect to any other database outside of that cloud, and they know that. I mean, and this is why I really strongly am starting to advocate this idea of this move towards data on Kubernetes is a way where open-source gets to take back the cloud. Because now we're deploying these virtual data centers and using open-source technology to create this portability. So we can use the compute network and storage, a Google, Amazon, Azure, OnPrem wherever, doesn't matter.

But you need to think of like, "All right. How is that going to work?" And that's why we're like, "If you rent your Cassandra from DataStax with Astra, you can also use the open-source Cassandra as well." And if we aren't keeping you happy, you should feel totally fine with moving it to an open-source workload. And we're good with that. One way or the other, we would love for you to use a database that works for you.

Jeremy: Right. And so this Stargate project that you're working on, is that the one that allows you to basically route to multiple databases?

Patrick: That's the dream. Right now it just does Cassandra, but there's been some really interesting ... There's some folks coming out of the woodwork that really want to bring their database technology to Stargate. And that's what I'm encouraged by. It's an open-source project, Stargate.io, and you can contribute any of the connectors for underlying data store, but if we're using GraphQL, if you're using GRPC, if you're using REST, the underlying data store is really somewhat irrelevant in that case. You're just doing gets and puts, or gets and sets. Gets and puts, yeah, that's right. Gets, sets, puts, it's a lot of words.

Jeremy: Whatever words. Yeah. Exactly.

Patrick: That's what I love about standard, Jeremy, there's so many to pick from.

Jeremy: Right, because there are ... Exactly, which standard do you choose? Yeah. So, because that's an interesting thing for me too, is just this idea of, I mean, it would be great to live in a perfect little cloud where you could say like, "Oh, well AWS has all the services I need. And I can just keep all my stuff there, whatever." But best of breed services, or again, the cost of hosting something in AWS maybe if you're hosting a Cassandra cluster there, versus maybe hosting it in GCP or maybe hosting it with you, you said you could host it cheaper than those could, or that we could host it ourselves.

And so I do think that there is ... and again, we've had this conversation about multi-cloud and things like that where it's not about agnostic, it's not about being cloud agnostic, it's about using the best of breed for any service that you want to use. And APIs seem to be the way to get you there. So I love this idea of the Stargate project because it just seems like that's the way where it could be that standard across all these different clouds and onto all these different databases, well I mean, right now Cassandra, but eventually these other ones. I don't know, that seems like a pretty powerful project to me.

Patrick: Well, the time has come. It's cloud native ... I work a lot with CNCF and cloud-native data is a kind of emerging topic. It's so emerging that I'm actually in the middle of writing a book, an O'Reilly book on it. So, yeah. Surprise. I just dropped it. This just in.

Yeah, because I can see that this is going to be the future, but when we build cloud-native, cloud applications, cloud-native applications, we want scale, we want elasticity, and we want self-healing. Those are the three cloud-native things that we want. And that doesn't give us a whole lot ... So if I want to crank out a quick REACT app, that's what I'm going to use. And Netlify's a great example, or Vercel, they're creating this abstraction layer. But Netlify and Vercel are both working, they've been partnering with us on the Stargate project, because they're seeing like, "Okay, we want to have that very light touch, developers just come in and use it," in building cloud-native applications.

And whenever you're building your application, you're just paying for what you use. And I think that's really key, not spinning up a bunch of infrastructure that you get a monthly bill for. And that bill can be expensive.

Jeremy: It seems crazy. Doesn't it seem crazy nowadays? Actually provisioning an EC2 instance and paying for it to run even if it does nothing. That seems crazy to me.

Patrick: There are start-ups around the idea of finding the instance that's running that's causing you money that you're not using.

Jeremy: Which is crazy, isn't it? It's crazy. All right. So let's go a little bit more into standards, because you mentioned standards. So there are standards now for a lot of things, and again, GraphQL being a great example, I think. But also from a database perspective, looking at things like TSQL and developers come into an organization and they're familiar with MySQL, or they're familiar with PostgreSQL, whatever it is. Or maybe they're familiar with Cassandra or something like that, but I think most people, at least from what I've seen, have been very, very comfortable with the TSQL approach to getting data. So, how do you bring developers in and start teaching them or getting them to understand more of that NoSQL feel?

Patrick: I think it's already happened, it's just the translation hasn't happened in a lot of minds. When you go to build an application, you're designing your application around the workflows your application's going to have. You're always thinking about like, "I click on this. I go there." I mean, this is where we wireframe out the application. At that point, your database is now involved and I don't think a lot of folks know that.

It's like, at every point you need to put data or get data. And I think this is where we've taught could be anybody building applications, which makes it really difficult to be like, "No, no, no, start with your data domain first and build out all those models. And then you write your application to go against those models." And I'll tell you, I've been involved in a few of these application boot camps, like JavaScript boot camps and things, they don't go into data modeling. It's just not a part of it.

Jeremy: Really?

Patrick: And I think this is that thing where we have to acknowledge like, "Yeah, we don't really need that anymore as much, because we're just building applications." If I build a React app, and I have a form and I'm managing the authentication and I click a button and then I get a profile information, I just described every database interaction that I need and the objects that I need. And I'm going to put my user profile at some point, I'm going to click my ID and get that profile back as an object. Those are the interactions that I need. At no point did I say, "And then I'm going to write select from where." No, I just need to get that data.

Jeremy: And I love thinking about data as objects anyways. It makes more sense, rather than rows of spreadsheets essentially that you join together, describing an object even if it's got nested data, like a document form or things like that, I think makes a ton of sense. But is SQL, is it still relevant do you think? I mean, in the world we're moving into? Should I be teaching my daughters how to write TSQL? Or would I be wasting my time?

Patrick: Yeah. Well, yes and no. Depends on what your kid's doing. I think that SQL will go to where it originally started and where it will eventually end, which is in data engineering and data science. And I mean, I still use SQL every once in a while, Bigtable, that sort of thing, for exploring my data. I mean for an analytics career or reporting data and things like that, SQL is very expressive. I don't see any reason to change that. But this is a guy who's been writing SQL for a million years.

But I mean, that world is still really moving. I mean, like a Presto and Snowflake and all these, Redshift, they all use Bigtable, they all use SQL to express the reporting capabilities. But ... And I think this is how you and I got sucked into this is like, well that was the database that we had, so we started using reporting languages to build applications. And how'd that work out?

Jeremy: Yeah. Well, it certainly didn't scale very well, I can tell you that, going back to sharding, because that is always something that was very hard to do. So I guess, I get the point that essentially if you're going to be in the data sciences and you actually need to analyze that data and maybe you do need to do joins, or maybe you need to work with big data in a way, that's a specialized aspect of it and I think people could dabble in that if they were just regular developers and they didn't want to go too deep.

But it sounds like the bigger, or the end goal here, maybe altruistic, is to just give people access to data. So even if they don't know SQL or they don't know something complex, just make it so that whatever data is there that anybody, with whatever level is, they can consume it.

Patrick: Yeah. And move fast with the thing that you're building. Actually, I use a Facebook term, but Facebook does do this. Internally there's a system called Occhio that provides gets and puts for your data, but it abstracts things like geographics and things like that. But the companies that are trying to move quickly, they understood this a long time ago. If you have to reason through, "Am I doing a full table scan? Is that an efficient interjoin?" If you have to reason through that, you're not moving fast anymore.

Jeremy: Right. Right. All right. Cool. All right, so let's talk about Astra a little bit more and this whole idea of, because Astra is the serverless version, the hosted version, the serverless version of Cassandra, right? Through DataStax?

Patrick: Right. And ...

Jeremy: Did I get that right?

Patrick: You got it right. And so it gives you full access. You could do Port 9042 if you still want to use a driver, but it gives you access via GraphQL, REST, and there's also a document API. So if you just want to persist your JavaScript API or JavaScript and then pull it back out your JSON, it does full documents. So it emulates what a MongoDB or DocumenDB does. But the important thing, and this is the somewhat revolutionary side of this, and again, this is something that we're looking to put into open-source, is the serverless nature of it.

You only pay for what you use. And when you want to create a Cassandra database, we don't even call it a Cassandra database on the Astra panel anymore. We just create a database. You give it a name. You click. And it's ready. And it will scale infinitely. As long as we can find some compute and network for you to use somewhere, it'll just keep scaling and that's kind of that true portion of serverless that we're really trying to make happen. And for me, that's exciting because finally, all that power that I feel like I've been hoarding for a long time is now available for so many more people.

And then if you do a million writes per second for 10 minutes and then you turn it off, you only pay for that little short amount of time. And it scales back. You're not paying a persistent charge forever.

Jeremy: I'm just curious from a technical implementation, because I'm thinking about PTSD or nightmares back of my days running Cassandra, and so I'm just trying to think how this works. Is it a shared tenancy model? Or is there a way to do single tenancy if you wanted that as a service?

Patrick: Under the covers, yes, it is multi-tenant, but the way that we are created ... so we had to do some really interesting engineering inside. So my RCO's going to kill me if I talk about this, but hey, you know what, Jeremy? We're friends, we can do this. He's like, "Don't talk about the underlying architecture." I'm talking about the underlying architecture. The thing that we did was we took Cassandra and we decomposed it into microservices mostly. That's probably, it's still Cassandra, it's just how we run it makes it way more amenable to doing multi-tenant and scale in that fashion where the queries are separated from the storage and things that are running in the background, like if you're familiar with Cassandra because it's a log structure storage, you ask to do compactions and things like that, all that's just kind of on the side. It doesn't impact your query.

But it gives us the ability to, if you create a database and all of a sudden you just hammer it with a million writes per second, there's enough infrastructure in total to cover it. And then we'll spin up more in the back to cover everything else. And then whenever you're done, we retract it back. That's how we keep our costs down. But then the storage side is separated and away from the compute side, and the storage side can scale its own way as well.

And so whenever you need to store a petabyte of Cassandra data, you're just storing, you're just charged for the petabyte of storage on disk, not the thousandth of a cluster that you just created. Yeah.

Jeremy: No. I love that. Thank you for explaining that though, because that is, every time I talk to somebody who's building a database or running some complex thing for a database, there's always magic. Somebody has to build some magic to make it actually work the way everyone hopes it would work. And so if anybody is listening to this and is like, "Ah, I'm just getting ready to spin up our own Cassandra ring," just think about these things because these are the really hard problems that are great to have a team of people working on that can solve this specific problem for you and abstract all of that crap away.

Patrick: Yeah. Well, I mean it goes back to the Dynamo paper, and how distributed databases work, but it requires that they have a certain baseline. And they're all working together in some way. And Cassandra is a share-nothing architecture. I mean you don't have a leader note or anything like that. But like I said, because that data is spread out, you could have these little intermittent problems that you don't want to have to think about. Just leave that to somebody else. Somebody else has got a Grafana dashboard that's freaking out. Let them deal with it. But you can route around those problems really easily.

Jeremy: Yeah. No, that's amazing. All right. So a couple more technical questions, because I'm always curious how some of these things work. So if somebody signs up and they set up this database and they want to connect to it, you mentioned you could use the driver, you mentioned you can use GraphQL or the REST API, or the Document API. What's the authentication method look like for that?

Patrick: Yeah. So, it's a pretty standard thing with tokens. You create your access tokens, so when you create the database, you define the way that you access it with the token, and then whenever you connect to it, if you're using JavaScript, there's a couple of collection libraries that just have that as one of the environment variables.

And so it's pretty standard for connecting the cloud databases now where you have your authentication token. And you can revoke that token at any time. So for instance, if you mistakenly commit that into your Git ...

Jeremy: Say GitHub. We've never done that before.

Patrick: No judging. You can revoke it immediately. But it also gives you our back, the controls over it's a read or write or admin, if you need to create new tables and that sort of thing. You can give that level of access to whatever that token is. So, very simple model, but then at that point, you're just interacting through a REST call or using any of the HTTP protocols or SQL protocol.

Jeremy: And now, can you create multiple tokens with different levels of permission or is it all just token gives you full access?

Patrick: No, it's multiple levels of protection and actually that's probably the best way to do it, for instance, if your CI/CD system, has the ability to, it should be able to create databases and tear them down, right? That would be a good use for that, but if you have, for instance, a very basic application, you just want it to be able to read and write. You don't want to change any of the underlying data structures.

Jeremy: Right. Right.

Patrick: That's a good layer of control, and so you can have all these layers going on one single database. But you can even have read-only access too, for ... I think that's something that's becoming more and more common now that there's reporting systems that are on the side.

Jeremy: Right. Right. Good.

Patrick: No, you can only read from the database.

Jeremy: And what about data backups or exporting data or anything like that?

Patrick: Yeah, we have a pretty rudimentary backup now, and we will probably, we're working on some more sophisticated versions of it. Data backup in Cassandra is pretty simple because it's all based on snapshots because if you know Cassandra the database, the data you write is immutable and that's a great way to start when you come to backup data. But yeah, we have a rudimentary backup system now where you have to, if you need to restore your data, you need to put in a ticket to have it restored at a certain point.

I don't personally like that as much. I like the self-service model, and that's what we're working towards. And with more granularity, because with snapshots you can do things like snapshot, this is one of the things that we're working on, is doing like a snapshot of your production database and restoring it into a QA cluster. So, works for my house, oh, try it again. Yeah.

Jeremy: That's awesome. No, so this is amazing. And I love this idea of just taking that pain of managing a database away from you. I love the idea of just make it simple to access the data. Don't create these complex things where people have to build more, and if people want to build a data access layer, the data access layer should maybe just be enforcing a model or something like that, and not having to figure out if you're on this shard, we route you to this particular port, or whatever. All that stuff is just insane, so yeah, I mean maybe go back to kind of the idea of this whole episode here, which is just, stop using databases. Start using these data services because they're so much easier to use. I mean, I'm sure there's concerns for some people, especially when you get to larger companies and you have all the compliance and things like that. I'm sure Astra and DataStax has all the compliance things and things like that. But yeah, just any final words, advice to people who might still be thinking databases are a good idea?

Patrick: Well, I have an old 6502 on a breadboard, which I love to play with. It doesn't make it relevant. I'm sorry. That was a little catty, wasn't it?

Jeremy: A little bit, but point well taken. I totally get what you're saying.

Patrick: I mean, I think that it's, what do we do with the next generation? And this is one of the things, this will be the thought that I leave us with is, it's incumbent on a generation of engineers and programmers to make the next generation's job easier, right? We should always make it easier. So this is our chance. If you're currently working with database technology, this is your chance to not put that pain on the next generation, the people that will go past where you are. And so, this is how we move forward as a group.

Jeremy: Yeah. Love it. Okay. Well Patrick, thank you so much for sharing all this and telling us about DataStax and Astra. So if people want to find out more about you or they want to find out more about Astra and DataStax, how do they do that?

Patrick: All right. Well, plenty of ways at www.datastax.com and astra.datastax.com if you just want the good stuff. Cut the marketing, go to the good stuff, astra.datastax.com. You can find me on LinkedIn, Patrick McFadin. And I'm everywhere. If you want to connect with me on LinkedIn or on Twitter, I love connecting with folks and finding out what you're working on, so please feel free. I get more messages now on LinkedIn than anything, and it's great.

Jeremy: Yeah. It's been picking up a lot. I know. It's kind of crazy. Linked in has really picked up. It's ...

Patrick: I'm good with it. Yeah.

Jeremy: Yeah. It's ...

Patrick: I'm really good with it.

Jeremy: It's a little bit better format maybe. So you also have, we mentioned the Stargate project, so that's just Stargate.io. We didn't talk about the K8ssandra project. Is that how you say that?

Patrick: Yeah, the K8ssandra project.

Jeremy: K8ssandra? Is that how you say it?

Patrick: K8ssandra. Isn't that a cute name?

Jeremy: It's K-8-S-S-A-N-D-R-A.io.

Patrick: Right.

Jeremy: What's that again? That's the idea of moving Cassandra onto Kubernetes, right?

Patrick: Yeah. It's not Cassandra on Kubernetes, it's Cassandra in Kubernetes.

Jeremy: In Kubernetes. Oh.

Patrick: So it's like in concert and working with how Kubernetes works. Yes. So it's using Cassandra as your default data store for Kubernetes. It's a very, actually it's another one of the projects that's just taking off. KubeCon was last week from where we're recording now, or two weeks ago, and it was just a huge hit because again, it's like, "Kubernetes makes my infrastructure to run easier, and Cassandra is hard, put those together. Hey, I like this idea."

Jeremy: Awesome.

Patrick: So, yeah.

Jeremy: Cool. All right. Well, if anybody wants to find out about that stuff, I will put all of these links in the show notes. Thanks again, Patrick. Really appreciate it.

Patrick: Great. Thanks, Jeremy.

View Details

About Mahdi Azarboon

Mahdi Aazarboon started working as a serverless specialist and evangelizing it through blog posts, conference talks and open source projects. He climbed up the corporate ladder, and currently works as Senior Manager - Cloud Presales at Cognizant. He helps big and traditional corporations to move into the cloud and improve their existing cloud environment. Having a hands-on background and currently working at the corporate level of cloud journeys, he has matured his overall understanding of serverless.

Linkedin: linkedin.com/in/azarboon/
Twitter: @m_azarboon

Watch this episode on YouTube: https://youtu.be/QG-N3hf1zqI

This episode sponsored by CBT Nuggets and Lumigo.

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm joined by Mahdi Azarboon. Hey, Mahdi. Thanks for joining me.

Mahdi: Hi. Thanks for having me.

Jeremy: So, you are a senior manager for cloud pre-sales in the Nordic region for Cognizant. So, I'd love it if you could tell the listeners a little bit about yourself, your background, and what it is that you do at Cognizant?

Mahdi: Yeah. Just a little bit of background, I started as a full stack developer, then I joined Accenture as a serverless specialist, and over there I started to play with AWS Lambda specifically. Started to do some geeky stuff, writing blog posts, and speaking at conferences and so on. Then, I was developing several solutions for multiple corporations in Finland, then I joined another consultancy company, Eficode, which are known for DevOps. It is very good, they have a good reputation for that in Nordic region. I was as a practice lead, AWS practice lead driving their business. Then, I joined my current company, Cognizant, and here I work as a pre-sales capacity. I'm not hands-on anymore, but basically I do whatever is needed to make our customers happy and make them to go to the cloud. So that means high-level solutioning, talking with the customer and as a senior architect, I comment about stuff, I make diagrams, And I translate business and technical stuff requirements, basically as an interface between the delivery and the customer side. Yeah, that's all.

Jeremy: Right. Awesome. All right. Well, so you mentioned in some of the blog posts that you were writing and some of that was a little while ago. And it's actually, I think there is some interesting perspective there. So I want to get into that in a little while, but I want to start by this idea or this post that you wrote about sort of what you need to know about Azure functions versus AWS Lambda and vice versa and it was sort of this lead-in to this concept of multi-cloud and not cloud-agnostic like being able to run the same workloads, but being able to understand the differences or maybe some of the nuances in Azure versus AWS and of course, that got extended to GCP and IBM cloud and some of these other things. But I'm curious why understanding different serverless services or different cloud services across clouds in this multi-cloud world we are living in now, why is that so important?

Mahdi: Yeah. That's a good question. First of all, I would like to clarify that whatever I'm telling in this podcast is just my personal opinion and doesn't reflect my employer. This is just to save myself.

Jeremy: Absolutely. Like a standard Twitter handle route.

Mahdi: Yeah.

Jeremy: Views are my own, right? Yeah.

Mahdi: I don't want to answer to my boss after this podcast. Answering to your question, the thing is that multi-cloud is inevitable and even AWS which was ... In the best practices, I remember like a few years ago, they were saying that, no, try to avoid that. They started to even admitting through their offerings that they are trying to embracing that multi-cloud with their Kubernetes offerings. The thing is that, well, whether AWS fans like it or not, Azure is gaining a lot of market share and it depends on the country. For example, in Finland at least AWS is really popular. But now I'm dealing, for example, in other countries like Norway or UK, Azure is very popular. I mean, you can just exclude yourself to be only with one cloud, but in my opinion, you are missing a lot of opportunities, both to learn and just as a company to embrace the capacities, because whether ...

Well, Azure provides some stuff which are better than AWS. I mean, I heard from a corporation that they really like AI capabilities of Azure much better than AWS and they do a lot of analytics. So it's inevitable whether many people like to admit it or not.

Jeremy: Right. Right. But so even the fact that it's inevitable and we talk about, multi-cloud is one of those terms ... I just talked to Rob Sutter about multi-cloud a couple of episodes ago and it's so expansive. I mean, everything from SaaS providers to, obviously the public cloud providers, to maybe even on-prem cloud, I know that sounds weird, but like your hybrid cloud and things like that. So the problem is that there are a lot of providers, there are a lot of SaaS products, things like that. I mean, are you advocating that people will try to become experts in multiple clouds or how do you sort of ... What level of knowledge do you think you need to have in order to work across multiple clouds?

Mahdi: I haven't met a single person who can claim to be expert in more than one cloud provider and I have talked with many experts because I have been running serverless in Finland and so I have been talking with many experts. None of them dared to claim that they knew it. I mean, even keeping up with one single cloud provider is a lot of work and I don't consider myself expert in any of them either, because I'm not hands-on anymore. The thing is that ... No, you don't have to be experts to work with different stuff. Of course, at some level you need some ... For example, you might need an Azure expert to work with Azure, AWS expert to work with AWS. But in my opinion, if you really want to keep up with the technology and so you need to be good in one provider, really good with that and then, know the fundamentals of the cloud, the best practices which are, I would say, it's irrespective of which cloud provider you are using there and be willing to learn.

For example, it happened to me. At that time, I mean, when I wrote that blog post, I was only working with AWS. Then they said to me that, okay, you have this project on Azure, go for it and I never touched Azure before. It was a lot of pain, but I learned a lot. So I mean, as I said, the fundamentals are same and now be expert in one and be willing to learn. In my opinion, that should be good enough.

Jeremy: Right. I'm curious, I think that's good advice to sort of be well-rounded. I mean, that's always good advice I think for technologists, going a mile wide and an inch deep is usually good enough. But like you said, being able to be an expert in a specific field or a specific technology or something like that can really help. So you think that's certainly a good career choice to sort of start to broaden your perspective a little bit?

Mahdi: Definitely. Actually, I was one of those AWS fans that really was following this Hero, Serverless Heroes, and so on, basically was parroting whatever AWS was telling and I was saying that I just want to come to work with AWS. Actually, it happened to be like that, but when I joined my current company, my manager said that most of the opportunities that you are filling, I mean, in my department, so is mostly Azure. So basically they said that it is as it is, and cope with it. And I felt very happy actually. When I, for example, see ... Well, I'm sure that anyone who is in the cloud gets many job offers from recruiters. I was thinking about it, at some point when I was AWS guy, at least in my experience, half of those job ads were irrelevant and ...

Jeremy: Right. Right.

Mahdi: ... depending on the country. For example in Finland, if you are Azure ... AWS is very popular at least and if you are Azure expert, you are going to miss a lot of opportunities. But at least in my experience, if you say that you are with that, you have worked with the other one, you know something, a lot of career opportunities opens up. This is my observation.

Jeremy: Right. Right. Yeah. And I think actually, you made a really good point and that's certainly, in terms of AWS heroes and so forth. I'm an AWS Serverless Hero and we get inside information but we spend a lot of time thinking about things the AWS way. AWS is very good at what they are doing with serverless and they have an interesting perspective in terms of what they believe serverless is supposed to be and what that roadmap looks like. But even just hosting this show and talking to so many different people in different clouds and different ways that they do it, getting that different perspective of how other people or other clouds think about serverless and how they are building it out. I think that's actually really good context to have.

Mahdi: Yeah, I agree. Actually, you are one of my heroes also, I was following you. But I should say that it has its own advantages and disadvantage was that I was in a kind of AWS bubble. But when I started to see that, okay, even AWS itself opens up having this multi-cloud offering and some serverless heroes start to write about that, I was like, okay, that's time for opening of your thing. But I mean, by that time actually, I already started to use Azure. So again, I mean, I would say that what you have been doing, actually heroes are doing a great job, really doing a great job.

Jeremy: Absolutely, totally agree.

Mahdi: Azure also have similar. If I remember correctly, they tried MVP, something like that.

Jeremy: I guess, that's MVP, yeah.

Mahdi: The thing is that, at least based on my observation, they have more or less same level of dynamics or a narrative between themselves. They also consider Azure more and AWS more and so. But I was lucky, maybe by the choice and so that somehow I had to join or use or attach to both communities. Yeah, it has been a very valuable experience.

Jeremy: Yeah. Yeah. So you went through that process, you were sort of an AWS convert or I guess, an Azure convert from AWS, and you stayed connected. But I know, that idea of transferring your skills and transferring the concepts and you mentioned sort of the pillars are the same as they are in AWS and you sort of have some of the general concepts, but as someone who went through that, what were the challenges that, what were some of the, I guess, challenges and the barriers that you faced going from AWS and that way of thinking into the Microsoft world?

Mahdi: That's a very good question. The thing is that in the department, at that time I was working at Accenture and actually all of us were big AWS fans because at least Accenture owned Avanade, so Azure was very separate, we were in an AWS bubble. Yeah. I'm sure that definitely AWS is much more mature in many aspects than Azure, no doubt. At least it was like that and I'm sure it's still like that. Their gap has been narrower, but that still might be the case. I remember at that time, many of my colleagues were really bashing down Azure, really bashing down and they were right. I mean, some of their services were really immature. But then again, I had the chance to ... Actually, it wasn't quite choice, they said to me, okay, this is an Azure project. Basically, it was a team, I would say quite junior, developed something on Azure, something that you never probably want to hear.

They developed everything in browser, infrastructure as a code nothing at all, they were junior, so they made quite many mistakes also, but they just made the app up and running. It didn't matter how or what, it was just running and that's all. So they told me that, okay, we need some little improvement, this was little improvement and that little improvement basically forced me to reverse engineer whatever they had done, and that required me to upgrade the whole application, because as I said, there was no infrastructure as a code, if I want to use it I had to use ... If I wanted to do local development, I had to use Windows, I had only Mac, so I had to change the complete platform. It was a very tedious process by itself. On top of that, I had to start to see how Azure functions work and that was another pain for that.

The thing is that I had AWS mindset and I was thinking that, okay, AWS is the best, they came out first with the cloud and Lambda, so Azure should be something like that. As I elaborated in the blog post, no, actually they are different and there are some small patches or nuances that makes some even days to find it out, but you need to find it out, otherwise, your app doesn't work. After a while when I reflected the things, I realized that, okay, of course, I was angry and pissed ... I was really bashing down Azure, it was fight of the dynamics over there, but after a while when I reflected through my whole process and actually I wrote in the blog post, I realized that part of the blame was on me because I was expecting Azure to work in the AWS way. No, that's not how it works.

I mean, when you look at, for example, authentication or the mindset, it's different. That requires a learning curve, I mean, you need to find out Stack Overflow and actually, the Azure community is really supportive. I really like it. They have their own community which is really supportive. So the pain basically was that ... Yeah, I had to find out how things work in Azure and what's different. But now that I'm working basically pre-sales in both of the cloud I can say that, again, fundamentals are same.

Jeremy: Right.

Mahdi: And these AWS architecture framework, there are five pillars. You can see that Azure has copied from AWS, it's obvious. Even they haven't changed the name. The naming is similar and you can find that it's just a bad copy. At least like few months ago that they had to implement for that. But at the end, I mean, Azure is catching up fast.

Jeremy: Right.

Mahdi: It's undeniable. And fundamentals are more or less same. I mean, if you want to make your app ... For example, you want to innovate, you should have shorter time to market. Basically, you need to use infrastructure as a code If you want to make your app really high-level appeal, you need to follow best practices, do maybe SRA. At the high level it's same, but when it comes to the detail level, it can be very different. Even the documentation was really confusing and it wasn't just me telling that.

Jeremy: Out of curiosity, was the documentation for AWS more confusing or was the documentation for Azure more confusing?

Mahdi: This is a million-dollar question. Actually, I thought that maybe it's me. I found the Azure doc very confusing, but I thought it's me, so I asked I think nine of my friends who are AWS experts that, "What's your opinion? Have you worked with Azure? Do you find documentation readable?" I think all of them said that it's confusing.

Jeremy: Yeah.

Mahdi: So I was like that, okay, then it's confusing. Then I talked with a few Azure experts who, they breathe in Azure, they are Windows guys and they never touched AWS and they said that, "No, documentation is good. Everything is fine." Actually, if I remember correctly one of them said that, "Actually, I find the AWS documentation confusing." It seemed like two different worlds, you know?

Jeremy: Right. I find them both confusing, actually.

Mahdi: Maybe now it has changed.

Jeremy: Right. Yeah. So, that's interesting. I mean, I think the documentation is a good ... Well, first of all having good documentation is important and I think they both have good documentation, but I do think it's organized differently, right?

Mahdi: Yes.

Jeremy: And again, it's organized more towards I think maybe that different mindset. But let's just talk a little bit about the maturity of those, because to be fair to Azure, I mean, Azure or Azure Functions, it has come a very, very long way. I remember way back in 2018, way back, I mean it seems like a long time ago at this point, seeing very early demos of Durable Functions and I remember thinking like, oh, that's just a mess, like that is not the way that you want to do that. Now fast forward three years, Durable Functions are pretty cool and they do a lot of really interesting things. It does take time to catch up. So certainly I would think your criticism of Azure Functions back then in terms of what it is now, that's probably there is a huge gap there.

Mahdi: Yeah. I'm sure that most of the criticism, the detailed one that I mentioned the blog post, I'm sure that many of them have either been fully addressed or they have been improved a lot. So that's why I don't want to focus that much on detail and I would focus more on the high-level things. Yeah.

Jeremy: Right. So speaking of the high-level things, let's go there for a second. So you mentioned like a well-architected framework, sort of this idea of their being something very, very similar, maybe even a carbon copy in Microsoft. But what about getting down, you said that your individual skills are kind of when you get into the weeds there, that is certainly different, so I mean for the most part though, event-driven, stateless computes, things like that, do those skills transfer over?

Mahdi: Yeah, they do. It's just a matter of implementation. For example, I can tell you, yes, those ones ... Well, there is some caveat. For example, I remember in Azure community, I was at that time, this probably has been changed, but I think it shows some kind of mindset. I was struggling to find out the observability tools of Azure, if I remember correctly it's what's called Application Insight, one of the tools, and they had some event driven insight, something like that which was, they call it near real-time. I remember that basically when I want to get the logs from the functions, it took three minutes to come up, three minutes. At the same time CloudWatch, for example, it was coming in 20 seconds, something like that, 10, 20 seconds and I mentioned it in their community.

If I remember correctly, it was a notable dude, either one of the product team, or he was a very notable dude and he said that three minutes time is, in my opinion, is near real-time. He said that and I remember we made a lot of joke out of that sentence with my colleagues about that.

Jeremy: I can imagine.

Mahdi: But that shows some kind of mindset. I mean, three minutes, I don't think is near real-time. Most probably this time has been reduced, but I just wanted to tell you their mindset about that. But, yeah, event-to-event stateless stuff, they are transferable. But when it comes to implementation, it's different. For example, as I mentioned that blog post, there was some stuff that you can do with an authentication with some, certain some, environmental variables in AWS, but that same thing in Azure, if I remember correctly, is done through something like service principles, it's different. So if you try to play with environmental variables, it turns out no, it doesn't work that way. It gets to very detailed stuff, that gets different. Yeah.

Jeremy: Yeah. Right. Right. Yeah. I'm curious to hear about like another sort of interconnectivity of what you would connect. I'm now trying to remember what they call bindings or triggers and bindings in Azure functions as opposed to events or actually event sources, I think we call them in the Lambda world. So would you look at the way that you connect to other services? Is that another thing that is similar between the two?

Mahdi: Okay. I should say that I don't remember that much of these details anymore, but as far as I remember, again, the high levels were more or less the same. Okay, they call it three gears, but I don't remember now what does AWS Lambda calls it. But it was more or less the same.

Jeremy: I can't even remember what it's called, it's like event sources or something like that.

Mahdi: Yeah. It was more or less same. Yeah, yeah, yeah.

Jeremy: Yeah.

Mahdi: And they had something like a bus, events bus in order to have a centralized event driven thing. It's same I would say.

Jeremy: Yeah.

Mahdi: Again, when it comes to poor person who has to implement it if he hasn't done it before. But the person who is doing the high-level architecture and so, I can easily see that, I mean, I don't see that much difference. But I know that if someone has to implement it and hasn't done it before, he will go through the most pain, because he has to find this small configuration things that, unfortunately, you need to make them. Otherwise, it doesn't work out. But high-level, it's same. It's event ...

Jeremy: Yeah.

Mahdi: Yeah.

Jeremy: I think the nuances are always those tough things. So thinking of the overall mindset here and sort of maybe the approach to serverless. So I know you went from AWS to Azure, but I'm curious, do you think it would be easier to go from Azure to AWS or easier to go from AWS to Azure?

Mahdi: Well, I came from this part of the river to the other one, so I can just speculate about the other part. But I would say it's more or less same, because again when I talk with a few Azure people who really have been breathing always in Azure and never touched or barely touched AWS, I felt that they are feeling same thing about AWS. So I would say it's more or less same. They need to go through the same pain, they will find AWS stuff very confusing, especially that they will not have that great community support of AWS, but they need to either do the Stack Overflow thing or have a enterprise support of AWS. I would say it's more or less same for them.

Jeremy: Yeah. I mean, I think that's interesting too just, that it is different enough that there is pain there, right? I mean, it would be nice if there was some standards and I know there's like the opening, the Cloud Computing Foundation is like open events and some of those things whatever, not that that's all working out for ... I think Kubernetes and Knative and those and some of those teams are implementing it or those projects, but I'm not sure the same things fall into AWS. But anyways, go ahead. You have any thoughts on that?

Mahdi: Actually, that Cloud Foundation, I was working at Eficode and they are really working that stuff. They are so good in Kubernetes. I find that also another world completely.

Jeremy: Yeah.

Mahdi: This Cloud Foundation stuff. I never had to implement any of that for any of our customers in any of the companies that I worked, that they were AWS or Azure. Yeah, some of them they used Kubernetes also, but that CNC or whatever it was ...

Jeremy: Yeah, CNCF.

Mahdi: Yeah, yeah. I found it, that's a different world for me also, I should say. Sometimes out of curiosity, I played with it, but I never ... Nobody ever asked me that, do you want to use that?

Jeremy: Right. Right. Yeah. No, that makes sense. All right. So we talked a lot about, we've been talking about the difficulties in switching between different cloud providers, but also the value of knowing those different cloud providers. And more so, so that you can build serverless applications. So let's talk about serverless in general. I know you are a little bit outside of the ... You're not in the developer role anymore. But this actually, could be really interesting to get your perspective on the management approach to this and how other companies are thinking about the value of serverless at a management level as opposed to ... I guess, even as a sort of planning level. So let me ask you this question then. Are you seeing companies looking at serverless and adopting serverless and that serverless mindset and then maybe a follow-up question would be, if they are not, why do you think they should be embracing serverless?

Mahdi: Okay. Firstly I'll answer the second part. Basically, the thing is that nowadays the world is fast changing. Many companies, many corporations basically, are benefiting from their existing market share or regulatories or the monopoly that they have. For now, it works. If they don't want to change basically if they have the mindset that things are working, what's the point for change. Most probably within a decade or so they are going to die, their business is going to die. Because the world is fast-changing and they need to have them to adapt to the market.

So ideally, they need to go through the pain and disrupt themselves. Disruption always brings pain. You cannot disrupt yourself and feel that everything runs smoothly. Ideally, they need to disrupt themselves, go through the pain and so become really agile in order to understand the customer feedback and deliver the value to the customer, what really the customer wants. They can either have this phase or they can ignore it and say that, okay, things are working, we are making money through our monopoly, regulatory, existing market share, whatever and then, their business is going to go away. These two choices, that's all. Yeah. Painful process to become more competitive and be ahead of customers or assume that everything is okay, and then at that time that's going to be very late.

Jeremy: Right. So let me go back to that first question then. So you are seeing people not doing that?

Mahdi: Okay. The thing is that what I'm telling is going to be biased because I'm working in a cloud team and whatever opportunity that they are going to bring to me, of course, you have the departments and the companies that they are interested in the cloud. So my mindset is a bit biased, but what I'm seeing is that it varies a lot and I mostly focus on corporations, because ... Yeah, of course, for startups it's much easier to go for that.

Jeremy: Right. Of course.

Mahdi: At least in Finland, my observation was that there are two ways. Either they are very ... it depends on the executive leadership. For example, a major bank in Finland, they say that, we want to go to the cloud and be, we want to go for that. And once, one of these big ones goes through that, there is going to be a domino effect on others. But there are some other ones say that, no, it's cloud, who is going to take care of the data? We are not going to do that and they don't touch it.

There are some other companies and their departments, I would say there are departments who are interested in trying things out and then, they have to fight internally with the more conservative departments. So I'm sure that there are three levels of that. But mostly, I work with the ones who are inclined toward using the cloud.

Jeremy: Right. Right. So then, the ones that are starting to dabble in the cloud, is that something where you see ... I mean, clearly there's lift and shift, right? Which I think we probably all understand at this point, it is not the best implementation or the best use of the cloud, right? That it is better to maybe use more native cloud services or cloud native services, I guess, to do that. So in terms of people just rehosting or maybe re-platforming, are you seeing this sort of rearchitecture, or I guess, this refactoring or is that something where companies are staying away from that?

Mahdi: First of all, I respectfully maybe have to disagree with you.

Jeremy: Okay.

Mahdi: Actually, I think rehosting is actually a good approach and that's what even AWS promotes for conservative companies who want to start working with the cloud and they want to get the fastest result in the shortest period of time, with the least amount of pain, it's better to do migration through the easiest one which is lift and shift. Easiest, everything is relatively.

Jeremy: Right.

Mahdi: And then, have a data-driven approach to see what really needs to be improved and then refactor or rearchitect or re-platform based on data. So in AWS terms, I'm sure you're right there with me, have that evolutionary architecture in a data-driven approach. So lift and shift, I don't consider bad at all. Actually, I consider it a very good cornerstone, stepping stone at the beginning, for the beginning.

Jeremy: Interesting. Okay.

Mahdi: Yeah. What was the other question?

Jeremy: No. I was just going to say, so you've got companies that are lift and shift, and, yeah ...

Mahdi: Oh, okay. Sorry. Sorry. Yeah. Sorry, I just remembered.

Jeremy: Yeah.

Mahdi: Sorry to interrupt you. Actually, I'm a bit careful about using the word cloud native. I remember, in a previous company that I was working, we had some philosophical fight about that and I'm sure that then everyone was dissatisfied and I had to have an authoritarian appearance that this is the definition of cloud native. I'm sure many of them hated me after that. But the word cloud native, I really struggle to find a consensus of what does it mean and if you spend some time, you realize that you will find a variety of definitions of that. So I'm picky for the word cloud native. There is a lot of fight can happen, what is exactly cloud native. Some consider Kubernetes cloud native. Some consider using AWS or Azure cloud native. So this is the picky ... this is a very controversial term, I would say. Yeah.

Jeremy: Well, let me interrupt you for a second. So when I think of cloud native, what I'm thinking of are services and components that are built specifically to run in the cloud, things like your API gateway at AWS or Azure functions or things that are like very much so built to run in the cloud environment where they do things. It's that serverless aspect. I think of it more serverlessly. I mean, I know containers and so forth fit in there as well. But that's how I think of cloud native. I think of cloud native as going beyond just your traditional VM and running everything on the VM and moving to the higher-level services that are more managed for you.

Mahdi: May I challenge you?

Jeremy: Absolutely.

Mahdi: So you just said that basically things that use cloud, like API gateway and so. And now I should ask more of a technical question. What is cloud?

Jeremy: Right. Well, that's another good question. Right.

Mahdi: Okay. I can tell you, based on these several definitions that I read and I reflected on them, I have this definition of cloud native, most probably many people I'm sure will disagree. So that's fine because it's very controversial. In my opinion, cloud native is very simple. If your application is architectured in a way that it can leverage the advantages of the cloud environment, then it's cloud native. Doesn't matter if it is on Kubernetes, if it's on AWS, if it's on Azure or so. If it can scale to zero and theoretically to infinity and you pay for only what you use, then it's cloud native. That's my definition of that and I read so many definitions, so I came up with this. But feel free to disagree with that, because many people disagreed with me. I'm fine with that.

Jeremy: That's all right. You are not the only one I'm sure, has differing opinions of what cloud native are. So let me ask this though because I think that's interesting, the way you explained the strategy of lift and shift of basically being able to say it's the, probably the lowest risk way to take an application that's on-prem and move that into the cloud and then to use data and so forth to kind of figure out what parts of the application might you want to migrate to, maybe again I don't want to overload the term, but more cloud native things. I think that's actually really interesting. I have found and I have seen many companies that seem to do this where it's more that they move things, they just rehost without really thinking through what that strategy is going to be and then they basically just end up having their on-prem in the cloud and not benefiting from some of those managed services and some of the benefits of the cloud that you get, they don't transfer on to them. That's what I have seen.

Mahdi: Well, you know it better than me. Your cloud environment is never perfect and it's always an ongoing operation. So I mean, going to the cloud ... Again, if you put your own frame in, put them I don't know, use EC2 or which VM or the AWS or Azure, that's a very good first step ...

Jeremy: Right. That's probably true.

Mahdi: ... but you need to be able to start leveraging that. At least get the data, which one is being used and hopefully, hopefully when you are going to the cloud, you have done some analysis and you have realized that some of the services even are not working with the cloud. Some of them need to retire, some of them cannot be rehosted. They must be rearchitected, because they are so legacy for that. But even again, assuming that you have done your homework and you have done rehosting, okay, you need to leverage that and go and see that all things that AWS or Azure provide, how much over-used or over-utilized or under-utilized are your CPUs and this kind of thing and according to that, do right sizing for that.

Jeremy: Right.

Mahdi: That's a good step for that. Then if they want, requires refactoring, try to I don't know, do refactoring and use more managed services for that. So again, rehosting is a good first step, but cloud is a long journey. I don't know who came up that cloud is cheap, I really don't know.

Jeremy: Right. No, I totally agree. You are right about the first step and I actually loved your point about which services might you be able to retire and not move at all because I think in a lot of these big companies, there are a lot of services that you probably don't need anymore or they are redundant or whatever and you could get rid of those moving to cloud. Good point. All right. I got a couple of more minutes and I want to go back to an article that you wrote. Now, this I think is like three years old and in terms of reading the article now, it's not relevant, because so many things have changed. But what's relevant is, what has changed and this was an article that was about the worrying and promising signals from the serverless community. I think this was an event you went to in Germany, they did this, and you have a couple of different points that you called out.

One of the points was that users have ignored security and that was a worrying sign for you. Where do you think sort of cloud security or more specifically serverless security is now? Do you think people are still thinking about it or have brought it front and center like it probably should be or do you think it's still a worrying factor?

Mahdi: Since I have implemented cloud solutions for I'll say mostly enterprises and a few startups, I haven't seen a single one of them using, having a cloud security specialist. Most of the corporations when they, at least in my experience, when they want to go to the cloud, they must address the security of it and typically because of the customer requirements, so they bring a security guy who has worked with this, let's put it this way, all their security stuff and he has to come in on the cloud part and it's funny that actually, sometimes I have to teach them basically. I remember they had a head of security for a customer. I really had to teach him and actually, I had the Lambda functioning in front of him and he was like, wow, is it really like that? I had to teach him what are the attack methods and it was funny. He had to sign off my solution that it is secure, but basically, I had to tell him what are the priorities.

Jeremy: You had to tell him what it was.

Mahdi: They address it from a traditional way. Yeah, they do some kind of a test, automated test and this kind of thing which is, yeah, definitely ... Again, I'm not a security expert, but as far as I understand, again they have some fundamentals which are safe, that's true, but when it comes to the cloud especially serverless and functional service, you will see that there is a lot of more attack vectors and unfortunately, these security experts, I have not seen any of them who have any expertise in that. I learned about it because I was curious about it and I started to work with basically professionals, some startups which provide professional security solutions for serverless. So that's how I got that, but again I had to go through the pain. It took few months to read so many stuff. But I haven't seen any security specialists who have been working on cloud projects who have done this.

Jeremy: Yeah.

Mahdi: So I would say customers, they consider it, but no, there is still a lot more way to mature.

Jeremy: They are not addressing it. Yeah. It's funny because I remember that in 2018, 2019, there were a couple of companies that were in the serverless security space and they were all acquired. So now they are part of larger platforms which is ...

Mahdi: Exactly.

Jeremy: ... great for them, don't get me wrong. All right. So then another thing you said and I think this is important, because the biggest complaint that I always hear about serverless is, just the workflows are not easy. So you had mentioned that DevOps was finding its way and that was sort of a promising signal, you think that we've ... I mean, we have got a lot of tools for serverless now. Speaking of Azure, the way to deploy an Azure function right through VS Code now with the plugins is really, really slick and Serverless Frameworks, SAM, CDK, all these are there, Terraform and so forth. I mean, have we gotten to some stability around serverless and sort of mixing in DevOps there?

Mahdi: Based on my experience, at least the ones that I have worked with, I can say that, yes, DevOps is now a part of a solution that's provided to the customer and maybe it's correct because personally, I went through the pain whenever I proposed any solution for the customer, so they are always using infrastructure as a code and always try to have a DevOps-centric viewpoint about your solution. So I try to push for that and, yes, I find customers receptive about that. It seems to me that, now DevOps is not one of those buzzwords for cool kids who just want to do this stuff, even the corporate guys are more receptive with that. Again, there is more way to really do the DevOps stuff, because you know that many companies claim that they are doing DevOps, but in reality, they are not. You know this better than me.

Jeremy: Right. Of course.

Mahdi: But, yeah, it's good. I'm happy for that. I mean, a few years ago DevOps was one of those buzzwords, but now I don't think it's buzzword anymore.

Jeremy: Yeah. Yeah. And I think that serverless has actually opened up a lot of making it easier for teams to do automation and things like that, there's a lot that you can do because you have that little bit of compute power that you can do something with. So I think that's definitely promising. So speaking of sort of compute power and other things that you can do with it, one of the things you mentioned was that you saw as a promising sort of signal was, that serverless-based prototypes were on the rise, meaning different services, so whether it was cues or whatever or I guess Lex and things like that, all kinds of services that allow you to do different things that are specializing in different capabilities. So how do you feel about that now? Because there are a lot of those APIs out there.

Mahdi: Yeah. Actually, I also find that even from these legacy corporations that I have been working with, I like the idea that now, they definitely when they want to do migration especially or this kind of thing or do anything cloud, first they do POC. Yes, I find it good. Honestly, I was sometimes impressed that, oh, from some people that I would never expect them to use this one, first let's do POC, then see what's come out. Oh, really? Yeah, it's good in my opinion. It's finding its way.

Jeremy: Yeah. Yeah. No, I like that too and I think you are right about proof of concept, because it's just one of those things where even if it's expensive to use the Google Vision API or something like that, it's a really good way to prove out how that fits into whatever the business use case you have for that and then like you said, you can certainly take a step further and create more sophisticated or I won't say sophisticated but maybe more integrated tools or something like that, that would work around that. So I think that's interesting, allow people to fail fast, learn quickly, and just build out their applications.

Mahdi: Yeah. When we say POC, I should say that I wouldn't exclude it only to this cool new serverless or what the AI stuff that AWS and Azure provide. Even for migration actually, POC is highly recommended. Again, I was working for some period of time for, I would say, one of the most conservative banks in Finland, small and conservative, for consultancy, but even then as we are trying to push the cloud and even then they said that, "Yeah, first let's do a POC of migration and see what's going to happen." Again, there really I was surprised. I would never expect it from them. But the idea of fail fast and learn fast, I think at least that it requires some level of maturity to reach that.

Jeremy: True.

Mahdi: That really needs more room for improvement, fail fast, learn fast. Yeah. Just something, I don't know, I would like to address about this cloud stuff if I can.

Jeremy: Yeah, absolutely.

Mahdi: Yeah. Basically, when companies or customers decide to go to the cloud, I'll recommend that don't look at only the technical aspect of it, because I see that there is, at least there is lot of debate for example ... At least it was like that. AWS, Azure, or this kind of thing, at the end I'll say that most of the things it doesn't matter that much. I mean, it depends on their, sometimes company policy, how much discount you can get, how much funding you can get from the cloud provider. So it's not really the technical people who decide, sometime it's the executive who decides.

Jeremy: Right.

Mahdi: But even then, when you go to the cloud, in my opinion as much as the technology and maturity of the cloud provider matters, the amount that your company is ready to change its operations is also important. This is my favorite example, that I developed and I would say at that time at least a state-of-the-art serverless solution, DevOps, or CI/DC stuff for a major bank in Finland and I was the first one who managed to do that among so many consultants that they have. It was really good. I'm proud of what I did and actually, I open-sourced that. It was really basically we could deploy multiple times per day and we went to their release manager and I said that, "Okay. It's like that. Everything is perfect. DevOps, CI/CD, we can release multiple times per day." And she said that, "No. It doesn't work like that. We need to release once per month," and we have to go through a very painstaking process, fill out so many useless documents.

It didn't matter how much I tried to convince her that, "Well, the idea is different. I mean, you need to do small deployment. This way actually you have less risk. You deploy once per month. Still every time something goes wrong, but when you do a more frequent deployment, your risk is lowered." She said, "No. We are a bank. It is as it is. Sorry." Most of that effort that I made at least at that time went to waste basically, because the process was legacy, even though the technology was good. I'm sure that by now, they have changed because I was among those innovators basically or the early adopters who made through that. But in my opinion, technology matters but operation and processes and release stuff also matters and everything needs to change. So basically it needs to be holistic approach of going to the cloud, not just implementing from technical viewpoint.

Jeremy: Mahdi, thank you so much for having this conversation with me. This was a lot of fun and then I love people who have sort of experienced, from moving from one cloud to another. It's a huge shift, but I think your advice here is great, just to sort of know those basics on those other platforms and do that. So if people want to reach out to you or find out more about, follow you on Twitter, things like that, how do they do that?

Mahdi: Well, I have a Twitter account, but nowadays I mostly put non-service stuff, but LinkedIn is a good option for me.

Jeremy: Okay. Great. And it's m_azarboon on Twitter and then, I will put LinkedIn and Twitter and that in the show notes as well. Thanks again, Mahdi.

Mahdi: Thank you very much for having me. Bye-bye. Thank you.

View Details

About Amy Arambulo Negrette

With over ten years industry experience, Amy Arambulo Negrette has built web applications for a variety of industries including Yahoo!, Fantasy Sports, and NASA Ames Research Center. One of her projects modernized two legacy systems impacting the entire research center and won her a Certificate of Excellence from the Ames Contractor Council. She has built APIs for enterprise clients for cloud consulting firms and led a team of Cloud Software Engineers. Currently, she works as a Cloud Economist at the Duckbill Group doing bill analyses and leading cost optimization projects. Amy has survived acquisitions, layoffs, and balancing life with two small children.

Website: www.amy-codes.com
Twitter: @nerdypaws
Linkedin: linkedin.com/in/amycodes

Watch this episode on YouTube: https://youtu.be/xc2rkR5VCxo

This episode sponsored by CBT Nuggets and Lumigo.

Transcript
Jeremy: Hi everyone, I'm Jeremy Daly, and this is Serverless Chats. Today, I'm joined by Amy Arambulo Negrette. Hey, Amy thanks for joining me.

Amy: Thank you, glad to be here.

Jeremy: You are a Cloud Economist at the Duckbill Group, so I'd love it if you could tell the listeners a little bit about yourself and your background and what you do at the Duckbill Group.

Amy: Sure thing. I used to be an application developer, I did a bunch of AWS stuff for a while, and now at the Duckbill Group, a cloud economist is someone who goes through cost explorer and your usage report and tries to figure out where you're spending too much money and how the best to help you. It is the best-known use of a small skill I have, which is about being able to dig through someone's receipts and find out what their story is.

Jeremy: Sounds like a forensic accountant, maybe forensic cloud economist or something to that effect.

Amy: Yep. That's basically what we do.

Jeremy: Well, I'm super excited to have you here. First of all, I have to ask this question, I've known Corey for quite some time, and I can imagine that working with him is either amazing or an absolute nightmare. I'm just curious, which one is it?

Amy: It is not my job to control Corey, so it's great. He's great to talk to. He really is fully engaged in any conversation you have with him. You've talked to him before, I'm sure you know that. He loves knowing what other people think on things, which I think is a really healthy attitude to have.

Jeremy: I totally agree, and hopefully he will subtweet this episode. Anyways, getting into this episode, one of the things that I've noticed that you've done quite a bit, is you create technical content. I've seen a lot of the talks that you've given, and I think that's something that you've done such a great job of not only coming up with content and making content interesting.

Sometimes when you put together technical content, it's not super exciting. But you have a very good way of taking that technical content and making it interesting. But then also, following up with it. You have this series of talks where you started talking about managing FaaS, and then you went to the whole frenemies thing with Fargate versus Lambda. Now we're talking about, I think the latest one you did was about Lambda and the container support within Lambda. Maybe we can just go back, or start at a point where, for people who are interested in maybe doing talks, what is the reason for even creating some of these talks in the first place?

Amy: I feel a lot of engineers have the same problem, just day-to-day where they will run into a bug, and then they'll go hit the all-knowing software engineer, which is the Google search engine, and have absolutely either nothing come up or have six posts that say, I'm having this problem, but you won't ever get an answer. This is just a fast way of answering those questions before someone has to ask.

Jeremy: Right. When you come up with these ... You run into this bug, and you're thinking to yourself, you can't find the answer. So, you do the research, you spend the time digging through, and finding the right way to solve it. When you put these talks together, do you get a sense that it's helping people and then that it's just another way to connect with the community?

Amy: Yeah. When I do it, it's really great, because after our talk, I'll see people either in the hallway, or I'll meet someone at a booth, and they'll even say, it's like, I ran into this exact same problem, and I gave up because it was such a strange edge case that it was too hard to fix, and we just moved on to another solution, which is entirely possible.

I also get to express to just the general public that I do, in fact, know what I'm talking about, because someone has given me a stage to talk for 30 minutes, and just put up all of my proofs. That's an actually fun and weirdly empowering place to be.

Jeremy: Yeah. I actually think that's really interesting. Again, for me, I loved your talks, and some of those things are ... I put those things at the back of my mind, but I know for people who give talks, who maybe get judged for other reasons or whatever, that it certainly is empowering. Is that something where you certainly shouldn't have to do it. There certainly should be that same level of respect. But is that something that you found that doing these talks really just sets the tone, right off the bat?

Amy: Yeah, I feel it does. It helps that when someone Googles you, a bunch of YouTube videos on how to solve their problem comes up, that is extremely helpful, especially ... I do a lot of consulting, so if I ever have to go onsite, and someone wants to know what I do, I can pull up an actual YouTube playlist of things that I've done. It's like being in developer relations without having to write all of that content, I get to write a fraction of that content.

Jeremy: Right. Unfortunately, that is a fact that we live with right now, which is, it is completely unfair, but I think that, again, the fact that you do that, you put that out there, and that gives you that credibility, which again, you should have from your resume, but at the same time, I think it's an interesting way to circumvent that, given the current world we live in.

Amy: It also helps when there are either younger engineers or even other younger professionals who are looking at the tech industry, and the tech industry, especially right now, it does not have the best reputation to be able to see that there are people who are from different backgrounds, either educationally or financially, or what have you, and are able to go out and see someone who has something similar being a subject matter expert in whatever it is that they're talking about.

Jeremy: Right. I definitely agree with that. That's that thing, where the more that we can amplify those types of voices and make sure that people can see that diversity, it's incredibly important. Good for you, obviously, for pushing through that, because I know that I've heard a lot of horror stories around that stuff that makes my blood boil.

Let's talk to some of these people out here who potentially want to do some of these talks, and want to use this as a way to, again, sell themselves. Because I can tell you one thing, once I started writing blog posts and doing talks and doing those sorts of things, clearly, I have a very different background, but it just gave me a bunch of exposure; job offers and consulting clients and things like that, those just become much easier to get when you can actually go out there and do some of this stuff.

If you're interested in doing that, I think one of the hard things for most people is, what even makes a good talk? You've come up with some really great talks. What's that secret sauce? How do you do that?

Amy: I think it can also be very intimidating since a lot of the talks that get a lot of promotion are always huge vendor events that they're trying to push their product, they're trying to push a solution. That usually takes up a lot of advertising real estate, essentially, where that's what you see, that's what you see all the threads and everything. When you actually get to these community conferences, or even when I would speak at AWS Summit, it was ... I had a very specific problem that I needed to solve. I ran into a bug, the bug was not in the documentation, because why would it be?

Jeremy: Why would you put that in there, right?

Amy: Of course. Then Google, three pages down, maybe put me on the path to finding the right answer, and it's the journey of trying to put all of the bug fixes in place to make it work for your specific environment and then being able to share that.

Jeremy: Right, yeah. That idea of taking these experiences that you've had, or trying to solve a problem, and then finding the nuances maybe in solving the problem as opposed to the happy path, which it's always great when you're following a blog post and it says, run this command, then run this command, then run this command. Well, what happens on that third command when the thing blows up, and you have no idea what to do? Then you end up Googling for five hours trying to find your way out of that.

You take this path of, find those bugs or find that non-happy path and solve it. Then what do you do around there? How do you then take that ... You got to make that interesting somehow.

Amy: Yes. A lot of people use gifs and memes. I use pictures of food and screencaps from Dungeons and Dragons. That's usually just different enough that it'll snap someone just out of their phone going, "Why is there a huge elf on my screen trying to attack people screaming elf errors." Well, that's because that's what they thought it would be great to call it. It's not a great error code. It doesn't explain what it is, and it makes you very confused.

Jeremy: Right. Part of that is, and again, there's that relatability when you create talks, and you want to connect with the audience in some way. But you also ... This is the other thing that I've always found the hardest when I'm creating talks, is trying to find the right level. Because AWS always does this thing where they're like, it's a 200 level, or it's a 400 level, and so forth. I think that's helpful, but you're going to get people of all different skill levels, and so forth. How do you take a problem like that, and then make it relatable, or understandable, probably? Find that right level?

Amy: The way I see it, there's going to be at least one person of these two types in the room that are not going to be your target audience, someone who doesn't know what you're talking about, but sees that a tool that they're considering is going to pose a problem, and they want to know how difficult it is to fix it. Or there's going to be a business person who has no technical background, and they just want to know if what they're evaluating is worth evaluating, if this error is going to be so difficult to narrow down and try to resolve that, yes, why would we go through something that my engineers are going to spend hours to try to fix something that's essentially a configuration issue?

When I write any section of a talk, I make sure that it addresses a person who may not have come into that with that exact problem in mind. For the people who have, they'll understand the ... In animation, it's called key images, where there are very specific slots where you understand the topic of what is happening and the context around it. I always produce more verbose notes that go with my presentation. I usually release it either at the end of the day, or later on that week, once everyone has had time to settle, and it provides a tutorial-esque experience where this is what you saw, this is how you would actually do it if you were in front of a screen.

Jeremy: Yeah.

Amy: There are people who go to technical talks with a laptop on their lap because they're also working while they're trying to do it. But most of the time, they're not going to have the console open while you're walking through the demo. So, how are you going to address that issue? It's just easier that way.

Jeremy: I like that idea too, of ... I try to do high-level bullet points, and then talk about the bullet point. Because one thing that I try to do, and I'd love to hear your thoughts on this as well. Here I am picking your brain trying to make my own talks better. But basically, I do a bullet point, and then I talk through it. I actually animate the bullet points coming in.

I'm not a huge fan of showing an entire slide with all the bullet points and then letting people read ahead, I bring a bullet point in, talk about the bullet point, bring another bullet point in. Is that something you recommend doing too? Or do you just present all the concepts and then walk people through it?

Amy: I think it depends. I tend to have very dense slides, which is not great for reading, especially if you're several rows back. I truly understand that. But the way I see it, because I also talk very fast when I'm on stage that I want there to be enough context around what's happening, so that if I gloss over a concept, then you visually can understand what's happening.

That said, if that's because the entire bullet block on my slide is going to be about a very specific thing that's happening. It's not something that you have to view step-by-step. Now, I do have a few where, especially in a more workshop scenario, where you're going, I want you to think about this first and then go on to this next concept. I totally hide stuff. I just discovered for a talk that I was constructing the other day, that there's an animation that drops them down like index cards, and that's now my favorite animation right now.

Jeremy: When you're doing that, like because this is the other thing, just for people who have ever ... If you're out there and you've ever written a talker or you've given a talk, the first iteration of it is never going to be the right one. You have to go through and you have to revise. It is sort of weird, and I don't know, maybe you felt this way too, in the pre-pandemic world, when you would give talks in person, most of the time, you'd give it to a relatively small audience, a couple of hundred people or whatever, as opposed to now, when we do talks, post-pandemic, and they're online, it's like, they're immediately available online.

It's hard to give the same talk over and over and over and over again, without somebody potentially having seen it. A lot of work goes into a single talk. Not being able to use the same time over and over again, is not great. But, how do you refine it? Is it that you tested it with a live audience, or do you use a family member or a friend, or a colleague? How do you test and refine your talks?

Amy: I'm actually an organizer at a meetup group, and specifically built around giving people of marginalized gender identities, and a place to stage and write technical content. It is a very specific audience.

Jeremy: I can imagine.

Amy: But it addresses that issue I had earlier about visibility, it also does help you ... If you don't have a lot of contacts in this industry, just as an aside, technical speaking is a way to do it, because everyone loves talking to each other after the stress has worn off, and you become the friendliest person after you've done that.

But also, there are meetup groups out there, specifically about doing technical feedback, or just general speaking feedback. If you want to do something general, Toastmasters is a great organization to do. If you want to do strictly technical, if you do any cloud-related stuff, the DevOps communities are super friendly, even if it's not specifically about DevOps. I'm not a DevOps person, but I have a lot of DevOps friends. Some of my best friends are DevOps people.

And you can get on a meetup or a Zoom call and just burn through your slides for about 10 or 15 minutes and see ... Your friends will be very honest with you, in a small group.

Jeremy: Right. One of the things I did notice, too, giving a speech in person or giving your talk in person versus giving a talk via Zoom call, is sometimes when you don't hear any laughs or chuckles from a little joke that you make in there, it can feel very lonely in that space after you're waiting for something in there, but. It's a little bit ...

Amy: It's worse when there are people in the room. I assure you, it is so much worse.

Jeremy: That is very true. If something falls flat, that's a good point. Just going back to more this idea of creating good talks, and what makes a good talk. Where do you find ... You mentioned, maybe it's a vendor conference or something and you maybe install the vendor stuff, and you find the bugs and so forth. But is there any other places that you get inspiration from? Are there any resources you use to sort build some of these talks?

Amy: Again, the communities help. The communities will tell you, really, it's like, I don't understand this thing, can someone hop on a call with me for real quick minute and explain why this concept is so hard? That's a very good place to base your talk off. As far as making them engaging, and interesting, I tend to clone video gaming videos, just because that's what I watch. I know, if it's going to be interesting to me, then it will probably be at least different than the content that's out there.

Jeremy: Right. That's a good way to think of things too, is if it's something that you find interesting, chances are, there are lots of other people that will find that interesting. All right, let's go back to just this idea of creating new talks. You had mentioned this idea of, again, finding the bugs and so forth. But one of the things that I think we see quite a bit is always that bleeding edge stuff. People always want to write content about something new that happened.

I'm guilty of this, I would think from a serverless standpoint where you're talking about things that are really, really bleeding edge. It's useful and they're interesting. Certainly, if you go to a conference about serverless, then it's really nice to see you have these talks and what might be possible. But sometimes when you're going to more practical type things. Again, even DevOps Days, and some of those other things, I think you've got attendees or talk listeners who are looking for very practical advice.

I guess the question is like, how do you take a new piece of content, one of these problems, whatever it is. I guess, how do you keep finding new content is probably the better way to ask that question?

Amy: Well, to just roll back just a little bit. My problem with bleeding edge content, I love watching it, but bleeding edge content will almost always be a product demo because it's someone who developed a new solution, and they want to share with everybody, which is just going to walk you through how it's used, which is great, except, and this is just a nature of what the cloud industry is like, all of this stuff, it changes day-to-day.

These tools may not be applicable in a few months, or they may become the new standard. There's no way to tell until you're already six months out, and by then, they've already gone through several product revisions. I once did a talk where I was talking about best practices, and AWS released their updated best practices the day before my talk, and I had to update three slides. It threw off my timing, it was great.

That's just one of those kind of pitfalls that you have to roll with. As far as getting new content, though, especially if you're dealing ... It depends who your audience is, because my audience tends to be either ICs or technical leads, and by then you're usually in a company ... If you're not developing these bleeding edge solutions, you're just using the tools that's out there already.

You had brought up my "Serverless Frenemies," which is still my favorite title of any talk that I've ever made, because when I did the managing containers one, and I love all my Devro friends, but they all got into my mentions about why don't you just use Fargate? If you're at the containerization stage, why don't you just use Fargate, because it's not even close to the same thing, it is closer to Kubernetes than it is to Lambda, and I'm looking for a Lambda-like solution. That's what that whole deal was about, and I was able to stretch that out into I think 30 minutes because Twitter will tell you what's wrong, whether or not it's accurate or not, and whether or not they're actually your friends. They are my friends, but come on.

Jeremy: Twitter can definitely be brutal. I think that, and maybe unpack a little bit what you were saying, is you're creating content around existing tools. One way to do it is, you're using existing tools, you're creating content around that, or you can create content around that. Looking at those solutions, you introduce a new solution to something, or you're even using an existing tool, nothing's perfect. You had mentioned that idea of bugs and so forth. But just, I guess new solutions, or just solutions, in general, maybe higher-level abstractions, everything creates some new type of problem that you have to deal with, and that's probably a pretty effective way to generate new content.

Amy: It is. If you ever have to write down an RCA, which, for those who have not had the pleasure of doing one is called a root cause analysis, where you took down production, and you had to explain why.

Jeremy: Yep.

Amy: Or you ever did this, hopefully, in stage, or hopefully, in development where you ran into a situation where ... I had a situation once where Lambda would not delete itself. I call it my Skynet problem where it just hit a stage where it was both trying to save and delete at the same time. It would lock itself and I had to destroy the entire stack and send that command several times just to force that command through.

If you ever have a problem like that, that is a thing that you write up instantly, and then you turn it into slide decks, and then you go to SlidesCarnival, you throw a very flashy background on it, and next thing you know, you have a TED talk, or a technical talk.

Jeremy: Right. The other thing too, is, I find use cases to be an interesting, just like ... Non-traditional use cases are kind of fun too, how can I use this in a way that it wasn't meant to be used, and do something like that?

Amy: I love those. Those are my favorite. I love watching people break away from what the tutorial says you have to do, and I'm going to get a little weird with it, and that to me is totally fascinating. When the whole, I fed these scripts into a computer meme came out, I thought that was super fascinating because that was something a company I had worked for did, they used analytics ... I used to work for Fantasy Sports, to write color commentary for your Fantasy Football team, and they would send it out.

If you did really well, you would get a really raving review, and if you did really poorly, you would get roasted by a computer, and then that gets sent to everyone in the league, and it's hilarious. But that is not a thing that you would just assume a computer would do, is just write hot takes on your Fantasy Football team.

Jeremy: That's ... Sure, go ahead.

Amy: It's so much fun. I love watching people get weird with the tools that are there.

Jeremy: There are times where you could do something like that, you could maybe create a content around some strange use case or whatever, and I love that idea of getting weird with that. The other part of it, though, is that, I guess, if you're sitting through a talk, and it's some super interesting problem that you're listening to, and again, I don't know, maybe it's some database replication thing, that you're just really into, whatever. That makes sense. But I think the majority of problems that developers have, are not that interesting, they're just frustrating.

Probably the worst thing to do is wanting to sit through a talk that talks about some frustrating issue you have. Is there a way to basically say, "Look, I have a problem that I want to talk about. It's not the most interesting problem, but how can you flip that and take a problem that's not interesting and make it interesting?

Amy: The batching containers and the frenemies talk was all based off of a bin library error from within the Lambda AMI. That, on paper is extremely boring, and should be a thing that you can easily look up, it is not. When I went around it trying to make tracking down library errors interesting, just saying it is very slow and can drain the energy out of your voice.

But, I put a lot of energy into my work in general, and that's just how I had to approach pulling these talk is like, I like what I do, just, generally. When I try to explain what I do to people, it sounds super boring, and I own that. Now I'm doing it with spreadsheets, which is much, much worse. But when I tell people, it's not about the error itself, it's about everything that happened to make this one particular error happen. The reason why this error happened was because Lambda uses AWS's very specific Linux AMI when they did not used to, and they left stuff out for either security or performance purposes.

Whether or not we as a group agree with that, that's a business decision that they made. How does their business decision affect your future business decisions and your future technical ones? Well, that becomes a way more interesting conversation, because it's like, we know this is going to break at this part, do we still want to use SSH? Do we still need it for this reason? You can approach it more from a narrative standpoint of, I wasted way too much time with this, did I need to? It's like, well, you shouldn't have, this should not have happened, but no bug should have happened, right?

Jeremy: Right.

Amy: You work through your process of finding a solution instead of concentrating on what the solution is because the solution they can look up in your show notes later.

Jeremy: Right. No, I love that idea of documenting your process as opposed to just the solution itself. You find the problem, you pull the thread and where does that take you? I think to myself, a lot of times I go down the rabbit hole on trying to find the solution to a problem that I have or a bug fix, whatever. Sometimes, the resolution is underwhelming. Maybe it's not worth sharing. But other times, there's a revelation in there. I think you're right, with a little bit of storytelling, you can usually take that and turn that into a really interesting talk.

Amy: One of the things it will also do, if you look at it from a process and from a narrative standpoint, is that when you take this video, and you send it to either a technical lead or a product manager, they'll understand what the problem was because you did not bog it down with code. There's very little live code in mine because I understand that people build things differently, just because every code is as different as every person. I get that and I've come to terms with it. This is the best way to share that information.

Jeremy: Absolutely. All right, let's wrap up the idea of building talks. What is your advice to someone who is starting out new? What's the best way for them to get started, or what's just some general advice for people starting to build talks?

Amy: The best content new engineers can do, and that's mostly because this is never the standpoint from which tutorials are ever written in, is that, as someone who knows very little of the way a language or a framework should work, write down your process, the entire thing on you getting either a framework onboarded, how you build, and a messaging system, things that people have written a billion times because chances are, one, you got that work from someone else's blog post or their documentation, and you can cite that. And two, when you do it that way, you not only get into the habit of writing, but you get in the habit of editing it in a way that makes it more palatable for people who are not in your specific experience.

When you do it this way, people can actually see, from an outsider's perspective, exactly what is hard about the thing that they built, or what people who do not have a different level of experience are going through. If a tutorial is targeted at engineers who know where the memory leaks in PHP are, that's the thing that comes with experience, that is not the thing that can be trained.

When a new engineer hits that point, and they found it in a new framework where you fix it, then you start knowing where to fix other problems. That way more senior engineers and more vetted people can learn from your experience, and then they will contact you and they will teach you how to find these issues, so you don't run into them again, and you end up with someone you can just bounce ideas off of. That's how you get pulled into these technical communities. It's a really self-healing process.

Jeremy: Yeah. I love that. I think this idea of you approaching something from a slightly different angle, your experience, the way that you do it, the way that you see it, the way that you perceive the word or the next prompt that comes back, or how you read an error message or any of those things, you sharing your experience around that is hugely valuable to the people that are building these things. But also, you may run into problems that other people like you run into, and it's just ... Sometimes, all it takes is just a tiny twisting of the words, rearranging a sentence in a way that now that clicks with somebody where the other time it didn't. I love that.

That's why I always encourage people, just even if somebody has written his content 100 times before, whatever slight difference there is in your content, that could have a powerful effect on someone else.

Amy: Yeah, it really can.

Jeremy: Awesome. All right, let me ask you a couple of questions about Lambda and Functions as a Service because I know that you spent quite a bit of time on this stuff. I guess a question, especially, maybe even from a cloud economist, what's next for Lambda and Functions as a Service? Because I know you've written about the Lambda containers, but what's maybe that next evolution?

Amy: What AWS did recently when they released Lambda Containers is basically put it at feature parity with Azure and GCP, which already had that ability, they had either a function service or a function to Json service where you could upload your own container. They finally released the base image, where, granted, if you knew where to look, you could get it before, but they actually released it, and announced it to the general public, so you don't have to know someone in order to be able to use it.

What I see a lot of people being able to do with this now is they really want to do local development testing, so they don't have to push anything to their account and rack up those charges, when all that you want to do is make sure that whatever one line update you made, actually worked and you didn't put the space or the cab in the wrong place, which is, I guess, how it works now and it takes down the entire stack, which again, we've all done at least once, so don't worry about it. If you've ever taken down production, don't worry, you're not the only one, I promise you. You can't throw a t-shirt into an empty conference room and not hit a dude who took down production. I'm going to save that for later.

Local development testing, live simulation is a really big thing. I've seen asked to do full-on data science just on Lambda containers, so they don't have to use Kubernetes anymore, because speaking of cost stuff, it's easier to track cost-wise than Kubernetes is, because Kubernetes is purely consumption-based, and you have to tie a bunch of stuff together in order to make that tracking work. That would be great.

I think from here on, and a lot of the FaaS changes, they're not going to be front ends anymore, it's all going to be optimizations by the providers, you're not going to see much of that anymore. It's not like before, where they would add three more fields and make a blog post about it. I think everything is just going to be tuning just from Lambda's perspective now. That and hooking it to more things, because they love their integrations. What good is Lambda if you can't integrate it yourself?

Jeremy: Right, if you can't hook it up to events. It's interesting, though, this move to support containers as a packaging format. You're right, I think this has been available in IBM, it's been available in Google, it's been available in Microsoft, these capabilities have existed for a while to use a container, and again, that's a very overloaded word, I know, but to use that as a packaging format. But moving to that, the parity there with the other cloud providers is one thing, but who's that conversation for? Whose mind does that change about serverless, or FaaS, I guess.

Amy: The security team.

Jeremy: Security, okay.

Amy: Because if you talk to any engineer, if it's a technical problem, they'll find a way to fix it. That's just the way, especially at the individual contributor level, that's how the brain works is like, oh, this is a small thing, I bet I can fix it with a few days, or a weekend. Weekend turns into a month, but that's a completely different problem. I've had clients who did not want to use Lambda because they could not control the containerization system. You would be pushing your code into containers that were owned by Amazon, and the way they saw that, they saw that as liability.

While it does have some very strong technical implications, because you're now able to choose the kind of runtime you do, easier than trying to hamstring layers together, because I know layers is supposed to fix this problem, but it's so hard. It's so hard for something that you should be able to download off of Docker and then play with it and then put it back. It's so unnecessarily hard, and it makes me so angry.

If you're willing to incur that responsibility, you can tweak your memory and you have more technical control, but also you have more control at a business level too, and that is a conversation that will go way easier as far as adoption.

Jeremy: Right. The other thing, in terms of, I guess the complexity of running K8s or running Kubernetes is one of those things where that just seems like a lot of complexity. You mentioned the billing aspect of it and trying to track cost. Not that everyone's trying to narrow down exactly how much this Lambda container ran them, maybe you have more insight into that than I do, but the idea of just the complexity.

It seems to me that if you start thinking about cost, that the total cost of ownership of running a container and a Lambda function or running it in Fargate, versus having to install and maintain ... I would say, even if you're using one of the managed services like EKS, or something like that, that the total cost of ownership of going down the serverless route has got to be better.

Amy: Yeah, especially if you're one of these apps that are very user generater based. You're tracking mostly events and content, and not even a huge amount of content, you're not streaming video, you're sharing pictures, or sharing ... If you were trying to rebuild Foursquare, you would just be sharing Geo data, which is comparatively an extremely small piece of data.

You don't need an entire instance, or an entire container to do that. You can do that on a very small scale, and build that out really quickly. That said, if you go from one of these three-person teams, and then there's interest in your product, for whatever reason, and it explodes, then not just your cost, but if you had to manage the traffic of that, if you had to manage the actual resources of that, and you did not think your usage would stick with your bill, that's not great.

Being able to, at least in the first few years of the company, just use Lambda for everything, that's probably just a safer solution, because you're still rapidly iterating, and you're still changing things very quickly, and you're still transmitting very small bits of data. That said, it's like there are also large enterprise companies that are heavy Lambda users, and even their Lambda bill compared to their Kubernetes bill, it is ... If you round it to down there Kubernetes bill, you would get their Lambda bill.

Jeremy: Right. Gotcha. I think that's really interesting because I do ... I actually would love to know your thoughts and whether you even see this. I don't know if we have enough data yet to know this, but this idea of using Lambda, especially early on in startups, or even projects within an enterprise, being able to have that flexibility and the low operational overhead and so forth, I think is really great. But do you see that, or is that something that you think will happen is, you'll get to a point where you'll say we've found some sort of stability point with this product, where we now need to move it over to something like Kubernetes, or a container management system because overall, it's going to end up being cheaper in the long run.

Amy: What usually happens when you're making that transition from Lambda to either even ECS or Fargate, or eventually Kubernetes is that your business logic has now become so complex, or your infrastructure requirements have become so complex that Lambda can't do it cleanly anymore. You end up maxing out on either memory or CPU utilization, or because you're ... Apparently Lambda has a limit on how many times you can invoke it at the same time, which some people have hit in real life.

Those are times when it stops being a cheaper solution, and it stops being a target solution because you can run your own FaaS environment within instances, and then you can have a similar environment to what you're building so you don't have to rebuild everything, but you don't have to incur that on-demand cost anymore. That's one path I've seen someone take, and that's usually the decision is that Lambda, before, when it was limited, can't hold it.

Now that you can put your own container, so long as it fits in that requirement, you can pad that runway out a little bit, and you can stretch out how long you have before you do a full conversion to ECS environment. But that is usually how it is because you just try to overload or you have, maybe, 50 Lambdas trying to support one application, which is totally a thing you can do, it may not be the best ... Even with Step, even with everything else. When that becomes too complex, and you end up just going through containers, anyway.

Jeremy: Right. I think that's interesting, and I think any company that grows to the point where that they need to start thinking about that next little infrastructure, it's probably a good thing. It's a good point to start having those conversations.

All right, I got just one more question for you, because I'm really interested. You mentioned what you do as a cloud economist, reading through people's bills and things like that. Now, I thought Corey just made this thing up. I didn't even know this thing existed until, Corey comes out, and he probably coined the term. But in terms of that ...

Amy: That's what he tells people.

Jeremy: He does tell people that, right. I think he did. So, I will definitely give him credit there. But in terms of that role, of being a cloud economist and having to look through people's bills, and trying to find them ways to save it, that's pretty insane that we need people like you to do that, isn't it?

Amy: Yes, it's a bananas job. I cannot believe this is a job that I'm actually doing. It's also a lot of fun. But if you think about it, that when I was starting out, and everything was LAMP stack, when I started. That was a hot new tech when I started, was the LAMP stack. The solution to all of those problems were we're going to throw more hardware at it. Then the following question was, why are we spending so much on hardware?

Their solution to that problem was, we're going to buy real estate to store all of the hardware on. Now that you don't have to do that, you still have the problem of, I'm going to solve this problem by throwing more hardware at it. That's still a mindset that is alive and well, and you still end up with the same problem, except now you don't have the excuse that at least we own the facility that data is in because you don't anymore.

Since you don't actually own the cases and the plates and everything, you don't have to worry about disposing of them and having to use stuff that you don't actually use anymore. A lot of my problems are, one of our services has gone out of control, we don't know why. Then I will tell you, who is spending that money. I will talk to that team to make sure that they know that it's happening because sometimes they don't even know what's happening. Something got spun up into their account, and maybe it was a testbed, maybe it was a demo, maybe they hired a vendor to load something into their environment and those costs got out of control.

It's not like I'm going out trying to tell you that you did something wrong. It's like, this is where the problem is, let's go find out what happened. Forensic cloud bill person, I'm going to workshop that into a business card, because that sounds way better than the title that Corey uses.

Jeremy: Forensic cloud accountant or something like that.

Amy: Yes.

Jeremy: I think it's also interesting that billing is, and the bills you get from AWS are a leading indicator of things that are potentially going wrong. Interesting, because I don't know if people connect this. Maybe I'm underestimating people here, but the idea that a bill that runs, or that you're seeing EC2 instances cost spiking, or you're seeing a higher load or higher bandwidth or things like that. Those can all be indicators of poorly written code, it can be indicators of the bad compression or missing compression settings, all kinds of things that it can jump out at you. Unless somebody is paying attention to those bills, I don't think most developers and most teams, they're not going to see that.

Amy: Yeah. The only time they pay attention when things start spiraling out of control, and ... Okay, this sounds like an intuitive issue, and first thing people will do, will go, "We're going to log everything, and we're going to find out where the problem is."

Jeremy: It'll cost you more money.

Amy: There is a threshold where cloud watch becomes very expensive.

Jeremy: Right, absolutely.

Amy: Then they hit that threshold, and now their bill is four times as much.

Jeremy: Right.

Amy: A lot of the times it's misconfiguration, it's like, very rarely does any product get to the point where they just can't ... It's built so poorly that it can barely hold itself up. That's never been the case. It's always been, this has been turned off, or AWS also offers S3 analytics. You have to turn them on per bucket, that's not a policy that's usually written in anyone's AWS config. When they launch it, they just launch it without any analytics. They don't know if the thing is supposed to be sending things to Glacier, if it's highly used data, there's no way to tell.

It's trying to find little holes like that, where it seems like it shouldn't be a problem, but the minute it becomes a problem, it's because you spent $20,000.

Jeremy: Right. Yeah. No, you can spend money very, very fast in the cloud. I think that is a lesson learned by many, many people.

Amy: The difference between being on metal and throwing hardware at a problem and being on the cloud and throwing hardware at a problem is that you can throw hardware at a problem at scale on the cloud.

Jeremy: Exactly. Right. There's no stopping point like we have to go by using servers ...

Amy: No one will stop you.

Jeremy: No one will stop you. Just maybe the credit card company or whatever. Anyways, Amy, you are doing some amazing work with that, because I actually find that to be very, very fascinating. I think, in terms of what that can do, and the need for it, it's a fascinating field, and super interesting. Good for Corey for really digging into that and calling it out. Then again, for people like you who are willing to take that job, because that seems to me like poring through those numbers can't be the most interesting thing to do. But it must feel good when you do find a way to save somebody some money.

Amy: Spreadsheets can be interesting. Again, it's like everything else about my job. If I try to explain why it's interesting, I just make it sound more boring.

Jeremy: Awesome. All right. Well, let's leave it there. Amy, thank you again, for joining me, this was awesome. If people want to find out more about you, or maybe they have horribly large AWS cloud bills, and they want to check out the Duckbill Group, how do they do that?

Amy: Honestly, if you search for Corey Quinn, you can find the Duckbill Group real fast. If you want to go talk to me because I like doing community engagement, and I like doing talks, and I like roasting people on Twitter just about different stuff, you can hit me up on Twitter @nerdypaws. If you want to be a professional, I'm also on LinkedIn under Amy Codes.

Jeremy: All right, and then you also have a website, Amy-codes.com.

Amy: Amy-codes.com is the archive of all my talks. It's currently only showing the talks from last year because for some reason, it's somehow became very hard to find a spot for the past year. Who knew?

Jeremy: A lot of people doing talks. But anyways, all right, Amy, thank you again. Appreciate it.

Amy: Thank you. Had so much fun.

View Details

About Julian Wood

Julian Wood is a Senior Developer Advocate for the AWS Serverless Team. He loves helping developers and builders learn about, and love, how serverless technologies can transform the way they build and run applications at any scale. Julian was an infrastructure architect and manager in global enterprises and start-ups for more than 25 years before going all-in on serverless at AWS.

Twitter: @julian_wood
All things Serverless @ AWS: ServerlessLand
Serverless Patterns Collection
Serverless Office Hours – every Tuesday 10am PT
Lambda Extensions
Lambda Container Images

Watch this episode on YouTube: https://youtu.be/jtNLt3Y51-g

This episode sponsored by CBT Nuggets and Lumigo.

Transcript
Jeremy: Hi everyone, I'm Jeremy Daly and this is Serverless Chats. Today I'm joined by Julian Wood. Hey Julian, thanks for joining me.

Julian: Hey Jeremy, thank you so much for inviting me.

Jeremy: Well, I am super excited to have you here. I have been following your work for a very long time and of course, big fan of AWS. So you are a Serverless Developer Advocate at AWS, and I'd love it if you could just tell the listeners a little bit about your background, so they get to know you a bit. And then also, sort of what your role is at AWS.

Julian: Yeah, certainly. Well, I'm Julian Wood. I am based in London, but yeah, please don't let my accent fool you. I'm actually originally from South Africa, so the language purists aren't scratching their heads anymore. But yeah, I work within the Serverless Team at AWS, and hopefully do a number of things. First of all, explain what we're up to and how our sort of serverless things work and sort of, I like to sometimes say a bit cheekily, basically help the world fall in love with serverless as I have. And then also from the other side is to be a proxy and sort of be the voice of builders, and developers and whoever's building service applications, and be their voices internally. So you can also keep us on our toes to help build the things that will brighten your days.

And just before, I've worked for too many years probably, as an infrastructure racker, stacker, architect, and manager. I've worked in global enterprises babysitting their Windows and Linux servers, and running virtualization, and doing all the operations kind of stuff to support that. But, I was always thinking there's a better way to do this and we weren't doing the best for the developers and internal customers. And so when this, you know in inverted commas, "serverless way" of things started to appear, I just knew that this was going to be the future. And I could happily leave the server side to much better and cleverer people than me. So by some weird, auspicious alignment of the stars, a while later, I managed to get my current dream job talking about serverless and talking to you.

Jeremy: Yeah. Well, I tell you, I think a lot of serverless people or people who love serverless are recovering ops and infrastructure people that were doing racking and stacking. Because I too am also recovering from that and I still have nightmares.

I thought that it was interesting too, how you mentioned though, developer advocacy. It's funny, you work for a specific company, AWS obviously, but even developer advocacy in general, who is that for? Who are you advocating for? Are you advocating for the developers to use the service from the company? Are you advocating for the developers so that the company can provide the services that they actually need? Interesting balance there.

Julian: Yeah, it's true. I mean, the honest answer is we don't have great terms for this kind of role, but yeah, I think primarily we are advocating for the people who are developing the applications and on the outside. And to advocate for them means we've got to build the right stuff for them and get their voices internally. And there are many ways of doing that. Some people raise support requests and other kind of things, but I mean, sometimes some of our great ideas come from trolling Twitter, or yes, I know even Hacker News or that kind of thing. But also, we may get responses from 10 different people about something and that will formulate something in our brain and we'll chat with other kind of people. And that sort of starts a thing. It's not just necessarily each time, some good idea in Twitter comes in, it gets mashed into some big surface database that we all pick off.

But part of our job is to be out there and try and think and be developers in whatever backgrounds we come from. And I mean, I'm not a pure software developer where I've come from, and I come, I suppose, from infrastructure, but maybe you'd call that a bit of systems engineering. So yeah, I try and bring that background to try and give input on whatever we do, hopefully, the right stuff.

Jeremy: Right. Yeah. And then I think part of the job too, is just getting the information out there and getting the examples out there. And trying to create those best practices or at least surface those best practices, and encourage the community to do a lot of that work and to follow that. And you've done a lot of work with that, obviously, writing for the AWS blog. I know you have a series on the Serverless Lens and the Well-Architected Framework, and we can talk about that in a little while. But I really want to talk to you about, I guess, just the expansion of serverless over the last couple of years.

I mean, it was very narrowly focused, probably, when it first came out. Lambda was ... FaaS as a whole new concept for a lot of people. And then as this progressed and we've gotten more APIs, and more services and things that it can integrate with, it just becomes complex and complicated. And that's a good thing, but also maybe a bad thing. But one of the things that AWS has done, and I think this is clearly in reaction to the developers needing it, is the ability to extend what you can do with a Lambda function, right? I mean, the idea of just putting your code in there and then, boom, that's it, that's all you have to do. That's great. But what if you do need access to lifecycle hooks? Or what if you do want to manipulate the underlying runtime or something like that? And AWS, I think has done a great job with that.

So maybe we can start there. So just about the extensibility of Lambda in general. And one of the new things that was launched recently was, and recently, I don't know what was it? Seven months ago at this point? I'm not even sure. But was launched fairly recently, let's say that, is Lambda Extensions, and a couple of different flavors of that as well. Could you kind of just give the users an over, the users, wow, the listeners an overview of what Lambda Extensions are?

Julian: I could hear the ops background coming in, talking about our users. Yeah. But I mean, from the get-go, serverless was always a terrible term because, why on earth would you name something for what it isn't? I mean, you know? I remember talking to DBAs, talking about noSQL, and they go, "Well, if it's not SQL, then what is it?" So we're terrible at that, serverless as well. And yeah, Lambda was very constrained when it came out. Lambda was never built being a serverless thing, that's what was the outcome. Sometimes we focus too much on the tools rather than the outcome. And the story is S3, just turning 15. And the genesis of Lambda was being an event trigger for S3, and people thought you'd upload something to S3, fire off a Lambda function, how cool is that? And then obviously the clever clubs at the time were like, "Well, hang on, let's not just do this for S3, let's do this for a whole bunch of kind of things."

So Lambda was born out of that, as that got that great history, which is created an arc sort of into the present and into the future, which I know we're also going to get on about, the power of event driven applications. But the power of Lambda has always been its simplicity, and removing that operational burden, and that heavy lifting. But, sometimes that line is a bit of a gray area and there're people who can be purists about serverless and can be purists about FaaS and say, "Everything needs to be ephemeral. Lambda functions can't extend to anything else. There shouldn't be any state, shouldn't be any storage, shouldn't be any ..." All this kind of thing.

And I think both of us can agree, but I don't want to speak for you, but I think both of us would agree that in some sense, yeah, that's fine. But we live in the real world and there's other stuff that needs to connect to and we're not here about building purist kind of stuff. So Lambda Extensions is a new way basically to integrate Lambda with your favorite tools. And that's the sort of headline thing we like to talk about. And the big idea is to open up Lambda to more effectively work mainly with partners, but also your own tools if you want to write them. And to sort of have deeper hooks into the Lambda lifecycle.

And other partners are awesome and they do a whole bunch of stuff for serverless, plus customers also have connections to on-prem staff, or EC2 staff, or containers, or all kind of things. How can we make the tools more seamless in a way? How can we have a common set of tools maybe that you even use on-prem or in the cloud or containers or whatever? Why does Lambda have to be unique or different or that kind of thing? And Extensions is sort of one of the starts of that, is to be able to use these kind of tools and get more out of Lambda. So I mean, just the kind of tools that we've already got on board, there's things like Splunk and AppDynamics. And Lumigo, Epsagon, HashiCorp, Honeycomb, CoreLogic, Dynatrace, I can't think. Thundra and Sumo Logic, Check Point. Yeah, I'm sorry. Sorry for any partners who I've forgotten a few.

Jeremy: No, right, no. That's very good. Shout them out, shout them out. No, I mean just, and not to interrupt you here, but ...

Julian: No, please.

Jeremy: ... I think that's great. I mean, I think that's one of the things that I like about the way that AWS deals with partners, is that ... I mean, I think AWS knows they can't solve all these problems on their own. I mean, maybe they could, right? But they would be their own way of solving the problems and there's other people who are solving these problems differently and giving you the ability to extend your Lambda functions into those partners is, there's a huge win for not only the partners because it creates that ecosystem for them, but also for AWS because it makes the product itself more valuable.

Julian: Well, never mind the big win for customers because ultimately they're the one who then gets a common deployment tool, or a common observability tool, or a HashiCorp Vault that you can manage secrets and a Lambda function from HashiCorp Vault. I mean, that's super cool. I mean, also AWS services are picking this up because that's easy for them to do stuff. So if anybody's used Lambda Insights or even seen Lambda Insights in the console, it's somewhere in the monitoring thing, and you just click something over and you get this tool which can pull stuff that you can't normally get from a Lambda function. So things like CPU time and network throughput, which you couldn't normally get. But actually, under the hoods, Lambda Insights is using Lambda extensions. And you can see that if you look. It automatically adds the Lambda layer and job done.

So anyway, this is how a lot of the tools work, that a layer is just added to a Lambda function and off you go, the tool can do its work. So also there's a very much a simplicity angle on this, that in a lot of cases you don't have to do anything. You configure some of the extensions via environment variables, if that's cooled you may just have an API key or a log retention value or something like that, I don't know, any kind of example of that. But you just configure that as a normal Lambda environment variable at this partner extension, which is just a Lambda layer, and off you go. Super simple.

Jeremy: Right. So explain Extensions exactly, because I think that's one of those things because now we have Lambda layers and we have Lambda Extensions. And there's also like the runtime API and then something else. I mean, even I'm not 100% sure what all of the naming conventions. I'm pretty sure I know what they do ...

Julian: Yeah, fair enough.

Jeremy: ... but maybe we could say the names and say exactly what they do as well.

Julian: Yeah, cool. You get an API, I get an API, everybody gets an API. So Lambda layers, let's just start, because that's, although it's not related to Extensions, it's how Extensions are delivered to the power core functions. And Lambda layers is just another way to add code to a Lambda function or not even code, it can be a dependency. It's just a way that you could, and it's cool because they are shareable. So you have some dependencies, or you have a library, or an SDK, or some training data for something, a Lambda layer just allows you to add some bits and bobs to your Lambda function. That's a horrible explanation. There's another word I was thinking of, because I don't want to use the word code, because it's not necessarily code, but it's dependency, whatever. It's just another way of adding something. I'll wake up in a cold sweat tonight thinking of the word I was thinking of, but anyway.

But Lambda Extensions introduces a whole new companion API. So the runtime API is the little bit of code that allows your function to talk to the Lambda service. So when an event comes in, this is from the outside. This could be via API gateway or via the Lambda API, or where else, EventBridge or Step Functions or wherever. When you then transports that data rise in the Lambda services and HTTP call, and Lambda transposes that into an event and sends that onto the Lambda function. And it's that API that manages that. And just as a sidebar, what I find it cool on a sort of geeky, technical thing is, that actually API sits within the execution environment. People are like, "Oh, that's weird. Why would your Lambda API sit within the execution environment basically within the bubble that contains your function rather than it on the Lambda service?"

And the cool answer for that is it's actually for a security mechanism. Like your function can then only ever talk to the Lambda runtime API, which is in that secure execution environment. And so our security can be a lot stronger because we know that no function code can ever talk directly out of your function into the Lambda service, it's all got to talk locally. And then the Lambda service gets that response from the runtime API and sends it back to the caller or whatever. Anyway, sidebar, thought that was nerdy and interesting. So what we've now done is we've released a new Extensions API. So the Extensions API is another API that an extension can use to get information from Lambda. And they're two different types of extensions, just briefly, internal and external extensions.

Now, internal extensions run within the runtime process so that it's just basically another thread. So you can use this for Python or Java or something and say, when the Python runtime starts, let's start it with another parameter and also run this Java file that may do some observability, or logging, or tracing, or finding out how long the modules take to launch, for example. I know there's an example for Python. So that's one way of doing extensions. So it's internal extensions, they're two different flavors, but I'll send you a link. I'll provide a link to the blog posts before we go too far down the rabbit hole on that.

And then the other part of extensions are external extensions. And this is a cool part because they actually run as completely separate processes, but still within that secure bubble, that secure execution environment that Lambda runs it. And this gives you some superpowers if you want. Because first of all, an extension can run in any language because it's a separate process. So if you've got a Node function, you could run an extension in other kind of languages. Now, what do we do recommend is you do run your extension in a compiled binary, just because you've got to provide the runtime that the extensions got to run in any way, so as a compiled binary, it's super easy and super useful. So is something like Go, a lot of people are doing because you write a single extension and Go, and then you can use it on your Node functions, your Java functions, your PowerShell functions, whatever. So that's a really good, simple way that you can have the portability.

But now, what can these extensions do? Well, the extensions basically register with extensions API, and then they say to Lambda, "Lambda, I want to know about what happens when my functions invoke?" So the extension can start up, maybe it's got some initialization code, maybe it needs to connect to a database, or log into an observability platform, or pull down a secret order. That it can do, it's got its own init that can happen. And then it's basically ready to go before the function invokes. And then when the extension then registers and says, "I want to know when the function invokes and when it shuts down. Cool." And that's just something that registers with the API. Then what happens is, when a functioning invoke comes in, it tells the runtime API, "Hello, you now have an event," sends it off to the Lambda function, which the runtime manages, but also extension or extensions, multiple ones, hears information about that event. And so it can tell you the time it's going to run and has some metadata about that event. So it doesn't have the actual event data itself, but it's like the sort of Lambda context, a version of that that it's going to send to the extension.

So the extension can use that to do various things. It can start collecting telemetry data. It can alter instrument some of your code. It could be managing a secret as a separate process that it is going to cache in the background. For example, we've got one with AppConfig, which is really cool. AppConfig is a service where you manage parameters external to your Lambda function. Well, each time your Lambda function warm invokes if you've got to do an external API call to retrieve that, well, it's going to be a little bit efficient. First of all, you're going to pay for it and it's going to take some time.

So how about when the Lambda function runs and the extension could run before the Lambda function, why don't we just cache that locally? And then when your Lambda function runs, it just makes a local HTTP call to the extension to retrieve that value, which is going to be super quick. And some extensions are super clever because they're their own process. They will go, "Well, my value is for 30 minutes and every 30 minutes if I haven't been run, I will then update the value from that." So that's useful. Extensions can then also, when the runtime ... Sorry, let me back up.

When the runtime is finished, it sends its response back to the runtime API, and extensions when they're done doing, so the runtime can send it back and the extension can carry on processing saying, "Oh, I've got the information about this. I know that this Lambda function has done X, Y, Z, so let me do, do some telemetry. Let me maybe, if I'm writing logs, I could write a log to S3 or to Kinesis or whatever. Do some kind of thing after the actual function invocation has happened." And then when it says it's ready, it says, "Hello, extensions API, I'm telling you I'm done." And then it's gone. And then Lambda freezes the execution environment, including the runtime and the extensions until another invocation happens. And the cycle then will happen.

And then the last little bit that happens is, instead of an invoke coming in, we've extended the Lambda life cycles, so when the environment is going to be shut down, the extension can receive the shutdown and actually do some stuff and say, "Okay, well, I was connected to my observer HTTP platform, so let me close that connection. I've got some extra logs to flush out. I've got whatever else I need to do," and just be able to cleanly shut down that extra process that is running in parallel to the Lambda function.

Jeremy: All right.

Julian: So that was a lot of words.

Jeremy: That was a lot and I bet you that would be great conversation for a dinner party. Really kicks things up. Now, the good news is that, first of all, thank you for that though. I mean, that's super technical and super in-depth. And for anyone listening who ...

Julian: You did ask, I did warn you.

Jeremy ... kind of lost their way ... Yes, but something that is really important to remember is that you likely don't have to write these yourself, right? There is all those companies you mentioned earlier, all those partners, they've already done this work. They've already figured this out and they're providing you access to their tools via this, that allows you to build things.

Julian: Exactly.

Jeremy: So if you want to build an extension and you want to integrate your product with Lambda or so forth, then maybe go back and listen to this at half speed. But for those of you who just want to take advantage of it because of the great functionality, a lot of these companies have already done that for you.

Julian: Correct. And that's the sort of easiness thing, of just adding the Lambda layer or including in a container image. And yeah, you don't have to worry any about that, but behind the scenes, there's some really cool functionality that we're literally opening up our Lambda operates and allowing you to impact when a function responds.

Jeremy: All right. All right. So let me ask another, maybe an overly technical question. I have heard, and I haven't experienced this, but that when it runs the life cycle that ends the Lambda function, I've heard something like it doesn't send the information right away, right? You have to wait for that Lambda to expire or something like that?

Julian: Well, yes, for now, about to change. So currently Extensions is actually in preview. And that's not because it's in Beta or anything like that, but it's because we spoke to the partners and we didn't want to dump Extensions on the world. And all the partners had to come out with their extensions on day one and then try and figure out how customers are going to use them and everything. So what we really did, which I think in this case works out really well, is we worked with the partners and said, "Well, let's release this in preview mode and then give everybody a whole bunch of months to work out what's the best use cases, how can we best use this?" And some partners have said, "Oh, amazing. We're ready to go." And some partners have said, "Ah, it wasn't quite what we thought. Maybe we're going to wait a bit, or we're going to do something differently, or we've got some cool ideas, just give us time." And so that's what this time has been.

The one other thing that has happened is we've actually added some performance enhancements during it. So yes, currently during the preview, the runtime and all extensions need to finish before we give you your response back to your Lambda function. So if you're in an asynchronous mode, you don't really care, but obviously if you're in a synchronous mode behind an API, yeah, you don't really want that. But when Extensions goes GA, which isn't going to be long, then that is no longer the case. So basically what'll happen is the runtime will respond and the result goes directly back to whoever's calling that, maybe API gateway, and the extensions can carry on, partly asynchronously in the background.

Jeremy: Yep. Awesome. All right. And I know that the plan is to go GA soon. I'm not sure when around when this episode comes out, that that will be, but soon, so that's good to know that that is ...

Julian: And in fact, when we go GA that performance enhancement is part of the GA. So when it goes GA, then you know, it's not something else you need to wait for.

Jeremy: Perfect. Okay. All right. So let's move on to another bit of, I don't know if this is extensibility of the actual product itself or more so I think extensibility of maybe the workflow that you use to deploy to Lambda and deploy your serverless applications, and that's container image support. I mean, we've discussed it a lot. I think people kind of have an idea, but just give me your quick overview of what that is to set some context here.

Julian: Yeah, sure. Well, container image support in a simple sort of headline thing is to be able to build and package your functions as a container image. So you basically build a function using a Docker file. So before if you use a zip function, but a lot of people use Serverless Framework or SAM, or whatever, that's all abstracted away from you, but it's actually creating a zip file and uploading it to Lambda or S3. So with container image support, you use a Docker file to build your Lambda function. That's the headline of what's happening.

Jeremy: Right. And so the idea of creating, and this is also, and again, you mentioned packaging, right? I mean, that is the big thing here. This is a packaging format. You're not actually running the container in a Lambda function.

Julian: Correct. Yeah, let's maybe think, because I mean, "containers," in inverted commas again for people who are on the audio, is ...

Jeremy: What does it even mean?

Julian: Yeah, exactly. And can be quite an overload of terms and definitely causes some confusion. And I sort of think maybe there's sort of four things that are in the container world. One, containers is an isolation mechanism. So on Linux, this is UNC Group, seccomp, other bits and pieces that can be used to isolate processes or maybe groups of processes. And then a second one, containers as the packaging mechanism. This is what Docker really popularized and this is about taking some code and the dependencies needed to run the code, and then packaging them all out together, maybe with some metadata to describe it.

And then, three is containers as also a design philosophy. This is the idea, if we can package and isolate software, it's easier to run. Maybe smaller pieces of software is easy to reason about and manage independently. So I don't want to necessarily use microservices, but there's some component of that with it. And the emphasis here is on software rather than services, and standardized tooling to simplify your ops. And then the fourth thing is containers as an ecosystem. This is where all the products, tools, know how, all the actual things to how to do containers. And I mean, these are certain useful, but I wouldn't say there're anything about the other kind of things.

What is cool and worth appreciating is how maybe independent these things are. So when I spoke about containers as isolation, well, we could actually replace containers as isolation with micro VMs such as we do with Firecracker, and there's no real change in the operational properties. So one, if we think, what are we doing with containers and why? One of those is in a way ticked off with Lambda. Lambda does have secure isolation. And containers as a packaging format. I mean, you could replace it with static linking, then maybe won't really be a change, but there's less convenience. And the design philosophy, that could really be applicable if we're talking microservices, you can have instances and certainly functions, but containers are all the same kind of thing.

So if we talk about the packaging of Lambda functions, it's really for people who are more familiar with containers, why does Lambda have to be different? You've got, why does Lambda to have to be a snowflake in a way that you have to manage differently? And if you are packaging dependencies, and you're doing npm or pip install, and you're used to building Docker files, well, why can't we just do that for Lambda at the same things? And we've got some other things that come with that, larger function sizes, up to 10 gig, which is enabled with some of this technology. So it's a packaging format, but on the backend, there's a whole bunch of different stuff, which has to be done to to allow this. Benefits are, use your tooling. You've got your CI/CD pipelines already for containers, well, you can use that.

Jeremy: Yeah, yeah. And I actually like that idea too. And when I first heard of it, I was like, I have nothing against containers, the containers are great. But when I was thinking about it, I'm like, "Wait container? No, what's happening here? We're losing something." But I will say, like when Lambda layers came out, which was I think maybe 2019 or something like that, maybe 2018, the idea of it made a lot of sense, being able to kind of supplement, add additional dependencies or code or whatever. But it always just seemed awkward. And some of the publishing for it was a little bit awkward. The versioning used like a numbered versioning instead of like semantic versioning and things like that. And then you had to share it to multiple places and if you published it as a SAR app, then you got global distri ... Anyways, it was a little bit hard to use.

And so when you're trying to package large dependencies and put those in a layer and then combine them with a Lambda function, the other problem you had was you still had a maximum size that you could use for those, when those were combined. So I like this idea of saying like, "Look, I'd like to just kind of create this little isolate," like you said, "put my dependencies in there." Whether that's PyCharm or some other thing that is a big dependency that maybe I don't want to install, directly in a Lambda layer, or I don't want to do directly in my Lambda function. But you do that together and then that whole process just is a lot easier. And then you can actually run those containers, you could run those locally and test those if you wanted to.

Julian: Correct. So that's also one of the sort of superpowers of this. And that's when I was talking about, just being able to package them up. Well, that now enables a whole bunch of extra kind of stuff. So yes, first of all is you can then use those container images that you've created as your local testing. And I know, it's silly for anyone to poo poo local testing. And we do like to say, "Well, bring your testing to the cloud rather than bringing the cloud to your testing." But testing locally for unit tests is super great. It's going to be super fast. You can iterate, have your Lambda functions, but we don't want to be mocking all of DynamoDB, all of building harebrained S3 options locally.

But the cool thing is you've got the same Docker file that you're going to run in Lambda can be the same Docker file to build your function that you run locally. And it is literally exactly the same Lambda function that's going to run. And yes, that may be locally, but, with a bit of a stretch of kind of stuff, you could also run those Lambda functions elsewhere. So even if you need to run it on EC2 instances or ECS or Fargate or some kind of thing, this gives you a lot more opportunities to be able to use the same Lambda function, maybe in different way, shapes or forms, even if is on-prem. Now, obviously you can't recreate all of Lambda because that's connected to IM and it's got huge availability, and scalability, and latency and all that kind of things, but you can actually run a Lambda function in a lot more places.

Jeremy: Yeah. Which is interesting. And then the other thing I had mentioned earlier was the size. So now the size of these container or these packages can be much, much bigger.

Julian: Yeah, up to 10 gig. So the serverless purists in the back are shouting, "What about cold starts? What about cold starts?"

Jeremy: That was my next question, yes.

Julian: Yeah. I mean, back on zip functional archives are also all available, nothing changes with that Lambda layers, many people use and love, that's all available. This isn't a replacement it's just a new way of doing it. So now we've got Lambda functions that can be up to 10 gig in size and surely, surely that's got to be insane for cold starts. But actually, part of what I was talking about earlier of some of the work we've done on the backend to support this is to be able to support these super large package sizes. And the high level thing is that we actually cache those things really close to where the Lambda layer is going to be run.

Now, if you run the Docker ecosystem, you build your Docker files based on base images, and so this needs to be Linux. One of the super things with the container image support is you don't have to use Amazon Linux or Amazon Linux 2 for Lambda functions, you can actually now build your Lambda functions also on Ubuntu, DBN or Alpine or whatever else. And so that also gives you a lot more functionality and flexibility. You can use the same Linux distribution, maybe across your entire suite, be it on-prem or anywhere else.

Jeremy: Right. Right.

Julian: And the two little components, there's an interface client, what you install, it's just another Docker layer. And that's that runtime API shim that talks to the runtime API. And then there's a runtime interface emulator and that's the thing that pretends to be Lambda, so you can shunt those events between HTTP and JSON. And that's the thing you would use to run locally. So runtime interface client means you can use any Linux distribution at the runtime interface client and you're compatible with Lambda, and then the interface emulators, what you would use for local testing, or if you want to spread your wings and run your Lambda functions elsewhere.

Jeremy: Right. Awesome. Okay. So the other thing I think that container support does, I think it opens up a broader set of, or I guess a larger audience of people who are familiar with containerization and how that works, bringing those two Lambda functions. And one of the things that you really don't get when you run a container, I guess, on EC2, or, not EC2, I'm sorry, ECS, or Fargate or something like that, without kind of adding another layer on top of it, is the eventing aspect of it. I mean, Lambda just is naturally an event driven, a compute layer, right? And so, eventing and this idea of event driven applications and so forth has just become much more popular and I think much more mainstream. So what are your thoughts? What are you seeing in terms of, especially working with so many customers and businesses that are using this now, how are you seeing this sort of evolution or adoption of event driven applications?

Julian: Yeah. I mean, it's quite funny to think that actually the event of an application was the genesis of Lambda rather than it being Serverless. I mentioned earlier about starting with S3. Yeah, the whole crux of Lambda has been, I respond to an event of an API gateway, or something on SQS, or via the API or anything. And so the whole point in a way of Lambda has been this event driven computing, which I think people are starting to sort of understand in a bigger thing than, "Oh, this is just the way you have to do Lambda." Because, I do think that serverless has a unique challenge where there is a new conceptual learning maybe that you have to go through. And one other thing that holds back service development is, people are used to a client's server and maybe ports and sockets. And even if you're doing containers or on-prem, or EC2, you're talking IP addresses and load balances, and sockets and firewalls, and all this kind of thing.

But ultimately, when we're building these applications that are going to be composed of multiple services talking together through using APIs and events, the events is actually going to be a super part of it. And I know he is, not for so much longer, but my ultimate boss, but I can blame Jeff Bezos just a little bit, because he did say that, "If you want to talk via anything, talk via an API." And he was 100% right and that was great. But now we're sort of evolving that it doesn't just have to be an API and it doesn't have to be something behind API gateway or some API that you can run. And you can use the sort of power of events, particularly in an asynchronous model to not just be "forced" again in inverted commas to use APIs, but have far more flexibility of how data and information is going to flow through, maybe not just your application, but your suite of applications, or to and from your partners, or where that is.

And ultimately authentications are going to be distributed, and maybe that is connecting to partners, that could be SaaS partners, or it's going to be an on-prem component, or maybe things in other kind of places. And those things need to communicate. And so the way of thinking about events is a super powerful way of thinking about that.

Jeremy: Right. And it's not necessarily new. I mean, we've been doing web hooks for quite some time. And that idea of, something is going to happen somewhere and I want to be notified of it, is again, not a new concept. But I think certainly the way that it's evolved with Lambda and the way that other FaaS products had done eventing and things like that, is just those tight integrations and just all of the, I guess, the connective tissue that runs between those things to make sure that the events get delivered, and that you can DLQ them, and you can do all these other things with retries and stuff like that, is pretty powerful.

I know you have, I actually just mentioned this on the last episode, about one of my favorite books, I think that changed my thinking and really got me thinking about how microservices communicate with one another. And that was Building Microservices by Sam Newman, which I actually said was sort of like my Bible for a couple of years, yes, I use that. So what are some of the other, like I know you have a favorite book on this.

Julian: Well, that Building Microservices, Sam Newman, and I think there's a part two. I think it's part two, or there's another one ...

Jeremy: Hopefully.

Julian: ... in the works. I think even on O'Riley's website, you can go and see some preview copies of it. I actually haven't seen that. But yeah, I mean that is a great kind of Bible talking. And sometimes we do conflate this microservices things with a whole bunch of stuff, but if you are talking events, you're talking about separating things. But yeah, the book recommendation I have is one called Flow Architectures by James Urquhart. And James Urquhart actually works with VMware, but he's written this book which is looking sort of at the current state and also looking into the future about how does information flow through our applications and between companies and all this kind of thing.

And he goes into some of the technology. When we talk about flow, we are talking about streams and we're talking about events. So streams would be, let's maybe put some AWS words around it, so streams would be something like Kinesis and events would be something like EventBridge, and topics would be SNS, and SQS would be queues. And I know we've got all these things and I wish some clever person would create the one flow service to rule them all, but we're not there. And they've got also different properties, which are helpful for different things and I know confusingly some of them merge. But James' sort of big idea is, in the future we are going to be able to moving data around between businesses, between applications. So how can we think of that as a flow? And what does that mean for designing applications and how we handle that?

And Lambda is part of it, but even more nicely, I think is even some of the native integrations where you don't have to have a Lambda function. So if you've got API gateway talking to Step Functions directly, for example, well, that's even better. I mean, you don't have any code to manage and if it's certainly any code that I've written, you probably don't want to manage it. So yeah. I mean this idea of flow, Lambda's great for doing some of this moving around. But we are even evolving to be able to flow data around our applications without having to do anything and just wire up some things in a console or in a terminal.

Jeremy: Right. Well, so you mentioned, someone could build the ultimate sort of flow control system or whatever. I mean, I honestly think EventBridge is very close to that. And I actually had Mike Deck on the show. I think it was like episode five. So two years ago, whenever it was when the show came out. I mean, when EventBridge came out. And we were talking and I sort of made the joke, I'm like, so this is like serverless web hooks, essentially being able, because there was the partner integrations where partners could push events onto an event bus, which they still can do. But this has evolved, right? Because the issue was always sort of like, I would have to subscribe to web books, I'd have to build a web hook to get events from a particular company. Which was great, always worked fine, but you're still maintaining that infrastructure.

So EventBridge comes along, it creates these partner integrations and now you can just push an event on that now your applications, whether it's a Lambda function or other services, you can push them to an SQS queue, you can push them into a Kinesis stream, all these different destinations. You can go ahead and pull that data in and that's just there. So you don't have to worry about maintaining that infrastructure. And then, the EventBridge team went ahead and released the destination API, I think it's called.

Julian: Yeah, API destinations.

Jeremy: Event API destinations, right, where now you can set up these integrations with other companies, so you don't even have to make the API call yourself anymore, but instead you get all of the retries, you get the throttling, you get all that stuff kind of built in. So I mean, it's just really, really interesting where this is going. And actually, I mean, if you want to take a second to tell people about EventBridge API destinations, what that can do, because I think that now sort of creates both sides of that equation for you.

Julian: It does. And I was just thinking over there, you've done a 10 times better job at explaining API destinations than I have, so you've nailed it on the head. And packet is that kind of simple. And it is just, events land up in your EventBridge and you can just pump events to any arbitrary endpoint. So it doesn't have to be in AWS, it can be on-prem. It can be to your Raspberry PI, it can literally be anywhere. But it's not just about pumping the events over there because, okay, how do we handle failover? And how do we handle over throttling? And so this is part of the extra cool goodies that came with API destinations, is that you can, for instance, if you are sending events to some external API and you only licensed for 1,000 invocations, not invocations, that could be too Lambda-ish, but 1,000 hits on the API every minute.

Jeremy: Quotas. I think we call them quotas.

Julian: Quotas, something like that. That's a much better term. Thank you, Jeremy. And some sort of quota, well, you can just apply that in API destinations and it'll basically store the data in the meantime in EventBridge and fire that off to the API destination. If the API destination is in that sort of throttle and if the API destination is down, well, it's going to be able to do some exponential back-off or calm down a little bit, don't over-flood this external API. And then eventually when the API does come back, it will be able to send those events. So that does just really give you excellent power rather than maintaining all these individual API endpoints yourself, and you're not handling the availability of the endpoint API, but of whatever your code is that needs to talk to that destination.

Jeremy: Right. And I don't want to oversell this to anybody, but that also ...

Julian: No, keep going. Keep going.

Jeremy: ... adds the capability of enhanced security, because you're not exposing those API keys to your developers or anybody else, they're all baked in and stored within, the API destinations or within an EventBridge. You have the ability, you mentioned this idea of not needing Lambda to maybe talk directly, API gateway to DynamoDB or to step function or something like that. I mean, the cool thing about this is you do have translation capabilities, or transformation capabilities in EventBridge where you can transform the event. I haven't tried this, but I'm assuming it's possible to say, get an event from Salesforce and then pipe it into Stripe or some other API that you might want to pipe it into.

So I mean, just that idea of having that centralized bus that can communicate with all these different things. I mean, we're talking about distributed systems here, right? So why is it different sending an event from my microservice A to my microservice B? Why can't I send it from my microservice A to company-wise, microservice B or whatever? And being able to do that in a secure, reliable, just with all of that stuff kind of built in for you, I think it's amazing. So I love EventBridge. To me EventBridge is one of those services that rivals Lambda. It's as, I guess as important as Lambda is, in this whole serverless equation.

Julian: Absolutely, Jeremy. I mean, I'm just sitting here. I don't actually have to say anything. This is a brilliant interview and Jeremy, you're the expert. And you're just like laying down all of the excellent use cases. And exactly it. I mean, I like to think we've got sort of three interlinked services which do three different things, but are awesome. Lambda, we love if you need to do some processing or you need to do something that's literally your business logic. You've got EventBridge that can route data from in and out of SaaS partners to any other kind of API. And then you've got Step Functions that can do some coordination. And they all work together, but you've got three different things that really have sort of superpowers in terms of the amount of stuff you can do with it. And yes, start with them. If you land up bumping up against any kind of things that it doesn't work, well, first of all, get in touch with me, I'll work on that.

But then you can maybe start thinking about, is it containers or EC2, or that kind of thing? But using literally just Lambda, Step Functions and EventBridge, okay. Yes, maybe you're going to need some queues, topics and APIs, and that kind of thing. But ...

Jeremy: I was just going to say, add DynamoDB in there for some permanent state or for some data persistence. Right? Yeah. But other than that, no, I think you nailed it. Honestly, sometimes you're starting to build applications and yeah, you're right. You maybe need a queue here and there and things like that. But for the most part, no, I mean, you could build a lot with those three or four services.

Julian: Yeah. Well, I mean, even think of it what you used to do before with API destinations. Maybe you drop something on a queue, you'd have Lambda pull that from a queue. You have Lambda concurrency, which would be set to five per second to then send that to an external API. If it failed going to that API, well, you've got to then dump it to Lambda destinations or to another SQS queue. You then got something ... You know, I'm going down the rabbit hole, or just put it on EventBridge ...

Jeremy: You just have it magically happen.

Julian: ... or we talk about removing serverless infrastructure, not normal infrastructure, and just removing even the serverless bits, which is great.

Jeremy: Yeah, no. I think that's amazing. So we talked about a couple of these different services, and we talked about packaging formats and we talked about event driven applications, and all these other things. And a lot of this stuff, even though some of it may be familiar and you could probably equate it or relate it to things that developers might already know, there is still a lot of new stuff here. And I think, my biggest complaint about serverless was not about the capabilities of it, it was basically the education and the ability to get people to adopt it and understand the power behind it. So let's talk about that a little bit because ... What's that?

Julian: It sounds like my job description, perfectly.

Jeremy: Right. So there we go. Right, that's what you're supposed to be doing, Julian. Why aren't you doing it? No, but you are doing it. You are doing it. No, and that's why I want to talk to you about it. So you have that series on the Well-Architected Framework and we can talk about that. There's a whole bunch of really good resources on this. Obviously, you're doing videos and conferences, well, you used to be doing conferences. I think you probably still do some of those virtual ones, right? Which are not the same thing.

Julian: Not quite, no.

Jeremy: I mean, it was fun seeing you in Cardiff and where else were you?

Julian: Yeah, Belfast.

Jeremy: Cardiff and Northern Ireland.

Julian: Yeah, exactly.

Jeremy: Yeah, we were all over the place together.

Julian: With the Guinness and all of us. It was brilliant.

Jeremy: Right. So tell me a little bit about, sort of, the education process that you're trying to do. Or maybe even where you sort of see the state of Serverless education now, and just sort of where it's evolved, where we're getting best practices from, what's out there for people. And that's a really long question, but I don't know, maybe you can distill that down to something usable.

Julian: No, that's quite right. I'm thinking back to my extensions explanation, which is a really long answer. So we're doing really long stuff, but that's fine. But I like to also bring this back to also thinking about the people aspect of IT. And we talk a lot about the technology and Lambda is amazing and S3 is amazing and all those kinds of things. But ultimately it is still sort of people lashing together these services and building the serverless applications, and deciding what you even need to do. And so the education is very much tied with, of course, having the products and features that do lots of kinds of things. And Serverless, there's always this lever, I suppose, between simplicity and functionality. And we are adding lots of knobs and levers and everything to Lambda to make it more feature-rich, but we've got to try and keep it simple at the same time.

So there is sort of that trade-off, and of course with that, that obviously means not just the education side, but education about Lambda and serverless, but generally, how do I build applications? What do I do? And so you did mention the Well-Architected Framework. And so for people who don't know, this came out in 2015, and in 2017, there was a Serverless Lens which was added to it; what is basically serverless specific information for Well-Architected. And Well-Architected means bringing best practices to serverless applications. If you're building prod applications in the cloud, you're normally looking to build and operate them following best practices. And this is useful stuff throughout the software life cycle, it's not just at the end to tick a few boxes and go, "Yes, we've done that." So start early with the well-architected journey, it'll help you.

And just sort of answer the question, am I well architected? And I mean, that is a bit of a fuzzy, what is that question? But the idea is to give you more confidence in the architecture and operations of your workloads, and that's not a goal it's in, but it's to reduce and minimize the impact of any issues that can happen. So what we do is we try and distill some of our questions and thoughts on how you could do things, and we built that into the Well-Architected Framework. And so the ServiceLens has a few questions on its operational excellence, security, reliability, performance, efficiency, and cost optimization. Excellent. I knew I was going to forget one of them and I didn't. So yeah, these are things like, how do you control access to an API? How do you do lifecycle management? How do you build resiliency into your application? All these kinds of things.

And so the Well-Architected Framework with Serverless Lens there's a whole bunch of guidance to help you do that. And I have been slowly writing a blog series to literally cover all of the questions, they're nine questions in the Well-Architected Serverless Lens. And I'm about halfway through, and I had to pause because we have this little conference called re:Invent, which requires one or two slides to be created. But yeah, I'm desperately keen to pick that up again. And yeah, that's just providing some really and sort of more opinionated stuff, because the documentation is awesome and it's very in-depth and it's great when you need all that kind of stuff. But sometimes you want to know, well, okay, just tell me what to do or what do you think is best rather than these are the seven different options.

Jeremy: Just tell me what to do.

Julian: Yeah.

Jeremy: I think that's a common question.

Julian: Exactly. And I'll launch off from that to mention my colleague, James Beswick, he writes one or two things on serverless ...

Jeremy: Yeah, I mean, every once in a while you see something from it. Yeah.

Julian: ... every day. The Besbot machine of serverless. He's amazing. James, he's so knowledgeable and writes like a machine. He's brilliant. Yeah, I'm lucky to be on his team. So when you talk about education, I learn from him. But anyway, in a roundabout way, he's created this blog series and other series called the Lambda Operations Guide. And this is literally a whole in-depth study on how to operate Lambda. And it goes into a whole bunch of things, it's sort of linked to the Serverless Lens because there are a lot of common kind of stuff, but it's also a great read if you are more nerdily interested in Lambda than just firing off a function, just to read through it. It's written in an accessible way. And it has got a whole bunch of information on how to operate Lambda and some of the stuff under the scenes, how to work, just so you can understand it better.

Jeremy: Right. Right. Yeah. And I think you mentioned this idea of confidence too. And I can tell you right now I've been writing serverless applications, well, let's see, what year is it? 2021. So I started in 2015, writing or building applications with Lambda. So I've been doing this for a while and I still get to a point every once in a while, where I'm trying to put something in cloud formation or I'm using the Serverless Framework or whatever, and you're trying to configure something and you think about, well, wait, how do I want to do this? Or is this the right way to do it? And you just have that moment where you're like, well, let me just search and see what other people are doing. And there are a lot of myths about serverless.

There's as much good information is out there, there's a lot of bad information out there too. And that's something that is kind of hard to combat, but I think that maybe we could end it there. What are some of the things, the questions people are having, maybe some of the myths, maybe some of the concerns, what are those top ones that you think you could sort of ...

Julian: Dispel.

Jeremy: ... to tell people, dispel, yeah. That you could say, "Look, these are these aren't things to worry about. And again, go and read your blog post series, go and read James' blog post series, and you're going to get the right answers to these things."

Julian: Yeah. I mean, there are misconceptions and some of them are just historical where people think the Lambda functions can only run for five minutes, they can run for 15 minutes. Lambda functions can also now run up to 10 gig of RAM. At re:Invent it was only 3 gig of RAM. That's a three times increase in Lambda functions within a three times proportional increase in CPU. So I like to say, if you had a CPU-intensive job that took 40 minutes and you couldn't run it on Lambda, you've now got three times the CPU. Maybe you can run it on Lambda and now because that would work. So yeah, some of those historical things that have just changed. We've got EFS for Lambda, that's some kind of thing you can't do state with Lambda. EFS and NFS isn't everybody's cup of tea, but that's certainly going to help some people out.

And then the other big one is also cold starts. And this is an interesting one because, obviously we've sort of solved the cold start issue with connecting Lambda functions to VPC, so that's no longer an issue. And that's been a barrier for lots of people, for good reason, and that's now no longer the case. But the other thing for cold starts is interesting because, people do still get caught up at cold starts, but particularly for development because they create a Lambda function, they run it, that's a cold start and then update it and they run it and then go, oh, that's a cold start. And they don't sort of grok that the more you run your Lambda function the less cold starts you have, just because they're warm starts. And it's literally the number of Lambda functions that are running at exactly the same time will have a cold start, but then every subsequent Lambda function invocation for quite a while will be using a warm function.

And so as it ramps up, we see, in the small percentages of cold starts that are actually going to happen. And when we're talking again about the container image support, that's got a whole bunch of complexity, which people are trying to understand. Hopefully, people are learning from this podcast about that as well. But also with the cold starts with that, those are huge and they're particular ways that you can construct your Lambda functions to really reduce those cold starts, and it's best practices anyway. But yeah, cold starts is also definitely one of those myths. And the other one ...

Jeremy: Well, one note on cold starts too, just as something that I find to be interesting. I know that we, I even had to spend time battling with that earlier on, especially with VPC cold starts, that's all sort of gone away now, so much more efficient. The other thing is like provision concurrency. If you're using provision concurrency to get your cold starts down, I'm not even sure that's the right use for provision concurrency. I think provision concurrency is more just to make sure you have enough capacity because of the ramp-up time for Lambda. You certainly can use it for cold starts, but I don't think you need to, that's just my two cents on that.

Julian: Yeah. No, that is true. And they're two different use cases for the same kind of thing. Yeah. As you say, Lambda is pretty scalable, but there is a bit of a ramp-up to get up to many, many, many, many thousands or tens of thousands of concurrent executions. And so yeah, using provision currency, you can get that up in advance. And yeah, some people do also use it for provision concurrency for getting those cold starts done. And yet that is another very valid use case, but it's only an issue for synchronous workloads as well. Anything that is synchronous you really shouldn't be carrying too much. Other than for cost perspective because it's going to take longer to run.

Jeremy: Sure. Sure. I have a feeling that the last one you were going to mention, because this one bugs me quite a bit, is this idea of no ops or some people call it ops-less, which I think is kind of funny. But that's one of those things where, oh, it drives me nuts when I hear this.

Julian: Yeah, exactly. And it's a frustrating thing. And I think often, sometimes when people are talking about no ops, they either have something to sell you. And sometimes what they're selling you is getting rid of something, which never is the case. It's not as though we develop serverless applications and we can then get rid of half of our development team, it just doesn't work like that. And it's crazy, in fact. And when I was talking about the people aspect of IT, this is a super important thing. And me coming from an infrastructure background, everybody is dying in their jobs to do more meaningful work and to do more interesting things and have the agility to try those experiments or try something else. Or do something that's better or even improve the way your build or improve the way your CI/CD pipeline runs or anything, rather than just having to do a lot of work in the lower levels.

And this is what serverless really helps you do, is to be able to, we'll take over a whole lot of the ops for you, but it's not all of the ops, because in a way there's never an end to ops. Because you can always do stuff better. And it's not just the operations of deploying Lambda functions and limits and all that kind of thing. But I mean, think of observability and not knowing just about your application, but knowing about your business. Think of if you had the time that you weren't just monitoring function invocations and monitoring how long things were happening, but imagine if you were able to pull together dashboards of exactly what each transaction costs as it flows through your whole entire application. Think of the benefit of that to your business, or think of the benefit that in real-time, even if it's on Lambda function usage or something, you can say, "Well, oh, there's an immediate drop-off or pick-up in one region in the world or one particular application." You can spot that immediately. That kind of stuff, you just haven't had time to play with to actually build.

But if we can take over some of the operational stuff with you and run one or two or trillions of Lambda functions in the background, just to keep this all ticking along nicely, you're always going to have an opportunity to do more ops. But I think the exciting bit is that ops is not just IT infrastructure, plumbing ops, but you can start even doing even better business ops where you can have more business visibility and more cool stuff for your business because we're not writing apps just for funsies.

Jeremy: Right. Right. And I think that's probably maybe a good way to describe serverless, is it allows you to focus on more meaningful work and more meaningful tasks maybe. Or maybe not more meaningful, but more impactful on the business. Anyways, Julian, listen, this was a great conversation. I appreciate it. I appreciate the work that you're doing over at AWS ...

Julian: Thank you.

Jeremy: ... and the stuff that you're doing. And I hope that there will be a conference soon that we will be able to attend together ...

Julian: I hope so too.

Jeremy: ... maybe grab a drink. So if people want to get a hold of you or find out more about serverless and what AWS is doing with that, how do they do that?

Julian: Yeah, absolutely. Well, please get hold of me anytime on Twitter, is the easiest way probably, julian_wood. Happy to answer your question about anything Serverless or Lambda. And if I don't know the answer, I'll always ask Jeremy, so you're covered twice over there. And then, three different things. James is, if you're talking specifically Lambda, James Beswick's operations guide, have a look at that. Just so much nuggets of super information. We've got another thing we did just sort of jump around, you were talking about cloud formation and the spark was going off in my head. We have something which we're calling the Serverless Patterns Collection, and this is really super cool. We didn't quite get to talk about it, but if you're building applications using SAM or serverless application model, or using the CDK, so either way, we've got a whole bunch of patterns which you can grab.

So if you're pulling something from S3 to Lambda, or from Lambda to EventBridge, or SNS to SQS with a filter, all these kind of things, they're literally copy and paste patterns that you can put immediately into your cloud formation or your CDK templates. So when you are down the rabbit hole of Hacker News or Reddit or Stack Overflow, this is another resource that you can use to copy and paste. So go for that. And that's all hosted on our cool site called serverlessland.com. So that's serverlessland.com and that's an aggregation site that we run because we've got video talks, and we've got blog posts, and we've got learning path series, and we've got a whole bunch of stuff. Personally, I've got a learning path series coming out shortly on Lambda extensions and also one on Lambda observability. There's one coming out shortly on container image supports. And our team is talking all over as many things as we can virtually. I'm actually speaking about container images of DockerCon, which is coming up, which is exciting.

And yeah, so serverlessland.com, that's got a whole bunch of information. That's just an easy one-stop-shop where you can get as much information about AWS services as you can. And if not yet, get in touch, I'm happy to help. I'm happy to also carry your feedback. And yeah, at the moment, just inside, we're sort of doing our planning for the next cycle of what Lambda and what all the service stuff we're going to do. So if you've got an awesome idea, please send it on. And I'm sure you'll be super excited when something pops out in the near issue, maybe just in future for a cool new functionality you could have been involved in.

Jeremy: Well, I know that serverlessland.com is an excellent resource, and it's not that the AWS Compute blog is hard to parse through or anything, but serverlessland.com is certainly a much easier resource to get there. So awesome. Julian, I will get all that stuff in the show notes. Thank you so much.

Julian: Oh, thank you very ... Oh, one more thing I didn't mention is Serverless Office Hours. Every Tuesday at 10:00 AM, Pacific Time, I'm in London, that's 6:00 PM. So Serverless Office Hours for an hour every week, we rotate about five different topics and bring any of your questions, anything. It's not just Lambda, it's Step Functions, API gateway, messaging, Lambda, serverless surprise as well. So have any questions, join us. And the links are also on Serverlessland and it's on Twitter and YouTube. That's another way you can get in touch. And yeah, just to finish up, Jeremy, thank you so much for inviting me. You've been a light in the serverless world and we really, really appreciate it, internally at AWS and personally about how you've created and talked about community and people, and just made the serverless thing such a cool place to be. So, yeah. Thank you for all you've done. And I really appreciate being able to share a little bit of time with you.

Jeremy: Well, thank you. It was great.

View Details

About Rebecca Marshburn

Rebecca's interested in the things that interest people—What's important to them? Why? And when did they first discover it to be so? She's also interested in sharing stories, elevating others' experiences, exploring the intersection of physical environments and human behavior, and crafting the perfect pun for every situation. Today, Rebecca is the Head of Content & Community at Common Room. Prior to Common Room, she led the AWS Serverless Heroes program, where she met the singular Jeremy Daly, and guided content and product experiences for fashion magazines, online blogs, AR/VR companies, education companies, and a little travel outfit called Airbnb.

Twitter: @beccaodelay
LinkedIn: Rebecca Marshburn
Company: www.commonroom.io
Personal work (all proceeds go to the charity of the buyer's choice): www.letterstomyexlovers.com

Watch this episode on YouTube: https://youtu.be/VVEtxgh6GKI

This episode sponsored by CBT Nuggets and Lumigo.

Transcript:
Rebecca: What a day today is! It's not every day you turn 100 times old, and on this day we celebrate Serverless Chats 100th episode with the most special of guests. The gentleman whose voice you usually hear on this end of the microphone, doing the asking, but today he's going to be doing the telling, the one and only, Jeremy Daly, and me. I'm Rebecca Marshburn, and your guest host for Serverless Chats 100th episode, because it's quite difficult to interview yourself. Hey Jeremy!

Jeremy: Hey Rebecca, thank you very much for doing this.

Rebecca: Oh my gosh. I am super excited to be here, couldn't be more honored. I'll give your listeners, our listeners, today, the special day, a little bit of background about us. Jeremy and I met through the AWS Serverless Heroes program, where I used to be a coordinator for quite some time. We support each other in content, conferences, product requests, road mapping, community-building, and most importantly, I think we've supported each other in spirit, and now I'm the head of content and community at Common Room, and Jeremy's leading Serverless Cloud at Serverless, Inc., so it's even sweeter that we're back together to celebrate this Serverless Chats milestone with you all, the most important, important, important, important part of the podcast equation, the serverless community. So without further ado, let's begin.

Jeremy: All right, hit me up with whatever questions you have. I'm here to answer anything.

Rebecca: Jeremy, I'm going to ask you a few heavy hitters, so I hope you're ready.

Jeremy: I'm ready to go.

Rebecca: And the first one's going to ask you to step way, way, way, way, way back into your time machine, so if you've got the proper attire on, let's do it. If we're going to step into that time machine, let's peel the layers, before serverless, before containers, before cloud even, what is the origin story of Jeremy Daly, the man who usually asks the questions.

Jeremy: That's tough. I don't think time machines go back that far, but it's funny, when I was in high school, I was involved with music, and plays, and all kinds of things like that. I was a very creative person. I loved creating things, that was one of the biggest sort of things, and whether it was music or whatever and I did a lot of work with video actually, back in the day. I was always volunteering at the local public access station. And when I graduated from high school, I had no idea what I wanted to do. I had used computers at the computer lab at the high school. I mean, this is going back a ways, so it wasn't everyone had their own computer in their house, but I went to college and then, my first, my freshman year in college, I ended up, there's a suite-mate that I had who showed me a website that he built on the university servers.

And I saw that and I was immediately like, "Whoa, how do you do that"? Right, just this idea of creating something new and being able to build that out was super exciting to me, so I spent the next couple of weeks figuring out how to do HTML, and this was before, this was like when JavaScript was super, super early and we're talking like 1997, and everything was super early. I was using this, I eventually moved away from using FrontPage and started using this thing called HotDog. It was a software for HTML coding, but I started doing that, and I started building websites, and then after a while, I started figuring out what things like CGI-bins were, and how you could write Perl scripts, and how you could make interactions happen, and how you could capture FormData and serve up different things, and it was a lot of copying and pasting.

My major at the time, I think was psychology, because it was like a default thing that I could do. But then I moved into computer science. I did computer science for about a year, and I felt that that was a little bit too narrow for what I was hoping to sort of do. I was starting to become more entrepreneurial. I had started selling websites to people. I had gone to a couple of local businesses and started building websites, so I actually expanded that and ended up doing sort of a major that straddled computer science and management, like business administration. So I ended up graduating with a degree in e-commerce and internet marketing, which is sort of very early, like before any of this stuff seemed to even exist. And then from there, I started a web development company, worked on that for 12 years, and then I ended up selling that off. Did a startup, failed the startup. Then from that startup, went to another startup, worked there for a couple of years, went to another startup, did a lot of consulting in between there, somewhere along the way I found serverless and AWS Cloud, and then now it's sort of led me to advocacy for building things with serverless and now I'm building sort of the, I think what I've been dreaming about building for the last several years in what I'm doing now at Serverless, Inc.

Rebecca: Wow. All right. So this love story started in the 90s.

Jeremy: The 90s, right.

Rebecca: That's an incredible, era and welcome to 2021.

Jeremy: Right. It's been a journey.

Rebecca: Yeah, truly, that's literally a new millennium. So in a broad way of saying it, you've seen it all. You've started from the very HotDog of the world, to today, which is an incredible name, I'm going to have to look them up later. So then you said serverless came along somewhere in there, but let's go to the middle of your story here, so before Serverless Chats, before its predecessor, which is your weekly Off-by-none newsletter, and before, this is my favorite one, debates around, what the suffix "less" means when appended to server. When did you first hear about Serverless in that moment, or perhaps you don't remember the exact minute, but I do really want to know what struck you about it? What stood out about serverless rather than any of the other types of technologies that you could have been struck by and been having a podcast around?

Jeremy: Right. And I think I gave you maybe too much of a surface level of what I've seen, because I talked mostly about software, but if we go back, I mean, hardware was one of those things where hardware, and installing software, and running servers, and doing networking, and all those sort of things, those were part of my early career as well. When I was running my web development company, we started by hosting on some hosting service somewhere, and then we ended up getting a dedicated server, and then we outgrew that, and then we ended up saying, "Well maybe we'll bring stuff in-house". So we did on-prem for quite some time, where we had our own servers in the T1 line, and then we moved to another building that had a T3 line, and if anybody doesn't know what that is, you probably don't need to anymore.

But those are the things that we were doing, and then eventually we moved into a co-location facility where we rented space, and we rented electricity, and we rented all the utilities, the bandwidth, and so forth, but we had Blade servers and I was running VMware, and we were doing all this kind of stuff to manage the infrastructure, and then writing software on top of that, so it was a lot of work. I know I posted something on Twitter a few weeks ago, about how, when I was, when we were young, we used to have to carry a server on our back, uphill, both ways, to the data center, in the snow, with no shoes, and that's kind of how it felt, that you were doing a lot of these things.

And then 2008, 2009, as I was kind of wrapping up my web development company, we were just in the process of actually saying it's too expensive at the colo. I think we were paying probably between like $5,000 and $7,000 a month between the ... we had leases on some of the servers, you're paying for electricity, you're paying for all these other things, and we were running a fair amount of services in there, so it seemed justifiable. We were making money on it, that wasn't the problem, but it just was a very expensive fixed cost for us, and when the cloud started coming along and I started actually building out the startup that I was working on, we were building all of that in the cloud, and as I was learning more about the cloud and how that works, I'm like, I should just move all this stuff that's in the co-location facility, move that over to the cloud and see what happens.

And it took a couple of weeks to get that set up, and now, again, this is early, this is before ELB, this is before RDS, this is before, I mean, this was very, very early cloud. I mean, I think there was S3 and EC2. I think those were the two services that were available, with a few other things. I don't even think there were VPCs yet. But anyways, I moved everything over, took a couple of weeks to get that over, and essentially our bill to host all of our clients' sites and projects went from $5,000 to $7,000 a month, to $750 a month or something like that, and it's funny because had I done that earlier, I may not have sold off my web development company because it could have been much more profitable, so it was just an interesting move there.

So we got into the cloud fairly early and started sort of leveraging that, and it was great to see all these things get added and all these specialty services, like RDS, and just taking the responsibility because I literally was installing Microsoft SQL server on an EC2 instance, which is not something that you want to do, you want to use RDS. It's just a much better way to do it, but anyways, so I was working for another startup, this was like startup number 17 or whatever it was I was working for, and we had this incident where we were using ... we had a pretty good setup. I mean, everything was on EC2 instances, but we were using DynamoDB to do some caching layers for certain things. We were using a sharded database, MySQL database, for product information, and so forth.

So the system was pretty resilient, it was pretty, it handled all of the load testing we did and things like that, but then we actually got featured on Good Morning America, and they mentioned our app, it was the Power to Mobile app, and so we get mentioned on Good Morning America. I think it was Good Morning America. The Today Show? Good Morning America, I think it was. One of those morning shows, anyways, we got about 10,000 sign-ups in less than a minute, which was amazing, or it was just this huge spike in traffic, which was great. The problem was, is we had this really weak point in our system where we had to basically get a lock on the database in order to get an incremental-ID, and so essentially what happened is the database choked, and then as soon as the database choked, just to create user accounts, other users couldn't sign in and there was all kinds of problems, so we basically lost out on all of this capability.

So I spent some time doing a lot of research and trying to figure out how do you scale that? How do you scale something that fast? How do you have that resilience in there? And there's all kinds of ways that we could have done it with traditional hardware, it's not like it wasn't possible to do with a slightly better strategy, but as I was digging around in AWS, I'm looking around at some different things, and we were, I was always in the console cause we were using Dynamo and some of those things, and I came across this thing that said "Lambda," with a little new thing next to it. I'm like, what the heck is this?

So I click on that and I start reading about it, and I'm like, this is amazing. We don't have to spin up a server, we don't have to use Chef, or Puppet, or anything like that to spin up these machines. We can basically just say, when X happens, do Y, and it enlightened me, and this was early 2015, so this would have been right after Lambda went GA. Had never heard of Lambda as part of the preview, I mean, I wasn't sort of in that the re:Invent, I don't know, what would you call that? Vortex, maybe, is a good way to describe the event.

Rebecca: Vortex sounds about right. That's about how it feels by the end.

Jeremy: Right, exactly. So I wasn't really in that, I wasn't in that group yet, I wasn't part of that community, so I hadn't heard about it, and so as I started playing around with it, I immediately saw the value there, because, for me, as someone who again had managed servers, and it had built out really complex networking too. I think some of the things you don't think about when you move to an on-prem where you're managing your stuff, even what the cloud manages for you. I mean, we had firewalls, and we had to do all the firewall rules ourselves, right. I mean, I know you still have to do security groups and things like that in AWS, but just the level of complexity is a lot lower when you're in the cloud, and of course there's so many great services and systems that help you do that now.

But just the idea of saying, "wait a minute, so if I have something happen, like a user signup, for example, and I don't have to worry about provisioning all the servers that I need in order to handle that," and again, it wasn't so much the server aspect of it as it was the database aspect of it, but one of the things that was sort of interesting about the idea of Serverless 2 was this asynchronous nature of it, this idea of being more event-driven, and that things don't have to happen immediately necessarily. So that just struck me as something where it seemed like it would reduce a lot, and again, this term has been overused, but the undifferentiated heavy-lifting, we use that term over and over again, but there is not a better term for that, right?

Because there were just so many things that you have to do as a developer, as an ops person, somebody who is trying to straddle teams, or just a PM, or whatever you are, so many things that you have to do in order to get an application running, first of all, and then even more you have to do in order to keep it up and running, and then even more, if you start thinking about distributing it, or scaling it, or getting any of those things, disaster recovery. I mean, there's a million things you have to think about, and I saw serverless immediately as this opportunity to say, "Wait a minute, this could reduce a lot of that complexity and manage all of that for you," and then again, literally let you focus on the things that actually matter for your business.

Rebecca: Okay. As someone who worked, how should I say this, in metatech, or the technology of technology in the serverless space, when you say that you were starting to build that without ELB even, or RDS, my level of anxiety is like, I really feel like I'm watching a slow horror film. I'm like, "No, no, no, no, no, you didn't, you didn't, you didn't have to do that, did you"?

Jeremy: We did.

Rebecca: So I applaud you for making it to the end of the film and still being with us.

Jeremy: Well, the other thing ...

Rebecca: Only one protagonist does that.

Jeremy: Well, the other thing that's interesting too, about Serverless, and where it was in 2015, Lambda goes GA, this will give you some anxiety, there was no API gateway. So there was no way to actually trigger a Lambda function from a web request, right. There was no VPC access in Lambda functions, which meant you couldn't connect to a database. The only thing you do is connect via HDP, so you could connect to DynamoDB or things like that, but you could not connect directly to RDS, for example. So if you go back and you look at the timeline of when these things were released, I mean, if just from 2015, I mean, you literally feel like a caveman thinking about what you could do back then again, it's banging two sticks together versus where we are now, and the capabilities that are available to us.

Rebecca: Yeah, you're sort of in Plato's cave, right, and you're looking up and you're like, "It's quite dark in here," and Lambda's up there, outside, sowing seeds, being like, "Come on out, it's dark in there". All right, so I imagine you discovering Lambda through the console is not a sentence you hear every day or general console discovery of a new product that will then sort of change the way that you build, and so I'm guessing maybe one of the reasons why you started your Off-by-none newsletter or Serverless Chats, right, is to be like, "How do I help tell others about this without them needing to discover it through the console"? But I'm curious what your why is. Why first the Off-by-none newsletter, which is one of my favorite things to receive every week, thank you for continuing to write such great content, and then why Serverless Chats? Why are we here today? Why are we at number 100? Which I'm so excited about every time I say it.

Jeremy: And it's kind of crazy to think about all the people I've gotten a chance to talk to, but so, I think if you go back, I started writing blog posts maybe in 2015, so I haven't been doing it that long, and I certainly wasn't prolific. I wasn't consistent writing a blog post every week or every, two a week, like some people do now, which is kind of crazy. I don't know how that, I mean, it's hard enough writing the newsletter every week, never mind writing original content, but I started writing about Serverless. I think it wasn't until the beginning of 2018, maybe the end of 2017, and there was already a lot of great content out there. I mean, Ben Kehoe was very early into this and a lot of his stuff I read very early.

I mean, there's just so many people that were very early in the space, I mean, Paul Johnson, I mean, just so many people, right, and I started reading what they were writing and I was like, "Oh, I've got some ideas too, I've been experimenting with some things, I feel like I've gotten to a point where what I could share could be potentially useful". So I started writing blog posts, and I think one of the earlier blog posts I wrote was, I want to say 2017, maybe it was 2018, early 2018, but was a post about serverless security, and what was great about that post was that actually got me connected with Ory Segal, who had started PureSec, and he and I became friends and that was the other great thing too, is just becoming part of this community was amazing.

So many awesome people that I've met, but so I saw all this stuff people were writing and these things people were doing, and I got to maybe August of 2018, and I said to myself, I'm like, "Okay, I don't know if people are interested in what I'm writing". I wasn't writing a lot, but I was writing a little bit, but I wasn't sure people were overly interested in what I was writing, and again, that idea of the imposter syndrome, certainly everything was very early, so I felt a little bit more comfortable. I always felt like, well, maybe nobody knows what they're talking about here, so if I throw something into the fold it won't be too, too bad, but certainly, I was reading other things by other people that I was interested in, and I thought to myself, I'm like, "Okay, if I'm interested in this stuff, other people have to be interested in this stuff," but it wasn't easy to find, right.

I mean, there was sort of a serverless Twitter, if you want to use that terminology, where a lot of people tweet about it and so forth, obviously it's gotten very noisy now because of people slapped that term on way too many things, but I don't want to have that discussion, but so I'm reading all this great stuff and I'm like, "I really want to share it," and I'm like, "Well, I guess the best way to do that would just be a newsletter."

I had an email list for my own personal site that I had had a couple of hundred people on, and I'm like, "Well, let me just turn it into this thing, and I'll share these stories, and maybe people will find them interesting," and I know this is going to sound a little bit corny, but I have two teenage daughters, so I'm allowed to be sort of this dad-jokey type. I remember when I started writing the first version of this newsletter and I said to myself, I'm like, "I don't want this to be a newsletter." I was toying around with this idea of calling it an un-newsletter. I didn't want it to just be another list of links that you click on, and I know that's interesting to some people, but I felt like there was an opportunity to opine on it, to look at the individual links, and maybe even tell a story as part of all of the links that were shared that week, and I thought that that would be more interesting than just getting a list of links.

And I'm sure you've seen over the last 140 issues, or however many we're at now, that there's been changes in the way that we formatted it, and we've tried new things, and things like that, but ultimately, and this goes back to the corny thing, I mean, one of the first things that I wanted to do was, I wanted to basically thank people for writing this stuff. I wanted to basically say, "Look, this is not just about you writing some content". This is big, this is important, and I appreciate it. I appreciate you for writing that content, and I wanted to make it more of a celebration really of the community and the people that were early contributors to that space, and that's one of the reasons why I did the Serverless Star thing.

I thought, if somebody writes a really good article some week, and it's just, it really hits me, or somebody else says, "Hey, this person wrote a great article," or whatever. I wanted to sort of celebrate that person and call them out because that's one of the things too is writing blog posts or posting things on social media without a good following, or without the dopamine hit of people liking it, or re-tweeting it, and things like that, it can be a pretty lonely place. I mean, I know I feel that way sometimes when you put something out there, and you think it's important, or you think people might want to see it, and just not enough people see it.

It's even worse, I mean, 240 characters, or whatever it is to write a tweet is one thing, or 280 characters, but if you're spending time putting together a tutorial or you put together a really good thought piece, or story, or use case, or something where you feel like this is worth sharing, because it could inspire somebody else, or it could help somebody else, could get them past a bump, it could make them think about something a different way, or get them over a hump, or whatever. I mean, that's just the kind of thing where I think people need that encouragement, and I think people deserve that encouragement for the work that they're doing, and that's what I wanted to do with Off-by-none, is make sure that I got that out there, and to just try to amplify those voices the best that I could. The other thing where it's sort of progressed, and I guess maybe I'm getting ahead of myself, but the other place where it's progressed and I thought was really interesting, was, finding people ...

There's the heavy hitters in the serverless space, right? The ones we all know, and you can name them all, and they are great, and they produce amazing content, and they do amazing things, but they have pretty good engines to get their content out, right? I mean, some people who write for the AWS blog, they're on the AWS blog, right, so they're doing pretty well in terms of getting their things out there, right, and they've got pretty good engines.

There's some good dev advocates too, that just have good Twitter followings and things like that. Then there's that guy who writes the story. I don't know, he's in India or he's in Poland or something like that. He writes this really good tutorial on how to do this odd edge-case for serverless. And you go and you look at their Medium and they've got two followers on Medium, five followers on Twitter or something like that. And that to me, just seems unfair, right? I mean, they've written a really good piece and it's worth sharing right? And it needs to get out there. I don't have a huge audience. I know that. I mean I've got a good following on Twitter. I feel like a lot of my Twitter followers, we can have good conversations, which is what you want on Twitter.

The newsletter has continued to grow. We've got a good listener base for this show here. So, I don't have a huge audience, but if I can share that audience with other people and get other people to the forefront, then that's important to me. And I love finding those people and those ideas that other people might not see because they're not looking for them. So, if I can be part of that and help share that, that to me, it's not only a responsibility, it's just it's incredibly rewarding. So ...

Rebecca: Yeah, I have to ... I mean, it is your 100th episode, so hopefully I can give you some kudos, but if celebrating others' work is one of your main tenets, you nail it every time. So ...

Jeremy: I appreciate that.

Rebecca: Just wanted you to know that. So, that's sort of the Genesis of course, of both of these, right?

Jeremy: Right.

Rebecca: That underpins the foundational how to share both works or how to share others' work through different channels. I'm wondering how it transformed, there's this newsletter and then of course it also has this other component, which is Serverless Chats. And that moment when you were like, "All right, this newsletter, this narrative that I'm telling behind serverless, highlighting all of these different authors from all these different global spaces, I'm going to start ... You know what else I want to do? I don't have enough to do, I'm going to start a podcast." How did we get here?

Jeremy: Well, so the funny thing is now that I think about it, I think it just goes back to this tenet of fairness, this idea where I was fortunate, and I was able to go down to New York City and go to Serverless Days New York in late 2018. I was able to ... Tom McLaughlin actually got me connected with a bunch of great people in Boston. I live just outside of Boston. We got connected with a bunch of great people. And we started the Serverless Days Boston for 2019. And we were on that committee. I started traveling and I was going to conferences and I was meeting people. I went to re:Invent in 2018, which I know a lot of people just don't have the opportunity to do. And the interesting thing was, is that I was pulling aside brilliant people either in the hallway at a conference or more likely for a very long, deep discussion that we would have about something at a pub in Northern Ireland or something like that, right?

I mean, these were opportunities that I was getting that I was privileged enough to get. And I'm like, these are amazing conversations. Just things that, for me, I know changed the way I think. And one of the biggest things that I try to do is evolve my thinking. What I thought a year ago is probably not what I think now. Maybe call it flip-flopping, whatever you want to call it. But I think that evolving your thinking is the most progressive thing that you can do and starting to understand as you gain new perspectives. And I was talking to people that I never would have talked to if I was just sitting here in my home office or at the time, I mean, I was at another office, but still, I wasn't getting that context. I wasn't getting that experience. And I wasn't getting those stories that literally changed my mind and made me think about things differently.

And so, here I was in this privileged position, being able to talk to these amazing people and in some cases funny, because they're celebrities in their own right, right? I mean, these are the people where other people think of them and it's almost like they're a celebrity. And these people, I think they deserve fame. Don't get me wrong. But like as someone who has been on that side of it as well, it's ... I don't know, it's weird. It's weird to have fans in a sense. I love, again, you can be my friend, you don't have to be my fan. But that's how I felt about ...

Rebecca: I'm a fan of my friends.

Jeremy: So, a fan and my friend. So, having talked to these other people and having these really deep conversations on serverless and go beyond serverless to me. Actually I had quite a few conversations with some people that have nothing to do with serverless. Actually, Peter Sbarski and I, every time we get together, we only talk about the value of going to college for some reason. I don't know why. It has usually nothing to do with serverless. So, I'm having these great conversations with these people and I'm like, "Wow, I wish I could share these. I wish other people could have this experience," because I can tell you right now, there's people who can't travel, especially a lot of people outside of the United States. They ... it's hard to travel to the United States sometimes.

So, these conversations are going on and I thought to myself, I'm like, "Wouldn't it be great if we could just have these conversations and let other people hear them, hopefully without bar glasses clinking in the background. And so I said, "You know what? Let's just try it. Let's see what happens. I'll do a couple of episodes. If it works, it works. If it doesn't, it doesn't. If people are interested, they're interested." But that was the genesis of that, I mean, it just goes back to this idea where I felt a little selfish having conversations and not being able to share them with other people.

Rebecca: It's the very Jeremy Daly tenet slogan, right? You got to share it. You got to share it ...

Jeremy: Got to share it, right?

Rebecca: The more he shares it, it celebrates it. I love that. I think you do ... Yeah, you do a great job giving a megaphone so that more people can hear. So, in case you need a reminder, actually, I'll ask you, I know what the answer is to this, but do you know the answer? What was your very first episode of Serverless Chats? What was the name, and how long did it last?

Jeremy: What was the name?

Rebecca: Oh yeah. Oh yeah.

Jeremy: Oh, well I know ... Oh, I remember now. Well, I know it was Alex DeBrie. I absolutely know that it was Alex DeBrie because ...

Rebecca: Correct on that.

Jeremy: If nobody, if you do not know Alex DeBrie, not only is he an AWS data hero, as well as the author of The DynamoDB Book, but he's also like the most likable person on the planet too. It is really hard if you've ever met Alex, that you wouldn't remember him. Alex and I started communicating, again, we met through the serverless space. I think actually he was working at Serverless Inc. at the time when we first met. And I think I met him in person, finally met him in person at re:Invent 2018. But he and I have collaborated on a number of things and so forth. So, let me think what the name of it was. "Serverless Purity Versus Practicality" or something like that. Is that close?

Rebecca: That's exactly what it was.

Jeremy: Oh, all right. I nailed it. Nailed it. Yes!

Rebecca: Wow. Well, it's a great title. And I think ...

Jeremy: Don't ask me what episode number 27 was though, because no way I could tell you that.

Rebecca: And just for fun, it was 34 minutes long and you released it on June 17th, 2019. So, you've come a long way in a year and a half. That's some kind of wildness. So it makes sense, like, "THE," capital, all caps, bold, italic, author for databases, Alex DeBrie. Makes sense why you selected him as your guest. I'm wondering if you remember any of the ... What do you remember most about that episode? What was it like planning it? What was the reception of it? Anything funny happened recording it or releasing it?

Jeremy: Yeah, well, I mean, so the funny thing is that I was incredibly nervous. I still am, actually a lot of guests that I have, I'm still incredibly nervous when I'm about to do the actual interview. And I think it's partially because I want to do justice to the content that they're presenting and to their expertise. And I feel like there's a responsibility to them, but I also feel like the guests that I've had on, some of them are just so smart, and the things they say, just I'm in awe of some of the things that come out of these people's mouths. And I'm like, "This is amazing and people need to hear this." And so, I feel like we've had really good episodes and we've had some okay episodes, but I feel like I want to try to keep that level up so that they owe that to my listener to make sure that there is high quality episode that, high quality information that they're going to get out of that.

But going back to the planning of the initial episodes, so I actually had six episodes recorded before I even released the first one. And the reason why I did that was because I said, "All right, there's no way that I can record an episode and then wait a week and then record another episode and wait a week." And I thought batching them would be a good idea. And so, very early on, I had Alex and I had Nitzan Shapira and I had Ran Ribenzaft and I had Marcia Villalba and I had Erik Peterson from Cloud Zero. And so, I had a whole bunch of these episodes and I reached out to I think, eight or nine people. And I said, "I'm doing this thing, would you be interested in it?" Whatever, and we did planning sessions, still a thing that I do today, it's still part of the process.

So, whenever I have a guest on, if you are listening to an episode and you're like, "Wow, how did they just like keep the thing going ..." It's not scripted. I don't want people to think it's scripted, but it is, we do review the outline and we go through some talking points to make sure that again, the high-quality episode and that the guest says all the things that the guest wants to say. A lot of it is spontaneous, right? I mean, the language is spontaneous, but we do, we do try to plan these episodes ahead of time so that we make sure that again, we get the content out and we talk about all the things we want to talk about. But with Alex, it was funny.

He was actually the first of the six episodes that I recorded, though. And I wasn't sure who I was going to do first, but I hadn't quite picked it yet, but I recorded with Alex first. And it was an easy, easy conversation. And the reason why it was an easy conversation was because we had talked a number of times, right? It was that in a pub, talking or whatever, and having that friendly chat. So, that was a pretty easy conversation. And I remember the first several conversations I had, I knew Nitzan very well. I knew Ran very well. I knew Erik very well. Erik helped plan Serverless Days Boston with me. And I had known Marcia very well. Marcia actually had interviewed me when we were in Vegas for re:Invent 2018.

So, those were very comfortable conversations. And so, it actually was a lot easier to do, which probably gave me a false sense of security. I was like, "Wow, this was ... These came out pretty well." The conversations worked pretty well. And also it was super easy because I was just doing audio. And once you add the video component into it, it gets a little bit more complex. But yeah, I mean, I don't know if there's anything funny that happened during it, other than the fact that I mean, I was incredibly nervous when we recorded those, because I just didn't know what to expect. If anybody wants to know, "Hey, how do you just jump right into podcasting?" I didn't. I actually was planning on how can I record my voice? How can I get comfortable behind a microphone? And so, one of the things that I did was I started creating audio versions of my blog posts and posting them on SoundCloud.

So, I did that for a couple of ... I'm sorry, a couple of blog posts that I did. And that just helped make me feel a bit more comfortable about being able to record and getting a little bit more comfortable, even though I still can't stand the sound of my own voice, but hopefully that doesn't bother other people.

Rebecca: That is an amazing ... I think we so often talk about ideas around you know where you want to go and you have this vision and that's your goal. And it's a constant reminder to be like, "How do I make incremental steps to actually get to that goal?" And I love that as a life hack, like, "Hey, start with something you already know that you wrote and feel comfortable in and say it out loud and say it out loud again and say it out loud again." And you may never love your voice, but you will at least feel comfortable saying things out loud on a podcast.

Jeremy: Right, right, right. I'm still working on the, "Ums" and, "Ahs." I still do that. And I don't edit those out. That's another thing too, actually, that one of the things I do want people to know about this podcast is these are authentic conversations, right? I am probably like ... I feel like I'm, I mean, the most authentic person that I know. I just want authenticity. I want that out of the guests. The idea of putting together an outline is just so that we can put together a high quality episode, but everything is authentic. And that's what I want out of people. I just want that authenticity, and one of the things that I felt kept that, was leaving in, "Ums" and, "Ahs," you know what I mean? It's just, it's one of those things where I know a lot of podcasts will edit those out and it sounds really polished and finished.

Again, I mean, I figured if we can get the clinking glasses out from the background of a bar and just at least have the conversation that that's what I'm trying to achieve. And we do very little editing. We do cut things out here and there, especially if somebody makes a mistake or they want to start something over again, we will cut that out because we want, again, high quality episodes. But yeah, but authenticity is deeply important to me.

Rebecca: Yeah, I think it probably certainly helps that neither of us are robots because robots wouldn't say, "Um" so many times. As I say, "Uh." So, let's talk about, Alex DeBrie was your first guest, but there's been a hundred episodes, right? So, from, I might say the best guest, as a hundredth episode guests, which is our very own Jeremy Daly, but let's go back to ...

Jeremy: I appreciate that.

Rebecca: Your guests, one to 99. And I mean, you've chatted with some of the most thoughtful, talented, Serverless builders and architects in the industry, and across coincident spaces like ML and Voice Technology, Chaos Engineering, databases. So, you started with Alex DeBrie and databases, and then I'm going to list off some names here, but there's so many more, right? But there's the Gunnar Grosches, and the Alexandria Abbasses, and Ajay Nair, and Angela Timofte, James Beswick, Chris Munns, Forrest Brazeal, Aleksandar Simovic, and Slobodan Stojanovic. Like there are just so many more. And I'm wondering if across those hundred conversations, or 99 plus your own today, if you had to distill those into two or three lessons, what have you learned that sticks with you? If there are emerging patterns or themes across these very divergent and convergent thinkers in the serverless space?

Jeremy: Oh, that's a tough question.

Rebecca: You're welcome.

Jeremy: So, yeah, put me on the spot here. So, yeah, I mean, I think one of the things that I've, I've seen, no matter what it's been, whether it's ML or it's Chaos Engineering, or it's any of those other observability and things like that. I think the common thing that threads all of it is trying to solve problems and make people's lives easier. That every one of those solutions is like, and we always talk about abstractions and, and higher-level abstractions, and we no longer have to write ones and zeros on punch cards or whatever. We can write languages that either compile or interpret it or whatever. And then the cloud comes along and there's things we don't have to do anymore, that just get taken care of for us.

And you keep building these higher level of abstractions. And I think that's a lot of what ... You've got this underlying concept of letting somebody else handle things for you. And then you've got this whole group of people that are coming at it from a number of different angles and saying, "Well, how will that apply to my use case?" And I think a lot of those, a lot of those things are very, very specific. I think things like the voice technology where it's like the fact that serverless powers voice technology is only interesting in the fact as to say that, the voice technology is probably the more interesting part, the fact that serverless powers it is just the fact that it's a really simple vehicle to do that. And basically removes this whole idea of saying I'm building voice technology, or I'm building a voice app, why do I need to worry about setting up servers and all this kind of stuff?

It just takes that away. It takes that out of the equation. And I think that's the perfect idea of saying, "How can you take your use case, fit serverless in there and apply it in a way that gets rid of all that extra overhead that you shouldn't have to worry about." And the same thing is true of machine learning. And I mean, and SageMaker, and things like that. Yeah, you're still running instances of it, or you still have to do some of these things, but now there's like SageMaker endpoints and some other things that are happening. So, it's moving in that direction as well. But then you have those really high level services like NLU API from IBM, which is the Watson Natural Language Processing.

You've got AP recognition, you've got the vision API, you've got sentiment analysis through all these different things. So, you've got a lot of different services that are very specific to machine learning and solving a discrete problem there. But then basically relying on serverless or at least presenting it in a way that's serverless, where you don't have to worry about it, right? You don't have to run all of these Jupiter notebooks and things like that, to do machine learning for a lot of cases. This is one of the things I talk about with Alexandra Abbas, was that these higher level APIs are just taking a lot of that responsibility or a lot of that heavy lifting off of your plate and allowing you to really come down and focus on the things that you're doing.

So, going back to that, I do think that serverless, that the common theme that I see is that this idea of worrying about servers and worrying about patching things and worrying about networking, all that stuff. For so many people now, that's just not even a concern. They didn't even think about it. And that's amazing to think of, compute ... Or data, or networking as a utility that is now just available to us, right? And I mean, again, going back to my roots, taking it for granted is something that I think a lot of people do, but I think that's also maybe a good thing, right? Just don't think about it. I mean, there are people who, they're still going to be engineers and people who are sitting in the data center somewhere and racking servers and doing it, that's going to be forever, right?

But for the things that you're trying to build, that's unimportant to you. That is the furthest from your concern. You want to focus on the problem that you're trying to solve. And so I think that, that's a lot of what I've seen from talking to people is that they are literally trying to figure out, "Okay, how do I take what I'm doing, my use case, my problem, how do I take that to the next level, by being able to spend my cycles thinking about that as opposed to how I'm going to serve it up to people?"

Rebecca: Yeah, I think it's the mantra, right, of simplify, simplify, simplify, or maybe even to credit Bruce Lee, be like water. You're like, "How do I be like water in this instance?" Well, it's not to be setting up servers, it's to be doing what I like to be doing. So, you've interviewed these incredible folks. Is there anyone left on your list? I'm sure there ... I mean, I know that you have a large list. Is there a few key folks where you're like, "If this is the moment I'm going to ask them, I'm going to say on the hundredth episode, 'Dear so-and-so, I would love to interview you for Serverless Chats.'" Who are you asking?

Jeremy: So, this is something that, again, we have a stretch list of guests that we attempt to reach out to every once in a while just to say, "Hey, if we get them, we get them." But so, I have a long list of people that I would absolutely love to talk to. I think number one on my list is certainly Werner Vogels. I mean, I would love to talk to Dr. Vogels about a number of things, and maybe even beyond serverless, I'm just really interested. More so from a curiosity standpoint of like, "Just how do you keep that in your head?" That vision of where it's going. And I'd love to drill down more into the vision because I do feel like there's a marketing aspect of it, that's pushing on him of like, "Here's what we have to focus on because of market adoption and so forth. And even though the technology, you want to move into a certain way," I'd be really interesting to talk to him about that.

And I'd love to talk to him more too about developer experience and so forth, because one of the things that I love about AWS is that it gives you so many primitives, but at the same time, the thing I hate about AWS is it gives you so many primitives. So, you have to think about 800 services, I know it's not that many, but like, what is it? 200 services, something like that, that all need to kind of connect together. And I love that there's that diversity in those capabilities, it's just from a developer standpoint, it's really hard to choose which ones you're supposed to use, especially when several services overlap. So, I'm just curious. I mean, I'd love to talk to him about that and see what the vision is in terms of, is that the idea, just to be a salad bar, to be the Golden Corral of cloud services, I guess, right?

Where you can choose whatever you want and probably take too much and then not use a lot of it. But I don't know if that's part of the strategy, but I think there's some interesting questions, could dig in there. Another person from AWS that I actually want to talk to, and I haven't reached out to her yet just because, I don't know, I just haven't reached out to her yet, but is Brigid Johnson. She is like an IAM expert. And I saw her speak at re:Inforce 2019, it must have been 2019 in Boston. And it was like she was speaking a different language, but she knew IAM so well, and I am not a fan of IAM. I mean, I'm a fan of it in the sense that it's necessary and it's great, but I can't wrap my head around so many different things about it. It's such a ...

It's an ongoing learning process and when it comes to things like being able to use tags to elevate permissions. Just crazy things like that. Anyways, I would love to have a conversation with her because I'd really like to dig down into sort of, what is the essence of IAM? What are the things that you really have to think about with least permission? Especially applying it to serverless services and so forth. And maybe have her help me figure out how to do some of the cross role IAM things that I'm trying to do. Certainly would love to speak to Jeff Barr. I did meet Jeff briefly. We talked for a minute, but I would love to chat with him.

I think he sets a shining example of what a developer advocate is. Just the way that ... First of all, he's probably the only person alive who knows every service at AWS and has actually tried it because he writes all those blog posts about it. So that would just be great to pick his brain on that stuff. Also, Adrian Cockcroft would be another great person to talk to. Just this idea of what he's done with microservices and thinking about the role, his role with Netflix and some of those other things and how all that kind of came together, I think would be a really interesting conversation. I know I've seen this in so many of his presentations where he's talked about the objections, what were the objections of Lambda and how have you solved those objections? And here's the things that we've done.

And again, the methodology of that would be really interesting to know. There's a couple of other people too. Oh, Sam Newman who wrote Building Microservices, that was my Bible for quite some time. I had it on my iPad and had a whole bunch of bookmarks and things like that. And if anybody wants to know, one of my most popular posts that I've ever written was the ... I think it was ... What is it? 16, 17 architectural patterns for serverless or serverless microservice patterns on AWS. Can't even remember the name of my own posts. But that post was very, very popular. And that even was ... I know Matt Coulter who did the CDK. He's done the whole CDK ... What the heck was that? The CDKpatterns.com. That was one of the things where he said that that was instrumental for him in seeing those patterns and being able to use those patterns and so forth.

If anybody wants to know, a lot of those patterns and those ideas and those ... The sort of the confidence that I had with presenting those patterns, a lot of that came from Sam Newman's work in his Building Microservices book. So again, credit where credit is due. And I think that that would be a really fascinating conversation. And then Simon Wardley, I would love to talk to. I'd actually love to ... I actually talked to ... I met Lin Clark in Vegas as well. She was instrumental with the WebAssembly stuff, and I'd love to talk to her. Merritt Baer. There's just so many people. I'm probably just naming too many people now. But there are a lot of people that I would love to have a chat with and just pick their brain.

And also, one of the things that I've been thinking about a lot on the show as well, is the term "serverless." Good or bad for some people. Some of the conversations we have go outside of serverless a little bit, right? There's sort of peripheral to it. I think that a lot of things are peripheral to serverless now. And there are a lot of conversations to be had. People who were building with serverless. Actually real-world examples.

One of the things I love hearing was Yan Cui's "Real World Serverless" podcast where he actually talks to people who are building serverless things and building them in their organizations. That is super interesting to me. And I would actually love to have some of those conversations here as well. So if anyone's listening and you have a really interesting story to tell about serverless or something peripheral to serverless please reach out and send me a message and I'd be happy to talk to you.

Rebecca: Well, good news is, it sounds like A, we have at least ... You've got at least another a hundred episodes planned out already.

Jeremy: Most likely. Yeah.

Rebecca: And B, what a testament to Sam Newman. That's pretty great when your work is referred to as the Bible by someone. As far as in terms of a tome, a treasure trove of perhaps learnings or parables or teachings. I ... And wow, what a list of other folks, especially AWS power ... Actually, not AWS powerhouses. Powerhouses who happened to work at AWS. And I think have paved the way for a ton of ways of thinking and even communicating. Right? So I think Jeff Barr, as far as setting the bar, raising the bar if you will. For how to teach others and not be so high-level, or high-level enough where you can follow along with him, right? Not so high-level where it feels like you can't achieve what he's showing other people how to do.

Jeremy: Right. And I just want to comment on the Jeff Barr thing. Yeah.

Rebecca: Of course.

Jeremy: Because again, I actually ... That's my point. That's one of the reasons why I love what he does and he's so perfect for that position because he's relatable and he presents things in a way that isn't like, "Oh, well, yeah, of course, this is how you do this." I mean, it's not that way. It's always presented in a way to make it accessible. And even for services that I'm not interested in, that I know that I probably will never use, I generally will read Jeff's post because I feel it gives me a good overview, right?

Rebecca: Right.

Jeremy: It just gives me a good overview to understand whether or not that service is even worth looking at. And that's certainly something I don't get from reading the documentation.

Rebecca: Right. He's inviting you to come with him and understanding this, which is so neat. So I think ... I bet we should ... I know that we can find all these twitter handles for these folks and put them in the show notes. And I'm especially ... I'm just going to say here that Werner Vogels's twitter handle is @Werner. So maybe for your hundredth, all the listeners, everyone listening to this, we can say, "Hey, @Werner, I heard that you're the number one guest that Jeremy Daly would like to interview." And I think if we get enough folks saying that to @Werner ... Did I say that @Werner, just @Werner?

Jeremy: I think you did.

Rebecca: Anyone if you can hear it.

Jeremy: Now listen, he did retweet my serverless musical that I did. So ...

Rebecca: That's right.

Jeremy: I'm sort of on his radar maybe.

Rebecca: Yeah. And honestly, he loves serverless, especially with the number of customers and the types of customers and ... that are doing incredible things with it. So I think we've got a chance, Jeremy. I really do. That's what I'm trying to say.

Jeremy: That's good to know. You're welcome anytime. He's welcome anytime.

Rebecca: Do we say that @Werner, you are welcome anytime. Right. So let's go back to the genesis, not necessarily the genesis of the concept, right? But the genesis of the technology that spurred all of these other technologies, which is AWS Lambda. And so what ... I don't think we'd be having these conversations, right, if AWS Lambda was not released in late 2014, and then when GA I believe in 2015.

Jeremy: Right.

Rebecca: And so subsequently the serverless paradigm was thrust into the spotlight. And that seems like eons ago, but also three minutes ago.

Jeremy: Right.

Rebecca: And so I'm wondering ... Let's talk about its evolution a bit and a bit of how if you've been following it for this long and building it for this long, you've covered topics from serverless CI/CD pipelines, observability. We already talked about how it's impacted voice technologies or how it's made it easy. You can build voice technology without having to care about what that technology is running on.

Jeremy: Right.

Rebecca: You've even talked about things like the future and climate change and how it relates to serverless. So some of those sort of related conversations that you were just talking about wanting to have or having had with previous guests. So as a host who thinks about these topics every day, I'm wondering if there's a topic that serverless hasn't touched yet or one that you hope it will soon. Those types of themes, those threads that you want to pull in the next 100 episodes.

Jeremy: That's another tough question. Wow. You got good questions.

Rebecca: That's what I said. Heavy hitters. I told you I'd be bringing it.

Jeremy: All right. Well, I appreciate that. So that's actually a really good question. I think the evolution of serverless has seen its ups and downs. I think one of the nice things is you look at something like serverless that was so constrained when it first started. And it still has constraints, which are good. But it ... Those constraints get lifted. We just talked about Adrian's talks about how it's like, "Well, I can't do this, or I can't do that." And then like, "Okay, we'll add some feature that you can do that and you can do that." And I think that for the most part, and I won't call it anything specific, but I think for the most part that the evolution of serverless and the evolution of Lambda and what it can do has been thoughtful. And by that I mean that it was sort of like, how do we evolve this into a way that doesn't create too much complexity and still sort of holds true to the serverless ethos of sort of being fairly easy or just writing code.

And then, but still evolve it to open up these other use cases and edge cases. And I think that for the most part, that it has held true to that, that it has been mostly, I guess, a smooth ride. There are several examples though, where it didn't. And I said I wasn't going to call anything out, but I'm going to call this out. I think RDS proxy wasn't great. I think it works really well, but I don't think that's the solution to the problem. And it's a band-aid. And it works really well, and congrats to the engineers who did it. I think there's a story about how two different teams were trying to build it at the same time actually. But either way, I look at that and I say, "That's a good solution to the problem, but it's not the solution to the problem."

And so I think serverless has stumbled in a number of ways to do that. I also feel EFS integration is super helpful, but I'm not sure that's the ultimate goal to share ... The best way to share state. But regardless, there are a whole bunch of things that we still need to do with serverless. And a whole bunch of things that we still need to add and we need to build, and we need to figure out better ways to do maybe. But I think in terms of something that doesn't get talked about a lot, is the developer experience of serverless. And that is, again I'm not trying to pitch anything here. But that's literally what I'm trying to work on right now in my current role, is just that that developer experience of serverless, even though there was this thoughtful approach to adding things, to try to check those things off the list, to say that it can't do this, so we're going to make it be able to do that by adding X, Y, and Z.

As amazing as that has been, that has added layers and layers of complexity. And I'll go back way, way back to 1997 in my dorm room. CGI-bins, if people are not familiar with those, essentially just running on a Linux server, it was a way that it would essentially run a Perl script or other types of scripts. And it was essentially like you're running PHP or you're running Node, or you're running Ruby or whatever it was. So it would run a programming language for you, run a script and then serve that information back. And of course, you had to actually know ins and outs, inputs and outputs. It was more complex than it is now.

But anyways, the point is that back then though, once you had the script written. All you had to do is ... There's a thing called FTP, which I'm sure some people don't even know what that is anymore. File transfer protocol, where you would basically say, take this file from my local machine and put it on this server, which is a remote machine. And you would do that. And the second you did that, magically it was updated and you had this thing happening. And I remember there were a lot of jokes way back in the early, probably 2017, 2018, that serverless was like the new CGI-bin or something like that. But more as a criticism of it, right? Or it's just CGI-bins reborn, whatever. And I actually liked that comparison. I felt, you know what? I remember the days where I just wrote code and I just put it to some other server where somebody was dealing with it, and I didn't even have to think about that stuff.

We're a long way from that now. But that's how serverless felt to me, one of the first times that I started interacting with it. And I felt there was something there, that was something special about it. And I also felt the constraints of serverless, especially the idea of not having state. People rely on things because they're there. But when you don't have something and you're forced to think differently and to make a change or find a way to work around it. Sometimes workarounds, turn into best practices. And that's one of the things that I saw with serverless. Where people were figuring out pretty quickly, how to build applications without state. And then I think the problem is that you had a lot of people who came along, who were maybe big customers of AWS. I don't know.

I'm not going to say that you might be influenced by large customers. I know lots of places are. That said, "We need this." And maybe your ... The will gets bent, right. Because you just... you can only fight gravity for so long. And so those are the kinds of things where I feel some of the stuff has been patchwork and those patchwork things haven't ruined serverless. It's still amazing. It's still awesome what you can do within the course. We're still really just focusing on fast here, with everything else that's built. With all the APIs and so forth and everything else that's serverless in the full-service ecosystem. There's still a lot of amazing things there. But I do feel we've become so complex with building serverless applications, that you can't ... the Hello World is super easy, but if you're trying to build an actual application, it's a whole new mindset.

You've got to learn a whole bunch of new things. And not only that, but you have to learn the cloud. You have to learn all the details of the cloud, right? You need to know all these different things. You need to know cloud formation or serverless framework or SAM or something like that, in order to get the stuff into the cloud. You need to understand the infrastructure that you're working with. You may not need to manage it, but you still have to understand it. You need to know what its limitations are. You need to know how it connects. You need to know what the failover states are like.

There's so many things that you need to know. And to me, that's a burden. And that's adding new types of undifferentiated heavy-lifting that shouldn't be there. And that's the conversation that I would like to have continuing to move forward is, how do you go back to a developer experience where you're saying you're taking away all this stuff. And again, to call out Werner again, he constantly says serverless is about writing code, but ask anybody who builds serverless applications. You're doing a lot more than writing code right now. And I would love to see us bring the conversation back to how do we get back there?

Rebecca: Yeah. I think it kind of goes back to ... You and I have talked about this notion of an ode to simplicity. And it's sort of what you want to write into your ode, right? If we're going to have an ode to simplicity, how do we make sure that we keep the simplicity inside of the ode?

Jeremy: Right.

Rebecca:
So I've got ... I don't know if you've seen these.

Jeremy: I don't know.

Rebecca: But before I get to some wrap-up questions more from the brainwaves of Jeremy Daly, I don't want to forget to call out some long-time listener questions. And they wrote in a via Twitter and they wanted to perhaps pick your brain on a few things.

Jeremy: Okay.

Rebecca: So I don't know if you're ready for this.

Jeremy: A-M-A. A-M-A.

Rebecca: I don't know if you've seen these. Yeah, these are going to put you in the ...

Jeremy: A-M-A-M. Wait, A-M-A-A? Asked me almost anything? No, go ahead. Ask me anything.

Rebecca: A-M-A-A. A-M-J. No. Anyway, we got it. Ask Jeremy almost anything.

Jeremy: There you go.

Rebecca: So there's just three to tackle for today's episode that I'm going to lob at you. One is from Ken Collins. "What will it take to get you back to a relational database of Lambda?"

Jeremy: Ooh, I'm going to tell you right now. And without a doubt, Aurora Serverless v2. I played around with that right after re:Invent 2000. What was it? 20. Yeah. Just came out, right? I'm trying to remember what year it is at this point.

Rebecca: Yes. Indeed.

Jeremy: When that just ... Right when that came out. And I had spent a lot of time with Aurora Serverless v1, I guess if you want to call it that. I spent a lot of time with it. I used it on a couple of different projects. I had a lot of really good success with it. I had the same pains as everybody else did when it came to scaling and just the slowness of the scaling and then ... And some of the step-downs and some of those things. There were certainly problems with it. But v2 just the early, early preview version of v2 was ... It was just a marvel of engineering. And the way that it worked was just ... It was absolutely fascinating.

And I know it's getting ready or it's getting close, I think, to being GA. And when that becomes GA, I think I will have a new outlook on whether or not I can fit RDS into my applications. I will say though. Okay. I will say, I don't think that transactional applications should be using relational databases though. One of the things that was sort of a nice thing about moving to serverless, speaking of constraints. Was this idea that MySQL or Postgres or whatever, really didn't have the scale or without, again, engineering a whole cluster and failover and sharding and all kinds of crazy things like that to make sure that you had the scale. Relational databases were just not the best choice when you were building things with serverless.

And so when I quickly realized that, I tried to find a solution. So I built something called Serverless MySQL, which sort of is a ... And again, I don't want the RDS proxy people to think that because RDS proxy sort of was trying to solve the same problem that Serverless MySQL was, that I have any problem with that. I'm actually glad. The fact that I had to use my Serverless MySQL was only because there wasn't a better solution for it. But I built that because I wanted to continue to use it. And even though I built that and it worked, there was just so many limitations. And it was one of those things where using NoSQL or no SQL just made so much more sense. And I forced myself into thinking that way because of the constraint. And that was huge, that changed my mind on how NoSQL works.

And I absolutely have to call out Rick Houlihan as well as Alex DeBrie. But Rick Houlihan, speaking about another sort of person who influenced so many of my thought processes and changed my mind so dramatically. When I saw Rick's 2018 talk at re:Invent about single table design. And I know that they're calling it something different now, but essentially single table design. This idea of ... I went back and watched 2017. And that was like, Okay, now my mind's moving around. And then watching 2017, watching 2018, then watching 2019. 2019, I was in his session watching it. And I could see his evolution of thinking. Of how he changed the way that he was approaching different problems and the patterns that he was using. And that clicked so much for me, that now I think about ... I feel like, this is going to sound strange, but if you've seen the movie The Matrix. You've probably seen the movie The Matrix.

Rebecca: Oh, yeah.

Jeremy: Oh, of course. Okay. When one of the guys is watching the green character scroll down the screen and he's like, "I don't even see the code anymore. I just see there's a blonde, there's whatever." That's how I feel looking at DynamoDB now. It just makes sense to me. It clicked. And I wish everybody could feel that way. I don't feel superior or anything like that, but it just works. It just works in my brain now. And that's one of the reasons why I built DynamoDB toolbox too, was I was trying to say, "How can I translate how I see this into a way that might be more relatable to other people and hopefully get them to sort of click." And now actually at Serverless Inc. one of the things with serverless clouds, we're building something called Serverless Data, which is a similar key-value store type thing.

Which again is me ... Is my manifestation of how I envision sort of the interface into these things and hopefully will make sense to people. So, but yeah. So to answer Ken's question. Serverless Aurora or Aurora Serverless v2, will definitely get me using it for anything that's got to be analytic or analytical processing or analytics process processing. And also probably using it as I sort of do now, as sort of a secondary, a separate secondary data store so that I can run queries on data, even though I want it to be more highly available on the front end through something like DynamoDB.

Rebecca: Oh my gosh. That moment where it clicks, I just have this mental image of two brain synapses extending their hands toward each other and finally touching. Yeah. Finally touching index fingers, being like, "Hey, we did it." All right. So from Matt Colter.

Jeremy: Oh.
Rebecca: Comes another one.

Jeremy: Love Matt too.

Rebecca: Right? There you got some fans and some friends. Some friends that you're fans of, or that are fans of you.

Jeremy: Speaking of Northern Ireland. Nor, Nor, Nor, Noren Ire ... what do they say? Noren Iron? Noren Iron.

Rebecca: I would... It would be trouble if I tried to pronounce it the way they say it. "So with the terms serverless-first," or Matt asks, "With the term serverless-first and cloud-native causing confusion as to what we know it means (code as a liability), do you have a less confusing name we could use?"

Jeremy: Ooh. So it's funny. I had this conversation with Jay Nair over breakfast one time where we were, I think we were at the serverless breakfast. Where were we? It must've been re:Invent I guess. But we were having this conversation and he was, he asked me a very similar question. And I said, "Look, it's as simple. Serverless is just the way." Right. So when somebody says, "So what's the term for serverless-first or for cloud-native or whatever." When you're building an application, it's just the way. That, it blows my mind. Now look, containers are great, by the way. I love containers. I know I don't talk about them very much. I'm not a fan of Kubernetes because I think it's overly complex for what it needs to do, but it works really well too. So I mean that if you know how to manage it, which not very many people do, that's why all the cloud providers are doing it now that, using containers it can be a very, very good way and a smart way to build your applications. You get portability around them. There're all kinds of reasons to do it. I love the fact that ... My skin was crawling a little bit when they announced that Lambda would support containers. Then when I realized it was just a packaging format, I was like, "Okay, that's much better. That I can deal with". But I think that's one of the things where it introduces a level of comfortability.

And if you put your, or if I should say, if I put my product manager hat on and I'm looking at that, I need to find familiarity, right? When I'm building a product, I need something that's familiar to people and not necessarily revolutionary, maybe evolutionary. I think serverless is revolutionary, which is part of the problem why it's not being adopted. I think as quickly as it could be because it's not quite as evolutionary as something like containers were. Containers are like the next step in VMs. Or it's this idea where you can now split things up into little, smaller chunks, and smaller chunks, and so forth, and then came orchestration and you had all the problems around that.

So I think that no matter how you're building your applications, whether you're building them in containers as a packaging format, as a runtime, whatever, how you serve those containers up, whether those are again, serverless, running in Lambda and Firecracker, or you're running them on IBM, or you're running them on Google Cloud ... Google Cloud Run, I think is a fascinating technology. No matter how you're doing it, I think the key here is to focus on the fact that you should be building your applications in a way that is going to be able to run on one of these modern types of infrastructures.

So that's the only thing I would say, if you're trying to write code to run directly on an EC2 instance, 1997 called, they want their ... I don't know their Pearl Jam shirt back, I don't know anyways. So ...

Rebecca: They want their HotDog shirt back.

Jeremy: They want their HotDog shirt back, right. I would say they want their Green Day shirt back, but I love Green Day. I was just listening to them the other day, anyways. So again, I don't think I have a better term for you, Matt, unfortunately, other than just to say, I don't think serverless is a very good term. I've never been a fan of the term. I've tried to defend it. And then I feel like what's the saying, if you argue with an idiot in public, people don't know that anyways. I don't know what the saying is. I was really ...

Rebecca: Who is the idiot?

Jeremy: Who's the idiot, right? Exactly. So it's one of those things. So the problem is, is that that term is not great. And things like serverless first, I liked that idea that a PR I mean, I love the sentiment of it, right? Like this idea of saying like, we're going to try to build everything serverless first, and then we're going to fall back to containers on the things we can't. And then maybe the things that we can't do, then we're going to fall back to VMs. And then in the worst-case scenario, you've got to build stuff on bare metal.

But so I love that sentiment. And I think that sentiment is always going to be the way, but I think there's just this idea of like, this is just modern app development. Like, this is just how you build apps now. And if you're building apps starting on the bottom up, unless you have a really compelling reason to do that, you are really handicapping yourself. And you're going to struggle because this is just the way the world's moving. So again, I just say, it's the way, it's the way to the way to build applications now.

Rebecca: Yeah. And as someone who had a large part in writing so many of those serverless messaging, serverless narratives, Lambda as packaged as containers, I'm sorry for making your skin crawl and to any listeners that I made crawl, I am @beccaodelay if you want to hit me up on Twitter about that, and I can apologize to you publicly. So lastly from Rob Sutter and, and maybe my personal favorite, because I can hear Rob's brand of humor in this question, since I have the great honor of knowing him personally as well, but he asks, "Do you ever go back and stand up instances just to remind yourself of everything you don't have to deal with anymore?"

Jeremy: Well, I just had Rob on episode 99, talking about FaunaDB, which is also a fascinating thing. So super exciting stuff that they're doing over there. So every once in a while, and I'm going to admit to this and hopefully people won't flood my website to crash it, but I have been so busy doing these other things that my personal website, jeremydaly.com is still a load-balanced WordPress site running on EC2 with an ELB in front of it, or maybe an ALB, whatever it is. And, and that is still out there running on multiple instances. And I have backup instances. I was hit on a Hacker News article one time that they shut my site down. So I had to spin up some additional instances, but that is still running there. So I am so afraid of going to touch that thing, that, because I just haven't done it in so long that I actually avoid it and I haven't done it yet.

So Off-by-none and Serverless Chat sites those are all static Jamstack sites. But yeah ... but so I don't do it often, but every once in a while, I do have to type in SSH and get into an old school server. The last startup that I was just at, we were running, we had actually inherited a Symphony app that was running on PHP.

So I did have to actually manage, manage a few clusters of servers, and have all the load balancing and all of the scale, the scale and groups and auto-scaling and things like that. And so I am not unfamiliar with all of that stuff, but if I do it for my own personal amusement, or I guess from my own personal suffering, I would not do that on purpose anymore. And, and that would be my advice to anybody is you maybe want to learn that stuff. And you probably do, if you want to get, you know, some of your AWS certs, but, but from a practical standpoint, I would not, I would not do that myself if, unless the absolute need arise or arose, I guess.

Rebecca: Yeah. I want to say I'm like, that sounds like fun dot, dot, dot. All right. So those are the listener questions. And I, I want to get back to a few of my own because this is me interviewing you. Right. But it's a time to reflect. Right. And so I just, I'm curious in the spirit of reflection in this 100th episode, if you were just starting Serverless Chats today, or just starting your newsletter or your blog or your conference talks, what advice would you offer to your own self?

Jeremy: Oh, well, I, one thing I have to absolutely call out because I don't, I certainly don't want people to think that these things just magically happen. I have a team that helps me with these now, when I first started, it was just me. I was writing the newsletter. I was reading the newsletter, I was copy editing. I was doing all of that stuff myself, but now I have two absolutely amazing people that help me out. Angela Milinazzo basically does a lot of social media stuff. She runs the newsletter for Serverless Chats Insiders, and she does all the guests reach out and are reaching out to the guests and coordination with the guests and so forth.

She does a lot of the marketing stuff and, and just so many things that can get taken off of my plate so that I can focus more on again, having the conversations and, and doing some of the more, hopefully important, the undifferentiated heavy-lifting stuff, I guess, that I don't have to do with, with some of those, some of the things that Angela does for me, which is amazing.

So thank you. And I've actually, I've been working with Angela for, I want to say like maybe eight or nine years, eight years now, something like that. And she worked at the same startup together. And then, and then we went separate ways is that we left that startup, but then we came back together to do this. So that was super exciting. And then also Melissa De orenzo, who was a time friend of mine, but all, she actually worked at my web development company with me, but she does copy editing for me, writes, helps with the research for the Serverless Stars and, and just copy edits the newsletter. And does all these other things, she does my accounting for me for all the stuff that we do. So I would not be able to do this stuff without having a really amazing team behind me.

And I think that is one of the biggest pieces of advice I would give anybody is teamwork. Like there is just ... I mean, I've been doing this if we want to, if we want to consider 1997 the sort of the start date of my career around this stuff, that's 24 years or so that I've been doing this. And I think any person is the sum of their experience. Like you just ... some things you remember, some things you forget, but whatever you experienced you're that's going to shape who you are. It's going to shape the way that you think. But you cannot ... there're just experiences that I will never have.

There are, there are backgrounds or situations that I will never experience. There are thoughts that I will never have because of who I am because of where I grew up, because of how, where I went to school, because of the experiences that I've had or whatever, and without having diverse teammates to sort of help you see things from a different angle, not only to help you, but I mean, not like help you, actually help you to work and accomplish things, but without having people around you that differ in ... whether it's in the slightest way or in a massive way ... without that different level of perspective you can't grow.

And so, I mean, that is one of the greatest things that I have ... I think the greatest thing for me, having gotten the chance to not only interview all these people for Serverless Chats, not only to read all of the posts that I've shared on Off-by-none and all of the articles there but to get to go to the conferences, to get, to give these conference talks, to have people question things that I've written in my blog posts. I should call out something really important here, and thank him for it. I wrote a blog post about serverless security, one of the, not my main post, but another one that was about like a SQL injection thing. And, and I used some language in there that was sort of general and sort of speculative, right. And Chris Munns actually called me out on that because he, and he actually kind of had to do a little tweetstorm.

I think he got some word from up above that he had to sort of call out, and, and make sure that he corrected the record on these things. And, and while I wasn't being critical of it, I was just being, I think I used some language and some general generalities that, that weren't accurate or did that maybe were misplaced in a way. And I don't think I did anything wrong other than the fact that I was too general and I wasn't clear enough. And when I got that feedback from him, I wasn't offended by it. Right. And that's another important thing. Like you've got to learn to take criticism, like absolutely, whether it's constructive or not, you need to learn to take it. But I remember getting that criticism from him and thinking to myself, I'm like, no, he's absolutely right.

Like, I can see how people could misinterpret what I said. And again, once you get a voice and once people start reading what you're writing, you have a responsibility and you absolutely have to just make sure. And that's one of the things that I did. I mean, sort of from that point forward, every post that I've ever written is highly researched. Right. And if I don't know something, I usually will say, I'm not a hundred percent sure about this. You know, I could be wrong about this or whatever. So I will try to call those things out if I'm not a hundred percent sure, but you can make, you can make a lot of assumptions when you're writing blog posts, because sometimes it's just easier to fill in a sentence here and there by adding some bit of flair to it or whatever, that will make an assumption.

And while it might seem right at the time those words can have an effect on somebody. And that's why I spend so much time now when I do write blog posts to try to heavily research those and make sure that the things I say are accurate and that I, again, I'm not using terms like "simple" and "easy" and things like that. 'Cause that's another thing what is simple to me or what is easy to me or what seems easy to me, or what seems easy to you or whatever could be wildly different from somebody else's perception. And that's another thing is just, I guess if we're moving on with the advice here is, know your audience again, going back to Jeff Barr, he just has this way to explain things that bring it down to a level that don't make you feel dumb, right.

That you're like, or that your eyes don't gloss over because you have no idea what he's talking about, but at the same time, doesn't feel like he's talking down to you either. And that's a hard thing to learn. I don't think ... I don't read a lot of blog posts that don't do that. And that's a hard thing to learn. So if you can find that level of humility where you say, I may know something, and I may know something really, really well, how can I communicate that to people in a way that doesn't talk down to them, but at the same time is accessible. Right? 'Cause accessibility is, is a huge thing as well.

Yeah. What else? I mean, going back again, I think the diversity thing, I don't want to harp on it too much, but that is those one of those things where me as a person personally, from 2018, when I first started writing these blog posts to the person that I am now, and the way that I think about things, politically as well as intellectually, the way I think about technology, all these different things.

I am a different person and I think differently now I have evolved my thinking dramatically because I was able just for a moment, just for a few minutes for 45 minutes, for 30 minutes on a conversation, or for five minutes in the hallway at a conference I was able to, sort of feel or be empathetic to somebody else's predicament or their perspective or there, sort of their ... and get a tiny taste into that background, that insight and so forth. And that has dramatically changed who I am. And I mean, I'm pretty happy with where I am now. I still think I've got a lot of work to do on myself in terms of continuing to open my mind, but I've met more people than I can ...

I don't remember all their names, unfortunately, but I have met so many people. And that is just, that is one of those things where I guess ... so my advice here is whatever you do if you're speaking at conferences or you're writing blog posts, or you're doing open source projects, or you're doing a podcast or whatever, open your channels up for feedback and talk to people. And, and if somebody has a problem with something, you say like, again, there are trolls, and then there are legitimate people who have concerns, or you know, who want to talk to you about something. And if you can take that criticism and you can open yourself up to that stuff, then I think, you just make yourself a better person. And yeah. And then the other thing, I guess I'll just say is, it's fine to make assumptions about products and find, to make assumptions about building things and software and whatever, never make assumptions about people.

Just the thing is, is that, you might read something somebody wrote and it might be wildly inaccurate. It might be way off base from your own personal thinking. And my biggest, I guess my, the advice that I give to a lot of people, especially younger people, is over the course of your career, you will hear a lot of good ideas and you'll hear a lot of bad ideas and you will not know the difference. It takes a long time to start understanding.

Most times when you hear something, you have no context, you don't have the context. You need to understand whether that's good or it's bad. Sometimes you do, but most times you don't. So just take everything with a grain of salt and know that the best thing you can do is continue to learn and keep an open mind and get as much context as you can. And hopefully, that turns into a better person and a better member of the community.

Rebecca: So I asked you for some serverless advice, but I think that's just really some great life advice. Most of those can be applied to a broader ...

Jeremy: Did I go off there?

Rebecca: No, I mean, thank you. I think let's, yeah, it's more than serverless, right? It's like the power of building a great team, the power of being able to receive critiques, and then in such a way where you want to improve your work and the way that you want to help others share it the way that you want to remain open, to understanding where your own blind spots are. Yeah. I'm a little floored, but yeah.

Thank you for sharing that with me. I think a huge, thank you, of course, to Angela and Melissa and your team, a huge thank you to Rob and Matt and Ken who submitted their questions that we asked on the air, and then a huge thank you to all of your listeners and to the serverless community. And so on this day, this very special day of your hundredth Serverless Chats episode. What do you hope remains with them, with your team, with your listeners, with the serverless community? What do you hope with them in their mind's ear as they drift off to sleep this evening after hearing you as your guest on Serverless Chats?

Jeremy: Oh man. Well, I mean, again, I think it's just one of those things where I'm wired in a way where I do what I do, I think, because I like to help people. And I think there are a lot of people who you just don't, you just don't know like, that's why I say don't make assumptions about people because you don't, you don't know, you have no idea what that person has gone through or what they're going through, or what's behind a smile or a frown, or whatever. You just, you don't know. And, and there's so much that, you know, human potential is amazing. And if you limit the possibility of human potential because you judge too soon, then I think that you do everything, everyone that disservice, including yourself and including, I guess, the evolution of mankind, which is something that is sort of passionate, that I'm passionate about.

So I guess the advice that I would leave is just look, don't make assumptions, be good to people, right? Like just, we're all in this together. So sort of like, I don't know, just go forth, be good, treat people, treat people well and build with serverless because that is definitely the future.

Rebecca: All right. I promise you as I drift off to sleep tonight, I will definitely think, be good to people, treat others well, build with serverless. Yeah. I love it. That's, it's three easy sentences. Good to remember, Jeremy, thank you so much for being our guest, but really your guest on Serverless Chats. And thank you for the honor of letting me be the asker. I can't wait to see you at episode 101. And I can't wait to see, did we, did I hear @vernor is going to be episode 101?

Jeremy: I don't know about 101, but I mean, there's plenty of episodes between now and 200, so.

Rebecca: Well, I can't wait. Is there anything else you want to leave us with?

Jeremy: No, I just, I want to thank you, Rebecca. I mean, you have been, you've been, I think another sort of what's the right word for it. Like you've been instrumental in me being able to do this, right. I mean, one of the things I had said earlier was, validation is extremely important. It's just what humans crave: validation. And it's not so much the notoriety or the popularity or any of those things that matter to me, what matters to me is that, that I help people. And when you put a lot of time and energy into something to try to help somebody, or you're thinking you try to help somebody, and it just doesn't get amplified. It is, it can be really frustrating. And, and again, like I said, there are so many ideas out there.

There's so many ideas out there that you do not know, that I do not know, that I've never heard before, they're perspectives that I've never heard before. And if you can't get to those, because you know, that person has five followers on Twitter, that's a shame, that's a horrible shame. And so I guess the ... what I want to say is that like you and the Heroes Program at AWS and, and the help that you gave me, I mean, just with, with dealing with like trying to do sponsorships and these other things that you, I mean, you helped coordinate some guests for me and like reach out to some people like that.

That's the kind of help, that's the type of thing where that, that feedback from you, that validation from you, the validation from AWS, that recognition that what I was doing was, was helping people and, and hopefully moving that needle, those are the kinds of things that people need. That was something that I needed. And I don't think I would, I don't think I'd be able to be doing what I'm doing now if I didn't have that encouragement and that support from you and other members of the community.

Rebecca: Well, I have no doubt that you would have achieved it regardless, but I am happy that I got to be the person to help you do so. And I think we share that ethos around and enthusiasm around the power of bringing community members together and being able to share their ideas and perspectives excitedly. So on the same plane here, I love it. Yeah.

Jeremy: Awesome. Awesome. Well, thank you, Rebecca. I appreciate it.

Rebecca: Yeah. Can't wait to see episode 100, happy episode 100 and, or, you know, hear it, I should say, 'cause I've seen it now this whole time, and then can't wait to, can't wait to see the next 100 episodes, really looking forward to it.

Jeremy: Awesome.

View Details

About Rob Sutter

Rob Sutter, a Principal Developer Advocate at Fauna, has woven application development into his entire career, from time in the U.S. Army and U.S. Government to stints with the Big Four and Amazon Web Services. He has started his own company – twice – once providing consulting services and most recently with WorkFone, a software as a service startup that provided virtual digital identities to government clients. Rob loves to build in public with cloud architectures, Node.js or Go, and all things serverless!

Twitter: @rts_rob
Personal email: rob@fauna.com
Personal website: robsutter.com
Fauna Homepage
Learn more about Fauna
Supported Languages and Frameworks
Try Fauna for FreeThe Calvin Paper

This episode sponsored by CBT Nuggets: https://www.cbtnuggets.com/

Watch this video on YouTube: https://youtu.be/CUx1KMJCbvk

Transcript
Jeremy: Hi, everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm joined by Rob Sutter. Hey, Rob. Thanks for joining me.

Rob: Hey, Jeremy. Thanks for having me.

Jeremy: So you are now the or a Principal Developer Advocate at Fauna. So I'd love it if you could tell the listeners a little bit about your background and what Fauna is all about.

Rob: Right. So as you've said, I'm a DA at Fauna. I've been a serverless user in production since 2017, started the Serverless User Group in Dubai. So that's how I got into serverless in general. Previously, I was a DA on the Serverless Team at AWS, and I've been a SaaS startup co-founder, a government employee, an IT auditor, and like a lot of serverless people I find, I have a lot of Ops in my background, which is why I don't want to do it anymore. There's a lot of us that end up here that way, I think. Fauna is the data API for modern applications. So it's a database that you access as an API just as you would access Stripe for payments, Twilio for messaging. You just put your data into Fauna and access it that way. It's flexible, serverless. It's transactional. So it's a distributed database with asset transactions, and again, it's as simple as accessing any other API that you're already accessing as a developer so that you can simplify your code and ship faster.

Jeremy: Awesome. All right. Well, so I want to talk more about Fauna, but I want to talk about it actually in a broader ... I think in the broader ecosystem of what's happening in the cloud right now, and we hear this term "multicloud" all the time. By the way, I'm super excited to have you on. I wanted to have you on for the longest time, and then just schedules, and it's like ...

Rob: Yeah.

Jeremy: You know how it is, but anyways.

Rob: Thank you.

Jeremy: No. But seriously, I'm super excited because your tweets, and everything that you've written, and the things that you were doing at AWS and things like that I think just all reinforced this idea that we are living in this multicloud world, right, and that when people think of multicloud ... and this is something I try to be very clear on. Multicloud is not cloud-agnostic, right?

Rob: Right.

Jeremy: It's a very different thing, right? We're not talking about running the same work load in parallel on multiple service providers or whatever.

Rob: Right.

Jeremy: We're talking about this idea of using the best services that are available to you across the spectrum of providers, whether those are cloud service providers, whether those are SaaS companies, or even to some degree, some open-source projects that are out there that make up this strategy. So let's start there right from the beginning. Just give me your thoughts on this idea of what multicloud is.

Rob: Right. Well, it's sort of a dirty word today, and people like to rail against it. I think rightly so because it's that multicloud 1.0, the idea of, as you said, cloud-agnostic that "write once, run everywhere." All that is, is a race to the bottom, right? It's the lowest common denominator. It's, "What do I have available on every cloud service provider? And then let me write for that as a risk management strategy." That's a cost center when you want to put it in business terms.

Jeremy: Right.

Rob: Right? You're not generating any value there. You're managing risk by investing against that. In contrast, what you and I are talking about today is this idea of, "Let me use best in class everywhere," and that's a value generation strategy. This cloud service provider offers something that this team understands, and wants to build with, and creates value for the customer more quickly. So they're going to write on that cloud service provider. This team over here has different needs, different customers. Let them write over there. Quite frankly, a lot of this is already happening today at medium businesses and enterprises. It's just not called multicloud, right?

Jeremy: Right.

Rob: So it's this bottom-up approach that individual teams are consuming according to their needs to create the greatest value for customers, and that's what I like to see, and that's what I like to promote.

Jeremy: Yeah, yeah. I love that idea of bottom-up because I think that is absolutely true, and I don't think you've seen this as aggressively as you have in the last probably five years as more SaaS companies have become or SaaS has become a household name. I mean, probably 10 years ago, I think Salesforce was around, and some of these other things were around, right?

Rob: Yeah.

Jeremy: But they just weren't ... They weren't the household names that they are now. Now, you watch any sport, any professional sport, and you see advertisements for all these SaaS companies now because that seems to be the modern economy. But the idea of the bottom-up approach, that is something where you basically give a developer or maybe you don't give them, but the developer takes the liberties, I would say, to maybe try and experiment with something new without having to do years of research, go through procurement, get approval to use some platform. Even companies trying to move to AWS, or on to Azure, or something like that, they have to go through hoops. Basically, jump through hoops in order to get them there. So this idea of the bottom-up approach, the developers are the ones who are experimenting. Very low-risk experiments, by the way, with some of these other services. That approach, that seems like the right marketing approach for companies that are building these services, right?

Rob: Yeah. It seems like a powerful approach for them. Maybe not necessarily the only one, but it is a good one. I mean, there's a historical lesson here as well, right? I want to come back to your point about the developers after, but I think of this as shadow cloud. Right? You saw this with the early days of SaaS where people would go out and sign up for accounts for their business and use them. They weren't necessarily regulated, but we saw even before that with shadow IT, right, where people were bringing their own software in?

Jeremy: Right.

Rob: So for enterprises that are afraid of this that are heavily risk-focused or control-focused top-down, I would say don't be so afraid because there's an entire set of lessons you can learn about this as you bring it, as you come forward to it. Then, with the developers, I think it's even more powerful than the way you put it because a lot of times, it's not an experiment. I mean, you've seen the same things on Twitter I've seen, the great tech turnover of 2021, right? That's normal for tech. There's such a turnover that a lot of times, people are coming in already having the skills that they know will enhance delivery and add customer value more quickly. So it's not even an experiment. They already have the evidence, and they're able to get their team skilled up and building quickly. If you hire someone who's coming from an AWS shop, you hire someone who's coming from an Azure shop on to two different teams, they're likely going to evolve that excellence or that capability independently, and I don't necessarily think there's a reason to stop that as long as you have the right controls around it.

Jeremy: Right. I mean, and you mentioned controls, and I think that if I'm the CTO of some company or whatever, or I'm the CIO because we're dealing in a super enterprise-y world, that you have developers that are starting to use tools ... Maybe not Stripe, but maybe like a Twilio or maybe they're using, I don't know, ChaosSearch or something, something where data that is from within their corporate walls are going out somewhere or being stored somewhere else, like the security risk around that. I mean, there's something there though, right?

Rob: Yeah, there absolutely is. I think it's incumbent on the organizations to understand what's going on and adapt. I think it's also imcu,bent on the cloud service providers to understand those organizational concerns and evolve their product to address them, right?

Jeremy: Right.

Rob: This is one thing. My classic example of this is data exfiltration in a Lambda function. Some companies get ... I want to be able to inspect every packet that leaves, and they have that hard requirement for reasons, right?

Jeremy: Right.

Rob: You can't argue with them that they're right or wrong. They made that decision as a company. But then, they have to understand the impact of that is, "Okay. Well, every single Lambda function that you ever create is going to run inside of VPC or is going to run connected to a VPC." Now, you have the added complexity of managing a VPC, managing your firewall rules, your NACLs, your security groups. All of this stuff that ... Maybe you still have to do it. Maybe it really is a requirement. But if you examine your requirements from a business perspective and say, "Okay. There's another way we can address this with tightly-scoped IAM permissions that only allow me to read certain records or from certain tables, or access certain keys, or whatever." Then, we assume all that traffic goes out and that's okay. Then, you get to throw all of that complexity away and get back to delivering value more quickly. So they have to meet together, right? They have to meet.

Jeremy: Right.

Rob: This led to a lot of the work that AWS did with VPC networking for Lambda functions or removing the cold start because a lot ... Those enterprises weren't ready to let go of that requirement, and AWS can't tell them, "You're wrong." It's their business. It's their requirement. So AWS built that for them to reduce the cold start so that Lambda became a viable platform for them to build against.

Jeremy: Right, and so if you're a developer and you're looking at some solution because ... By the way, I mean, like you said, choosing the best of breed. There are just a lot of good services out there. There are thousands and thousands of SaaS companies, and I think ... I don't know if we made this point, but I certainly consider SaaS companies themselves to be part of the cloud. I would think you would probably agree with that, right?

Rob: Yeah.

Jeremy: It might as well be cloud providers themselves. Most of them run on top of the cloud providers anyways, but they found ...

Rob: But they don't have to, and that's interesting to me and another truth that you could be consuming services from somebody else's data center and that's still multicloud.

Jeremy: Right, right. So, anyway. So my thought here or I guess the question I have is if you're a developer and you're trying to choose something best in breed, right? Whatever that is. Let's say I'm trying to send text messages, and I just think Twilio is ... It's got all the features that I need. I want to go with Twilio. If you're a developer, what are the things that you need to look for in some of these companies that maybe don't have ... I mean, I would say Twilio does, but like don't necessarily have the trust or the years of experience or I guess years under their belts of providing these services, and keeping data secure, and things like that. What's the advice that you give to developers looking to choose something like that to be aware of?

Rob: To developers in particular I think is a different answer ...

Jeremy: Well, I mean, yeah. Well, answer it both ways then.

Rob: Yeah, because there's the builder and the buyer.

Jeremy: Right.

Rob: Right?

Jeremy: Right.

Rob: Whoever the buyer is, and a lot of times, that could just be the software development manager who's the buyer, and they still would approach it different ways. I think the developer is first going to be concerned with, "Does it solve my problem?" Right? "Overall, does it allow me to ship faster?" The next thing has to be stability. You have to expect that this company will be around, which means there is a certain level of evidence that you need of, "Okay. This company has been around and has serviced," and that's a bit of a chicken and an egg problem.

Jeremy: Right.

Rob: I think the developer is going to be a lot less concerned with that and more concerned with the immediacy of the problem that they're facing. The buyer, whether that's a manager, or CIO, or anywhere in between, is going to need to be concerned with other things according to their size, right? You even get the weird multicloud corner cases of, "Well, we're a major supplier of Walmart, and this tool only runs on a certain cloud service provider that they don't want us to use. So we're not going to use it." Again, that's a business decision, like would I build my software that way? No, but I'm not subject to that constraint. So that means nothing in that equation.

Jeremy: Right. So you mentioned a little bit earlier this idea of bringing people in from different organizations like somebody comes in and they can pick up where somebody else left off. One of the things that I've noticed quite a bit in some of the companies that I've worked with is that they like to build custom tools. They build custom tools to solve a job, right? That's great. But as soon as Fred or Sarah leave, right, then all of a sudden, it's like, "Well, who takes over this project?" That's one of the things where I mentioned ... I said "experiments," and I said "low-risk." I think something that's probably more low-risk than building your own thing is choosing an API or a service that solves your problem because then, there's likely someone else who knows that API or that service that can come in, and can replace it, and then can have that seamless transition.

And as you said, with all the turnover that's been happening lately, it's probably a good thing if you have some backup, and even if you don't necessarily have that person, if you have a custom system built in-house, there is no one that can support that. But if you have a custom ... or if you have a system you've used, you're interfacing with Twilio, or Stripe, or whatever it is, you can find a lot of developers who could come in even as consultants and continue to maintain and solve your problems for you.

Rob: Yeah, and it's not just those external providers. It's the internal tooling as well.

Jeremy: Right.

Rob: Right?

Jeremy: Right.

Rob: We're guilty of this in my company. We wrote a lot of stuff. Everybody is, right, like you like to do it?

Jeremy: Right.

Rob: It's a problem that you recognized. It feels good to solve it. It's a quick win, and it's almost always the wrong answer. But when you get into things like ... a lot of cases it doesn't matter what specific tool you use. 10 years ago, if you had chosen Puppet, or Chef, or Ansible, it wouldn't be as important which one as the fact that you chose one of those so that you could then go out and find someone who knew it. Today, of course, we've got Pulumi, Terraform, and all these other things that you could choose from, and it's just better than writing a bunch of Bash scripts to tile the stuff together. I believe Bash should more or less be banned in the cloud, but that's another ... That's my hot take for another time. Come at me on Twitter if you don't like that one.

Jeremy: So, yeah. So I think just bringing up this idea of tooling is important because the other thing that you potentially run into is with the variety of tooling that's out there, and you mentioned the original IAC. I guess they would... Right? We call those like Ansible and those sort of things, right?

Rob: Right.

Jeremy: All of those things, the Chefs and the Puppets. Those were great because you could have repeatable deployments and things like that. I mean, there's still work to be done there, but that was great because you made the choice to not building something yourself, right?

Rob: Right.

Jeremy: Something that somebody else could jump in on. But now with Terraform and with ... You mentioned Pulumi. We've got CloudFormation and even Microsoft has their own ... I think it's called ARM or something like that that is infrastructure as code. We've got the Serverless Framework. We've got SAM. We've got Begin. You've got ... or Architect, right? You've got all of these choices, and I think what happens too is that if teams don't necessarily ... If they don't rally around a tool, then you end up with a bunch of people using a bunch of different tools. Maybe not all these tools are going to be compatible, but I've seen really interesting mixes of people using Terraform, and CloudFormation, and SAM, and Serverless Framework, like binding it all together, and I think that just becomes ... I think that becomes a huge mess.

Rob: It does, and I get back to my favorite quote about complexity, right? "Simplicity before complexity is worthless. Simplicity beyond complexity is priceless." I find it hard to get to one tool that's like artificial, premature optimization, fake simplicity.

Jeremy: Yeah.

Rob: If you force yourself into one tool all the time, then you're going to find it doing what it wasn't built to do. A good example of this, you talked about Terraform and the Serverless Framework. My opinion, they weren't great together, but your Terraform comes for your persistent infrastructure, and your Serverless Framework comes in and consumes the outputs of those Terraform stacks, but then builds the constantly churning infrastructure pieces of it. Right? There's a blast radius issue there as well where you can't take down your database, or S3 bucket, or all of this from a bad deploy when all of that is done in Terraform either by your team, or by another team, or by another process. Right? So there's a certain irreducible complexity that we get to, but you don't want to have duplication of effort with multiple tools, right?

Jeremy: Right.

Rob: You don't want to use CloudFormation to manage your persistent data over here and Terraform to manage your persistent data over here because then you're not ... That's like that agnostic model where you're not benefiting from the excellent features in each. You're only using whatever is common between them.

Jeremy: Right, right, and I totally agree with you. I do like the idea of consuming. I mean, I have been using AWS for a very, very long time like 2007, 2008.

Rob: Yeah, same. Oh, yeah.

Jeremy: Right when EC2 instances were a thing. I guess 2008. But the biggest thing for me looking at using Terraform or something like that, I always felt like keeping it in, keeping it in the family. That's the wrong way to say it, but like using CloudFormation made a lot of sense, right, because I knew that CloudFormation ... or I thought I knew that CloudFormation would always support the services that needed to be built, and that was one of my big complaints about it. It was like you had this delay between ... They would release some service, and you had to either do it through the CLI or through the console. But then, CloudFormation support came months later. The problem that you have with some of that was then again other tools that were generating CloudFormation, like a Serverless Framework, that they would have to wait to get CloudFormation support before they could support it, and that would be another delay or they'd have to build something custom, which is not always the cleanest way to do it.

Rob: Right.

Jeremy: So anyways, I've always felt like the CloudFormation route was great if you could get to that CloudFormation, but things have happened with CDK. We didn't even mention CDK, but CDK, and Pulumi, and Terraform, and all of these other things, they've all provided these different ways to do things. But the thing that I always thought was funny was, and this is ... and maybe you have some insight into this if you can share it, but with SAM, for example, SAM wasn't extensible, right? You would just run into issues where you're like, "Oh, I can't do that with SAM." Whereas the Serverless Framework had this really great third-party plug-in system that allowed you to do some of these other things. Now, granted not all third-party plug-ins were super stable and were the best way to do something, right, because they'd either interact with APIs directly or whatever, but at least it gave you ... It unblocked you. Whereas I felt like with SAM and even CloudFormation when it didn't support something would block you.

Rob: Yeah. Yeah, and those are just two different implementation philosophies from two different companies at two different stages of their existence, right? Like AWS ... and let's separate the reality from the theory here. The theory is that a large company can exert control over release cycles and limit what it delivers, but deliver it with a bar of excellence. A small company can open things up, and it depends on its community members for contributions to solve problems. It's very much like this is the cathedral and the Bazaar of cloud tooling, right?

AWS has that CloudFormation architecture that they're working around with its own goals and approach, the one way to do it. Serverless Framework is, "Look, you need to ... You want to set up a stall here and insert IAM policies per function. Set up a stall. It will be great. Maybe people come and maybe they don't," and the system inherently sorts or bubbles up the value, right? So you see things like the Step Functions plug-in for Serverless Framework. It was one of the early ones that became very popular very quickly, whereas Step Functions supporting SAM trailed, but eventually came in. I think that team, by the way, deserves a lot of credit for really being focused on developers, but that's not the point of the difference between the two.

A small young company like Serverless Framework that is moving very quickly can't have that cathedral approach to it, and both are valid, right? They're both just different strategies and good for the marketplace, quite frankly. I have my preferred approach, which is not about AWS or SAM vs Serverless Framework. It's the extensibility of plug-in frameworks to me are a key component of tooling that adapts as quickly as the clouds change, and you see this. Like Terraform was the first place that I really learned about plug-ins, and their plug-in framework is fantastic, the way they do providers. Serverless Framework as well is another good example, but you can't know how developers are going to build with your services. You just can't.

You do customer development. You talk to them ahead of time. You get all this research. You talk to a thousand customers, and then you release it to 14 million customers, right? You're never going to guess, so let them. Let them build it, and if people ... They put the work in. People find there's value in it. Sometimes you can bring it in. Sometimes you leave it up to the community to maintain, but you just ... You have to be willing to accept that customers are going to use your product in different ways than you envisioned, and that's a good thing because it means customers are using your product.

Jeremy: Right, right. Yeah. So I mean, from your perspective though ... because let's talk about SAM for a minute because I was excited when SAM came out. I was thinking to myself. I'm like, "All right. A simplified tooling that is focused on serverless. Right? Like gives me all the things that I think I'm going to need." And then I did ... from a developer experience standpoint, and let's call out the elephant in the room. AWS and developer experience are not always the same. They don't always give you that developer experience that you would want. They give you tons of tools, right, but they don't always give you that ...

Rob: Funny enough, you can spell "developer experience" without AWS.

Jeremy: Right. So I mean, that's my ... I was disappointed when I started using SAM, and I immediately reverted back to the Serverless Framework. Not because I thought that it wasn't good or that it wasn't well-thought-out. Like you said, there was a level of excellence there, which certainly you cannot diminish, It just didn't do the things I needed it to do, and I'm just curious if that was a consistent feedback that you got as being someone on the dev advocate team there. Was that something that you felt as well?

Rob: I need to give two answers to this to be fair, to be honest. That was something that I felt as well. I never got as comfortable with SAM as I am with the Serverless Framework, but there's another side to this coin, and that's that enterprise uptake of SAM CLI has been really strong.

Jeremy: Right.

Rob: Enterprise it, it does what they need it to do, and it addresses their concerns, and they liked getting tooling from AWS. It just goes back to there being a place for both, right?

Jeremy: Right.

Rob: Enterprises are much more likely to build cathedrals. They want that top-down, "Okay, everybody. This is how you define something. In fact, we've created a module for you. Consume it here. Thou shalt not write new S3 to web server configuration in your SAM templates. Thou shalt consume this." That's not wrong, and the usage numbers don't lie with SAM. It's got a lot of fans, and it's got a lot of uptake, but that's an entirely different answer from how I feel about it. I think it also goes back to I'm not running an enterprise. I've never run an enterprise. The biggest I've got in terms of responsibility is at best a small company, right? So I think it's natural for me to feel that way when I try to use a tool that has such popularity amongst enterprise. Now, of course, you have the switch, right? You have enterprises using Serverless Framework, and you have small builders using SAM. But in general, I think the success there was with the enterprise, and it's a validation of their strategy.

Jeremy: Right, right. So let's talk about enterprises for a second because this is where we look at tools like the CDK and SAM, Serverless Framework, things like that. You look at all those different tools, and like you said, there's adoption across some of those. But at the end of the day, most of those tools are compiling down to a CloudFormation or they're compiling down to ... What's it called? The Azure Resource Manager Language or whatever the heck it is, right?

Rob: ARM templates.

Jeremy: ARM templates. What's the value now in CloudFormation and those sort of things that the final product that you get to ... I mean, certainly, it's so much easier to build those using these frameworks, but do we need CloudFormation in those things anymore? Do we need to know those? Does an individual developer need to be able to understand those, or can they just basically take a step back and say, "Look, CDK does it for me," or, "Pulumi does it for me. So why do I need to know what's baked into those templates?"

Rob: Yeah. So let's set Terraform aside and talk about it after because it's different. I think the choice of JSON and YAML as implementation languages for most of this tooling, most of this tooling is very ... It was a very effective choice because you don't necessarily have to know CloudFormation to look at a template and define what it's doing.

Jeremy: Right.

Rob: Right? You don't have to understand transforms. You don't have to understand parameter replacement and all of this stuff to look at the final transformed template in CloudFormation in the console and get a very quick reasoning about what's happening. That's good. Do I think there's value in learning to create multi-thousand-line CloudFormation templates by hand? I don't. It's the assembly language of the cloud, right?

Jeremy: Right.

Rob: It's there when you need it, and just like with procedural languages, you might want to look underneath at the instructions, how it unrolled certain loops, how it decided not to unroll others so that you can make changes at the next level. But I think that's rare, and that's optimization. In terms of getting things done and getting things shipped and delivered, to start, I wouldn't start with plain CloudFormation for any of these, especially of anything for any meaningful production size. That's not a criticism of CloudFormation. It's just like you said. It's all these other tooling is there to help us generate what we want consistently.

The other benefit of it is once you have that underlying lingua franca of the cloud, you can build visualization, and debugging, and monitoring, and like ... I mean, all of these other evaluatory forensic. "Evaluatory?" Is that a word? It's a word now. You heard it here first on this podcast. Like forensic, cloud forensic type tooling that lets you see what's going on because it is a universal language among all of the tools.

Jeremy: Yeah, and I want to get back to Terraform because I know you mentioned that, but I also want to be clear. I don't suggest you write CloudFormation. I think it is horribly verbose, but probably needs to be, right?

Rob: Yeah.

Jeremy: It probably needs to have that level of fidelity there or that just has all that descriptive information. Yeah. I would not suggest. I'm with you, don't suggest that people choose that as their way to go. I'm just wondering if it's one of those things where we don't need to be able to look at ones and zeroes anymore, and understand what they do, right?

Rob: Right.

Jeremy: We've got higher-level constructs that do that for us. I wouldn't quite put ... I get the assembly language comparison. I think that's a good comparison, but it's just that if you are an enterprise, right, is that ... Do you trust? Do you trust something like CDK to do everything you need it to do and make sure that it's covering all its bases because, again, you're writing layers of abstraction on top of a layer of abstraction that then abstracts it even more. So yeah, I'm just wondering. You had mentioned forensic tools. I think there's value there in being able to understand exactly what gets put into the cloud.

Rob: Yeah, and I'd agree with that. It takes 15 seconds to run into the limit of my knowledge on CDK, by the way. But that knowledge includes the fact that CDK Synth is there, which generates the CloudFormation template for you, and that's actually the thing, which is uploaded to the CloudFormation service and executed. You'd have to bring in somebody like Richard Boyd or someone to talk about the guard rails that are there around it. I know they exist. I don't know what they are. It's wildly popular, and adoption is through the roof. So again, whether I philosophically think it's a good idea or not is irrelevant. Developers want it and want to build with it, right?

Jeremy: Yeah.

Rob: It's a Bazaar-type tool where they give you some basic constructs, and you can write your own constructs around it and get whatever you need. But ultimately, that comes back to CloudFormation, which is then subject to all the controls that your organization puts around CloudFormation, so it is ... There's value there. It can't be denied.

Jeremy: Right. No, and the thing that I like about the CDK is the idea of being able to create those constructs because I think, especially from a ... What's the right word? Compliance standpoint or something like that, that you can write in these constructs that you say, "You need to use these constructs when you deploy a microservice or you deploy this," or whatever it is, and then you have those guard rails as you mentioned or whatever, but those ... All of those checkboxes are ticked because you can put that all into one construct. So I totally think that's great. All right. So let's talk about Terraform.

Rob: Yeah, so there's ... First, it's a completely different model, right?

Jeremy: Right.

Rob: This is an interesting discussion to have because it's API calls. You write your provider, whatever your infrastructure is, and anything that can be an API call can now be a Terraform declarative resource. So it's that mapping between declarative and imperative that I find fascinating. Well, also building the dependency graph for you. So it handles all of those aspects of it, which is a really powerful tool. The thing that they did so well ... Terraform is equally verbose as CloudFormation. You've got to configure all the same options. You get the same defaults, et cetera. It can be terribly verbose, but it's modular. Every Terraform file that you have in one directory is concatenated, and that is a huge distinction between how CloudFormation wants everything in one template, or well, you can refer to something in an S3 bucket, but that's not actually useful to me as a developer.

Jeremy: Right, right.

Rob: I can't mount an S3 bucket as a drive on my workstation and compose all of these independent files at once and do them that way. Sidebar here. Maybe I can. Maybe it supports that, and I haven't been able to discover it, right? Whereas Terraform by default, out of the box put everything in its own file according to function. It's very easy to look in your databases.tf and understand what's in there, look in your VPC.tf and understand what's in there, and not have to go through thousands of lines of code at once. Yes, we have find and replace. Yes, we have search, and you ... Anybody who's ever built any of this stuff knows that's not the same thing. It's not the same thing as being able to open a hundred lines in your text editor, and look at everything all at once, and gain an understanding of it, and then dive into the next level of detail a hundred lines at a time and understand that.

Jeremy: Right. But now, just a question here, without... because the API thing, I love that idea, and actually, Serverless Components used an API thing to do it and bypass CloudFormation. Actually, I believe Architect originally was using APIs, and then switched to CloudFormation, but the question I have about that is, if you don't have a CloudFormation template, if you don't have that assembly language of the web, and that's not sitting in your CloudFormation tool built into the dashboard, you don't get the drift protection, right, or the detection, and you don't get ... You don't have that resource map necessarily up there, right?

Rob: First, I don't think CloudFormation is the assembly language of the web. I think it's the assembly language of AWS.

Jeremy: I'm sorry. Right. Yes. Yeah. Yeah.

Rob: That leads into my point here, which is, "Okay. AWS gives you the CloudFormation dashboard, but what if you're now consuming things from Datadog, or from Fauna, or from other places that don't map this the same way?

Jeremy: Right.

Rob: Terraform actually does manage that. You can do a plan against your existing file, and it will go out and check the actual existing state of all of your resources, and compare them to what you've asked for declaratively, and show you what the changeset will be. If it's zero, there's no drift. If there is something, then there's either drift or you've added new functionality. Now, with Terraform Cloud, which I've only used at a basic level, I'm not sure how automatic that is or whether it provides that for you. If you're from HashiCorp and listening to this, I would love to learn more. Get in touch with me. Please tell me. But the tooling is there to do that, but it's there to do it across anything that can be treated as an API that has ... really just create and retrieve. You don't even necessarily need the update and delete functionality there.

Jeremy: Right, right. Yeah, and I certainly ... I am a fan of Terraform and all of these services that make it easier for you to build clouds more easily, but let's talk about APIs because you mentioned anything with an API. Well, everything has APIs now, right? I mean, essentially, we are living in a ... I mentioned SaaS, but now we're sort of ... This whole idea of the API economy. So just overall, what are your thoughts on this idea of basically anything you want to do is available now via an API?

Rob: It's not anything you want to do. It's everything you want to do short of fulfilling your customer's needs, short of solving your customer's specific problem is available as an API. That means that you get to choose best-in-class for everything, right? Your customer's need isn't, "I want to spend $25 on my credit card." Your customer's need is, "I need a book."

Jeremy: Right.

Rob: So it's not, "I want to store information about books in a database." It's, "I need a book." So everything and every step of the way there can now be consumed from an API. Again, it's like serverless in general, right? It allows you to focus purely or as close to purely as we can right now on generating customer value and addressing customer problems so that you can ship faster, and so that you have it as a competitive advantage. I can write a payment processing program. I know I can because I've done it back in 2004, and it was horrible, and it was awful. It wasn't a very good one, and it worked. It took your money, but this was like pre-PCIDSS.

If I had to comply with all of those things, why would I do that? I'm not a credit card payment processor. Stripe is, and they have specialists in all of the areas related to the problem of, "I need to take and process payments." That's the customer problem that they're solving. The specialization of labor that comes along with the API economy is fantastic. Ops never went away. All the ops people work at the cloud service providers now.

Jeremy: Right.

Rob: Right? Audit never went away. All the auditors have disappeared from view and gone into internal roles in payments companies. All of this continues to happen where the specialists are taking their deep, deep knowledge and bringing it inside companies that specialize in that domain.

Jeremy: Right, and I think the domain expertise value that you get from whatever it is, whether it's running a database company or whether it's running a payment company, the number of people that you would need to hire to have a level of specialization for what you're paying for two cents per transaction or whatever, $50 a month for some service, you couldn't even begin. The total cost of ownership on those things are ... It's not even a conversation you would want to have, but I also built a payment processing system, and I did have to pass PCI, which we did pass, but it was ...

Rob: Oh, good for you.

Jeremy: Let's put it this way. It was for a customer, and we lost money on that customer because we had to go through PCI compliance, but it was good. It was a good experience to have, and it's a good experience to have because now I know I never want to do it again.

Rob: Yeah, yeah. Back to my earlier point on Ops and serverless.

Jeremy: Right. Exactly.

Rob: These things are hard, right?

Jeremy: Right, right.

Rob: Sorry.

Jeremy: No, no. Go ahead.

Rob: Not to interrupt, but these are all really hard problems that people with graduate degrees and post-graduate research who have ... They're 30 years old when they start working on the problem are solving. There's a supply question there as well, right? There's just not enough people, and so you and I can like ... Well, I'm not going to project this on to you. I can stumble through an implementation and check off the requirements just like I worked in an optical microscopy lab in college, and I could create computer programs that modeled those concepts, but I was not an optical microscopist. I was not a PhD-level generating understandings of these things, and all of these, they're just so hard. Why would you do that when customer problems are equally hard and this set is solved?

Jeremy: Right, right.

Rob: This set of problems over here is solved, and you can't differentiate yourself by solving it better, and you're not likely to solve it better either. But even if you did, it wouldn't matter. This set of problems are completely unsolved. Why not just assemble the pieces from best-in-class so that you can attack those problems yourself?

Jeremy: Again, I think that makes a ton of sense. So speaking about expertise, let's talk about what you might have to pay say a database administrator. If you were to hire a database administrator to maintain all the databases for you and keep all that uptime, and maybe you have to hire six database administrators in order for them to ... Well, I'm thinking multi-region and all that kind of stuff. Maybe you have to hire a hundred, depending on what it is. I mean, I'm getting a little ahead of myself, but ... So if I can buy a service like Fauna, so tell me a little bit more about just how that works.

Rob: Right. Well, I mean, six database engineers in the US, you're over a million dollars a year easily, right?

Jeremy: Right.

Rob: I don't know what the exact number is, but when you consider benefits, and total cost, and all of that, it's a million dollars a year for six database engineers. Then, there are some very difficult problems in especially distributed databases and database scaling that Fauna solves. A number of other products or services solve some of them. I'm biased, of course, but I happen to think Fauna solves all of them in a way that no other product does, but you're looking ... You mentioned distributed transactions. Fauna is built atop the Calvin paper, which came out of Yale. It's a very brief, but dense academic research paper. It's a PHC research paper, and it talks about a model for distributed transactions and databases. It's a layer, a serialization layer, that sits atop your database.

So let's say you wanted to replicate something like Fauna. So not only do you need to get six database engineers who understand the underlying database, but you need to find engineers that understand this paper, understand the limitations of the theory in the paper, and how to overcome them in operations. In reality, what happens when you actually start running regions around the world, replicating transactions between those regions? Quite frankly, there's a level of sophistication there that most of the set of people who satisfy that criteria already work at Fauna. So there's not much of a supply there. Now, there are other database competitors that solve this problem in different ways, and most of the specialists already work at those companies as well, right? I'm not saying that they aren't equally competent database engineers. I'm just saying there's not a lot of them.

Jeremy: Right.

Rob: So if you're thing is to sell books at a certain scale like Amazon, then that makes sense to have them because you are a database creator as well. But if your thing is to sell books at some level below that, why would you compete for that talent rather than just consuming it?

Jeremy: Right. Yeah, and I would say unless you're a horse with a horn on your head, it's probably not worth maintaining your own database and things like that. So let's talk a little bit more, though, about that. I guess just this idea of maybe a shortage of people, like if you're ... You're right. There's a limited number of resources, right? I'm sure there's brilliant database engineers all around the world, and they have the experience where, right, they could come in and they could probably really help you maintain your database. Even if you could afford six of them and you wanted to do that, I think the problem is it's got to be the interestingness of the problem. I don't think "interestingness" is a word either, but like if I'm a database engineer, wouldn't I want to be working on something like Fauna that I could help millions and millions of people as opposed to helping some trucking company maintain their internal database or something like that?

Rob: Yeah, and I would hope so. I hope it's okay that I mention we're hiring. So come to Fauna.com and look at our roles database engineers.

Jeremy: You just read that Calvin paper first. Go ahead.

Rob: But read the Calvin paper first. I think it's only like 12 pages, and even just the first page is enough. I'm happy to talk about that at any length because I find it fascinating and it's public. It is an interesting problem and the ... It's the reification or the implementation of theory. It's bringing that theory to the real world and ... Okay. First off, the theory is brilliant. This is not to take away from it, but the theory is conceived inside someone's mind. They do some tests, they prove it, and there's a world of difference between that point, which is foundational, and deploying it to production where people are trusting their workloads on it around the world. You're actually replicating across multiple cloud providers, and you're actually replicating across multiple regions, and you're actually delivering on the promise of the paper.

What's described in the paper is not what we run at Fauna other than as a kernel, as a nugget, right, as the starting point or the first principle. That I think is wildly interesting for that class of talent like you talked about, the really world-class engineers who want to do something that can't be done anywhere else. I think one thing that Fauna did smartly early was be a remote-first company, which means that they can take advantage of those world-class engineers and their thirst for innovation regardless of wherever Fauna finds them. So that's a big deal, right? Would you rather work on a world-class or global problem or would you rather work on a local problem? Now, look, we need people working on local problems too. It's not to disparage that, but if this is your wheelhouse, if innovation is the thing that you want to do, if you want to be doing things with databases that nobody else is doing, this is where you want to be. So I think there's a strong argument there for coming to work in a place like Fauna.

Jeremy: Yeah, and I want to make sure I apologize to any database engineer working at a trucking company because I'm sure there are actually probably really interesting problems with logistics and things like that that they are solving. So maybe not the best example. Maybe. I don't know. I can't think of another example. I don't want to offend anybody who's chosen a more local problem because you're right. I mean, there are local problems that need to be solved, but I do think that there are people ... I mean, even someone like me. I want to work on a bigger problem. You know what I mean? I owned a web development company for 12 years, and I was solving other people's problems by building them a website or whatever, and it just got to a point where I'm like, "I'm not making enough of an impact here." You're not solving a big enough problem. You want to work on something more interesting.

Rob: Yeah. Humans crave challenge, right? Challenge is a necessary precondition for growth, and at least most of us, we want to grow. We want to be better at whatever it is we're doing or just however we think of ourselves next year that we aren't today, and you can't do that with challenge. If you build other people's websites for 12 years, eventually, you get to a point where maybe you're too good at it. Maybe that's great from a business perspective, but it's not so great from a personal fulfillment perspective.

Jeremy: Right.

Rob: If it's, "Oh, look, another brochure website. Okay. Here you go. Oh, you need a contact form?" Again, it's not to disparage this. It's the fact that if you do anything for 12 years, sometimes mastery is stasis. Not always.

Jeremy: Right, and I have nightmares of contact forms, of building contact forms, by the way, but...

Rob: It makes sense. Yeah. You know what you should do is just put all of those directly into Fauna and don't worry about it.

Jeremy: Easy enough. Easy enough.

Rob: Yeah, but it's not necessarily stasis, but I think about craftsmen and people who actually make things with their hands, physical builders, and I think a lot of that ... Like if you're making furniture, you're a cabinet maker. I think a lot of that is every time, it's just a little bit wrong, right? Not wrong, but just a little bit off from your optimum no matter how long you do it, and so everything has a chance to evolve. That's there with software to a certain extent, but the problem is never changing.

Jeremy: Right.

Rob: So, yeah. I can see both sides of it, but for me, I ... You can see it when I was on serverless four years ago and now that I'm on a serverless database now. I like to be out at the edge, pushing that edge out for everyone who's coming behind. It can be challenging because sometimes there's just no way forward, and sometimes everybody is not ready to come with you. In a lot of ways, being early is the same as being wrong.

Jeremy: Right. Well, I've been ...

Rob: Not an original statement, but ...

Jeremy: No, but I've been early on many things as well where like five years after we tried to do something, like then, all of a sudden, it was like this magical thing where everybody is doing it, but you mentioned the edge. That would be something ... or you said on the edge. I know you mean this way, but the edge in terms of the actual edge. That's going to be an interesting data problem to solve.

Rob: Oh, that's a fascinating data problem, especially for us at Fauna. Yeah, compute, and Andy Jassy, when he was at AWS, talked about how compute was bifurcating, right? It's either moving all the way out to the edge or it's moving all the way into the cloud, and that's true. But I think at Fauna, we take that a step further, right? The edge part is true, and a lot of the work that we've done recently, announcements with CloudFlare workers. We're ready for that. We believe in it, and we like pushing that out as close to the user as possible. One thing that we do uniquely is we have this concept of user-defined functions, and anybody who's written T-SQL back in the day, who wrote store procedures is going to be familiar with this, but it's ... You bring that business logic and that code to your data. Not near your data, to your data.

Jeremy: Right.

Rob: So you bring the compute not just to the cloud where it still needs to pass through top-of-rack and all of this. You bring it literally on to the same instance as your data where these functions execute against them there. So you get not just the database, but you get a compute layer in there, and this helps for things like filtering for things like the equivalent of joins, stuff that just ... If you've got to load gigabytes of data and move it somewhere, compute against it, reduce it to something, and then store that back, the speed of light still matters. Even if it's the speed of light across a couple switches, it still matters, and so there are some really interesting things that you can begin to do as you pull more and more of that logic into your data layer, and you also protect that logic from other components of your application.

So I like that because things like GraphQL that endpoints already speak and already understand, just send it over, and again, they don't care about the architectural, quite frankly, genius--I can say that because I didn't create it--the genius behind all of this stuff. They just care that, "Look, I send this request over and I get it back," and entire workflows, and complex processes, and everything are executing behind the scenes just so that the endpoints can send and retrieve what they need more effectively and more quickly. The edge is fascinating. The thing I regret the most about the edge is I have no hardware skills, right? So I can't make fun things to do fun things in my house. I have to buy them, but you can't do everything.

Jeremy: Yeah. Well, no. I think you make a good point though about bringing the compute to the data side, and other people have said there's no ... Ben Kehoe has been talking about this for a while too where like it just makes sense. Run the compute where the data is, and then send that data somewhere else, right, because there's more things that can be done with data after that initial bit of compute. But certainly, like you said, filtering it down or getting the bits that are relevant and moving a small amount of data as opposed to a large amount of data I think is hugely important.

Now, the other thing I just want to mention before I let you go or I want to talk about quickly is this idea of going back to the API economy aspect of things and buying versus building. If you think about what you've had to do at Fauna, and I know you're relatively recent there, but you know what they've done and the work that had to go in in order to build this distributed system. I mean, I think about most systems now, and I think like anything I'm going to build now, I got to think about scale, right?

I don't necessarily have to build to scale right away, especially if I'm doing an MVP or something like that. But if I was going to build a service that did something, I need to think about multi-region, and I need to think about failover, and I need to think about potentially providing it at the edge, and all these other things. So you come down to this thing, and I will just use the database example. But even if you were say using like MySQL, or Postgres, or something like that, that's going to scale. That's going to scale pretty well to get to a certain point, and then you're going to have to start sharding, right? When data gets hard, it's time to shard, right? You just have to start sharding everything.

Rob: Yeah.

Jeremy: Essentially, what you end up doing is rebuilding DynamoDB, or trying to rebuild Fauna, or something like that. So just thinking about that, anything you're building ... Maybe you have some advice for developers who ... I know we've talked about this a little bit, but I just go back to this idea of like, if you think about how complex some of these SaaS companies and these services that are being built out right now, why would you ever want to take that complexity on yourself?

Rob: Pride. Hubris. I mean, the correct answer is you wouldn't. You shouldn't.

Jeremy: People do.

Rob: Yeah, they do. I would beg and plead with them like, "Look, we did take a lot of that on. Fauna scales. You don't need to plan for sharding. You don't need to plan for global replication. All of these things are happening." I raise that as an example of understanding the customer's problem. The customer didn't want to think about, "Okay, past a thousand TPS, I need to create a new read replica. Past a million TPS, I need to have another region with active-active." The customer wanted to store some data and get that data, knowing that they had the ASA guarantees around it, right, and that's what the customer has.

So get that good understanding of what your customer really wants. If you can buy that, then you don't have a product yet. This is even out of software development and into product ideation at startups, right? If you can go ... Your customer's problem isn't they can't send text messages programmatically. They can do that through Twilio. They can do that through Amazon. They can do that through a number of different services, right? Your customer's problem is something else. So really get a good understanding of it. This is where knowing a little ... Like Joe Emison loves to rage against senior developers for knowing not quite enough. This is where knowing like, "Oh, yeah, Postgres. You can just shard it." "Just," the worst word in computer science, right?

Jeremy: Right.

Rob: "You can just shard it." Okay. Now, we're back to those database engineers that you talked about, and your customer doesn't want to shard a database. Your customer wants to store and retrieve data.

Jeremy: Right.

Rob: So any time that you can really focus in, and I guess I really got this one, this customer obsession beaten into me from my time at AWS. Really focus in on what the customer is asking you to do or what the customer needs even if they don't know how to express it, and build for that.

Jeremy: Right, right. Like the saying. I forgot who said it. Somebody from Harvard Business Review, but, "Your customers don't want a quarter-inch drill. They want a quarter-inch hole."

Rob: Right, right.

Jeremy: That's basically true. I mean, the complexity that goes behind the scenes are something that I think a vast majority of customers don't necessarily want, and you're right. If you focus on that product ideation thing, I think that's a big piece of that. All right. Well, anyway. So I have one more question for you just to ...

Rob: Please.

Jeremy: We've been talking for a while here. Hopefully, we haven't been boring people to death with our talk about APIs and stuff like that, but I would like to get a little bit academic here and go into that Calvin paper just a tiny bit because I think most people probably will not want to read it. Not because they don't want to, but because people are busy, right, and so they're listening to the podcast.

Rob: Yeah.

Jeremy: Just give us a quick summary of this because I think this is actually really fascinating and has benefits beyond just I think solving data problems.

Rob: Yeah. So what I would say first. I actually have this paper on my desk at all times. I would say read Section 1. It's one page, front and back. So if you're interested in it, you don't have to read the whole paper. Read that, and then listeners to this podcast will probably understand when I say this. Previously, for distributed databases and distributed transactions, you had what was called a two-phase commit. The first was you'd go out to all of your replicas, and you say, "Hey, I need lock." When everybody comes back, and acknowledges, and says, "Okay. You have the lock," then you do your transaction, and then you replicate it, and then you say, "Hey, everybody. I'm done. Release the lock." So it's a two-phase commit. If something went wrong, you rolled it all the way back and you said, "Hey, everybody. Forget it."

Calvin is event-sourcing for databases. So if I could distill the entire paper down into one concept, that's it. Right? Instead of saying, "Hey, everybody. Give me a lock, I'm going to do something," you say, "Hey, everybody. Here's what we're going to do." It's a deterministic application of the transaction so that you can ... You both create the lock and execute the transaction at the same time. So rather than having this outbound round trip, and then doing the thing in an outbound round trip, you have the outbound round trip, and you're done.

They all apply it independently, and then it gets into how you structure the guarantees around that, which again is very similar to event-sourcing in that you use snapshotting or checkpointing. "So, hey. At this point, we all agree. So we can forget all of our old transactions, and we roll forward from here." If somebody leaves the cluster, they come back in at that checkpoint. They reapply all of the events that occurred prior to that. Those events are transactions. They'd get up to working speed, and then off they go. The paper I think uses Paxos. That's an implementation detail, but the really interesting thing about it is you're not having a double round trip.

Jeremy: Yeah.

Rob: Again, I love the idea of event-sourcing. I think Amazon EventBridge is the best service that they've released in the past couple years.

Jeremy: Totally.

Rob: If you understand all of that and are already building serverless applications that way, well, that's what we're doing, just event-sourcing for database. That's it. That's it.

Jeremy: Just event-sourcing. It's easy. Simple. All right. All the words you never want to hear. Simple, easy, just. Right. Yeah. Perfect.

Rob: Yeah, but we do the hard work, so you don't have to. You don't care about all of that. You want to write your data somewhere, and you want to retrieve your data from an API, and that's what Fauna gives you.

Jeremy: Which I think is the main point here. So awesome. All right. Well, Rob, listen. This was great, and I'm super happy that I finally got you on the show. Congratulations for the new role at Fauna and on what's happening over there because it is pretty exciting.

Rob: Thank you.

Jeremy: I love companies that are innovating, right? It's not just another hosted database. You're actually building something here that is innovative, which is pretty amazing. So if people want to find out more about you, follow you on Twitter, or find out more about Fauna, how do they do that?

Rob: Right. Twitter, rts_rob. Probably the easiest way, go to my website, robsutter.com, and you will link to me from there. From there, of course, you'll get to fauna.com and all of our resources there. Always open to answer questions on Twitter. Yeah. Oh, rob@fauna.com. If you're old-school like me and you prefer the email, there you go.

Jeremy: All right. Awesome. Well, I will get all that into the show notes. Thanks again, Rob.

Rob: Thank you, Jeremy. Thanks for having me.

View Details

About Daniel Kim

Daniel Kim (He/Him) is a Senior Developer Relations Engineer at New Relic and the founder of Bit Project, a 501(c)(3) nonprofit dedicated to making tech accessible to underserved communities. He wants to inspire generations of students in tech to be the best they can be through inclusive, accessible developer education. He is passionate about diversity & inclusion in tech, good food, and dad jokes.

Twitter: @learnwdaniel
Volunteer with Bit Project: bitproject.org/volunteer
Learn Serverless with Bit Project: bitproject.org/course/serverless

Watch this video on YouTube: https://youtu.be/oDdrbDXQG6w

This episode sponsored by, CBT Nuggets.

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Daniel Kim. Hey, Daniel. Thanks for joining me.

Daniel: Hi, Jeremy. How's it going?

Jeremy: It's going real ...

Daniel: I'm glad to be here.

Jeremy: Well, I'm glad that you're here. So, you are a Senior Developer Relations Engineer at New Relic, but you're also the founder of Bit Project. So, I would love it if you could tell the listeners a little bit about yourself and your background and what Bit Project is all about.

Daniel: That sounds great, Jeremy. My name is Daniel. I'm a Senior Developer Relations Engineer at New Relic, which means I get to help the community and go find developers and help them become better developers. And I got into developer relations because I founded a school club and now it's a nonprofit, but it started as a school club, called Bit Project, where me and my friends gathered together to teach each other awesome web technologies. And yeah, that's how I got my start. And I am still running Bit Project as a nonprofit to help students around the world build and ship projects using awesome technologies and help them learn and become better developers.

Jeremy: Right. And one of those awesome technologies is serverless. And that's what I want to talk to you about today because this is a really great program that you're running here that helps make Serverless more accessible to more people, which is what I'm all about, right? So, I absolutely love this. So, let's go back and talk a little bit about Bit Project and just get into how it got started. You mentioned it was a project you were doing with some college friends, but how did it go from that to what it is now?

Daniel: Yeah. So, I started this, I think, late freshman year when I was still in school at UC Davis. I was not a computer science major, actually. I was an electrical engineering major, but as I got into technology and seeing all the possibilities of things you can build with cool tech, I was like, "I really need to get into web development because this is so awesome. I can make changes on the fly. I can see awesome things. I can build awesome things with my hands." Well, with my computer. So, yeah, I got a couple of friends together because I'm a very social person so I like to build and learn things together with my friends. So, I got a couple of them together. We rented a lecture hall and then we just taught each other everything we knew to each other. For example, I was super into Gatsby and React, so I was teaching my friends React. Other friends were super into backend development, so they were teaching me things like how to design APIs and how to connect a frontend to a backend, like really awesome things to each other.

And it started like that until I decided to scale the program so I could help more and more of my fellow students. So, instead of doing four-person meetups, I would organize a workshop. And those workshops turned into sponsored workshops with funding, which meant a lot of free food, which meant more people, and it just ballooned into this awesome student organization where we always had the best food. We had free Boba, free pizza, and we would share with each other all these awesome technologies and tools that we learned how to work with using in our projects. So, that's how it started.

Jeremy: Right. And then, so once you got this thing rolling, obviously you're seeing some success with it, then you get into developer relations?

Daniel: Yeah, definitely. So, that's when I understood what I wanted to do with them for the rest of my life. I didn't want to be that production engineer on-call all the time. I wanted to be that engineer that helped other engineers become more successful and find the joy in programming. I love seeing when developers find that "aha moment" when they're learning something new and help them become better developers. And I found that out when I was teaching my friends how to program because I got more joy out of seeing other people succeed than me succeeding myself. So, I was like, "Developer relations is the path for me." So, that's why I directly entered developer relations right out of college, because I was like, "This is what I'm meant to do." Because one of my favorite things to do is figure out how to break down really complex ideas and concepts into more fun, easy-to-understand chunks so everyone can succeed and have a good time. That's my thing.

Jeremy: No, I love that. I love that because I feel like, especially people who are maybe not traditional tech people or don't have a traditional tech background, sometimes it just takes a little bit of twisting of the presentation for them to really understand that. And I love that idea of just reaching out and trying to help more people because I'm on the total same page with you here. So, now you go and so you get into developer relations and you've got this Bit Project thing. And so is this something that you wanted to keep as a side project? What was the next evolution of that?

Daniel: Yeah, definitely. So, I think Bit Project is an extension to the advocacy work I do at New Relic. Because at New Relic, my job is not to push New Relic the product. We have amazing product marketing managers and other folks who do that. My job is to make it easy for people to level up the community, like the people in the community to level up as developers and help the community. And one way I do that is through Bit Project. So, a lot of the work I do at New Relic mirrors or is parallel to the work I'm doing at Bit Project, where I help make complex ideas more accessible to developers. So, in a way, it's not more of a side project. It's like a parallel project of what I'm doing at New Relic, what I'm doing at Bit Project.

Jeremy: Right. And so in terms of the things that you're teaching at Bit Project too because that's the other thing too. I think leveling up developers is one of those things where, I mean, if somebody wants to go learn HTML or CSS or one of those things, there's probably plenty of resources for them to go and do that. There's probably nine million YouTube tutorials out there, right?

Daniel: Definitely.

Jeremy: But for concepts like Serverless, right? And I mean even Serverless with Azure and AWS and some of these other things, these are newer things. I've actually interviewed quite a few candidates for a recent position that I'm trying to fill and not a lot of them are learning this stuff in college.

Daniel: Definitely. Something that we really wanted to instill to our students was that this is not your average bootcamp or course. We're not promising any six-figure salary after our bootcamps. That's not what we're promising. What we're promising is the opportunity to learn a concept that is foreign to many developers, even seasoned developers, because it's a relatively new technology, and we teach you the tools we give you and teach you the ways to become successful. So, we won't teach you everything you need to know, but we will teach you how to find the things you need to know to become successful developers. So, we help establish a good foundation for developers to learn new things and then build things on their own.

Jeremy: Right. And this is ...

Daniel: That's the focus of our program.

Jeremy: Yeah. And this is completely free, right?

Daniel: It's completely free. We're run thanks to the generosity of our corporate sponsors. So, shout out to them. Yeah. So, it's completely free for all students. So, please go and apply if you're interested and you are a student.

Jeremy: So, one of the major things that you focus on, and I know that you have different courses or different workshops that you're going through. And I know some of the other ones are a little bit earlier like the DevOps one. But you have a pretty robust serverless. I mean, that's the main thing, right? Teaching people to build serverless applications on Microsoft Azure. So, I'm curious, especially having somebody jump in from maybe a non-traditional tech background or no tech background at all, and also students of all ages, right? We're not just talking about high school or college kids here, that jumping into something like serverless, what makes serverless such a good, I guess, jumping in point for the types of candidates that you're looking for?

Daniel: Yeah. This is actually a great question because I have this conversation a lot with my colleagues at New Relic because when seasoned engineers hear about serverless, they jump straight into the, "How is this scalable for my enterprise use case? How is this going to integrate with my 70,000 other microservices?" They get into those questions immediately. But if you really boil down what serverless is, it's basically running code without thinking about infrastructure. That's the crux of what serverless is. And if you think about it from that perspective and not worry about all the other technical hurdles into implementing it in scale, it becomes a lot easier to digest for students. And it becomes a really friendly medium to get started with coding a project because you just have to code a small JavaScript or Python function that you just deploy to the cloud.

It just magically works. We try not to overwhelm students with all the infrastructure talk and more focus on the code that they're writing. And I was really inspired because one of my mentors for my career is Chloe Condon from Microsoft, and I remember her writing a lot of blogs around getting started with serverless. She built this fake boyfriend app with a Twilio and serverless. And I was like, "Hey, this is not that unapproachable for students to get started with serverless functions," because it was only maybe 40, 50 lines of code. It integrated multiple APIs. So, I was like, "This is the perfect medium," because it's relatively simple to understand the idea of just writing code and deploying it to a magical kingdom where the magical kingdom controls everything, you know?

Jeremy: Right.

Daniel: So, that's my inspiration for using serverless as a medium to teach people how the modern full-stack app works, if that makes sense.

Jeremy: Yeah. No, I totally agree, and I use this quite a bit where I tell people when I was a kid when I first started programming in the late 1990s, everything was CGI bins, right? So, we were just uploading code using FTP, but it was seriously magical. Now, again, it wouldn't scale, right. But it was magical in terms of how that happened. But even if it didn't scale, the point where you can get to that, what do we call it? The "aha moment," right? Where you're like, "Oh, this is how that works," or, "Oh, I get it now." I think you just get there faster with serverless.

Daniel: Exactly. I think that's one of the reasons I love serverless is that we have students spin up a serverless function day one of the camp. We don't wait until day three or day four to teach them how to build with serverless. We're like, "Hey, this is the environment that you're going to work in," and then we have them write their own serverless functions based on a boilerplate code that we have written already. So, we try to make the barrier to entry as low as possible, so students don't get intimidated by the word "serverless."

Jeremy: Right. Right. Yeah. And I think also it's probably a good place to get people started thinking about just what the cloud is and how the cloud works in general.

Daniel: Yeah. Definitely. Some of our students have never even heard of what an API is. So, we really take students from zero to understanding how different services work on the internet and how we can take advantage of services and other code that other people have written to write our own applications. Because a lot of students, especially junior developers, don't realize how little you have to code to actually get an app working. Because most likely there's someone in the world who's coded something that you're looking for to implement already. So, it's more like a jigsaw puzzle than trying to build something yourself.

Jeremy: Right. It's that whole Lego concept, just sticking those building blocks together. So, you mentioned, though, some of your students they've never even heard of an API or they don't know what an API is. So, I'm just thinking from the perspective of an absolute beginner, how do you scope a project for an absolute beginner to get them to somewhere where they actually have something that gets them to that "aha moment," makes them feel like, "Hey, I've actually done something interesting here," but not overwhelm them with things like open API spec 3.0? You know what I mean? All this kind of stuff.

Daniel: Yeah. I think one of the most important things when you're designing a curriculum is understanding the pain points of the student. So, this curriculum was designed by a bunch of students. I'm not the only one that wrote this curriculum. This curriculum had a lot of contributors from all over the world who are high school and college students. We knew that we didn't want to go too in-depth from the beginning because we have a lot of students from non-traditional backgrounds that don't have a lot of previous knowledge. So, what we try to do is set up guide rails and have boilerplates and things like that to ensure that they're successful. Because the worst thing you can do when you're working with a junior developer is just overwhelm them with information and have really, really hard assignments that lead to frustration.

So, we try to make that path really, really easy. But instead, what we try to do is have stretch goals or have extra-curricular assignments where they can apply what they have learned. So, if they're a little bit more advanced and they're getting the concepts and they're understanding at a deeper level how things work, they're able to practice and hone those skills. So, what we do is we try to work with our mentors, our fabulous mentors who are engineers in the industry, to help students code those stretch goals and help them understand at a deeper level if they have the capacity to do so. So, we try to customize the experience for every student based on their previous experience.

Jeremy: Right. And I think another important thing is setting expectations with the students as well. I mean, you mentioned earlier that this isn't a bootcamp that you're guaranteeing $100,000 salaries when you walk out. I think that that is something, to me, where I think that level of honesty and truth is really important because I think there are a lot of these eight to 12-week boot camps that over-promise. And I don't know. I mean, I've been doing this for 24 years, and I don't feel like I'm an expert on anything and I've been doing it for a very long time. So, eight weeks doesn't get you to be an expert in anything, but if you can become productive, that's pretty exciting.

Daniel: Yeah. Our goal is not to get you a six-figure job. Because that would be nice, but I feel like that's straight-up lying. Because I don't know all the students before they start personally, and I can't promise them a six-figure job. That's just ridiculous to me. But what I can promise is that you will ship an app. That's what I can promise. And I feel like when you're shipping an app and you're writing code to build an actual app that will work, you learn so much. You learn how to plan for a software project, how to ask questions, how to look for things on Google. So, that's the things we promise is the experiences, not necessarily the shiny six-figure salary. Even though I wish I could promise that. That would be amazing.

Jeremy: Right. Yeah. And I think probably the greatest skill you can teach anyone as a developer is how to Google and how to use stack overflow.

Daniel: Definitely.

Jeremy: All right. So, you mentioned something about customizing, trying to make sure that the curriculum is adapted for the particular student. So, tell me a little bit more about that because that sounds really interesting.

Daniel: Yeah. So, one of the reasons that I find our content and curriculum really special is that it's open-ended. It's not like they're programming exactly what every other student programs. So, for the first four weeks, we teach how serverless functions work, how to set up your development environment, everything through pair programming. So students, instead of having lectures, we have senior engineers actually pair program with junior developers, younger students, or students with less experience, so they can ask questions in the chat to learn as they are doing it with a mentor. And during the last four weeks, we actually have the students apply the things they've learned in the first four weeks through pair programming into their own applications. So, we teach them, "Hey, by week one, you should have this part of your project done. Week two, you should have this part of your project done." But we don't really specify exactly what their project should be. So, at the end of the camp, every single student has a different project they have built based on the interests they have, which has been really awesome to see.

Jeremy: Well, that's also great, too. It's one of those things where, when your English teacher forces you to read Romeo and Juliet and you're not interested in Shakespeare, it's really hard to excel in that sometimes. So, letting people pick and choose where they go, I think is, again, is just a really good motivator and an excellent way. And again, just not over-promising. Just teaching people some of the basics, and then you have something to work on, something to iterate on, something to go a little bit deeper on and start understanding. If you just know that there are headers when you call an API, then you can maybe start doing some research as to what the other headers are and what I can do with those. And I think that level of curiosity would be really great for somebody and again, would excite them and get them going down that path.

Daniel: Definitely. And I think the best way that students learn is actually trying to implement the things they have in their head. Because some of these projects that students have built for their capstone projects have been very, very complicated using serverless functions. One of the students actually built a Dropbox clone using serverless functions, and it was actually amazing. I couldn't do that, honestly, but she built it in three weeks, I think. So, I think it's the creativity that really, really I find impressive and amazing every cohort we have, is the variance in projects that we have for every single student.

Jeremy: Right. Yeah. So, what are some of those projects? Because I think that'd be really interesting. Just give some examples of the sort of things you can build, right? Because the "Hello, World" tutorials are out there. People can go and probably cobble something together, but it sounds to me like the students that you have are building something that is actually, maybe not-production ready, but it is something that solves a real problem and it's a real solution to that. So, what are some of those different projects?

Daniel: Definitely. One of our students, Bo, built an IoT heart rate monitor that connected to a serverless function. So, every time that the heartbeat went over a certain number, it would send a Twilio text message to the family members of whoever was subscribed to that particular heartbeat monitor. And he built that because his grandfather was suffering with some heart issues, and it was really important to his family that they knew that he was doing okay. They got alerted every time his heartbeat got too fast. So, he actually built this whole thing using a Raspberry PI. He had a heartbeat sensor that was attached to a bracelet and it actually connected to a Serverless function. And he demoed it and he actually did jumping jacks to get his heart rate up. It actually worked, which was super awesome. We got to demo during our demo day.

Another student built a face mask detector. So, she would have someone take a picture on her website of someone wearing a face mask. And it would tell, using some cognitive APIs, if someone was wearing a mask or not. And she designed that because she knew a lot of local businesses who didn't have staff directly in the entrance of the business, and she wanted to make sure there was a solution where the owners could make sure someone was wearing a mask before they entered the establishment. So, that was a really cool project. There was another student who was actually in his forties who was a mining engineer who wanted to make a career change. So, he actually built this awesome serverless function that sent out earthquake notifications based on the data from the government, which was really, really awesome as well. So, there's so many projects that students have built with serverless functions, ranging everything from IoT to big data and so many things that I've learned actually, by watching all these projects being built.

Jeremy: Yeah. That's amazing. And actually I think that something that's really interesting, you mentioned the gentleman with the career change, is that developers, I think, especially career developers, I mean, we get narrowly focused on solving software problems, right?

Daniel: Exactly.

Jeremy: And we maybe don't think so much about some of these other real-world problems that exist. So, that idea of taking your existing life experiences and problems that you've been dealing with and have a solution maybe in your head, but you can't express that. That's really frustrating, right? So, being able to do something like this and being able to express that, I think that's absolutely amazing.

Daniel: Definitely. And I think this is one of the reasons why I find this program really rewarding for both students and the people who actually run the program because they see folks who have zero experience getting to the point where they can build the things that are in their head, which I think is magical.

Jeremy: Yeah. No, I totally agree. Also, I think you said there's some other case studies on the blog?

Daniel: Yeah. So, if you go to bitproject.org and go to the blog, we have a bunch of case studies that are still being uploaded. So, every week we're going to have new student projects that are going to be uploaded there. So, if you want to see some of the cool stuff that our students have built, feel free to go check it out.

Jeremy: Awesome. All right. So, you just mentioned that this is a really rewarding thing. And I know for me, I do a lot of open source projects. I try to help as many people as I can. I don't run a nonprofit that runs courses. Maybe someday. But I do get exactly what you're saying because it is great to get that feedback, to see someone be successful because you've helped enable that. So, I know you're looking for mentors, right?

Daniel: Yeah, definitely. We're looking for mentors who have previous experience or passion with serverless to mentor students, to get them to that point where they can build their own apps. So, we'd love to have you if you are interested and have a couple of hours per week to spare.

Jeremy: Right. What's the requirement or the time commitment? It's just a few hours a week?

Daniel: We recommend four to five hours a week to just work directly one-on-one with the student, and previous experience in serverless or just regular full-stack development is quite encouraged because we want to make sure that you are able to answer some of the technical questions that students might have around the content.

Jeremy: Right. And you mentioned that, again, just going back to the mining example, but it sounds like that gentlemen was a little bit older. So, what's the age range of the students that you have in this program?

Daniel: We don't have a minimum or maximum age that we accept. We just care about passion and the willingness to complete the program. Because the program is completely free, the standard that we set for our applicants is not of experience, but more of passion and desire to learn and become a successful developer.

Jeremy: Right, right. Yeah. So, what about for mentors or people who are looking to do this? Again, I know it's rewarding to work with people and to help people. You know it's rewarding. What can you tell people, though, that might be interested in this? What are some of the other benefits, I guess, of being a mentor?

Daniel: Yeah. Some of the really cool benefits I've seen is that we've been working directly with the Azure Functions team at Microsoft to mentor our students because they are using Azure Functions as the platform to host their serverless functions. And we've actually had PMs that are building as their functions, work with our students to get new ideas for product features, as well as engineers getting direct feedback on the features they worked on only a couple of weeks prior. Which I think is quite magical because I've seen these older PMs who are building that product the students are using and the students are very blunt, let me tell you. So they're like, "This feature makes no sense." So, a couple of weeks later it's magically fixed for some reason. I don't know how that could have happened, but things get resolved quite quickly when the student feedback comes in.

Jeremy: Yeah. And also the other thing is, is that again, it's feedback that's, I guess, untainted from the experience of being a developer, right?

Daniel: Definitely.

Jeremy: So, it's like that childhood honesty that is what probably every product management team needs to figure that stuff out. So, all right. Well, so where are you going with this? What do you hope to do with Bit Project? I mean, is this something you want to grow or you want to add more courses? What's the future?

Daniel: Definitely. That's a great question. So, as I work for New Relic, we are pivoting to create more content and more courses and more interactive learning materials and experiences in the DevOps field. So, right now we're creating content around observability, around container orchestration, things like that, that are more niche skills that students could learn to better their chances of getting a job as a site reliability engineer or a DevOps engineer. But most importantly, right now what we're trying to do is make sure that we're ready to scale as soon as possible because we feel like we have something really special here where we're teaching students how to ship apps, not to learn specific concepts like HTML or CSS. I think we have a really unique model here of how we're teaching students and how we're working with industry, leveraging cloud advocates and engineers who want to volunteer their specialized skill to better the community. So, right now I see the future as us helping make and lead more engineers of the future so we can have better services and better internet, hopefully, in a couple of years or a couple of decades.

Jeremy: Well, it's a very noble goal. And so what about data science? I know there's a thing on the site about data science and I think you're doing some work with universities around that, right?

Daniel: Yeah, definitely. So, we have a program called Bit University where we create these really easy-to-integrate data science courses for humanities classrooms, because there's a huge demand right now for humanities students to get data science experience, to get research opportunities as well as job opportunities. But a lot of them actually don't have access to data science courses because they're a humanities major, especially at smaller schools. So, what we do is we partner with universities like Cal State Fullerton and Sacramento State University to provide data science courses specifically tailored for humanities majors at these schools partnering with professors. So, yeah, that's the program and it's been super successful and we've had so many humanities students learn the basic skills they need to get these internships and these research opportunities, which has been really rewarding.

Jeremy: Yeah. That's awesome. So Daniel, is there anything else you want to tell the listeners about Bit Project?

Daniel: Definitely. Yeah. So, if you or your company want to help us make more technical content, like let's say you work in DevOps or you even work in Serverless functions that you want to extend the work we're doing, especially if you're an advocate, please reach out to me. I'm on Twitter. I'm on email. So, please reach out to me to work together on more technical content because my job is to make things more assessable. So, if you want to make anything, whether it's your area of expertise or something you think could be more accessible, I'd love to work with you to make sure that happens. And that is a free resource that's available to the community. That's my plug.

Jeremy: That's awesome.

Daniel: Reach out to me.

Jeremy: That's awesome. No, I love it. Daniel, I appreciate, one, you being here and sharing this with everybody, but also the work that you're doing with the community is just amazing. The more people we can get into serverless and the more people we can get to understand this next generation of, I don't know, applications, I guess you want to call it, is absolutely a very, very noble goal. So, you mentioned Twitter. So, it's just learnwdaniel, right?

Daniel: Definitely. Yeah.

Jeremy: And then also the Bit Project has a Twitter, just bitPRJ. And then if you're interested in volunteering, you go to bitproject.org/volunteer. And students, if students want to sign up, how do they do that? They just go to bitproject.org?

Daniel: Yeah. You can go directly to apply at bitproject.org, or if you want more information about the program, just go to bitproject. There's a huge banner at the top that will lead you directly to that website.

Jeremy: Awesome. All right. Well, I will make sure I get all that into the show notes. Thanks again, Daniel.

Daniel: Thank you.

View Details

About Gojko Adzic

Gojko Adzic is a partner at Neuri Consulting LLP. He one of the 2019 AWS Serverless Heroes, the winner of the 2016 European Software Testing Outstanding Achievement Award, and the 2011 Most Influential Agile Testing Professional Award. Gojko’s book Specification by Example won the Jolt Award for the best book of 2012, and his blog won the UK Agile Award for the best online publication in 2010.

Gojko is a frequent speaker at software development conferences and one of the authors of MindMup and Narakeet.

As a consultant, Gojko has helped companies around the world improve their software delivery, from some of the largest financial institutions to small innovative startups. Gojko specializes in agile and lean quality improvement, in particular impact mapping, agile testing, specification by example, and behavior driven development.

Twitter: @gojkoadzic
Narakeet: https://www.narakeet.com
Personal website: https://gojko.net

Watch this video on YouTube: https://youtu.be/kCDDli7uzn8

This episode is sponsored by CBT Nuggets: https://www.cbtnuggets.com/

Transcript
Jeremy: Hi everyone, I'm Jeremy Daly and this is Serverless Chats. Today my guest is Gojko Adzic. Hey Gojko, thanks for joining me.

Gojko: Hey, thanks for inviting me.

Jeremy: You are a partner at Neuri Consulting, you're an AWS Serverless Hero, you've written I think, what? I think 6,842 books or something like that about technology and serverless and all that kind of stuff. I'd love it if you could tell listeners a little bit about your background and what you've been working on lately.

Gojko: I'm a developer. I started developing software when I was six and a half. My dad bought a Commodore 64 and I think my mom would have kicked him out of the house if he told her that he bought it for himself, so it was officially for me.

Jeremy: Nice.

Gojko: And I was the only kid in the neighborhood that had a computer, but didn't have any ways of loading games on it because he didn't buy it for games. I stayed up and copied and pasted PEEKs and POKEs in a book I couldn't even understand until I made the computer make weird sounds and print rubbish on the screen. And that's my background. Basically, ever since, I only wanted to build software really. I didn't have any other hobbies or anything like that. Currently, I'm building a product for helping tech people who are not video editing professionals create videos very easily. Previously, I've done a lot of work around consulting. I've built a lot of product that is used by millions of school children worldwide collaborate and brainstorm through mind-mapping. And since 2016, most of my development work has been on Lambda and on team stuff.

Jeremy: That's awesome. I joke a little bit about the number of books that you wrote, but the ones that you have, one of them's called Running Serverless. I think that was maybe two years ago. That is an excellent book for people getting started with serverless. And then, one of my probably favorite books is Humans Vs Computers. I just love that collection of tales of all these things where humans just build really bad interfaces into software and just things go terribly.

Gojko: Thank you very much. I enjoyed writing that book a lot. One of my passions is finding edge cases. I think people with a slight OCD like to find edge cases and in order to be a good developer, I think somebody really needs to have that kind of intent, and really look for edge cases everywhere. And I think collecting these things was my idea to help people first of all think about building better software, and to realize that stuff we might glance over like, nobody's ever going to do this, actually might cause hundreds of millions of dollars of damage ten years later. And thanks very much for liking the book.

Jeremy: If people haven't read that book, I don't know, when did that come out? Maybe 2016? 2015?

Gojko: Yeah, five or six years ago, I think.

Jeremy: Yeah. It's still completely relevant now though and there's just so many great examples in there, and I don't want to spent the whole time talking about that book, but if you haven't read it, go check it out because it's these crazy things like police officers entering in no plates whenever they're giving parking tickets. And then, when somebody actually gets that, ends up with thousands of parking tickets, and it's just crazy stuff like that. Or, not using the middle initial or something like that for the name, or the birthdate or whatever it was, and people constantly getting just ... It's a fascinating book. Definitely check that out.

But speaking of edge cases and just all this experience that you have just dealing with this idea of, I guess finding the problems with software. Or maybe even better, I guess a good way to put it is finding the limitations that we build into software mostly unknowingly. We do this unknowingly. And you and I were having a conversation the other day and we were talking about way, way back in the 1970s. I was born in the late '70s. I'm old but hopefully not that old. But way back then, time-sharing was a thing where we would basically have just a few large computers and we would have to borrow time against them. And there's a parallel there to what we were doing back then and I think what we're doing now with cloud computing. What are your thoughts on that?

Gojko: Yeah, I think absolutely. We are I think going in a slightly cyclic way here. Maybe not cyclic, maybe spirals. We came to the same horizontal position but vertically, we're slightly better than we were. Again, I didn't start working then. I'm like you, I was born in late '70s. I wasn't there when people were doing punch cards and massive mainframes and time-sharing. My first experience came from home PC computers and later PCs. The whole serverless thing, people were disparaging about that when the marketing buzzword came around. I don't remember exactly when serverless became serverless because we were talking about microservices and Lambda was a way to run microservices and execute code on demand. And all of a sudden, I think the JAWS people realized that JAWS is a horrible marketing name, and decided to rename it to serverless. I think it most important, and it was probably 2017 or something like that. 2000 ...

Jeremy: Something like that, yeah.

Gojko: Something like that. And then, because it is a horrible marketing name, but it's catchy, it caught on and then people were complaining how it's not serverless, it's just somebody else's servers. And I think there's some truth to that, but actually, it's not even somebody else's servers. It really is somebody else's mainframe in a sense. You know in the '70s and early '80s, before the PC revolution, if you wanted to be a small software house or a small product operator, you probably were not running your own data center. What you would do is you would rent it based on paying for time to one of these massive, massive, massive operators. And in fact, we ended up with AWS being a massive data center. As far as you and I are concerned, it's just a blob. It's not a collection of computers, it's a data center we learn something from and Google is another one and then Microsoft is another one.

And I remember reading a book about Andy Grove who was the CEO of Intel where they were thinking about the market for PC computers in the late '70s when somebody came to them with the idea that they could repurpose what became a 8080 processor. They were doing this I think for some Japanese calculator and then somebody said, "We can attach a screen to this and make this a universal computer and sell it." And they realized maybe there's a market for four or five computers in the world like that. And I think that that's ... You know, we ended up with four or five computers, it's just the definition of a computer changed.

Jeremy: Right. I think that's a good point because you think about after the PC revolution, once the web started becoming really big, people started building data centers and collocation facilities like crazy. This is way before the cloud, and everybody was buying racks and Dell was getting really popular because people buying servers from Dell, and installing these in their data centers and doing this. And it just became this massive, whole industry built around doing that. And then you have these few companies that say, "Well, what if we just handled all that stuff for you? Rather than just racking stuff for you," but started just managing the software, and started managing the networking, and the backups, and all this stuff for you? And that's where the cloud was born.

But I think you make a really good point where the cloud, whatever it is, Amazon or Google or whatever, you might as well just assume that that's just one big piece of processing that you're renting and you're renting some piece of that. And maybe we have. Maybe we've moved back to this idea where ... Even though everybody's got a massive computer in their pocket now, tons of compute power, in terms of the real business work that's being done, and the real global value, and the things that are powering global commerce and everything else like that, those are starting to move back to run in four, five, massive computers.

Gojko: Again, there's a cyclic nature to all of this. I remember reading about the advent of power networks. Because before people had electric power, there were physical machines and movement through physical power, and there were water-powered plants and things like that. And these whole systems of shafts and belts and things like that powering factories. And you had this one kind of power load in a factory that was somewhere in the middle, and then from there, you actually have physical belts, rotating cogs in other buildings, and that was rotating some shafts that were rotating other cogs, and things like that.

First of all, when people were able to package up electricity into something that's distributable, and they were running their own small electricity generators next to these big massive machines that were affecting early factories. And one of the first effects of that was they could reuse 30% of their factories better because it was up to 30% of the workspace in the factory that was taken up by all the belts and shafts. And all that movement was producing a lot of air movement and a lot of dust and people were getting sick. But now, you just plug a cable and you no longer have all this bad air and you don't have employees going sick and things like that. Things started changing quite a lot and then all of a sudden, you had this completely new revolution where you no longer had to operate your own electric generator. You could just plug in and get power from the network.

And I think part of that is again, cyclic, what's happening in our industry now, where, as you said, we were getting machines. I used to make money as a Linux admin a long time ago and I could set up my own servers and things like that. I had a company in 2007 where we were operating our own gaming system, and we actually had physical servers in a physical server room with all the LEDs and lights, and bleeps, and things like that. Around that time, AWS really made it easy to get virtual machines on EC2 and I realized how stupid the whole, let's manage everything ourself is. But, we are getting to the point where people had to run their own generators, and now you can actually just plug into the electricity network. And of course, there is some standardization. Maybe U.S. still has 110 volts and Europe has 220, and we never really get global standardization there.

But I assume before that, every factory could run their own voltage they wanted. It was difficult to manufacture for these things but now you have standardization, it's easier for everybody to plug into the ecosystem and then the whole ecosystem emerged. And I think that's partially what's happening now where things like S3 is an API or Lambda is an API. It's basically the electric socket in your wall.

Jeremy: Right, and that's that whole Wardley maps idea, they become utilities. And that's the thing where if you look at that from an enterprise standpoint or from a small business standpoint if you're a startup right now and you are ordering servers to put into a data center somewhere unless you're doing something that's specifically for servers, that's just crazy. Use the cloud.

Gojko: This product I mentioned that we built for mind mapping, there's only two of us in the whole company. We do everything from presales, to development testing supports, to everything. And we're competing with companies that have several orders of magnitude more employees, and we can actually compete and win because we can benefit from this ecosystem. And I think this is totally wonderful and amazing and for anybody thinking about starting a product, it's easier to start a product now than ever. And, another thing that's totally I think crazy about this whole serverless thing is how in effect we got a bookstore to offer that first.

You mentioned the world utility. I remember I was the editor of a magazine in 2001 in Serbia, and we had licensing with IDG to translate some of their content. And I remember working on this kind of piece from I think PC World in the U.S. where they were interviewing Hewlett Packard people about utility computing. And people from Hewlett Packard back then were predicting that in a few years' time, companies would not operate their own stuff, they would use utility and things like that. And it's totally amazing that in order to reach us over there, that had to be something that was already evaluated and tested, and there was probably a prototype and things like that. And you had all these giants. Hewlett Packard in 2001 was an IT giant. Amazon was just up-and-coming then and they were a bookstore then. They were not even anything more than a bookstore. And you had, what? A decade later, the tables completely turned where HP's ... I don't know ...

Jeremy: I think they bought Compaq at some point too.

Gojko: You had all these giants, IBM completely missed it. IBM totally missed ...

Jeremy: It really did.

Gojko: ... the whole mobile and web and everything revolution. Oracle completely missed it. They're trying to catch up now but fat chance. Really, we are down to just a couple of massive clouds, or whatever that means, that we interact with as we're interacting with electricity sockets now.

Jeremy: And going back to that utility comparison, or, not really a comparison. It is a utility now. Compute is offered as a utility. Yes, you can buy and generate compute yourself and you can still do that. And I know a lot of enterprises still will. I think cloud is like 4% of the total IT market or something. It's a fraction of it right now. But just from that utility aspect of it, from your experience, you mentioned you had two people and you built, is it MindMup.com?

Gojko: MindMup, yeah.

Jeremy: You built that with just two people and you've got tons of people using it. But just from your experience, especially coming from the world of being a Linux administrator, which again, I didn't administer ... Well, I guess I was. I did a lot of work in data centers in my younger days. But, coming from that idea and seeing how companies were building in the past and how companies are still building now, because not every company is still using the cloud, far from it. But not taking advantage of that utility, what are those major disadvantages? How badly do you think that's going to slow companies down that are trying to innovate?

Gojko: I can give you a story about MindMup. You mentioned MindMup. When was it? 2018, there was the Intel processor vulnerabilities that were discovered.

Jeremy: Right, yes.

Gojko: I'm not entirely sure what the year was. A few years ago anyway. We got a email from a concerned university admin when the second one was discovered. The first one made all the news and a month later a second one was discovered. Now everybody knew that, they were in panic and things like that. After the second one was discovered, we got a email from a university admin. And universities are big users, they need to protect the data and things like that. And he was insisting that we tell him what our plan was for mitigating this thing because he knows we're on the cloud.

I'm working on European time. The customer was in the U.S., probably somewhere U.S. Pacific because it arrived in the middle of the night. I woke up, I'm still trying to get my head around and drinking coffee and there's this whole sausage CV number that he sent me. I have no idea what it's about. I took that, pasted it into Google to figure out what's going on. The first result I got from Google was that AWS Lambda was already patched. Copy, paste, my day's done. And I assume lots and lots of other people were having a totally different conversation with their IT department that day. And that's why I said I think for products like the one I'm building with video and for the MindMup, being able to rent operations as a utility, but really totally rent ops as a utility, not have to worry about anything below my unique business level is really, really important.

And yes, we can hire people to work on that it could even end up being slightly cheaper technically but in terms of my time and where my focus goes and my interruptions, I think deploying on a utility platform, whatever that utility platform is, as long as it's reliable, lets me focus on adding value where I can actually add value. That makes my product unique rather than the generic stuff.

Jeremy: You mentioned the video product that you're working on too, and something that is really interesting I think too about taking advantage of the cloud is the scalability aspect of it. I remember, it was maybe 2002, maybe 2003, I was running my own little consulting company at the time, and my local high school always has a rivalry football game every Thanksgiving. And I thought it'd be really interesting if I was to stream the audio from the local AM radio station. I set up a server in my office with ReelCast Streaming or something running or whatever it was. And I remember thinking as long as we don't go over 140 subscribers, we'll be okay. Anything over that, it'll probably crash or the bandwidth won't be enough or whatever.

Gojko: And that's just one of those things now, if you're doing any type of massive processing or you need bandwidth, bandwidth alone ... I remember T1 lines being great and then all of a sudden it was like, well, now you need a T3 line or something crazy in order to get the bandwidth that you need. Just from that aspect of it, the ability to scale quickly, that just seems like such a huge blocker for companies that need to order provision servers, maybe get a utility company to come in and install more bandwidth for them, and things like that. That's just stuff that's so far out of scope for building a business to me. At least building a software business or building any business. It's crazy.

When I was doing consulting, I did a bit of work for what used to be one of the largest telecom companies in the world.

Jeremy: Used to be.

Gojko: I don't want to name names on a public chat. Somewhere around 2006, '07 let's say, we did a software project where they just needed to deploy it internally. And it took them seven months to provision a bunch of virtual machines to deploy it internally. Seven months.

Jeremy: Wow.

Gojko: Because of all the red tape and all the bureaucracy and all the wait for capacity and things like that. That's around the time where Amazon when EC2 became commercially available. I remember working with another client and they were waiting for some servers to arrive so they can install more capacity. And I remember just turning on the Amazon console. I didn't have anything useful to running it then but just being able to start up a virtual machine in about, I think it was less than half an hour, but that was totally fascinating back then. Here's a new Linux machine and in less than half an hour, you can use it. And it was totally crazy. Now we're getting to the point where Lambda will start up in less than 10 milliseconds or something like that. Waiting for that kind of capacity is just insane.

With the video thing I'm building, because of Corona and all of this remote teaching stuff, for some reason, we ended up getting lots of teachers using the product. It was one of these half-baked experiments because I didn't have time to build the full user interface for everything, and I realized that lots of people are using PowerPoint to prepare that kind of video. I thought well, how about if I shorten that loop, so just take your PowerPoint and convert it into video. Just type up what you want in the speaker notes, and we'll use these neurometrics to generate audio and things like that. Teachers like it for one reason or the other.

We had this influential blogger from Russia explain it on his video blog and then it got picked up, my best guess from what I could see from Google Translate, some virtual meeting of teachers of Russia where they recommended people to try it out. I woke up the next day, the metrics went totally crazy because a significant portion of teachers in Russia tried my tool overnight in a short space of time. Something like that, I couldn't predict it. It's lovely but as you said, as long as we don't go over a hundred subscribers, we're fine. If I was in a situation like that, the thing would completely crash because it's unexpected. We'd have a thing that's amazingly good for marketing that would be amazingly bad for business because it would crash all our capacity we had. Or we had to prepare for a lot more capacity than we needed, but because this is all running on Lambda, Fargate, and other auto-scaling things, it's just fine. No sweat at all. It was a lovely thing to see actually.

Jeremy: You actually have two problems there. If you're not running in the cloud or not running on-demand compute, is the fact that one, you would've potentially failed, things would've fallen over and you would've lost all those potential customers, and you wouldn't have been able to grow.

Gojko: Plus you've lost paying customers who are using your systems, who've paid you.

Jeremy: Right, that's the other thing too. But, on the other side of that problem would be you can't necessarily anticipate some of those things. What do you do? Over-provision and just hope that maybe someday you'll get whatever? That's the crazy thing where the elasticity piece of the cloud to me, is such a no-brainer. Because I know people always talk about, well, if you have predictable workloads. Well yeah, I know we have predictable workloads for some things, but if you're a startup or you're a business that has like ... Maybe you'd pick up some press. I worked for a company that we picked up some press. We had 10,000 signups in a matter of like 30 seconds and it completely killed our backend, my SQL database. Those are hard to prepare for if you're hosting your own equipment.

Gojko: Absolutely, not even if you're hosting your own.

Jeremy: Also true, right.

Gojko: Before moving to Lambda, the app was deployed to Heroku. That was basically, you need to predict how many virtual machines you need. Yes, it's in the cloud, but if you're running on EC2 and you have your 10, 50, 100 virtual machines, whatever running there, and all of a sudden you get a lot more traffic, will it scale or will it not scale? Have you designed it to scale like that? And one of the best things that I think Lambda brought as a constraint was forcing people to design this stuff in a way that scales.

Jeremy: Yes.

Gojko: I can deploy stuff in the cloud and make it all distributed monolith, so it doesn't really scale well, but with Lambda because it was so constrained when it launched, and this is one thing you mentioned, partially we're losing those constraints now, but it was so constrained when it launched, it was really forcing people to design things that were easy to scale. We had total isolation, there was no way of sharing things, there was no session stickiness and things like that. And then you have to come up with actually good ways of resolving that.

I think one of the most challenging things about serverless is that even a Hello World is a distributed transaction processing system, and people don't get that. They think about, well, I had this DigitalOcean five-dollar-a-month server and it was running my, you know, Rails up correctly. I'm just going to use the same ideas to redesign it in Lambda. Yes, you can, but then you're not going to really get the benefits of all of this other stuff. And if you design it as a massively distributed transaction processing system from the start, then yes, it scales like crazy. And it scales up and down and it's lovely, but as Lambda's maturing, I have this slide deck that I've been using since 2016 to talk about Lambda at conferences. And every time I need to do another talk, I pull it out and adjust it a bit. And I have this whole Git history of it because I do markdown to slides and I keep the markdown in Git so I can go back. There's this slide about limitations where originally it's only ... I don't remember what was the time limitation, but something very short.

Jeremy: Five minutes originally.

Gojko: Yeah, something like that and then it was no PCI compliance and the retries are difficult, and all of this stuff basically became sold. And one of the last things that was there, there was don't even try to put it in a VPC, definitely, you can but it's going to take 10 minutes to start. Now that's reasonably okay as well. One thing that I remember as a really important design constraint was effectively it was a share nothing platform because you could not share data between two Lambdas running at the same time very easily in the same VM. Now that we can connect Lambdas to EFS, you effectively can do that as well. You can have two Lambdas, one writing into an EFS, the other reading the same EFS at the same time. No problem at all. You can pump it into a file and the other thing can just stay in a file and get the data out.

As the platform is maturing, I think we're losing some of these design constraints, and sometimes constraints breed creativity. And yes, you still of course can design the system to be good, but it's going to be interesting to see. And this 15-minute limit that we have in Lamdba now is just an artificial number that somebody thought.

Jeremy: Yeah, it's arbitrary.

Gojko: And at some point when somebody who is important enough asks AWS to give them half-hour Lamdbas, they will get that. Or 24-hour Lambdas. It's going to be interesting to see if Lambda ends up as just another way of running EC2 and starting EC2 that's simpler because you don't have to manage the operating system. And I think the big difference we'll get between EC2 and Lambda is what percentage of ops your developers are responsible for, and what percentage of ops Amazon's developers are responsible for.

Because if you look at all these different offerings that Amazon has like Lightsail and EC2 and Fargate and AWS Batch and CodeDeploy, and I don't know how many other things you can run code on in Lambda. The big difference with Lambda is really, at least until very recently was that apart from your application, Amazon is responsible for everything. But now, we're losing design constraints, you can put a Docker container in, you can be responsible for the OS image as well, which is a bit again, interesting to look at.

Jeremy: Well, I also wonder too, if you took all those event sources that you can point at Lambda and you add those to Fargate, what's the difference? It seems like they're just merging into two very similar products.

Gojko: For the video build platform, the last step runs in Fargate because people are uploading things that are massive, massive, massive for video processing, and just they don't finish in 15 minutes. I have to run to Fargate, and the big difference is the container I packaged up for Fargate takes about 40 seconds to actually deploy. A new event at the moment with the stuff I've packaged in Fargate takes about 40 seconds to deploy. I can optimize that, but I can't optimize it too much. Fargate is still order of magnitude of tens of seconds to process an event. I think as Fargate gets faster and as Lambda gets more of these capabilities, it's going to be very difficult to tell them apart I think.

With Fargate, you're intended to manage the container image yourself. You're responsible for patching software, you're responsible for patching OS vulnerabilities and things like that. With Lambda, Amazon, unless you use a container image, Amazon is responsible for that. They come close. When looking at this video building for the first time, I was actually comparing code. I was considering using CodeBuild for that because CodeBuild is also a way to run things on demand and containers, and you actually can get quite decent machines with CodeBuild. And it's also event-driven, and Fargate is event-driven, AWS Batch is event-driven, and all of these things are converging to each other. And really, AWS is famous for having 10 products that do the same thing effectively and you can't tell them apart, and maybe that's where we'll end.

Jeremy: And I'm wondering too, the thing that was great about Lambda, at least for me like you said, the shared nothing architecture where it was like, you almost didn't have to rely on anything other than the event that came in, and the processing of that Lambda function. And if you designed your systems well, you may have some bottleneck up front, but especially if you used distributed transactions and you used async invocations of downstream functions, where you could basically take some data that you needed to pass into it, and then you wouldn't necessarily need that to communicate with anything other than itself to process that data. The scale there was massive. You could just keep scaling and scaling and scaling. As you add things like EFS and that adds constraints in terms of the number of transactions and connections that, that can make and all those sort of things. Do these things, do they become less reliable? By allowing it to do more, are we building systems that are less reliable because we're not using some of those tried-and-true constraints that were there?

Gojko: Possibly, but every time you add a new moving part, you create one more potential point of failure there. And I think for me, one of the big lessons when I was working on ... I spent a few years working on very high throughput transaction processing systems. That's why this whole thing rings a bell a lot. A lot of it really was how do you figure out what type of messages you send and where you send them. The craze of these messages and distributed transaction processing systems in early 2000s, created this whole craze of enterprise service buses later that came. We now have this... What is it called? It's not called enterprise service bus, it's called EventBridge, or something like that.

Jeremy: EventBridge, yes.

Gojko: That's effectively an enterprise service bus, it's just the enterprise is the Amazon cloud. The big challenge in designing things like that is decoupling. And it's realizing that when you have a complicated system like that, stuff is going to fail. And especially when we were operating around hardware, stuff is going to fail badly or occasionally, and you need to not bring the whole house down where some storage starts working a bit slower. You create circuit breakers, you create layers and layers of stuff that disconnect things. I remember when we were looking originally at Lambdas and trying to get the head around that and experimenting, should one Lambda call another? Or should one Lambda not call another? And things like that.

I realized, let's say for now, until we realize we want to do something else, a Lambda should only ever talk to SNS and nothing else. Or SQS or something like that. When one Lambda completes, it's going to track a message somewhere and we need to design these messages to be good so that we can decouple different parts of the process. And so far, that helps too as a constraint. I think very, very few times we have one Lambda calling another. Mostly when we actually need a synchronized response back, and for security reasons, we wanted to isolate something to a single Lambda, but that's effectively just a black box security isolation. Since creating these isolation layers through messages, through queues, through topics, becomes a fundamental part of designing these systems.

I remember speaking at the conference to somebody. I forgot the name of the person who was talking about airline. And he was presenting after me and he said, "Look, I can relate to a lot of what you said." And in the airline community basically they often talk about, apparently, I'm not an airline programmer, he told me that in the airline community, talk about designing the protocol being the biggest challenge. Once you design the protocol between your components, the message is who sends what where, you can recover from almost any other design flaw because it's decoupled so if you've made a mess in one Lambda, you can redesign that Lambda, throw it away, rewrite it, decouple things a different way. If the global protocol is good, you get all the flexibility. If you mess up the protocol for communication, then nothing's going to save you at the end.

Now we have EFS and Lambda can talk to an EFS. Should this Lambda talk directly to an EFS or should this Lambda just send some messages to a topic, and then some other Lambdas that are maybe reserved, maybe more constrained talk to EFS? And again, the platform's evolved quite a lot over the last few years. One thing that is particularly useful in that regard is the SQS FIFO queues that came out last year I think. With Corona ...

Jeremy: Yeah, whenever it was.

Gojko: Yeah, I don't remember if it was last year or two years ago. But one of the things it allows us to do is really run lots and lots of Lambdas in parallel where you can guarantee that no two Lambdas access the same kind of business entity that you have in the same type. For example, for this mind mapping thing, we have lots and lots of people modifying lots and lots of files in parallel, but we need to aggregate a single map. If we have 50 people over here working with a single map and 60 people on a map working a different map, aggregation can run in parallel but I never ever, ever want two people modifying the same map their aggregation to run in parallel.

And for Lambda, that was a massive challenge. You had to put Kinesis between Lambda and other Lambdas and things like that. Kinesis' provision capacity, it costs a lot, it doesn't auto-scale. But now with SQS FIFO queues, you can just send a message and you can say the kind of FIFO ID is this map ID that we have. Which means that SQS can run thousands of Lambdas in parallel but they'll never run more than one Lambda for the same map idea at the same time. Designing your protocols like that becomes how you decouple one end of your app that's massively scalable and massively parallel, and another end of your app that we have some reserved capacity or limits.

Like for this kind of video thing, the original idea of that was letting me build marketing videos easier and I can't get rid of this accent. Unfortunately, everything I do sounds like I'm threatening someone to blackmail them. I'm like a cheap Bond villain, and that's not good, but I can't do anything else. I can pay other people to do it for me and we used to do that, but then that becomes a big problem when you want to modify tiny things. We paid this lady to professionally record audio for a marketing video that we needed and then six months later, we wanted to change one screen and now the narration is incorrect. And we paid the same woman again. Same equipment, same person, but the sound is totally different because two different equipment.

Jeremy: Totally different, right.

Gojko: You can't just stitch it up. Then you end up like, okay, do we go and pay for the whole thing again? And I realized the neurometric text-to-speech has learned so much that it can do English better than I can. You're a native English speaker so you can probably defeat those machines, but I can't.

Jeremy: I don't know if I could. They're pretty good now. It's kind of scary.

Gojko: I started looking at one like why don't they just put stuff in a Markdown and use Markdown to generate videos and things like that? All of these things, you get quota limits still. I thought we were limited on Google. Google gave us something like five requests per second in parallel, and it took me a really long time to even raise these quotas and things like that. I don't want to have lots of people requesting stuff and then in parallel trashing this other thing over there. We need to create these layers of running things in a decent limit, and I think that's where I think designing the protocol for this distributed system becomes an importance.

Jeremy: I want to go back because I think you bring up a really good point just about a different type of architecture, or the architectural design of decoupling systems and these event-driven things. You mentioned a Lambda function processes something and sends it to SQS or sends it to SNS to it can do a fan-out pattern or in the case of the FIFO queue, doing an ordered pattern for sequential processing, which those were all great patterns. And even things that AWS has done, such as add things like Lambda destination. Now if you run an asynchronous Lambda function, you still have to write some code or you used to have to write some code that said, "When this is finished processing, now call some other component." And there's just another opportunity for failure there. They basically said, "Well, if it succeeds, then you can actually just forward it off to one of these other services automatically and we'll handle all of the retries and all the failures and that kind of stuff."

And those things have been added in to basically give you that warm and fuzzy feeling that if an event doesn't reach where it's supposed to go, that some sort of cloud trickery will kick in and make sure that gets processed. But what that is introduced I think is a cognitive overload for a lot of developers that are designing these systems because you're no longer just writing a script that does X, Y, and Z and makes a few database calls. Now you're saying, okay, I've got to write a script that can massively scale and take the transactions that I need to maybe parallelize or that I maybe need to queue or delay or throttle or whatever, and pass those down to another subsystem. And then that subsystem has to pick those up and maybe that has to parallelize those or maybe there are failure modes in there and I've got all these other things that I have to think about.

Just that effect on your average developer, I think you and I think about these things. I would consider myself to be a cloud architect, if that's a thing. But essentially, do you see this being I guess a wall for a lot of developers and something that really requires quite a bit of education to ramp them up to be able to start designing these systems?

Gojko: One of the topics we touched upon is the cyclic nature of things, and I think we're going back to where moving from apps working on a single machine to client server architectures was a massive brain melt for a lot of people, and three-tier architectures, which is later, we're not just client server, but three-tier architectures ended up with their own host of problems and then design problems and things like that. That's where a lot of these architectural patterns and design patterns emerged like circuit breakers and things like that. I think there's a whole body of knowledge there for people to research. It's not something that's entirely new and I think you can get started with Lambda quite easily and not necessarily make a mess, but make something that won't necessarily scale well and then start improving it later.

That's why I was mentioning that earlier in the discussion where, as long as the protocol makes sense, you can salvage almost anything late. Designing that protocol is important, but then we're going to good software design. I think teaching people how to do that is something that every 10 years, we have to recycle and reinvent and figure it out because people don't like to read books from more than 10 years ago. All of this stuff like designing fault tolerance systems and fail-safe systems, and things like that. There's a ton of books about that from 20 years ago, from 10 years ago. Amazon, for people listening to you and me, they probably use Amazon more for compute than they use for getting books. But Amazon has all these books. Use it for what Amazon was originally intended for and then get some books there and read through this stuff. And I think looking at design of distributed systems and stuff like that becomes really, really critical for Lambdas.

Jeremy: Yeah, definitely. All right, we've got a few minutes left and I'd love to go back to something we were talking about a little bit earlier and that was everything moving onto a few of these major cloud providers. And one of the things, you've got scale. Scale is a problem when we talked about oh, we can spin up as many VMs as we want to, and now with serverless, we have unlimited capacity really. I know we didn't say that, but I think that's the general idea. The cloud just provides this unlimited capacity.

Gojko: Until something else decides it's not unlimited.

Jeremy: And that's my point here where every major cloud provider that I've been involved with and I've heard the stories of, where you start to move the needle at all, there's always an SA that reaches out to you and really wants to understand what your usage is going to be, and what your patterns were going to be. And that's because they need to make sure that where you're running your applications, that they provision enough capacity because there is not enough capacity, or there's not unlimited capacity in the cloud.

Gojko: It's physically limited. There's only so much buildings where you can have data centers on the surface of Earth.

Jeremy: And I guess that's where my question comes in because you always hear these things about lock-in. Like, well serverless, if you use Lambda, you're going to be locked in. And again, if you're using Oracle, you're locked in. Or, you're using MySQL you're locked in. Or, you're using any of the other things, you're locked in.

Gojko: You're actually not locked in physically. There's a key and a lock.

Jeremy: Right, but this idea of being locked in not to a specific cloud provider, but just locked into a cloud in general and relying on the cloud to do that scaling for you, where do you think the limitations there are?

Gojko: I think again, going back to cyclic, cyclic, cyclic. The PC revolution started when a lot more edge compute was needed on mainframes, and people wanted to get stuff done on their own devices. And I think probably, if we do ever see the limitations of this and it goes into a next cycle, my best guess it's going to be driven by lots of tiny devices connected to a cloud. Not necessarily computers as we know computers today. I pulled out some research preparing for this from IDC. They are predicting basically from 18.3 zettabytes of data needed for IOT in 2019, to be 73.1 zettabytes by 2025. That's like times three in a space of six years. If you went to Amazon now and told them, "You need to have three times more data space in three years," I'm not sure how they would react to that.

This stuff, everything is taking more and more data, and everything is more and more connected to the cloud. The impact of something like that going down now is becoming totally crazy. There was a case in 2017 where S3 started getting a bit more latency than usual in U.S. East 1, in I think February of 2018, or something like that. There were cases where people couldn't turn the lights on in their houses because the management software was working on S3 and depending on S3. Expecting S3 to be indestructible. Last year, in November, Kinesis pretty much went offline as far as everybody else outside AWS concerend for about 15 hours I think. There were people on Twitter that they can't go back into their house because their smart lock is no longer that smart.

And I think we are getting to places where there will be more need for compute on the edge. First of all, there's going to be a lot more demand for data centers and cloud power and I think that's going to keep going on for the next five, ten years. But then people will realize they've hit some limitation of that, and they're going to start moving towards the edge. And we're going from mainframe back into client server computing I think. We're getting these products now. I assume most of your listeners have seen one like all these fancy ubiquity Wi-Fi thingies that are costing hundreds of dollars and they look like pieces of furniture that's just sitting discretely on the wall. And there was a massive security breach yesterday published. Somebody took their AWS keys and took all the customer data and everything.

The big advantage over all the ugly routers was that it's just like a thin piece of glass that sits on your wall, and it's amazing and it looks good, but the reason why they could do a very thin piece of glass is the minimal amount of software is running on that piece of glass, the rest is running the cloud. It's not just locking in terms of is it on Amazon or Google, it's that it's so tightly coupled with something totally outside of your home, where your network router needs Amazon to be alive now in a very specific region of Amazon where everybody's been deploying for the last 15 years, and it's running out of capacity very often. Not very often but often enough.

There's some really interesting questions that I guess we'll answer in the next five, ten years. We're on the verge of IOT I think exploding because people are trying to come up with these new products that you wouldn't even think before that you'd have smart shoes and smart whatnot. Smart glasses and things like that. And when that gets into consumer technology, we're no longer going to have five or ten computer devices per person, we'll have dozens and dozens of computing. I guess think about it this way, fifteen years ago, how many computer devices were you carrying with yourself? Probably mobile phone and laptop. Probably not more. Now, in the headphones you have there that's Bose ...

Jeremy: Watch.

Gojko: ... you have a microprocessor in the headphones, you have your watch, you have a ton of other stuff carrying with you that's low-powered, all doing a bit of processing there. A lot of that processing is probably happening on the cloud somewhere.

Jeremy: Or, it's just sending data. It's just sending, hey here's the information. And you're right. For me, I got my Apple Watch, my thermostat is connected to Wi-Fi and to the cloud, my wife just bought a humidifier for our living room that is connected to Wi-Fi and I'm assuming it's sending data to the cloud. I'm not 100% sure, but the question is, I don't know why we need to keep track of the humidity in my living room. But that's the kind of thing too where, you mentioned from a security standpoint, I have a bunch of AWS access keys on my computer that I send over the network, and I'm assuming they're secure. But if I've got another device that can access my network and somebody hacked something on the cloud side and then they can get in, it gets really dangerous.

But you're right, the amount of data that we are now generating and compute that we're using in the cloud for probably some really dumb things like humidity in my living room, is that going to get to a point where... You said there's going to be a limitation like five years, ten years, whatever it is. What does the cloud do then? What does the cloud do when it can no longer keep up with the pace of these IOT devices?

Gojko: Well, if history is repeating and we'll see if history is repeating, people will start getting throttled and all of a sudden, your unlimited supply of Lambdas will no longer be unlimited supply of Lambdas. It will be something that you have to reserve upfront and pay upfront, and who knows, we'll see when we get there. Or, we get things that we have with power networks like you had a Texas power cut there that was completely severe, and you get a IT cut. I don't know. We'll see. The more we go into utility, the more we'll start seeing parallels between compute and power networks. And maybe power networks are something that you can look at and later name. That's why I think the next cycle is probably going to be some equivalent of client server computing reemerging.

Jeremy: Yeah. All right, well, I got one more question for you and this is just something where it may be a little bit of a tongue-in-cheek question. Because we talked it a little bit ... we talked about the merging of Lambda, and of Fargate, and some of these other things. But just from your perspective, serverless in five years from now, where do you see that going? Do you see that just becoming the main ... This idea of utility computing, on-demand computing without setting up servers and managing ops and some of these other things, do you see that as the future of serverless and it just becoming just the way we build applications? Or do you think that it's got a different path?

Gojko: There was a tweet by Simon Wardley. You mentioned Simon Wardley earlier in the talk. There was a tweet a few days ago where he mentioned some data. I'm not sure where he pulled it from. This might be unverified, but generally Simon knows what he's talking about. Amazon itself is deploying roughly 50% of all new apps they're building on serverless. I think five years from now, that way of running stuff, I'm not sure if it's Lambda or some new service that Amazon starts and gives it some even more confusing name that runs in parallel to everything. But, that kind of stuff where the operator takes care of all the ops, which they really should be doing, is going to be the default way of getting utility compute out.

I think a lot of these other things will probably remain useful for specialists' use cases, where you can't really deploy it in that way, or you need more stability, or it's not transient and things like that. My best guess is first of all, we'll get Lambda's that run for longer, and I assume that after we get Lambdas that run for longer, we'll probably get some ways of controlling routing to Lambdas because you already can set up pre-provisioned Lambdas and hot Lambdas and reserved capacity and things like that. When you have reserved capacity and you have longer running Lambdas, the next logical thing there is to have session stickiness, and routing, and things like that. And I think we'll get a lot of the stuff that was really complicated to do earlier, and you had to run EC2 instances or you had to run complicated networks of services, you'll be able to do in Lambda.

And Lambda is, I wouldn't be surprised if they launch a totally new service with some AWS cloud socket, whatever. Something that is a implementation of the same principle, just in a different way, that becomes a default we are running computer for lots of people. And I think GPUs are still a bit limited. I don't think you can run GPU utility anywhere now, and that's limiting for a whole host of use cases. And I think again, it's not like they don't have the technology to do it, it's just they probably didn't get around to doing it yet. But I assume in five years time, you'll be able to do GPUs on-demand, and processing GPUs, and things like that. I think that the buzzword itself will lose really any special meaning and that's going to just be a way of running stuff.

Jeremy: Yeah, absolutely. Totally agree. Well, listen Gojko, thank you so much for spending the time chatting with me. Always great to talk with you.

Gojko: You, too.

Jeremy: If people want to get in touch with you, find out more about what you're doing, how do they do that?

Gojko: Well, I'm very easy to find online because there's not a lot of people called Gojko. Type Gojko into Google, you'll find me. And gojko.networks, gojko.com works, gojko.org works, and all these other things. I was lucky enough to get all those domains.

Jeremy: That's G-O-J-K-O ...

Gojko: Yes, G-O-J-K-O.

Jeremy: ... for people who need the spelling.

Gojko: Excellent. Well, thanks very much for having me, this was a blast.

Jeremy: All right, yeah. And make sure you check out ... You mentioned Narakeet. It's a speech thing?

Gojko: Yeah, for developers that want to build videos without hassle, and want to put videos in continuous integration, and things like that. Narakeet, that's like parakeet with an N for narration. Check that out and thanks for plugging it.

Jeremy: Awesome. And then, check out MindMup as well. Awesome stuff. I've got all the stuff in the show notes. Thanks again, Gojko.

Gojko: Thank you. Bye-bye.

View Details

About Alexa Abbas

Alexandra Abbas is a Google Cloud Certified Data Engineer & Architect and Apache Airflow Contributor. She currently works as a Machine Learning Engineer at Wise. She has experience with large-scale data science and engineering projects. She spends her time building data pipelines using Apache Airflow and Apache Beam and creating production-ready Machine Learning pipelines with Tensorflow.

Alexandra was a speaker at Serverless Days London 2019 and presented at the Tensorflow London meetup.

Personal links

Twitter: https://twitter.com/alexandraabbas
LinkedIn: https://www.linkedin.com/in/alexandraabbas
GitHub: https://github.com/alexandraabbas

datastack.tv's linksWeb: https://datastack.tv
Twitter: https://twitter.com/datastacktv
YouTube: https://www.youtube.com/c/datastacktv
LinkedIn: https://www.linkedin.com/company/datastacktv
GitHub: https://github.com/datastacktv
Link to the Data Engineer Roadmap: https://github.com/datastacktv/data-engineer-roadmap

This episode is sponsored by CBT Nuggets: cbtnuggets.com/serverless and Stackery: https://www.stackery.io/

Watch this video on YouTube: https://youtu.be/SLJZPwfRLb8

Transcript
Jeremy: Hi, everyone. I'm Jeremy Daly, and this is Serverless Chats. Today I'm joined by Alexa Abbas. Hey, Alexa, thanks for joining me.

Alexa: Hey, everyone. Thanks for having me.

Jeremy: So you are a machine learning engineer at Wise and also the founder of datastack.tv. So I'd love it if you could tell the listeners a little bit about your background and what you do at Wise and what datastack.tv is all about.

Alexa: Yeah. So as you said, I'm a machine learning engineer at Wise. So Wise is an international money transfer service. We are aiming for very transparent fees and very low fees compared to banks. So at Wise, basically, designing, maintaining, and developing the machine learning platform, which serves data scientists and analysts, so they can train their models and deploy their models, easily.

Datastack.tv is, basically, it's a video service or a video platform for data engineers. So we create bite-sized videos, educational videos, for data engineers. We mostly cover open source topics, because we noticed that some of the open source tools in the data engineering world are quite underserved in terms of educational content. So we create videos about those.

Jeremy: Awesome. And then, what about your background?

Alexa: So I actually worked as a data engineer and machine learning engineer, so I've always been a data engineer or machine learning engineer in terms of roles. I also worked, for a small amount of time, I worked as a data scientist as well. In terms of education, I did a big data engineering Master's, but actually my Bachelor is economics, so quite a mix.

Jeremy: Well, it's always good to have a ton of experience and that diverse perspective. Well, listen, I'm super excited to have you here, because machine learning is one of those things where it probably is more of a buzzword, I think, to a lot of people where every startup puts it in their pitch deck, like, "Oh, we're doing machine learning and artificial intelligence ..." stuff like that. But I think it's important to understand, one, what exactly it is, because I think there's a huge confusion there in terms of what we think of as machine learning, and maybe we think it's more advanced than it is sometimes, as I think there's lower versions of machine learning that can be very helpful.

And obviously, this being a serverless podcast, I've heard you speak a number of times about the work that you've done with machine learning and some experiments you've done with serverless there. So I'd love to just pick your brain about that and just see if we can educate the users here on what exactly machine learning is, how people are using it, and where it fits in with serverless and some of the use cases and things like that. So first of all, I think one of the important things to start with anyways is this idea of MLOps. So can you explain what MLOps is?

Alexa: Yeah, sure. So really short, MLOps is DevOps for machine learning. So I guess the traditional software engineering projects, you have a streamlined process you can release, really often, really quickly, because you already have all these best practices that all these traditional software engineering projects implement. Machine learning, this is still in a quite early stage and MLOps is in a quite early stage. But what we try to do in MLOps is we try to streamline machine learning projects, as well as traditional software engineering projects are streamlined. So data scientists can train models really easily, and they can release models really frequently and really easily into production. So MLOps is all about streamlining the whole data science workflow, basically.

And I guess it's good to understand what the data science workflow is. So I talk a bit about that as well. So before actually starting any machine learning project, the first phase is an experimentation phase. It's a really iterative process when data scientists are looking at the data, they are trying to find features and they are also training many different models; they are doing architecture search, trying different architecture, trying different hyperparameter settings with those models. So it's a really iterative process of trying many models, many features.

And then by the end, they probably find a model that they like and that hit the benchmark that they were looking for, and then they are ready to release that model into production. And this usually looks like ... so sometimes they use shadow models, in the beginning, to check if the results are as expected in production as well, and then they actually release into production. So basically MLOps tries to create the infrastructure and the processes that streamline this whole process, the whole life cycle.

Jeremy: Right. So the question I have is, so if you're an ML engineer or you're working on these models and you're going through these iterations and stuff, so now you have this, you're ready to release it to production, so why do you need something like an MLOps pipeline? Why can't you just move that into production? Where's the barrier?

Alexa: Well, I guess ... I mean, to be honest, the thing is there shouldn't be a barrier. Right now, that's the whole goal of MLOps. They shouldn't feel that they need to do any manual model artifact copying or anything like that. They just, I don't know, press a button and they can release to production. So that's what MLOps is about really and we can version models, we can version the data, things like that. And we can create reproducible experiments. So I guess right now, I think many bits in this whole lifecycle is really manual, and that could be automated. For example, releasing to production, sometimes it's a manual thing. You just copy a model artifact to a production bucket or whatever. So sometimes we would like to automate all these things.

Jeremy: Which makes a lot of sense. So then, in terms of actually implementing this stuff, because we hear all the time about CI/CD. If we're talking about DevOps, we know that there's all these tools that are being built and services that are being launched that allow us to quickly move code through some process and get into production. So are there similar tools for deploying models and things like that?

Alexa: Well, I think this space is quite crowded. It's getting more and more crowded. I think there are many ... So there are the cloud providers, who are trying to create tools that help these processes, and there are also many third-party platforms that are trying to create the ML platform that everybody uses. So I think there is no go-to thing that everybody uses, so I think there is many tools that we can use.

Some examples, for example, TensorFlow is a really popular machine learning library, But TensorFlow, they created a package on top of TensorFlow, which is called TFX, TensorFlow Extended, which is exactly for streamlining this process and serving models easily, So I would say it TFX is a really good example. There is Kubeflow, which is a machine learning toolkit for Kubernetes. I think there are many custom implementations in-house in many companies, they create their own machine learning platforms, their own model serving API, things like that. And like the cloud providers on AWS, we have SageMaker. They are trying to cover many parts of the tech science lifecycle. And on Google Cloud, we have AI Platform, which is really similar to SageMaker.

Jeremy: Right. And what are you doing at Wise? Are you using one of those tools? Are you building something custom?

Alexa: Yeah, it's a mix actually. We have some custom bits. We have a custom API, serving API, for serving models. But for model training, we are using many things. We are using SageMaker, Notebooks. And we are also experimenting with SageMaker endpoints, which are actually serverless model serving endpoints. And we are also using EMR for model training and data preparation, so some Spark-based things, a bit more traditional type of model training. So it's quite a mix.

Jeremy: Right. Right. So I am not well-versed in machine learning. I know just enough to be dangerous. And so I think that what would be really interesting, at least for me, and hopefully be interesting to listeners as well, is just talk about some of these standard tools. So you mentioned things like TensorFlow and then Kubeflow, which I guess is that end-to-end piece of it, but if you're ... Just how do you start? How do you go from, I guess, building and training a model to then productizing it and getting that out? What's that whole workflow look like?

Alexa: So, actually, the data science workflow I mentioned, the first bit is that experimentation, which is really iterative, really free, so you just try to find a good model. And then, when you found a good model architecture and you know that you are going to receive new data, let's say, I don't know, I have a day, or whatever, I have a week, then you need to build out a retraining pipeline. And that is, I think, what the productionization of a model really means, that you can build a retraining pipeline, which can automatically pick up new data and then prepare that new data, retrain the model on that data, and release that model into production automatically. So I think that means productionization really.

Jeremy: Right. Yeah. And so by being able to build and train a model and then having that process where you're getting that feedback back in, is that something where you're just taking that data and assuming that that is right and fits in the model or is there an ongoing testing process? Is there supervised learning? I know that's a buzzword. I'm not even sure what it means. But those ... I mean, what types of things go into that retraining of the models? Is it something that is just automatic or is it something where you need constant, babysitting's probably the wrong word, but somebody to be monitoring that on a regular basis?

Alexa: So monitoring is definitely necessary, especially, I think when you trained your model and you shouldn't release automatically in production just because you've trained a new data. I mentioned this shadow model thing a bit. Usually, after you retrained the model and this retraining pipeline, then you release that model into shadow mode; and then you will serve that model in parallel to your actual product production model, and then you will check the results from your new model against your production model. And that's a manual thing, you need to ... or maybe you can automate it as well, actually. So if it performs like ... If it is comparable with your production model or if it's even better, then you will replace it.

And also, in terms of the data quality in the beginning, you should definitely monitor that. And I think that's quite custom, really depends on what kind of data you work with. So it's really important to test your data. I mean, there are many ... This space is also quite crowded. There are many tools that you can use to monitor your distribution of your data and see that the new data is actually corresponds to your already existing data set. So there are many bits that you can monitor in this whole retraining pipeline, and you should monitor.

Jeremy: Right. Yeah. And so, I think of some machine learning like use cases of like sentiment analysis, for example... looking at tweets or looking at customer service conversations and trying to rate those things. So when you say monitoring or running them against a shadow model, is that something where ... I mean, how do you gauge what's better, right? if you've got a shadow... I mean, what's the success metric there as to say X number were classified as positive versus negative sentiment? Is that something that requires human review or some sampling for you to kind of figure out the quality of the success of those models?

Alexa: Yeah. So actually, I think that really depends on the use case. For example, when you are trying to catch fraudsters, your false positive rate and true positive rate, these are really important. If your true positive rate is higher that means, oh, you are catching more fraudsters. But let's say your new model, with your model, also the false positive rate is higher, which means that you are catching more people who are actually not fraudsters, but you have more work because I guess that's a manual process to actually check those people. So I think it really depends on the use case.

Jeremy: Right. Right. And you also said that the markets a little bit flooded and, I mean, I know of SageMaker and then, of course, there's all these tools like, what's it called, Recognition, a bunch of things at AWS, and then Google has a whole bunch of the Vision API and some of these things and Watson's Natural Language Processing over at IBM and some of these things. So there's all these different tools that are just available via an API, which is super simple and great for people like me that don't want to get into building TensorFlow models and things like that. So is there an advantage to building your own models beyond those things, or are we getting to a point where with things like ... I mean, again, I know SageMaker has a whole library of models that are already built for you and things like that. So are we getting to a point where some of these models are just good enough off the shelf or do we really still need ... And I know there are probably some custom things. But do we still really need to be building our own models around that stuff?

Alexa: So to be honest, I think most of the data scientists, they are using off-the-shelf models, maybe not the serverless API type of models that Google has, but just off-the-shelf TensorFlow models or SageMaker, they have these built-in containers for some really popular model architectures like XGBoost, and I think most of the people they don't tweak these, I mean, as far as I know. I think they just use them out of the box, and they really try to tweak the data instead, the data that they have, and try to have these off-the-shelf models with higher and higher quality data.

Jeremy: So shape the data to fit the model as opposed to the model to fit the data.

Alexa: Yeah, exactly. Yeah. So you don't actually have to know ... You don't have to know how those models work exactly. As long as you know what the input should be and what output you expect, then I think you're good to go.

Jeremy: Yeah, yeah. Well, I still think that there's probably a lot of value in tuning the models though against your particular data sets.

Alexa: Yeah, right. But also there are services for hyperparameter tuning. There are services even for neural architecture search, where they try a lot of different architectures for your data specifically and then they will tell you what is the best model architecture that you should use and same for the hyperparameter search. So these can be automated as well.

Jeremy: Yeah. Very cool. So if you are hosting your own version of this ... I mean, maybe you'll go back to the MLOps piece of this. So I would assume that a data scientist doesn't want to be responsible for maintaining the servers or the virtual machines or whatever it is that it's running on. So you want to have this workflow where you can get your models trained, you can get them into production, and then you can run them through this loop you talked about and be able to tweak them and continue to retrain them as things go through. So on the other side of that wall, if we want to put it that way, you have your ops people that are running this stuff. Is there something specific that ops people need to know? How much do they need to know about ML, as opposed to ... I mean, the data scientists, hopefully, they know more. But in terms of running it, what do they need to know about it, or is it just a matter of keeping a server up and running?

Alexa: Well, I think ... So I think the machine learning pipelines are not yet as standardized as a traditional software engineering pipeline. So I would say that you have to have some knowledge of machine learning or at least some understanding of how this lifecycle works. You don't actually need to know about research and things like that, but you need to know how this whole lifecycle works in order to work as an ops person who can automate this. But I think the software engineering skills and DevOps skills are the base, and then you can just build this knowledge on top of that. So I think it's actually quite easy to pick this up.

Jeremy: Yeah. Okay. And what about, I mean, you mentioned this idea of a lot of data scientists aren't actually writing the models, they're just using the preconfigured model. So I guess that begs the question: How much does just a regular person ... So let's say I'm just a regular developer, and I say, "I want to start building machine learning tools." Is it as easy as just pulling a model off the shelf and then just learning a little bit more about it? How much can the average person do with some of these tools out of the box?

Alexa: So I think most of the time, it's that easy, because usually the use cases that someone tries to tackle, those are not super edge cases. So for those use cases, there are already models which perform really well. Especially if you are talking about, I don't know, supervised learning on tabular data, I think you can definitely find models that will perform really well off the shelf on those type of datasets.

Jeremy: Right. And if you were advising somebody who wanted to get started... I mean, because I think that I think where it might come down to is going to be things like pricing. If you're using Vision API and you're maybe limited on your quota, and then you can ... if you're paying however many cents per, I guess, lookup or inference, then that can get really expensive as opposed to potentially running your own model on something else. But how would you suggest that somebody get started? Would you point them at the APIs or would you want to get them up and running on TensorFlow or something like that?

Alexa: So I think, actually, for a developer, just using an API would be super easy. Those APIs are, I think ... So getting started with those APIs just to understand the concepts are very useful, but I think getting started with Tensorflow itself or just Keras, I definitely I would recommend that, or just use scikit-learn, which is a more basic package for more basic machine learning. So those are really good starting points. And there are so many tutorials to get started with, and if you have an idea of what you would like to build, then I think you will definitely find tutorials which are similar to your own use case and you can just use those to build your custom pipeline or model. So I would say, for developers, I would definitely recommend jumping into TensorFlow or scikit-learn or XGBoost or things like that.

Jeremy: Right, right. And how many of these models exist? I mean, are we talking there's 20 different models or are we talking there's 20,000 models?

Alexa: Well, I think ... Wow. Good question. I think we are more towards today maybe not 20,000, but definitely many thousands, I think. But there are popular models that most of the people use, and I think there are maybe 50 or 100 models that are the most popular and most companies use them and you are probably fine just using those for any use case or most of the use cases.

Jeremy: Right. Now, and speaking of use cases, so, again, I try to think of use cases or machine learning and whether it's classifying movies into genres or sentiment analysis, like I said, or maybe trying to classify news stories, things like that. Fraud detection, you mentioned. Those are all great use cases, but what are ... I know you've worked on a bunch of projects. So what are some of the projects that you've done and what were the use cases that were being solved there, because I find these to be really interesting?

Alexa: Yeah. So I think a nice project that I worked on was a project with Lush, which is a cosmetics company. They manufacture like soaps and bath bombs. And they have this nice mission that they would like to eliminate packaging from their shops. So they asked us, when I worked at Datatonic, we worked on a small project with them. They asked us to create an image recognition model, to train one, and then create a retraining pipeline that they can use afterwards. So they provided us with many hundred thousand images of their products, and they made photos from different angles with different lightings and all of that, so really high-quality image data set of all their products.

And then, we used a mobile net model, because they wanted this model to be built-in into their mobile application. So when users actually use this model, they download it with their mobile application. And then, they created a service called Lush [inaudible], which you can use from within their app. And then, people can just scan the products and they can see the ingredients and how-to-use guides and things like that. So this is how they are trying to eliminate all kinds of packaging from their shops, that they don't actually need to put the papers there or put packaging with ingredients and things like that.

And in terms of what we did on the technical side, so as I mentioned, we used a mobile net model, because we needed to quantize the model in order to put it on a mobile device. And we used TF Lite to do this. TF Lite is specifically for models that you want to run on an edge device, like a mobile phone. So that was already a constraint. So this is how we picked a model. I think, back then, like there were only a few model architectures supported by TF Lite, and I think there were only two, maybe. So we picked MobileNet, because it had a smaller size.

And then, in terms of the retraining, so we automated the whole workflow with Cloud Composer on Google Cloud, which is a managed version of Apache Airflow, the open source scheduling package. The training happened on AI Platform, which is Google Cloud's SageMaker.

Jeremy: Yeah.

Alexa: Yeah. And what else? We also had an image pre-processing step just before the training, which happened on Dataflow, which is an auto-scaling processing service on Google Cloud. And after we trained the model, we just saved the model active artifact in a bucket, and then ... I think we also monitored the performance of the model, and if it was good enough, then we just shipped the model to developers who actually they manually updated the model file that went into the application that people can download. So we didn't really see if they use any shadow model thing or anything like that.

Jeremy: Right. Right. And I think that is such a cool use case, because, if I'm hearing you right, there were just like a bar soap or something like that with no packaging, no nothing, and you just hold your mobile phone camera up to it or it looks at it, determines which particular product is, gives you all that ... so no QR codes, no bar codes, none of that stuff. How did they ring them up though? Do you know how that process worked? Did the employees just have to know what they were or did the employees use the app as well to figure out what they were billing people for?

Alexa: Good question. So I think they wanted the employees as well to use the app.

Jeremy: Nice.

Alexa: Yeah. But when the app was wrong, then I don't know what happened.

Jeremy: Just give them a discount on it or something like that. That's awesome. And that's the thing you mentioned there about ... Was it Tensor Lite, was it called?

Alexa: TF Lite. Yeah.

Jeremy: TF Lite. Yes. TensorFlow Lite or TF Lite. But, basically, that idea of being able to really package a model and get it to be super small like you said. You said edge devices, and I'm thinking serverless compute at the edge, I'm thinking Lambda functions. I'm thinking other ways that if you could get your models small enough in package, that you could run it. But that'd be a pretty cool way to do inference, right? Because, again, even if you're using edge devices, if you're on an edge network or something like that, if you could do that at the edge, that'd be a pretty fast response time.

Alexa: Yeah, definitely. Yeah.

Jeremy: Awesome. All right. So what about some other stuff that you've done? You've mentioned some things about fraud detection and things like that.

Alexa: Yeah. So fraud detection is a use case for Wise. As I mentioned, Wise services international money transfer, one of its services. So, obviously, if you are doing anything with money, then a full use case is for sure that you will have. So, I mean, in terms of ... I don't actually develop models at Wise, so I don't know actually what models they use. I know that they use H2O, which is a Spark-based library that you can use for model training. I think it's quite an advanced library, but I haven't used it myself too much, so I cannot talk about that too much.

But in terms of the workflow, it's quite similar. We also have Airflow to schedule the retraining of the models. And they use EMR for data preparation, so quite similar to Dataflow, in a sense. A Spark-based auto-scaling cluster that processes the data and then, they train the models on EMR as well but using this H2O library. And then in the end, when they are happy with the model, we have this tool that they can use for releasing shadow models in production. And then, if they are satisfied with the performance of the model that they can actually release into production. And at Wise, we have a custom micro service, a custom API, for serving models.

Jeremy: Right. Right. And that sounds like you need a really good MLOps flow to make all that stuff work, because you just have a lot of moving parts there, right?

Alexa: Yeah, definitely. Also, I think we have many bits that could be improved. I think there are many bits that still a bit manual and not streamlined enough. But I think most of the companies struggle with the same thing. It's just we don't yet have those best practices that we can implement, so many people try many different things, and then ... Yeah, so I think it's still a work in progress.

Jeremy: Right. Right. And I'm curious if your economics background helps at all with the fraud and the money laundering stuff at all?

Alexa: No.

Jeremy: No. All right. So what about you worked in another data engineering project for Vodafone, right?

Alexa: Yeah. Yeah, so that was a data engineering project purely, so we didn't do any machine learning. Well, Vodafone has their own Google Analytics library that they use in all their websites and mobile apps and things like that and that sense Clickstream data to a server in a Google Cloud Platform Project, and we consume that data in a streaming manner from data flows. So, basically, the project was really about processing this data by writing an Apache Beam pipeline, which was always on and always expected messages to come in. And then, we dumped all the data into BigQuery tables, which is data warehouse in Google Cloud. And then, these BigQuery tables powered some of the dashboards that they use to monitor the uptime and, I don't know, different metrics for their websites and mobile apps.

Jeremy: Right. But collecting all of that data is a good source for doing machine learning on top of that, right?

Alexa: Yeah, exactly. Yeah. I think they already had some use cases in mind. I'm not sure if they actually done those or not, but it's a really good base for machine learning, what we collected the data there in BigQuery, because that is an analytical data warehouse, so some analysts can already start and explore the data as a first step of the machine learning process.

Jeremy: Right. I would think anomaly detection and things like that, right?

Alexa: Yeah, exactly.

Jeremy: Right. All right. Well, so let's go on and talk about serverless a little bit more, because I know I saw you do a talk where you were you ran some experiments with serverless. And so, I'm just kind of curious, where are the limitations that you see? And I know that there continues ... I mean, we now have EFS integration, and we've got 10 gigs of memory for lambda functions, you've even got Cloud Run, which I don't know how much you could do with that, but where's still some of the limitations for running machine learning in a serverless way, I guess?

Alexa: So I think, actually, from this data science lifecycle, many bits, there are Cloud providers offer a lot of serverless options. For data preparation, there is Dataflow, which is, I think, kind of like serverless data processing service, so you can use that for data processing. For model training, there is ... Or the SageMaker and AI Platform, which are kind of serverless, because you don't actually need to provision these clusters that you train your models on. And for model serving, in SageMaker, there are the serverless model endpoints that you can deploy. So there are many options, I think, for serverless in the machine learning lifecycle.

In my experience, many times, it's a cost thing. For example, at Wise, we have this custom model serving API, where we serve all our models. And if they would use SageMaker endpoints, I think, a single SageMaker endpoint is about $50 per month, that's the minimum price, and that's for a single model and a single endpoint. And if you have thousands of models, then your price can go up pretty quickly, or maybe not thousands, but hundreds of models, then your price can go up pretty quickly. So I think, in my experience, limitation could be just price.

But in terms of ... So I think, for example, if I compare Dataflow with a spark cluster that you program yourself, then I would definitely go with Dataflow. I think it's just much easier and maybe cost-wise as well, you might be better off, I'm not sure. But in terms of comfort and developer experience, it's a much better experience.

Jeremy: Right. Right. And so, we talked a little bit about TF Lite there. Is that something possible where maybe the training piece of it, running that on Functions as a Service or something like that maybe isn't the most efficient or cost-effective way to do that, but what about running models or running inference on something like a Lambda function or a Google Cloud function or an Azure function or something like that? Is it possible to package those models in a way that's small enough that you could do that type of workload?

Alexa: I think so. Yeah. I think you can definitely make inference using a Lambda function. But in terms of model training, I think that's not a ... Maybe there were already experiments for, I'm sure there were. But I think it's not the kind of workload that would fit for Lambda functions. That's a typical parallelizable, really large-scale workloads for ... You know the MapReduce type of data processing workloads? I think those are not necessarily fit for Lambda functions. So I think for model training and data preparation, maybe those are not the best options, but for model inference, definitely. And I think there are many examples using Lambda functions for inference.

Jeremy: Right. Now, do you think that ... because this is always something where I find with serverless, and I know you're more of a data scientist, ML expert, but I look at serverless and I question whether or not it needs to handle some of these things. Especially with some of the endpoints that are out there now, we talked about the Vision API and some of the other NLP things, are we putting in too much effort maybe to try to make serverless be able to handle these things, or is it just something where there's a really good way to handle these by hosting your ... I mean, even if you're doing SageMaker, maybe not SageMaker endpoints, but just running SageMaker machines to do it or whatever, are we trying too hard to squeeze some of these things into a serverless environment?

Alexa: Well, I don't know. I think, as a developer, I definitely prefer the more managed versions of these products. So the less I need to bother with, "Oh, my cluster died and now we need to rebuild a cluster of things," and I think serverless can definitely solve that. I would definitely prefer the more managed version. Maybe not serverless, because, for some of the use cases or some of the bits from the lifecycle, serverless is not the best fit, but a managed product is definitely something that I prefer over a non-managed product.

Jeremy: Right. And so, I guess one last question for you here, because this is something that always interests me. Just there are relevant things that we need machine learning for. I mean, I think the fraud detection is a hugely important one. Sentiment analysis, again. Some of those other things are maybe, I don't know, I shouldn't call them toy things, but personalization and some of the things, they're all really great things to have, and it seems like you can't build an application now without somebody wanting some piece of that machine learning in there. So do you see that as where we are going where in the future, we're just going to have more of these APIs?

I mean, out of AWS, because I'm more familiar with the AWS ecosystem, but they have Personalize and they have Connect and they have all these other services, they have the recommendation engine thing, all these different services ... Lex, or whatever, that will read text, natural language processing and all that kind of stuff. Is that where we're moving to just all these pre-trained, canned products that I can just access via an API or do you think that if you're somebody getting started and you really want to get into the ML world that you should start diving into the TensorFlows and some of those other things?

Alexa: So I think if you are building an app and your goal is not to become an ML engineer or a data scientist, then these canned models are really useful because you can have a really good recommendation engine in your product, you could have really good personalization engine in your product, things like that. And so, those are, I think, really useful and you don't need to know any machine learning in order to use them. So I think we definitely go into that direction, because most of the companies won't hire data scientists just to train a recommender model. I think it's just easier to use an API endpoint that is already really good.

So I think, yeah, we are definitely heading into that direction. But if you are someone who wants to become a data scientist or wants to be more involved with MLOps or machine learning engineering, then I think jumping into TensorFlow and understanding, maybe not, as we discussed, not getting into the model architectures and things like that, but just understanding the workflow and being able to program a machine learning pipeline from end to end, I think that's definitely recommended.

Jeremy: All right. So one last question: If you've ever used the Watson NLP API or the Google Vision API, can you put on your resume that you're a machine learning expert?

Alexa: Well, if you really want to do that, I would give it a go. Why not?

Jeremy: All right. Good. Good to know. Well, Alexa, thank you so much for sharing all this information. Again, I find the use cases here to be much more complex than maybe some of the surface ones that you sometimes hear about. So, obviously, machine learning is here to stay. It sounds like there's a lot of really good opportunities for people to start kind of dabbling in it and using that without having to become a machine learning expert. But, again, I appreciate your expertise. So if people want to find out more about you or more about the things you're working on and datastack.tv, things like that, how do they do that?

Alexa: So we have a Twitter page for datastack.tv, so feel free to follow that. I also have a Twitter page, feel free to follow me, account, not page. There is a datastack.tv website, so it's just datastack.tv. You can go there, and you can check out the courses. And also, we have created a roadmap for data engineers specifically, because there was no good roadmap for data engineers. I definitely recommend checking that out, because we listed most of the tools that a data engineer and also machine learning engineer should know about. So if you're interested in this career path, then I would definitely recommend checking that out. So under datastack.tv's GitHub, there is a roadmap that you can find.

Jeremy: Awesome. All right. And that's just, like you said, datastack.tv.

Alexa: Yes.

Jeremy: I will make sure that we get your Twitter and LinkedIn and GitHub and all that stuff in there. Alexa, thank you so much.

Alexa: Thanks. Thank you.

View Details

About Jason McGee

Jason McGee, IBM Fellow, is VP and CTO at IBM Cloud Platform. Jason is currently responsible for technical strategy and architecture for all of IBM’s Cloud Platform, across public, dedicated, and local delivery models. Previously Jason has served as CTO of Cloud Foundation Services, Chief Architect of PureApplication System, WebSphere Extended Deployment, WebSphere sMash, and WebSphere Application Server on distributed platforms.

  • Twitter: @jrmcgee
  • LinkedIn: https://www.linkedin.com/in/jrmcgee/
  • IBM Cloud Code Engine: Learn more during this live virtual event on April 14th (also available on-demand after April 14th)
  • Read more: https://www.ibm.com/cloud/code-engine
  • Get started today: https://cloud.ibm.com/docs/codeengine?topic=codeengine-getting-started

Watch this episode on YouTube: https://youtu.be/yH_mgW2kGzU

This episode sponsored by IBM Cloud.

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm joined by Jason McGee. Hey Jason, thanks for joining me.

Jason: Thanks for having me.

Jeremy: So you are an IBM fellow and the VP and CTO of the IBM Cloud platform. So I'd love it if you could tell our guests a little bit about yourself and what it is that you do at IBM.

Jason: Sure. I spend my day at IBM worried about developers and platform services on our public cloud. So I'm responsible for both the technical strategy and the delivery of our Kubernetes and OpenShift platforms, our serverless environments, and kind of all the things that surround that space, logging, and monitoring and other developer tools that kind of make up the developer platform for IBM Cloud.

Jeremy: And what about yourself? What's your background?

Jason: Been a software, kind of middleware guy, my whole life. I used to be the chief architect for WebSphere app server. So I spent the last 20 plus years working on enterprise application platforms and helping companies be able to build mission-critical business systems.

Jeremy: Awesome. So I had Michael Behrendt on the show not too long ago and it was great. We talked about a whole bunch of different things. IBM's point of view of serverless. We talked a little bit about the future of serverless and we talked about the IBM Cloud Code Engine, which I want to get into, but for the benefit of our listeners and just because I'm so fascinated by some of the things that IBM is doing now with serverless, it's just super interesting. So could you sort of give me your point of view or IBM's point of view on serverless and just sort of refresh the listener's memory sort of about how IBM is thinking about serverless and how they're probably thinking about it maybe differently than some of the other cloud providers?

Jason: Yeah, sure. I mean, it's such a fascinating space and it's really changed a lot, I think, over the last five years or so from its kind of maybe beginnings in being very aligned with serverless functions and kind of event-driven computing and becoming a more general concept about how developers especially can consume cloud platforms. I think if you look at the IBM perspective on serverless, there's a couple layers to the problem that we think about. First is we've been pretty clear that we think Kubernetes and distributions of Kubernetes like OpenShift are kind of the key foundation compute environment for developers to use going forward. And we've done a ton of work in kind of building out our Kubernetes and OpenShift platforms and delivering them as a service on our public cloud. And that's an incredibly flexible platform that you can really build any kind of application. I think over the last five years, we've proven we can run anything on Kubernetes databases and AI and stateless apps and whatever you want.

Jeremy: Right.

Jason: So very, very flexible. However, sometimes flexible also means complicated and it means that there's lots to manage and there's lots of concepts to get your head around. And so we've been thinking a lot about, well, how do you actually consume a platform like Kubernetes more easily? How does the developer stay more focused on what they're really trying to do, which is like build application logic, solve problems? Now they don't really want to stand up coop clusters and configure security policies. They just want to write code and run code and they want to get the power of cloud to do that. Right? And so I think serverless has kind of morphed to be, for us, more about the experience that we can build on top of that container platform that's more oriented around how developers get work done and allows them to kind of more easily take advantage of the scale and power of public clouds without having to kind of take on the burden of a lot of that kind of work and management.

And so the work that we've been doing is really aligned in that direction, that we've been working in projects like Knative, in the open source community to build simpler abstractions on top of Kubernetes. And we've been starting to deliver those in our cloud through things like Code Engine.

Jeremy: Yeah. And I think that's interesting too because I always have, this is probably the wrong way to say it, but it's sort of a chip on my shoulder about Kubernetes because it just got so complicated. Right? It's just so many things that you have to do, so hard to manage. And as a serverless guy myself, I love just the simplicity of being able to write some code and just get it out there, have it auto scale, tie into all those events. So I think that a lot of cloud providers have sort of moved that way to say like, "Well, we're going to manage your Kubernetes cluster for you." Right? Which essentially is just, I think moving backwards, but also moving forwards at the same time, if that makes sense. But so in terms of the use cases that this opens up because now you're not necessarily limited to a sort of bespoke implementation of some serverless platform, you have a lot more capabilities. So what types of use cases does this open up?

Jason: Yeah. I mean, I may have a couple of comments on that. I mean, so I think with Kubernetes, you have the complexity of managing the Kubernetes environment, but even if that's totally taken care of for you, and even if you're using a managed Kubernetes service like the things we offer on IBM Cloud, you still have that kind of resource burden of using Kubernetes. You have services and pods and replica sets and namespaces and all kinds of concepts that you have to kind of wrap your head around and know how to use in the right way. And so there's a value in like, "Can we abstract that? Can we move away from that?" And it's not like this idea hasn't been tried before. I mean, we've had paths platforms, like kind of Cloud Foundry style, Heroku, very opinionated paths environments in the past and they definitely simplify the user experience. However, they came with this negative, which is if you don't fit within the box of the opinion ...

Jeremy: Right.

Jason: ... then you can't do what you want to do. And the cost of going outside the box was super high. Maybe you had to completely switched platforms. You were completely blocked. You to switch to some other approach. And so part of what's informing us and as we think about this is how do you have more of a continuum? You have a simple model. It's aligned around what you're doing. Just run my source code, just run my container image. I want to run a batch job, but it's all running on one platform. They're running next to each other. You can drop down a layer into Kubernetes if you want to. If what you're trying to accomplish needs some of that flexibility, you should have access to it without having to kind of start over. And so that's kind of how we've approached the problem a little bit differently is bringing this all together into kind of one unified serverless environment on top of Kubernetes.

And that lets us handle different use cases. That lets those handle kind of stateless, data processing and functions. That lets us handle simple web apps. That lets us handle very data-intensive, high-scale computation and data processing, async processing like batch all in one combined way.

Jeremy: Right. Yeah. And I think it's interesting because there are artificial limitations may be put in place sometimes on serverless platforms. If you think about AWS Lambda, for example, you get 15 minutes of compute and they bumped things up. So now, and again, I've just sort of grew up in the AWS environment, but they have things like 10 gigs for a function or something like that. And so they've increased these things, but they are sort of artificial limits that I think, depending on the type of workload that you're doing, they can really get in your way, especially if, like you said, you're doing these data-intensive things. So from an IBM perspective, I mean that's sort of gone, right?

Jason: Right. Exactly. That's a great, very concrete way to look at the problem. The approaches that have been taken in some of the other cloud environments is these different use cases like serverless functions, single containers, batch processing, they're different services. And every service has its own kind of limitations or rules about what you can and cannot do. How long your thing can execute, how big your code can be, how much data you can transfer. We've taken a different approach to say, "Let's eliminate all those limits and let's have one logical service, one environment that supports all those styles." We can still expose a simplified kind of consumption model for the developer like just give me your source code or just give me your image, but I can run it in a way that doesn't have those computational limits, and therefore I can do more. Right? I can run more kinds of workloads. I don't run up against some of those walls that kind of stopped me from getting my work done.

Jeremy: Right. Right. Yeah. And I like that approach too because I'm a big fan of managed services. I think that if you have a service that does image recognition for you, that's great. And do you have a service that does queuing for you? That's great. But in some cases, you start stringing together so many different services and I feel like you lose a lot of that control. So I like that idea of just basically being able to say, "Look, I've got the compute. I can do whatever I need to do with it. It will scale to whatever I needed to scale to." And I think that's where this idea of IBM Cloud Code Engine comes in, which just became GA so I'd love it if you could tell the listeners exactly what that is.

Jason: Yeah, absolutely. So, so Code Engine is the new service that we launched that makes some of these concepts I've been talking about real. It is a service that allows developers to deploy functions, containers, source code, batch jobs, into IBM Cloud. The entire environment behind that application is managed for you. So we handle you don't manage clusters, you don't provision infrastructure. You can scale all the way to zero. So you can literally only pay for what you're using. You can scale up to thousands of cores that are in parallel processing your application and we manage that entire runtime environment for you. So you can think of it as a multi-tenant shared Kubernetes-based runtime environment that you can run your workloads on that presents to you the personality that you need for different workloads. And because it's all in one service, if you have an application that's like a mix of some single containers and batch jobs, they can actually talk to each other, they can talk to each other over a private network connection. They can work together instead of being kind of siloed in these completely different environments.

Jeremy: Right? Yeah. And so from the developer, I guess, perspective, you had mentioned that you can deploy just code or you could deploy a container if you want to. So what does that developer experience look like? So is this something where I could just say, "Look, I don't need to have a whole ops team now managing this for me. If I just want to write code, deploy it into these things, I'm sure there's some things I need to know," but for the most part, what does that developer experience look like?

Jason: Yeah. So you absolutely could do it without a whole ops team. The experience right now, there's like maybe kind of three basic entry points. You can give me source code and we will take care of compiling that source code, combining with a runtime, executing it for you, giving it a web end point, scaling it. You can give me some hints about kind of how much resource you think you need and things like that and we can scale that up and down and manage it for you, including all the way down to zero. That's nice if you're coming from maybe a historical paths background or it's just like, "Here's my code, run it for me." You can have that experience with Code Engine. You could also start with a container image. So lots of developers now, because of things like Kubernetes and Docker, are very familiar and comfortable with packaging up their application as a container image, but you don't want to then deal with creating a cluster and dealing with Kubes.

So you can just say like, "Here's my image, run it for me." And one of the advantages we have with Code Engine is we can really do that with any container image. You don't have to have a container image that follows some particular framework that's built in a very special way. We can take any container image and you can just literally point me at the image and say, "Run this for me," and Code Engine will execute it and scale it and manage it for you. Or you can start with a batch job interface. So like a more of an async kind of parallel job submission model. So maybe I'm doing Monte Carlo simulations or data processing and I want to parallelize that across a whole bunch of machines and cores, Code Engine gives you an interface for that. So as a developer, you kind of start with one of those three entry points and let Code Engine take care of how to run that and scale it and keep it highly available and things like that.

Jeremy: Right. So I love the idea of the batch jobs. I want to talk about that a little bit more, but let's go back to some of the use cases here. So what if I was building just like a REST API, that seems to be a very popular, serverless use case, what would I do for that? Do I need to have some sort of an API type gateway type thing in front of it? Or how does that work?

Jason: No, Code Engine provides all that for you. So you would literally either just take your implementation and package it in a container or point us at your source code directory. If you have source code, we use things like Paketo Buildpacks to build a runtime around that source code. And so you can use different languages. So you can either point us, with our CLI tool, you point us at the source code directory and we'll build it and package it in a runtime and run it for you. Or you point us out a container image that you've uploaded to our container registry or to your container registry of choice and then Code Engine will execute that for you. It will give you that web end point, right? So it'll give you a HTTP end point that you can use to access that service. And it will watch the demand on that system and scale it up and down as needed. And by default, we'll just scale it to zero. So it'll just be kind of registered in the system and it'll take care of scaling it up as needed to handle the demand on the app.

Jeremy: All right. Cool. And then what about these batch jobs? So I talked a little bit about this with Michael and this idea of being able to run massively parallel execution. So how does that all work?

Jason: Yeah. So similar, obviously with batch, there's a little bit more kind of metadata that you have to provide to describe the job and what you want to execute and how things relate to each other. So there's some input data you provide along with the implementation of the batch job, which itself could just be like a container image and you submit that job. So the CLI interface is a little bit different. You're not standing up a long-running REST end point, you're submitting a job to Code Engine for execution, and it will go take that job and execute it and parallelize it for you. You can also use Frameworks on top. One of the things we've been doing a lot of work on, maybe Michael talked about it a little bit when he was here, is some work we're doing around Ray. Ray is a really interesting new project that lets you do kind of distributed computing, especially around data workloads in a really easy way.

And so you can actually stand up Ray on top of Code Engine and so Ray acts as kind of the application interface for the developer to be able to easily parallelize their code, particularly Python code, and then Code Engine acts as the runtime below it. And you can take a simple function in Python, mark it as Ray remote and it'll now execute on the cloud and distribute itself across a thousand cores. And you get your answer back 20 times faster than you would have running it locally. And so you can have those kinds of async environments as well.

Jeremy: Awesome. And so what about some customers? So do you have customers that are having success with this now?

Jason: Yeah, we have a number. I mean, we have the European Microbiology Laboratory, which is using it to do science processing and provide access for scientists to the large-scale compute environments of the cloud. We have some airlines that are leveraging this. The airline scenarios, I think, the scenario is actually kind of interesting because it shows the power of combining REST end points, more interactive workloads with batch workloads. In their case, they're exploring using it to do dynamic pricing. So if you think about how you do dynamic pricing, there's kind of two dimensions. It's like, there's a very interactive, somebody is getting a price on a ticket or a route, and you want to be able to present them with dynamic price information as part of that web interaction. But then there's like a data processing angle.

You're looking at all kinds of data coming from your backend systems from route data, from the fleet and historical information. And you're trying to decide what the right price table is for that route. And so you're doing batch processing in the background, and then you're doing this interactive processing. You can implement both halves on serverless with Code Engine and they scale as needed. If you're getting a lot of traffic on the web front end, it scales up as needed without you having to do anything. So they can kind of combine both halves in one environment.

Jeremy: Right. Right. And so in terms of, I think we kind of talked about this a little bit, but when you see all these different services, right, and no matter what it is, whether it's Google's Kubernetes engine that they run or it's EKS on AWS or something like that, I think a lot of people look at these and like, "Oh, it's just another managed Kubernetes cluster." Right? So what are the major differences? I know we talked about it a little bit, but maybe you could just be a little bit more succinct and sort of talk about why is it so different than other sort of previous generations of tools or some of the other competing products out there.

Jason: Yeah. So if you look kind of behind the curtain on Code Engine, you'd see a couple of things. One is there is Kubernetes there, there is a Kubernetes environment there. The differences that Kubernetes environment is completely managed by the Code Engine service. So we're not, if you look at, in IBM Cloud, we have the IBM Cloud Kubernetes service and our Red Hat OpenShift service. So in those services, we're managing a cluster on your behalf, but we give you the cluster. It's like, "Here's your Kube cluster. We'll manage its life cycle, but you have direct access to it." With Code Engine, we have Kube cluster there, we completely manage it in all respects. You have no kind of direct access to it. That allows us to manage scale and capacity. We run that in a multi-tenant way. I mean, we have security and isolation between tenants, but logically you can think of it as like a big Kube cluster that lots of users are sharing, which is how the pay as you go model ultimately works because we're keeping track of what you're actually running and just charging you for that.

So one part of it is fully managing that runtime environment. We've layered on top of that things like Knative so that we have that developer abstraction like a simpler way to define services, to do the source code and image stuff that I talked about. That's coming through largely through things like Knative, which again, we're completely running for you, but it gives you some of that simple interface now that we talked about, and we're doing that in an open-source way with the community. So it's not like proprietary to IBM Cloud. And then on top of that, we built kind of the batch processing system. So batch scheduling and some of these unique interfaces, the command line interface and the user experience to get into that environment for the different workflows that I talked about. And one of the cool things is, because we built it on top of that Kubernetes layer, we can also expose the Kubernetes API if we want.

So like the Ray example I gave you, Ray doesn't really know anything about Code Engine, but Ray knows how to deploy and leverage a Kube cluster. So we're able to actually hand Ray the Kubernetes API server end point inside of Code Engine for your instance. And that framework can use Kubernetes to stand itself up. And then you can use the kind of simple abstractions on top, and that's still all in Code Engine. It's still pay as you go and it still scales to zero. And so that's what I meant by this you can kind of blend the lines and drop down to or the framework can drop down to something like Kubernetes as needed to give you that flexibility.

Jeremy: Yeah, that's awesome. So you mentioned you have a fully managed Kubernetes service and then you also have a bunch of other serverless services that run within the IBM Cloud. So OpenWhisk or, I guess, IBM Cloud functions now. And then also, I mean, you mentioned Cloud Foundry, which is sort of a pass, but it also sort of an easy-to-use serverless environment in a sense. Right? And so I guess, is this like an evolution? Is this where you suggest people go?

Jason: Yeah. Yeah. So I think the simplest way to think about it is yes, Code Engine is the evolution of those ideas. It doesn't necessarily have a direct technical lineage, always, between those projects, but the problem that functions with IBM Cloud functions that Whisk was trying to solve and the problem that Cloud Foundry was trying to solve with source code, start from source code paths, are both represented in what we're doing in Code Engine. So Code Engine will be the kind of natural evolution path for those workloads and for the problems that those users are using those platforms for. The Cloud Foundry one, I think, is super interesting, in the sense that with the rise of Kubernetes has clearly pivoted many people who were doing Cloud Foundry into doing Kubernetes.

Jeremy: Yeah.

Jason: And people are using Kubernetes as their foundation and the Cloud Foundry project, which we're deeply involved in, has done a lot of work to kind of realign Cloud Foundry with Kubernetes in a better way. But what never went away, what people always still saw value in with Cloud Foundry was the simple push my source code developer experience. Right? And so that still carries forward. And with Code Engine, we're taking that same experience that we had in Cloud Foundry, and we're bringing it into this new service and bringing it onto Kubernetes seat, so the developer still gets that similar experience, but without the boundaries that we talked about. The challenge with Cloud Foundry was always like, oh, as soon as you want to do stateful things, or you want to do async jobs, Cloud Foundry didn't solve that problem. Go use a Kube cluster or go use some completely different environment. And so it's kind of the same experience with the boundaries removed and that's where we would see people go.

Jeremy: Right. So if I'm in one of those services, now, if I've got things written in Cloud Functions or in Cloud Foundry, and I've hit some of those limits, or I just want to take advantage of some of the cooler things that Code Engine does, is there a simple migration path for those?

Jason: Yeah. In general, yes. For Cloud Foundry, for sure. It's pretty straightforward to take the same source code directory that you have and just push it to Code Engine instead. Right? So I think the path for a Cloud Foundry, I mean, there's edge cases with everything obviously, but the base of workflow is the same. You can use the same source input directories. We mapped to Paketo Buildpacks, which Cloud Foundry, a lot of that stuff came out of Cloud Foundry. And so that has a really clean path. For Cloud Functions. There's a little bit of a timing thing in general, yeah, you can take your same functions. You can run them on Code Engine. OpenWhisk has some advantages still that we haven't quite gotten built into Code Engine yet. It's got faster startup times, for example, right? The runtime model behind Code Engine, we're still starting a container, like a full container.

In OpenWhisk we had done a bunch of work on warm start of containers and container pooling so we can get like small number of milliseconds startup times on those functions. And some of that hasn't worked its way into Code Engine yet. So there are still some cases with Cloud Functions where it has some capability that doesn't quite exist in Code Engine yet, but over time that will get filled in and there'll be a simple path there to move all those workloads over to Code Engine as well.

Jeremy: Right. So with Code Engine, because you mentioned this idea of sort of like the cold starts. So does Code Engine keep containers warm for a certain amount of time or is it always a cold start?

Jason: It is, in general, a cold start. It can keep some of them, like in the scale up scale down cycle, it may keep them around for a while, so it doesn't be overly aggressive about scaling them down and bringing them right back. But it's not doing some of the warm start tricks yet that OpenWhisk was doing where we have a pool of primed container instances, and then we're injecting code into them and running them. That's work-in-progress. There's work to do both in Knative to improve that stack and then stuff to do in Code Engine. There's a balancing act there too ...

Jeremy: Yeah, definitely.

Jason: ... on things like network isolation and getting on customer VPC networks and other things which are harder to do in that warm start model.

Jeremy: Yeah, definitely. All right. So if somebody wanted to get started with Code Engine, what's the best way for them to do that, just sign up and start writing some code or how do they do that?

Jason: Yeah, kind of. I mean, obviously, we've been talking a lot about how developers use these things. And so I always think the best way to get started is either to build something on it or to try out some specific source code project. We have a lot of things that we've done to try to make that easy. So there's a Code Engine landing page on IBM Cloud. It has some great examples to guide you through those three starting points I talked about, start from source code, start from image and do batch. We have some really nice tutorials, like specific text analysis tutorials, for example, that'll show you how to build applications on Code Engine. And we actually have a pretty cool Git repo, which will take you through tons of samples of how to use Code Engine to solve all kinds of problems.

So there's a lot of really good code assets out there that a developer could go to and actually try something real on Code Engine and the getting started experience is super easy. You've got IBM Cloud, you log in and you go to Code Engine, you create a project, you push an image and then a couple of minutes you'll have something up and running that you can play with.

Jeremy: Amazing. All right. So I love watching the evolution of things and again, just this different way that, that IBM is thinking about serverless and, again, trying to make it easier. Because I always look back and I think of Lambda when it first came out, I was like, "Oh, it's so easy. You just put some code there and it's just done for you." And then we got more and more complex and more and more complex. And not that we didn't need to, I mean, some of this complexity is absolutely necessary, but I'm just curious, seeing the evolution and where things have gone, I talked to a bunch of people earlier about, Roger Graba, for example, who was one of the first people involved with the IBM or the OpenWhisk project, I guess it was Apache OpenWhisk or it became Apache OpenWhisk, whatever what it was, seeing that evolution and seeing the changes that these different cloud providers have gone through, seeing the changes that IBM has gone through and where you sort of are now with Cloud Code Engine.

I'd love to get your perspective here on where you think this is going, not just maybe what the future is for IBM, but what you think the future of serverless is and just cloud computing maybe in general. I know that's a lot of question.

Jason: I'll give you a long answer.

Jeremy: Perfect.

Jason: So that brings to mind two things. First, let me talk about the complexity thing for a second. Managing complexity is always hard. You are so right. That many things start out with a value prop of like, this is easy. And then as people use, the more you add more, and then three years later, we're like, "We need a new thing that's easy because that other thing is too hard now." And there's no magic pill for that. That's always a hard problem to manage. However, one of the things I like about the approach that we're trying to take with Code Engine is because we've layered it on Kubernetes, It gives us a way to kind of decide where we want that complexity to show up. When we had a Cloud Functions OpenWhisk stack and we had a Cloud Foundry stack and you had a Kubernetes stack, you had to try to solve all problems within each stack.

So each stack was getting more complex because you were trying to like, "Oh, I need storage. And I need like private networking. And I need all these things." With Code Engine, I think we have an opportunity to say, once you cross some line, we're just going to ask you to drop down a layer and go use it directly in Kubernetes, right? You can push some of the complexity down and that allows us to hold a harder line on complexity in the developer layer on top. So it's the balancing act we're trying to play is because we built it on a common platform, we don't have to solve all problems in Code Engine directly.

Jeremy: Right.

Jason: So that's kind of my viewpoint on the complexity problem. On the evolution, it's really interesting. So one of the other things that my team's working on and launched recently is this thing called IBM Cloud Satellite, which is about distributing cloud outside of cloud data centers so you can kind of consume cloud services anywhere you want. So cloud computing in general, and this is not just an IBM thing, in the industry cloud computing is diversifying to be kind of omnipresent. You can consume cloud on-prem, at the edge, in our cloud data centers, wherever you want. There's a programming model dimension to that problem, too. As you specially go to the edge, you kind of want some of these simple to consume, easy to deploy, scale to zero, resource-efficient, you need some kind of model like that because at the edge, especially, you don't have 2000 cores worth of compute to go deal with.

You have one box in a retail store, or you have two servers in the back of the distribution center. And so I think things like Code Engine layered on top of distributed cloud and in our case, things like Satellite, is actually a really powerful combination. I think we're going to see serverless become the dominant application development and deployment model, especially for these edge use cases, because it combines ease of deployment and management with efficiency and scale to zero footprint, which are all really attractive when you get outside of a mega data center like you have in cloud.

Jeremy: Right. Right. So I love this idea, too, about sort of expose the complexity when the complexity needs to be exposed. I love this idea of sort of creating same defaults, right? If you could default Kubernetes to do all the optimal things that you would need it to do for use case X, if you could just do that for me and then if I say, "Oh, I want to tweak this one thing," then be able to kind of go down to that level. But I love this idea of you mentioned about edge too because that's one of those things that I think, from a programming model, as you said, how do you write code that's sort of, I guess, environment-aware? How does it know what's running at the edge versus running in a data center versus running maybe in a hybrid cloud and partially in your own private cloud or your own private data center? That model, just wrapping your head around it from a developer standpoint, I think is incredibly complex right there.

Jason: Yeah. It is. And sometimes it's like, how do they know? And then sometimes it's like, how do I just operate at a high enough level of abstraction that like the differences between those environments can get handled below me? If I'm consuming Kubernetes clusters directly, the shape of that Kubernetes cluster in like a retail store or a telco data center in Atlanta somewhere or in the cloud are going to all be different because you have a different amount of capacity. You have a different networking arm. So you're going to have to deal with the differences. If I'm giving you a container image and saying, "Run this," the developer doesn't have to deal with those differences. The provider might have to deal with those differences but the developer doesn't have to deal with those differences. So that's where I think things like serverless and approaches like Code Engine really come to be much more valuable because you're just dealing at this higher level of abstraction and then Satellite and Code Engine and other services can kind of magically deal with the complexity for you.

Jeremy: Yeah. And so I know we talked a lot about Kubernetes and what's running underneath a lot of these services. Is that something you see, though, as being that sort of common format across all these different services, or do you think that something will evolve beyond Kubernetes to become a standard?

Jason: Right now, I really think that Kubernetes will become the base platform. What Kubernetes is will probably keep evolving. And I'm not saying it's Kubernetes forever, but I don't think we should underestimate the power of the kind of industry-wide alignment that exists around containerization and Kubernetes as the next infrastructure platform, if you will, because that's kind of really what it is. And I told you at the beginning, I used to build webs for apps servers. So I was like very involved in the whole Java app server era, the late 90s and early 2000s. And at that time, the industry kind of aligned around two platforms, Java and .net, as the two dominant, at least enterprise, application platforms. We have everyone aligned on Kube. Literally, there's nobody in the industry who's not like, "Kubernetes is the platform." So I think it will be the abstraction for infrastructure in all these environments. The question will be, how do you consume it? Who manages it? How's it delivered? How does it optimize itself? And then at what level do you consume?

And I don't think Code Engine is the end of it at all. I think there's lots of room for improving the consumption experience on top of Kubernetes for these developer use cases.

Jeremy: Yeah. Yeah. And that's actually was going to be my next question, sort of where do you see, what's the next evolution of Code Engine, right? So is that going to be kind of driving into specific use cases more and trying to solve those or becoming more flexible? How do you see the developers, I don't know, in five years, maybe this probably a hard question, but in five years, how are we going to be writing cloud applications?

Jason: Yeah. It's a great and super hard question, but I think projects like Ray, I think, are an interesting forward look into where this might go. One of the things that I've always felt like, if I look at the whole history of paths in particular over the last five, six, seven years, paths has always been about simplifying the experience for the developers, but fundamentally, most paths environments don't change anything about how you write the code. They change how you package the code, how you deploy the code, how the code is executed, and how the dependencies of the code are satisfied. But the actual code you write probably wasn't any different. Right? And that's where I think there's the next step is like, how do we actually get into the languages, into the code structure itself to be able to take advantage of cloud capacity, to be able to take advantage of scale and there's lots of projects that have taken attempts at that.

Ray, as an example, I think is a particularly interesting one, because there's some good examples where you can take a Python function, you literally add like one annotation to it in the language, and now it becomes remotely executable and horizontally scalable for you.

Jeremy: Right.

Jason: It's that kind of stuff that I think three or four years from now, there'll be a lot more of, where we're actually changing how code is written because that code can assume there's some containerized, scalable fabric out there somewhere that it can go execute on top of.

Jeremy: Right. Yeah. And I think that that pendulum swing for developers, especially, well, developers in the cloud, who's they used to be writing a bunch of code, whether it was JavaScript or Python or Java, whatever it was and then all of a sudden now they have to switch context and be like, "All right, now I have to write a YAML file in order to configure my cloud resources," and that sort of back and forth. So yeah, that marrying of basically saying like a programming language for the cloud is a really interesting concept.

Jason: And I think the distributed cloud notion, funnily enough, is a big enabler of that. Because, I don't know, the other tension I see right now is like, let's say you wanted to use Lambda or you want to use serverless functions. That only works in your cloud environment, but you're also running something at the edge or you're running something in your data center, so you're forced to kind of use different approaches, which tends to force you to kind of some common denominator models.

Jeremy: Right. Right.

Jason: And so you're kind of holding back from really adopting some of these newer models because of the diversity. Well, if cloud goes everywhere and those services go everywhere, then now I can just say, "Well, I'll use the serverless model everywhere. And so I can really deeply adopt it." So I think the distributed cloud thing will open up the opportunity to embed these approaches more deeply in kind of day-to-day development activities.

Jeremy: Yeah. No, I love that. I'm all for that approach because I think this split-brain sort of approach to it is getting very complex and it's not super easy. So is there anything else that you'd like to let the listeners know about IBM Cloud Code Engine?

Jason: No. I mean, I think we touched on a lot of the motivation behind it and the kind of core capabilities. I would just encourage you to go check it out, go check out the space, go give it a try and love to hear people's feedback as they do that.

Jeremy: Awesome. Well, first of all, I got to make sure I thank IBM Cloud for sponsoring this episode because just the team over there and everything that all of you are working on is amazing stuff and I appreciate the support. We appreciate the support in the community for what you're doing. So if people want to find out more about you or more about Cloud Code Engine, how do they do that?

Jason: Yeah. And you can find me on Twitter, JRMcGee, or LinkedIn. For me personally, I love to talk to people. For Code Engine, I think the best place to start is the product page, which is ibm.com/cloud/code-engine. And from there, you can get to all of the code examples I talked about.

Jeremy: Awesome. All right. Well, I will put all that stuff in the show notes. Thanks again, Jason.

Jason: Yeah. Great. Thanks, Jeremy.

View Details

About Denis Bauer

Dr. Denis Bauer is an internationally recognized expert in artificial intelligence, who is passionate about improving health by understanding the secrets in our genome using cloud-computing technology. She is CSIRO’s Principal Research Scientist in transformational bioinformatics and adjunct associate professor at Macquarie University. She keynotes international IT, LifeScience, and Medical conferences and is an AWS Data Hero, determined to bridge the gap between academe and industry. To date, she has attracted more than $31M to further health research and digital applications. Her achievements include developing open-source bioinformatics software to detect new disease genes and developing computational tools to track, monitor, and diagnose emerging diseases, such as COVID-19.

  • Twitter: https://twitter.com/allPowerde
  • LinkedIn: https://www.linkedin.com/in/denisbauer/
  • Webpage: https://bioinformatics.csiro.au/

Watch this episode on YouTube: https://youtu.be/5MGxgYd93Jw

This episode sponsored by New Relic.

Transcript:

Jeremy: Hi everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm chatting with Denis Bauer. Hey, Denis, thanks for joining me.

Denis: Thanks for having me. Great to be on your show.

Jeremy: So you are a Group Lead at CSIRO and an Honorary Associate Professor at Macquarie University in Sydney, Australia. So I would love it if you could explain and tell the listeners a little bit about your background and what CSIRO does.

Denis: Yeah. CSIRO is Australia's government research agency and Macquarie University is one of Australia's Ivy League universities. They've been working together on really translating research into products that people can use in their everyday life. Specifically, they worked together in order to invent WiFi, which is now used in 5 billion devices worldwide. CSIRO has also collaborated with other universities, for example, has developed the first treatment for influenza. And on a lighter note has developed a recipe book, the Total Wellbeing Diet book, which is now on the book bestseller list alongside Harry Potter and The Da Vinci Code. From that perspective CSIRO really has this nice balance between product that people need and product that people enjoy.

Jeremy: Right. And what's your background?

Denis: So my background is in bioinformatics, which means that in my undergraduate, I was together with the students that did IT courses, math, stats, as well as medicine and molecular biology and then in the last year of the study all of this was brought together and sort of a specialized way of really focusing on what bioinformatics is. Which is using computers, back in the days it was high-performance compute, in order to analyze massive amounts of life science data. Today, this is of course, cloud computing for me at least.

Jeremy: Right. Well, that's pretty amazing. Today's episode ... I've seen you talk a number of times all remotely, unfortunately. I hope one day that I'll be able to see you speak in-person when we can start traveling again. I've seen you speaking a lot about the scientific research that's being done and the work the CSIRO doing and more specifically, how you're doing it with serverless and how serverless is sort of enabling you to do some of these things in a way that probably was only possible for really large institutions in the past. I want to focus this episode really on this idea of serverless for scientific research. We're going to talk about COVID later, we can talk about a couple of other things, but really it's a much broader thing. I had a conversation with Lynn Langit before, we were talking about Big Data and the role that plays in genomics and some of these other things and how just the cloud accelerates people's ability to do that. Maybe we can start before we get into the serverless part of this. We could just kind of take step back and you could give me a little bit more context on the type of research that you and your organization has been doing.

Denis: Yeah. So my group is the Transformational Bioinformatics Team. So again, it's translating research into something that affects the real world. In our case that usually is medical practice because we want to research human health and improve disease treatment and all this management going forward and for that data is really critical. It's sort of the one thing that separates a hunch from actually something that you can point to and say, "Okay, this is evidence moving forward," and from there you can incrementally improve and you know that you're going in the right direction rather than just exploring the space.

Jeremy: Right. And you mentioned data again. Data is one of those things where, and I know this is something you mentioned in your talks, where the importance of data or the amount of data and what you can do with that is becoming almost as important, if not just as important, as the actual clinicians on the frontline actually treating disease. So can you expand upon that a little bit? What role does data play? And maybe you could give us an example of where data helped make better decisions.

Denis: Yeah. So a very recent example is of course with COVID, where no one knew anything really at the beginning. I mean, coronaviruses were studied, but not to that extent. So the information that we had beginning of a pandemic were very basic. From that perspective, when you know nothing about a disease, the first thing you need to do is collect information. Back then, we did not have that information and actions were needed. So some of the decisions that had to be made back then were based on those hunches and those previous assumptions that were made about other diseases. So for example, in the UK they define their strategy based on how influenza behaved and how it spread and we now know that it's vastly different, how influenza is spreading and how coronavirus is spreading. So therefore in the course of the action more research was done and based on that, they adjusted, probably the whole world adjusted how they managed or interfered with the disease. We now know that whatever we did at the beginning was not as good as what we're doing now, so therefore data is absolutely critical.

Jeremy: Right. And the problem with medical data, I would assume, is one, that it's massive, right? There's just so much of it out there. When we're going to start talking about genomics and gene sequencing and things like that, I can imagine there's a lot of data in every sample there. And so, you've got this massive amount of data that you need to deal with. I do want to get into that a little bit. Maybe we can start getting into this idea of sort of genome editing and things like that and where serverless fits in there.

Denis: Yeah, absolutely. So my group researches two different areas. One is genome analysis where we try to understand disease genes, predict risk, for example, of developing heart disease, diabetes, in the future, but the other element is around doing something, treating actual patients with newer technology, and this is where genome editing or genomic surgery comes in, where the aim is to cure diseases that previously thought to be incurable genetic diseases. The aim of genome engineering is to go into a living cell and make a change in the genome, at a specific location, at a specific time, without any interference of accidentally editing other genes. And this is a massively complicated task on a molecular level, but also on a guidance level, on a computational level, which is where serverless comes in.

Jeremy: Right. Now, this is that CRISPR thing, right?

Denis: Exactly. So CRISPR is the genome engineering or genome editing machinery. It's basically a nano machinery that goes into yourself, find right location in the genome, and makes that edit at that spot.

Jeremy: Right. So then how do you find the spot that you're supposed to edit?

Denis: Mm-hmm. So CRISPR is programmable, so as IT people we can easily relate to that, in that it basically is a string set. It goes through the genome, which is 3 billion letters, and it finds a specific string that you program it with. Therefore, this particular string needs to provide the landing pad for this machinery to actually interact with the DNA because you can't interact at any location.

Jeremy: Right.

Denis: From that perspective, it's like finding the right grain of sand on a beach. It has to be in the right shape, the right size, and the right color, for this machinery to actually be able to interact with the genome, which of course, it's very complicated. But it doesn't stop there because we want it to be only editing a specific gene and not accidentally editing another correct gene. Therefore, this particular landing pad or the string needs to be unique enough in the 3 billion letters of the genome in order to not accidentally veer it away. This particular string needs to be compared to all the other potential binding sites in the genome to make sure that it's unique enough to attract faithfully this machinery. This particular string is actually very short, therefore, when you think of the combinatorics, it's a hugely complicated problem that requires a lot of computational methods in order to get us there.

Jeremy: Yeah, I can imagine. So before CRISPR can go in and even identify that spot, I'm assuming there's more research that goes into understanding even where that spot is, right? Like how you would even find that spot within the sequence genome.

Denis: Yeah, of course. The first thing you need to find out is what kind of gene do you actually want to edit, where's the problem and this the first part of my microbes research, finding the disease genes of really identifying and even within the gene because it has a complicated structure. Even within the gene, you need to find the location that is actually most beneficial for the machinery to interact with and this is where we developed the search engine for the genome. It's a webpage where researchers can type in the gene that they want to edit and the computational then goes in and finds the right spot, right shape, color, and size, binding side, but also makes sure that it's unique enough compared to all the other sites.

Jeremy: Right. And so, this search engine, exactly how does this work explain this. Like what's the architecture of it.

Denis: Yeah. So in order to build the search engine for the genome we wanted to have something that is always online, that researchers can go in at any time of the day and trigger off or kick off this massive compute. In order to do that in the cloud, you would have the option of having massive EC2 instances running 24/7, which of course would have broken the bank. Or, we could have used an autoscaling group where it would eventually scale out to the massive amount of compute in order to serve that task. Researchers tend to not have a lot of patience when it comes to online tools and online analysis, therefore it needed to be something that could be done within seconds. Therefore, an autoscaling group wasn't an option either, so therefore the only thing that we could do was use serverless. This search engine for the genome is built on serverless architecture and back then, we built it like four years ago, that was one of the first real-world architectures that did something more complicated than serve an Alexa scale.

Jeremy: Right. You obviously can't fit 4 billion letters into a single Lambda function, so how do you actually use something like Lambda, which is stateless, to basically load all that data to be able to search it?

Denis: Yeah, exactly. That was the first problem that we actually ran into and back then, we weren't really aware of this problem. Back then the research requirements were even less. It wasn't only the memory issue, but it was also the timing out issue. We figured, "Okay, well, how about rather than processing this one task in one go, we could break it up into smaller chunks, parallelize it."

Jeremy: Right.

Denis: And this is exactly what we've done with a serverless architecture in that, we used SNS topic in order to send the payload of which region in the genome a specific Lambda function should analyze. And then from there the result of that Lambda function was then put into a DynamoDB database sort of in an asynchronous way of collecting all the information and after all of this was done the summary was sent back to the user.

Jeremy: So like a fan-in fan-out pattern?

Denis: That's exactly right.

Jeremy: Right. Cool. So then, where were you storing the genome data, was that in like S3?

Denis: Exactly. This particular one is in S3. We did experiment with other options, like having a database or having Athena work with that, but the problem was that the interaction wasn't quite as seamless as S3. Because in biometrics we do have a lot of tricks around the indexing of large flat files and therefore any other solution that was in there in order to shortcut this wasn't as efficient as this purpose-built indexing approaches. So, therefore, having the files just sit on S3 and query from there was the most efficient way of doing things.

Jeremy: Right. And so, are you just searching through like one sequence or there are like thousands of sequences that you're searching through as part of this? And then how were they stored? We're you storing like 4 billion letters in one flat file or are they all multiple files, how does that work?

Denis: Yeah, so it is for 3 billion letters in one flat file.

Jeremy: Did I say 4 billion, sorry, 3 billion.

Denis: 3 billion letters in one flat file and indexing in order to, not start from the beginning, but jump in straight where that letter is. It depends on the application case as well, like if you're searching one reference genome, which is basically what it's called when you search a specific genome for a specific species, for example, human. For human, it typically is one genome, but if you search bacterial data or viral data, there can be multiple organisms in one file, so it really depends on the application case.

Jeremy: Awesome. Yeah. I'm just curious of how that actually works, because I can see this being a solution for other big data problems as well, like being able to search through a massive amount of texts in parallel and breaking that up, so that's pretty cool. Basically, what you're doing is you're using Lambda here sort of as that parallelized supercomputer, right? And sort of as high-performance compute. From a cost standpoint, you mentioned having this running all the time would be sort of insane to run all the time. How do you see the cost differ? I mean, is this something that is like dramatically different where like anybody can use this or is it something where it's still somewhat cost-prohibitive?

Denis: Anyone can use it for sure. Not for this application, but for another application, we've made a side-by-side comparison of running it the standard way in the cloud with EC2 instances and databases and things like that. The task that we looked at was around $3,000 a month, and this was for hosting human data for rare disease research, whereas using serverless, we can bring that down to $15 a month ...

Jeremy: Wow.

Denis: ... which is like less than a cup of coffee to advance human research. So to me that's absolutely a no-brainer to go into this area.

Jeremy: I would say. What are the tricks might have you been using, or might you have been using to speed up some of this processing? Like in terms of like loading the data and things like that, were there anything that you could use serverless to power that?

Denis: Well, we look at parquet indexing as one of the solutions and that worked for the super-massive files in the human space really well. But again, it comes down to indexing S3 and there was nothing really special around the serverless access. In saying that, one of the big benefits of serverless, again, is being able to paralyze it, which means the data doesn't have to be in one account. It can be spread over multiple accounts and you just point the Lambda functions to the multiple accounts and then collect back the results. And this is something that we've done, for example, for the COVID research where we did the parallelization in a different way, so by now. Genome research there's always the problem of having to deal with large data. Serverless is our first. We were always going to serverless first and therefore, we came against this problem of running out of resources in the Lambda function very frequently.

Jeremy: Right.

Denis: Therefore, we came up with this whole range of different parallelization patterns that serve anything from completely asynchronous with the GT scan is where you reserve the data back in a DynamoDB. Synchronous approach is where you don't necessarily have to collect, sorry, asynchronous approach is where you don't actually have to collect the data back, to completely synchronous approaches where you basically have to monitor everything that you do and make sure that everything is running in CSIRO in order to collect the data back together.

Jeremy: Right. Let's get into the COVID response here because I know there was quite a bit of work that your organization did around that. Before we get into the pattern differences, what exactly was the involvement of CSIRO in the Australian government's response to coronavirus?

Denis: Yeah, so we were fortunate in walking together with CEPI, which is the international consortium sponsored by the Gates Foundation, which way back when was preparing for disease X, pandemic X to come and it was curious that only a year later COVID hit. So all of this pre-work in setting up this hypothetical disease in the future only a year later it actually was needed. So, therefore, CSIRO and CEPI had already put everything in place in order to have rapid response should the pandemic hit, being able to test the vaccine development, so the efficacy in animal models, that was the part that CSIRO was tasked to do. But in order to do that, because with pathogen RNA viruses in this particular case, we know that they mutate, which means it changed slightly the genome and every replication cycle. Also, we've heard about the England strain or the South African strain, being slightly different.

So with every mutation, there is a risk that the vaccine might not be working anymore, might not be effective anymore. Therefore, the first task we needed to find out was, where is this whole global pandemic heading? Is it mutating away in a certain direction? Like, is that direction something that we should put the future disease research on, rather than focusing on the current strains that are available. So, therefore, we've done the first study around this particular question of how the virus is mutating and whether the future of vaccine development is actually deputized by that. Good news was that coronavirus is mutating relatively slowly and therefore the changes that we've observed back then and likely nowadays as well, is probably not going to affect vaccine efficacy dramatically.

Jeremy: Right. You had mentioned in another talk that you gave something about being able to look at those different variants and trying to identify the peaks or whatever that were close to one another, so you could determine how far apart each individual variant was or something like that. Again, I know nothing about this stuff, but I thought that was kind of fascinating where it was like, I don't know if it was, you could look at the different strains and figure out if different markers had something to do with whether or not it was more dangerous or they were easier to spread and things like that, so I found that sort of to be really interesting.

Denis: Yeah. There are different properties, again, with those mutations. We don't know what actually could come out of this because again, coronaviruses are not studied to that extent to really be confident to say a change here would definitely cause this kind of effect.

Jeremy: Right.

Denis: Therefore, coming back to a purely data-driven approach and that's what we've done. So we've converted each virus with its 20,000 letters in its sequence into a KMO profile. So KMO are being little strings, little rods, and be collected how often the specific rod appeared in that letter. So basically, sterilizing it, or [inaudible] coding, if you want to. And with that kind of information, we were running a principal component analysis in order to put it on a 2D map. And then from there, each distance between a dot, which represents a particular virus strain, to the next dot represents the evolutionary distance between those two entities. And from there, we can then overlay the time component to see if it's moving away from its origin. And we do know that this is happening because with every mutation it gets passed on to the next generation of viruses and mutates then and so on, so it does slightly drift away from the first instance that we recorded.

And this is what we've done with machine learning in order to identify and create this 2D map for researchers to really have a sort of an understanding and a way of monitoring how fast it's actually moving and whether that pace is accelerating or not. Currently, they have 500,000 instances of the viruses collected from around the world. So 500,000 times 20 thousand the lengths of the genome, that is 10 billion data points that we need to analyze in order to really monitor where this whole pandemic is going.

Jeremy: Right. And so are you using a similar infrastructure to do that, or is that different?

Denis: We are. Although in this particular case we had to actually give up on serverless in that, the actual compute that we're doing is not done on serverless. We're using EC2 instance, but the EC2 instance is triggered by serverless and the rest of this whole thing is handled and managed by a serverless instance. Eventually, we're planning on making it serverless, but it requires some re-implementation of the traditional approaches which we just didn't have time for it at the moment.

Jeremy: Right. Is that because of the machine learning aspect?

Denis: It's not necessarily the machine learning aspect, it's more of the traditional methods of generating these distances if you want. There's another element to it, which is around creating phylogenetic trees which is, basically, a similar way of recording the genetic distances between two. So you can think of this like the tree of life, where you have humans and the apes, and so on. A phylogenetic tree is basically that, except for only the coronavirus space. And in order to create that we needed to use traditional approaches, which use massive amounts of memory and there was no way of us parallelizing it in one of those clever ways to bring it down into the memory constraints of a Lambda function yet.

Jeremy: But you say yet, so you think that it is possible though that you could definitely build this in a serverless way?

Denis: Yeah, absolutely. I mean, it's just a matter of parallelizing it with one of our clever parallezation methods that we developed now. Another COVID approach, for example, which we implemented from scratch, we're using serverless parallelization in a different way. So here we're using recursion in order to break down these tasks in a more dynamic way, which would basically be required in the tracking approach as well. With this one, the approach is around being able to trace the origin of infection. So imagine someone comes into a pathology lab and it is not quite clear where they got the infection from therefore the social tracing is happening, interviews where they've been, who did they get in contact with, and so on. Also, molecular tracing can happen, where you can look at the specific profile, the mutation profile that that individual has and compare it to all the 500,000 virus strains that are known from around the world and the ones closest to it are probably close to the origin where someone got it from.

And therefore, being able to quickly compare this profile with 10 billion entities that are online that you can compare with was the task and there for doing that serverless was what we developed. It's called the Path Beacon Approach because Beacon is a protocol in the human health space that we adopted and it's completely serverless. What it does is it breaks down the task of recording all those 10 billion elements out there. It breaks it down into dynamic chunks because we don't necessarily know how much mutations are in each element of the genome and therefore sometimes there might be two or three mutations and sometimes there might be thousands of them.

Jeremy: Right.

Denis: Therefore, first paralyzing it in larger chunks and then if necessary, and a Lambda function would be running out of time, we can split off to new Lambda functions that handles some tasks and so on. So if we can process down the recursion in order to spin more and more Lambda functions that all individually deposit their data. So here's another asynchronous approach because we don't have to go back to the recursion tree in order to resolve the whole chain, but each Lambda function itself has the capability of recording, handling, and shutting down the analysis.

Jeremy: Let's say that I'm an independent lab somewhere, I'm a lab in the United States or whatever, and I run the test and then I get that sequence. Is this something I can just put into this service and then that service will run that calculation for me and come back and say, "This strain is most popular or occurs most likely in XYZ?"

Denis: That's exactly right. That's exactly the idea. And this is so valuable because the pathology labs they might have their own data from their local environment, like from the local country, which they don't necessarily are in a position of sharing with the world yet. And therefore being able to merge these two things of the international data with the local data, because serverless allows you to have different data sources in different accounts, I think is going to be crucial going forward. Especially around with a vaccination status and things like that, where we do want to know if the virus managed to escape, should it escape from the vaccine. All of this is really crucial information to keep monitoring the progression going forward.

Jeremy: Right. Now you get some of the data, was it GISAID, or something like that, where you get some data from. And I remember you mentioning something along the lines of, you were trying to look at different characteristics, like maybe different symptoms that people are having, or different things like that, but the reporting was wildly inaccurate or it was very variant. It varied greatly. I think one of the examples you gave was, like the loss of smell, for example, it was described multiple ways in free-text, so that's the kind of thing. So what were you doing with that, what was the purpose of trying to collect that data?

Denis: Yeah. GISAID is the largest database for genomic COVID virus data around the world. They originally came from influenza data collecting and then very quickly moved towards COVID and provided this fantastic resource for the world and the pathology labs of the world to deposit their data. In that effort, in order to make that data, to collect the crucial data, the genomic data for tracing and tracking made that available. They not necessarily implemented the medical data collection part in a way that enables the analysis that we would want to do. Partly because of the technical aspects, but mainly because it requires a lot more ethical and data responsibility and security consideration in order to get access to that kind of data. All they had was a free-text field with every sample to sort of have, if the pathology lab had that information, to quickly annotate how the patient was doing.

This clearly was a crude proxy for what we actually would have needed to have the exact definition of the diseases, ideally, annotated in an interoperable way using technologies and this is basically what we've developed. So we're using FIRE, which is the most accepted terminology approach really around the world, which allows you to catalog certain responses. Instead of saying anosmia, which is the loss of sense smell, it has a specific code attached to it. This code is universal and it's relatively straightforward to just type in the free-text and then the tool that we've developed automatically converts that into the right code and this should be the information that is recorded. Similarly, in the future, what kind of vaccines a person has received and so on. And then from there we can identify, or we can run the analysis of saying, 20,000 letters in the SARS-CoV-2 genome, so the COVID virus genome, any one of those mutations is it associated with how relevant or how infectious a certain strain is or whether it has a different disease progression, or it might be whether it's resistant to a certain vaccine.

All of these is really critical, but because there 20,000 letters these associations can be very [spiries 00:32:49]. In order to get to a statistical significant level, we do need to have a lot of data and currently, this data is just not available. Like we went through and we looked for the annotations where we had good quality data of how the patient was going. I think we ended up with 500 instances out of the 200,000 that was submitted back then that were good enough annotated in order to do this association analysis of saying that mutation is associated with an outcome. And while we found some association specifically in the spike protein that would be affecting how virulent or what kind of disease this particular strain could cause, it definitely was not statistically significant. So we definitely need to repeat that once we have more data and better-annotated data.

Jeremy: Yeah. But that's pretty amazing if you could say someone's loss of smell for example, is associated with particular variants of the disease or that certain ones are more deadly or more contagious or whatever. And then if you were able to track that around the world, you'd be able to make decisions about whether or not you might need a lockdown because there was a very contagious strain or something like that. Or, maybe target where vaccines go in certain areas based off of, I guess, the deadliness of that strain or whatever it was. That's pretty cool stuff.

Denis: Yeah, exactly. So rather than shutting down completely, based on any strain, it could be more targeted in the future and probably will be more targeted in the future.

Jeremy: All right. Now, is this something where everything you've built, all of this information you've learned, that when the next pandemic comes because that's another thing I hear quite a bit. It's like, the next pandemic is probably right around the corner, which is not comforting news, but unfortunately probably true. Is this the kind of thing though where with all this stuff you're putting into place that the next round of data is just going to be so much better and we're going to be so much better prepared?

Denis: Absolutely. That is definitely the aim. I mean, you do have to learn from the past, and having this instance happen firmly puts it from the theoretical space where everyone was talking about before to, "Oh, yes. Is actually happening." There was a paper published in Nature last month, sorry, last year. It was around, how much money have you lost through this particular pandemic, I mean, the lives lost obviously, are invaluable. Looking at the pure economics of it, so how much money have we lost and how much will this damage go on to the future. Therefore, they did a cost-benefit analysis of saying, "How much are we willing to invest in order to prevent anything like this from happening in the future?"

Jeremy: Right.

Denis: The figures that they came up with, and this was way back when we didn't really even know what the complete effect was, and we still don't know. But even back then the figures were astronomical. So I think there's going to be huge shift in order to see the value of being prepared, the value of the data, the value of collecting all this information, the value of making science-based decisions, I think it's going to ...

Jeremy: It will be nice. A change of pace at least here in the United States.

Denis: ... I'll be very optimistic going forward, we're much more prepared than we ever were in the past.

Jeremy: That's awesome. All right. So you are part of this transformational bioinformatics group and so, you have sort of the capabilities to work on some of these Serverless things and build some other products or some other solutions to help you do this research. But I can imagine there are a lot of small labs who, nevermind having the money to pay for, or small research groups that don't necessarily have the money to pay for all this compute power, but also maybe don't have the expertise to build these really cool things that you've built that are obviously, incredibly helpful. What have you done in terms of making sure that the technical side of the work that you've done you've made that accessible to other researchers?

Denis: Yeah, absolutely. My group, the Transformational Bioinformatics Group, is very privileged in that we do have a lot of support from CSIRO in order to build the latest news, tours with the latest news to compute. As you said, other researchers around the world are not as privileged, therefore the tools that we developed we want to make as broadly applicable and as broadly accessible as possible so that other people can build on those achievements that we had. If COVID has taught us anything, it's working together to really move into the right direction together, what is not only rewarding, but it's also necessary in order to keep up with the threats that are all around us. So with that, the digital marketplaces, from my perspective, are the way to do this. Typically, digital marketplaces you think that it's an EC2 instance that is spun up with a Windows machine or something like that, while subscribed to a specific service that is set up for a fixed consumption.

But from my perspective, because it allows you to spin up a specific environment with a specific workflow in there, that you have access to because it's in your account, you can build upon. Therefore, this is the perfect reproducible research and collaborative research approach where someone, like us, can put in the initial offering and other people can build on top of that. This is what we've done with VariantSpark, which is our genome analysis technology, so in order to find associations between disease genes and certain diseases. This is a hugely complicated workflow because you first have to normalize stuff, you have to quality control things, you then have to actually run a variance bug and then visualize the outcomes.

So typically, being able to describe all of that and for other people to set it up in their account from scratch without us helping them, it is complicated. And this is basically the bane of the existence of biomedics research, in that the workflows are so complicated that reproducing them is typically impossible. Whereas now, we can just make a Terraform or CloudFormation or ARM template or whatnot, put it into the marketplace for other people to ascribe to, to spin it up in the way that we intended to, that we optimized to and then from there they have this perfectly reproducible base in order to build upon. Unfortunately, this whole thing ... Variance bug is an Elastic MapReduce offering. The marketplaces are currently only looking at EC2 instances as sort of their basis, the virtual machines as their basis.

Jeremy: Right.

Denis: What we definitely need is a serverless marketplace.

Jeremy: Right. I totally agree with that. So you mentioned something about the data. Have your organization run this in your AWS account for example, and then have other people just send their data to you?

Denis: That certainly would be an option. The problem typically with medical data is that there's a security and a privacy concern around it. Genomic data is the most identifiable data you can think of, you only have one genome, and encodes basically your future disease risks and everything that's ... Basically, it is a blueprint of your body. From that perspective, keeping that data as secure as possible is the aim of the game. Nevermind that it's so large, you totally want it to shift that around, but I think the security element is what really sells me to the idea of bringing the compute to the data, bringing that compute and the structure to the securely protected data source of the researchers or the research organization that have and hold the data and is responsible for the data. It also allows dynamic consent, for example, where people that consent for their data to be used for research, they can revoke that, so it's a dynamic process. Being able to have the data in one place and handle the data in one place directly allows this to be executed faithfully, robustly, and swiftly, which I think is absolutely crucial in order to build the trust so that people can donate or lend their data to genomic research.

Jeremy: Yeah. That is certainly something from a privacy standpoint where you think about ... You're right. Everything about you is encoded in your DNA, right? So like there's a lot of information there. But now, I'm curious if somebody else was running this in their environment after they do the processing on this, and again, I'm just completely ignorant as to what happens on the other end of this thing, but. The data that they get out of the other end of this thing, is that something that can be shared and can be used for collaboration?

Denis: Yeah. Typically, the process is you run the analysis, you get the result out, you publish that, and it sort of ends there. I think in order for genomic data to be truly used in the clinical practice and to inform anything from disease risk to what kind of treatments someone should receive, what kind of adverse drug reactions they are at risk of, it really needs to be a bit more integrated. So, therefore, the results that comes out of it should somehow feed back into the self-learning environment. That's one avenue. The other avenue is that the results that are coming out they really need to be validated and processed. Therefore, typically there are wet labs that investigate that this theoretical analysis is correct in order to move forward.

Jeremy: Interesting. Yeah. I'm just thinking, I know I've seen these companies that supposedly analyze your DNA and they try to come up with like, are you more susceptible to carbohydrates, those sorts of things there. Now while that may be a lofty endeavor for some, I'm thinking more like, people who are allergic to things or environmental exposures that may trigger certain things. Tying all that information together and knowing if that, I mean, I'm assuming that has to be encoded in your DNA somewhere like your, I guess your allergies, I keep using that example. So how does that information gets shared? Is that just something that is like way out of scope because you've got people testing just their own group of samples and doing specific analysis on it, but then not sharing that back to a larger pool where like everybody can sort of look at that?

Denis: Definitely, that is the aim going forward. The Global Alliance for Genomics and Health is putting things in place in order to enable this data sharing on a global scale. The serverless beacon that we've developed is moving along the line as well to make it more efficient for individual research labs to light their own beacon in order to share their results with the rest of the world, like the $15 per month in order to share data with the world. I don't think we're quite there yet in terms of the trust, in terms of the processes to make this actually a reality within the next, I don't know, five years.

Ultimately, it definitely is the easy aim, and ultimately this is the need. An element to that is also that, the human genome is incredibly complex and therefore there is no real one-to-one relationship between mutation and an outcome. We do know that, for example, for cystic fibrosis, it's one mutation that causes this deadly devastating disease, but typically it's a whole range of different exacerbation factors, resilience factors, that work together and it's very personal with the kind of risk that it generates. In order to quantify this risk, we need to have massive amounts of data, massive amounts of examples of which kind of combination is causing what kind of outcome.

Jeremy: Right.

Denis: In order to do that probably putting all the data in the same place it's not going to happen ever. Therefore, sharing the models that were created on individual sub-parts and refining the models on a global level, like sharing machine learning compute models, I think is probably going to be the future. And this is a really interesting and exciting space and a new space as well where it's sort of a combination of secret sharing and distributed machine learning in order to build models that truly capture the complexity of the human genome.

Jeremy: Yeah. Well, it's certainly amazing and fascinating stuff and I am glad we have people like you that are working on this stuff because it is really exciting in terms of where we're going just to mean, not only just tracking and tracing diseases and creating vaccines but getting to the point where we can start curing other diseases that are plaguing us as well. I think that's just amazing. I think it's really cool that serverless is playing a part.

Denis: Absolutely. So my goal is really to bring the world together and see the value of scientific research and bring that scientific research into industry practices.

Jeremy: Awesome. All right. Well, Denis, thank you so much for sharing all this knowledge with me. I don't think I understood half of what you said, but again, like I said, I'm glad we have people like you working on this stuff. If people want to reach out to you or find out more about CSIRO and some of the other research and things that you're doing, or they want to use some of your tools, how do they do that?

Denis: Yeah. The easiest is to go to our web page, which is bioinformatics.csiro.au, or find me on LinkedIn, which is allPowerde, and start the conversation from there.

Jeremy: All right. That sounds great. I will get all that information in the show notes. Thanks again, Denis.

Denis: Fantastic to be here.

View Details

About Aaron Turner

Aaron Turner is a senior engineer at Fastly. They were previously doing rad stuff at Google and various startups and agencies. In their spare time, they are hacking on various WebAssembly projects on the web, cooking up some dope beats, and shredding local skateparks.

  • Twitter: @torch2424
  • Website: https://aaronthedev.com/

Links related to the content in the episode:

  • Getting Started
    • WasmByExample
    • MadeWithWebAssembly
    • Wasi.dev
  • Choosing a Wasm Language
    • AssemblyScript
    • Website
    • Mentioned Production AssemblyScript article: Micrio article by Marcel Duin
    • Emscripten
    • Compiling to Wasm
    • Rust
    • Wasm Book
  • Great WebAssembly Talks
    • WasmSummit
    • Youtube Channel
    • WasmSF
    • Youtube Channel
    • Patrick Hamann | WebAssembly – To the browser and beyond! | performance.now() 2019
    • WebAssembly
    • for Javascript Developers, by Aaron Turner
    • Simulating
    • Sand: Building Interactivity With WebAssembly by Max Bittker | JSConf EU 2019
    • Robert
    • Aboukhalil :: Level up your web tools with WebAssembly :: #PerfMatters Conference 2019
  • Keeping up with WebAssembly
    • Fastly Blog
    • Bytecode
    • Alliance Blog
    • WasmWeekly
  • Twitter and Newsletter
  • WebAssembly In the Future
    • Lin Clark’s Wasm Summit Keynote
    • WebAssembly
    • Specification Proposals
    • WASI
    • Specification Github

Watch this episode on YouTube: https://youtu.be/Ef1iE9KaAd8

This episode sponsored by New Relic and Epsagon.

Transcript
Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm joined by Aaron Turner, hey Aaron, thanks for joining me.

Aaron: Yeah. Thanks for having me. I'm really excited to be here.

Jeremy: Awesome. So you are a Senior Software Engineer at Fastly. So I'd love it if you could tell the listeners a little bit about yourself and your background and what you do at Fastly?

Aaron: Yeah, totally. So what is it about me? So what I do is I work a lot of WebAssembly. We've been doing that for about two and a half years. I started getting really involved in the community and through that work, I was going through a lot of meetups and things, and I ran into Tyler who was the CTO Fastly and they were working on this new edge WebAssembly thing. And just the timing lined up and our interests and both were passionate about the moment it did. So I joined the company and it's been going great so far. And there, what I'm working on is a lot of WebAssembly work, both in terms of bringing on new languages to the platform, but also it's a lot of community work, participating in a lot of events, still doing podcasts and things and just hanging out with people and having a good time.

Jeremy: Awesome. So actually I had Tyler on the show not that long ago. And we talk mostly about computed edge, but we got into WebAssembly a little bit and I'm finding this whole thing fascinating. Because I remember way back in the day Java applets and of course Flash if anybody remembers that, I think that's still around. But this idea of trying to bring compute and more complex applications in a bytecode form, bring those to the browser. And this a really interesting thing, I don't think it worked out really well in the end, but it seems like WebAssembly is a better shot at doing that. And I find that really, really interesting. So I'd love to just pick your brain for a little while here and talk about WebAssembly, but I think maybe for the benefit of the audience why don't we start with what exactly WebAssembly is?

Aaron: Yeah, totally. So WebAssembly I like to describe it or really what it is is bytecode for the web like you were alluding to. And what that means it has a few implications. Two of my personal favorites are predictable performance. So when you look at something like JavaScript it's an interpreted language, but we got really good at running it really fast. So we built a just-in-time compilers. So what that's going to do is go ahead and read your JavaScript and compile it over and over many times. It makes some assumptions about what your code is doing, which can get really, really fast if it assumes the right things. But if it makes the wrong assumptions they can get really slow. Whereas if you're running a bytecode it's always predictably performing. So a really bad analogy people tell me not to say, but I like to think of it as like, if you're driving on the freeway, it's like your WebAssembly. Taking their freeway depending on what you're trying to do most of the time it's going to be faster.

There might be sometimes you're taking the streets driving around, doing the neighborhood might be a little faster, but nine times out of 10, if it's far away in a few miles, we'll just use the freeway. You know what I'm saying? That sort of thing. So that's what I like to describe the predictable performance. At least that's how it works in my head. And when kids ask me or my little brother is like, what are you doing? I'm like, that's it. And then another thing about is that it's very portable. So as the nature of the web it's great for distributing logic wherever it may be. So both portable in terms of, it runs in all major browsers which is a huge solid, because you start shipping in. It's very bondable so if you wanted to throw WebAssembly into an MPM package, for example, and use it there that can run in the background and do some of that predictable performance for you in your JavaScript ecosystem.

It's also very language-agnostic. So if I guess parent language compiles down to WebAssembly the actual bytecode then you can use any language essentially and have an interface with this JavaScript WebAssembly API. And I'll get into it's very portable and the fact that multiple runtimes support it. So for example, Nodewell it has its own adapters there, but also people have built their own runtimes for WebAssembly itself like WASMtime which has extended long run time and as well what Fastly built, which is called Lucet. And it's an ahead of time compiled WebAssembly runtime type thing, but yeah, that's probably how I would describe it best.

Jeremy: All right. So there's a lot to unpack there. And so what we can do is we'll go through some of those things. But just in terms of like, I think this is where we look at things that run in the browser. And I know we'll get into this, that you can run a WebAssembly in more places than just the browser. But running in the browser one of the things that I think we think a lot about is security, what does it have access to? Can it do network calls? Can it access local resources, things like that? What are the capabilities of WebAssembly?

Aaron: Yeah. Thank you very much for asking that, because it's always like, I'm always performance, portability. Awesome. Then it's like, oh yeah, there's also these other great things. So yeah, total insecurity. The one thing that's really nice about WebAssembly is that it has this concept of linear memory. So the idea really is that you are given a heap or just the way I like to think of it from a JavaScript background, self-taught here is, just this one really big array and you can't go out of the array and you can't go before the array. And that's all you can access and it's sandbox. So because of that, you can't escalate out of that memory and do things with the host. So WebAssembly has a really nice feature in which this linear memory of sandbox. So you can't really escalate out of it just because WebAssembly don't allow that. And on the topic of capabilities one thing that's nice about WebAssembly as well as this concept of host calls.

So essentially you can say, hey but we talked about the host here. It can be JavaScript like you mentioned, or one of those standalone runtimes. You're like, hey, host I know you have access to this function let's say, if I call this, I want you to go do this on my behalf. So for example, a common one in WebAssembly that people often use is like, let's say you want to use console.log, for example, you can go ahead and import and say, hey, look, I want to import console.log. And when I call it, I'm going to pass it this value and then JavaScript can then say, oh, hey, you called this function that I gave you access to, this host call that I provided. Cool, I'll go do some work for you.

In this case, log out a number let's say, and then let you continue executing. So this opens up all types of cool use cases of depending on what the host wants to provide, you can start doing some really cool things and have this security where the user can give you some code that you don't really know what it's doing, but only has as much power as you will let it have a doing those calls really. And that, since that memory is sandboxed they can't work out of there either. So starts to build this very secure. You can start trusting code that people are giving us because of these two features, which is really exciting.

Jeremy: Yes. So you mentioned use cases and I think use cases are probably the best way to communicate to people what the capabilities are, right? So great performance, linear memory, sandbox, security, that all sounds awesome. But if you can't explain what you can do with it, right, it's hard sometimes to visualize. So you mentioned this idea of calling JavaScript or being able to do those host calls and stuff, but what are some of the practical use cases that you would use it for? And maybe even more importantly, what's the use case and then why is WebAssembly better for it?

Aaron: Yeah, totally. So WebAssembly started off as something that's for browser, it's starting to evolve more into serverless use cases and things, but just taking a step backward started the major use case here was speeding up JavaScript in the browser. So JavaScript was in this interesting place where it had this unwilling monopoly, I will say, on the browser, where you had to use JavaScript. So because of that lots of different companies at different interests in it and started pulling in different ways that it really wasn't ever designed for it. So WebAssembly was like hey, look, that whole performance thing we're trying to a JavaScript yeah, it's good for that. But let's take a little bit of a weight off of JavaScript and give something else. So I bring all this up to say that speeding up JavaScript really. So if you have let's say a loop that maybe it's doing image blurring, for example, that isn't supported by the browser natively. I don't know if there was a CSS thing proposed, but let's pretend there isn't but it'll be, go ahead.

Jeremy: Even if there was, it's not going to work the same way in all browsers I'm sure, so ...

Aaron: Yeah. So that's really computationally intensive. You're going to be looping around trying to figure out what pixels need to be duplicated, which ones don't, and WebAssembly is really good at those tasks because of that whole predictable performance thing. So what you would do is take that block of JavaScript and instead replace it with a WebAssembly module that you then pass and say, hey, look, here's the pixels that I want you to go ahead transform, turn back in your linear memory, what the new images, and I'll go ahead and display that. And having access to that predictable performance then lets you do those things on the browser a lot easier, and they're not so taxing. You can also imagine this same computationally intensive speeding up JavaScript is great for game engines. I've seen a few game engines already take some of their physics engines and start replacing pieces with WebAssembly modules, because it's just built and designed for doing these computationally intensive math operations and things that game engines really need.

And then another one is probably just general business logic. So there's all types of different times where we're just like, hey, we got this data structure and we just going to make it look like this now for the server. So sometimes JavaScript's good at that or will be. It definitely depends on the use case. So that's one thing I'll note here is that what does sound most of the time works is I'm sure there's that one use case it's like, okay, fine. But you know what I mean? But nine times out of 10, let's say we have to take some JSON. It's huge and we need to maybe convert these objects into an array let's say, I don't know. I'm just making things, you get the point ...

Jeremy: I get the point.

Aaron: That type of business logic where you're trying to translate things and just stuff that's really tedious on the CPU. You can start to put that on WebAssembly modules and from the portability too, you can start sharing it on different platforms. And it's like, oh, this is really exciting. Yeah.

Jeremy: So I'm curious about this too. You said if JavaScript isn't fast enough and a lot of companies have to use JavaScript sometimes make it do something, sometimes it shouldn't do maybe. So if I'm a developer now, I love Node on the backend, I love to know because I'm just so familiar with JavaScript that it's just really easy to go back and forth between the front end and the backend using this one single language. And pretty much anytime you have to write something for the front end, it has always been JavaScript. You had to basically do that. So is this something though where ifw... We will get into the other languages you can use ... could I compile an entire application or maybe a front end app or what do they call it a single page app or something like that, that I might write in JavaScript now, could I do something like that compile it down to WebAssembly and then do a similar thing?

Aaron: Yeah. So that's actually a really interesting question. I'm glad you asked that. So the short answer is no, but we'll get into why. And it becomes really down to two different reasons. The first being WebAssembly and its nature is binary format depends on strict typing. So you have to say, hey, I have an i32, I have a i64 integer, i64, so on and so forth. JavaScript is dynamically typed and that's part of why we need to interpret and do that just in time thing to figure out those types on the fly. So JavaScript just isn't quite designed to compile down to WebAssembly. And then two I'll say, is that there are some alternatives. Let's say you wanted to write a Rust step that is a Full SPA. There's a few products out there that do that, I've seen them on GitHub. I don't know any off the top of my head, but they totally exist. And that's something you're into then totally feel free to do it.

But I will make the point that JavaScript as we mentioned, has been being pulled in a lot of different ways. But one thing that was sure was designed for was interacting with DOM and building UIs. So that's where I think JavaScript really excels where I would say you really want to just compile a straight SPA into WebAssembly because WebAssembly really get those computationally intensive things. But SPAs, for example, when you see those react demos of like, look, we got 10,000 triangles that are rendering every second. JavaScript is amazing at that because we've been iterating on making JavaScript good at that for so long. So yeah, I hope does that answer it. Does that make sense?

Jeremy: It does. No. No, it does. And actually, to extend that though, there are things that you might want to do in a browser that are fairly I would say insecure especially if you have to do anything with crypto or you have to sign a call or some of these other things call an API that you maybe you need to have a secret in that API call. And obviously, you don't want that available via your JavaScript by just viewing source. So are those some of the things that you could potentially do with WebAssembly where you could compile down something that did cryptographical signing or something like that. And then with the bytecode could you reverse engineer that bytecode? Could you come back and find the actual source on that? Or is that something that might be secure where you could use something like that to do some of those more complex and maybe things that would add a little bit of security to your app?

Aaron: Yeah. So that's a really interesting question. I'm glad you asked it. So I will first iterate and say again, we're talking about that cryptography is going to use a lot of math operations very performance-intensive, which makes it a great fit for WebAssembly. That being said on the security side of things, it gets really interesting because there is a text format for WebAssembly. It's not the same as JavaScript where you can go and view source and it's just like, oh yeah, this is totally what's running on the page. Nowadays with mangling and the translation, it gets more complicated, but WebAssembly is more of a binary format where we can do some funny things. You'd be like, hey, look, it's moving this memory here and there, but to see the actual source code is a little bit more difficult. That being said, I won't pretend I'm a security expert especially in cryptography space ...

Jeremy: I'm not either. So yeah.

Aaron: I've definitely seen some projects that are like, hey, look, we can make this things secure with WebAssembly. Here's our white paper on why, but I wouldn't be able to confidently sit here and be like, oh yeah, totally just throw secrets and WebAssembly, no one will ever know what it is. You know what I mean? So I would definitely suggest maybe I always get back to you about that or we can do ...

Jeremy: Yeah. No, I'm just curious. I'm thinking through the places where it really fits in. Because I think that's a problem, security in the browser, it was with so many people using APIs now usually, and again, a good use case for serverless is essentially setting up a function that all it does is just adds that secure key or whatever it is that access token into a third-party API calls so that you can just pass it through from your from browser. So if you could eliminate that step, you know what I mean? And be able to do some of that I think that would be interesting.

Aaron: Yeah, totally. Yeah. And I definitely agree. For example, whenever we're importing a bunch of packages, and let's say they're both accessing global scope just cross their fingers and hope they don't do something they're not supposed to. That is one benefit again about WebAssembly is that all that sandboxing, that linear memory and dances that only really has access to what you give it to. So if I were to have three WebAssembly modules, they couldn't go and talk to each without JavaScript being in between hey, you're telling this person cool, here you go. Let make sure you're not doing anything funny between one another in the current state of WebAssembly today, so.

Jeremy: Awesome. All right. So let's move on. Let's talk about WASI. What is WASI?

Aaron: Oh yeah. So WASI is an acronym for the WebAssembly System Interface. And this is where in my head, WASI is the node of WebAssembly. I know it's a very loaded term, so please take it with a grain of salt, but essentially it's a standardized system interface for WebAssembly. So you get things a lot of positive Slack calls, if you're familiar with a lot of ... it's getting low level, but you can imagine stuff like fd write, fd read. So reading file descriptors and things, you get access to those things. So you can imagine a Node, you have file system and that's how you would let's say, make generate files on a server. He used the module fs. WASI offers that lower-level primitive that allows you to do those things in WebAssembly like create files, read them and move them around your file system, which I'm talking in circles, but I think I get the point. So what actually ...

Jeremy: So you wouldn't use WASI in the browser, right?

Aaron: So it gets funny there because there have been some ... you can use things IndexedDB if your familiar to create a pseudo file system and start to port, maybe let's say you wanted to compile something in the browser. If you want to bring a C compile into the browser, you could use the WASI things and mock out some of these system-level resources as a browser equivalent and get into this funny world where it does make sense. But WASI itself one of its goals right now isn't really to run in the browser if people are bringing it there. So if you want to, again, bring a compiler but you totally could, which is exciting and really cool. And I think that there's a talk by Ben Smith they did this exactly for their class. They taught I think I forgot what university, but yeah, they made like a C compiled WebAssembly compiler that link took C source compiled it. And then there's something else is I wanted to say about that?

Jeremy: Well, I'm just curious, so then if it's not really for the browser, I get it. Everybody loves to do those things like, oh, can I take this thing that wasn't built for this and make it work in that. But so what would you say are the primary use cases then for using WASI?

Aaron: So a lot of it is probably bringing WebAssembly to the server. Yeah, probably for the server or just even standard command line applications. WASI is really exciting. There's a lot you can do. And you'll hear me say that a lot about WebAssembly, there's a lot you can do with it. So really I like to think of WASI as like all the benefits of WebAssembly, so that sandboxing that you get and that host calling interface. Because really what WASI is using is it host call just like, hey, you have access to the file system please, do what you want with it. So because you get those benefits. There's an infamous tweet. I know I'm talking, but ...

Jeremy: I think an infamous.

Aaron: Infamous, that's not the right word, but a tweet that went viral by Solomon Hykes. If you're familiar, the co-founder of Docker about how WASM plus WASI existed. And I think in 2008 is when they made Docker that they maybe had not needed to make Docker. And the reasoning there is that like, if you could take different applications and compile them down and give them access using a system interface that's standardized and have access to things like file systems, you can start to imagine this world where you get container-like functionality where you're just like, hey, I'm gonna compile my whole app and all of its dependencies down to WebAssembly, give it access to WASI and they can start to do things, and operate in a sandbox way where you don't have to worry about it messing with the parent operating system or completing with other apps and things like that.

And then of course we're on Serverless Chats. So a lot of serverless use case there because WebAssembly is very I guess lightweight and things of that sort. Sandbox, it makes it a great contender for serverless because you can just instantiate the WebAssembly module start running immediately and then close it down. And even do things like snapshotting of the memory and saving state and things like that. But we get to in the future there's still some kinks to figure out there just in general in the whole ecosystem. But yeah, and then a lot of standalone applications. So we were mentioning if you want to compile a C compiler to WebAssembly, give it access to WASI and run it on your local machine. Now you can start compiling things on your actual let's say standalone runtimes. If you just wonder if for some reason it gives C source code to just a WebAssembly module and have it compiled.

Yeah, sure. Cool. And I'm sure there will be use cases for it because then you could imagine a world where it's like, I won't use the exact same file that I used my exact same compiler binary, because I use Mac windows and Linux all do recompile for each architecture at each operating system. So again, that portability and yeah, it was probably, those are really good use cases I can think of off the top of my head right now.

Jeremy: Right. And this might be a stupid question, but I guess the idea of it being the Node of WebAssembly in that sense where basically that's what it is. It's its own container essentially that can do anything, it can interact with the file system, it's got all of the capabilities, HTTP networking calls, it can do all that stuff. And I know the V8 engine is pretty popular with edge computing and things like that, is that something where that can run on like a V8 engine or is there another type of underlying, I guess, container management system or something that would need to run those?

Aaron: Yeah. So I'm glad you brought this up. So one thing I would just say as a quick note so HTTP is still being standardized. So all this is really young. So I wouldn't want somebody to like, oh, HTTP is in there? Cool. All right, let me just close the tab and start crying. You know what I mean? So there's a lot of standardized API still in the works. And we can get into them that too if you'd like, but to answer your question, I'm like, hey, could you use WASI inside of VA? Again, we can probably polyfill some things and then you can totally it in there and that JavaScript way. But what gets interesting about that is that you can run it if you can imagine a world where maybe you don't really need the JavaScript as well that V8 provides, then you're instantiating a JavaScript runtime and a WebAssembly runtime. Whereas if you use some of these other runtimes that only support WebAssembly you get a lighter weight output from it and things like that.

So yes, you can use V8, but it's depending on your use case if you want to have that solid WebAssembly JavaScript relationship at all times, V8 is probably the right answer, but if you just want to use WebAssembly, then the standalone runtime is going to be lighter weight for, and be able to optimize specifically for WebAssembly a little bit better.

Jeremy: Right. So there are standardized runtimes for WebAssembly?

Aaron: Yes.

Jeremy: Okay. That makes sense. All right. So you mentioned some standardized APIs and we talked a little bit about crypto and how that's a good one to deal with. But you've got crypto, you've got machine learning, all kinds of these complex things that are really hard to do especially in JavaScript. And again, I know we're now more towards the server side of things, but what are those types of ... You mentioned they were standardized APIs, what can those do? What are those capabilities?

Aaron: Yeah, totally. So I will say these are also still in flux, they're still being developed. And if you want to participate there's totally community group for WASI that you could totally hop in and join, but this is me, I attend the meetings, having for a little bit. So it's like sharing this little things is going on here and there. So yeah, one of them is definitely crypto, a colleague of mine, Frank Dennis, is working on that where there's a lot of common crypto applications that we would want the host to say, hey, the host runtime we know that you're running just raw bytecode, there is no layer between you and the kernel essentially. So could you please provide SHA256 and make sure that it's working correctly, you won't expose the right thing.

So probably like common crypto functions that people are often using. I don't know if SHA256 is one of them I heard ... you can imagine. If you're a crypto person out there you know, exposing those common ones that folks use. And then machine learning, I'm a little bit more familiar with that because the Bytecode Alliance is a group of different companies that are working together on WASI and WebAssembly and all these specs and they recently announced they're working on a WASI. So a neural network and WASI and providing the primitives there of, if you want to build neural networks on top of WASI what are some of the host calls that we'll need? What are some of the functionalities that we'll need? So we start to build our own neural networks in WebAssembly, which is really exciting.

Jeremy: So are these standardized APIs? Again, this may be a stupid question just because I don't know enough about this stuff yet, but are those NPM packages or Python packages from PyPI or something. Whatever it, is that the idea behind some of these APIs where you say, I need a crypto package, or I need a machine learning package, or I need an image manipulation package, is something where there will be an ecosystem where people can write and contribute packages that other people could just pull in?

Aaron: Yeah. So this is more a little level above that I would say, you can imagine 10 portal, for example, as a new date thing was coming into JavaScript. It was already a few libraries out there, hey we're playing around with this new date API. Here is my version of it and your version of it, but a lot of the community is working together on deciding, okay, well, my organization has these needs for temporal, my organizations as these, what are the compromises that we'll need to make from the entire community to have access to this one standardized API? Eventually for example, Chrome or your whoever's implementing your JavaScript runtime will support natively you wouldn't have to use a library anymore. So I would say it's a little level above, and I'm sure as these APIs develop people will develop like, hey, look, here's the current version of WASI and Rust let's say. And you can include it in your Rust program it then compiles all the way down to WebAssembly. And so, yeah, I hope that ...

Jeremy: Yeah. No, I'm just wondering here because one of the things that I think makes Node and JavaScript, the reason why those ecosystems grew so much was because people contribute to those things. So if there's a way for people to do that and know oh, if this doesn't do this now I don't have to write this whole thing myself. Somebody else may have done this for me and I can just go ahead and bring that in. Yeah. So I don't know if they'll eventually be like an NPM for WASI or whatever or for, I guess WebAssembly in general. But yeah, anyways, I'm just thinking that would be a cool thing to have. So if there's not one of those things out there, then my suggestion is you create one of those things and make it so that people could contribute code.

Aaron: Yeah. So if I could on that note like I had mentioned a little bit earlier before I grazed over, is that in the browser at least, I think in WASI too you could ship just an NPM package of what saying today. And if they used WASI like we mentioned, so it's a little part in there depending on what you want to do. So just a registry in today's world is probably MPM and I have seen some smaller WebAssembly package manager type things that are popping up here and there but right now I think NPM is the ... for a lack of better word the king of the packages right now. We'll see if something else comes up may do something that's more WebAssembly focused rather JavaScript with the WebAssembly if that make sense.

Jeremy: Yeah. Cool. So let's into the toolchains here and maybe we just focus on the popular ones because I'm sure they're a lot of people that are going down this road but you mentioned earlier that WebAssembly was language-agnostic, so you can just write it in any language you want or? There's got to be some limitations here.

Aaron: Yeah. So there are some. Pretty much the limitation is that, is your language willing to support WebAssembly? So especially if it's a strictly type language. So there is, for example, Go is working on something. Let me think of another, Swift has a WebAssembly implementation, Zig is an up-and-coming language that has WebAssembly implementation. There are some languages where it gets funny. So for example, these dynamically type like JavaScript and Python, they're in a world where it's well, just the language doesn't quite line up with what WebAssembly needs when it comes to being compiled. So I mentioned those, but probably the biggest three right now is Emscripton, which is a toolchain for compiling CNC plus to WebAssembly. A lot of folks use it. Google I know it has a lot of folks working on Emscripton and they highlight a lot of projects using it. If I'm not mistaken there's a talk from WebAssembly SF in which the Google Earth team talked about how they're using Emscripton to like, again, take all that business logic and make it more portable across where they need to run Google Earth which is really exciting.

Another big one is Rust. I had mentioned earlier, it's a language that's grown to become quite popular. It's got another systems-level programming language, but Rust has a little flavor of both taking these older C applications I guess not really porting, but building these one-off modules to go off and do maybe me to put JavaScript, serverless type stuff, whatever it may be, Rust is really starting to shine in that area. And a lot of folks really passionate about it. The community is really cool and there's a lot of great documentation. I guess what I'm trying to say really is that Rust has a really solid WebAssembly support. They are going all-in on it, which is really exciting.

Jeremy: Well, just thinking back to things that I've heard, I feel like whenever I hear Rust I just think WebAssembly, is that the wrong way to think of it?

Aaron: Rust does a lot of different things, but it's not maybe the wrong thing just because they spend a lot of time really building out a lot of the tooling, a lot of the community around it and things. So Rust definitely also has lots of great use cases that I've seen at Rust conferences where they were in Rust on no server, they do it in games, so on and so forth. It's a systems level approach, so programming language. So you can imagine you can run see there. Nine times out of 10 I think you could run Rust there, but just there WebAssembly ... What's the right word? Involvement, there's another word starts with an "I" that means what I'm trying to say, but they spend a lot of time working on creating a great developer experience on WebAssembly and Rust. So yeah, because right now I will say like some languages like Go and things are still very young in their WebAssembly implantation. So some of the tooling is like, good luck, but I'm sure that will get better over time as more of the community works on it and things.

And then if I could transition too, there's one more tool chain that's really popular right now and it's called AssemblyScript, I'm a member of the team on that. And what a AssemblyScript is it a very TypeScript-like, not TypeScript exactly, but if you can read TypeScript, the Typescript-like language that compiles to WebAssembly. So the target. There's a lot of these JavaScript developers that we mentioned earlier, JavaScript can't compile WebAssembly. So it's like, come on, we want some of the fun too. So we're hoping to maybe fill that gap and getting developers the closest thing to JavaScript that we can provide it for them that allows them to access all the benefits of WebAssembly. So yeah.

Jeremy: Yeah. Well, so the AssemblyScript thing because I was doing some research before this and when I saw that, I was like, okay, now you might have me because I'm thinking some of these other things. I spend a lot of time in Node and in TypeScript and JavaScript. So for me, I was like, oh, this would be really great. So I do have a bunch of questions on this though and since you work on the AssemblyScript team, you're the perfect person to answer these. So what are the differences between TypeScript and AssemblyScript? Is it almost exactly the same or are there some significant things that I'd have to worry about?

Aaron: Yeah. So one thing is I'll definitely point out is that we mentioned earlier about SPAs, but just imagine not even ... I'm trying to think. So if you take a TypeScript React app or TypeScript Node app, you can't just grab the AssemblyScript compiler and be like, hey, I get WebAssembly for free. Cool. That's not how ... There's a lot of small fundamental differences where you have to actually take the time to port things. But the porting comes down to like, okay, you know this number type number isn't again, that strictly type integer of whatever many bits. So you have to take your numbers and convert them to i32s, you'd use a float there, you need to specifically say like, I want to float here. And things of that sort. But for the most part, I actually wrote an article in the Facet Blog about porting TypeScript to AssemblyScript and what that looks like.

Yeah, I think it might even have been a JavaScript application, but yeah, just like, hey, look, there's this variable here. It's a number we know that but JavaScript does the work for us. Let's just explicitly say this is a number. And yeah, pretty much that's like a lot of the big fundamental differences. There are some "gotchas" to some of the script. One of them being is that since it's young, it's only about a two-year-old language, it's been getting popular though thankfully, is closures is a big one. So if you're doing a lot of callbacks, that callbacks, I call it callbacks. As of right now, we're still working on getting that working in WebAssembly memory and things. So you have to pull those functions out to separate functions which is a little annoying, but we'll get there. You know what I mean? And in most use cases you have that. And then ...

Jeremy: So in terms of the workflow for some of these things, like I'm writing TypeScript now you got to compile it, right? And then I usually run tests against my TypeScript and then compile or compile and run the tests. So what does that workflow do I have to go from ... Or I guess I just writing in AssemblyScript which has to be very similar to TypeScript and then just compile it down to Rust, but how does the testing work and some of that other stuff?

Aaron: Yeah. So that's actually probably the best closest thing. So AssemblyScript is an NPM package that you MPM install into a project and you can scaffold out in AssemblyScript project essentially. Another thing is AssemblyScript it pretty much uses the .TS file extension, so if you'll put the .TS code it's like, oh, hey, this is a TypeScript. And it comes with a TSConfig. So the TSConfig will say like, hey, i32 was an alias for number for your Linters and things. So you open up VS Code and it's like, oh, this is just TypeScript. Cool. Awesome. So you can just hit the ground running in that aspect. If you're used to writing TypeScript, the amount of workflow difference should be near nothing to my experiences. When you start doing some really specific TypeScript things or if you have small things here and there, it gets funny.

So good examples is that even for documentation, for all the AssemblyScript packages I've been writing, I just use TypeDoc. TypeDoc can look through AssemblyScript and be like, cool. Yeah. This is, this goes there, this goes there and there's almost no actual small things I need to do there. And in terms of testing AssemblyScript has its own testing libraries, but I can promise you it's extremely solid. It's called Aspect. It's a just library written by Joshua Tenor. And if you've written Jest, it looks like JavaScript testing. It's like, describe it does that, yeah, expect to be, it offers all of that in AssemblyScript so.

Jeremy: Right. So you write something in AssemblyScript, I know we were talking, so that is compiling it down, is that using WASI? And then what can I do? Can I pull in Node packages in there or does everything I do need to be AssemblyScript?

Aaron: Yeah. I'm glad you asked. So it ends up compiling to it. So this is really technical but it uses a compiler backend called Binaryen. And it pretty much sits at the same level as LOVM if you're familiar. So it uses their immediate representations you then compile to WebAssembly. And then on the note of using separate NPM packages, it would have to be AssemblyScript all the way down. So that is one thing that some people are like, oh, but my favorite image thing is it there? And it's like, we'll have to port the dependencies to you which is annoying, but the ecosystem is really growing. For example, I've been running a lot of URL packages lately for URL parsing. And I saw we have the testing library, there is someone who recently wrote a JSON parser. So there's lots of little small things popping up. The community is growing and it's all an NPM.

So if it says like, hey, look, I'm an AssemblyScript package the VM can install it, and then it should work in your MPM project. I think you asked one more thing.

Jeremy: No, I was just wondering again, if you have the same access to some of the low-level things. So if you're compiling down to WebAssembly using AssemblyScript, are you then able though to do things like HTP calls and access the file system and those things?

Yeah. So that's what you'd ask about WASI. So yeah, essentially you would say, hey, import WASI, and then it'll give you access to those fd read, fd write and stuff, you don't have to directly access through file descriptors. Another community project as-WASI, is kind of like sitting at the same level of Node where you import file system as a whole word. And then it has a create file or maybe not create file, read file, write file, whatever the API names are that I can't remember in my head. But we do into this interesting place. We have a project, an AssemblyScript project, that we're waiting for HDP to be standardized on WASI for that is just called AssemblyScript/Node, because ideally since the APIs are so similar there's nothing stopping us from creating a very as close as to Node as possible API where you can start to just maybe copy paste notes then they should just work, because the languages are so similar.

Aaron: One last ramble if you don't mind ...

Jeremy: No, go ahead.

Aaron: ... is that AssemblyScript is so similar to a TypeScript that even though you can't take AssemblyScript code and run it through the TypeScript compiler, you can take AssemblyScript code, do some small tweaking here and there and get it to run through the TypeScript compilers. You can have the same source code, and let's say you're running in an environment for whatever reason, doesn't support WebAssembly you could just compile it then to JavaScript and stab and call it a day. Which is I think as a testament to how similar these two languages are. So yeah.

Jeremy: Well that's portability too, right? That's really actually, that's cool. So you mentioned a little bit about the ecosystem and the community around it. That is a huge thing where again, things grow when stack overflow is a lot of people's friends, right? So if you can't get your questions answered there or you can't Google for some of these things. So I know you said you've got the ecosystem is growing and you've written a bunch of blog posts and there's a bunch of other people working in this. But if I'm out there and I'm working on AssemblyScript, am I going to be able to find a lot of blog posts on this or is this still very, very early?

Aaron: Yeah. So I think it's probably a little 50/50 not that it's that early. So I guess the reason I'm saying 50/50 is because there's not swaths of people in stack overflow, they're all the AssemblyScript experts that can answer all your questions. The community's a little tighter. So we have a discord server and we have a help channel that's very active. If you want to have a question asked by the person that writes AssemblyScript. Yeah. They're there almost ready to answer any questions, I'm there. Our teams about four to five people, I'm sorry, a sixth person, but we're all there. And I check at least maybe every couple ... whenever I have downtime at work, I'll check and see if anyone asks me any questions. So it's a tight community on our discord.

In terms of blog posts, there's a lot of blog posts. The only reason why I'm 50/50 there is because the project has grown really fast. So there are some blog posts even that I've given that's like, hey, here's how you use pointers in AssemblyScript. Yeah. Not everyone wants to do that. Now it's like, we have more mature runtime, garbage flexes to someone's stuff. So your mileage may vary depending on what block ... There's a lot out there, yes, but maybe 50% of them are still very relevant to the AssemblyScript you would write today. If that makes sense. Or be there like, here's how you write your own JSON thing, but now there's a package for it. So they don't do it? So yeah.

Jeremy: That makes sense. So have people been using this? Are there some success stories here? Have people been successfully building applications with AssemblyScript?

Aaron: Yeah. So just going chronologically in my head. The first one if I may a little self-plug here, is that the reason why I got involved in the project was because I was really excited about WebAssembly when I heard about virtual performance and I was like, oh, I got to get on this asap. So one thing I like to do was build emulators because I think they're really good at testing any new technology because you need graphics, you need audio, you need to make sure that it runs fast enough or things of that sort. So I built in Game Boy emulator called WASMBoy in AssemblyScript. And it's really early days where I was just pointers, memory, stuff's moving, but because it was a Game Boy emulator that's what you have to do anyways. And from that, we found a lot of bugs in the project and we worked through them together.

Yeah, that was probably I guess maybe got no one in with this only community early on, it was like you have to do the build the Game Boy thing, nice! So that's probably the first one. Probably the most recent one I can think of as of recently in terms of just general community is there's someone named Marshall Duin. I said their name right. But they work on a storyboarding application. I think it's called Micrio. And they wrote a whole article about how they were using Javascript and think Canvas at the time and things like that. And they had taken all those hot paths, those things that were computationally intensive and rewritten them in AssemblyScript and started using WebGL and stuff just updated the application to modern-day. But AssemblyScript was a huge part of that. And they wrote an article about like, yeah, I could read it, it made sense, compared to alternatives I didn't have to learn a whole new language they had a really good experience with it.

And I think I'm 99% sure she's on the article up right now, but they shipped it to production and their users are happy and they saw huge performance increases from just the nature of WebAssembly because they used it in the right places and things. So that's really exciting. And then probably more recently today Fastly has been using AssemblyScript that's one of our supported languages on computed edge. So it's still in data and we've had some customers try it out and had some really good feedback about it. And some folks really like it again because it's like, hey, look, I know JavaScript this isn't too much of a transition for me. Cool. Thank you. So that's really exciting.

And then I know Shopify publicly announced they've been playing with AssemblyScript a lot. I know from being on the team we chat with them a bit, but I don't want to get too into their business, but I'm very happy for them. But I would redirect you over to what they've said, just I don't see any wrong. Yeah. But they're trying us out as well, which is really exciting. And shipping stuff if I'm not mistaken. So, yeah.

Jeremy: Awesome. So it sounds like WASM and WASI, AssemblyScript, all of these things are coming together. It's growing, it's becoming more solid like you said, you wrote some blog posts where it was really low-level stuff, and then you started building ways that make that easier. So I guess there's still some limitations in here, it's probably not the right choice for everybody, but I'm excited about it because I think that it could change a lot of things, but I guess maybe since you work on this team and you're part of this ecosystem, what's the future? Where do you think this is going to go? And are we going to get to a point where we can use this for pretty much just everyday stuff?

Aaron: Yeah, totally. So the first one that I'd probably be most excited to bring up is this idea of nano processes. So that sounds really flashy. So I'm not the one championing this, I'm not the main person behind this. I would very much redirect you to Lin Clark's talk who gave a talk about this idea WASM summit of last year. But from how I interpreted when nano processes has this idea of that there's this concept of shared nothing linking. So for example, in the MPM ecosystem, if you have an MPM module that requires another MPM module, and this top MPM module was like, hey, look, I want access to everything that MPM or Node provides. And in this bottom module, and for some reason, this module is like, going to throw some things on global because I feel like it, and the bottom module is like, oh, that's cool. I'm going to figure out some way to get required by this person. That way I can start accessing the things that I didn't have to require.

Therefore, I look okay from a security perspective, but I'm taking advantage of some of the JavaScript type things that MPM allows. And we've had a lot of security, I guess, scares that they were valid scares because things in the ecosystem, because of this problem. So the idea of shared nothing linking is that in the future we're hoping that WASI modules can import other WASM modules. From that you have to declare again from the host call or whatever it may be like, I want these specific things so we can figure out like, okay when we run in the runtime, we know you only need access to fast. You don't need access to machine learning you, whatever crypto, you don't need access to that. Your parent module has access to that, but you don't need it. So we're only going to make sure you only have access to those things, not everything that your parent needs only what you need. So this creates a really good thing so that way you're not escalating out of MPM WASM modules to new WASM modules, because you're really only linking what you need and not the whole everything. If that makes sense.

And another big one is interface types. Again, so we're already on this topic of WASM modules importing other WASM modules and linking together and things. So not every language is the exact same. You talked about this language on let's say WebAssembly. So to see a problem where it's like, well Python strings are different from JavaScript strings that are different from Rust strings that are different from Java strings, what do we do? So this idea of interface types. So pretty much a specification for like, hey, look, if we know your language uses UTF-8 and this language is UTF-16 let's figure out a way so that when you y'all talk to each other, there's not a huge performance cost that needs to be paid of reencoding every single time y'all go back and forth between each other. Or if you just take a string and you just pass it to a third WebAssembly module, we shouldn't have to reencode it from like UTF-8 to UTF-16 back to UTF-8, it should just go straight through, you know what I mean?

So trying to specify that and find a where we can standardize that so that WebAssembly stays fast in that regard of passing memory around between other WebAssembly modules and leading towards that. And take it with a grain of salt, but that MPM-ness, everyone is connecting these packages, like LEGOs that build a bigger picture and a better structure using community and open source and things. So, yeah. Let me try to think, and then ...

Jeremy: Was there anything on performance? I know the performance is pretty good, but any updates or ideas on performance?

Aaron: So glad you brought that up. Thank you. So yeah, one of the big ones is SIMD. One of the champions for it, one of the most involved people is Thomas Lively and they give her a really funny talk, in which ... SIMD what it stands for is, Single instruction, multiple data. And I can't explain it over video, but the idea is that pretty much they're like, hey, let's say we have this four different array of four numbers and we want to add them all to another four numbers. If we were to do that in JavaScript, we'd be like, okay, well array zero, plus array zero of this one, because the new array zero array one of array one there goes a new one. And then in their slides, they're like, but we have SIMD we just add them directly and it equals any result. And you know it's faster because it took less slides to explain that then if you're going array zero to array zero.

So that's really the idea is that if you had a vector by the mathematical terms, but if you had an array of numbers and you want to add them to another array or do the same single instruction to multiple data, it just does it all once. And it's lots of performance benefits out of that. So getting that working in WebAssembly it has a ridiculous number of different instructions because you have to think, okay, well, if I have an i32 that's six long versus it just creates a lot of WebAssembly instructions. I'm probably rambling a little too much. And then I had one more note and threads. Thread is another ...

Jeremy: Yeah. Threads, I was curious about that because I didn't ask you that earlier. And I'm just wondering is when WASM runs, is it single-threaded multi-threaded, can it do to do multiple threads? How does that work?

Aaron: Yeah. So as of today WebAssembly is all synchronous and it's all single-threaded. In JavaScript ... I've given a lot of talks about this. If you use Web Workers which unlocks multiple threading on the web with WebAssembly, that is the performance, like chef's kiss. Yes, if anyone out there is building a computationally-intensive of intensive web app, Web Workers and WebAssembly is a match made in heaven. But for a lot of folks out there that wanted maybe reporting over a larger C application, for example, threads is launching off threads and WebAssembly itself is something that lot of folks are excited about. If I remember correctly last Google IO or Chrome Dev Summit or whatever it may be, VLC has been talking about a port to the web and what they're doing there.

And they're really building a lot of their port on top of WebAssembly threads, which a lot of folks at Google are doing a lot of specification work there. So it just shows like, hey, look, we can watch whatever random video format, the browser, only the supported ones using a VLC port. Yeah. And also, again, I'm not an expert on the work at VLC. So we can do some Google searches I can send a link or provide more research later, but yeah, those are super exciting. So yeah.

Jeremy: And I mentioned earlier this idea of an MPM type thing, but is it possible? Because this is another thing where I like about serverless is that, you can write one function in Python or another function in JavaScript and another function or Node, and then another function in Java or whatever. And of course, they're isolated, so they don't have to necessarily run together so you can use those different runtimes. But is that something that's possible? If somebody wrote a WebAssembly script in Rust and compiled it down and then somebody else wrote one in AssemblyScript, are those able to work together?

Aaron: So that's the future we're definitely headed towards and we're really excited about. There is a specification I mentioned earlier about the WebAssembly module for another WebAssembly module. We're hoping with this event called a Module Imports, which more of we're working on it. So essentially you wouldn't have to have all AssemblyScript or all Rust once they compile it at WebAssembly then they can say, okay, well we're both WebAssembly now. So let's import each other, it doesn't matter what our source code was. So module imports is hoping to solve that problem. And once we get more of these things standardized out while things are still young, that's the future we're headed towards of which my dream thing seems for you as well is like, oh, I would ... Personally I love all the work that Python does in the machine learning community, but Python's not for me, but I would love to access all the cool things that doing over there. But as a primarily JavaScript developer, I'm like, oh I'm going to play with Dom ...

Jeremy: And that's actually ...

Aaron: ... not to downplay myself, but you know what I mean.

Jeremy: No, no, no, no. No, but that's interesting because that's the thing where, think of a global NPM, you know what I mean? Where it's just some functionality was built and if you can pull that in to maybe you're AssemblyScript person because that's what you're familiar with, but you can't make AssemblyScript, do some complex crypto thing or whatever that may be some other language could and compile down because it's either further along or whatever it is. Being able to share those across applications and reuse those. That sounds pretty exciting.

Aaron: Yeah. It sounds super exciting. Definitely, I think that's the future we're headed towards. As a community, we'll see what gets announced here and there and things. And again, there are some folks trying it out already, this idea, but no, we'll see. Really it began also, as you are very bullish on this idea of like, yes, we can totally have this cross. It's no longer are we bounded to our languages. It's like, we just write code and we can all be one large ecosystem. It's going to be awesome. I'm looking forward to it.

Jeremy: Cool. So I want to ask you one more thing though. So let's say that I'm a developer I'm listening to this. And I say, WebAssembly sounds amazing, how would I convince my development team or more importantly, probably my boss, how would I convince them, hey, this is something we should start investing in?

Aaron: Yeah. I'm so glad you asked. So the thing about WebAssembly that's really interesting, and I get asked this often, it really depends on what you're doing. We've maybe covered like, oh, you can do this, you could do that, you could do this, so it depends what your area is. For web applications definitely if you're building an application that does computationally-intensive things, maybe you're working on an online photo that, or whatever it may be. Maybe you're working on a spreadsheeting tool where you have to do complex math functions. Just understanding really where WebAssembly fits into an application that's where you can start to make the pitch to your boss. So if you're really excited about this right now, I have some resources that I built called WASM By Example which is how you can get started with little bite-size examples of WebAssembly, but another one is Made with WebAssembly.

So wasmbyexample.dev and then madewithwebassembly.com. WebAssembly is just a showcase of a bunch of folks using WebAssembly in production and side projects pushing like, okay, well what's what WebAssembly good at, these are some examples. So if you can go on madewithwebassembly and you're like, oh, my app does something similar to that, maybe you're on the right track. Maybe it's worth considering well, what your team needs and things. So yeah, I guess really just understanding what WebAssembly is good at and then see, how can we fit this into our application is probably your best bet, because then you can make a solid argument for why.

Jeremy: Awesome. All right. Well, I'm going to put all of that information in the show notes. I know there's a bunch of links that you have and all really good documentation and just good ideas, like you said if you want to convince your boss or whatever to use it. So I'll make sure I get all that in the show notes, but Aaron, thank you for sharing this. This was super exciting because I just don't know enough about this stuff. So anytime I can, I just love learning this stuff. And I'm super excited about the computed edge stuff. I think WebAssembly is such a great packaging format for that. But I'll put some of those other links you mentioned in the show notes, but if they just want to find out more about you how do they do that?

Aaron: Yeah, totally. Probably the best way to get ahold of me is on Twitter. So my username is torch2424, just as a quick, people ask me why? I got to be in it really young when I was like six, Torch was my imaginary friend. And then 24 is my favorite number when my kid Brian was 24, 24 twice. So ...

Jeremy: There you go.

Aaron: Yeah. Hopefully that story helps drive with my name, but yeah, I'm on Twitter most of the time. Yeah, probably reach out to me on Twitter, probably the best. I'm trying to think. If you like ...

Jeremy: I've got your LinkedIn here, I've got your GitHub which is torch2424 as well, check out Fastly.com and everything that your team is working on over there. And then I think, yeah, you mentioned The Bytecode Alliance too, right? So that's just bytecodealliance.org. Probably some great information there and then I'll get everything else in the show notes so people can check this stuff out. Give them some weekend reading to do, but thanks again, Aaron. I really appreciate it.

Aaron: Yeah. Thank you very much for having me. I super appreciate it. And yeah, looking forward to keeping in touch and things.

View Details

About Anahit Pogosova

Anahit is an AWS Community Builder and a Lead Cloud Software Engineer at Solita, one of Finland’s largest digital transformation services companies. She has been working on full-stack and data solutions for more than a decade. Since getting into the world of serverless she has been generously sharing her expertise with the community through public speaking and blogging.

  • Twitter: @anahit_fi
  • LinkedIn: https://www.linkedin.com/in/anahit-pogosova/
  • Solita: https://www.solita.fi/en/
  • "Mastering AWS Kinesis Data Streams, part 1”: https://dev.solita.fi/2020/05/28/kinesis-streams-part-1.html
  • "Mastering AWS Kinesis Data Streams, part 2”: https://dev.solita.fi/2020/12/21/kinesis-streams-part-2.html
  • AWS Community Day Nordics 2020: https://youtu.be/gtE2o8qsq-4

Watch this episode on YouTube: https://youtu.be/7pmJJcm0sAU

This episode sponsored by New Relic and Stackery.

Transcript

Jeremy: So you mentioned poll-based versus stream and things like that. So when you connect Kinesis to Lambda, this is the other thing too, I think that confuses people sometimes. You're not actually connecting it to Lambda directly for pretty much all of these triggers in these integrations. There's another service that is in between there. So what's the difference between the Lambda service and the Lambda function itself?

Anahit: That's a great one because I think it's, again, one of those very confusing topics, which are not explained too well in the documentation. And the thing is that when you're just starting dipping your toes in the Lambda world, you just think that, "Okay, I write my code, and I upload it and deploy it, and everything just works. And this is my Lambda," right? But you don't really know how much of the extra magic is happening behind the scenes, and how many components are actually involved into making it a seamless service. And there is a lot of components that come into ... so you can think of a Lambda function as the function that we actually write and deploy and invoke. But then the Lambda service is what does all the triggering, invoking and batching and error handling.

And it really depends on the way the Lambda works, or the way long the service works. It really depends on the invocation model, is you prefer to the poll based, not poll based. So again, one thing that is not too clearly explained, in my opinion, is that there is actually three different ways you can work with Lambda or communicate with Lambda. So you can invoke a Lambda synchronously. So request response traditional way, and the best example, I think, is API gateway, which does that so it requests something from Lambda, it waits for the response. Then there is the async way, which is one of the most common. So you just send something to Lambda and you don't care about what happens next.

Jeremy: Which uses an SQSQ behind the scenes to queue ...

Anahit: Exactly. Yes. That's also like fun facts that you learn along the way. But the point is that like services like SNS, for example, or S3 notifications, they all use the async model, because they don't care about what happens with the identification. They just invoke Lambda and that's it. But then there is this third, gray area or a third totally different way of invoking the Lambda function, and it's called poll-based. And that's exactly how Kinesis operates with Lambda. And it's meant for streaming event sources, so it's both Kinesis data, DynamoDB streams. Also, Kafka currently uses poll-based model. And it also works with the queue of event sources like SQS.

Jeremy: Right. SQS, yeah.

Anahit: And Amazon MQ, I think they also use them, the poll-based method. And what poll-based invocation or the component that is most essential in the poll-based model, it's called the event source mapping. One of the misunderstood components or one of the hidden heroes, I would say, we find in Lambda, because it's an essential service or essential part of the Lambda service. And event source mapping actually takes care of all that extra things that Kinesis plus Lambda combination is capable of. So it's responsible for batching, it's responsible for keeping track of this point in the stream and where a shard, where it's ...

Jeremy: A shard iterator, because anybody wants to know the ...

Anahit: Yes, exactly, shard iterator.

Jeremy: ... technical term for it.

Anahit: Yes, thank you. And, yeah, the most important for me, it handles the errors and retries behind the scenes.

Jeremy: Right.

Anahit: And basically, if you don't have event source mapping, you can't have batching. So it takes care of accumulating, or in case of standard, consistent consumer, it pulls your Kinesis stream, on your behalf, it accumulates batches of records, and then it invokes your Lambda function with that batches of records that it accumulated. Again, in case of enhanced fan-out, of course, it doesn't poll, it gets the records from the Kinesis stream directly. But then from the perspective of your Lambda function doesn't matter, it just gets triggered by the event source mapping, because as you've said yourself, it's not the Lambda that you connect to Kinesis stream, it's the event source mapping that you connect to the stream, and then you point your Lambda to that event source mapping, so.

Jeremy: Right. So you can connect a Lambda function or the Lambda service directly to the Kinesis stream itself, or you can use enhanced fan-out and push it to the Lambda function. Although, for all intents and purposes, it's pretty much the same thing.

Anahit: Yeah. And for your Lambda function, it doesn't really matter how that data ended, or how those records ended up there, you just get a batch of records, and then you deal with it. And I mean, all the rest is pretty much the same from the perspective of a Lambda function, because it's nicely abstracted behind the event source mapping, which hides all that magic that happens behind the scenes.

Jeremy: Right. So you mentioned some aggregations stuff in there and about like Windows and time windows and things like that. So tumbling windows, that's something you can do in Kinesis, as well. Can you explain that?

Anahit: Yeah, it's a feature that actually came out very, very recently. In the end of the re:Invent, I would even say, and I think it was like one day before I was going to publish my second part of my blog post that was already finally ready to submit it, and then in the evening I get this and I was like, "Okay, I have to write a whole new chapter now." But it is a very interesting aspect, you can use it with both Kinesis and DynamoDB streams, actually, so it's available for both. And it's a totally different way of using streams, which wasn't there before. So with Lambda function you know that you can retain state between your function executions unless you are using some external data source or database.

And here, what you're allowed to do with this tumbling window is that you can persist the state of your Lambda function within that one tumbling window. So tumbling window is just a time window, it can be at maximum of 15 minutes, and all your invocation within that 15 minutes max interval, they can pass the state and aggregate the state of passing to the next Lambda invocation. So you can do cool things like real time analysis that you could previously do only with Kinesis data analytics, for example. Here you can do right inside your Lambda. And then when the interval is ending, the 15 minutes, for example, interval is ending, you can send that data somewhere, let's say to a database or somewhere else. And then the next interval is starting, and then you're accumulating again.

And so it's pretty fascinating in the sense that it allows you to do something that wasn't there before. It's a completely different way of using the Lambda basically, with the streams. But of course, there are limitations with that, you can only aggregate the data on the same chart because one Lambda is processing one shard at a time.

And then there is also this thing called paralyzation factor, which we haven't talked about. But which basically means that instead of having one Lambda reading for one shard at a time, you can have actually up to 10 Lambdas that are reading from that same shard. So you can boost the power of reading, because if one Lambda, for example, if Lambda execution takes too long, and you can't keep up with your stream, then you can either add more shards, for example, to your stream, but it's expensive, that takes time and has some limits. Or then you can immediately just throw more Lambdas at it, just say like more horsepower, and they will take care of it. But if you have more than one Lambda reading from a chart, you can't use this new tumbling window features, which makes sense, of course.

Jeremy: Right. And that depends on what you're doing because I mean, the idea of the parallelization factor, such a hard word to say. But the whole point of that is to say you're reading up to 1000 records per second off of this stream. And if for some reason it takes more than a second to process one of those records or whatever, then you're going to see the problem with not being able to process enough records quickly, because you're backing up your stream if you're writing to it. So again, it's just one of those trade-offs.

Anahit: Yeah, but again, this new feature, I think it's going to be developed still, maybe someday it's going to have some kind of support for it, though I can't see really how under the hood, it would function between different Lambdas in the central. But anyhow, this, I think is a very cool new thing that I'm actually eager to try out in production if I just figure out a case for that because it just looks so cool. And it's so simple to do.

Jeremy: Yeah. Well, I mean, and the other thing is, is that depends on what you're doing with it. So the use case that I've seen, and actually I started playing around with like SQS batch windows to try to do something similar. I know they're different, but they're same idea where when you're doing aggregations, if you're just reading off a stream, and you're trying to aggregate, you have to grab that data from somewhere, like you said, because Lambdas are stateless.

So you have to query a DynamoDB database or something like that, or table and then pull back what the last aggregations were. And then you read in the data from the stream, and then you write, you do your aggregation there, and then you write it back to the DynamoDB. If you're doing that hundreds of times a second, that is pretty inefficient, where if you just did it, set your tumbling windows to one minute even, and you could read thousands of records, and then be able to just write that back to the database one time, just the efficiency gain there is huge.

Anahit: Exactly, exactly. If you have a use case that is like that, because I personally don't, that's why it's ... I'm trying to come up with one in the future ...

Jeremy: Come up with ...

Anahit: Yes.

Jeremy: Find the problem for the solution, right?

Anahit: Yes, exactly. But yeah, it can be very, very helpful. And again, it's pretty straightforward to using it. So I can see a lot of people loving it really.

Jeremy: Yeah. Awesome. All right. So another thing I just want to mention, because we keep talking about Lambda, and we mentioned the concurrency thing and some of those other bits. In terms of provisioning shards and having one Lambda per shard, and then potentially, if you do the parallelization factor, 10 Lambdas per shard, if you had 100 shards, because you had a lot of data coming in, and you had the parallelization factor turned on, then you've got 1000 concurrent Lambdas being run at once, which I did ...

Anahit: And guess what happens next?

Jeremy: So what happens next, yeah. And the people don't know the soft limit in any region is 1000 concurrent executions, for your Lambda concurrency. So, that's just something that people need to think about.

Anahit: Yeah, for sure. And that's something I bring up quite often, because we've been there, done that, but actually 100 shards is not even too much. There are apparently streams with 1000s of shards. So we have something like 40 shards in our stream. So it's a really quite, quite decent amount. But yes, as you said, exactly, so if you have, for example, 100 shards, and then you have a parallelization factor of 10, you will have 100 times 10, 1000 Lambdas, running or consuming that stream at all time. So there will be constantly 1000 Lambdas, concurrent Lambda invocations. And you probably won't run into any problems until there is some other Lambda in that same region in the same account, that is probably very business-critical, that does something very important, and then it starts to fail for some unknown reason. And that reason is not even that Lambda, the reason is your stream, which is consuming the entire budget that you have allocated for Lambda.

So yeah, it's something people oversee quite often that though Lambda scales endlessly, potentially. In reality, all the services, they come with the safety mechanisms of soft limits, there is no service in AWS, I think that that is ... it comes out of the box with no limits, just use it as it is. So basically for your own safety, there are some soft limits. And on the other hand, though, they are soft, which means that you can increase them by submitting a ticket to support. It will take some time I warn you, especially if you go higher than normal.

But though you can do that, there still is going to be a limit. There is always going to be limit. And you just need to know that it exists because one day, you're probably going to hit it. And so you have to monitor that all the time. And yeah, that's one thing to keep in mind. But that's that's a common thing with SQS as well, for example, and maybe SNS. So all the services that can scale Lambdas pretty much like out of hands, then you're faced with that concurrency, Lambda concurrency limits that you have to be careful with.

Jeremy: Right. One limit that I love that has nothing to do with Kinesis but with SNS is for, if you send SMS messages with SNS, I think the default limit is $1 spend per month. So if you send like 200 text messages or something like that, it ends up cutting you off. Maybe not 200, you can probably send more than that. But it is a very, very low limit. And I think it's just because they don't want people, I don't know, spamming SMS, or something like that. But anyway ...

Anahit: Can you increase it? Is it soft limit?

Jeremy: No, no, you can increase it, yeah. But you have to submit the ticket. But basically, I remember, I set up a new account for something and we were doing all these alarms, and it was like, within two days, I get a message saying ...

Anahit: "That's it. That's enough."

Jeremy: ... "No, you can't... no, you've exceeded your limit." And I was like, "Well, that was fast." So and if you're using something like control tower, or any of these things to provision hundreds of accounts, some of these soft limits that are in there can affect you. So, whether it's Lambda or some of these other ones, but ...

Anahit: It's surprisingly easy to reach all of the soft limits, as long as ... I mean, as soon as you go to the real world cases from those who, "Hello, World!" cases. And yeah, it's not a problem reaching the limits, per se, the problem is many people don't know that they are there.

Jeremy: Yeah, yeah good point.

Anahit: That's when the problem starts.

Jeremy: Good point. All right. So speaking of limits, there are limits, obviously to Kinesis. And some of these things that maybe even go beyond some of the soft limits. I mean, there's just limitations in distributed systems, and there's limitations in network throughput and some of those other things. And so as you hit some of those limits, or maybe let's just talk about errors in general, as you start to run up against problems, whether they're caused by limits or whether they're caused by something else, what are some of the things that I guess, could go wrong when you're using Kinesis?

Anahit: Right. Well, that's my favorite topic, really. But I mean, with every service, as you said, I mean, nobody says it better than Werner Vogels who says that, "Everything fails all the time." And I love that phrase because that's true.

Jeremy: Very true.

Anahit: And it's not because you want to be pessimistic, but rather, because you want to be prepared, and you want to sleep better at night. Because if you're not prepared, then surprising things will happen eventually. And for me personally, with Kinesis, or any other service, really, when I start working with a new service, first thing basically that I ask is that, "What are the ways in which it fails? What are the possible errors? What are the possible limits? Are they hard limits? Are they soft limits?" All this. And even what's the built-in functionality for retries, for example? What are the default timeouts? And that kind of thing?

So those are very common that things that you start to question after you have got a lot of headache with one of the services. And then you start to question those specific questions when you start working within your service. And with Kinesis, you can probably separate the errors for writing to Kinesis stream, and reading to Kinesis stream.

So for writing, well, first, it's nice that AWS SDK has a built-in functionality for retries for all the failures or system failures that happen and timeouts as well. And it's not documented or it used to be not documented too well, because personally, I learned about the built-in retries, when I was developing unit tests, and then they were behaving weirdly, and I was like, "Something's going on here. What's that?" And then I realized, oh, it retries three times by default for every system error. Wonderful, that's a wonderful news.

But maybe not so wonderful news is the thing called partial failure. And it's actually very common for all the services that are using batching. So what it means is that when you, for example, write a batch of records to Kinesis, it's not an atomic operation, it's not either the entire batch succeeds or entire batch fails, and you get an error code back from Kinesis. The reality is that you almost always get a success code back from Kinesis, and it's very misleading, because parts of that batch could have failed, and you don't know about that. And what you should do, instead of just waiting for an error to come back, what you should do instead is to look at the response that comes back from Kinesis. And to see if there is this field called a failed error count or something like that, which basically tells you where they're actually failures within that batch that didn't go through to the Kinesis. And that can happen, for example, because of throttling. So some of the records just didn't made it, they didn't make it to Kinesis stream.

So, that's that's one of the basically main issues that we have had with Kinesis streams. And you have to take care of those partial failures manual and you have to do some smart retries and backups, and random detours and things like that. And then there are the timeouts of course, which always happen and you need to know the kind of the default settings for the timeouts. Because in case of Kinesis for example, the service times out after two minutes. So, and actually, there're two timeouts. There is a timeout for a new socket connection, and then there is a timeout for sending a request.

So first, you will wait two minutes to create a socket connection, and then you will wait another two minutes for sending the request, then it will be retried three times. And then like 10 minutes in and you're still waiting for one batch to go through, in a pessimistic scenario. And again, those are things you don't really see in the documentation right away, and those are the things that you end up finding out because you have some problems. And the other point is that you almost or you always have to set the timeouts to a lower value than the default two minutes.

Jeremy: Yeah, those defaults ...

Anahit: It's crazy.

Jeremy: ... defaults are not great.

Anahit: Yeah. No, no, not at all. So those are the main things with writing. So like partial failure, sometimes timeouts and that kind of things. But with reading, things get even more interesting, because there's so many options. And one of the things that is very common, it's called poison pill record.

Jeremy: Yeah, the poison pill.

Anahit: Oh, the poison pill, yes. And nowadays, it's actually pretty avoidable, but let's get back to it later. But the idea of poison pill is that if you have a Lambda function attached to your Kinesis stream, and it's reading from the shard, and everything is fine, until there is some corrupt record in your shard for some reason. And your Lambda function tries to read that record and it fails, and then you don't have proper error handling because well who needs error handling, and then your entire Lambda function fails, right? But when your entire Lambda function fails, what happens is that Lambda returns or as we know, event source mapping actually returns the entire batch back to the stream. And then it retries or makes Lambda retry with that entire batch that just failed.

Jeremy: And speaking of defaults, it retries 10,000 times I think by default?

Anahit: No, you're too optimistic. By default, it retries forever.

Jeremy: Oh, forever.

Anahit: Yes, it retries until the data expires, which means from 24 hours to up to seven days or one year. But let's explain why it's not good, it's not a good thing. Well, first of all, you don't want to have all these unnecessary Lambda invocations that don't do anything. They just send to the same records, and then they ...

Jeremy: They keep failing ...

Anahit: Yes. They keep failing at the same point of the batch, and then they start all over again. And it's pointless. But the problem is that in, well, let's say 24 hours, let's take the optimistic scenario, so in 24 hours, the batch expires finally, and Lambda can forget about it. So the batch gets deleted from the stream, and then the next batch comes in. But the problem here is that by that moment in time, probably you're streaming or your shard is probably filled with records that were written around the same time as the records that you were trying to process, which means that they are expiring around the same time, like the previous batch.

And if your Lambda is not quick enough, you might end up in a situation when records end up falling from your stream. This overflowing sink analogy that I had in my blog post, when you basically pour water to the sink more quickly than you can drain it. And then the water ends up on the floor. So, that's the exact situation. So basically, what ended up happening is just because of having one bad record, and no proper error handling, you ended up losing a lot of, or you can potentially end up losing a lot of records. So hence the poison pill because that one bad record poison, poisoned the entire shard basically.

Jeremy: And I'm actually curious, something I've never tested this, but let's say that you get a batch of records, it's only say 100 Records, because there's only 100 records in the stream. So it sends that batch to Lambda and then Lambda fails, because there's a poison pill and it sends those 100 records back. If another 100 records come in, because it's still under whatever your threshold was for batches, would it then send in like the 200 the next time and then will it keeps sending in up to the full batch amount as it retries those batches?

Anahit: I would imagine it should really. That would make sense. We just never had the situation because we usually have the complete batch.

Jeremy: Right, you have a full batch.

Anahit: But I would imagine that's how it should work yet, because it accumulates entire batch. But it doesn't matter because it will stop at the exact same record. It processes them in order, it will stop that exact same record. And then well of course, if you process them in order, you can process all the records in parallel, if you want to. But then you want to have the ordering. But yeah, it's very funny situation, and it's very easy to end up in it, and been there, done that once again, but luckily ... Well, first of all proper error handling in your Lambda function where you don't allow the entire function to fail just because of one record that didn't go through. And then there are different ways to approach that.

And then the things that I was talking about a lot or mentioning a lot is the error handling that comes out of the box with event source mapping. And nowadays, and actually, it's developing. And each year, they are adding new functionality and new possibilities that weren't there before. So what you said about 10,000 retry attempts, it's a totally new feature, it wasn't there. They added this maximum retry attempt settings to the event source mapping. But again, by default, it's minus one, which means that it does it infinitely. So, but you can set it to up to 10,000 if you want to. And then you can set the maximum age of the record that Lambda will accept. So if the records get older than some specific age, I think it can be up to one week even, you can ... your Lambda will keep the records in one process them.

And then there is on-failure destinations where you can send the information about your failed record if everything fails. Then I think that one of the fun possibilities is the cold batch bisecting. So it's when you basically split your problematic batch in two and then Lambda tries to send these two parts separately, and then hopefully, the other one succeeds. And then it continues with the failed one and splits it recursively further until hopefully, you end up with just one bad record.

Jeremy: Just the one.

Anahit: Yes, but on the way there, you actually end up sending same records over or processing same records over and over and over again. So it's not optimal. And then there was one more announcement around the same time, because of which I had to update my workflows. I think it's called custom checkpoints.

Jeremy: Custom checkpoints, yep.

Anahit: Yeah. It's basically common sense. Instead of just failing your Lambda saying, "Well, no can do. There was a batch, I don't know, something bad happened." Instead of that, you can return the exact sequence number of the records back to the stream, the record that caused the problem. So if you went on with your batch, you processed your record, and then you return that and back to event source mapping, and it knows that, "Okay, next time I retry, I will start from that end, rather than starting, again, from scratch." So.

Jeremy: And that should eliminate the need to do the bisecting?

Anahit: Yeah, that's ...

Jeremy: The bisecting. Yeah, right. So if you're ...

Anahit: That's what I'm thinking.

Jeremy: ... have an existing system that is using bisecting, you don't have to change it. But AWS likes to do that, where you keep the old functionality in, but there's a better way to do it. The same thing with dead-letter queues and Lambda destinations, right?

Anahit: Yes, exactly. But for the sake of it, if you like the idea of just kind of splitting your batch, and sending them separately behind the scenes without doing anything, well, you can have that. But yes, of course, this new functionality would be so much better, because you would avoid all this unnecessary read processing of the same records. Yeah.

Jeremy: Right. So what are some of those other common issues? I mean, you mentioned timeouts, and maybe like network issues are obviously happened, but what are some of maybe the other distributed network things that pop up when you're using Kinesis?

Anahit: Yeah, I think the timeouts and network problems are really the core of it, most of the times really. And the other one that I've mentioned several times already that at least once a guarantee, so-called a response guarantee, so that it then prevents the ... Basically, with Kinesis, you are not guaranteed to get your data exactly once, it's at least once. So you will have duplicates in your stream. And it's because of, for example, the retry functionality that we just discussed, both with sending and receiving the records. But also, the fact that, for example, the network issues also contribute to that because you might have sent a batch of records to Kinesis, but never heard back from it. You just didn't get the message so to speak. And then you will retry it because you don't know either it went through or not. And then maybe it did go through and then you end up writing the same batch all over again.

These are the things that happen pretty much all the time. And the only thing or the only way to deal with them, it's just to know that they happen and to be prepared for them, with at least once guarantee your downstream systems must be resilient in the sense that they won't change if the same data comes over and over again. So they need to be able to handle that repeating records in your stream. And then with the network problems, well, there's not much you can do about network problems. Of course, if you have a producer that is running inside VPC, creating a Kinesis VPC endpoint is a good idea, so the traffic won't leave your VPC. But pretty much, that's the only thing you can do about those.

But on the other hand, you can handle those issues with ... or let's say timeouts are also network issues in some way or quite often. And the thing that we were discussing before that default timeouts are really not that great, most of the time you need to adjust those with Kinesis, especially, not especially, but that's a good example, maybe. But actually, one fun thing I remembered about the timeouts is related to DynamoDB, which are probably familiar to you, in a sense, because the DynamoDB also has some ridiculous default timeout, like a minute, two minutes, something like that.

And when a couple of years ago, at re:Invent, I was speaking with one of DynamoDB guys, and was asking that, "Okay, we have this API that needs to retrieve data from DynamoDB, and it needs to be very, very quick. So latency should be very low." And we used to have Lambda in between, so Lambda was doing calls to DynamoDB. And the first thing he said was, "Reduce the timeouts." Because apparently, DynamoDB can timeout pretty frequently. So it's much better to drop the connection sooner rather than later. So you set the timeout to, I don't know, 1000 milliseconds, and then you let the SDK handle the retry, instead of waiting for, like forever. But that was funny. That was the first thing that they recommended me to do. "Okay. "

Jeremy: Yep. Even though they set those defaults pretty high, but ...

Anahit: Yeah, exactly.

Jeremy: All right. So then, in terms of monitoring this, though, I mean, that's one thing that I really like about Kinesis is that you do get quite a few metrics where you can look and see how your shards are doing, how quickly they're being drained, how backed up they are, and stuff like that. What are some of those, I guess, the most important metrics that you want to keep your eyes on?

Anahit: Right. So of course, there are separate ones for writing to the stream and for reading to the stream. So I would say for writing, what is it, right throughput exceeded exception, is like the metrics that tells you that you exceeded the throughput of your stream basically. So that's the one that was pretty much eye-opening for us, because well, the thing is, I think, with metrics in general is that they are at best minute-based. So they are aggregate metrics, or aggregate values over one minute time, right? And with Kinesis, as we have mentioned several times, all the limits are per second. So it's 1000 records per second one, one megabyte per second. And that's the information you don't get from the metrics. So you don't see the picture, per second picture.

So there is a metric that tells you how many records come in and how much open data comes in. And you might look at those and think, "Okay, the threshold is still far, far away, I for sure have no issues with the stream." And then you notice that there is this provisioning throughput exceeded exception metric that is being popping up, and you figure out that, "Okay, apparently, bad things can happen even in this situation." Of course, it's because of, for example, the network's issues that we discussed before, or spike in traffic, because the records ...

Jeremy:
Lots of traffic.

Anahit:
Yep. The records arrive to your stream on uniformly in a way, so it might be that one second, it's like 5000 records. And the next second, it's just like three direct records. And you can see that in metrics, or even one metric. You have to observe the metrics that tell you what goes wrong in a way. That's the key, I guess. And same goes to reading from the stream, really. There is this rich provision throughput exceeded, which is basically only for a shared iterator case. So the standard reading, consuming the stream. So when you exceed, for example, two megabytes, or you exceed this five requests per second, which we don't even go into. Read my blog post, you will know what I'm talking about.

But you get those, and then there is, I think the most important one is the iterator age when it comes to reading from the stream, because that's the one that tells you that kind of age of the record, meaning how long they have been in that stream. And apparently, if the age increases, it means that you can't consume them fast enough. So then you might have a problem. They are with your consumer, for example, or you have too many consumers, and then you have to have the enhanced fan-out and things like that.

But they're basically like two, three metrics that you have to keep an eye on. And if you see any issues with those, then you have to dig deeper, maybe enable the enhanced metrics, which are not stream level, but they are shard level metrics. For each shard, you can have the same or similar information so you can diagnose it more precisely.

Jeremy: Right, yeah. And if it was only serverless, or I should say fully serverless, and do this automatically for us, that would be much better.

Anahit: Yes.

Jeremy: Well, so just like your blog post, this episode turned out to be quite lengthy. And but I hope people got quite a bit of knowledge from this, and are not afraid of using Kinesis, because it's an amazing service. Yes, it has all of those caveats that we talked about, but it's still an amazing service. But if you've got a few more minutes, I'd love to just pick your brain for a second, because I think there are a lot of common misconceptions about building serverless applications, and again, whether Kinesis is serverless or not, we'll put that aside. But just all of these different services, even Lambda, and having to build in the retries, and know about either bisecting or using the custom checkpoints or doing some of these other things, there's a lot that goes into it. So what are some of the ... and maybe just even from your own perspective, when you're building serverless applications or using fully managed services? Like what are just some of those misconceptions that maybe people have?

Anahit: Yes, I've noticed those, well, few of them actually when working with serverless. And people usually have strong opinions about serverless. It's either they go both ways. But I think many people assume that it's either very easy, or then and you don't have to do anything, everything is done for you, or then it's way too complicated. And I think, again, Yen Cui had a nice blog post lately about the complexity of serverless, or perceived complexity of serverless. And what he was saying is that serverless is not complex, it just reveals the underlying complexity of the systems that we used to build before. So all those things that were built in and hidden from everybody's eyes, but there was still there. Now, they are more obvious with using all the different components, and you connect them to each other, and you have all that ecosystem living there, but ...

Jeremy: Which gives you more control over the individual components as well.

Anahit: ... it gives you more observability, it gives you more control and all these nice things why we love serverless. So I'm all for it. But on the other hand, I think it's a simplistic view to think that fully managed and serverless, it means that you basically just deploy your code, and you have to worry about nothing. Because as we discussed with you several times already, yeah, you will probably get away with that on the "Hello, World!" level, it will be pretty much okay. But then when you get to the real world and real world scale, you actually do need to know in quite some detail how each and every service that you are using, how they work, and how they fail. Because once again, they will fail at some point, and you basically ... you need to know how they fail and what can happen just to sleep at night.

Jeremy: Yeah. And I also think just this idea, that again, they said it and forget it for simple things, like you said, yes, but just ongoing management, right? I mean, and optimizations and shards, refactoring code and with the shards thing, with monitoring that and saying, "Hey, we're starting to creep up to this next level, or maybe we're not processing fast enough, or maybe our shard iterator keeps pushing over a certain amount of time during certain times of the day."

Anahit: I'm getting anxious, now.

Jeremy: All right. You want to go back and look at all those metrics, right?

Anahit: Exactly! But that's exactly right, but maybe will sound scary. Well, we'll put it that way, but on the other hand, the ... again, well, it reveals the complexity your systems do have anyway. And the good news here, I think is that in case of AWS, there is a lot of commonalities in how services work.

Jeremy: Yeah, true.

Anahit: And once again, I think understanding of one service through and through will help you to understand all these issues with the distributed systems and under errors and built-in retries and whatnot. So you don't really need to remember every single thing by heart, and it's not as overwhelming as we make it sound at the moment. It does require some work, but I think it's well worth it.
Jeremy: I totally agree. Well, Anahit, thank you so much for taking the time to talk with me and educate the masses about Kinesis. If people want to find out more about what you do or want to contact you, how do they do that?

Anahit: Well, first of all, they need to read the blog. It's long, but I hope it's worth it, and it has some nice pictures, so some benefits. Then they can reach me on LinkedIn, first name, last name. And Twitter, again, first name, last name. And yeah, I think that's about it.

Jeremy: Awesome. And then the blog at solita.fi. And then you've got a really good talk that you gave, I think it was at AWS community day, maybe Stockholm. So, then that.

Anahit: Oh my God, it's been over a year already. That was the last trip that I made before ... it's horrible.

Jeremy: Isn't that crazy? I know. It's been a year, it's been a year.

Anahit: It's been a year.

Jeremy: We just celebrated or, celebrated I guess ... there was just a year passed for ServerlessDays Nashville which was the last conference that I went to in person. So I am looking forward to doing that again and bumping into people and talking to people about this in the hallway because those are the best conversations. So-

Anahit: For sure.

Jeremy: ... anyways, I will take all of this stuff, your Twitter, LinkedIn, blog, the two blog posts that you wrote about this, as well as that video talk from community at Stockholm. I will put all that into the show notes. Anahit, thank you again so much.

Anahit: Thank you so much, Jeremy. It was so much fun.

View Details

About Anahit Pogosova

Anahit is an AWS Community Builder and a Lead Cloud Software Engineer at Solita, one of Finland’s largest digital transformation services companies. She has been working on full-stack and data solutions for more than a decade. Since getting into the world of serverless she has been generously sharing her expertise with the community through public speaking and blogging.

  • Twitter: @anahit_fi
  • LinkedIn: https://www.linkedin.com/in/anahit-pogosova/
  • Solita: https://www.solita.fi/en/
  • "Mastering AWS Kinesis Data Streams, part 1”: https://dev.solita.fi/2020/05/28/kinesis-streams-part-1.html
  • "Mastering AWS Kinesis Data Streams, part 2”: https://dev.solita.fi/2020/12/21/kinesis-streams-part-2.html
  • AWS Community Day Nordics 2020: https://youtu.be/gtE2o8qsq-4

Watch this episode on YouTube: https://youtu.be/U4snzWHMrtU

Thanks to our episode sponsor, Epsagon.

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Anahit Pogosova. Hi, Anahit, thanks for joining me.

Anahit: Hi, Jeremy. Thanks so much for having me.

Jeremy: So you are an AWS community builder and also a lead cloud software engineer at Solita. So I would love it if you could tell the listeners a little bit about your background, and what it is you do at Solita.

Anahit: Right. So yes, so I have been working at Solita for pretty long time. So it's a digital transformation company. It was originated in Finland over 25, 26 years ago, and out of those years, I have been on-board for 11 years. Which sounds extraordinary nowadays, I suppose, because everybody gets surprised. But during those years, I've had several roles as a backend and full stack developer. And then I moved to the cloud, to AWS and started doing all the cool stuff with serverless. And I have been also working as a data engineer for several years with one of our customers, so a lot of different stuff.

And we actually have offices in six countries in Europe, of course, they are empty at the moment. And I'm based here in Finland. And yeah, we focus on software development and cloud integration services, analytic services, some consultancy, and service design. So if you're interested, we are hiring. And yeah, that's about Solita and me.

Jeremy: Well, any company that can retain someone for 11 years, sounds like a good place to work at.

Anahit: Right? I think so too. No, apparently, it sounds suspicious to many people. Why exactly?

Jeremy: I don't know. That's a conversation for another podcast, I think, about the job-hopping thing. But anyways, well, I'm glad that you're here. And thank you very much for taking the time to talk to me. I'm super, super excited about this topic, actually, because I came across this blog post that you wrote. Now, this was actually the first version of this that you wrote was, or the first part of this, I think was maybe almost a year ago now or something like that.

Anahit: Yeah, something like that.

Jeremy: But then you had a second part of it that came out in maybe November. And this was two posts, they were called "Mastering AWS Kinesis Data Streams." And now the cool thing about Kinesis is, it's a super powerful service. I think we learned from a recent outage at AWS that Kinesis, pretty much powers everything, every backend service at AWS is powered by Kinesis, which is pretty cool, but also scary at the same time. But, but it's a fascinating service. And I want to warn the listeners, because I want to get super technical with you. I want to get into some of these different details about how this service works, some of the limitations, some of the use cases for it and things like that.

And I would absolutely suggest that people read the two posts that you wrote, now they are very, very long, it took me a long time to get through them. But they are excellent, they're really well written. And it reads a lot easier than the documentation, and you give some good examples in there and some good reasoning behind it, which the documentation doesn't always do. So first of all, I want to start with why you wrote this post in the first place because there is a lot of documentation out there. But why did you write these two posts?

Anahit: Yeah, these two very long posts, as you said. So maybe to give some background, I've been working with Kinesis a bit over three years now with one of my customers, who is at the Finnish National Broadcasting Company called YLE. I always bring this example, you can think of it as BBC in Finland, highly respected accompanied with a lot of content and a lot of viewers as well. So our team is responsible for streaming the user interaction data to the cloud. And at the moment, we have something over 0.6 terabytes of data per day. In the moment of writing the first blog, it was half a terabyte, so it's growing constantly.

And yeah, so we did with Kinesis. And when I started like three-plus years ago, I basically had no production experience with it, just like the "Hello, World!" kind of a thing. And most of the things I learned, or most of the things that are in the blog post, I actually learned the hard way, so by making the mistakes, and by seeing the failures, and that kind of things. And I actually wish that blog post, or two blog posts, like that would exist back then when I started, because as you said that there's a lot of documentation on AWS, of course, but for example, in the case of Kinesis and Lambda, you have to read the Kinesis documentation, and then you have to read the Lambda documentation, then you have to marry them together. And it's a lot of reading and not necessarily too clear.

So I wrote this in a short way, I wrote it to myself three years ago, that kind of thing. And I hope it will help others not to make the same mistakes that I had to make myself. So maybe it will help somebody who has already started their Kineses journey or just thinking about it. And the thing is that while I was writing those blog posts, or before working with Kinesis, I have learned so much when I started to dig under the hood of how the service actually works. So I have learned so much about how the AWS services work in general. So like digging deep or understanding deeply, just one service, in my opinion, gives you a wider understanding of all the other services. So even if you're not that interested in using Kinesis, I would still recommend reading my blog post.

And I actually point out some of the common issues or things that are common for other services as well and distributed services in general things like idempotency, and timeouts, and error handling and that kind of stuff. And to tell the truth, I still use or I do use my own blog post as a reference manual, pretty often myself, because I have a horrible memory, especially when it comes to exact numbers. So it's nice to have a one place where I go to look for stuff. And yeah, so to help myself and to help others is the short answer to your question.

Jeremy: Well, no, I think that's great that, first of all, that you did that to help others, but the fact that you did it to help yourself, that is not an uncommon thing. I know, for me, most of the blog posts that I wrote were just ways for me to make sure that I wrote something down, and it would actually live out there that I would be able to go back and reference myself, just like you said. Because I figured out things my own way, and then it's really helpful for me to go back and see how I did it, as opposed to try to find a needle in a haystack somewhere else. So yeah, so awesome. So again, I think that's amazing. And I'll say it again, I read those blog posts, and I learned so much about Kinesis that I thought I already knew, but just seeing it in that different way was really, really helpful to me.

Anahit: Oh, great to hear that, especially from you, because I assume you do know quite a bit about Kinesis already.

Jeremy: I know a little bit. Yeah, no, I've used it quite a bit, but I mean, just in terms of like failure modes and some of these other things and the different caveats you run into, which is something that the documentation doesn't capture as well as it needs to. And that's one thing I find about AWS documentation, but documentation in general is, it's very easy to get to that, "Hello, World!" phase, like you mentioned, but then to get over that hump and bring it into production, I mean, that's a whole other beast.

Anahit: Yeah. And maybe the simplicity of the serverless, and the managed services nowadays is also quite deceiving in that sense, because nobody reads the documentation from start to finish anymore. You just go skim through it, and then it's like, "Okay, I will try out and see how this works." And then you try out.

Jeremy: And you can get going.

Anahit: Yeah, you get going. And I said, "Okay, this thing is working. I know how it works." Yeah, you do until something fails, because it will.

Jeremy: Exactly, exactly. All right, well, so let's start. Let's take a step back, because I know what Kinesis is, you know what Kineses is, but I'm not sure everyone knows exactly what Kinesis is. So let's start there. Why don't you give a quick overview of what is Kinesis, and why would you use it?

Anahit: Yeah, so Kinesis is massively scalable, and fully managed service in AWS, which is meant for streaming data, huge amounts of data, really. And what they say is that it actually scales pretty much endlessly, not unlike Lambda functions. And it has a lot of service integrations, like other services can send events to Kinesis, for example, AWS IoT Core has that functionality, CloudWatch events, events and blogs. Even some more exotic options with database migration service also has some sort of integration with Kinesis. So it's pretty common to use them in that combination.

And then as you mentioned, in the beginning, it's actually a pretty crucial service in AWS itself. And not everybody realizes that, that a lot of services use Kinesis under the hood, like the CloudWatch events themselves, use it under their hood, the logs, IoT services use it and even Kinesis Firehose use Kinesis as their underlying service. And as far as I know, they're one of the biggest customers for the Kinesis team, so it's cross pollination in that sense. And yeah, that outage last November, it actually showed. I would say that many people we don't know about the Kinesis before the power outage I suppose, or not too much, at least.

And yeah, so Cognito failed, CloudWatch failed. And then there was this chain of failures that they experienced for entire day because Kinesis didn't work the way it was supposed to work. So pretty important service, no matter do you use it or not in your everyday life.

Jeremy: Right, right. Yeah. So in terms of what it actually does, you mentioned it's a data streaming service for high volumes of data. And AWS is famous for creating a bunch of services that do very similar things. We've got SQS EventBridge exists now, SNS is a pub-sub type thing, which I guess you could think of Kinesis that way as well. So, I guess maybe why not use SQS or EventBridge or SNS? What specific reasons would you use Kinesis over those?

Anahit: Yeah, that's a really great question. And I think it's a question a lot of people struggle with, especially when they just start their AWS journey or messaging service journey or whatnot. Because there are so many services that look alike, and it's very difficult to distinguish which one of them do you actually need to use and how to actually choose from them. I have a feeling that those services have been converging lately. I think they are becoming even more close together than they used to be. For example, with like SQS an SNS FIFO support that they recently. So those are more similar than they used to be, than back in the day when they added SQS Lambda trigger that wasn't there. So it used to be SQS, SNS, Lambda pattern, and now you can do it directly. So that went to that direction as well.

And now, especially when they added, I think it was before re:Invent this year, or at the re:Invent, I don't remember anymore, they added to SQS, that batch window support exactly the same actually as Kinesis has. So in that sense, they are exactly the same now, so the same amount of, or the same time, or the same amount of records that you can batch before reading them to a Lambda function, which is quite cool. But then they are getting even more closer. And the question is, what would you actually choose?

And I think that the truth of the matter is that in many cases, you can go with many of those services. It wouldn't necessarily be a wrong choice. But there probably is going to be one particular service that is going to be better tuned for your particular use case. And in that case, you basically ... what it comes down to, there are like several questions that you need to ask, for example, the throughput requirements. So how much of throughput are you going to handle? Is it like individual events every now and then? Or is it a stream of events and huge volumes of events? And then again, what's the size of the events? So for example, as SQS can support, or SNS can support two big overheads, it's like 256 kilobytes or something. And with Kinesis it's one megabyte, so that kind of thing.

Then you should think about the data retention requirements, because like some services can store data for a longer time and others can't. Ordering: how do you want to write data to the stream? Do you want to batch the records, do you want to write individual records, do you want to have direct integrations or custom code that writes to the source, or how do you want to consume the record. So do you want to do the pops up, as you said, or do you want to call? What do you want to do? Or do you want to batch this once again, or do you want to be ... So several questions you can go through before deciding it.

And actually, Kinesis in that sense stands separately from all the other services, because it's not even in the same part of the service least in the console. It's considered to be an analytic service, as opposed to application integration service. So you made that distinguished quite a lot. And basically, with Kinesis, as I said, you have virtually limitless scaling possibilities, using the shards, so you can have more shards. And you can scale more, depending on how much data you need to accommodate. And one record can be as much as one megabyte. So it's a huge chunk of data that you can pretty much send to any other service for that matter.

Jeremy: And so, I want to talk about shards, but let me interrupt you for a second. The thing that is interesting about, like you mentioned with SQS, is SQS right now, with FIFO, The first in first out, you can do ordered records. So that's one of the things that Kinesis has always done. I know that one of the big differences, I think, though, is that SQS can really only have one subscriber. Once you take the message off of that queue, it's gone. Whereas we can, excuse me, as with Kinesis, you can actually have multiple subscribers. And as you said, with the data retention, you can go back in time, right? So I think that's another big thing. But it's funny, you mentioned the analytics versus application integration, because I know way back in the beginning, Kinesis was a really great choice for application integration and people were using it almost as like EventBridge essentially, do like eventing and stuff like that, or as a common thing, but of course you had to have multiple subscribers and it was sort of a pain.

Anahit: That's an interesting piece of information. I didn't even know about it actually. Because now the distinguish ... they are trying to make the difference I think bigger now between Kinesis and the other services now. Like the analytic services, they stand separately from the AWS point of view. But of course, it doesn't mean you can't use it. And actually, you can pretty successfully as a messaging service.

Jeremy: Right, yep.

Anahit: And, yes, so I was, you actually mentioned yourself that the big difference with SQS and Kinesis is that you can have multiple consumers for the same stream. But then again, SNS has that as well, and I think EventBridge as well?

Jeremy: Right.

Anahit: But for Kinesis, you can't actually even do the filtering that you can do with SNS and EventBridge, so it's ...

Jeremy: It's also true.

Anahit: You have to send all the events or the same events to ...

Jeremy: Can't they build one service that just does everything for me?

Anahit: Right? That's what I'm thinking. And this message retention is actually pretty funny that you mentioned because they have announced, I think, once again, before re:Invent, this extended message retention. So before, it used to be that you can keep your messages in Kinesis from 24 hours to up to seven days if you need to. And now you can it have up to one year, which makes a database out of it all of a sudden. And I think it will bring all sorts of new use cases with it, because if you just can put your data in this Kinesis and then do whatever you want with it for inside here, in many, many cases, you don't even need to deliver it to any destination after that, it's just fine like that. Of course, you have to pay extra for that, but that's a different conversation.

But yeah, that's a pretty big difference to pretty much any other of the messaging services, because you can't do that with that. And even with ordering, though SQS, and SNS also have ordering. But at least with SQS, the FIFO queues, they have lower throughput than the normal queues. So there is already this limit. And we can use if you don't have that, because ordering comes pretty much out of the box.

I think the main difference for me personally is how they work with Lambda functions, because I think Lambda has a wonderful support, and it's improving every year. And this year, they again added new possibilities or the functionality there. It has a great support for handling Kinesis records or batches and errors, which is always an interesting aspect for me. So, that's a big difference. But of course, like big pink elephant in the room here, is the cost. That's what everybody is concerned about. And I have heard so many times that Kinesis is too expensive. And I think it's still a bit more of an enterprise product rather than smaller company startup thing, because I think mainly because it doesn't have free tier, that's my opinion. Because you just start to pay immediately from the get-go like SQS at least have those three messages per month. And we had this interesting conversation with Yan Cui a while ago who was talking about the sweet spot between SQL and Kinesis, that there is ...

Jeremy: Yes, I remember that.

Anahit: ... actually a point here after which Kinesis actually cost you less than SQS. If you have the big enough amount of incoming data or your data is large enough in its volume, then SQS will start to cost you much, much more than Kinesis not to speak about how difficult it will be to manage really like the consumption and all that things, so. Yeah, but here are few differences for you to consider, but I think each service has its stronger suit. And as you said, we don't have one service that has all of the features that we would like them to have. So every one of them is suited better for a particular use case, I'd say so.

Jeremy: Right. So with Kinesis, another thing, again, that I think separates it very much so from your SQS in your EventBridge is that you do have to set up the shards. So you have to actually provision something in order to send data to so it's not like just an endpoint where you send data and it'll accept as much as you want. So explain shards and then partitions because this is something we could go super deep on this, but FIFO queues and SQS, for example, have a group ID or a message group or whatever that allows you to do sharding there as well, but without provisioning it. But let's keep the conversation focused on Kinesis here. So shards and partition keys, what are those all about?

Anahit: Yeah, so as you said, unlike Kinesis, or unlike SQS, I'm sorry, Kinesis does need to have provisioning. And you can think of a shard as some sort of order queue within the stream. So your Kinesis stream is basically combined of set of these queues, and each queue comes with its own throughput limitations. So you can send 1000 records or one megabyte of data per second to each shard and then on the out, you can get like two megabytes per second. So if you have more data, you basically need to add more shards to your stream and that's the way your stream is going to scale. So of course, each shard is going to cost you, so that's why you have to consider how much shards you are actually adding to your stream.

And the way your data is spread across the shards in the string is by using the partition key that you mentioned. So it's basically just a string that you add to every single data payload that you send to your stream. You just add a separate stream called partition key. And what Kinesis does is it calculates a hash function of that string, and based on that hash function, it decides which shard the record belongs to. So each shard is assigned a range of hash values which don't overlap. So basically, when you send your records to a stream, it ends up in exactly one shard in that stream. So that's the mechanism, it's pretty simple mechanism, but it's pretty powerful as well. And yeah, and the records, as I said, they are ordered inside each of the shards. So you have this ordering out of the box on the shard level.

Jeremy: Right. And then the sharding itself, so if you have five streams set up, the algorithm will split that into ... then, of course, the partition keys have to be different enough, right?

Anahit: Yes.

Jeremy: So that it can actually split them, but you can't send just like one or whatever, just send like a single-digit or something ...

Anahit: No, you can, but you probably shouldn't.

Jeremy: Right. And actually, you could probably control which shard it goes into by doing that. But then if you want to expand, so let's say that you're writing 4995 records per second across five different shards, and you say, "Okay, now I need to add a sixth shard, or a seventh shard," or whatever and keep adding shards, how does that rebalancing work?

Anahit: Yeah, so you can do it by doing so-called resharding, so you can add new shards. And there's actually two ways to add shards. You can split the existing shards as far as I remember, and then you can add a separate shard. So when you split a shard, the partition keys are split between those two shards. It's more or less equally, because the idea is that as you said, you have to have, or it's better to have a random partition key, because in that case, your records will be distributed equally or uniformly across all the shards instead of sending all of them to the first shard and overwhelming the first shard and then the rest will be just idle, and not using the capacity it could have been using. So a random enough distribution of partition keys is very important.

But then if for some reason, for example, you have a shard, which is overwhelmed, so it has more records coming in than the others, then you can split that particular shard and make it into two and then can you just have to take care of spreading the records between those two based on the partition key.

Jeremy: Right, right. And the fact that you have to do that manually, that you have to say, "Okay, this is a hot shard, or my numbers are going up, I have to add something separately." The big question is, is this really serverless?

Anahit: Yeah, that's a big question indeed. And my blog post, I actually argued that it's not entirely. So ...

Jeremy: I have this ongoing thing with Chris Munns at AWS where I think it's serverless, and he thinks it's not but ...

Anahit: Okay, so I'm more on ...

Jeremy: So kind of serverless?

Anahit: ... his side. Yeah, but getting there. In my blog post, I actually compare it to DynamoDB in a sense, in early days, because DynamoDB also started out without auto scaling, without on-demand capacity. So your provision capacity, and not unlike shards, and then you pay for what you provision, even if you don't use it at all. So it's pretty much the same. And then you have to use API calls to add some capacity and remove capacity, but everybody was not too happy about it. But still, it was assumed to be a serverless service, right? DynamoDB ...

Jeremy: Right.

Anahit: ... always from the get-go. So in that sense, Yeah, kind of, but here as well, we have the same fully managed service. But again, we need to take care of the throughput ourselves. So there is no mechanism that would take into account the incoming and outgoing records and decide, "Okay, now I scale up." You can build it. And there is actually a blog post about Kinesis auto scaling, which uses like five other components to do that. So you can automate it, but it's still something not supported by the service itself. Though everybody's holding their breath for it to come any moment now. I actually was hoping it will come at re:Invent, but well, what can you do? I guess the outage was a bigger thing to concentrate on.

Jeremy: That was a bigger thing they had to deal with. Yeah, maybe they were pushing the auto scaling functionality and they broke it, but ...

Anahit: Actually, that's exactly what I thought when they broke it. I was like, "Yes, auto scaling is coming," but then it turned out there were some other issues with that.

Jeremy: Right, right. Yeah. So well, anyways, so alright, so Kinesis though in terms of getting data into it, right, there's a number of different ways to send data into Kinesis. And another thing that's fascinating too, I think just about the Kinesis service in general is Lambda. We think of Lambda because its function as a service as a very serverless service, it sits perfectly in the serverless ecosystem. So whether or not Kinesis is 100%, serverless or not, it is used in a lot of applications, right, applications that have nothing to do with serverless applications or anything like that. Just it is a really good service that powers a lot of things, as we said. So, what are some of the different ways that you can get data into Kinesis? Because you mentioned batching, and some of those other things?

Anahit: Yeah, sure. So, of course, it wouldn't be too useful if we couldn't get data into it, right?

Jeremy: Right.

Anahit: So ...

Jeremy: And quickly.

Anahit: Yeah, that as well. So there's actually many different ways to do that. And, for example, one useful way, if you are going to stream your data from outside the cloud, to the cloud is the Amazon Kinesis agent, which is a standalone application that you run on your server, for example, that can stream files to Kinesis. So for example, if you want to stream your logs from your server to the cloud, that that can be done with the Kinesis agent. So, that's one way.

Then as I said, there are some direct integrations of some services, actually can push events directly to Kinesis, like CloudWatch, and stuff like that. One interesting service of those is, of course, API gateway, because it does require you some work because it basically acts like a proxy over the API calls for Kinesis. So you need to do some VTL magic and stuff. But it comes with some throughput limitations, of course, as with API gateway in general, but it's very useful for many cases.

And then there is tons of community-contributed tools and libraries that you can use to do it, but I think the mainstream, or the most common ways to write data to the stream is actually, either using communities produced from write library, so KPL, in short. And it's basically another level of abstraction above the API calls. And it gives you some extra functionality, but it also runs asynchronously in the background. So you need to have a C++ daemon, right, running on your system all the time. But it will collect the records and send them to Kinesis synchronously, which means that you might have some delay, or latencies that come with it, so it won't push them immediately. But the biggest issue with it is that it's actually only available in Java. So a bit limited use case.

And then the most favorite of mine, because it gives you the most flexibility when you need to, how to write data and how to handle the errors and stuff is the AWS SDK, which is basically the API calls. And luckily, there is a lot of SDK language support. So you don't have to be bound to just Java. Though I don't have anything against Java, I worked with for like, eight, seven years. I don't remember anymore.

Jeremy: Yeah, I'm not a big fan of Java anymore. But I write a lot of Lambda functions, so every time you ... I've never been able to get them to boot up quickly. The cold start has always been horrible with Java. So I've stuck to mostly Node and Python, just to ...

Anahit: Same.

Jeremy: ... keep things simple. But so all right, so you mentioned the Kinesis Producer Library, which I actually remember way back in the day, we had like a Ruby ETL thing, and we were using in the consumer library, and it was a mess, it was just a lot of things that had to happen. So it's easier if you can just have a nice simple SDK, or even better, have the service just natively push it into Kinesis for you, and then have another native service consume off of that, which is super easy. But so there are a lot of use cases with Kinesis. I think people can probably use their imagination for high throughput streaming data, click tracking, ad network type stuff for all kinds of things that you would need to see. Your sensor data, you mentioned IoT integrations and some of those things. So I think that makes a lot of sense. But what are some of the less common use cases? I know you have some ideas around how you can manipulate the system in a way to use it for your benefit, that's not super fast, high streaming data?

Anahit: No, we are mostly using, it's really not with my customer. We do use it mainly for the big data and streaming all that user interaction data, like the classical way of using it. So, that's what we do. And then there is actually one more use case nowadays, you can use it with DynamoDB as the events trace. They edited just again, just recently. So that's another cool thing, because I think DynamoDB string had some extra limitations with Kinesis doesn't, so.

Jeremy: Yep.

Anahit: But yeah, actually, as I mentioned in the beginning, the difference between the service integration services and analytic services, it doesn't necessarily exist, it's more likely in our head, is we don't really have to use Kinesis with the huge amount of data. And one use case that I personally found extremely useful is that when you, for example, have a Lambda function that needs to consume events from some stream or queue, and you want to invoke exactly one Lambda function at all times. So you want to process the events in order or basically consequently, not in parallel.

So with this SQS, what you can do is to use the Lambda reserve concurrency for that purpose. So you can say that, "Okay, I only allow one Lambda execution of this particular Lambda at all times." But it will mean that all the others will be throttled, then you have to take care of SQS visibility timeout and make sure that the retry attempts are big enough, so your valid messages don't end up in a dead letter queue and all that kind of extra worrying, I would even say, that is not necessary.

And what I found very useful is that with Kinesis, the way Lambda works with Kinesis is that it gives you one concurrent Lambda execution per shard. So, if you have a Kinesis stream, which is attached to a Lambda function, there is going to be as many concurrent Lambda executions at any given time as you have shards. So each Lambda will be reading from each dedicated shard.

Jeremy: Right.

Anahit: So basically, if your throughput requirements are okay, and you can have a stream with just one shard where you push all your events, then you have a Lambda consuming from it, then out of the box, you are getting a situation when just one Lambda function is reading from the stream at all times, and you don't have concurrent executions, you don't have to take care or worry about all the throttling and stuff. And then out of the box, you get all this nice functionality for error handling that Kinesis comes with. So I actually love it for that use case and it won't cost you millions, it probably will cost you like couple of hundreds per year. And I think it's pretty much well worth it if you think of all the management costs that you are avoiding that way.

Jeremy: Right. Yeah. And actually, the SQS, the reading off of the queue, the Lambda trigger for that, I believe that you need to set a minimum of five, concurrent ...

Anahit: Yep, yep.

Jeremy: ... for that, because that works that way. But I think and again, I could be wrong about this, because again, how can you possibly know all the services in AWS? But I believe if you use SQS FIFO queues with a single message group ID, that will also only invoke one Lambda function. I'm not 100% sure of that, but yeah, but either way, no matter which service you use to do that, that is a really cool use case. Because I can think of some cool things like, I don't know, maybe you were billing, you were doing like shipping labels, and something where you needed it to be like one after the other, they needed to be sequential, there could be some cool, definitely some cool use cases for that type of stuff.

Anahit: Yeah, we have found it very useful in one of our use cases. And the fun part was that I was struggling with SQS, like, "How do I do this properly? And I don't like this," and like ... then it was like, "Okay, I have been talking about Kinesis for like two years now to everybody around, so why didn't I think about it in the first place?" But yeah, it's a fun way to do that.

Jeremy: Right. All right. So let's move on to consuming data off of the stream. So there are a bunch of different ways to do this. I mentioned the Kinesis Consumer Library, which I think is also Java-based, and you need to run it in. But anyways, the easiest way to consume data off of a Kinesis stream, you've mentioned this, I think most people would agree would just be to use Lambda because it is a really, really cool integration. So what's the Lambda Kinesis story?

Anahit: Yeah, so I've mentioned it several times, because it's really my favorite way of consuming data from Kinesis. You don't need all that extra headache of keeping track of where you are exactly in each shard, and each stream on every given moment of your life. So it's very nice. And I think it takes care of a lot of heavy lifting on your behalf from reading or reading from the stream.

Jeremy: Right.

Anahit: And well, as I said, like, error handling is one thing, one big thing that Lambda makes also much easier for you with Kinesis stream. Then batching, so Lambda can read batches of records from the stream up to 10,000 batches of records in a single batch. So yeah, those are keeping track of ... that's the most important probably, keeping track of where actually where exactly you are in the stream because otherwise, you have to have some external ways to do it. And, for example, can this consumer library uses a DynamoDB table to do that, which it actually spins up behind the scenes without you even probably knowing about it.

Jeremy: And it's provisioned, too.

Anahit: And it's provisioned, and ...

Jeremy: Not on demand.

Anahit: ... it's pretty low. And then one day somebody from your team comes knocking on the door and saying, "Hey, I'm getting this weird DynamoDB provision throughput exceeded errors. We don't have a DynamoDB table." Hmm, where does that one come from? So yeah, it's much easier but then there is, of course, other services that we want actually to mention here is Kinesis Firehose and Kinesis Analytics, because those two are the other services in the Kinesis family. And they both have a very nice integration with Kinesis streams, so they both can be attached to a Kinesis stream as a stream consumer. In case of Kinesis Analytics, it actually can be a string producer as well.

So, Firehose is a service that is used for streaming data to a destination. So if Kinesis streams is just for streaming the data, and then you have to consume it somehow, the entire purpose of Firehose is to deliver data to the destination. So you can connect the two, to stream the data and then deliver it to the destination and the destination can be S3, Redshift, Elasticsearch. And I think that the coolest one recent one is the HTTP endpoint. So basically, you can deliver it anywhere you want. And then the Firehose has also some pretty neat features like batching, and transforming the data, and converting the format.

Jeremy: Yeah, transforming.

Anahit: Converting from like, for example, JSON to Parquet, that's what we use a lot, in our case, compressing the data ...

Jeremy: And then you can query it from Athena, for example.

Anahit: Yep, from Athena spectrum, and all that things. Yes, so it's very, very useful. And you can connect it directly to Kinesis streams, and it is truly serverless because you don't have to provision it. So it scales.

Jeremy: But there are no limits, though, right? I know I can just get a Firehose, like if you're choosing between Kinesis data streams and Kinesis data Firehose, there's an upper limit to the Firehose, right?

Anahit: Yes, that's true. I don't remember the exact limits. But then, the scary thing about Firehose was several years ago, that there was no mentioning anywhere that any of the operations can fail like rising to the stream, or to Firehose can actually fail. Because Kineses had all these metrics with exceeding the throughput, for example. So you see, you have a metric that says that something bad happens, so you know that something bad can happen. With Firehose, they didn't even have the metric that would tell you that, "Hey." So what they had is documentation that says that it's endlessly scaling service or something like that. And then there is the fine print with, "But if you use it with this, and this and this ..."

But as far as I know, if you're using it with Kinesis streams, it actually adjusts to the throughput of the Kinesis stream. So those limits don't apply anymore. So there's ...

Jeremy: Ah, interesting.

Anahit: ... yeah, there's this separation.

Jeremy: Interesting.

Anahit: Yeah. And then the other service from the Kinesis's family was that the Kinesis data analytics, which is one of my favorite ones, really, because it seems small, but you can do a lot of neat things with that. So what you can do is that you can analyze your streaming data in near real time. So you can basically write SQL queries with Kinesis data analytics, and it will perform joins and under different filters that aggregates over some, for example, time-based window. And then it can send the results of that aggregates to either another stream or another Firehose, or actually, it can send it to Lambda, so you can do whatever you want with it. So there's a lot of cool use cases that come with Kinesis analytics. And they both integrate very nicely with stream, you need to stream, but the "Got you," moment here, which apparently not many people realize is that both Firehose and Kinesis analytics, they act as a normal consumer for the stream.

So I mentioned that there is this throughput limit for each charge, right? So there is only one megabyte per second that you can write, and two megabytes per second that you can read. So this in practice means that you can have two consumers reading from each shard at the same time. And Kinesis analytics and Firehose are both considered consumers. So you can if you have a Kinesis analytics application, and the Firehose attached to the same stream, and then you want to add a Lambda function, for example, then you might exceed, end up exceeding that throughput. So you have to be careful about that, so yeah.

Jeremy: Yeah. But so then with that, though, so again, that makes sense. You can have two consumers, but they added something called enhanced fan-out. So how does that come into play?

Anahit: Right. So I enhanced fan-out is funny in the sense that it's very difficult to understand what actually happens by reading the documentation. I think that part took me actually the longest time to figure out because I personally don't use it at work, so it was a research project for me, mostly. I'm trying to figure out what is happening there because like all the combination, like enhanced fan-out and how it works with other features. But what it basically is, is that instead of sharing this two-megabyte throughput, outgoing throughput with all the other consumers, instead, you can have a separate, king of your dedicated elite highway that you get with the stream and then you get your own two megabytes per second of throughput. And you can have up to 20 consumers at the moment, I think, that each of them will get the actual megabytes. So you can basically consume a lot of data with that.

And the nice part is that the latency here is also much lower than with the standard throughput. So I think they claim it's 70 milliseconds of latency versus minimum of 200 milliseconds for the standard throughput, which is a big, big deal. And it actually stays the same in contrast with the standard throughput with where it goes up with each added consumer. So it's a really nice feature. And how partly how they achieve it is by using a HTTP two persistent connection, instead of HTTP. And the consumer, actually, instead of polling, as it does with standard throughput, instead of polling for records, Kinesis actually pushes the records through that persistent connection to the consumer. So in that way, we avoid all the limitations that come to polling the records from the stream, we can't get records API and that kind of thing. So that removes all the headache, but you have to pay for it.

Jeremy: Of course, of course.

Anahit: So, that's the problem. And the thing is that it sounds very cool, and you might think, like, "Why wouldn't you use it all the time?" Well, you have to pay for it. The truth of the matter is, in most cases, you don't need it. So if you have just up to three consumers for your stream, you're probably going to be just fine with a normal shared throughput model.

View Details

About Buddy Brewer

Buddy Brewer is the Field CTO for New Relic in the Americas. In this role, he helps customers get long-term value out of New Relic. Buddy has over 20 years of experience leading engineering and product management teams building tools to help developers and operations professionals deliver better digital experiences. A former entrepreneur in the observability space, Buddy has helped companies across every geography and industry in the world improve their software’s speed, quality, and user experience.

LinkedIn: https://www.linkedin.com/in/bbrewer/
Twitter: @bbrewer
Personal Website: BuddyBrewer.com
New Relic Free Tier: https://newrelic.com/signup
New Relic Explorer: https://newrelic.com/platform/full-stack-observability

Watch this video on YouTube: https://youtu.be/Y4n3fE8g9Ec

This episode is sponsored by New Relic.

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Buddy Brewer. Hey Buddy, thanks for joining me.

Buddy: Hey Jeremy. Thanks for having me.

Jeremy: You are a Field CTO at New Relic so I'd love it if you could tell the listeners a little bit about yourself and what's new with New Relic.

Buddy: Yeah. Been with New Relic for a couple years now and in this Field CTO role I get to spend lots of time with our customers to help them get long-term value out of our observability platform. I'm an engineer by trade. Started my career as a software developer like many of our customers in New Relic. Spent substantially all of my career in product development in various capacities. Engineering, leading engineering teams, product management. And like I said, now I spend most of my time with customers helping them tackle their own observability challenges in their businesses. We're doing a lot right now with New Relic to help people make sense out of the volume of data and to help people pull all of the different types of metrics, events, logs, and traces that go into all this observability into views that they can actually use to help their customers get better experiences in a world where software architectures are just ... They're just becoming more complex by the month.

Jeremy: Right. Well, awesome. First of all, I want to thank New Relic for sponsoring this episode and for the amazing amount of support that they give to us here at Serverless Chats and what we do. So thank you very much for that. Now, you mentioned these tools that you're working on to be able to observe modern applications. And the new tool that was recently launched is the New Relic Explorer. I've looked at this thing. This is absolutely fascinating. It does all kinds of really great things. But I'd love it if you could tell the listeners a little bit more about that product.

Buddy: Yeah. It's part of our full stack observability product in the New Relic One platform. So it's an in-place upgrade that everyone who uses full stack observability today gets. And what it does is it takes all of the information across all of the different dimensions that people are used to seeing in New Relic One, it pulls them together into new views that help people make sense at a macro level of what's going on in the health of their software across all of the dimensions that matter today. So infrastructure, front end, the application logic. All of that stuff in single views. And there's another part of New Relic Explorer that helps people understand in realtime what the key changes are that are happening in a way that requires zero configuration, which is really important to our customers today because the software architectures and the underlying containers and everything that serve those are changing so fast that people just don't have time to manually configure things today like they used to be able to.

Jeremy: Yeah, right. And one of the things too with cloud infrastructures, you've got all this telemetry data coming in from all these different places and most of the time ... I mean, I know at least what I had been doing is using a bunch of different dashboards and basically jumping between different things trying to figure out what's healthy, what's not healthy. And I love these new views that are in the New Relic Explorer because it actually shows you the changing ... If a problem is getting worse and worse and worse, it gives you this growing bubble. So these visualizations are really, really helpful. So I think that's Lookout right? That does that?

Buddy: That's right. Yeah, that's Lookout. The way that I think of Lookout is imagine if you could take something like the Unix diff command and apply it to all of your telemetry data comparing now versus any point in the past. Whereas the Unix diff command is a text console rendering, what Lookout does is it renders all of this in a visual display in a web browser so that you can see ... Like you said, you had these bubbles that really display two dimensions at the same time. The volume of data, whatever it is that you're looking at for a piece of data. A lot of people use this to visualize changes in errors or throughput or latency but it could also be order volume or really any metric that you want. That's the first dimension. And then the second dimension is the magnitude of changes. Right?

Jeremy: Right.

Buddy: What it helps you do is to zero-in, not just on the things that are red ... Because in environments where folks have thousands, or even tens of thousands for some of our enterprise customers, containers running on any given day, the nature of that design and the fault tolerance inherent in that architecture ensures that on any given day there's going to be stuff that's red. Right?

Jeremy: Right.

Buddy: So if a customer calls in about a problem, you log in, you see some things that are red. Well, some of that stuff was red yesterday. What Lookout helps you do is to focus specifically on those things that changed from healthy to not healthy around the same time as a customer-impacting problem. And then you can see all of the different pieces that also correlate to those changes so you could pull it all out and focus just on the things that matter.

Jeremy: Yeah. That's super helpful because, again, it's one of those things where ... I mean, I've worked as an SRE in the past and you get these constant errors sometimes that keep coming up and they're just kind of there. But sometimes it's the severity of the errors. It's how bad was it yesterday versus how bad is it today? Of course, we wouldn't leave a problem that long. But seeing those changes over time and seeing that growing bit of it, I think is just incredibly helpful from that sort of global view standpoint.

And then the other thing that's part of this, which I think is another really cool representation, is the Navigator piece. And this basically uses a red, yellow, and green sort of ... What is it? A hexagonal or an octagon or something like that. But basically shows these little blocks that show you what's healthy and what's not healthy and then you can dive down into each one of those to see more detail.

Buddy: That's right. And what we did was we designed that view to pack an order of magnitude more information density into a screen compared to the view that we had prior to this. Those views continue to be part of the product, but again, this New Relic Explorer is an in-place upgrade that everyone gets that you can use in addition to all the things that you already have with New Relic. But the first piece is that order of magnitude more information density on a screen. The other thing that it does is it summarizes all of this in a way that you can use it as ... You think of it like the new mission control for New Relic. So for your game day dashboard on the major event that you knew was coming and you wanted to be 100% situationally aware about everything going on in your software when that happens. Whether it's a major advertising campaign that you expect to bring a lot of traffic to your site or it's a big calendar event or major event in the news if your media or something like that. That mission control that allows you to see everything in a single view.

And then we have some flexibility where you can customize that view or create multiples of them that align to specific teams. We call those workloads. So you can take the different workloads that are running in your architecture, all if its constituent pieces, from the front end components, the back end components, the infrastructure, all of it, you can visualize in a single view that aligns to specific teams that work on them. So everyone can get their own tailored mission control. And then another piece of all of this ... Because everything that I've talked about so far has been this extremely high altitude look at what's going on in the software, which you need. But as soon as you notice something that requires you to take action, you very quickly need to switch to a lower altitude. So one of the things that we built in to New Relic Navigator is this concept of related entities.

An entity for us is any component in your architecture that makes your software go. Whether it's a Docker container or some other piece of infrastructure or it's a web application that's built in JavaScript or it's back end logic written in Node or Go or PHP or anything. All of those individual discreet components, we call those entities. They all have a health condition, an alert status. They can have events, logs, and traces that are associated with all of them. But one of the things that is a really critical piece of data that we have at New Relic that this release helps to expose is the relationships between all of those. So it's not just a simple linear relationship or even a tree structure. It's a connected graph of all of these different components. So if one thing is red, the thing that you click on might not be the root cause. And the impact of it being red might not be limited to just the pieces that are connecting to that piece. There could be a number of other services that depend on it that are being upstream or downstream impacted.

So every time you click on something we show you all of the upstream and downstream relationships so you can follow it. What you used to have to do is say I'm going to go see what's happening in the application tier, and then you sort out what you can sort out and then you had to pop out of that and then go into a different view and look at all of your front end and just stitch this together in your head. In the worst case, developers responding to problems had to load up five or six different tabs in their web browser and click back and forth between all of these different things. The related entities, New Relic Explorer, the hexagons in Navigator, all of that stuff is designed to help people with that problem so that you can go straight to what the root cause is by just navigating the shortest path through that graph instead of having to pop out and start your search over again.

Jeremy: Right. Yeah. And I wish I only had to open four or five tabs. I mean, you see it's a lot more than that. And you're searching through logs and trying to find that. Now, the other thing that's really cool ... And there are a lot of distributed tracing products out there now and it's very, very cool, where you can go and see how data is moving through different components in your applications, which is really, really helpful. But what's crazy, I think, about this visualization in Navigator and Lookout and everything that New Relic has done, it gives you the ability to connect, like you said, through that graph multiple services that might be sharing things or different data coming from different places. And all of that stuff is instrumented pretty much automatically. I mean, depending on which service you're using. But all of that stuff ... It's not like you have to go in and instrument all of these little tiny bits. This data's just being collected, these traces are being done. And then this really cool service just visualizes all of it for you.

Buddy: That's right. Full stack observability, which is where this new functionality lives. Like I mentioned, this isn't a new product, it's an enhancement to an existing one. Full stack observability exists on top of another part of our platform which we call the Telemetry Data Platform. We've been building that for so many years. We only in last July exposed it as its own product. Priced really simply just on ingest. And we have a free tier by the way that anybody can sign up for and you can ingest 100 gigabytes per month for free with no charge. And one of the great things about the Telemetry Data Platform is it's a high volume, scalable place to put all of your telemetry data, agnostic to whether it's in a metric or an event or a log or a trace. You can just put it all into this single platform. Which is the first step that you have to do if you ever want to have a shot at tearing down all these silos between all the different pieces of information.

So take traces for an example. Having the Telemetry Data platform enables us to do things like show logs in context. Because the logs are in the same data store as all of the trace data. So if you click on a trace, and even if you click on a span inside of the trace, then if you generated any log events that happened just during the context of that span of that trace, we can display it inline. What you used to have to do is you had to go into a different tab in your web browser and start over again and maybe try to use some kind of a trace ID or span ID or something like that that you hope was also indexed in your logging tool and that you could go find it there. Having all of it inside of one data store means that if you're looking at something that's a particular type of data like a trace, you can see other types of related data like logs in the same context and in the same view.

Jeremy: Right. Yeah. And I think an important question would be the simplification of this. Everybody wants things to be simpler and have these really simple views and good ways to represent and visualize their data. But one of the things is that if you were an SRE or you were an ops person in the past, you probably were familiar with all these different tools that you were using and you knew exactly what you needed to see and how different things ran and stuff like that. But it's not just ops people or SREs or people who are just always worried about the infrastructure that are impacted now by a lot of these changes because I think we've made a big shift in the way that we develop applications and who's responsible for the lifecycle of those applications. I'd love to talk about that evolving role. I guess we would call them modern developers maybe? That modern developers building for the cloud and building these complex systems. What kinds of responsibilities do you think have shifted to them?

Buddy: It's interesting how the nature of the role of software development has changed. I started my career 100% front end developer. And specifically building tools for front end developers to reason about the health of the front end experience. And then you had back end developers. And you could meet someone at these networking events and talk about which parts of the stack that you work on. But those lines are fading. There was a report that came out last year. I think it UBS. Compared how many developers identify as different types of developers, front end engineer, back end engineer, et cetera. And the thing that was remarkable about it is specifically people who identified as a full stack engineer, 55% of the respondents identified as full stack engineers. So more than half.

Now, in 2015, five years ago, it was only 29%. So it's the majority and also the fastest growing cohort of engineering role. And that's how you end up with the situation we were talking about before where you've got so many tabs open in your browser is because all of these tools have been built for specific slices of the application architecture. Logging tools, front end analysis tools, back end analysis tools, infrastructure analysis tools. All of that. Full stack observability, the product that we offer with New Relic aims at being a full stack analysis tool. So again, in that single tab you can see the relationships between all of these different tiers. And we did that specifically in response to what we saw as this broader trend, both from the analysts but also talking to our own customers and realizing ... New Relic's been in this business now for 13 years. Started in 2008. So we've seen a lot of this evolution firsthand among our customers and they were asking us for this. They wanted simpler views that connected all of these different pieces together because increasingly what happens, somebody gets a notification that there's a problem that they have to solve and it's not just in a slice of the architecture.

They're on a team that is designed to do everything that it takes to deliver a particular part of the customer experience. So if something goes wrong with that experience, whether it's in the infrastructure tier or the application tier or the front end tier, they're accountable to finding it and fixing it. So we're building tools to help people do that better.

Jeremy: Yeah. And I think that's interesting in terms of that evolution where even when ... Let's say AWS started with EC2s and things like that, the virtual machines, back in 2008, 2007, somewhere around there. You started building applications that way and I think you had very traditional ops people setting up the networking for people and setting up an EC2 and I don't think a lot of people were doing CI/CD, at least not like they are now. So you take that code that a developer would write and someone would set up that instance or that environment for you to dump the code into. And then as we moved towards things like containers, developers are now responsible for packaging their own containers and requiring the resources they need or the packages they need, things like that. And then moving even further down the line to serverless where in most cases you don't even have an ops person involved right?

Buddy: Yeah.

Jeremy: I mean, there's nothing for them to set up sometimes. So that change in how we're building applications ... Do you think that that change of how developers are getting closer and closer to the infrastructure, that that's sort of one of those things that prompts a need for this full stack observability?

Buddy: Oh yeah. Yeah, for sure. And it's affecting everyone. It has crossed the chasm. This is not just cloud native startups that are adopting this. Substantially every large enterprise that I talk to, and I speak to usually multiple per week, are somewhere along this journey of cloud migration. And that includes shifting workloads from data centers or traditional monoliths decomposing into microservices, moving from data centers into cloud. Orchestrating all of this with containers on Kubernetes. And increasingly across the board, cloud native and traditional enterprise, like you said, moving to serverless because there's certain economies and efficiencies that you get out of being able to take advantage of that layer of abstraction that companies of all sizes and across all industries ... Not just gaming and super high tech and media and commerce but also financial services and travel. Just everybody is moving toward this and adopting it and they're looking for tools that can help reason about the connections between all of these different pieces that they're now responsible for.

There's another component to this that we have been and continue to work hard on at New Relic, which was a point that you touched on earlier. This notion of when you have so many components that you have to manage and all of these things are moving are changing so rapidly, you don't have time to undertake these expensive manual tasks to create all of this instrumentation. So we've been for over a year now progressively opening up our platform to accept other types of third-party data, not just our own agent technology that we've building since 2008, but things like Prometheus and Open Telemetry. You can just point exporters at our endpoints and make it easier to get that data on board. Taking all of our instrumentation logic and making it easy to wrap that in automatable frameworks like things like Terraform scripts and stuff so you could actually build observability in as code and deploy it at scale.

We've been talking about New Relic Explorer which is our release that we're talking about today that's all about the visualization. But it's enabled by a tremendous amount of work that we've done and continues to be underway to simplify the instrumentation. Because as the architectures themselves become more complex, it obviously gets harder and harder to keep up. Frankly, a lot of our customers have issues where they can't instrument fast enough to keep up with the change that's happening in their infrastructure. So as a result they have all of these dark areas of their application that are critical to delivering the experiences to their customers, but they don't have observability into what's going on inside of it. So we've been doing a lot of work to give people tools to get leverage on that problem too so that they can add instrumentation to all the pieces that matter.

Jeremy: Right. And I know that New Relic has done a ton of work on instrumenting Lambda functions for example. Like being able to instrument these things where you can't necessarily run the agents. And I know there's been a lot of really cool innovations in the serverless space around some of that stuff. But I'm curious just from a developer perspective, and maybe you have some experience here of seeing some of your customers do this, how much is your average developer who's maybe building a cloud application working on a team, how much is that developer actually going in and using these observability tools to see what's going on? Is that something where they need to be heavily involved in that or are you still seeing a good separation between the ops team in that regard?

Buddy: It's evolving. We're seeing it change in a couple of dimensions. Developers use observability data and monitoring tools far more than they did years ago. Although, there's always been a segment who needed to do that. New Relic, one of the things that we're known for as a company is being the monitoring platform that is the most developer-friendly. Our CEO, Lou, is a programmer at heart who still writes code on the weekends even as the CEO of a public company. It's part of our DNA. And I think we've always had a natural affinity through that to the types of adopters who are developers, who both write the code and they deploy the code. What we're seeing is, and as our company has grown, that cohort of developers who are responsible for both writing the code and deploying and managing it are exploding in size and scale and the number of those people that are out there. So we've just been riding that wave, if you will, of developers who continue to be responsible for how all of that stuff actually materializes. And it's now becoming essentially the standard way of operating, like I mentioned before, not just for cloud native startups but at large enterprises as well.

And that blur between ops and dev is fading to the point that it's really difficult to see as we sit here in 2021. That's on the role side. One of the other things that's changing, I think, that's really interesting about just the way people are using telemetry data is the set of use cases that it's relevant to. Historically all of this data about what's happening in your software, the canonical use case for when you need that is when something's on fire. Right?

Jeremy: Right.

Buddy: The mean time to resolution. How fast can I get a problem solved? How quickly can I take something that's red and turn it green? It's the classic use case for New Relic and for any other tool in this space. But what we're seeing that's evolving is people are using this telemetry data outside of that context more and more frequently as part of their day-to-day software development. For example, how do you choose where to tune and target your reduction of technical debt? If we can present observability to you that helps you understand not just the parts that are the slowest ... Because sometimes things are slow but they're asynchronous, they don't matter, or whatever. But what are the parts that are the slowest that are actually impacting customer experiences in a way that damages your business or damages your brand? So we're seeing people use that telemetry data. Nothing's on fire. But they want to use it in order to better plan and prioritize their development work.

Or another example is we've seen a lot of development in this field of chaos engineering. And testing resiliency not just by looking at the data and evaluating the architecture or doing things like load tests, but actually intentionally breaking things and then seeing how the system reacts in response to that. A tool like New Relic Navigator is really good for being able to see ... Or actually Lookout might even be the best of the features that we're releasing now that help people with this. Where you can spot these changes that maybe you didn't anticipate so you didn't set up threshold alerting or something on. But you can in realtime see how all of the different pieces of your application change when you go in and you test the resilience of your system by breaking it.

So we're seeing continued convergence of the roles, which brings more and more folks into looking at observability data. But we're also a widening of the number of use cases beyond just the traditional fire fighting. It's a cliché, but it's true. As more and more businesses essentially become digital businesses, the data that describes how your digital experiences and working are taking on more and more strategic value to those companies, at least the forward-thinking ones. And so they're looking for ways to leverage that data and new and creative ways beyond traditional fire fighting.

Jeremy: Yeah. I want to talk to you about resiliency and a little bit about chaos engineering, but before we move on from this role, I'm really curious, you've been doing this for quite some time and as you see this evolve, where's the line for developers? How far should we push them down this getting into the ops role? I know we said it's very blurred, but is it at setting up their own automation? Is it at doing networking or actually touching infrastructure? How far do you think a modern developer really needs to go down that path?

Buddy: Well, it's different for different organizations and there's not a single pattern that you can apply. It's a complicated enough problem that everybody kind of needs to tailor their approach to fit the dynamics of the environment that they operate in. Sometimes things like regulatory compliance come into play and all the rest of that. But I think probably at the highest altitude, the broad trend is we're seeing developers become accountable by default to all of it until they reach the point where they can trade off the management to a third party like a cloud provider. So for example, the networking and things like that. If you can trade that off to your cloud provider, but when it comes to defining all of the infrastructure let's do that in code in an immutable way so that I can automate and deploy and do all of that and handle it as a developer. So developer and operations, we've been saying this for over 10 years now, but it's gotten to the point now as I sit here in 2021 where companies of all sizes and across all industries ... Increasingly I go in and I talk to people and you used to kind of ... In the prep, it's like is this going to be a meeting with the development group or is this the operations group?

We just don't talk about it that way anymore. People are accountable to all of it. There might be specialties within the teams where maybe somebody has a time bias. They spend a little bit more of their time in one category versus the other. But the fact is that most people I talk to today, they've got some level of accountability across all of that stuff. Which is one of the reasons why ... There's only so much that a human being can do at any given time.

Jeremy: Right. Exactly.

Buddy: So the way that a lot of people are gaining leverage on that is by trading some of those pieces off to third parties like the cloud providers.

Jeremy: Yeah. No, I think that makes a ton of sense. Another thing I think ... You mentioned something about complexity in there. And one of the things we're seeing quite a bit of now, which is a very popular way to develop software, is to go down the microservices route and get rid of those old monoliths. So as people are building more microservices and you have multiple teams, which means they probably don't always follow the same standards and some might be written in different languages, some might be running in different environments, the complexity of the data that's coming from that and all of that information, trying to organize all of it. This is just one of those things where something like New Relic I think captures ... It kind of captures that perfectly right? Where it's like you've got all this chaos and you try to make sense of it. So just your thoughts on microservices and the role of some of these tools now to make sense of all that data.

Buddy: Yeah. It's been, I think, really great for engineering teams who've moved to these architectures that it allows them to decouple things and move faster. As more and more of businesses move their revenue toward digital they necessarily have to scale up their headcount of people in engineering which means you now have an organizational problem of how do you keep all of these people in this increasingly large organization productive without creating so many dependencies on each other that they all grind to a halt. So microservices are really great for that. Of course, it's also true that the magnitude of data and complexity that engineers are responsible for and accountable to is scaling at a rate that's faster than headcount. So engineers are having to take on more today and they're having to do more with less. But microservices help them at least manage the dependencies versus the old monolith architecture. This is the reason why organizations are moving away from monoliths and toward microservices today.

It does, like you said, create a whole new set of problems for observability platforms like New Relic One to solve for. The relationships, there's orders of magnitude more of these relationships, orders of magnitude more components. All of that data has to be tracked. It all has to be managed and visualized in a way that allows folks to look at all of this at multiple altitudes so that you can see what's happening overall. The mission control kind of thing that we were talking about earlier. But it can't stop there. You have to be able to get very quickly to the individual metrics, events, the logs, the traces, and all of that stuff that are happening right around where a problem comes up. For New Relic, it's required us, in order to continue to serve our customers in the face of all this change to change almost everything about our platform. We went from having probably a dozen different discrete products ... We had a real user monitoring product, a synthetic monitoring product, a mobile app monitoring, APM, infrastructure, logs. We had all of these different discrete products. In order to keep pace with all this and continue to serve our customers, like we talked about earlier with the evolving role of the engineer toward full stack responsibilities, to bring all of that together into a single product.

That change happened because of changes that are happening in the organizations that we serve and in the broader application architectures. Another massive change that we made is we decoupled ... Well, we stopped counting hosts for one thing. We used to price all of this by units that just don't really make sense anymore in the modern era. How do you count up how many hosts that you've got in a world where it's going to be different an hour from now? So we stopped. We switched to ... You do know how many engineers you have. So full stack observability is priced by the seat. And then we decoupled the data. Because, like I said a minute ago, the amount of data that organizations are having to manage is scaling at a rate that's much faster than their headcount. So we took the data and we actually carved that out as a separate thing in our Telemetry Data Platform. Priced it very aggressively. 25 cents a gigabyte and then we give people 100 gigabytes a month for free if they sign up for our free tier. So that you can, in an economically feasible way, track all of this data across all of these different services.

That goes back to the point that I was making a few minutes ago about the big problem that a lot of organizations have is that they just don't have observability across all of their application. There's a couple of reasons for that. One, if the instrumentation is too complex they can instrument fast enough to keep up and so we're working on that. We've done a number of things to make things simpler and we continue to make investments there. The other is sometimes it's just not economically feasible. So we built our whole pricing model and packaging around making it actually feasible for people to be able to generate all this telemetry. Now, of course, once it shows up in the database, it's incumbent on us to help our customers make sense of all of that, hence things like New Relic Explorer which we're talking about today.

Jeremy: Yeah. And that's one of the things I was going to say. Abstraction is very hard. When you're trying to abstract anything it's hard to find the right level. So how do you approach all of this data without oversimplifying it?

Buddy: Yeah. It's hard. We do some of the things that you would expect. We have, and we've always had, curated views that are informed by our 13 years of experience helping thousands of companies manage their own data. We also happen to be a provider of software at fairly large scale. So we have a lot of experience living in the same problem space obviously as our customers do. So we work hard to give people out-of-the-box views that help them understand what's going on in their software. And then of course we've got the ability for you to create custom dashboards like you would expect. We have a query language that allows you to interact directly with the high cardinality events that we store on behalf of our customers. Not everyone can do this, but every organization that we work with usually has a small number or sometimes a lot, but usually at least a few power users who understand the query language and can get in there with a scalpel and pull exactly what they need out.

The thing that we do that I think is unique to New Relic, though, among observability platforms, we also added a programmability layer about a year and a half ago. And what that allows you to do is to move beyond just dragging and dropping widgets from the palate to create custom dashboards and it moves beyond query languages and working with raw data toward the ability to actually write your own code in ReactJS, interact with our data model using GraphQL. So standards that lots of people know. And you can build truly tailored bespoke visualizations. So we've had customers do everything from combine operational data with weather data, geographic data for people who have physical points of presence in stores and things like that. You can build your own. We also have an app catalog and an ecosystem where you can go in and you can install things.

So that's how we manage it at New Relic. We try to bring a point of view. Every company in the world who has this problem, and it's a common problem, it's a balancing act that you will never be done with. You're always working on it. But we try really hard to provide people without a box curated views that allow them to be immediately productive. But at the same time affording the flexibility, not just through custom dashboards and things like that, but actually a platform you can build applications on top of so that you can visualize any sort of way that you want but not make that ecosystem so convoluted to navigate and all of that stuff that it's impossible to find the pieces that solve 80% of the problem. So you can imagine if we took everything that anybody ever did custom and we made all of those first-class objects in the system, it would be such a huge haystack that you wouldn't be able to find the pieces to solve 80% of the problem. So we promote that to people as part of our out-of-box experience. But then we give you the flexibility if you want to, and many of our long-time customers have adopted this, to create truly custom applications to see the data exactly how you need to see it.

Jeremy: Right. Yeah. And I think when you have all this data coming in and you're collecting metrics and logs and traces, that's great to be able to look at all of that stuff independently but you just want to get to that root cause analysis. You want to be able to figure out what that root cause was and be able to jump in. So having those predefined views, I always find those to be helpful because if you just gave me like, "Hey, here's all the data. Just set up the alerts and the graphs and everything that you want," that usually doesn't get you very far until you can spend days and days and days digging into that. So having that top-level stuff and letting you dig in, I think, is a really good way to approach it.

Buddy: Yeah. Like I said, we try to bring that point of view. A little bit of a sidebar from our core discussion but for those in your audience who are interested in sort of historical trivia you may recall that New Relic got its start in 2008 building ... Our founder, Lou, built an APM product on top of Ruby on Rails, which was setting the world on fire back in 2008. Twitter was based on Rails. Some very major apps were based on Rails. One of the things that was a very defining characteristic of Rails was that it had a very strong point of view. You do not put your controllers in that directory, you put your controllers in this directory.

I think some of New Relic's early design intent, given the fact that it was built from within that Rails community, was to start with a point of view so that people could get productive as quickly as possible. It's one of the things people loved about Rails was you got that really fast out-of-box experience. So over time, as our business has grown ... And of course, now New Relic does way more than Ruby on Rails, although you still can monitor your Rails application with New Relic. You have all of this additional flexibility. But where the company got its start and one of the things that we were known for really early on was that developer productivity that came from bringing a point of view of once you get the ... Drop in the instrumentation and you log in and you immediately have insights. So we try to hold onto that even as we give people more tools to create all these custom applications and everything on top.

Jeremy: Yeah. And I think an opinionated approach to certain things with some flexibility, sometimes it can steer you in the wrong direction but you're right, it just gets people productive so much faster.

All right. I want to go back to the resiliency and some of that chaos engineering stuff because it seems like Lookout and Navigator, these are those perfect tools like you said for doing those chaos days or things like that. So what are some of your thoughts on building resiliency into these systems and how can New Relic Lookout and the other services underneath that, how can that help you make sure that you're building resilient systems?

Buddy: Yeah. A lot of what goes into New Relic Navigator and Lookout is having the ability to see in realtime what's changing in your application. So like I said, many of our early adopters for example, when we first started opening this up to a small set of customers before we reached our general availability launch these features, in addition to using the new features for the production events that were coming from real customer traffic and things like that, they were also using it as a way to reason about what was happening in their software architecture when they're intentionally making changes. That was another one of the core use cases that we saw people using this for. And in particular with Lookout, because it's zero config, it's really good at helping people reason about these unknown unknowns in software which is one of the differentiating characteristics that gave rise to this notion of observability in the first place.

A lot of the distinctions that people draw between monitoring and observability is that monitoring was defined by this characteristic of, "I know all of the failure modes, I'm going to instrument them all with threshold-based alerting, and then I want to get a page when something breaks." And in observability, one of the defining characteristics of it was in contrast to the way that people used to do things. More and more of the way that software fails today, oftentimes it fails in a unique fashion because of all of the ... There are more variables in the equation anymore than you can count because of microservices and all that stuff. So you have to have a model that allows you to see what is changing and not just where the changes are but actually direct you toward the ones that are causing a customer impact without relying on you having analyzed all of the possible failure modes in advance.

So since Lookout isn't reliant on prior configuration or thresholds, it's just looking for changes and then correlating all of those changes to each other so you can see where the clusters are. Which is a lot of what the problem space that people are solving in AI ops for example. This is an exploratory realtime versus a lot of the AI ops is about sending you notifications, which we also do. But this is about seeing it when you're actually logged in and exploring what's happening in your software. Makes it highly useful for those situations when you're doing chaos engineering. And it's something that we're seeing. We're far from a state where everybody's doing that today. But again, we're seeing a lot of growth and increasingly companies who are solving for those types of use cases and it was one of the things that we designed New Relic Lookout to help people do.

Jeremy: Yeah. Well, I think if you are at the point where you need to start doing chaos engineering, you probably have a lot of applications and a lot of services talking to one another. And I think just convincing some team, "Hey, by the way, we are going to break something in production to test it," if you're going to do that and you can actually convince some team members to let you do that, you better have a pretty good tool that's going to be able to capture and be able to observe what is actually breaking. And especially even if you have to revert quickly, at least be able to see the history of that and be able to go in and see okay, when this broke this particular service was no longer responding or something like that. So, are there any surprising things you found as people started adopting this stuff?

Buddy: Yeah. One of the things that I thought was most interesting when I was going through our feedback from our early access program ... We designed this for engineers to use. For people who are in the work every single day, to help them do their jobs better. It's common for us to see managers and directors and executives engage in our telemetry data but it's almost always rolled up to a summary that people are using to track things like SLOs and SLAs and maybe correlate that to some sort of a business outcome like conversion rate for a commerce company or ad impressions for media or something like that. Things that are at a higher altitude. One of the surprising things that we saw with the early access program for New Relic Explorer was we saw a use case where a manager who historically was unable because of all of the stuff that we've talked about so far, all of the complexity and everything, it's impossible for them to reason about it and do all of their other responsibilities as a manager.

So their job typically was air traffic control to get managerial leverage on a larger problem. So it's like, "Here's something going on over here. I'm going to send this to the person on my team who's responsible for it, ask them to look into it as part of their job as day-to-day manager." What we found in the early access program ... One of our use cases in specific that comes to mind was someone who hadn't actually rolled up their sleeves and done the root cause analysis in quite a while. Because the complexity required and everything, just didn't have time to do it. And he discovered an issue using New Relic Explorer. But before sending it on to the person on their team responsible for it, they went ahead and clicked in and said, "Let me just see if I can figure out what's going wrong here."

And for the first time in a long time they were actually able to perform root cause analysis on the thing and send it directly to the engineer outside of their team who was responsible for doing the work to actually file the ticket and fix it and all that stuff without having to task it out to an individual on their team to do the investigation. So it probably saved them, what? A day? Two days maybe? At least a day.

Jeremy: That's a lot of time.

Buddy: Yeah. So that was something that we didn't necessarily expect was that it was going to unlock the ability of folks who don't ordinarily do root cause analysis and detailed work to be able to actually navigate to what the root cause was in a way that they'd never been able to do before. It was actually one of the more, I think for me personally, hugely validating points that we had achieved what we had set out to in terms of making an interface that people could use efficiently. When not only the people in your target audience, but also people who weren't necessarily in your target audience were still able to diagnose a problem in realtime because the connections were there in the right place. I mean, the data's always been there. Collecting data's not hard. What's hard is making it all accessible at scale and connecting all of it and delivering insights. Not just piling a bunch of data into a data lake somewhere or something like that.

So when we got that story back that someone was able to, who doesn't ordinarily do this day-to-day, actually get in and diagnose a problem, it was surprising and it was also really validating for us and for the team.

Jeremy: Yeah. And I think that's amazing. No matter what level you are at, whether you're a developer or you're a manager or you're somewhere in between, not only do you reduce mean time to recovery and you can find those problems faster and figure out what the issue is, but that saves a lot of time. I can't tell you how many times I spent days looking through logs and all kinds of things trying to figure out exactly why every 100th time this thing runs something goes wrong. And being able to go and trace that and find that information quickly saves you time, saves you money, saves you mental anguish, I would think, for a lot of these things. So that's pretty cool.

Buddy: Yeah. Sure. We thought so.

Jeremy: Awesome. All right. Well, listen Buddy, I really appreciate you being here and sharing all this stuff about the New Relic Explorer and the New Relic One platform. So if people want to find out more about you, ask you some questions maybe, or they want to find out or sign up for New Relic One and use this new New Relic Explorer, how do they do that?

Buddy: Yeah. Well, for me personally, I'm most active these days on LinkedIn of the social platforms so you can find me there. Just Buddy Brewer. I'll be the one that pops up working at New Relic. And for New Relic, like I mentioned earlier, we have a free tier that is really easy and really the best place for someone to get started who's had no exposure to New Relic. You just go on our website. It's up at the top right. Click on sign up. And what you'll get is 100 gigabytes a month of ingest that you can put into the New Relic platform and one seat license for all of this stuff that we just talked about today. So you can actually ingest 100 gig of your own data every month and just go use New Relic Explorer and all the other parts of full stack observability.

Jeremy: Awesome. And you can find that at newrelic.com. Thanks again, Buddy.

Buddy: Thanks, Jeremy.

View Details

About Sarjeel Yusuf

Engineer turned product manager, Sarjeel Yusuf is greatly interested in how the move to cloud computing and the rise of DevOps is revolutionizing the way we manage and release our software systems. Ex Thundra, and currently at Atlassian, Sarjeel is focused on bringing DevOps enabling solutions from the perspective of incident investigation and resolution in Opsgenie. By leveraging his past experience in Serverless monitoring and debugging at Thundra, he believes that there is a great opportunity in how serverless can unlock the potential of DevOps teams.

In his free time, Sarjeel loves to write about new advancements in the fields of serverless, DevOps, and more recently, product management strategies. His writings can be found on his personal medium account as well as other publications. He would love to get in touch with anyone who would love to brainstorm ideas in pushing existing technologies to build amazing products.

  • Twitter: @SarjeelY
  • Linkedin: https://www.linkedin.com/in/syedsarj/
  • Website: sarjeelyusuf.me
  • Opsgenie: https://www.atlassian.com/software/opsgenie

Watch this video on YouTube: https://youtu.be/T7eUUUBRZQQ

This episode is sponsored by Epsagon.

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm joined by Sarjeel Yusuf. Hey, Sarjeel, thanks for joining me.

Sarjeel: Hey, Jeremy, thank you so much for having me. I just want to say it's pretty exciting to be here. I've been watching the show for quite a while now, and it's just exciting to be here with you and talk about everything serverless, I guess.

Jeremy: I'm excited to have you here. So, just to introduce yourself. So, you are a product manager at Atlassian. So, I'd love it if you could tell the listeners a little bit about your background and what you do at Atlassian.

Sarjeel: Sure. So, yeah, as you've mentioned, I'm a product manager at Atlassian. Actually, a very new product manager. Just a year ago, I was a software developer within Atlassian, within Opsgenie, and now I'm a product manager at Opsgenie. So, I made the switch to product management very recently, actually.

And so, for those who don't know what Opsgenie is, Opsgenie is basically an on-call incident management tool. It allows you to route your alerts to the right person, make sure that everybody is aware of incidents that may occur. And it helps you all the way from incident awareness to incident investigation and retribution. And my specific role at Opsgenie is basically helping DevOps practicing teams to better their entire DevOps flow, especially considering incident management in the DevOps pipeline.

Jeremy: Right. So, that's actually what I want to talk to you about today, is just about DevOps. It's such an interesting discipline. And as teams sort of evolve and start using the cloud, it's almost like it's sort of necessary, I think, in order for you to adopt some sort of a DevOps culture.

And working at Atlassian, obviously, Atlassian has Jira, and Opsgenie, and all these other services that help with software development, and the software development lifecycle and things like that. But I think there's a major confusion out there about what exactly we mean by DevOps. And especially when you see companies labeling tools as like, "Hey, here's a DevOps tool." Or you've got DevOps engineers and things like that, that just seems really weird to me, because I don't think of DevOps that way. And maybe we could start there and sort of just set a baseline for the listeners here, and have you explain what exactly is DevOps, and what do we sort of mean by as a practice or as a culture as opposed to a set of tools or engineers?

Sarjeel: Yes. Yeah, that's it, right? DevOps right now, the reality that DevOps has ... The word DevOps has become a buzzword. Actually, quite interestingly, I think it was yesterday or a few days ago, I saw a tweet by Patrick Debois who was saying that just because ... It goes along the line of something like this. Just because an idea has become a buzzword doesn't mean that you should shy away from it. You should still go into it and explore what it is, and you learn from it.

That's the problem right now. The industry has been capitalizing on DevOps. Especially a lot of new startups are capitalizing on DevOps, marketing themselves as DevOps tool. So much so that the promise of DevOps is kind of lost or not fulfilled when you have all of these DevOps tools or DevOps engineers or DevOps certifications coming up in the industry.

Let's try to understand what exactly DevOps is. I think the best person who explains this or who captured this is Jez Humble. He basically describes DevOps as a set of practices, a cultural mindset, not exactly a set of tools. Yes, you can have tools to help with your DevOps practices. I'm not saying that, "Oh, any tool that says is associated with DevOps, that's definitely a lie." No, it's not like that.

So, you can have tools to help with your DevOps practices, your DevOps culture. Harboring that culture in your company or in your team. But at the end of the day, it comes down to how you and your team and your entire organization are going from the ideation phase all the way to the release to production and then maintaining of your product. For example, that's where we, at Opsgenie, operate incident management. How you maintain your product, and then how you learn from that and then go through that loop again.

So, traditionally, what we saw was that we had all these separate teams where you had different roles associated to a separate state in your development flow. For example, you had ideation. The first one would be ideation where you would see more involvement of product managers and designers and sometimes engineering managers. I'm just talking very generally. You would have build, you would have tests, release, monitoring, incident management, feedback. All of these were siloed.

And the problem became that when your product, when your software would go from one stage to another stage, when those involved in one stage would throw it over the wall to those involved in the next stage, the people receiving it in the next stage, there was some communication gap. And what that resulted in was that things just went slower, especially when you would scale your product, and especially when things would go wrong. That's what we see as an incident management tool.

Especially for our customers, when our customers are using Opsgenie and the responders are not necessarily the people who were responsible for building the code, it takes them longer to resolve the incident. That's expected. You are trying to resolve something that you didn't build, that you don't know the nitty gritty details about, and you're trying to find what went wrong. That's what DevOps aims to solve. So, I would say that with DevOps, what you can achieve is that you can go faster. You can increase your velocity while maintaining stability. That's the entire promise of DevOps.

Jeremy: Yeah. I like, basically, that quote of just because it's a buzzword doesn't mean you don't need it. And I feel like the same thing has happened with serverless as well, where everybody just starts slapping the term serverless on their product, or say we do something with serverless. I think it just confuses things more and more. And so, when you say things like DevOps, we need a DevOps tool, or we need a DevOps engineer, it sort of perverts the underlying principles, I guess, of what you're trying to achieve. And so, maybe let's go there for a second. From a principle standpoint or a cultural philosophy, as you had said, what are sort of the main objectives here? What are we trying to achieve with DevOps? Because you mentioned this idea of throwing it over the wall. And that happened all the time, right? I wrote some code, I give it to my ops team. My ops team tries to put it into production. And I'm going back a way. I know you're actually much younger than I am, so good for you. But that, actually, I think, is good, because it gives you a fresh perspective on seeing how things should be working, as opposed to old people like me saying to ourselves like, "Well, we used to do it this way. So, maybe we should keep doing it this way."

So, that idea of throwing things over the wall and having something not work, and then having to just kind of kick it back as opposed to just have a flow that this whole thing gets taken care of. So, what are sort of those principles that sort of enable you to break down those walls or break down those silos and just kind of have your software flow all the way from ideation through to production, and deployment, and then to even monitoring, and troubleshooting, and incident response?

Sarjeel: Yeah, that's actually a very good question. What exactly is the solution? If we say that, "Okay, all those tools that are coming out, or all the certifications that are coming out isn't exactly the solution." Then what can we do to break down those silos? I believe that there are two things that we can do. One is to try to involve everybody across that stream in mostly all the stages. Even as a product manager, I try my best to get involved in all the stages. And then also, even within the ideation phase, get the technical side involved within the ideation phase.

So, it's not only a product PM group only, like get everybody on the same table and understand how we can go from ideation to production. And that is one culture, that is one practice that you really need to incorporate in your team. Stop thinking about people as just fixed roles and allow more flexibility and allow the flow of ideas more. That one way is how we can really break down the silos.

Another way is that, "Okay, now that you have everybody involved in everything." The responsibility of the groups. I mean, it's, it's almost impractical to have a single or a group of engineers building everything and also making sure everything runs and also maintaining the systems and getting everything deployed while ensuring its stability. It becomes very difficult. If we still look at traditional practices, having one team do everything would become very difficult. So this is where I believe automation comes in, and automation is key.

Also, while we're talking about automation, we should also try to think of this left shift culture. Bringing everything closer to either the development team or the ops team, or whoever else, but basically bringing it closer to the build stage. Right now, we are seeing this trend. A lot of people, and including I, would say that CI/CD is kind of the backbone of DevOps, because CI/CD is now looking at a lot of automation. And we see a lot of automation features coming up over there. When you're looking at automation, you're also looking at incident resolution. You think that entire incident resolution that would sit over here, coming closer to your CI/CD. And eventually, we're also seeing CI/CD tests, and all the automated tests coming closer to the developers themselves. You see debugging and having all these integrations in the IDE. Being able to locally test your cloud apps and things like that.

Yeah. It's pretty great. We are seeing a left shift, we are seeing an increase in automation. So, it's not only a buzzword, but even though it is perceived that way, but the reality that we are seeing, these improvements happen. And we are seeing an increase in DevOps practices and successful practices, actually.

Jeremy: Yeah. I think automation is a good point, because that's one of those things where sort of like automate all the things. It sounds really, really good. But then it also scares people too. A lot of ops people say, "Wait a minute, if you automate away my job, then what am I supposed to do?" And the answer to that is there's a million more things that you can do, especially around security, around speeding up the pipeline. Again, minimizing your time to recovery, or just things that you can work on. But the idea of automation is a key principle, I think, in DevOps, because it just gets ... It's the idea of getting things from somebody's IDE into production as quickly as possible. And then being able to sort of understand how that change maybe impacted the overall system or whatever, and be able to resolve those things much more quickly.

I remember the days where we used to work for months on a software release, and then we would put the software release out there, and then 80 things would be broken. So, we would decide, "All right, is it bad enough that we have to roll back the whole thing? Or is it okay where we can live with some of these bugs, and then just set out the QA team to start doing some bug hunting?" And you don't want to do that. That's just not the way that rapid software development and modern software development works. So, this idea of deploying very quickly and being able to see if there's any impact that is negative or whatever, and be able to roll back those changes quickly, I think, is super important.

And then the other thing you mentioned about sort of shifting left, or this idea where the developers become more responsible for the code that they write. I think that's actually a really, really good thing, where it's like, "If I'm going to put a piece of code out there that is going to use too many cycles, or it's slowing things down, or it's affecting the latency or whatever it is." I shouldn't rely on some other engineer that's running my system to say, "Hey, I found this problem in your code, can you go fix it?" It should basically be as a team, you're saying, "Okay, I released this code. We're noticing these high latency warnings, or errors, or whatever. I'm the one who is responsible for that. I should go in and I should be the one that fixes that."

Sarjeel: Yeah. That's absolutely true. Okay. You mentioned that at some point, you used to write code, and then you used to interact with the QA engineers and things like that. In that sense, Jeremy, I have been lucky that when I started my career ... I started my career around 2018. Not that way back then. When I started my career in 2018, the first company that I joined was Thundra, actually. You probably heard of Thundra. I believe you have had ...

Jeremy: Absolutely right.

Sarjeel: You have had Emrah Şamdan over here also talking about serverless observability and debugging, and things like that. I joined Thundra. And then after Thundra, I joined Opsgenie. And both of these companies practiced building software, along the principles of DevOps. So, I have never actually seen QA engineers or a specific team just to resolve incidents. For us, it was always like, "Okay, you write the code. You wrote the code. If something goes wrong, you're on call" ... And if you're on call, or even if you're not on call, the person on call would alert you that whatever changes you made, something was wrong. Then they would pull you in as a responder.

And then I look at our customers. Some of our customers still do have these practices. Especially when you're a large enterprise customer, it's a bit harder to change the entire culture. It's a bit slower. When I talk to these customers, and they tell me about these problems, it becomes very difficult for me to relate to them, essentially, because I have never ... But coming back to your point about like, if you build it, you run it. And I think that's exactly what I see serverless and serverless offerings as a great opportunity, especially when you're new to DevOps and you're trying to look at DevOps, or you're thinking of adopting DevOps, or your team is thinking of adopting DevOps. I believe this is where serverless comes into play. If we go back to what I previously said about like, "Okay, we want to try to reduce ops. We want to see a left shift of you build it, you run it." So, as things coming closer to the people who are building things. We also want to see automation. This is where I believe serverless comes into play.

Jeremy: Right.

Sarjeel: The reason why I say this is because ... So, when I graduated from university and I got my first job as a junior developer, Thundra gave me a perspective of both ... This is cloud computing, right? And I just graduated, and I had seen, "Okay, this is cloud computing now." I had always heard about it in university. Right in university, you hear about the latest trends and things like that. "Oh, my god, I'll get to work on AWS." I've never interacted with any AWS service before. And I was presented containers, EC2 containers, and I was presented AWS Lambda. And with AWS Lambda, I just got to it, wrote my first lines of code, got it and uploaded it, and I was able to trigger the lambda function. With EC2, I spent quite a while trying to understand, getting over that learning curve, to a point where I was like, "Oh my god, if I don't get it done by this week, I'll probably be fired."

Jeremy: Well, it's funny, though, that you mentioned the idea of where serverless fits in, in DevOps. I totally agree with you here. And I'll give you a history lesson. And so, I hope I don't sound like an old man yelling at clouds. But essentially, how it used to be was that you would need to maintain a server somewhere. And usually, it was a physical server that ... We weren't even talking about VMs and things like that. It was a physical server, and there was networking, and there's all these other things you had to do with it.

And that was something where there was a clear line between someone who was a developer and was writing code to between someone who was actually installing software patches, and doing the networking and actually plugging in cables in a data center somewhere. So, a lot of that changed in the late aughts, 2008, 2009, when EC2 started to become more popular with AWS, and so forth. And that made it a little bit easier, but you were still thinking about VPCs and trying to do networking and that kind of stuff. It was easier, but still something that you wanted someone to set up for you so that as a developer, I would just have an environment that I could use.

What serverless has changed is that now you just have an environment. And so, you don't have to set up an environment, you just need an AWS account or a Google Cloud account, or IBM or whatever, that you can just go and just upload some code and have it immediately execute within that environment. And so, that's one of the things for me, where if you try to say to a developer, "Hey, I need you to take responsibility for all of this stuff. And oh, by the way, we're running on EC2 instances, and VPCs, and you need to know the security groups, and you need to understand how all of these things might be able to affect you." That is too much, in my opinion, to ask somebody. But to say, look, and you're throwing your code into ... Even if it's a container in Fargate or something like that, or you're doing a Lambda function, that's pretty isolated environment. It's pretty easy for you to reason about if something is not working. "I'm not able to connect to a service. It's running too slow. It's timing out." Things that are easy, I think, for you to understand and debug, and that just becomes ...

I don't think that's too much of an ask. So, I do think that you're asking developers now to go all the way through that spectrum, and to understand a little bit of the operational aspect of it, but they don't have to understand the deep networking stuff or how packets are routed and some of that stuff. They just need to understand some of the basic cloud principles. I think serverless enables that and really is this huge enabler of companies accepting DevOps.

Sarjeel: Yeah, exactly. That whole point about a lot of the underlying infrastructure being abstracted away to the cloud vendor and becoming the responsibility of the cloud vendor. That in itself is just extremely helpful to anybody trying to practice DevOps, any team trying to practice DevOps. Because all of a sudden, you no longer have to worry about your ENIs, or your security groups as you mentioned. All of that is managed by the cloud vendor that you're using, whether it be AWS or Google Cloud provider. What that allowed you to do is, as we have seen quite a bit, as one of the well-known benefits of serverless is it actually allows you to focus on your business logic more. It not only allows you to focus on your business logic, but another hidden gem, I would say, is that it also allows you to connect and communicate, focus on the communication and sharing of code, and getting over that learning curve when the other teams are involved.

So, even though you didn't write the code yourself, if you look at somebody else's Lambda function, you can focus on ... Or if you look at somebody else's FaaS functions or Lambdas, let's say, it's easier to understand. It's easier to collaborate on a code base. And, on top of that, it just makes it easier for an entire team going through that spectrum to manage that pipeline, the DevOps pipeline going from ideation to ... In fact, I say that, especially as a product manager, I would say that all product managers should also learn how to deploy Lambda functions, especially when you're trying to ideate through an idea.

It's become so easy. You can write throwaway code. It becomes very easy to write. You just write throwaway code. Just code that works, just to test whether an idea works or not. And especially when you're trying to find that perfect product market fit, just write a bunch of Lambda functions with your engineering manager or your lead engineer, and show that to the test group of customers, see if it works, go back and ideate it. It's so easy to do that because, one, serverless functions are cheap, or serverless services are cheap. The pay-as-you-go model. They're very lightweight, they're very easy to get up and running with. You don't need to worry about all that infrastructure that we already talked about. So, even there, just in the ideation phase, it's very easy to go forward.

Jeremy: Yeah. I think there are a lot of benefits to just using serverless to do some of these DevOps practices. And I know we haven't really mentioned all of the principles, I guess. We mentioned a couple of the main ones, but I think one of the things a part of the DevOps culture, or at least a part of what you need to do to fully embrace it is this idea of building microservices, right?

I mean, microservices allow individual teams or small groups of people to work on parts of the application independently. And when you start dealing with some massive monolith, and you've got a bunch of different teams all contributing to the same code base, it gets really, really messy. So, being able to break those up into smaller things is super important.

Serverless, I think, has a bunch of really cool things baked in, especially with intercommunication between microservices, and you don't have to set up things like Kafka, or RabbitMQ, or some of these other things that's just another thing to manage. So, what are your thoughts on that? What are some of the tools or the services available as part of the serverless ecosystem that just help with microservices?

Sarjeel: As you mentioned, microservices, we're all familiar with the benefits of microservices.

Jeremy: I hope we are.

Sarjeel: Hopefully. Believe me, I have dealt with a monolith, especially like when you look at front end as a monolith. In many cases, front ends can be considered a monolith you have this one big front end code base, and it just becomes very difficult. You really do see the benefits of microservices. It's actually the idea of microservices that really plays well with the entire DevOps culture and practice, where you can have each team working on something, you can go fast on that. Especially when you look at the stability.

I know we're going off on a tangent over here. We haven't started talking about how serverless is baked into the benefits of building microservices. I just wanted to mention that one point that is really amazing that I have seen dealing with monoliths and microservices is that when you're looking at it from a DevOps perspective, and you're looking at stability of your system, just having one part break and not affecting the other part. That in itself, I believe, is taken granted for. It's pretty amazing. Being able to decouple all these different aspects or all these components of your entire system. And looking at them individually where one component's failure does not necessarily result to another component failure. That, in itself, is pretty amazing with microservices.

However, what does that lead to is that when you're thinking of microservice architectures, then you also need to think about communication overhead, as you mentioned. Yes, you did decouple all of these, but now you still need all of these to communicate with one another.

Jeremy: And reliably.

Sarjeel: And reliably. Yes, exactly. Reliably. As you mentioned, there's a lot of overhead over there. I personally haven't dealt with Kafka or RabbitMQ.

Jeremy: Consider yourself lucky.

Sarjeel: Yeah. We saw EventBridge, and I think a lot of people would agree with me over here, that EventBridge is definitely the next best thing after AWS Lambda. That's because of all the use cases that it has enabled, and how powerful of a service it is. And it really allows you to think about serverless architectures and event-driven architectures from a whole new perspective.

One of the best things that it allows you to do is reduce all those ops that you would otherwise have to deal with. All that overhead with communication. One of the things like even marshalling and de-marshalling, you're literally just communicating in the form of events. Being able to leverage other capabilities of EventBridge, such as routing of events based on rules. That in itself also just enabled a lot of use cases within your microservice architecture and also as ancillary services supporting your DevOps pipelines.

Yes, one is definitely EventBridge. Again, when we were mentioning all of these services, it's also good to point out that age-old myth about serverless equating to only Lambda function. A lot people, even I, when I began, looking at, "Okay, what is this word, serverless?" I started thinking, "Okay, yeah. Serverless equals Lambda functions." Then I realized, no, it's actually a whole set of tools. It's a whole set of services that are available out there.

We mentioned EventBridge. We should also give credit to DynamoDB. Especially when you're looking at it from the point of scalability, you can have your entire microservice architecture built using serverless service. But if your data layer isn't scalable, then what's the point?

Jeremy: What's the point? Right.

Sarjeel: Exactly. Having that incorporated also. And then also, having a lot of the responsibility being abstracted away to the cloud vendor, that in itself allows a lot of teams trying to adopt DevOps to go faster.

When you're looking at EventBridge, DynamoDB, then of course, you have your AWS Lambda functions, your Fargate, basically your containers as a service. If you find FAS services a bit limiting, you can always look at containers as a service. We are seeing a rise in popularity with containers as a service. So, you have this whole set of tools in your cupboard that you can just basically bring in plug and play. And that's what serverless allows you to do. It lets you bring in a service and let you plug it in and play it in the entire way your microservice architecture operates.

Jeremy: Right. Yeah. I think you bring up a point too. You mentioned DynamoDB, which has global tables, and all kinds of things that allow you to replicate data to other regions.

The other thing that's cool about DynamoDB, or EventBridge, or Lambda functions, is that it runs in multiple availability zones, even if you're running it in a single region. And that gives you redundancy and resiliency and all these backups that ... Again, speaking of Kafka or RabbitMQ or something like that, where you'd have to have multiple services or multiple systems running in multiple regions or multiple availability zones that were subscribing to all these events and trying to manage all of that complexity.

EventBridge just kind of does that for you. You don't even have to think about it. Same thing with DynamoDB. But DynamoDB global table is actually something where this could get us to, maybe not an easy way to get to it, but certainly possible to start thinking about active-active regions, where you can actually have your systems running in Europe, and you have them running in the US, and maybe you have them running maybe in Australia or something like that.

So, what are your thoughts on that and where serverless helps get teams to deliver ... Not only to deliver software faster, but to deliver software to more places or more regionally.

Sarjeel: Right. If we take a step back, and if we look at ... You mentioned active-active. Yes, active-active architectures in itself is a whole different topic. And you have one of the great personalities in this field, Adrian Hornsby, who talks about this quite well. He has a great set of resources, blogs, and talks about that. Anybody who's interested and wants to learn anything what active-activity, they can definitely go and refer to that.

But if you look at that architecture from a DevOps point of view, from the fact that what do we actually want to achieve with this type of architecture. You trace back a lot of its motivation and its origins, not origins per se, because it's been there in academia for quite a while now, but a lot of the motivation to adopt such an architecture. A lot of it comes from the fact that, that one horrific story of the Netflix outage.

I just want to mention as a side note. We've been talking about this Netflix outage for quite a while. I'm just waiting for the next big outage because I think we have been overusing this Netflix outage story quite a bit now. Working in an incident management tool, like an incident management company. We do hear about a lot of outages, but none of them compared to what we saw with Netflix, or on that scale, but regardless. So, we see a lot of that motivation for active-active coming from the outage of Netflix, and we saw Netflix kind of start pushing the idea of resilient architectures. I'm not saying that it wasn't there before. Of course, it was, but we saw Netflix, one of the big tech companies starting to push and really think about it from a whole new perspective. And the whole point over here is to maintain stability. As we mentioned earlier, that actually is one of the goals of DevOps. Now, when you're thinking about active-active, it's easier said than done.

Jeremy: That's very true.

Sarjeel: Right. It's actually easier said than done. When you start thinking of how serverless tools can come and help with setting up such an architecture, we do see a lot of burden lifted off. So, for example, you mentioned DynamoDB global tables. Then there's also Route 53. So, we have geo routing that they recently announced. I believe it's one of the more recent capabilities with Route 53. We have DNS failover with Route 53. Then you have API gateway, which came up with custom domains, which allows you to target now regional endpoints. So, we can see, we can actually build this entire active-active service with serverless services, where a lot of that responsibility again, gets abstracted away to the cloud vendor.

I know I've said this statement, being abstracted away as a cloud vendor quite a bit. Simply because I want to stress the fact of how important it is to try to reduce the ops to eventually move towards a very successful DevOps practicing team. Again, having a lot of these activities, let's say, being automated by these managed or services, especially when it comes to scalability. We talk about serverless in the sense that a lot of times when you talk about the limitations of serverless, you look at it, one of them is that it's stateless. A lot of people have difficulties in thinking about stateless. How to think of stateless, whole business logic, and how would you have a business logic translated to a stateless architecture. But with active-active, it actually becomes an advantage to have stateless architectures. To have stateless compute services, because you don't want to hold the state in too long in an active-active. And you want to keep on switching between nodes.

Another benefit where serverless really shines is the fact that it's pay as you go model. So, if you have nodes that aren't being used, why pay for them? So, that's another advantage where you can see the benefits of serverless come into play.

Now, there's that, and there's also the scalability. So, all of a sudden, you have a lot of traffic being routed to a specific node. You may not have handled your routing rules very well, considering your traffic or things like that, but it's okay. It's okay. You DynamoDB is going to scale. If you have, let's say, a Lambda function over there, or maybe Fargate instance, it will scale. So, auto-scaling, pay as you go, statelessness, all of these come together to really help you build that active-active architecture and start thinking of how you can build that active-active architecture.

Now, by the way, another thing I want to mention, which I think we missed upon was when we're talking about the characteristics of serverless functional or serverless in general, and how we are using these serverless functions, or serverless tools to build microservices. One of the characteristics is that when you think of a serverless architecture, it's event-driven, and that plays very well when you're dealing with microservices. All of a sudden, you now have to start thinking about event-driven architectures. Again, we're coming back to EventBridge where your EventBridge may be triggering a lot of your Fargate instances or Lambda functions. Just having that constraint of having functions or having your compute resources being triggered by events allows you to make sure that you think about this architecture in an event-driven fashion.

There is a possibility, though. I must point this out that there is a possibility for you to fall into an anti-pattern. Especially when you're trying to adopt serverless, and you're moving to this granular architecture, you're moving to serverless architecture, a microservices from your monolith. It is easy to fall into an anti-pattern where you try to replicate your entire logic, your entire business logic on monoliths exactly into your serverless architecture where you would have one Lambda sitting before another Lambda function, which sits before another Lambda function. All of a sudden, you have this anti-pattern, where you can't do things as synchronously. And asynchronicity is something that is, again, another benefit of having Lambda function, but just thinking about microservices. So, there's this anti-pattern you may fall into, but as long as you think about it in an event-driven fashion, as long as you know what you're doing, as long as you do your research before building its architecture, you should be good.

Jeremy: Yeah. I think that anti-pattern is very prevalent where people just end up, unfortunately, trying to stack too much logic, or try to chain functions together in a way that is definitely slower.

We talked a lot about the benefits, I think, from a DevOps perspective, or from a DevOps culture of building things with serverless, and I think that makes a lot of sense. There's still ops work to be done. We still have to clean up development environments, or maybe run some audits or some of these other things. So, I guess from that perspective of ... And maybe this falls more on operations, but I think it's sort of part of the full cycle. Where does serverless fit in there? And what are some of the tools that are available for you to sort of just kind of run the infrastructure beyond just trying to deploy code that is maybe client-facing?

Sarjeel: Yeah. Actually, this is a pretty great question, because when I look at serverless and how serverless can aid in DevOps practices. We talked about serverless functions and serverless technologies inside your main code base, inside your main infrastructure itself. But yes, then there's a whole set of other use cases that we can come to where your serverless function to serverless technologies can act as ancillaries, helper functions or ancillary services, aiding you to get through that DevOps pipeline.

So, as you mentioned, cleaning up your environment, or even just thinking about deployments, how we're looking at automated deployments throughout the CI/CD stage, and also automated tests. Then again, monitoring and debugging and identifying root causes of incidents and remediation and all of that. All of that can actually be done with several functions. There are a lot of tools out there that are trying to help you achieve these things. Atlassian itself is building a lot of tools that helps you achieve this, helps you automate through this. But having those tools, and having those third-party tools and having serverless technologies integrated with those tools really does give you that extra boost to go faster while maintaining the stability that we're always talking about, that we're really trying to go for. I can give you an example.

Jeremy: Absolutely.

Sarjeel: There are actually many examples that we can talk over that I would actually like to point out. One of the examples that I really like is, for example, we recently ... Actually, not recently. About a year ago or so, we built an integration with EventBridge. And basically, the use cases were such that the way we saw customers using the Opsgenie, EventBridge integration was, okay, they get an alert from either Datadog and New Relic about some form of configuration drift in their infrastructure. Once they identify this infrastructure drift in the AWS setup, Opsgenie would send you an alert. And that alert acts as a trigger through EventBridge into your AWS infrastructure that can run automated playbooks to correct that configuration drift. I think that was just an amazing use case that we saw some of our customers using.

Another thing was like, for example, security compliance, or when you see some suspicious activity in your account, you can use AWS CloudTrail, or you can use any other security monitoring or audit logging tool. You integrate that with Opsgenie. Opsgenie gets that alert. And upon that alert, using ... Again, this is where I've seen customers leverage the event routing capability of EventBridge. Depending on what the content of the alert is, they're routed to the right area of their infrastructure to immediately remediate that. All of this being done automatically.

So, what we're actually seeing is kind of a reduction in the need for SRE teams, the need for infrastructure maintenance, and basically all of ops. The developers themselves can't set this up, because it's so easy. It's so easy to get up and running with EventBridge and Lambda functions or serverless in general, that you can have your development teams set this up, and take responsibility of that ops part also.

Jeremy: Right. Yeah.

Sarjeel: I think in the beginning, you mentioned like sometimes ops can get scared, like, "Oh, what's the point? What are we needed for?" Why not? Let's come together. That's the whole point, of coming together, and helping. If the development team can't set it up, that's where I feel that we need to start thinking of ops in a whole different way, especially with the advent of serverless and all of these third-party tools. We need to start thinking of ops in a whole different way of how we can leverage this new technology in the best way possible to accelerate according to the team's development practices, according to the team's cultural practices in building software. Because, again, every team is different. That's also a reason why. You can't really say that there's one solution or one tool that fits that would solve all the DevOps problems of the industry. No. Every team is different. Every team is different within an organization. Every organization is different.

Regardless, you can have like third-party tools to try to help you bolster your DevOps solutions. But as soon as you see that, "Okay, it's not working." That's where you can fill in the gaps with serverless services I think that in itself is just pretty amazing.

We are looking at customers do this with Opsgenie wasting a lot of automation come up. For example, in Opsgenie itself, we're using Lambda functions to replicate customer traffic. You have synthetic monitoring Lambda function, and you have transactional monitoring Lambda function. So, these synthetic monitoring functions, they're hitting our APIs. Then the transactional monitoring Lambda functions are receiving the input and processing it and sending it to New Relic, and the other monitoring tools that we're using. Whenever something is wrong that Lambda function will automatically send an alert to Opsgenie, a surprise, we use Opsgenie ourselves internally. We get an alert. So this way, we manage to track or we managed to catch errors or incidents before our customers can even get it.

And remember, this is, again, where you can leverage the characteristics of serverless services or tools, because, again, it's pretty easy to set up, so developers can set this up. I remember going around and playing with a few monitoring Lambda functions myself when I was a developer. So, I set it up, and then getting that connected to New Relic, and doing the whole ... Again, we try to play by that motto. You build it, you run it as much as possible. So, for example, when something goes wrong in Opsgenie, we get that alert. We try to investigate it ourselves. So, all of that is made possible because at some point, we are using Lambda functions to send a lot of monitoring data and generate a lot of data and send that monitoring data over to New Relic.

Jeremy: Yeah. I think you hit the nail on the head in terms of where the SRE team members go after some things become easier. And you mentioned this idea of CI/CD pipelines. So, if it's super easy to set up a CI/CD pipeline, and it's just a matter of a couple of clicks in a dashboard, or it's just you have to deploy maybe another cloud formation template or something. If I was an SRE, which I've done roles similar to SRE in the past, I would be really, really tired of setting up another CI/CD pipeline for somebody. If that was my job, just, "Oh, we got to set up another one of these. Set up another one of these." That is just wasted human capital where you could be spending that time, like you said, writing a lambda function that sends synthetic traffic or getting into chaos engineering.

If you have people who know the ops side of things, and can say, "Hey, what happens if this service can no longer communicate with that service? How does your service react? How does the other service recover, and so forth?" And again, becoming chaos engineers around that, I think, is a hugely important thing that larger teams have got to start doing maybe even smaller teams. But you've got to start doing to understand the nature of distributed systems, and what happens when one thing breaks down. So, I do think that there's an evolution here, where it's like the more you can automate, the more sort of your developers can own some of that stack, it just frees up people to do more important work than things that can just easily be automated.

Sarjeel: Yeah. No, I definitely agree with you. Once you have a lot of automation ... Again, we get back to the same point where you have automation, you can start thinking of the business logic and start thinking about how you want your company to perform to scale and basically work for your customers. Now, when we look at SRE, SRE is now free to start basically looking at the resiliency of the system. Performing more tests, making sure that we are at the amount of nines that we want in terms of resiliency and availability. And SRE gets freed up because of that, right?

Also, when you're talking about automation, you mentioned CI/CD. I wanted to sidetrack to this upcoming concept of GitOps. So, we've heard quite a bit of it recently. There's a company we've worked, I believe, they're really pushing the needle on this. They're looking at this quite a bit. Even that, when we think about automation. So, a lot of that automation, when you're managing ... This entire idea of GitOps, the motivation comes behind like the rise of Kubernetes, Kubernetes becoming popular, how you would manage your Kubernetes infrastructure. And they're looking at the property of Kubernetes to kind of ... Because it's a defined architecture. I forgot the term. I'm so sorry. You define it in your kubectl and everything. You'd push the infrastructure changes or your code changes. That's when GitOps would basically automate. You'd go from continuous delivery to continuous deployment. It's a push from continuous delivery to continuous deployment.

Whenever I look at continuous deployment, and so whenever I look at GitOps in general, I find it very scary, because, okay, you pushed something. And all of a sudden, all of these things are happening automatically. Your entire infrastructure is just about to change. Your entire code base is about to change. It's always very nice. You have somebody in the middle, in a staging environment. You first push to a staging environment, somebody in the middle.

Jeremy: You test it. Right Yeah, exactly.

Sarjeel: You test it. Exactly. And then there's a little button that says, "Okay, deploying." And it goes to production, everything is cool. But you're reducing that. At the end of the day, that's what the idea of DevOps is, to try to reduce the manual labor and push for automation. And then I look at GitOps, and I'm like, "Did we just go crazy? Are we going too far?" Then that's where I see, "Okay, you can have ..." GitOps shouldn't only be thought about in terms of, okay, yeah, automating your deployment, but you should also think of it from a perspective of observability. And, again, when we're talking about observability, that's where I believe we can leverage these Lambda functions, because one, your Lambda functions are very light, and they're just being used for monitoring. So, you're continuously monitoring the actual state, as compared to the desired state. And whenever you see the actual state had drifted away from the desired state, again, you can either use Opsgenie as your alert consolidation tool. Or you can just trigger an event through EventBridge from your lambda function, send an event to EventBridge and, again, go and remediate that. Yeah. Basically, just go and remediate that drift away from the desired state.

So, this is one way. We are looking at automation, and we are trying to find ways to go faster and faster. And this happens that, okay, as we go faster, we still need to remember, maintain stability, maintain availability. And this is where we see the benefit of, especially, Lambda functions, or Azure functions or whatever form of FaaS functions you're using. Because when we talk about using FaaS functions in production for your actual code base, there are a lot of limitations that everybody talks about. A lot of edge cases that aren't covered by these services. But in this case, in this regard, it fits perfectly. One, it's cost-effective, it's easy to spin up, and use it scalable. At Opsgenie, when we want to increase the traffic on a certain API, simply have several concurrent Lambda functions, just bombarding that API with requests and different kinds of requests. This is scalable. It's easy to spin up.

This is exactly where one of the benefits lies. But again, yeah, as I mentioned, in production, you may have some limitations. When you're thinking about it, in terms of microservices and active-active architecture as we talked about before. But when you're thinking about it as ancillary services, and just helping you go through that DevOps pipeline, when you're thinking of it as a glue code, especially, it's really beneficial to use serverless functions.

Jeremy: Yeah. I think all that ties together too. I mean, GitOps is something that just ... CI/CD continuous deployment is one of those things where, yes, it scares a lot of people, because it's just going through FaaS. It allows you to move so quickly and make changes so quickly. And I think that if you embrace the whole culture, if you embrace the idea of microservices and serverless deploying very small units of code. Just this idea of test-driven development or being able to have the tests that you need in there, and so forth, the ability for you to roll back quickly, adding in things like chaos engineering to know what happens if we put something out there and it breaks that we know that the other things will degrade gracefully. Having that capability and kind of following that whole thing. I mean, that's sort of the holy grail of doing this stuff, because it's okay if you break something sometimes, but it should go through a test process and there should be a development environment where you're testing these things against other things. But if something does break, you're isolating it, you're minimizing, you're creating those bulkheads there that are minimizing the impact that it has on a larger scale.

So, we're running out of time and so before we finish, though, I do want to talk about ... I mean, we've been talking a lot about serverless, and EventBridge, and active-active, and DynamoDB and all these great things. It's not like a team can just go ahead and shift tomorrow and start using all this stuff, right? There are a number of barriers to adoption, some of those being just the cultural change in a company, first of all. But also, just this idea of the learning of these tools, and then maybe even the limitations of some of these tools. So, what are your thoughts on some of the barriers that might exist to people who want to adopt, not only DevOps, but maybe DevOps with serverless?

Sarjeel: I can best answer this with a story of mine, or something that I experienced, especially when I switched over to Opsgenie from Thundra. Remember, I was in Thundra serverless. At that time, Thundra was a serverless monitoring tool. Now, it's become much more, of course. At that time, we were focused on serverless. And I was just like, "Oh my god, it's an amazing technology." Then I switched over to Opsgenie and I see, "Okay, we aren't using it that much." And in fact, when I switched over to it, a major functionality of ours, that was initially being built on serverless architecture, the senior engineers rolled back on the decision and went back to EC2 and other forms of container services.

And I asked, "Why did this happen?" I didn't have that much experience, and I really wanted to know what happened. So, I remember, one of the co-founders actually sat me down. He was a pretty cool guy. He actually took me to a whole different meeting room. He sat me down, like, "Okay, I'm going to teach you something now." I'm like, "Okay." He told me that the service is great, and it is definitely, in some way, the future. But in its current state, we do see a lot of issues. And this is back in late 2018, let's say. Around that time, the maximum run time was five minutes, I believe, for a Lambda function.

Jeremy: Five minutes. Yes.

Sarjeel: Yeah. It is in that same year or a bit later that we saw 15 minutes then. So, that is a huge improvement, I would say. At that time, we weren't ready. Our use case was not the best use case. Or the way we were looking at our use case was not in the best manner to adopt serverless. And I think it's very important to understand the limitations of what you can and cannot build, and how you can get around these barriers, because I feel that there's always a way to get around these barriers. Are you just willing to invest in it? It's not like Opsgenie gave up on serverless. We continued. We still use a lot of serverless components in a lot of areas in Opsgenie. Especially in our DevOps pipeline, for example, when we want to spin up emergency instances, we use Fargate. We were using Fargate. I'm not sure if we are still. We're using Fargate for our SRE, in our logging, and other operations, because it's easy to spin up, and it's cost-effective for that use case. But there are definitely limitations.

What I have noticed, Jeremy, even in the small period from 2018 since I graduated to now, I have seen ... We have all seen major leaps of improvement. Just mentioning five minutes to 15 minutes runtime, that was a major improvement. I remember there was this conversation, this whole conversation that I had with Emrah Şamdan and how we were looking at, "Okay, we need to think about runaway cost." We now see that the billing has become more granular for a lot of serverless services, for example, Lambda functions, which was 100 millisecond. Now, it's one millisecond. That in itself is just a huge improvement. I think we see the same thing definitely for a lot of serverless services.

And then there are other things. For example, being able to debug your serverless infrastructure. That in itself is problematic, but we do see a lot of improvements in the industry, for example. And also, a lot of third-party tools are coming up. So, for example, I've been following Thundra's growth as they went from ... They kind of started encapsulating all of ... enabling cloud developers. Recently, they came up with Thundra Sidekick, which is, I think, a very cool feature. If people haven't seen that, I recommend they go check it out. We're definitely looking at it. And this whole community that's coming up to fill in those gaps, and there's still a long way to go. But I still feel that even what we have right now is pretty amazing.

Jeremy: Yeah. No, I agree. And I think that there are limitations. Serverless is not a silver bullet. You're going to run into limitations, but I do see ... I mean, I would recommend to anybody, if you're trying to establish a really good DevOps practice within your organization, or you're just building applications, the services that are serverless, and have those serverless qualities are going to be the ones that make the most sense for you to choose if you can. If you can't, then don't. But if you can, choose those, because that just gives you all of those benefits we've been talking about through this entire episode. And just that ability for you to really own your code, get those CI/CD pipelines to the point where you're delivering multiple releases per day and things like that.

So, Sarjeel, listen, thank you so much for joining me and spending this time and sharing your knowledge on DevOps and serverless. If people want to get ahold of you or find out more stuff that you're working on, how do they do that?

Sarjeel: Well, I'm a pretty open guy. You can just contact me with whatever channel you find. Twitter is great. So, Jeremy, I think you are putting ...

Jeremy: Yes, I'll put the stuff in the show notes. Yeah.

Sarjeel: Yeah. You can contact me through Twitter or on my email. You can find my email on my website, which I think is also going to be in the show notes.

Jeremy: Yep, sarjeelyusuf.me, right?

Sarjeel: Yeah, exactly. So, you can contact me through my Gmail or Twitter, anywhere, and I would just love to talk to anybody. As you realized, Jeremy, I'm also pretty new to this field. There's a lot of learning that I need to do, and I want to do. And so, I would really love for people to reach out, and I would really love to reach out to people and, I guess, brainstorm on a whole bunch of ideas, and especially use cases, because I think it's the use cases that are really driving all of these improvements in serverless.

Jeremy: I totally agree. Well, listen, fresh blood and new ideas, new perspectives are always good things. So, also, if people want to check out Opsgenie, opsgenie.com. But otherwise, we'll put all this stuff in the show notes. Thanks again, Sarjeel.

Sarjeel: Thank you so much, Jeremy. Thanks so much for having me.

View Details

About Jeff Hollan

Jeff Hollan is the Principal PM Manager for Serverless Azure Functions. He started his career at Microsoft in IT and spent a few years managing and building enterprise applications. He is always developing and shipping solutions on the latest tech and is an active member of the serverless tech community.

  • Twitter: https://twitter.com/jeffhollan
  • Email: jeff.hollan@microsoft.com
  • Blog: https://hollan.io
  • Azure Functions: https://azure.microsoft.com/en-us/services/functions/
  • GitHub Durable Task Framework extension for Azure Functions: https://github.com/Azure/azure-functions-durable-extension

Watch this video on YouTube: https://youtu.be/ZDVB0AsYDcs

This episode is sponsored by New Relic. Sign up for free at newrelic.com.

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm joined by Jeff Hollan. Hey, Jeff, thanks for joining me.

Jeff: I'm thrilled to be here. Thanks for the invite.

Jeremy: So you are a Principal Product Manager at Microsoft Azure. And I'd love it if you could tell the listeners a little bit about yourself and what you do as a principal product manager at Azure.

Jeff: Sure. So I've been at Microsoft now for a little over seven years. About five years ago, I switched to focusing on serverless. So I was one of the original members when Azure was like, "Hey, we want to try to go bigger in serverless." So spent some time in a different product called Logic Apps, which has serverless workflows. And then for the last three or four years, I've been running the Azure Functions Team. And so my day-to-day entails understanding a little bit about how the products being used, talking to customers, and then helping formulate the backlog with our engineering team and deliver features to hopefully make people's lives easier with serverless.

Jeremy: Awesome. Well, so I'm super excited to have you here because I think I talked to you a year ago at ServerlessDays Nashville.

Jeff: Yes.

Jeremy: I was talking about having you on the show because Azure Functions and what Microsoft is doing with serverless is, is absolutely fascinating. If there was anybody else who's in the space race against AWs when it comes to the advancements in serverless, I would think that would be Microsoft Azure. And it's pretty exciting because I feel like you are doing things differently. And I've had a conversation with people from IBM Cloud and Google, and of course, AWS, and everybody is doing things slightly differently. So I'd love it if you could just maybe give a quick overview of what Azure Functions are and the general serverless offering that Microsoft has right now.

Jeff: Sure. Yeah, so I guess the best place to start is Azure Functions. And you can in many ways think of it like AWS Lambda. To your point, there are some differences here and there and I'm sure we might even highlight them as we go.

Jeremy: Sure.

Jeff: But at its core, hopefully it is the same. I want to write some event-driven compute. Here's my language of choice. Go ahead and publish it and have it, do its serverless scale option. I think some of the things that folks notice from the get-go, there's a few application concepts that are a little bit different. We enable you to develop and write in what's called the Function App. And so you can actually create four or five different functions that are one deployment thing. And then those four or five functions can scale with each other.

But the other one that I always tend to talk about a lot is just the other supporting products that are around. So you've likely heard, and people who've listened to this have likely heard by CAF, serverless is more than just FaaS. But when you think about the supporting pieces of technology, whether that's serverless workflows with Logic Apps, whether that's Stateful Functions with Durable Functions, going into, I guess the NoSQL database Cosmos DB has a serverless skew. So that's oftentimes where we end up talking a lot more as saying, "Hey, FaaS and functions are going to play a critical role, but it's all these other supporting pieces too that you'll start to see those differences as well."

Jeremy: Right. Yeah, and I think that, again, serverless, at least the evolution of it and what I always think about is it's event-driven, like you said. And so you're getting these events. And in a Microsoft Azure or Azure Functions, they're called Triggers. And again, if people don't ... I'm hoping that people listening to this podcast, they know what serverless is. They know event-driven compute. At least they get the idea of that. But basically it's something gets triggered, a queue is written in it. And that the triggers that Azure Function a or database record is written in it triggers that, or somebody uploads something to Blob Storage. So those are your triggers. But something that's really interesting, and I'd love to know more about is this idea of bindings. So what's the difference, because I understand triggers, but what's the deal with bindings?

Jeff: Yeah, bindings are ... There's two different types of bindings. So there's input bindings where it passes data into your functions, and an output bindings where it's going to write some data. So in the same way that you have this big list of triggers, like I want a trigger on a queue, I want a trigger on a storage account, you can have bindings that talk to these different services too. And in a similar experience to triggers, you don't write that code. So like the best example is, let's say I want an HTTP trigger, so I want my function to trigger on an HTTP request.

Jeremy: Yeah.

Jeff: But maybe that HTTP request has something in the path where it has like a customer ID. So it's like when they call it, the path's going to have a customer ID. And that customer ID has a customer record in my database. And rather than having the first few lines in my function be like, okay, parse out the customer ID, connect to the database, pull in that customer details, you can define what's called an input binding where you're like, "Okay, my trigger is HTTP. I want to pull in data from my database." And the data that you should pull in maps to the path of the HTTP trigger. So you can do this metadata mapping. You say, talk to Cosmos DB, the NoSQL serverless database in Azure. And what will happen is your function triggers. And it's going to automatically go grab that data from the database, pull it in and stick it into your function for you. So it just injects it in for reference data for whatever else. So that's an input binding.

More commonly, we see people using output bindings, which would be, I guess the opposite of that. You can almost kind of connect it. It's like, hey, when this HTTP request is done, I want to write a record to an event stream like Kinesis or Event Hubs is that Azure flavor or a database. Same idea, you set the value in some variable. And then through metadata, through like this JSON file, you're like, "Hey, when my function is done, whatever the value is of this variable, I want you to go write that to a queue message or to an event message or something else. So, they're totally optional. You don't have to use input bindings or output bindings, but in the spirit of serverless, people are like, "Oh, less code. That's great. If I'm pulling data in from a database or sending something out, maybe I could use these bindings instead."

Jeremy: Yeah, and I love this because I have been asking for this type of functionality from another cloud provider for quite some time. But I love that idea because functions as a service generally are supposed to be at least stateless. So they're not supposed to ... You're supposed to be able to spin up thousands of these things. And every time you spin up a new one, there's nothing in there. It's just your code. So you have to go and retrieve data somehow. So if you do pass in an ID for a customer, usually your first bit of code is that boilerplate that has to go and look up that customer record, download that data into the function, then do what you need to do.

And then oftentimes, you want to send that event off. You want to queue that for some additional processing. And then maybe you also want to return event, something back to the HTTP request so that the customer gets something. That's a lot of extra code that you have to write. So that's really cool that you can do that just with configuration basically. And I guess one of the questions I have is, I get being able to write to maybe your own services, like write to a queue, or write to an event bus, but what about to third-party SaaS services?

Jeff: Yeah, we have a few of those, not as many, but there are a few output bindings for services like Twilio is one that I use for a few of my stuff where, same idea, but instead of saying like, "Hey, whatever I set to this variable to a database," we have a Twilio binding or a SendGrid binding that's like, "Hey, this is the variable that will give you the details of a text message that I want sent to a mobile device." And it will integrate with that as well. So you can pull in these different extensions, is what it's called, the view, trigger and binding functionality. And so there's some dozen or so extensions today, including things to like Microsoft Services and elsewhere like Twilio that you can use that can just reduce that code.

Jeremy: Oh, that's amazing. So another thing I noticed about with serverless in general, especially with FaaS. And you said FaaS is much, or serverless is much more than FaaS. But I often see when people are new to it that they say, "I'm going to take my application. I'm going to stick it into one function. I'm just going to let it all run in that one function." And they don't really take advantage of some of the other trickery that is available to them. So I'm curious with the ... And I want to talk to you about composition in a minute, but I'm curious which bindings are people using? This idea of input and output bindings are super powerful, reduces code dramatically, but are people using that? What are the common ones that people use or is there a lot of fall-back to the whole monolithic functions as a service?

Jeff: In general, I am always surprised with how many people use bindings and tell me, "We love bindings."

Jeremy: Okay.

Jeff: A part of it is that as convenient as that sounds to have a variable that you set in something else to storage, somewhere that sometimes folks will hit the boundaries of bindings is, what if you want a little bit more control over for the thing you're doing? So let's imagine we don't have an output binding today for SQL, but let's say you were talking to a SQL relational database. Yeah, it'd be cool to set a variable and it goes to SQL relational database, but what if you wanted to execute something like a stored procedure? Or what if you want it to have a little bit more control or stream the data into a Blob instead of just sticking it in a variable? You can't do that with output bindings. And so we usually just tell people, "That's okay. Just use the SDK." But people are like, "Oh, we love output binding."

So, the most popular ones are probably queues. Queues are just such an important part of serverless when you're distributing things using that message broker. So I think queues takes the cake for us. Storage is probably just another useful one, storing some whatever here and there. And then the final one would be database. But I would guess, and I haven't looked at the binding data specifically. I can think of our trigger data a little bit more clearly. But I would guess that like queues is twice as popular as the next thing down when it comes to what people are integrating with from their functions.

Jeremy: Yeah, well queues, any ways you can use those in multiple directs, especially if you're just trying to minimize downstream pressure and things like that. There's all kinds of reasons why you would use that. But actually speaking of something like downstream pressure, so what happens if there's an error? Because obviously you do a lot of error handling in your code. That's something that a lot of people do. Now, the talk that I gave at Nashville at ServerlessDay Nashville was about not putting error-handling in your code and instead using the features of the cloud to handle some of those errors for you. So, what are the error-handling capabilities in inputs and bindings?

Jeff: Yeah, this is another one where there's this give and take because the way that it ends up working behind the scenes, since you're almost forced to go into this world where you can't do the error-handling because the way the platform is ease, that it's like, your execution is done. I see you have this value in this local variable. I'm going to go take it from here. But if there was some issue that happened, there would be log messages and metrics that got admitted, but your execution finished. You can't really go in and try and catch it and redo it. So you end up doing this types of compensation type stuff where you are getting that alert, or you're getting that metric, and that event that says, hey, the output binding failed.

But that's also a reason where we see, again, bindings are super convenient, but I don't want folks who are using Azure Functions to think, "Oh, well, if I'm talking to a queue, I have to use bindings," because some people are like, "I actually want to have a try-catch and maybe retry it a few times in code." We have some retry policies that will let you provide, like retry something three or four times. But for the most part, you end up being like, "I've got to make sure I'm keeping an eye on those logs when I'm using something like output binding." So definitely a consideration where it's like, "Okay, maybe the convenience might not work for this scenario if I need a little bit more fine-grain control."

Jeremy: Right, Now with the retry capabilities is that something, though, if I try to write to a queue with a binding, is that going to try multiple times or if it fails, I'm going to get a warning?

Jeff: We'll try multiple times and then when it fails, you'll get that warning.

Jeremy: Got you. Okay. Interesting. All right. So are there any best practices though? You mentioned this idea of, if you need more fine-grain control, then just write the code yourself. I would be super happy if it was like, we'll add more fine-grain control for you, so you don't have to write the code yourself. But what are the best practices? When would you say use a binding versus writing your own code? What is the fine-grain? What's that line, I guess, for that fine-grain control?

Jeff: Yeah, I think people usually bump it pretty early on when they're trying to ... like bindings ... All of the binding info is metadata. What it writes to, how it writes to the name of the stuff that it's creating, the content it's pulling from your own variable, but all of the details for a storage blob, what's the name for the file that I'm creating in your storage account? That's usually defined through metadata. And I mentioned at the beginning, you can pass some of that through. You can be like, "Oh, well, the thing in the path parameter make that the file name," but you're still limited in all the things you can control. So usually once you start bumping up against that, and there's even patterns where people do these gymnastics, is what I would almost call it to get everything to work. But once you start bumping up against those types of limits, if you end up find yourself being like, "Oh, I really want to do this thing with binding, but it's not super convenient," you're almost going to be better off at that point to just use the SDK.

Now, to your point, I wish ... there actually should be a little bit of a cleaner wrap to saying, "Okay, well, can you at least get me started with getting that SDK?" Because there are some best practices to, I think, shared the same with Lambda. Connection reuse is the one I see the most often biting people in the butt, is where they're connecting to a database, and they're creating a new connection for every single execution. And what you want to do is move that connect. Yeah, don't do that, no matter your provider, where we want to reuse things. So that's the big one.

But again, to your point, there's a little bit of a cognitive leap there for folks, where they're doing all this through metadata. They're not really thinking about the database, they're just setting a variable. Now they've hit a bump, whether it's around error-handling or around configuration, where you got to make sure you use those SDKs the right way. And candidly, it's like, hopefully you've read the docs, or else you might end up moving from bindings to shooting yourself in the foot.

Jeremy: Right. Yeah, well, reading the docs is always good advice. And ...

Jeff: And everyone does it, right? Yeah.

Jeremy: Right. Exactly. Just like you read your iTunes terms of service.

Jeff: Exactly, right.

Jeremy: Yeah, so another thing about serverless that I think gets a lot of criticism is the idea of cold starts. And certainly functions as a service, it's on demand. That's the greatest thing about it, is that it'll just scale and scale and scale. But there is that penalty in the beginning when a new function, a trigger comes in, it needs to warm up a container or whatever it is that's running in the background. So how does Azure handle, or how does Azure Functions handle cold starts and how much of an impact do you see that affecting your customer's use cases, I guess?

Jeff: Yeah, cold start is the final boss of serverless, it almost feels like. And I really want to see better progress on cold start across the board. And this is like an area of incredible innovation, too. If you look at some of the numbers of AWS Lambda, Google Cloud, and Azure Functions, it's pretty impressive what they can do, but it's still a challenge. And honestly, we're constantly working to get our numbers down.

So I guess there's two answers. One is, how do we help get the numbers down? And then the other one is, what happens if it's just too much to handle? And you have no tolerance for it. So we employ a few things. I think a lot of these are fairly similar, though. I don't actually know how Lambda Google Cloud Functions run behind the scenes, or other providers who might be doing. But a few things we do, we employ this concept called Placeholders, where rather than like seeing file new VM, or file new container whenever there's a function, we actually have this pool of containers that are already running the language, they're already running all of our bits. And then we just hurry and pop your code and we mount it as a zip, and we try to start it up as fast as possible.

In the last six months, we actually have been rolling out some machine learning, too. So we've got some folks in Microsoft Research who worked at looking at a bunch of historical data for functions. It's actually all open source. It's anonymized. But if you go to GitHub, you can actually see a bunch of Azure Functions anonymized data.

Jeremy: Awesome.

Jeff: And they trained a bunch of models. So that hopefully, Jeremy, if you were using Azure Functions and it's Monday at 8:00 AM, that our model, hopefully, would get smart enough over time to say, "Oh, there's a 70% chance that at Monday at 8:00 AM. Jeremy's about to hit this thing. We're actually just going to warm it up before he even executes it." So that's something that we've been rolling with for a while. But even then the ... And then just trying to make progress on the underlying technology, the underlying platform. There's a lot of components to building a multi-tenant secured service that all add a little bit of a national latency.

So something we're aware of. And then I guess to the second part of that question is, we do have some options to fully mitigate it or partially mitigate it. The one is the fateful pinger. We have folks, I mentioned, you can create this Function app concept. You can have multiple functions in there. One thing that even I have done, and I would say don't quote me on this, but I'm on a podcast. Now, my name's right there. You can create another function in that same app that triggers on a timer. So a timer is a first-class concept in Functions. Just have that thing trigger once every 10 minutes, and your whole app is going to get poked every 10 minutes by us. You don't even have to poke it. We'll poke it ourselves on that interval and keep it warm.

Jeff: And then the final option, though, I say this last for a reason because there are cost implications too, is you can deploy your function in the Premium Skew, which lets you preallocate and prewarm, where we will keep it warm, not just by poking it. We'll actually just keep the process running 24/7. But you're paying now for that consistent compute across that time.

Jeremy: Right, Well, I'll tell you one thing. If anything that came out of this conversation is, I am now going to code name cold starts "Bowser." That's what I'm going to ... I'm going to use that now. And if anybody asks me, I'm just going to say Bowser. Awesome. All right. So what are some of the other Serverless services that are in Azure? Because a big part of it way beyond FaaS, this idea of managed services, whether it's databases or Blob storage or things like that, there's such a blurry line now in some cases. Some things are sort of serverless. People put serverless on it so that it sounds good, I guess. But what are some of the other major ones that are available in Azure? And that, again, it's a first-class citizens, I guess, with the Functions?

Jeff: Yeah, and you alluded to it at the beginning too there's almost. And I imagine your listeners would fall in this as well. There's almost the purist view of serverless, and we'll start with those services. And then there's that if you went to the azure.com/serverless marketing page, you'll see a lot of services that we could have a very good conversation on like, "Well, how serverless are they really?" But in terms of the traditional definition of serverless, the only pay when you use it, those things, the service that I see paired with Functions the most and for good reasons is something called Azure Logic Apps, which you can think of it in some ways like AWS Step Functions, if you're familiar with that. Same underlying concept where there's a declarative workflow definition being created, and it's going to help go and orchestrate something for you. I think that the things to check out if you haven't before with Logic Apps is the first, it's got a visual designer that you can do in the portal, you can do in Visual Studio Code. So that instead of crafting that JSON, which everyone loves writing JSON,

Jeremy: Yeah, we like it.

Jeff: Almost as much as they love writing YAML, instead of writing that JSON, you're just saying like, "Do this, add a parallel step here, do that, and the other." But the other one that a lot of folks find use from is, it's got all of these connectors. So there's like 300 plus. I think we might've just crossed 400 connectors where maybe I'm calling function, function, function, and then I want to drop some data in Salesforce, or I want to update a Google sheet. There are connectors for all of these different services that you pop in that workflow too. So Logic Apps is the one that there's a tight pairing with, especially when you want to integrate with other things, or orchestrate stuff that works out really nice.

Jeremy: And what about the Cosmos DB and things like that. And then you have Azure, I think it's just called Azure Blob storage. A good naming there because that's what it is.

Jeff: Yeah, when I was learning the cloud, though, I was so confused when I'd be like, "Get a blob." And I was like, "What is a blob?"

Jeremy: Right.

Jeff: It's just a file. It's just some random bit of binary data.

Jeremy: Right.

Jeff: Oh, yeah. Blob, I think stands for something I don't know.

Jeremy: Yeah, it probably does. Yeah.

Jeff: I really do think it's an acronym, which I should know. Yeah, so Cosmos DB, you can think of it. Again, I know a lot of folks listening are familiar with AWS. So this is similar to DynamoDB. Obviously, there's going to be differences here and there, but there is a serverless skew in that one so that you pay for your read writes on-demand and you get some free tier. Azure storage itself doesn't have a free tier. I would imagine similar to S3, is how you can think of that. Though, it's just so inexpensive that you're paying some fraction of a penny here and there. Another one too, if I'm not to ... So API Management, similar to API Gateway, that's got a consumption serverless to you. And then Azure Event Grid, which is similar to, is it EventBridge?

Jeremy: EventBridge. Yes. Yeah.

Jeff: Is that AWS EventBridge?

Jeremy: Yeah. Not to be confused with Alibaba's EventBridge ...

Jeff: Oh, wow.

Jeremy: ... who has something that integrates with AWS EventBridge, which is very confusing. Yes. So Event Grid is the Azure one.

Jeff: Yes, that's right. Event Grid lets you do some sub stuff. And that it also pulls in events from other providers as well to trigger your serverless stuff. So those are the core traditional way. I guess SQL also has a serverless skew as well. So if you just want to SQL database, there's a flavor that will go auto-scale to zero, and you only pay per transaction.

Jeremy: Awesome. And they all have ... They're all integrated with Azure Functions. There's all bindings and triggers and things like that.

Jeff: Yeah. Yep. Exactly. So hopefully, it's not too hard to use them together in building a solution. Between all of those building blocks, you can put together some pretty impressive things that if they're not being called don't charge any money.

Jeremy: Awesome. Cool. All right. So I want to move on to composition. So function composition. This is something that ... This is maybe whatever the level is before the final boss. Because there are a lot of attempted solutions to this. And by "function composition," I mean this idea of breaking individual functions into very discrete pieces of logic. So maybe I have one piece logic that just calls an SDK or an API somewhere and downloads that data. Maybe I have one that just does some encryption algorithm for me. Maybe I have one that pulls data from another data store or writes data to a queue or something like that if I wasn't using bindings.

So I might have all these different functions that do very specific pieces of business logic for me. And I want to compose them together because I want to reuse them. And the step functions you mentioned, which is the AWS concept for this, they just recently released this Synchronous Express Workflows so you can actually glue them all together in a synchronous pattern and have it return data immediately, which is cool. Of course, there's a cold start issue and some of those other things that come into play there. But what are some of the options in Azure, because you mentioned Logic Apps, which is a really cool feature, and it's almost like a no code or very, very low code solution to that. But what else do you have because you've experimented and have some other products that do this composition?

Jeff: Yeah, and if there's one really interesting piece of tech that's being cooked up in Azure Land that I think all folks should just look at and pay attention to because I do think these types of tools are going to spread beyond, it's durable functions or stateful functions. And there's even some open source tech too that plays very similar roles. In fact, some of the people who built the tech for Durable Functions are building things Temporal Workflow or Cadence Workflow. But what this lets you do, it's almost bizarre how it does it. And so this might be something that you want to go look it up.

You can write ... So to your point, Jeremy, let's say I've got my six different functions that are all doing their own thing. One of them is if I'm doing an order processing pipeline, one of them is get product details. The other one is, create a shipping label. The other one is charge the customer, blah, blah, blah, all these individual pieces of unit, and I want to compose them together. I can create this special type of function called a Durable Function where I write using code that process that I want.

So I could be writing in JavaScript code, in Python code, in .NET code and say, "Okay, first thing, call the function that gets the details. Once that's done, call this function." In the same way, just the same way I would almost call this API, call this one. So I'm writing in code to call these different pieces. I can write loops. I can tell it in code to do the loop in parallel. I can tell it to do the loop sequentially. And so I can orchestrate these processes that can run for weeks at a time, for months at a time. Maybe they only take 15 seconds. And it will go ahead and compose it for you.

And behind the scenes, what it's doing is more or less the same thing that you're doing by hand if you're not using Durable Function. It's storing things in storage. It's storing things in queues. It's queuing these things up. So it's not ... We're not letting you create a function that can actually run for 45 minutes or that actually waits there and double charges you for those calls. It's just this special function that's doing all of this state management for you behind the scenes to let you actually compose this in code. So you end up having something similar to a logic app or a step function. But in this case, it's written entirely in code. The same qualities, though, of a function, it only charges you when a step's actually executing.

So for example, one of the ways I used to Durable Functions is to manage my resources. I spin up new stuff all the time. I'm like, "Oh, a new product got announced. I'm going to go give this a ride." I'm not as good at deleting them after the fact though.

Jeremy: Right. Yeah. Deleting them.

Jeff: Sometimes I get a nasty scare where I'm like, "Oh yeah, I was trying out this new database thing, and now I have this bill." So I have spun up a durable function in my subscription where whenever I create a resource, it triggers this dribble function, and it sets a timer on itself. And it's like, "Wait for a day. And after a day, send me a text message and see if I want to extend it longer. But if not, just delete the resource for me automatically."

Jeremy: Right.

Jeff: And so in theory, this is a function that is running for a day or longer, but I only pay for the few seconds at the beginning where it triggers and sets the timer. And then I pay for a few seconds a day later when it wakes itself back up and sends me that text message. So you can do these really interesting patterns where you're composing, or managing state, or doing things more long running with this Durable Functions product. That's also, you can see all the code on GitHub too. So again, it's very interesting. A little bit complicated when you try to figure out what's actually happening here, but worth keeping an eye on.

Jeremy: Awesome. And so can you do synchronous and asynchronous with those?

Jeff: Yes. So same idea, you alluded to it with step functions, I imagine, I don't know the ins and outs. You can say, "Hey, start this orchestration." And then synchronously returned back to me a response 5, 10, 20 seconds later. Or the default behavior is async where you kick this thing off and it's going to immediately return to you back like, "Hey, we started your thing. Check this end point for the status." And then you would just pull that endpoint so you can control do you want it to hang out and wait, and then send you back the eventual response, or just go run its thing off in the background?

Jeremy: Now, when you're doing the Durable Functions, are you calling other functions that you've already written? And so then do you have control over resource management of those individual functions? Or how does that work?

Jeff: Yeah, so you can call some function anywhere in your subscription as long as it's got an HTTP endpoint. So if it's like an HTTP trigger, there's just like call this HTPP function. Or you can write what are a special type of function. We have a trigger called an Activity Trigger, which are functions that are intended to only be called from a durable function. And so you might have functions that are like, "This is always going to be called from a durable function." There's no HTTP endpoint. It's not listening to its own queue. It's just called an activity trigger, and you trigger them off too. And it's almost the same way as, you don't have any more control over it necessarily than if you were just composing things yourself. We'll just, hopefully, make it easier to reuse them across your account.

Jeremy: Right. So is this something that's a replacement for Logic Apps, or is there ... How would you choose between the two?

Jeff: The guidance that I give the most, and that I see most people falling not for, I'm not trying to trick them, I really want people to be successful, but falling into, it's personal preference. It really becomes a ... There's almost ... We joked about this when we were talking about this podcast. When you think about things of, in AWS, writing CloudFormation templates or writing CDK stuff.

Jeremy: Right.

Jeff: And there are strong opinions on both sides of like, "Hey, the eventual YAML is the best thing. And no, it's so much more convenient to express things in code." It's a similar type of ... It's like, do you want to describe your composability through code and through JavaScript code? Or do you really want to do it in some declarative WorkFlow-y state, the machine language thing? Whichever one of those you want, you can do very similar things across both. There's differences. Logic Apps has all those connectors. So if you want to use one of those connectors, that might tip the scales to Logic Apps. But in general, it's the personal preference is what I end up telling folks to choose.

Jeremy: Right. And are there limitations to this? If I use Logic Apps, would I run into some limitations maybe for latency, or maybe what I can do, or same thing with Durable Functions? Or is it something also where I'd be better off maybe if I only had like two or three steps, maybe just composing those all into a single function themselves rather than adding that extra latency from a Durable Function or Logic App?

Jeff: Yeah, there's a few here that pop to the top of mind. Logic Apps ... The pricing is a little bit different. In dollar for dollar, I would think Logic Apps would be a little bit more expensive than Durable Functions. So if you're optimizing for price, the Durable Function one will be the one you want to go. In terms of scale, they're pretty related. The bottleneck that folks end up running into with Durable Functions, and this is probably deeper down the line I imagine most folks who are listening to this are just getting introduced to it. But I mentioned behind the scenes, it's storing a bunch of state and pulling state for you.

Jeremy: Yeah.

Jeff: You give it details of like, this is my storage account, and this is my queue that I want you to use. And we'll go use that to store the state. If you end up doing really high volumes, like thousands and thousands of these durable orchestrations a second, the underlying storage account will actually start to be like you're reading and writing a whole lot of stuff. You only have a certain amount of limits here. So you end up having to either move to a premium skew or having to rearchitect in a way where you can chart a little bit better.

And Logic Apps has some stuff because it's this managed workflow service. I think you would actually get higher, in theory, scale numbers in a Logic App than a Durable Function. But again, the real bottleneck's going to be that underlying state and that underlying storage account. So a few things to consider there as well. Latency, I think is pretty similar between the two. And that both of them are writing data between steps. So there's a little few milliseconds between each of those functions as it coordinates itself in store state. I think they're going to be pretty comparable though.

Jeremy: Right. And I was going to actually ask you about the billing for that, because I know some of the other services will charge you, and some of the other clouds will charge you every step you take, you get charged there, then you get charged for the execution of the function or whatever the services that it's executing. So how does the billing compare? You said that Logic Apps were maybe a little bit more expensive. That might be because they bill per step ...

Jeff: Exactly. Yeah.

Jeremy: ... whereas Durable Functions just bill execution time?

Jeff: Yeah, exactly. Logic Apps is a per action charge. So every step is a charge. Durable Functions is charging you for the gigabit seconds that you use when a process is actually happening. So the thing to note here is that, when no ... The Durable Function is just in charge of coordinating what step do I do next. So during that period of time where it's deciding what's the next step I do, you're paying for the gigabit seconds. When another step is actually running, or if you've added something a delay, there's no gigabit seconds running at all. So it's a little bit closer to serverless price, or I guess, Azure Functions' pricing, or AWS time to pricing. And then the only other thing is, that underlying storage account, that queue, or that blob store that you've connected to it that it's using to store and retrieve state, there's going to be some cost associated with that as well.

Jeremy: Right. And the billing increment. Is it still 100 milliseconds?

Jeff: So, yeah, it's the minimum of 100. And then above that, it's milliseconds.

Jeremy: Oh, right.

Jeff: So it's the lowest that you'll pay for an execution a 100 milliseconds. If it only lasted 50 milliseconds, you'd be charged 100. But everything on top of 100 is by millisecond. So you might be charged 104 milliseconds, but you would never be charged for 79 milliseconds.

Jeremy: Awesome. Cool. All right. So let's move on to operational, the operational aspect of this because that's another thing I think that hangs people up just when they're getting started with serverless is, they're like, "Wait a minute. Where do I FTP my code to? Or where do I put my container?" It's a different mindset, I think when it comes to writing and packaging serverless applications. So you have a bunch of really cool tools. I think the biggest thing you've got going for you is, you created VS Code, right? So it's one of the most ...

Jeff: Not me personally, but yes.

Jeremy: Not you personally. Yes, I know, but ...

Jeff: Folks I worked closely with.

Jeremy: Right. So, you own a majority of the ecosystem in a sense in terms of the IDEs that are out there. And I love VS Code. And I honestly fought it for a couple of years because I was just Sublime. I just used to use Sublime and it was like, I don't know. And then I switched from Sublime to ADOMD. And then I was using ADOMD, and I was like, "ADOMD is great. And then next thing, people are like, "Oh, VS Code, VS Code." I finally used it, I'm like, "Why did I ever use anything else?" So tell your colleagues there, great job on that because I do enjoy VS Code. But you've built in a bunch of things into the VS Code and of course, there's extensions and other people have done this too. AWS has some extensions as well.

Jeff: Yep.

Jeremy: But just how easy is it to write a serverless application in Azure now?

Jeff: Yeah, this is one area where I feel ... Similar to Durable Functions. This is actually a pretty slick story, comparatively speaking. Yeah, we care a lot about things like Visual Studio Code, and Visual Studio for our friends in .NET land. Obviously, we want to make sure we support whether you're using PyCharm or Sublime or whatever else. But that's where we're focusing a lot of our optimal experiences on. So to your point, there's extensions for pretty much every cloud provider in VS Code now, which is one of the reasons that's been a runaway success is. Whether you're developing on Lambda, or you're developing on Azure Functions, VS Code can still find value. That said though, for the getting started experience.

So yeah, you pop in the Azure Functions extension, it's going to help you create that first project with a set of templates. But the thing here that folks might notice when they're coming from other providers is that once I get that first thing set up, I have my project, you push a five, and you'll see a bunch of crazy stuff happening in the console where the functions runtime will spin up on your machine and give you this really light debugging experience. And I don't mean light, and that it's a subset of the features. I mean light in that, it's just like a little CLI tool that powers this debug experience. You're not dealing with Docker containers. There's also no emulation involved here. So one of the things with Azure Functions is our runtime. The thing that actually triggers your code that runs all those triggers and bindings. It's open source on GitHub. It's cross-platform, it runs on every OS, it can run in a container.

And so that local debug experience is actually powered by the same runtime that we're using in the service to trigger your code. And so it's whatever, a couple of megabytes and installs through the VS Code extension. And now you're just debugging. You go drop a message in that storage account, and you'll see your trigger there on your VS Code box execute, and it will pass in the data and it will hit your breakpoint. And so that really tight development loop that feels a little bit closer to developing an Express app or a Console app or whatever else you might be doing, is something that a lot of folks enjoy about that functions development experience.

Jeremy: Right. And then right from there, you can just publish it, right, into production?

Jeff: Yeah, publish it or stick it in a container and publish it, pop it up there and then we'll just run the same code in the cloud for you and do the serverless stuff.

Jeremy: So what about the CICD process for that? Because one of the things, I think you and I have talked about this before, infrastructure is code for Azure. There is a solution for that, but when you're building a serverless function, you would use a different way to do that?

Jeff: So yeah, you could. The infrastructure is code. The final step, the last thing is Azure Resource Manager or ARM, which you can think of it like cloud formation. It's pretty, and I don't know how unique this is to Azure or not. It's pretty not pleasant to look at.

Jeremy: Well, Azure is not alone because other infrastructures code management systems are just the same, right?

Jeff: So most of the tooling, in fact, all the tooling will generate those things for you. And then you can use them to help deploy your systems for CICD. Still something good to know about and realize that it's there. Though, often folks are either using like our own tooling and VS Code or things like Serverless Framework, so that you can do this with Serverless Framework and have a different abstraction.

In terms of CICD though, one of the ways that I've started to do this recently, and I know it's not fully in the deploying the function everything too, but at least in the CICD process, there's this new experience we just released a few weeks ago where, let's say I go through that local experience. And instead of publishing to Azure, I check it into Git or GitHub specifically. So I just checked my code in to GitHub, either a private or a public repo. I can then just go party over into the Azure portal and say, "Hey, go to this GitHub Repo. This is actually the source that I want for my function. And go connect it to a function that I actually want running in dev or prod or whatever else."

And it will go and create, GitHub Actions for you, set up a CICD pipeline. And so that now rather than manually doing the publisher, the builder, the test steps, I just check code and open a pull request and merge it into the main branch. And it's going to actually go and build and deploy and publish my function app using GitHub Actions. But you didn't have to go figure out how to go, how do I create a CICD pipeline and GitHub Actions? We'll just go wire one up for you that's got all those best practices baked in. So ARM templates have to be aware of third-party services like Serverless Framework that can make your life easier or Terraform. And then finally things like GitHub Actions. We see a lot of people moving towards now as well.

Jeremy: Right. Yeah, and I know the Serverless Framework worked very closely with a team from Azure to build in a lot of that functionality. So that's definitely a cool way to do it. And also another smart strategic move from Microsoft probably was buying GitHub.

Jeff: Sure.

Jeremy: There's a lot of tooling that can be done with GitHub Actions, which are really cool. And actually I've been meaning to dive into those more, but that's awesome. So, all right. So then in terms of the hosting options, you had mentioned earlier, I guess the professional level or the professional tier there. So what are the hosting options? Because I know you have the on-demand, but then you mentioned the professional tier. What does that mean? And what are some of the other options?

Jeff: And I'm realizing that we should have called it the professional tier. We should Azure Functions Home and Azure Functions Professional.

Jeremy: Oh, is it premium.

Jeff: It's premium. Yes.

Jeremy: Oh, it's premium. Sorry.

Jeff: Professional sounds so much better. No. I was like, "Oh yeah, Microsoft, we love having professional skews." We really should have done that.

Jeremy:
Right. Right.

Jeff: Yeah, you're right. So the skew that the majority of our users are on and for good reason that most folks who go build functions is what's called Consumption or Consumption/Serverless.

Jeremy: Yeah.

Jeff: And then, yes, the next click-up is this premium tier, which as the name would entail, comes with a few premium features, including no cold start functions that can execute indefinitely. There's no capital limit on how long they can execute, and you get some bigger beef in your hardware. But it also costs more because it's a premium offering. So that's why a lot of folks choose that Consumption one.

There's a few other nuanced hosting options, which I'll just mention and happy to go into them. One is that, there's another service in Azure called App Service Web Apps or Web Apps, which has been around for a long time. It's for hosting websites, similar to like Elastic Beanstalk in AWS or App Engine in Google Cloud. We actually have a hosting option where we've got a lot of folks who are using functions alongside their websites to do background processing. So you can actually choose to deploy your function into one of those plans for your website. And it will run in the spare compute of your web server. So in essence, it's free, your functions are free. You're still paying for your website though. So you're paying for your website thing, and then we just run your functions in the cracks of that. So that's called hosting it in an app service plan if you want to even cut down costs more, or just keep it warm alongside your web code.

And then finally, there's, I mentioned Azure Functions Runtime is open source. You could stick in a container. We've got a few folks who, and it's an interesting reason of why, and this is an evolving story, but a few folks who grabbed that function, they stick in a container and they'll actually go and deploy it into a Kubernetes' cluster. So we've got a few folks, not very many, I think in terms of a ton of people in Consumption, a few people in Premium, a little bit of people running in an App Service plan, and then a small fraction of folks who are like, "I'm actually going to go run this in Kubernetes." I think that captures the big spectrum.

Jeremy: Yeah, and actually that App Service plan you mentioned is fascinating because one of the things that ... I had a long conversation with Paul Johnson actually about the greenness of serverless, running less servers because you don't have to power as many machines. And that idea of running things in the cracks of other people's or of your own servers, I guess, is really interesting. And it reminds me of spot pricing or spot instances that you can do on AWS. But I'm going to give you a free idea here. You should let people who are running their own web apps rent out their free compute to other people who want ...

Jeff: Oh, sure.

Jeremy: ... right. So then you could potentially lower your costs for hosting your own app.

Jeff: Wow, that's cool that. Yeah, it's almost what my internet provider tries to do without me getting money from it where they'll let other people ...

Jeremy:
Right, exactly.

Jeff: Yeah, no. There's another element to that we've talked about as a team a few times where it's like, I would love a world where whatever compute I'm using, whatever I'm paying for, just be like, go ... If I have a function that needs to execute, go see if there's a crack in my database compute, in my web server compute. Maybe even at one point, there was some experimentation of what if I actually had machines in an office, that I could just say like, "Go run it on some compute somewhere." But you're right. That's an angle I hadn't thought about, which is super interesting, which is, maybe I'm running a web app and I'm like, "Look, I'm at 80% CPU all the time. I'm fine to rent out 10% of my cores for someone." And then they run their function, they pay for that thing. And I get a little bit of chunk of that back to bring down the cost of even the 80% that I am using. That's a pretty slick idea.

Jeremy: I don't know, it just came to me. So #AzureWishlist, I guess. All right. So are there any limitations though that people need to think about when you're building these apps and deploying them? Again, you mentioned different ways to package those, and things like that, but are you limited in any sense in terms of what you can do? Can you deploy other resources as part of these apps? Or where do you bump up against the rough edges there?

Jeff: A few... There's always limitations. I'm just thinking of the ones that I think are the big gotchas for folks. So, one to be aware of is, I mentioned at the beginning, we have this function app concept where you can create multiple functions as the part of the same deployment package. So think about you have a customer's API, you might have "Get Customers, Add Customer, and Delete Customers" that you all deployed in the same app, and it's like, "Oh yeah, that makes sense."

However, where people get into some trouble is, they'll be like, "Oh, I'm going to have 50 functions in the same app. And I'm going to have my customer API. And I'm also going to have my sales API. And I'm also going to have my, whatever orders API, all in the same app." The challenge with that is, behind the scenes, your function app is the unit of deployment. So if you version one thing in your app, you version the whole app. But it's also the unit of scale, which means if something needs to scale or something needs more resources, we do it as one big chunk for your whole app.

So I always tell folks if you're worried about it, just do one function to one app. Then you're writing in a very way to AWS Lambda, where every Lambda adds something. You can do that, that's fine. You're managing more resources, you're managing more deployments. But you know this, Jeremy, you can survive in that world. It's not the end of the world. If you do want to add a few more things, keep it to three to six. Don't start going above 10. It just becomes a little bit unwieldy and you'll start to see things scale a little bit more slowly, or I guess in general, because they're scaling in these big blocks, bigger deployment package, all those things.

The run duration is another one, especially if you're in that Consumption tier, 10 minutes is the max for a function unless you're in that Premium tier. So be aware of that limitation. Yeah, just other ... If you're interested in scenarios, I think a lot of these span serverless providers, but things like machine learning, where you're trying to do really heavy computation, especially one execution, you want to do multithreading, a lot of stuff, you're not going to have a whole lot of success. We don't give you heavy compute power per execution. Really the thing with serverless as we want to scale you out to as many instances as we can.

And so unless you can break it into thousands of executions, don't try to do heavy compute in one execution, even if you're on the Premium tier, even. I just don't think it ... Not that that's a bad pattern in general. I don't know enough about machine learning to tell you if that's anti-pattern or not, but you might want to go look at some other options to find a product that's going to give you a better experience if those are the types of workloads you're trying to run. That's a few that come up. I'm sure there's probably more.

Jeremy: Yeah, no, and I liked the idea too of fine-grained control over each component of my application. So I want to be able to scale this more than I want to be able to scale something else. I want to be able to control of my memory, and some of those other things. So yes, so that makes a lot of sense of breaking them up. Also this idea of, I guess, from a deployment standpoint, things like Canary deployments, or just like moving things through stages and stuff like that, I know there's something called Deployment Slots. I think they're called ... What are those all about?

Jeff: Yeah, so every function you can create at least one slot for if you want, it's most valuable for HTTP traffic. And the way that it would work is, imagine that your function is scaled out across whatever hundreds of cores, it's got an HCPM point directly. You don't have to use API Management or API Gateway with Functions. We also just expose an HTTP endpoint as part of the function.

So let's say you're getting all these HTTP requests that are hitting this thing, and it's scaled out, you want to version it as you mentioned. You could deploy to the slot. And so this thing is still processing like crazy. Your production functions processing like crazy. You deploy to the slot, you could test it out there. It's got its own endpoint at like slash slot or something, make sure things are working. And then you're ready, the real value of it comes when you're ready to swap it into production, because without a slot, what ends up happening is you redeploy on top of the production bits and then it has to restart and you'll see a little bit of a jitter on your execution count as it updates the code. A slot, we just gradually start to route all of your incoming traffic to the slot.

And so there's this graceful handoff where first, it's 10% of your traffic, then it's 20%. We'll just automatically do that balancing for you until eventually your slot becomes the thing in production. And the thing that was in production is now in your slot. And if something goes terribly wrong and you're like, "Oh shoot, I didn't actually test this," you just swap them back. And you can swap them back and forth and just keep deploying your bits into the staging slot and swap them when you're ready to go.

Jeremy: Oh, that's awesome. Actually, another thing I noticed too about when you go to the dashboard and you're watching your functions, you're watching traffic coming to your functions, it actually gives you ... It tells you what server it's running on. And you can actually see that. It's just a level of insight you probably don't need, but I just appreciated when I saw it.

Jeff: There is ... Yes, you are right. In the monitoring story, and by default, we pair with this offering or this monitoring solution called App Insights. You could think of it, again, if you're in AWS land like x-ray and those types of experiences, App Insights does that. As part of App Insights, there's this feature called Live View, and you can see what instances you're running on, how many real-time executions are coming in. It just gives this real time glimpse at your app.

The one reason, though, where in Azure land that actually comes in a bit of value is, one of the differences ... A little bit late to call this one out, but I failed to mention this at the beginning, one of the differences just in how Azure Functions works is that, we don't by default do a single concurrency per instance. And so in AWS Lambda, you have one execution happening on one instance at a time. And that execution has full access to the memory and course that are available. And Functions will deploy you on a server behind the scenes. We'll route multiple triggers to the same instance as long as we can see that it's healthy, and then you can actually grow into memory. And we will only charge you for the memory that you actually consume. So if you only consume 128 megs, we'll charge you for 128 megs. If you have a few more executions or your function app gets more complicated, and you start consuming 512, we'll charge you for 512.

But that does mean that sometimes when you're debugging, there's pros and cons to that approach around concurrency and pooling or whatever else, that sometimes that Live View comes in handy because then you can actually see like, "Oh yeah, behind the scenes in the Azure Data Center, I'm on 100 servers, and this is how many requests are being routed to each one of those servers. And maybe this is why this instance had a time is because maybe I've got too many functions in that app, and there's just too much stuff running there." So some concepts that like, I want to make sure we're improving the system so people don't have to think about it, but it's something that it doesn't hurt to be aware of.

Jeremy: Yeah. It's super interesting because I've watched a video that you've done or one of your demos where you show it scaling. You're just running your artillery against it or whatever and you see it scaling and all these things being added. And actually, maybe that's a good question because this is something I know there are limitations in other clouds. How fast does it scale? So once you're provisioned on a single server or a single server, does it scale rapidly, or is there a ramp-up period that have to abide by?

Jeff: There's some ramp-up period, but hopefully, it's aggressive enough that it works for folks. And in general, we've got metrics from, especially our larger customers who want to do 300,000 requests per second. And they're doing some big promotion. We're like, "Okay, we've got to make sure we do this." Keep in mind the notion that an instance in Azure Functions is different than a single instance in AWS Lambda. But we will add one, at most, one instance every second. So we will add a new, full-powered server behind the scenes for every second. Now, each one of those instances might handle 100 requests. But it does mean when you run a load test or when I run those load tests, you often see a little bit of a spike for those first few seconds where latency's a little bit higher.

If I go throw 10,000 requests all at once, the first 15, 20 seconds, you're going to get a little bit higher latency while we hurry and add and add and add and add those servers. And then usually after that, you'll see it level off. So that one second, at most one second, adding a server is the number to be aware of, but in general, because concurrency can be different. There's not a whole lot you can design around it. But it is something to note of, if you're worried about load, probably run a load test, and you'll see that hill and valley that I talked about from that.

Jeremy: Right. And then you also have the problem of downstream resources that maybe aren't as Serverless as Azure Functions would be, and having to deal with that problem as well, which is where maybe queues and some of those other things certainly could come into play.

Jeff: Yes. That tends to be the problem more than not. Folks ... I've been on more than a few engagements with customers where they're like, "We really want functions to scale like this." We validate, we make sure they can, and then they come to us a week later and they're like, "Okay, you were right, but it turns out our database behind the scenes did not handle this as well. So can I actually cap the scaling? I need to slow it down a little bit." And we're like, "Okay. Yeah, we'll work with you here too."

Jeremy: Can you do that though? Can you cap?

Jeff: We can.

Jeremy: You can do function concurrency type of thing?

Jeff: Yes. So you can set a maximum to say like, "Hey, I actually ..." You can't do in terms of scale slower. You can't tell us to scale slower, but you can tell us where to stop, and we'll stop at a certain point for you.

Jeremy: Perfect. All right. So the other thing you mentioned too is, you said there's an action versus a regular function or an ...

Jeff: Activity function?

Jeremy: Yeah, an activity function. I'm sorry. And so do all functions get an HTTP address? And then what's the ... what did you just call it? API Management was the name of the service there? So how does that all interrelate?

Jeff: Yeah, so every function does not get an address. Only functions that you say like, "I want this to be triggered through HTTP, or through a webhook," will get an address.

Jeremy: Okay.

Jeff: And then the other ones won't. So like an activity trigger doesn't, an Event Hub trigger, which is like Kinesis doesn't. But if you say like this is intended to be HTTP triggered, it's going to be an HTTP request, or yeah, an HTTP request, it will get an endpoint. And I mentioned, you could just use that endpoint. Like it's HTTPS, it's got a certificate, it's free. However, a lot of folks will actually go and add a layer on top of it, like you mentioned, API Management. They'll do that for a number of reasons. One is that, you might have 20 different functions, and having a single API Management where you can do authentication and monitoring, all those things in a single spot is great.

The other pattern that this is helpful with too though is, if you use something like API Management, you can just swap the implementation details of these APIs out really easily. So maybe you start with, "Hey, we have a big Python Flask app that's hosting our APIs." Maybe I start by hosting that in that web server offering I was telling you about. You front it with API Management, you expose all those APIs. And then little by little, you could actually break those APIs and turn them into functions.

Jeremy: Yeah, you'll start with that one. Yeah.

Jeff: And your users still call the same API. They're still using the same auth. It just so happens behind the scenes. You're becoming a little bit more efficient and you're moving to Serverless, but you're ... you have this layer of separation between your implementation and the actual surface of your API.

Jeremy: Yeah, no, that strangler pattern that is, it's like I tell everybody, don't try to shift everything over to serverless because it's just that it'll take you too long and you'll never do it. So just start breaking off those pieces that you can do that and having something like, again, API Gateway or API Management in Azure, having the ability to now just pick off those routes and send them to different places, I think is super important. But that actually brings up, I think, another topic that I'm curious to get your input on, is just this idea of hybrid applications in general.

In a perfect world and maybe my perfect world, maybe not everybody's perfect world, but things would be running almost entirely serverless. Meaning, that you weren't babysitting application servers, you weren't worrying about databases. You weren't trying to provision more things so that it will scale. But there are a lot of enterprises and a lot of businesses and people maybe don't even want to move fully to serverless that are going to continue to run these hybrid apps. So whether they're running a Kubernetes cluster for some container stuff where they've got a bunch of legacy servers running, or they've got a bunch of SQL Server 2000 still running somewhere in their data center. What's the approach to hybrid at Azure?

Jeff: Yeah, this is ... I would be interested to get your thoughts on this one too. I almost start by saying the majority of folks and the majority of organizations, I feel in this current time, it may be actual, I'll finish the answer by saying the only rub on this is where things might be headed. But like most folks, I don't know if you have a great enough excuse to do anything hybrid. If you're working at a small, mid, you're probably going to find a lot more benefits, not that you don't have, I'm sure you have great excuses and please tweet them at me, I'm fine to read them. But you'll probably find more benefits by moving fully to the cloud.

There are still a subset of folks. And historically, this is an area that Microsoft is really tried to invest heavily in with full on-prem stacks. I tend to sympathize with these folks a bit when there's like, "Hey, we're meeting with a big financial institution." And they're like, "We have legal requirements that all of the compute processing has to happen within this state or this country boundary. And you don't have a data center in this state or this country. So does that mean I can't use Azure Functions?" And we don't want to tell them no.

So there's the world where hybrid does still have merit, and even folks who have huge footprints with hybrid data centers that are almost to the strangler pattern it's the other way. It's like they can't just shut those off all at once. And so that is where we are investing in tools to help make that easier. There's a few different things like Azure called Azure Arc that lets you manage resources that are on-prem through the cloud. Azure Functions plays a role in that too. I mentioned you can take functions and you can run them anywhere on-prem or in the cloud.

So there's worlds where it happens. And if you're in a world where you're like, "I really want to use functions. All these things sound great. I like the programming model. I like how things become very purpose-built, but I might not be running in the fully managed service. Is there still value for me in functions if I'm not getting function pricing?" I think the answer is yes. And I think we have users who are in that mode today who are telling us, "No, there's still value in the serverless scale, the serverless resource utilization even if it's using my resources."

The only other aspect of this that I'm interested to see where it goes is, we have been in a few engagements where if you think about like IoT, Satya, our CEO says like, "Oftentimes with IoT, you want the compute to happen as close to the data as possible." And so I was in this engagement once with a sport's stadium. And they wanted to build a really smart sports stadium with thousands and thousands of sensors. And they wanted to be able to process millions of events every second in the stadium. And to them they're like, "Look, does it make sense for us to send all of these millions of events to the data center, pay for ingress and egress, have the processing happen there and then send the data back, adding potential latency and cost? Or does it make sense for us to actually have those functions running in the cracks of all of the compute that we already have in the stadium?"

And that is where there's a world where IoT might start to shift this as well, where maybe it will make sense in some of these worlds to have that function running closer to the actual data that it's processing. So it's early days in that. You could go look at a tutorial of how to run an Azure Function on an IoT device right now, but it's still super early, and I don't know where that's going to go.

Jeremy: Yeah. Now, I think the programming model for serverless is, or for at least functions as service, is a useful model regardless of whether you're running in the cloud or not. So I totally agree with you. Even if you are running on-prem, but your developers have that level of, I guess, abstraction where they don't have to think about, I have to deploy something to this server or I have to packages a container and then I've got to put it into a pot. I do all that stuff. If it's just, I'm writing a function that needs to react to some piece of logic, I think that makes a lot of sense.

In terms of, I think your stadium example's another great example, which is also probably why compute at the edge is another thing that makes a lot of sense. Maybe you don't need to be sending all that data back to Northern Virginia or Oregon or something like that, but you can instead just send it to the local cell phone tower that can do some processing. And then make a decision as terms of what data has to be synced back to some home run, or some other data centers, something like that. So yeah, I agree with you there. I think that is interesting, but I certainly would be against, just maybe from a conservation standpoint, setting up servers in your own data center just to run functions. I love that idea of running them in the cracks. That just seems to be a smarter move, I think.

Jeff: Yeah. Yeah. And I hope folks don't hear those kinds of scenarios and use them as an excuse to maybe do like, be honest with yourself. Look at those things. I'm not going to tell you everything belongs in the cloud, but by and large, to your point around efficiency costs even environmentalism, there are benefits of these economies of scale as well.

Jeremy: Yeah. And again, unless you are a massive sports stadium or a huge international bank, you probably can put your stuff in the cloud and it's going to save you a lot of money. It's going to save you a lot of headache. I managed the data center for a while many, many, many, many years ago. And it was, "No, thank you. I would never want to do that again." So, all right. So let me ask you this because we're running out of time here, but I'd love to just get your thoughts. You've been on the serverless train very, very early. I'm just curious, where are we going with this? I ask a lot of my guests this question like, what's next for serverless or what's the future of serverless? Where is serverless in five years? You can answer any one of those questions. But I'd love to get your thoughts on where this train is going?

Jeff: Branching from our last talk, the first thing that I feel more confident about moving forward, which was an open question that I've had for a few years is that, the value of functions in an application development pattern, regardless of necessarily cloud provider, just the value of that concept of event-driven compute, highly abstracted, highly productive, I think it's going to become more mainstream than it is. I still cringe a little bit whenever serverless trends in some Hacker News posts. And I opened the comments, knowing that there's going to be a lot of people who are skeptical and they're like, "Oh no, this is like the biggest vendor lock-in scam you've ever seen."

Jeremy: Haters gonna hate.

Jeff: Haters gonna hate. And I think that the value of that model will just grow. And even just the whole notion of, this is a useful pattern, in the same way you think about things like microservices or service-oriented architecture, it's like, "Oh yeah, this pattern has a lot of value." I think that functions being an essential part of that application, not the whole application, but an essential part is a big thing.

Jeff: I expect that we're going to continue to see over the next few years, a lot of innovation around state functions. Traditionally, whenever you're doing a functions overview, it's like functions are stateless and they're short-lived and all those things. I mentioned how we're doing some stuff here in Durable Functions. Cloudflare just recently entered this space with this thing called Durable Entities, I think is what they call it. I don't expect that to be the last of them. There's some startups like Temporal that are making a lot of noise. I would expect, and I don't know, but I would expect that Amazon and Google will continue to innovate either through workflows or other things to just make you managing state a little bit easier with functions. The other one, and I don't actually know where this will go, but I'm trying to keep an eye on it. It's this notion of, and I hesitate to use the word "containers" because I containers has been kind of inflated to mean too much.

Jeremy: Right. Yeah. Very overloaded term.

Jeff: So that there's a notion of could I take an existing application and make no changes to have it conform to the runtime APIs of Azure Functions or at AWS Lambda? Can I take an application that certain however I want and have it run in a serverless way? So some something similar to what Google Cloud run offers. And again, I don't think it necessarily has to be married to containers. I don't think containers is the only way to make that happen. But just making things a little bit more flexible. Making it so that ... And I don't know, we talked too about, there's value in you breaking things into functional pieces. So it's this balance of, hosting a monolith as a function is not where I want things to go, but somewhere in that, I just think we're going to continue to evolve of, I don't know, maybe there's, I don't know, I don't know. I'm thinking out loud with you here, but I think there's stuff in that space of, flexibility for your deployment, but while still getting many of the benefits of serverless.

Jeremy: Yeah. Yeah, now, I fully agree with that. I think the biggest change to serverless is the paradigm shift for how you build apps. And that's just something that, there's so much momentum for building, even absent just using containers, and again, a super overloaded term. But just using Kubernetes or any sort of Docker container type thing and building apps that way, that seems to be ... Has a lot of gravity right now. And so getting people to shift to the single-purpose function is probably not an easy thing to do. But yeah, that's an interesting thought. I thought very much so in that direction as well, like how do you just say to people like, let me take your existing app and I'll make it serverless for you without having to change the programming model that you're used to.

Jeff: Yeah, and I do think even when we're talking about composability, I still see a lot of architecture diagrams, which are well architected, that have so many different function dots all around the architecture. And one of the pain points that I always hear back is like, "Can you make it easier for me to deal with these tiny little pieces that are running everywhere?" Yeah, I get the benefits and I'm agile, but I do expect, I don't know if it's going to ... The serverless application model, whether it's SAM or Serverless or whatever Serverless Framework, I don't know what it is, but I just... There's still room that we've got to go in this space, in serverless to make it a little bit easier.

I think that's part of the appeal that people have with Kubernetes sometimes is they're like, "I have a cluster," and I think about a cluster. And in serverless then, it's like, well, hey, good news, you don't have to think about a cluster. The bad news, now you're dealing with 50 different pieces here, and you've got to version them all and deploy them all and monitor them all as a single app. So yeah, I'm curious to see what we can end up doing in this space and what different providers, whether they're cloud providers, or startups or whatever else do to try to make this a bit easier.

Jeremy: Yeah. Well, I appreciate what you're doing over there at Azure to try to solve those problems. And maybe one day we'll defeat Bowser and save the princess. Is that what it is? I'm sorry ...

Jeff: Look, I do think that there will be a world and as the numbers go down and down where it's like, I think technology will get to a spot looking at where cold start doesn't come up anymore when you talk about serverless. I don't know when that's going to happen, but it will happen.

Jeremy: Awesome. All right. Well, we'll leave it there. Jeff, thank you so much for spending the morning with me here. If people want to find out more about you, or contact you, or find out more about Azure and what's happening over there, how do they do that?

Jeff: Yeah, Twitter is the place that I am the most active and checking things. So @jeffhollan, just my first and last name. I'm fine if folks want to email me as well though. So jeff.hollan@microsoft.com. Shoot me an email. I've got a blog that I'll post some stuff. I was just thinking this weekend. I haven't posted for quite a while. But love blogging of like, "Hey, here's how to do in order processing," or "Here's how to do a stream error handling or whatever else." So hollan.io is my website. And then functions in general, if you're interested to learn more about Azure Functions, azure.com/functions. I think those are four places. Pick your own if you want to have a conversation.

Jeremy: Awesome. Well, I will put all that in the show notes. Thanks again, Jeff.

Jeff: Thanks so much, Jeremy. Great chatting with you.

View Details

About Nofar Asselman

Nofar Asselman is the Head of Business Development at Epsagon, where she initiated Epsagon’s partnership with Amazon Web Services (AWS) and developed growth opportunities for the company. Nofar leads Epsagon’s business development department strategy, through revenue-generating channels and creating new alliances.

Nofar is a key figure at the AWS Partner Community and founded the first-ever AWS Partners Meetup Group. The group is focused on sharing joint AWS go-to-market strategies that successfully affect AWS Partners’ ecosystem and growth.

Nofar is passionate about her work with AWS cloud communities, organizes meetups regularly, and participates in conferences, events, and user groups. Nofar is a Founding Member of the Multi-Cloud Leadership Alliance (MCLA) and she loves sharing insights and best practices about her AWS experiences in blog posts on Medium.

  • Twitter: @AsselmanNofar
  • Medium: nofar-asselman.medium.com/
  • LinkedIn: www.linkedin.com/in/nofar-asselman-074b8134/
  • Multi-Cloud Leadership Alliance (MCLA): https://www.linkedin.com/groups/13779838/
  • Epsagon: https://epsagon.com/

Watch this episode on YouTube: https://youtu.be/FL9XLtW57Ms

This week's episode is sponsored by Off-by-none, the weekly newsletter that focuses on the technical details of building modern applications in the cloud, driven by the serverless community. Visit us to subscribe, provide feedback, submit your articles, and nominate people who are contributing to the serverless community to become a Serverless Star.

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today, I'm speaking with Nofar Asselman. Hey, Nofar, thanks for joining me.

Nofar: Hi, Jeremy. Thanks for having me.

Jeremy: So you are the VP of Business Development at Epsagon, so I'd love it if you could tell the listeners a little bit about your background, what you do at Epsagon, and what Epsagon is all about?

Nofar: Yeah, absolutely. So actually I started in my past life, I was an attorney. I work in a big law firm, and this is where I understood that it's fun to work with startups from the legal side, but I wanted to move to the business side. And this is where I start my startup journey and today I work at Epsagon, where we help microservices customers to adopt microservices in confidence while providing them really seamless experience in monitoring their microservices architectures.

Jeremy: Awesome. All right. So one thing that's really interesting about what Epsagon does is Epsagon has built a business around sort of the serverless ecosystem providing a solution for people who use serverless or are trying to build things with serverless. And I think that's really fascinating because serverless has become obviously quite a buzzword over the last few years. And a lot of people will stick the word "serverless" in their product title or in their description somehow. But what I would love to talk to you about is this idea of actually building a business sort of for serverless, right? So building something for the serverless ecosystem, whether that's a tool, whether that's some sort of a thing that makes it easier for you to monitor or build new things or whatever it is. But something that is for the serverless ecosystem. And so you have a ton of experience in this, so I'd love to get your perspective, but maybe we could start just sort of like what is the current state of the serverless market from a business perspective?

Nofar: Yeah, so I think that serverless is definitely growing. I'd say that it's not growing as fast as we thought it would grow, but I think we see more and more companies are leveraging serverless technologies to really achieve business agility and go to market faster. But I think if 2020 was a year where we did see an uptick in customers that are leveraging serverless, in 2021, we'll see that in a higher scale, just because I think that last year there were still some challenges around tooling and expertise that was still missing from lots of organizations. And there are so many great tools out there now that helping these customers to leverage their technology using serverless and really meeting market demands and meeting their customer's needs smoothly with serverless. So I think this year will be significant in terms of serverless growth.

Jeremy: Right. Yeah. And I think that is something that we've seen a lot of is there's been a lot of complaints that serverless is not easy to adopt as sort of a change in the mindset in terms of how you, again, maybe need new tooling, maybe we need new monitoring tools, maybe we need other solutions that help us do serverless. Certainly need education and training. So do you think, though, that with the growth of the serverless market that there are opportunities for tools and solutions and things like that?

Nofar: Yeah, absolutely. Absolutely. I do think that building solutions around serverless is super important and if we're looking long-term, that's definitely the technology that will be the ground of many companies that will use serverless architectures. But I think today ... And that's what we did in Epsagon, so I'm not very objective, you can see what we've done in Epsagon, where we started to build the serverless solution, serverless monitoring solution and then we expanded the product to also include containers and microservices. Because in the end of the day, when customers are moving to microservices and to more modern applications they would most likely to use serverless, but I wouldn't count on that they're going to use 100% serverless. So you do want to provide them with seamless experience and solution that they can use as a one-stop-shop to their distributed applications.

Jeremy: Right. Yeah. No, and I think that that's a really good point because you haven't seen many shops that are like 100% serverless, right? There's a lot of this hybrid stuff that's going on. I mean, I've talked to several of them, whether it's Liberty Mutual, that's trying to go 100% serverless or LEGO who I think is 100% serverless at this point. And so if you're building tools to support the 100% serverless companies then are you sort of limiting your opportunities there? I mean, like you said with Epsagon shifting over to Kubernetes or expanding I should say to support that. There's been a few monitoring companies that have expanded in that direction as well, went from serverless as a start then they went into the broader container market. But then there are other services and solutions and tools that have stayed just focused on serverless. And then you've had a bunch of other monitoring solutions and other tools that were not serverless that have worked their way back to serverless. So I guess, is the market big enough for serverless-specific tools or do you really think you have to take that approach where you broaden and look at that hybrid market?

Nofar: Yeah, I think it's a great question because at the end of the day, if you're only doing serverless you're probably going to do it better than others. But if you're expanding so the market is larger than it used to be. So I do think that if you're expanding your offering you definitely going to increase the markets that you can approach. But if you are keeping or want to focus on serverless, I think you would need to pay attention also to product integrations and finding partners that can complement your solution. In order again, to provide the customer with a better experience of using one tool or one interface instead of multiple platform if it's monitoring or security or deployment. So I think that's super important because in the end of the day when customers wants to leverage serverless and they want to benefit from business agility and so on, you don't want them to spend so much time on tooling and all of those things. You want to help them to bring value faster.

Jeremy: Right. No, no, no, no, no, I totally agree with that. And I think that's one of those things where some of the companies that are coming top-down for this, they are sort of adding ... It's an add-on, right? And so again, there's a lot of this traditional mindset where it's like, well, here's how you do traditional development. Here's how you do monitoring and deployments. And here's how some of these other things work. Whereas when you take it from the other side of the equation, when you look at it from just the serverless specific. There are things that are done in serverless that are so different than the way that we do them with our traditional application that I do agree. I think that focusing on serverless could give you sort of a leg up, but, you do need to identify the right customer because they need to be someone that's sort of willing to experiment with new tooling.

Nofar: Right, right. Definitely.

Jeremy: All right. Awesome. So what I want to do is I want to use your expertise here and pick your brain a little bit because the point of, I think, this episode is there are a lot of people who are building products with serverless, but they're just building them with serverless. So again, I want to focus on this idea of building products for serverless like things that layer on top of it, things that other solutions, other services that might aid you in your development of serverless or your ability to maintain and manage serverless, or even to educate people on serverless. So, let's start with maybe this idea of how do you find the right contact person in a company you're trying to sell to? Now, I'm not a sales person. I don't think you're a salesperson either. But, I hate selling things. I'm not a huge person trying to sell things. So, this is an awkward conversation for me.

But I do find that if you can find the right person in the company and maybe that's not the right specific person, but just the right level of person. And you can strike up an authentic dialogue with that person then selling becomes less about selling and more about helping, that may be a little altruistic from the sales side of things. Anyway, so if you trying to get into a company that was building with serverless, so who would you start with? Would you try to target developers? Would you try to target senior leadership like the CIO, CTO level? Would you try to go like maybe VP of engineering like mid-level? Who's the right person to target for a serverless solution?

Nofar: Yeah, I think it really depends on the customer serverless journey or where the customer is in which stage of adoption the customer is. Because it's really, you know, early stages, let's say if we look on AWS serverless where customers using are Lambda, but not more than let's say $1,000 per month. So there you would see that most of the customers or most of the users would be developers. This is going to be the main point of contact for you. The motivation there is more of tooling and really improve the developer's experience. But then if you're looking at the company that is a bit further in that journey and looking from $1,000 and let's say up to $5,000 or more or less. This is where the company is starting to scale on serverless and then the needs are different. Then the needs are more to reduce total cost of ownership in the short term, dev ops needs and so on. And there is the main point of contact that I would approach would be mid-level, director of engineering, people in or cloud architect leads, this kind of personas. And then if you were going up then more than $5,000 per month, I'd say this is where the CTO is getting involved. This is where the total cost of ownership in the longterm is more the main motivation for this adoption.

Jeremy: Right. I love that level breakdown. That's actually kind of helpful. So, you mentioned though like where customers are in their journey. So, I think that's another question sort of like I've been involved in a lot of startups. I've unfortunately had to be involved in some of the sales processes for startups, because as a technical person I would come in and sort of do evaluations and things like that. And I know for one, trying to sell into an enterprise can be very, very long and lengthy process. Whereas sometimes startups are a little bit faster to adopt a technology and try something new. Mid-market or midsize companies there's always a sort of a crapshoot there in terms of how quickly they will move and how many layers of red tape that they have. But I'm curious like what would be the size company that you would think would be best to target? Like should you, if you're building a new product for serverless should I start with the startups and try to target those companies? Or should I shoot for the moon and go right for the enterprise customer?

Nofar: Yeah, definitely. I would say that it really depends again on the consumption of serverless on the customers, but where we see most or when we see the need of monitoring for serverless, we obviously see it's among startups, right? They are cloud native, born in the cloud usually they use serverless from day one, and this is really our bread and butter. So definitely start-ups is the main audience for the cloud native in the cloud native space. But then you also see enterprise customers that are starting to leverage serverless technologies, really again, to improve agility and to reduce costs and to reduce overhead from their teams. And you see that they're starting to have more and more projects of where they use serverless and also containers, but serverless in particular.

And once they see success there, this is where you see more projects that are involving serverless as well. I can give you an example that we had a win with one of our customers with one team in an enterprise company. And then within three months, we already had two upsells, two different teams that just saw the success and replicated that. So that's exactly how enterprise ... They are moving slow, but when they are moving it's at scale. So this is why I love also working with enterprises because when they see the value and you have one success with one team it's so much easier to work with other teams in the company as well.

Jeremy: Yeah, no, I think that makes a ton of sense. And that's the other thing with enterprises too, is that now with all these different business models or organizational models that are used, you have these little project teams that are a little bit independent, they're like startups within a larger enterprise. And if those can start adopting a specific tool or some other solution that getting the rest of the company to pile on or to adopt that is a good possibility. So that's always a good tactic, but again, it's about identifying companies and figuring out which companies are doing which. But that actually brings me to another question I think is kind of interesting when it comes to this, not a ton of people or I shouldn't say this ... A lot of people are using serverless, but a lot of people are sort of dabbling in serverless. They're not full on serverless, using it every single day or whatever.

And I guess my question is if you're building a new serverless company and you're trying to target people with this, do you want to try to sell something to someone that isn't using serverless? Or do you want to try to sell something to someone that is using serverless? Like if you're already using serverless, it seems like it's an easier sale to sell them a serverless solution as opposed to trying to sort of create a market.

Nofar: Yeah. I think creating a market is super important, but if you are building a company around serverless and you do want to make profit out of it, I would say that waiting for customers to adopt serverless and sort of educating them and convincing them to move and adopt serverless that would take you a lot of time to see any value out of it. I do think that education is super important, but I think that targeting customers that even if they've just started, but they're already sold on the serverless idea and they already see the value. I think this is where, in my opinion, to keep most of your focus on to educate them how they can take these just initial adoption and scale that. So I would focus on these kinds of customers.

Jeremy: Right. Yeah. I totally agree. I see a lot of blog posts of people who are starting to use serverless, and it's amazing that they're either trying to do things themselves and they're not aware of the larger ecosystem. And again, there's so much content out there about how to use serverless, how to build, how to monitor, how to do distributed tracing, how to do canary deployments, all those kinds of things. And I see a lot of people reinventing the wheel. And I think a big part of that, like you said, is sort of this lack of education, which I mean that's another thing where if you're a startup and you're in the serverless sphere here.

And again, I think this goes for any company that's in any new emerging technology market or something like that, participating in the community, being someone who's out there whether you're just doing webinars or you're posting videos, or you're writing blog posts, or you're going to conferences or contributing to open source, things like that. I think that's really important. I'm just curious your perspective from I guess a business development standpoint, how important is that that as a new company in some ecosystem that you get involved in the community and that you contribute to the community and try to help educate?

Nofar: I think it's super important. I'm a huge believer in communities. I run a community of AWS partners. I write blog posts about things that I'm doing when it comes to cloud alliances and how you can grow your business with cloud alliances. So I think that in the serverless space, there are so many ... You know the community in serverless spaces is just amazing. You are playing a huge part of it and Farrah Campbell as well, doing such a great job in really educating the market and showing that you don't need to be an expert in order to leverage serverless and that's awesome. And I think that going to conferences as a customer that wants to leverage serverless or is a company that builds a solution around serverless. Going to these conferences like ServerlessDays, in the past life before COVID, but ServerlessDays, AWS re:Invent and all these old ... any summit that is really technical-oriented where there is a huge focus on serverless that's definitely where I would go. And also this is a great platform for you to meet potential customers, to learn about challenges that they experience and how you can actually be more helpful to these kinds of companies.

Jeremy: Yeah, no, I totally agree. And I can't stress that enough. One of the things that's been great about Epsagon and a lot of the other companies in this space is they've helped create the community. They've built, they've run the conferences, they've done the webinars, they've flooded things with blog posts. And the other thing I love is, I think AWS, Microsoft, even Google, they do a great job of writing articles and creating content around their particular products and services. But I know for a fact that any article that comes out of one of those places goes through multiple layers of review and everything has to be just right, and you don't get the criticism. I don't think you get the breath of fresh air where someone says like, "Oh, well, this is great. But..." And you don't always get that when you don't have sort of independent voices.

And even if an independent voice is selling their own product, even if you're selling your own monitoring service or your own deployment pipeline or whatever it is, your own CI/CD service. Regardless of whether you're doing that or not, at least when it comes to the technology itself, the overall technology being serverless, I think you've got more honesty there. Which to me is super important because again, I love serverless, I build as much as I possibly can with it, but it's not a silver bullet.

Nofar: Yeah, absolutely. Absolutely. Something will go wrong, but you do need to have the right tools in order to overcome them. Absolutely. Yeah.

Jeremy: Absolutely. Right. Totally agree.

Nofar: And I think what's great about Epsagon as well, is that we are talking in every single conference that we can and our CTO, Ran, is an AWS Serverless Hero. And when he talks about observability within serverless applications, he will talk about Epsagon but in the end he will first educate you about the challenges, how to solve them with open source tools. So, I think that market education is very, very important for every company that is building product in this space.

Jeremy: Yeah, Ran and Nitzan do a great job and your team does an excellent job for the community. So that is definitely much appreciated. All right. So we talked about education as being one of the challenges, and this is just sort of a larger problem I think with just people understanding serverless or what they may need for serverless and things like that. But there's also education in terms of educating people on your own products. So, like you said, with Epsagon you'll educate on the problem itself and then educate on how your solution helps that. So that can be done through blog posts and those sorts of things as well. But what about things like brand recognition? Like how big of a challenge is that for a new startup in this space?

Nofar: It is a huge challenge because when you just starting your business specifically, especially like two years ago nobody really knew what serverless is and then you're coming with a product that is helping to monitor a serverless applications. So that's definitely challenging. You want to gain brand recognition to get credibility so this is where we also use the community and we open our product for a free trial. So companies and customers will be able to see the value for themselves. But I would say that the number one strategy that we had in order to improve our brand recognition would be, of course, doing marketing events, but also working with strategic partners like AWS. And one of the examples that I always give is that about few months after we went GA with the serverless solution.

So we just started the AWS partnership and we were launch partners of Lambda Layers that was announced in 2018, I think. And then all of a sudden in re:Invent we see Epsagon's logo in Mark Vogel's keynote. One out of nine names out there. So all of a sudden from a small company that nobody knows we are on the main keynotes at re:Invent and that just changes everything. So I think partnerships plays a huge role when it comes to brand recognition. And, of course, AWS wouldn't just do that without validating our solution and vouching for us. So that's definitely a great way to improve the brand awareness and recognition.

Jeremy: Right. Yeah. And I guess this probably ties in, because I think definitely the route that you went with being an Amazon partner or part of the Amazon partner network or APN that was a very, very wise strategic move. Because the other thing, I mean, this is just other challenges I think about. Like if I'm starting a technical company, one of the things that you probably need to do if you're dealing with the cloud and you're dealing with serverless, because you're likely not building your own serverless solution, you're building something on top of... Sort of standing on the shoulder of giants, as they say.

So you might be building something on top of AWS, which means you might need to access data within their AWS Cloud, or you may ask them to create roles or security tokens that allow you to access their systems and manipulate things, maybe even deploy things to their systems. So that's another thing like, and again, maybe APN is part of it, that the partner network is part of it, but how do you build that trust for one? And then maybe to extend that even further, how do you deal with people that are asking for you to be compliant and have all these global compliance issues that would probably pop up, especially if you're dealing with something like enterprises?

Nofar: Yeah. I think that's in order to be enterprise ready you have to comply and achieve all sort of certifications like HIPAA and ISO and SOC 2 and so on. So that's definitely needed. I don't know an enterprise that would work with you without these kind of certificates. And of course, meeting the AWS best practices and well architected that's another thing that's really helped us to improve the way that we operate.

Jeremy: Yeah. Yeah, no, I agree. And that's the thing too, you're right with enterprises it would be really hard for them to say, "We're not PCI compliant. We're not... " And again, if you use a lot of those other tools, you leverage the existing services that AWS has or Microsoft Cloud or Microsoft Azure, some of those other companies have, then you get a lot further than you would if you were just building through your own data center. But still, definitely a lot to think about there. So, maybe we could talk about the partner thing for a minute, because that was, I think a really good and brilliant move by Epsagon early on is to say, "We want to be part of this larger network so that we get our name injected here and that we get these certifications. So that when it comes to asking, "Who do we choose for this?" that at least your name is on the list, whether it's at the top of the list or not, at least it's on that list. So how did that happen? How did you go about that? I get the thinking behind it, but what was sort of the process that you used to get to that point?

Nofar: Yeah, I think that's what ... The process was that I joined Epsagon, we were like three months before our product was GA, and I was thinking, "Okay, how can we get more exposure? How can we get customers?" And I was starting to think like, "Okay, if we're going to approach different companies they don't know us, right?" So that's really difficult to convince them, "Hey, you should listen to us. Hey, we have an awesome product." Then I just tried to think my potential customer. Who would have these kinds of customers in common? Who would have the same customers that I'm looking to have? So, of course the first name that pop up to my head is AWS. Because before we expanded our offering to Kubernetes and containers and to Microsoft as well so we supported AWS Lambda. So naturally our customers would be AWS customers, right?

Jeremy: Right.

Nofar: So that was okay. That makes sense. But then, okay how do I start to actually partner with a giant like AWS? I mean, how do I navigate these forests? And I just started to see what kind of programs I can leverage and try to also, I think the main thing that I did was to connect with the product teams there, really to understand if they do see the value in Epsagon, and really to build a "better together" story as really Epsagon as enabler of adopting serverless applications. So I think that as enabler we were able to see the commitment from AWS as well and support that helped us through our journey and growing our business in the first years.

Jeremy: Yeah, no, and I mean, and that was one thing too, that was very interesting from the beginning. And I know AWS provided a lot of support for monitoring companies and for that being able to do distributed tracing and things like that. Because AWS just didn't do an excellent job at that and there's still a lot of gaps in their solution. So, having the ability to sort of expand that to their partners I think was very helpful for them as well. So if you can ever find a way that you can help AWS it seems as though that many of these cloud providers will take advantage of your offer to help and being able to kind of fill some of those gaps is definitely a way to do it.

Nofar: Yeah. And it's a bit tricky as well because lots of companies will have a product that most likely will compete with a product that AWS has, because AWS has lots of products. So I wouldn't worry too much about that, because in the end of the day you need to understand what is the focus of AWS. For us, it was okay, the focus of AWS is the monitoring piece or the serverless computing piece. And obviously it's the serverless computing piece so if they do provide some sort of monitoring, we definitely can provide a solution that can complement this solution and not necessarily compete. So that's definitely something that I recommend to take in mind because lots of founders have like, "Oh, I don't want them to steal my idea. I don't want them to think that I'm competing with them." I wouldn't worry too much about that.

Jeremy: Right. Yeah. And actually that is a great point. And I'm glad you brought that up because this is one of those things where I've talked to so many... I mean, I've been doing this for a very, very, very long time. So I see all these new things pop up on Product Hunt or whatever. And I'm like none of these are new ideas. It's just somebody figured out a way how to execute them. And that's one of those things where like, if you have a good idea, you know everyone has good ideas. I'm sure a hundred other people have the same exact idea as you it's just whether or not you can execute it. And that's the other good point about just about dealing with or competing with AWS. AWS is likely going to build a solution so that their customers have a one-stop shop to sort of solve maybe 80% of the problem. But for most customers, there's always a better solution out there and so I do think, and maybe this again, I think you maybe have already said this, but your advice is if you think you're competing with AWS, try to figure out exactly why AWS might be competing with you. And if it's something you can compete with them on which is again, better developer experience, just better tooling, easier to use things like that. Your advice is go for it, right?

Nofar: Yeah. It really depends on the product, but most cases, yeah, I would say go for it. Because in the end of the day you're going to do it much better because this is your entire focus, right? So you really need to ask yourself, is that the focus of AWS or not?

Jeremy: Right. Yeah. No great advice. All right. So speaking of advice, here's what I'd love to get from you. So anybody who is ... Because first of all, I want the market to be flooded with serverless products, right? Deployment engines or monitoring solutions or whatever they are. I mean, maybe not monitoring solutions, we don't want to get in Epsagon's way but you know just things that ... More solutions out there. Because I think the more tooling you have, the more product offerings you have. Now, again, not all of them are going to succeed but the more that you have out there, it just sort of "the rising tide raises all ships" sort of thing. Like the more people hear about it, the more these solutions advertise, the more people figure out why serverless is amazing and hopefully more people adopt it. Which again, start buying the tools and then it keeps raising those ships. So, what would your advice be to people who are just getting started? What's the first step? So I just came up with an idea or I'm in the process of building this product right now. What's the first step that I do to try to start up my sales cycle and get some adoption of my product?

Nofar: So I think collaborating with partners would be something that I would do right at the beginning, identifying the right partners that will help you to get in there, really put the foot in the door in companies that you probably would not have access to, especially the enterprise customers. Because startups would probably adopt your solution even if they don't really know the company. And even if they say, "Okay, maybe they will not be here tomorrow, but I don't mind. It's not such a huge commitment for me." But if you're looking at the enterprise customers, I think this is where you would need a partner to get in, especially the beginning. And really identifying who's this partner? I mean, it can be a cloud vendor, it can be an ISV, really depends on the product that you're selling and trying to think of okay, let's think about my next customers, which kind of tools are they going to use? Which kind of vendors are they going to use? And really to start crafting your partnership strategy.

Jeremy: Yeah, no, and I think a big part of this too, just my own advice, is relationships. And I know that's probably a little cliched to say, but you start making friends, go to these conferences, whether they're online or whatever, start following the people who are in the space, get involved in the community, write some blog posts. Again, we talked about that earlier, but once you start making friends and you start making contacts within the industry, the other thing is, is that you can leverage those partners pretty extensively. I mean, if you find a consultant for example, that you can sort of sell on your product and show them why your product is amazing or whatever. Then hopefully for their consulting clients, they would say, "There's this product that I use or that I really like and that's something that I want to use." So again, value of relationships, again, we're both not salespeople so it's sort of hard to give you overall sales advice, but I guess, development or business development strategy advice, I think the idea of partnerships and just making connections is a huge deal.

Nofar: Definitely, definitely. And always think how you bring value rather than, okay, how do I sell my product? Right.

Jeremy: Right, yes.

Nofar: About like a year ago or so I was thinking, okay, how can I help customers? If I was a customer that is just starting using serverless what would I need to know? And I came up with the idea, okay, let's do a session about best practices. And then I talked to Stackery back then, and I talked to PureSec back then providing security for serverless and planning for serverless and said, "Okay, let's do an event with 45 minutes about monitoring, deployment, and security. And let's see if that will be interesting for serverless customers." And we've done that with AWS and all of a sudden we had more than 100 people registering to the event and we had to stop the registration. And you've seen that, okay, there is a demand for this kind of content. So that was a great way to understand how you can actually bring value to the customers.

Jeremy: Right. Right. Yeah. And I think that's one thing about content creation that is kind of funny. I think we get into this thing where like we discover something really interesting, some really cool way to do X, Y, Z and then we write a blog post about that. And for people who are in the ecosystem who are already using serverless, they might come across it. Maybe it's a problem they're having, and they find it. But from a more global education standpoint, there are a lot of really good, but also a lot of really bad "what is serverless" posts and videos out there. So, yeah having a nice sort of like holistic overview of the whole thing I think is super important. So one thing that I actually wanted to ask you about and this just again goes into, I think, setting expectations.

So it's a long journey, right? If you're building a serverless product, if you're building any product don't expect to be financially independent overnight and be buying a private island. It's going to take a long time. It's a long journey. You've really got to work at it. You've got to make those partnerships. You've got to put the work in, you've got to do that, to grow those relationships and grow those customers. But just maybe another expectation that I think maybe somebody who's like, "Oh, well, I'll just sell this to enterprises." An expectation that not everybody has or an expectation maybe they do have is that it's like selling to anybody else, but enterprise sales are a long and sometimes painful slog. Sometimes it goes faster, but I'd love to just sort of talk about the enterprise sales cycle for a service like this, especially a software service and what that typically looks like, and maybe from your experience, how long that typically takes?

Nofar: Yeah, definitely enterprise sales is a bit more tricky than I'd say a startup sale, obviously. The challenges are different, it takes a lot of time from one call to another, to a PLC to someone to sign the contract it takes a lot of time. In some companies you need to be onboarded as a new vendor then that can take like another nine to 12 months or so. So you need to be patient and you need to find your champion inside the organization, that's for sure. You want someone to be your advocates internally there and I think that another thing that you can definitely leverage, something that we need was to use marketplaces. So if this customer is already paying to Microsoft or paying to AWS, you can save a lot of time instead of onboarding, to be onboarded as a vendor, you can just leverage one of these kinds of marketplaces. And for us, it saved a lot of time.

We had one particular customer that said to me, the champion there told me, "If it wasn't for the marketplace, we probably would not move forward with you because it would take about 12 months to move forward with you. And I would probably just ditch you guys, sorry." So we were able to move pretty fast. And when I say fast, I mean about two months, three months for an enterprise deal, I'd say. And it can be longer ...

Jeremy: Only two.

Nofar: ... only, yeah. It can be longer than that. Specifically for security tools it can be like six months a year. But for us, I think the average would be a few months, I'd say the longest was I think five months, if I'm not mistaken. It's something like that.

Jeremy: Yeah. No, because that's just one of those things where there's a whole process to it, and once it's sort of ... And it depends on how big your tool is, what it is. I mean, again, some companies have the ability to just sort of make those self-service purchases and then those get boiled up. But if you're adopting for an entire company, you know depending on what it is you're building it can get really complex. But you mentioned the idea of identifying a champion. This goes back to relationships, right? You find one person that sees you speak at a conference, or just happens to read one of your blog posts and comments on your blog posts and you can start a back and forth with them. Having somebody in a company though that is advocating for you is a hugely important step because enterprises have no problems sort of pumping the brakes on things and slowing things down if somebody internally is not pushing on the gas.

Again, I'm not criticizing enterprises at all because they are big, heavy machines that have a lot of moving parts that make it sort of difficult to navigate sometimes. So having that champion I think is hugely important. And then the other thing you said, this is another really important thing, is onboarding, like how easy is it to get your solution into that company? And I know that when Epsagon first launched and other monitoring tools, Lambda was limited, there was no such thing as layers, right? You didn't have custom run times, you didn't have the new extensions API and things like that.

Jeremy: Which meant that in order for you to instrument your functions, you had to actually put code within the function itself. You had to put a wrapper in there. And some of that got easier over time I know the serverless framework did something where they would automatically wrap your functions and things like that, which was great. But then it went to just layers. So you add a layer and then suddenly you get this capability and now with the extensions API it's even more amazing. But yeah, I think that is a huge thing to think about is to say, how hard is it for me to enable or turn my product on for a customer?

Nofar: Yeah, definitely. I think that for every partnership, I always divide it into three: co-sell, co-market and co-develop. So there are lots of things you can do under the co-develop category where you actually can bring more value to your customers and extension and Lambda Layers are great examples for that. I remember that when we did the Lambda Layers integration and all of a sudden the number of users that actually completed the onboarding was just amazing. It went up within a week just because it was so much easier for them to onboard to the platform. So those kinds of integrations are super valuable if it helps you really to cut the onboarding time significantly. So that's another thing that a partner can really help you to achieve while working with serverless customers.

Jeremy: Yeah, no, I love it. That's just good advice there so pay attention. No, I mean, if you're building a business that is absolutely huge. So all right. While I have you for a few more minutes, I have a couple of questions that we actually got from the listeners. And if you are a Serverless Chats insider, you can go to serverlesschats.com/insiders, and you can ask questions of our guests and we will try to read them and get the insights of brilliant guests like Nofar here. And one of the questions that we got was about just whether or not serverless might've been a trend? And several years ago, I know you joined a few months before Epsagon went GA, but how did you deal with the possibility that serverless might just be a trend? I mean, how did Epsagon deal with this maybe? That you say, "Well, what if this goes away in three years?" I mean, you made a massive commitment to back serverless.

Nofar: Right. Yeah. It is definitely seems like a trend sometimes where every single company is saying that hey, we support serverless as well. And in the end of the day, when you're starting to use the product you understand that the support for serverless is very, very limited. But I think that's the way that you can use serverless technology and the value it brings to customers. I think that's so significant that more and more companies will definitely adopt serverless and as we move forward and as time pass, more and more companies get more expertise, you see more and more jobs, where the description is saying, "Hey, we need expertise in serverless."

So I'm really happy to see that because you see that the market is going to this direction. I also had a conversation with one of our customers that said that they actually, one of the reasons that they migrated to serverless it was because they wanted to attract brilliant developers to the company. So that was an awesome thing to hear. So I do think it's here to stay and it's here to grow hopefully significantly. But I do understand why it might seems like a trend at least at the beginning where you hear serverless in every direction.

Jeremy: Right. Yeah. I mean, and that was funny too. The first time I sort of came across it was with Lambda in early 2015. And I remember just thinking when I first saw it though, I was like this is something. There's something there, there and it was just one of those things where for me I was like, "There's no way this can be a trend because if this works the way that it's supposed to work and it continues to develop out, I just feel like this could be the way that we just build things from now on or at least how we build things in the cloud." All right. So another really interesting question here was, if you go back to the origins of Epsagon why did you decide that Lambda was the right target to start with?

Nofar: I think that Ran and Nitzan the founders of Epsagon they a very strong technical background. Both of them were in one of the elite intelligence units in Israel, in the IDF. So, Ran is really very geeky, were one of these guys that started to code when they were four or something like that. And both of them recognize that there is something new, serverless, and it's actually very, very different to understand what's going on in your application when you are developing in serverless. You know the most natural thing to understand how different events are correlated that's something that is very difficult to achieve. And you have to do a lot of manual work to have these kinds of understanding. So they investigated and saw this problem also in container-based applications and microservices in general, but they saw that in serverless, it's so much more painful because everything is managed and you don't really have control and cannot simply install an agent and see what's going on.

So this is where they started to play and in the end of the day, they built Epsagon entirely on serverless. So they definitely understood the problem is as they moved forward and say, "Okay, I think we know how to solve it." And today we're actually I'm happy to update that we also have a patent around how we actually provided capability on the tracing, a U.S. patent on that. So that's huge, right? So you see that there is a way that you can actually make things much more easier to manage and you can adopt serverless much easily with these kinds of tools.

Jeremy: Yeah, no, that's a good point. All right. So last question, if you could do it all over again, and I don't know if you can answer this question, but if you could do it all over again, would you stay with initial targeting for serverless and then expand to containers, or would you maybe have done it the other way around or done both at the same time? How would you have done that if you were able to go back in time and do it again?

Nofar: I think it's a good question, because in the end of the day you do want to have one product for more than applications. So whether you're starting with serverless and containers, I don't think it's necessarily make much of a difference if in the end of the day, you do provide a solution that is a one-stop shop for more than applications. So even if we would start with containers and move to serverless, then I don't think it would make much of a difference. And we've seen that pretty early back then when we launched our serverless product, we immediately saw that there is a need for Kubernetes, for ECS, for microservices in general. And, we just implemented the same technology on these kind of environments and launched another product really to provide a complete solution for these kind of environments.

Jeremy: Awesome. All right. Well, Nofar thank you so much for joining me and sharing all this information. If you are thinking about starting a serverless company that is targeting serverless users, hopefully you got a lot of value out of this. But there was a lot of good advice, a lot to unpack there if you are doing this. So, if our listeners want to maybe reach out to you, Nofar, and maybe pick your brain a little bit more, or they want to find out more about Epsagon how do they do that?

Nofar: LinkedIn or Twitter just ping me and I'll be happy to chat.

Jeremy: Awesome. All right. And then you can go to epsagon.com and I will put your Twitter. You also have a blog nofar-asselman.medium.com. So I will put all of that in the show notes. Thanks again, Nofar.

Nofar: Thank you, Jeremy. That was fun. See you.

View Details

About Our Guests

Gillian Armstrong

Gillian Armstrong is a Solutions Engineer at Liberty IT and an AWS Machine Learning Hero.

  • Twitter: @virtualgill
  • LinkedIn: https://www.linkedin.com/in/gillian-armstrong/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/gillian-armstrong/

Luca Bianchi
Luca Bianchi is the CTO of Neosperience, co-founder and co-organizer of many serverless community events in Italy, and an AWS Serverless Hero.

  • Twitter: @bianchiluca
  • LinkedIn: https://www.linkedin.com/in/lucabianchipavia/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/luca-bianchi/

Sheen Brisals

Sheen Brisals is the Senior Engineering Manager at The LEGO Group, a serverless speaker, and an AWS Serverless Hero.

  • Twitter: @sheenbrisals
  • LinkedIn: https://www.linkedin.com/in/sheen-brisals/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/sheen-brisals/

Farrah CampbellFarrah Campbell is the Alliances & Ecosystem Director at Stackery, a co-organizer of ServerlessDays Virtual as well as several other serverless community events, and an AWS Serverless Hero.

  • Twitter: @FarrahC32
  • LinkedIn: https://www.linkedin.com/in/farrahcampbell/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/farrah-campbell/

Serhat Can

Serhat Can is the Technical Evangelist at Atlassian and an AWS Community Hero

  • Twitter: @srhtcn
  • LinkedIn: https://www.linkedin.com/in/serhatcan/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/serhat-can/

Yan Cui
Yan Cui is an Independent Consultant, Developer advocate at Lumigo, Host of the Real World Serverless podcast, theburningmonk, and an AWS Serverless Hero.

  • Twitter: @theburningmonk
  • LinkedIn: https://www.linkedin.com/in/theburningmonk/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/yan-cui/

Ben Ellerby

Ben Ellerby is the VP of Engineering at Theodo, Editor of Serverless Transformation, and an AWS Serverless Hero.

  • Twitter: @EllerbyBen
  • LinkedIn: https://www.linkedin.com/in/benjaminellerby/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/ben-ellerby/

Ran Ribenzaft

Ran Ribenzaft is the CTO at Epsagon and an AWS Serverless Hero

  • Twitter: @ranrib
  • LinkedIn: https://www.linkedin.com/in/ran-ribenzaft/
  • AWS Heroes Page: https://aws.amazon.com/developer/community/heroes/ran-ribenzaft/

Watch this episode on YouTube: https://youtu.be/tENIWFp3uj8

This week's episode is sponsored by New Relic and Epsagon.

Transcript

Jeremy: Hi everyone, I'm Jeremy Daly, and this is Serverless Chats. Today we have an absolutely amazing episode for you. re:Invent 2020 is finally done. We're in February of 2021 and unless they decide to maybe stick another week of videos in, re:Invent 2020 is finally done. So, I figured why not do an episode where we can get the best input and the best insights from some of the most amazing people in serverless.

So, today, I have eight AWS heroes with me, and we're going to talk about all the amazing things that happened at re:Invent 2020. So, I'm just going to go quickly around the horn here and introduce everybody. So, first up, is AWS Serverless Hero, Independent Consultant, Developer Advocate at Lumigo, host of the "Real World Serverless Podcast," The Burning Monk himself, Mr. Yan Cui.

Yan: Hey guys.

Jeremy: All right, next we have an AWS Community Hero, he's a Technical Evangelist at Atlassian, and the guy who once helped me stalk Werner Vogels, just so we could get a photo with him, Mr. Serhat Can.

Serhat: Hey folks, happy to be here.

Jeremy: And, next, is another AWS Serverless Hero, he's the CTO of Neosperience, co-founder and co-organizer of just about every Serverless community event in Italy, the Italian Stallion of serverless, Mr. Luca Bianchi.

Luca: Hello.

Jeremy: All right, next, we are joined by yet another AWS Serverless Hero, he's also the CTO at Epsagon, and way too smart for his own good, the man they call Mr. Obervability, Ran Ribenzaft.

Ran: Hey, everyone.

Jeremy: All right, moving on, we have another AWS Serverless hero, he's the VP of Engineering at Theodo, editor of "Serverless Transformation," and an excellent Nashville, Tennessee, drinking buddy, Ben Ellerby.

Ben: Hey Jeremy, thanks for having me.

Jeremy: All right, next up, another AWS Serverless Hero, she is the Alliances and Ecosystem Director at Stackery, a co-organizer of Serverless Days Virtual, and several other Serverless Community events, my good friend, the amazing, Farrah Campbell.

Farrah: Hi everybody, thanks for having me, Jeremy.

Jeremy: All right, next we have an AWS Serverless Hero, and Senior Engineering Manager at The Lego Group, he's a serverless speaker, an amazing writer, and all-around awesome guy, Mr. Sheen Brisals.

Sheen: Hey, everyone. Thank you, Jeremy.

Jeremy: And, finally, we have an AWS Machine Learning Hero, to round out the panel. She's a Solutions Engineer at Liberty IT, and co-conspirator in the Werner stalking incident with Serhat and me, the absolutely brilliant, Gillian Armstrong.

Gillian: Thanks, Jeremy, it's good to be here.

Jeremy: All right, so, we have eight amazing people right now, all kinds of knowledge that they can drop. So, what we're going to do is, we're going to go through each person. I'm going to just give you the floor, I want you to introduce something interesting from AWS. Whether AWS re:Invent, whether it was an announcement, or something that happened, just tell me about it. So, let's start with... let's start with, Yan. What's your favorite thing that happened at, re:Invent?

Yan: Yeah, sure, I think probably the biggest one, the biggest announcement at re:Invent for me was the Aurora Serverless V2. Which is really a bit of marvel when it comes to engineering excellence. Some of the things you can do with serverless, Aurora Serverless V2, in terms of the instance scaling and really fine grain steps of how quickly and how much to scale up. It solves a lot of problems that Aurora Serverless V1 had, that people actually want to use in production. And, things just takes forever to scale up and you go in the big double increase in size. Which in terms of cost, it's probably not the most, I guess, the pay as you go, because you're not doubling your costs straight away.

Jeremy: Right.

Yan: So, I'm quite excited about what they're going to do once it goes GA on the Aurora Serverless V2.

Jeremy: Yeah, and what do you think about the cost aspect of that? Because that was one of the things where you look at those ACUs and they've actually doubled the cost of those. But, do you think that just the benefit of having all those features and how quickly that's going to scale. Because it does, it scales instantaneously. Do you think that's going to be a good trade-off from a financial standpoint?

Yan: I think, I guess that we still have to see what happens once people start using it in production. But, because of the fact that they've reduced how, I guess the units that you scale up by. Because, before you have to scale from 4 to 8, to 16 to 32, and that means you end up spending a lot of money on ACU that you're not going to be using. Because it's just over that 8 ACU, and you just need 8.5, but yet that'd be 16. And, the fact that it takes longer for them to scale down, also means that you end up paying for overhead that you don't need for longer as well. So, besides all the extra features you get, the fact that it's more expensive per ACU doesn't necessarily mean that you end up paying more overall. Because you cut down a lot of the waste that you have.

So, I guess it's still wait, we still have to see in production, what actually transpires. But, given the fact that it now can scale up faster, which unlocks use cases that you couldn't do before. Before, it just takes too long to scale up. And, some extra features they talk about in terms of, was it? I've got some notes here, some nicer features. Like, right, so there's some stuff that was missing before, like a global database, IAM auth and the Lambda triggers, so all of those now are become available in the Aurora Serverless V2, where they were missing on the V1. So, between that and the fact that you, you cut out a lot of your waste, I suspect that it's going to make a lot more financial sense to go through V2 going forward.

Sheen: I'm curious by the way, why it's named V2, and not just, "Hey, we updated the Aurora and it's now better." Like, it's the first time I've seen that they nicknamed it V2. Is it something that happens behind? Or, like what's the reason?

Luca: Or, maybe Aurora Truly Serverless, or something like this.

Yan: I think that it's because it's a complete set of company offering. And, the fact that you can use V1 and V2 side-by-side. So, it's not a complete replacement of the Aurora Serverless, but it's a new, new imagining of how Aurora Serverless should work. Which is why it's V2. And, you can actually have the same cluster with both V1 and V2. I'm not sure how that's actually, how well that's going to work in production. In terms of capability, everything else, like okay, some stuff will work with the new features, like Lambda triggers and some won't. I'm not sure how that actually plays out in the real world, but at least you have that option of running the two versions side-by-side. Which, might make migration easier, so that you can gradually move stuff over, as opposed to you know, stop one day, the downtime, and then bring up a new V2 and then, see what happens.

Jeremy: Right, and thinking about some of those use cases, did anybody have any thoughts on maybe when people could say, "With this level of scalability now, what use cases could I maybe start migrating to serverless, that I couldn't before?"

Ran: I would think that the lowest scales would really be the Aurora Serverless, because honestly, you don't care about as much of the infrastructure, and you do care for every bit that you're paying. And, you do care about the instantaneous scalability. So, I think for the lower scales, it makes sense for them. On the other end, the more, bigger scales, billions or hundreds of billions of requests or rows, every day, it's something that probably, serverless, wouldn't be the best fit for. That's my opinion.

Yan: I think we have to factor into account how much engineering expertise you have to actually do that yourself. Because one thing that always gets left on the table is, "Well if I have to hire somebody to do this for me, to run on containers, that's going to cost me $20,000 a month, how much am I actually saving? If I save $10,000 a month on the AWS cost and spend double that on the staffing cost." So, again, if you've got expertise already, then yeah, I think you're absolutely right, but if you have to hire expertise from externally, then you have to factor into account the total cost ownership, not just, the cost of your AWS bill.

Ran: The other question is when should I choose RDS and when should I choose Aurora now?

Yan: Yeah.

Ran: Except for if I want an Oracle database or something which is traditionally old, but for MySQL or + works, why should I even bother with RDS anymore?

Jeremy: It's a good question, does anybody have an answer to that?

Yan: I mean, Aurora does give you some really amazing features, that you have to build yourself. Which are quite hard to configure and setup and all of that. Things like the global database, the fact that you've got Lambda triggers, you've got IM authentication, which I still have a problem with that, because you still need to have a root access. You still need to have a root user, so you still have that security overhead of having to maintain that. And, deal with that. So, but still, you've got a bunch of unique features that Aurora has over I guess a normal RDS? What do you call it? Just a non-Aurora RDS?

Jeremy: Right.

Yan: And Aurora has also got some pretty crazy performance, the last time I saw some benchmarks that you can ... Someone's able to get up to 10,000 ops per second, or something like that on Aurora, which is quite significant. It's not easy to get to that level of throughput on the RDS.

Jeremy: Right.

Serhat: I think if you look at it from a different angle. So, when Serverless Aurora V2 becomes popular, it will give you an extra option to DynamoDB, for those who are more to a SQL, and there are use cases now. They just stopped everything in Dynamo, because this the go-to option they have. With V2, that will kind of get loosened up. And, I mean, earlier this week an engineer came to me with that design. Where she wanted to keep a certain audit data in DynamoDB, but when I looked at the query that she had to perform on the data, I knew that it's too much for DynamoDB.

Jeremy: Right.

Serhat: So, these are our situations, not necessarily millions or billions of transactions and things. But, even that sort of flexibility and the ease of using, without compromising the performance. So, those aspects will come into play, eventually when it becomes swapped out.

Jeremy: Yeah, and I love using RDS, or even Aurora Serverless as one of the major ways that I've used it recently, was just to store data off of streams from DynamoDB as a way that I could query it. But, it was more for the admin side of things, and less for a sort of a front-end use case. But, yeah, I think that's a super interesting thing that you can do is say, "Look, if I can mix and match those two databases together, get the operational performance and that stability of Dynamo, but then also be able to expand my query capabilities with something like Postgres or MySQL," I think that's an interesting way to do it.

Ben: Especially, and I think with the glue elastic views, which came out at re:Invent, as well, we'll be able to have the same data replicated across multiple data stores. I've not played around with it yet, but I think we'll be able to have application data from DynamoDB replicated to relational database or other DynamoDB databases, but we'll be able to have data scientists interacting with those relational database directly. At the moment, I'm particularly using S3, as sort of Serverless Data Lake, and then using Athena to query on top of that. But, it's nowhere near as powerful as having a data scientist having direct access to a relational database.

Jeremy: Right, right.

Yan: There's also the fact that you can stream stuff from DynamoDB streams to Elasticsearch as well, which is such a common use case. I built this like three times in the last six months for different projects. And, it's just so much easier, if I can just point this elastic view to between DynamoDB and elastic search, that'd be awesome.

Jeremy: Yeah, the problem with Elasticsearch is that it's not Serverless yet, right? So, we're still managing those clusters. And, 99 percent of it's done for you, but you still have to sort of think about it. So, any other thoughts on Aurora Serverless V2?

Yan: I guess one thing that's worth mentioning is that, at least right now, before GA days, you don't have a data API yet. Which means, from a serverless point of you, you still have to worry about managing socket connections, pulling, all of that, and maybe bringing in a database proxy. But, I imagine, by the time they go GA, they should have added the support for the data API, because that was quite a big game-changer for, Aurora Serverless.

Jeremy: Right, definitely. All right, let's move on. Let's go to Luca, what was your favorite, re:Invent, announcement?

Luca: Yeah, my favorite announcement is the container image formal support for Lambda. Because, this is something that was very game-changing for machine learning practitioner, but also for a number of use cases. Because, now we have the possibility to drop and define all the dependencies of Lambda in a specific file format, or image, that can be deployed on the Lambda. It is not a direct support for container images. So, you cannot deploy containers, but you can describe in the same format everything that you want to be packaged inside with your Lambda, and this is great. It's great because it makes you able to go over a lot of constraints. It makes you able to implement tree-shaking or dependency removing. And, so you can optimize also the package that is deployed to Lambda. And, it's something that you could have done before, but you needed some kind of serverless-specific plugin, or to do this by yourself using CDK or whatever, and it was not straightforward. Right now, it's completely direct. And, you specify the format.

The other thing that is very important is that, choosing a standard format. We have a lot of Kubernetes of container dev ops engineer, that now can bring their workloads to Lambda using exactly the same format. Not every feature of the language of the format is supported. For example, you cannot open up a connection to all ports, to add services or manager connection with the image within the container, but the subset, which is supported, is something which is really great. Because you can run even runtime scripts on the packaging machine when you are wielding the package.

Moreover, the container support and everything that you can package to make your image working, and I'm referring about the Lambda runtime environment that can be packaged on the container makes you able to do two things. The first one is that you can choose your own image format, so you can start from any kind of available doc or image, you can choose Ubuntu, you can choose Redhat, you can choose Fedora, or whatever kind of image. And, you can package, you can start from say a python image. So, an image which has already a different kind of the python runtime, bundle with that or say some kind of Linux-based utilities, such as ImageMagick or FMPG, or whatsoever, so you can choose your base image, and you can bundle with that you runtime environment client. Which has been released open source, so you can also use that feature to test locally your development work cycle. Which, is great, because you can run a docker on your machine, deploy that image, make Lambda calls to that runtime environment. And then, test even before pushing back to the cloud.

Jeremy: Yeah, and I think that's a good distinction too, that it's a packaging format. You're not actually running a container in the Lambda function itself. But, yeah, I mean that opens up a whole bunch of tooling options. So, has anybody seen some feedback on how this helps people that have existing workflows start using serverless?

Luca: Yeah, I think so. And, I think that Gillian has some nice use cases because we have discussed about them before.

Gillian: Yeah, well, I'm mainly looking at that because we're doing machine learning. Obviously, the models are quite vague. Shoving them into traditional Lambda's not going to work out for you. So, having the extra space is really nice and it's definitely allowing us to do a bit of experimentation and see what we can do. The system I'm working on, it's got lots of different data extraction. So, we need different models for different times. We don't want them all running, like just burning cycles when they're not being used. Because some of them will be used very infrequently. So, being able to put those in a Lambda, it's a massive advantage. But then, I do have some concerns around the containerized Lambdas that you're going to see companies coming, and the first time anyone's using a Lambda, it's in a container and then they never come out of the container, instead of saying, "Well, if I can build this Lambda without the container, I'll start there." And, only move into a container when the use-case doesn't work for just a straightforward, simple Lambda, where you get a lot more just done for you. And, that can be right for some companies, but I'm wondering if some people might come from their container worlds, go into the container Lambdas, and never actually just go and build a plain Lambda, and that's actually all they need.

Yan: So, Gillian, I've got a question for you. Have you actually tried loading a large machine learning model in the container image of a Lambda function?

Luca: I tried to load an image model of more than 400 megabytes, and it's a good approach if you don't need to update the model. And, if you want to have some kind of immutable packaging. And, it works well, it requires a bit of time, when the engine is optimizing your container. Because, once you push, the workflows specifically that you push the image into an Amazon container, register it, and then you update the Lambda, or you create the Lambda refer pointing to that image in the container registry. And, when it happens, the Lambda starts pulling the image and optimizing the image, and that phase is dependent on the size of the package of the image.

Yan: So, my question for you guys, in that case, is have you measured the latency for reading that model?

Gillian: It's not fast.

Yan: I was talking to someone from IKEA.com, they had the exact use case, that Gillian, you're thinking about. They're trying to load a machine image, a container image, that's 1.5 gig, so they're trying to load a machine learning model that's 1.5 gig, and it took them 4 minutes to do that, in 100 megabyte chunks, which means that if you want to load 110 gigs, like full 10 gig container image, you won't have time. Lambda's going to time-out before you can load it. So, that's something that I'm trying to figure out, is that just something that they were running into, or is it just more of a platform problem that they haven't figured out yet.

And, the trade-off, they made the design for this container image is that, the container image itself, it doesn't matter how big it is, it gets broken up into small chunks, into sparse file system. So, that's how they're able to limit the cold start penalties for loading a container image. But then, that means, if it actually need to load lots of files, a large file from your container image, well good luck. It's going to be pulling small chunks from the whatever ... file power system.

Gillian: I want to talk to someone who gets 10 gig in their Lambda, I think you get a prize if you manage to get a 10 gig Lambda, and it's like successful and working well.

Jeremy: I'm also wondering, how does the increased cores, now that you can have 10 gigabytes of memory in there, does that help at all with loading these larger things? Or are you still just limited by the network?

Yan: Yeah, I do wonder that as well, because even before the 10 gig image, sorry, Lambda functions, you also had the full CPU when you're running the initialization as well. I don't know how well that applies to the container images as well. And, also, I guess one thing that's also worth mentioning with container image, is that now you are responsible for the security and the updates of, and patching of, the OS. Which is something that 95 percent of us don't want to do.

Jeremy: Right, and are we blurring the lines even more, I don't know, Farrah, with you, what's going on with Stackery? I know you're working with a lot of customers that are building using SAM, but also cloud formation and serverless framework, and some of these other things. Is this something you're seeing though, where people are super excited about it, because they think, "Hey, now I can just use containers on Lambda functions," or is this just getting too confusing?

Farrah: I think it's something that definitely excites people. I mean, I think what it does bring, is it provides the opportunity to be able to reuse images that you've already done that will validate, build, and deliver Lambda functions that previously, you would have to set up a whole tool chain for.

Jeremy: Right.

Farrah: I think it also helps, it's helping people, you don't have to jump into serverless head first. So, you can make these incremental approaches to starting to try to modernize your application. But, also, it's already fitting into, you see it really fitting into how people are already working. So, I think that we see Amazon really trying to figure out ways to integrate with tooling that's already there. With workflows and patterns that are already there. I mean, you see EventBridge has over 140 SaaS integrations now. The Lambda extensions did that API. I really just see ... while I think it's confusing, there's a lot of confusion, when you should use this or maybe where to use Fargate. How is this going to be filled? Does this support extensions? Does it support layers? So, I think there's still a lot of questions, but I do think it's really moving into, really trying to figure out, how do you help? Help developers with their current workflows. And, how do you help speed that along and make that a little more seamless.

Yan: It doesn't work with layers, I've checked. It doesn't work with layers, unfortunately.

Jeremy: It doesn't work with layers?

Yan: It doesn't.

Luca: Nope, not yet.

Jeremy: Well, I'm sure it will eventually, right?

Yan: Probably not, because the layers is a file system attaching to your file system that you already have.

Jeremy: That's a great, very good point.

Yan: But, if the file system is a container, then where are you going to attach it?

Jeremy: Right. Well, I think the other thing that's important to remember here too, is this is not ... Using a container as a packaging format is not an AWS innovation. IBM is already doing this, they're doing this as Azure. So, a lot of these other cloud providers have done this before, but it's certainly, as you said, Farrah, it certainly does help people sort of move in that direction. But, then, I also fear what Gillian said, that maybe people just get stuck in that. But, I think it's hard to fight gravity of the popularity of containers right now.

Yan: Yeah, and also I think that 90, 95 percent of people just don't need to use containers. It's like all of these new features they're adding. They keep adding EFS and extensions, all this other stuff. They're medicines, as opposed to just specific symptoms and problems. It's not something you should just go out every day.

Jeremy: Totally agree, totally agree. All right, so let's move on, Serhat, what was your, re:Invent announcement?

Serhat: So, one of my most favorite re:Invent announcements was one millisecond billing. And, I know a lot of people are really excited about this. And, I can't stress the importance of this change enough. This is really important, and when you look at it, AWS is probably going to lose a ton of money. Probably they lost a ton of money overnight. And, I know, from my friends, they save a lot of money overnight, and this shows how customer-focused AWS is. And, probably, it's not about just the money. Because, from our previous use cases, we had to run functions with more CPU and RAM, and then we are seeing like 100 milliseconds execution time, but we are paying for 100 milliseconds, now we don't have to. So, that means we can run our functions faster and cheaper.

So, that also means now people are thinking about moving to Lambda is going to cost them much more. And, they're going to lose some performance, now they can choose the highest memory if they want to. And, they're going to pay just the amount of execution time they spent. So, this enables a lot of more use cases. And, because now, there are more use cases you can run on Lambda, then in many cases you don't need another container management service, or EC2 or whatever to be able to run your whole services.

Because, there are definitely cases where you need to be fast, and then you start thinking about cold starts, cost issues, a lot of other things. And then, you start using containers, EKS, whatever, along with AWS Lambda, then your whole operation become a mess, right? So, that also means, it's not just about Lambdas getting cheaper, it's also about enabling more use cases.

Sheen: I'm curious, by the way, if there is any case that it doesn't become cheaper. Like, is there any mathematical behind this thing that I know it will get more expensive?

Jeremy: Maybe, only if you move more workloads on there. I think one of the things that I noticed with the one millisecond billing, is that it sounds really great in theory, and I am a huge fan of it. I think when we were out at AWS, maybe two years ago, I said, "Could you make it, maybe, a 50 millisecond?" Even that would be better than the 100 millisecond, and they went all the way down to one millisecond, which is great, but you do pay an invocation cost, right? So, even if your Lambda function is only running for 10 milliseconds, you are paying that invocation cost, so I wonder, if you're invoking more functions because now you don't have to worry about squeezing multiple operations into a single function, if that was how you're trying to do some sort of optimization, that if you're calling more functions that you are still, maybe, paying a little bit more because of those invocation costs.

Yan: I guess you could yeah, I guess if you're doing patching before, but now you're doing just one record at a time, you end up paying more for that 20 cents per million requests, as opposed to the whatever ... for millisecond billing.

Jeremy: I'm sure you could just turn your couch over and find some change that'll pay for that bill anyways, because it is so incredibly low. But yeah, no, I think for a lot of different things, they are, the 1 millisecond for me is, if you're doing some of those operations where, maybe you're polling, or you need to call an API for example. Being able to call an API, and having to pay an extra, maybe it takes 34 milliseconds to call the API, having to pay that extra whatever it is 66 milliseconds just seems like an excessive amount of time when you don't need to. And, if you do that millions and millions of times, it does start to add up.

Yan: I do think that, I guess in my worry for the millisecond billings, more that now there's more excuse for people to prematurely optimize because they want to cutdown 15 milliseconds of execution time. When, in fact, over the course of a month, they're paying 0.02 cents for the whole thing.

Jeremy: Right.

Yan: You're spending hundreds of dollars of engineering time on something you're never get your money back.

Jeremy: Right.

Yan: So, that's kind of more my concern about this millisecond billing. Before, there's no point, because whatever you do, you're going to pay 100 milliseconds anyway, but now that's their argument. "Yeah, we can save you some money."

Jeremy: So, Sheen, I know over at LEGO, your team has been a fan of using sort of the, not Lambda lifts, but sort of a fat Lambda, is how you referred to them in the past. Optimizing those, because they're doing multiple synchronous things together, have you seen a reduction in costs now that you're getting that one millisecond billing?

Sheen: Yes, I mean, things change. I was going to say actually, because I myself and many people said, "Oh, with millisecond billing, the Lambda functions are going to become single purpose, et cetera, et cetera." But, when you have a team of engineers doing Lambda functions all day, they don't really look at things that way, they just continue as usual. From the early days of having the sort of fat Lambda, or Lambda lift, I think that's changed. Now, it's more linear and single-purpose. But, even then, these days when you write a Lambda function, it's not just simply doing a few things and quitting. You have structure logging in place, you have bunch of conflict things getting loaded, bunch of things from parameters stored, and you have layers and this and that.

So, ultimately, there's so much on top of a simple Lambda function, so, that's the other side of this argument. Yes, it does benefit, but I don't think engineers are looking at that way in their day-to-day double-upping Lambda functions, in my opinion.

Jeremy: Right, and I think actually what Yan said too, about prematurely optimizing. I think most developers aren't thinking about costs still, even as we move to this serverless world, people are just kind of building their applications, and when it comes to speed, like maybe getting the latency down, and things like that, those are decisions they'll make. But saying, "Oh, well, we want to shave 8 cents off of our monthly Lambda bill." I think that's still outside the normal view of your average developer right now.

Sheen: Yeah, the counter to the optimizing, pretty much the optimization is like a ... Here's another thought now, because it costs less, why don't we up the RAM a bit, to get a bit more performance? So, that means, they will end up paying more, yeah? So, there's always these two sides to this.

Jeremy: Yeah, well, and I think the other thing too is that, like Alex Casalboni's optimizer tool, it just got a lot more complicated, because there's so many different options. But, all right, any other thoughts on the one millisecond billing?

Ben: Yeah, I'd just like to add to that actually, because our development teams are a bit weird, Jeremy. We do put a focus on reading our billing dates every week. More from an application understanding point of view, and to validate as things scale, the cost is still going to be in hand. So, all of our teams get this every sprint, as part of the scrum process. So, in the review they look at the AWS bill and make decisions based off that. And, for us, Lambda is never top of the list. It's optimizing things like API gateway, it's probably a better use of time. Although, Lambda, can be higher costs, depending on your use case.

Sheen: That's good to know, Ben, because I think I can take that back to our teams. I think that's very important, because often engineering teams, they never get to see the production billing or anything. That's something that's really useful, yeah.

Yan: Yeah, I think in practice, I see most people spent more money on the things like Gateway and CloudWatch. CloudWatch and x-ray stuff, well maybe not x-ray, but definitely CloudWatch, and anything the CloudWatch is usually pretty high up in the list of things that cost a lot of money.

Jeremy: All right, awesome, so, all right, let's go to, Gillian. What was your favorite announcement?

Gillian: Well, there's lots of cool announcements, but step functions are definitely my favorite serverless orchestrator. I use them for lots of different things. And, although, Lambda is the duct tape that you can use to fix any problem. You can stick it between anything and get anything stuck together, I do like to see things simplifying. So, seeing things like that synchronous express workflows. Seeing things like being able to automatically straight from an API gateway to a step function, or straight to API gateway from a step functions. And, I know you can get a nice circular thing going on there. So, being able to not having to put Lambda in between, obviously, you could've used a Lambda to call API gateway from a step function, but now if you can put it straight in, best code is no code. So, being able to just really simplify what you're putting together, simplify workflows, make your applications much, much easier, and much less code and much less phases, I think that's pretty cool.

Jeremy: Right, and as they always say, the most dangerous part of your application is the code that you write. Right? Like, everything else is sort of battle-tested, is there. And, that synchronous workflow ties in very nicely, I think, with the 1 millisecond billing. Because, now with the synchronous express, you can do that function composition, and you can actually have several Lambda functions that run back-to-back-to-back-to-back as part of a synchronous workflow, and now you're not paying 100 milliseconds every time. Or that, exorbitant transition fee that you do with normal step functions.

Yan: I think you're going to have a much bigger problem if you do that with cold starts though because the idea of using synchronous express workflows is attach them to the API gateway and stuff like that. If there's an API and it's user-facing, you definitely don't want to have like 5 LAMBA functions cold starting one after another. That's not going to be great for your user experience.

Jeremy: That's probably true. I'm wondering though if it's one of those things where it's like regular Lambda functions, once you get them warmed up, does it allow you to do certain things? But, also, even if you have, I guess, well, I guess if it's happening behind the scenes asynchronously, then asynchronous wouldn't be that big of a deal. But, if you potentially do need to have multiple things, multiple APIs called to bring back single response or something like that. I don't know, I think that if they can speed it up, they could be the solution to the function composition problem.

Yan: What I'd like to see is that they introduce some kind of a scripting language. Some kind of a retail, well maybe not retail, because that one is retail ... Some kind of a templating thing, where they can essentially just execute a script without it being a separate Lambda function. So, that would remove a lot, I guess, performance concerns that I would have. Like, I said, cold starts may not be that big an issue, but it does make your worst-case performance a lot worse when you've got them stacking up after one another. Because it's all on the same workflow.

Gillian: It'd be great if it just warmed everything up at the start. And know all those Lambda were in the step functions, just warmed them up right away.

Jeremy: Right.

Yan: I mean, you could provisional currency, but then that becomes the interesting cost-wise, it becomes another thing you've got to worry about.

Jeremy: Right, and expensive. And, actually, I think that's actually a good point though, just this idea, and maybe going back to, Sheen, what you were saying about the fat Lambda, and some of these other things. With the single-purpose function, I love single-purpose functions, I think it makes a ton of sense, but then on the other side of things, if you do have multiple steps that have to happen, having those all run at a single Lambda function, sometimes makes a lot of sense too. So, I think AWS is pushing people towards individual Lambda functions. And, look, having a Lambda function that does one simple thing is great, because you can compose those, they can be reused and things like that. But, without being really well-coordinated with step functions and understanding how all that stuff works.

And then, on the other side, like you said, Yan, paying that penalty of cold starts, if that's not a solution that hopefully gets better over time. So, I guess, maybe, a question for everybody, where are you? Are we still on the fat Lambda if need-be camp? Or, have we all moved towards the single-purpose?

Luca: I've seen two different approaches from people. On one side, you have people with the power refraining from serverless, preferring to embrace managed services and whatsoever, and they tend to embrace fat Lambda or Lambda leads as much as they can. On the other side, you have a lot of people that are enthusiastic from the possibility to package varied bunches, small pieces of code within a Lambda, and they would Lambda-ize everything. And, we are trying to have some difficulty in tying the balance between them, because my problem is not about having a huge Lambda or having a fat Lambda, but is ready to the fact that it pushes developer to adopt some kind of bad practices about ... okay, having just one crude service with everything inside, and you package everything within the Lambda, and maybe you just use 3 meg of that server, but it's super simple to package everything within your Lambda and who cares about that?

But, when you go into production and you measure cold start, you are hitting hard in the head by the cold starts. And, it's something about already too the behavior of the developer. And, I think I more shifted about having smaller Lambdas, because it encourages you to adopt some things like good patterns and good architectures, but it's not a dogma. Rather something that I like more.

Jeremy: Well, I'm curious to get your perspective, Ran, I mean you're the one who's looking at these functions being run. Is it easier to observe a single-purpose function, where only one thing's happening. Or trying to parse through those stack traces when you get errors on the Lambda list?

Ran: It's some and some for that. Is there use-cases for the simple exception that you're having, so having a monolithic Lambda is very easy to troubleshoot, because all-encompassing single location. You can see everything from the beginning all the way to the end. I mean, that's kind of single-node, but on the other side, when it's getting complex, when you're having pipelines, of 3, 4, 10 functions with different services, different third-parties and API calls, these problems tend to be more complex. And, you kind of want to see just an encapsulated problem. Like, the call to strike was earnest or something in the build was changed throughout the course of time spent. This thing that you can't see in a monolithic, because all the logic is internally. Everything happens internally in a single function. There is no outbound call that says, "Okay, I analyzed some data, I'm passing it to the other service that's responsible to charge that specific user. But, everything happens internally.

So, the upside to having more microservices approach, or more event-driven, and breaking the functions into smaller pieces is that if you're having the right tool, you will have greater visibility into what happens. Because, you know that at this point when it tried to charge the user, the input was X and something was missing from this specific input. Unlike a monolithic, by the way, at Epsagon we're having both cases. We're still stuck with some Flask data function that is having, kind of our fat Lambda, but except for that we got, I don't know, 600, 700 functions, different services, all try to be as small as possible. Someday, we will migrate our fat Lambda.

Yan: I think there's still limitations, things like I find with Kinesis. That's probably what Kinesis had done to these streams. That's the one place where I can't really quite follow single-responsibility functions as much as I'd like. Because of the fact that you have to contend with constraints on how many subscribers you can have. And, at the same time, there's no filtering, so you end up having to do a lot of filtering your own code, if you want to be handling just one type of event, when you've got an event process that funnels everything to you in one stream. So, I think, besides that, most other cases I found single-purpose functions have definitely been ideal. I've seen some clients that just go fat functions, I tell you, like, Ran, said, old errors you see happen in one function, you have no idea what happened to be going on.

Jeremy: Lot of console.logs.

Yan: Yeah, everything's in that one log. All the layers, one function. Everything got one message.

Sheen: I agree with what Ran was saying, because initially many teams, when they start, they want to put everything together, cost and so many things, but then, when it gets to production and running more, that's when the observability problem hits. They want to see what's going on, that's when the reality hits and they wish everything was grander, and it's latent, so it'll get more visibility. And, that is the case actually, because when you have production environment, you need to know what's going on. Especially now with all the different features now we have. We have all the destinations and keeping track of ERS and things like that, so, yeah that's an important thing.

Jeremy: Right.

Yan: And, it does an example from Lucas's example as well. We have got a client that had this one API that's doing one import, it's doing the service rendering, everything else just a simple get and put from DynamoDB corrects the stuff. Every single function, every single cold start is at 1.5 seconds, because of the React, because of that one endpoint that the services are rendering. So, that's where that one function is easy, but then you end up paying the cold start for every single endpoint.

Jeremy: Right, and I'm curious from your perspective too, Farrah, because I know you're working with a lot of companies that are doing this stuff. Is that something that you're seeing as well? Is it sort of a mix and match of the single-purpose versus the Lambda-lith/fat Lambda people? And, I don't want to fat-Lambda shame anybody, I certainly don't want to do that. So, but I mean, again, I use them sometimes too, but I'm just curious, what's your experience?

Farrah: Yeah, it think we definitely see a mix of both, but I think the goal for our companies, people want their environments to become more flexible. And, that's the whole goal of modernizing. And, if you have a big fat Lambda, is your architecture flexible? I might have to say, it probably is not. So, I think we really try to work towards, I'd say, more single-purpose functions. But, definitely, you see a combination of both.

Jeremy: Awesome, all right, so let's move on. Sheen, what was your exciting announcement from, re:Invent 2020?

Sheen: A few things, but one of my favorites with EventBridge, now we can archive and replay events. And, I'll tell you a reason why I like this. Because, EventBridge, we started to adopt EventBridge as soon as it came out. And then, at two or three occasions, in different use cases, when I spoke to teams to use EventBridge, there were a resistance. The simple reason being, what happens if I lose an event? What do I do? Especially when it comes to critical events that say, carries customer order data, or payments details, and things like that. So, in such scenarios, situations, you can't just go without any proof. So, I had to back off in those situations and use EventBridge in other scenarios. But, with the archive and replay mechanism, we get a sort of a confidence.

So, okay, you have your events here, if something gone wrong, or you are the consumer of a target Lambda, for example, have issues, it gives us the flexibility to replay those events as, and when, we need. So, that's an important thing which was missing until the announcement. Now, I had a brief look at the archive and replay setup. It's not completely clear to someone who is coming in. Because you may end up replaying, hitting all the targets, and you need to be careful, you are archiving for that particular pattern, or that rule that you have. Otherwise, it's not going to make much sense.

And, also, the other important thing many people miss is that when you replay the event, the event comes with an extra attribute, "Oh, I am a replay event," or something like that. So, that, again, is something important to look into when we build our patterns, different rules, and things like that. So, that's why this is one of my favorites, especially when it comes to EventBridge.

Jeremy: Yeah, I was of the very early adopter of EventBridge, and I remember the first thing I did was, you create a rule that captures every event, and you just send that somewhere so that you have that backup. And, so, adding this in, it probably seems like a relatively small thing, but it really does help. And then, like you said, the ability for you to have that little extra flag in there that says it's a replay event. That's super helpful, because if you're building in item potency, and some of the other things that you have to do when you maybe reprocess an event that already happened, that's a really good thing to have.

Now, Ben, I know you're a huge fan of EventBridge, as well, what are your thoughts on those new capabilities?

Ben: Sure, yeah, I mean, we're using EventBridge, on nearly all of our projects these days. And, actually, just a couple of days ago, an article went live on the AWS Game Tech blog, which talks about our use EventBridge in the E-Sports space. And that's, Gamercraft, is the company, so feel free to look at the use case. Archiving is something we've been doing ourselves for a while, so it's great to get that just done, out of the box. And, especially as a lot of our clients are in the regulate space. It's great to have things like, archive history out-of-the-box as well. As, we're starting to get things like, last year, the encryption at REST support, these can start to be used in more regulated industries. Replay's also great, and just before re:Invent, the instruction of retry policies, and dead letter queues, means we're getting a lot more robust than straight out-of-the-box.

We're already using, again, in a lot of our projects, and last thing I'm particularly using it for, but if you're using an event source space architecture, obviously archive and replay can be crucial parts for you.

Jeremy: Yeah, and also the thing that is getting more and more popular with just the way people are building their applications, is splitting up accounts. So, you have maybe a microservice, separate account for each microservice for example. And, you might have a separate account for each microservice, and then each stage of that microservice and things like that. So, cross-communication between accounts with EventBridge is kind of a complicated thing. I don't know if anybody watched, Steven Ledig's EvenBbridge talk during, re:Invent, but really, really, interesting and he just recently released some of those patterns on a GitHub repository, too. But, just thoughts on that, I mean people who've had experience with this, it's kind of clunky right now, but hopefully getting better.

Yan: It's gotten a lot better already. There's still some problems with it, like the fact that it still only delivers to the default bus on the destination account and stuff like that. But, at least you can now use resource policies to control which account can access, so you don't have to make a change on both the account where the bus is as well as the account where the destination is. Now, you can just do everything from the destination account, when you add a new subscriber, so that was quite nice change, and make things a lot easier for people doing that multi-account pattern.

I think, one thing I think, Sheen, you mentioned that you talk about the archive in replay, one thing that's missing is that when you replay, it just dumps as much events at you as quickly as they can. They don't respect the event ordering, so I was talking to the guys at Mahan, this big Swedish grocery shopping company. So, they build some tooling around EventBridge, and they actually build this EventBridge CLI2 they've got and they implemented with respecting the timestamps, so it gives them at the right time, as opposed to just, "Here you go, here's a million events, boom."

Gillian: We're looking at, EventBridge, definitely very interested in starting to use it more and more. So, I'll ask the people who are using it, so how do you feel about the observability? That does seem to be a little bit lacking in EventBridge. Logging, even with the new archive, I don't know if you can really use it to query, and find out what's happening, and if events haven't been picked up by any rules, you don't really know that they haven't.

Jeremy: If we only had someone here who knew about observability.

Serhat: Yep.

Jeremy: Oh, wait we do.

Yan: So support for EventBridge a while back, and I think Epsagon has support for it now, okay. So, with Lumigo, you can definitely just go to the explore page and then just query any data that, Lumigo, captures for you including stuff that traverses your EventBridge, and you can see your trace goes through EventBridge, Lambda, EventBridge, Lambda, and so on. Whereas, X-Ray, doesn't suppose that yet. So, I think Lumigo, Epsagon, and I want to say Thunder as well.

Serhat: Thunder does as well, yeah.

Yan: Yeah, cool, so yeah the oldest sort of, serverless, focus is the observability is they added support for, EventBridge, a while back.

Ben: And, if you go just back to the multi-account use case, we've been doing that on a lot of projects. So, we have one AWS account per service, and then one account per environment. Which is great from a sort of blast-radius security point of view. We also have clients who have legacy architectures, and actually, just last week, I found myself writing some .NET in a legacy system, which, I wouldn't advise you to do, but this was then sending events to EventBridge in the new system. And, that was a cross-account, EventBridge, which allows us to do a really nice sort of strangle-pan style migration to serverless. Through a progressive, sort of minimum viable migration-style, rather than a big flip-the-switch-style migration.

Sheen: Ben, questions about this cross-account event sharing. When I looked at a while ago, one thing I didn't quite like is, I lose the control transforming what I sent to the other account. It was kind of forcing me to send the original event to the other account, which is not going to work in every scenario. Where, the source of the event account needs to control what it provides to the other account. Is that still the case? Do you see issues around this?

Ben: Yeah, I think it's still the case, and in our use case, it's really two trusted systems. So, we didn't really have to think too much about limiting data that's coming from the events. From a legacy system point of view, sending data to other AWS accounts, maybe we could reduce sort of the cost of doing that, by reducing the data loads. But, yeah, those untrusted systems, we still have a sort of issues, around how we can try and filter the data at source, rather in the target AWS account.

Yan: Sheen, in that case, why not just do the transformation at the destination account between EventBridge and whatever eventual processing thing you've got? Because, I think the transmission's there, right?

Sheen: You mean between the two event busses or ... ?

Yan: No, not between the two event bus, but your event bus goes from the central account to the microservice account, and then you're going to process it, so can you not just do the transformation there, as opposed to between the busses?

Sheen: Yeah, I could do that, yeah, yeah. My point was, even the point from the source event when it goes out. That's where I would prefer to have the control.

Yan: Why is that?

Sheen: Say if you're sending payments data from a service that captures the payment data and tokens and things like that. I may have so many PII or data that you want to send to another account, that is dealing with order, or something else. So, there are certain scenarios where, it depends, again, even if it's within the same department, organization, it's fine. If, it's going to a different department, or different organization then, it can become an issue, exposing everything from the original event.

Jeremy: Okay, Ran, do you have any other thoughts on the observability of EventBridge?

Ran: You know, you're going to have this because if you're using this as an event hub to all of your services, you do have to need something in place, otherwise, it will go chaotic. Especially, if you're using EventBridge, it means that your system is definite microservices oriented. So, either make sure you're following message IDs and what happened to them or choose a solution off-the-shelf, otherwise it will get chaotic very fast.

Jeremy: Definitely, all right, so let's move on to your favorite announcement, Ran, and I know this wasn't necessarily ... I don't think this was announced at re:Invent, but it was pre-re:Invent, but I think it had a major impact on what you do.

Ran: Yep, so it's the Lambda log extensions. Obviously, I like almost all the announcements, but the one that I really cared about was the Lambda log extensions. That, you're right, pre:Invent, maybe a week or two before, re:Invent, this time. Basically, what they did, earlier this year is to provide the Lambda extensions with the runtime. So, you can provide your own runtime, and your own set of extensions on top of Lambda functions of, the runtime API, and do some more things while your Lambda is running, and before your Lambda is running, and after your Lambda is running.

Now, the third part of it, if I'm looking at the extensions API, runtime API, the third thing is the logs API. As we all know, in order to get a service from most of the solutions out there, that are doing either monitoring security observability or any other thing through a Lambda, they require some log analysis. They rely on ingesting logs, and comprehend, in order to generate meaningful insights. So far, right before, re:Invent, or actually up to September or October, there was just one destination for CloudWatch logs. Means that it was kind of competition between how do we solve, for a customer, having several solutions to listen for these logs.

Now, for, re:Invent, I do have kind of a question to everyone. For, re:Invent, they announced the logs API that allows me to choose a custom destination to ship or query or get or analyze these logs. Which is fantastic, it's amazing, but I think it was kind of one month later than needed, because around the end of September, the CloudWatch team, the part of the logs team announced two destinations.

Jeremy: Two?

Ran: Exactly.

Jeremy: Two destinations.

Ran: Like, exactly what we needed for so long time. So, it feels like it's kind of, I can still use the old or traditional way of streaming logs with the built-in integrations, by the way, that CloudWatch destinations got to Kinesis, or to S3, or a Lambda function, which makes good sense. Or, start to use the log's API. So, again, this is a great extension, this is a great capability, but it seems that somebody, somewhere else solved this problem for us. At least, for the meantime. Someday it will be, "Hey, we need three, we need four, we need five." So, the Lambda extensions for logs, is I want to say, unlimited, or at least not limited by a small number. But, that's my take on this one.

Yan: You are limited to five extensions per function.

Ran: Yeah.

Yan: So, there's other things to keep in mind by extensions as well, is that it runs at the same time as your function location. You don't have that background processing time after the customers, after your Lambda functions call is finished running. Which means, in practice, what people end up doing for the extension is they're doing a subversion model, whereby they're buffering things and then sending them in batches. Otherwise, you're going to have to add delay to every single function vocation at the end, and there's no trigger, there's no signal for you to know when the Lambda function call has finished.

So, you don't actually know when in your extension call you can actually save, to say, "Okay, the invocation is finished, I'm going to spend 10 milliseconds to send the logs to whatever destination." Which means, you have this weird batching into the next invocation, which means in cases whereby there's a gap of idle time between vocations on the worker, that means your logs it's not going to go anywhere. Until, either it does invocation, or the work itself gets garbage collected.

And this gets worse when it comes to provision concurrency, because guess what? That thing's going to be sitting there for a long time, before it gets garbage collected after 8 hours. Which means, if you are running provision concurrency, there's a chance you're not going to see your logs for a very log time. Unless, there's a regular invocation on those provisions concurrency. So, I agree.

Farrah: When you say a long time, I'm curious how long do you mean by that? What's the timeframe that people could expect for that?

Yan: Up to 8 hours. So, a Lambda worker has got a lease for up to 8 hours, which means, for provision concurrency, that gets kept around from the moment it's created. If there's one invocation at the start, okay, the logs are batched, I got buffered in the buffer in the extension. Nothing happens for 8 hours, so you're going to see the logs at the end of the lifetime of that worker, when it gets garbage collected. Because, that's the time when the extension gets a signal that Lambda function's terminated, it's shut down. So, you can now clean up and as part of the cleanup you can then send the logs to one of the third-party services they're using. Which means, you're just going to get a weird things of, "Okay, where's my logs? I don't see it for like 10 minutes." Because, there's no activity on those provisions with concurrency.

Ran: Yeah, and you're probably the one that read all 100 percent of the fine print of serverless. Anything that is written on AWS, you know exactly how many hours there are for logs to stream from a provision concurrent Lambda.

Yan: As part of the release, I had to read into a law of the extensions documentation, to figure out how it works and had some chat with Santiago, one of the guys that runs the team that work on that feature. So, I probably learned a bit more, too much about how that works than I should.

Jeremy: Hopefully, it's one of those things where most people don't need to know how that works. They just use Epsagon or Lumigo or Thunder or something and they just plugin it in. But, I have seen some people doing real-time log streaming with that. Especially, in a development environment. Which is really cool. I always looked at the Cloudwatch sort of the attaching listeners to your CloudWatch logs as sort of a lazy-logging type thing. Because, it's always delayed and it always take extra time, so I think if you need real-time logs, that extensions API certainly gets you much closer than you were before. But, yeah, I hadn't heard about that buffering problem of logs getting stuck in there, that's kind of interesting.

Ran: Yeah, I would that I know that you have that problem of continuously loading logs from CloudWatch. Like, you try to refresh, refresh, and it seems to get longer than expected. I do say, that when you're subscribing logs, it comes much earlier. If you're doing realtime processing, it will come faster than you'll see that on the CloudWatch console itself. So, that's one thing, but when you're having the log extension API, it's realtime, I would say in a matter of milliseconds or less. It's more for, I would say, realtime analysis, or gathering or batching of data. You might want to batch or gather some data, reduce, or down sample, or do something meaningful instead of all of the logs, just send the metrics out of these logs. Or alert when a five-window batch time seem to have some anomalous error.

And, again, it's a matter of, if you can wait probably one second to get your log that's probably an overkill, but if you're doing something that requires tons of data, or something that is more realtime, that's probably the case for it.

Yan: So, Jeremy, there's one problem that does make it really difficult to actually, what you're talking about, streaming, just because in your extension code, you're polling the log API. Sorry, you're registering for events from the Lambda logs API, but you don't have an event to tell you when the function invocation's finished. As in the actual module code has finished. So, you don't know when it's safe for you to terminate to stop your extension, because it's got a same, sort of similar to as a polling model to cut some extensions. So, it cuts the runtime, where you run your extensions when the invocation starts, and then you have to say when you're ready to yield, and to give up. So, if you're not careful you end up just running it for longer. The function finished here, but your extensions still running, so you end up causing actual delays to the Lambda invocation time itself.

Jeremy: Interesting.

Ben: One thing we actually did to get the logs in realtime for development environments, we did this earlier, actually last year. We created our own custom run center with node.js, this is before the extension support came out, and then we overwrote consoles.log, created an API gateway web software directly to the developer's computer, and then we could shoot console logs directly to the developers computer. Before the function finished executing, and in sort of almost realtime. So, if you do want to get direct feedback in your development environments, and you could do this, I suggest you do this in production. It was a bit of a hassle to set up, it's definitely not.

Yan: Serhat, did you guys, Thunder, didn't you guys have something a long time ago that let you do realtime debugging against a node runtime, essentially doing something similar that you are pushing, you're running a node debug on Lambda and then you're pushing events to someone's ID who's listening to that endpoint?

Serhat: I know, Sercan was doing some crazy stuff. Yeah, I don't know the details, I don't want to know the details.

Jeremy: All right, let's just hope that somebody figures it out and it works. All right, let's move on, so, Farrah, what was ... You got a favorite announcement?

Farrah: I watched a lot of customer stories for Stackery, I was writing about companies doing serverless, success. And, what really excites me, as if you're starting a product or company today, you literally have instant access to the compute power that enterprises are using. You have security performance that will get you to scaling that you need. But, I saw all these talks that are talking about dealing with hundred terabytes of data that they need to import, or imported or uploaded somehow. And, so I really just feel like watching all these, you really see the raw power of AWS. I know there's simply no way that people could have done these types of things prior to utilizing the cloud without spending, I don't know how much money, but it'd be astronomical. Those types of things are really, really exciting. At a time where I feel like, I felt pretty stagnant, a lot of us, we're not traveling, we're kind of stuck. To really see that innovation is still happening.

And actually, it's even moving faster, and in times of COVID, companies were able to scale and respond to their different needs. I saw a lot of stories from Liberty Mutual, Sheen, and your thoughts from, LEGO. But, there was tons of them, from like Volkswagen to AutoDesk, all of that was, I guess you'd say pretty powerful and incredible. This kind of didn't make me feel so stagnant.

Jeremy: Yeah, I think that the idea of the number of people that have been able to build companies in the cloud, using serverless technology, without having to spin up hundreds of thousands of dollars worth of equipment to do things. Especially, even just some of the machine learning stuff that is getting baked in. I was at a small startup before and we were barely doing anything in terms of ... we were a small customer of AWS, and I think our bill every month was like, $18,000. And, this was what? Seven or eight years ago. If we had serverless, at the time, our bill probably would have been $2000 a month. And, so what else could you have done?

And, it also means, what else can someone else do? What can the single developer do, or the small development team? Or someone who's just interested in maybe experimenting with something. I mean, it opens up a lot of doors.

Farrah: It definitely does, we're seeing that with startups and, in fact, soon I hope to have a couple case studies out about this. But, you really just see teams and the power that they have and the speed, and how that extends their runway, and their delivering on their roadmap a lot sooner. And just what that feels like. And, that, to me, it's really exciting, because I feel like we all need something to kind of hold onto right now to keep us moving and engaged in what we're doing. And so, those types of things really help me.

Sheen: I think that's an important point that Farrah mentioned. The startups, they don't just become successful simply because of cost-factor. I recently wrote a blog post, all the sort of flexibility and tooling that we get for free or whatever, for nothing. That allows us to move fast. That's an important thing. I watch quite a few of these real use-cases at re:Invent. It amazes me the different use cases, the way they use serverless. The one I liked was the Scottish Land Registry. There were tons and tons of records, real documents with S3 and serverless. Amazing, amazing stuff.

There was another thing which I never thought. I was watching DynamoDB-related, the talk from this famous entertainment company. Anyway, so, their approach is to keep it on-demand when they launch a product because they don't know the volume of traffic and the capacity. Then, once they study the traffic pattern, then they set to the provision capacity mode. Which I never really thought, because usually provision at-table and that's it, you're done. You never go back and change this, so amazing stories and really cool tips to take home.

Jeremy: Right, and I think that's a good point you make too, about setting the provisioned throughput for DynamoDB. There's actually several services where it doesn't always have to be on-demand. There's a few things where there's provision currency for Lambda, or obviously, provision throughput, where you can set certain sort of baseline, where you want to be. And, there's always a lot of flexibility in that, so if it goes above it, it's still going to scale, whatever. But, you can really optimize your costs as well. So, even in those situations where you might say, well, it's all on-demand, it's going to be really expensive, and I have a really flat sort of pricing model. There are ways to do that. Including savings plans for Lambda. So, even if you are using Lambda quite a bit and it's not as spiky a workload, you can still find savings in there. You've just got to do a little bit of digging, but it's possible.

Sheen: Yeah, it was Disney actually, that DynamoDB talked about.

Jeremy: Oh, yeah, Disney+, right yes. That was a good talk, that was a good talk.

Sheen: Yeah, it was, yeah.

Gillian: And, they made the cost anomaly detection service, it was GA during re:Invent, and I believe it's free. So, you should definitely turn it on, because it saved me right after re:Invent. Because I went in and started turning things on and trying things out.

Jeremy: Yeah.

Gillian: And left the server on.

Jeremy: How many people using a table, and they're wondering why. "Why am I paying this money for a provision table?" All right, so, I want to finish up with you, Ben. So, we already talked about EventBridge, which I know is a big thing for you. But was there anything else at, re:Invent, that really stuck out for you?

Ben: Sure, yeah, well EventBridge was the highlight for me, but just taking a second to think just sort of what my second topic would be. I'm actually working at the minute on a big, serverless, data lake project. Where we have data going into an S3 data lake, and then we're querying that with Athena, and visualizing it in quick sites for business intelligence. And, that's going really well, but what we need to do now is more realtime insights. And, we were previously doing this with some stuff going into Dynamo and querying off of that. But, what came out at re:Invent, and hasn't been probably the most talked about feature is the tumbling window support for Lambda.

Jeremy: Yes.

Ben: So, previously, let's say we had data in a DynamoDB table, which was then, had a DynamoDB stream into Kinesis, we could do some processing on that data, but we couldn't really build aggregate statistics based off previous data. With this support, we can now have the stream of data coming into Kinesis, and for each sort of batch of data, we can have as one of the inputs the states of the output of the previous batch of computation.

Jeremy: Right.

Ben: So, we could, for instance, calculate the day's sales by having the output of the previous batch and keep adding on to that. We had a slightly different use case, and it's a little bit more complicated. But, this has really helped us have realtime data coming, visualized to the user.

Yan: So, Ben, could you not do that with, Kinesis analytics before? Or, is that not possible?

Ben: Yeah, so, Kinesis, I think is definitely sort of the other way to do it. And, Kinesis analytics actually had some new stuff out at, Reinvent, I think. When we tried to do it we found Kinesis analytics a little bit restricted. Because we weren't just doing the simple summation, we needed a bit more complex state coming from our previous execution. But, yeah, there might be a way to do it with Kinesis analytics. It was the tumbling window support was a little bit more flexible for our use case.

Yan: Okay, sure.

Jeremy: Right, yeah and there was a couple of other, I mean, besides tumbling windows, there were custom checkpoints that were added in. The SQS batch windows, again, just all different ways that you can have just a little bit more control over how you're processing your data. I think that was a big win, and I had talked to Ajay Nair right after re:Invent and kind of went through some of these different things, and, it seems to be the goal of AWS to basically take all these little objections, of well, I can't, or "I don't have enough control over this, or I don't have enough control over that," or whatever. And, just keep adding those in and adding those in. And then, I think that goes back to your point, Gillian, where it's like eventually, it's just going to be a lot of configuration, and maybe no code at all.

Yan: I think the SQS batch scares me a little though.

Sheen: Why is that?

Jeremy: Why is that?

Yan: Because the dealing with partial failures is already kind of tricky.

Jeremy: You don't want to deal with 10,000 partial failures?

Yan: Yeah, yeah. Okay, imagine you've got a batch of 10,000, two records failed, how do you then know which ones to delete yourself and which ones you don't? And, let it retry. If you don't then you have to do item potency, make sure it's done right. And, how do you track 9,998 records in-process previously and two didn't.

Jeremy: Right, right.

Sheen: That's interesting because someone asked me the same question. I mean, not 100,000, I'm sorry, 10,000, but even 250 messages batch, they pose the same concern. You know what I did? I pointed that person to your recommendation, you talk about, Yan, to manually kind of do the dealing, so that when it fails you won't get everything sort of reprocessed. I mean, that's why I like the one that ... What is it called? Custom ...

Jeremy: Custom checkpoints.

Sheen: Yeah, checkpoint, yeah, that helps. So that you don't kind of get into this mess, just ...

Jeremy: Yeah, I mean, because the bisecting of batches was great, but then you would still end up reprocessing a bunch of things. Now, you can just say if it fails at this point, I'm going to start again at that point. So, that's a pretty good thing. So, all right, anyone else have any other thoughts or big things that happened at re:Invent, something you want to share? The floor is yours? Luca.

Luca: Yeah, I think that we had great announcement shifting machine learning more towards dev ops, because, AWS announced that Sagemaker Pipeline, and so, which is great to manage machine learning model are shifting from development to production. And, it's something that was missed and it's something that is filling the gap between data scientist and the developer. I was having a very nice discussion with a friend a couple of weeks ago, and he told me, this means that data scientists are not anymore wizard with magical books full of spells, but they are becoming engineers and they are bringing things into production. And, situational pipelines goes in that direction.

Jeremy: Yeah, no, I think machine learning ... There were a lot of announcements with machine learning. I wish we had more time, we've already been talking for a while. I try to stay away from machine learning, but it's maybe just because I watched The Terminator too many times when I was a kid, and I'm just very nervous of Skynet becoming self-aware. Anyway, let's leave it there, everyone, thank you so much.

Sheen, Serhat, Gillian, Yan, Ran, Farrah, Ben, and Luca, this has been absolutely amazing. Nine people may have been too much. I don't know, we might have to cut it down next time, but I appreciate all your insights. I'm going to put in the show notes links to Twitter and information on all of these amazing people that were on the show today. Thank you for watching, thank you all for participating, and we will see you next time.

Serhat: Thanks, everyone.

Farrah: Bye.

Gillian: Thank you.

Luca: Thank you very much.

View Details

About Michael Behrendt

Michael Behrendt is a Distinguished Engineer in the IBM Cloud development organization. He is responsible for IBM’s technical strategy for offerings around serverless & Function-as-a-Service. Before that, he was the Chief Architect of IBM's core cloud platform and was one of the initial founding members incubating it, led the development of IBM's Cloud Computing Reference Architecture, was a worldwide field-facing cloud architect for many years, and drove key product incubation & development activities for IBM's cloud portfolio Michael has been working on Cloud Computing for more than 15 years and has 37 patents. He is located in the IBM Research & Development Laboratory in Boeblingen, Germany.

  • Twitter: @Micheal_BEH
  • LinkedIn: Michael Behrendt
  • IBM Cloud Code Engine: https://ibm.biz/codeengine
  • Cloud Functions: https://cloud.ibm.com/functions/
  • IBM Cloud Functions Tutorial: https://cloud.ibm.com/functions/
  • IBM Cloud Code Engine Getting Started: https://cloud.ibm.com/docs/codeengine?topic=codeengine-getting-started
  • IBM Cloud Free Tier: https://www.ibm.com/cloud/free

Watch this episode on YouTube: https://youtu.be/t3KHoCAVazU

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly, and this is Serverless Chats. Today I'm speaking with Michael Behrendt. Hey Michael. Thanks for joining me.

Michael: Hey, Jeremy. Thanks for having me.

Jeremy: So you are a distinguished engineer, chief architect, Serverless IBM Cloud at IBM. So why don't you tell the listeners a little bit about your background and what you do at IBM?

Michael: Sure. Thank you. So I've been working at IBM in various technical roles over the last 15 to 20 years. I have been in product development, product incubation, I've been working in the field as a workload architect. And for the last 10 years as well I've been working in the Cloud division in itself, working on various topics, incubating it and so on. And since about six years now, I'm really focused on serverless as a topic as a whole. So that's what I'm doing most of my time. Working with customers, working on product development, making architectural decisions, technology decisions, and so on.

Jeremy: Awesome. All right. First of all, I want to thank IBM for sponsoring this episode. So that's great continuing to support the community and continuing to invest in serverless. And when it comes to serverless at IBM, you are the guy. You were there right back in the beginning. I had Rodric Rabbah on the show a couple of weeks ago. And we were talking about how it all got started. But I know you have a bunch of stories as well. So what if we go all the way back and start that sort of six years ago and talk about how did it begin? How did serverless at IBM sort of get kicked off?

Michael: There is some interesting stories there. So long a while ago now, I've been looking into the serverless market as it was evolving, what was happening in the field, what customers are doing. And I felt like we need to do something in the serverless space as well. And by purpose, I thought we shouldn't be starting this as a right off the bat product development effort, but rather since it was such a new space do some exploratory stuff first and have it really open-ended in terms of what we are going to end up with from a technology perspective.

So I was in Beijing for a business trip and I had a call with a VP for research at IBM for Cloud. And I still remember it was 10:00 PM at night. And we talked about we need to do something in that space. So we agreed on that call, let's do something in that space. And he basically then brought in a team from the research side, Rodric was part of the team to kick off that whole effort.

Jeremy: Right. So I don't think I've ever heard a story that starts 10:00 PM in Beijing, ever heard a story that didn't end, or it didn't have an exciting ending to it. So all right. So you brought in this team to kind of start working on it. And so what did you do first? What was the initial goal? I mean, you were surveying the market, doing the research, as you said. So sort of, how did you sort of take those first steps?

Michael: So we put together this team of really talented people in research, and we basically set up our goal. It's what do we want to accomplish from a workload perspective? Which kind of workloads do we want to support? We want to allow composition of functions, something we are talking about these days as well, but it was like a new concept back then. We wanted to be able to be very flexible in terms of which kinds of workloads people can run. Should it only be functions or should it be more cost in the workloads as well? So we went into different directions.

We looked at non-functionals like, how quickly should it be possible to deploy a new function or update a function to have a very quick interloop development cycle. And that drove lots of technology and design decisions. And we've been running that with playbacks every week I believe, where the team played back to a broader group of people like what they were doing, their findings and so on. And we iterate it all way towards into that. And one of the big milestones that was at the end of this first wave was OpenWhisk, as an open source project.

Jeremy: Right. So what were some of those early use cases? Because that was one of the things when serverless sort of first started coming out. And again, OpenWhisk is a fast, functions as a service similar to Lambda or a Google Cloud Functions, things like that. But what were those early use cases? Because I remember way back in the beginning, it was very, very limited.

Michael: Yeah. So I think one of the first use cases was, and that is a bread and butter use case these days as well, still it's those HDP endpoints. That was a very broadly applicable horizontal in many industries applicable use case. Another one that I still remember the specific customers we were working with in these days was data processing, like objects or photos in particular that had to be processed in a certain way, like auto cropping, auto sharpening, object detection, storing metadata.

And I still remember we had one of our very first customers, they went GA while we were still in beta. And so because they felt good with what they had. And I still remember talking to the CEO one time and he said, in the early days, their operations guy talked to him and asked whether our billing engine was broken because the bill was so low. And they came in from a past background. So they moved from a past to function as a service and what they saw was 10X performance increase in combination with 90% cost reduction. And that was just astonishing to them which they had never seen before.

Jeremy: Right. Oh, that's amazing. So we can't talk about functions as a service without sort of talking about the 800 pound gorilla in the room, which is AWS Lambda. But I know that, and this is something I actually really appreciate about what's happening in the serverless movement right now, is that people are looking at it slightly differently. So while everyone's trying to come up with a definition, sort of how people are applying it and how they're looking at it, there's a lot of diversity there, whether it runs on Kubernetes or whether it runs on its own VMs or whatever it is, or the V8 engine, if there's something like CloudFlare Workers. So there's a lot of I guess, difference of opinion, but in a good way. So I'm curious, looking at something like AWS Lambda which I know talking to Rodric that's sort of triggered like, "Hey, we need to do something as well." Just what's the different philosophy there, I guess?

Michael: Yeah. So I've been talking to many, many customers over the last years. And many of them are using serverless, but many of them are still not using serverless yet. And one of the biggest inhibitors I heard frequently was, we would love to use serverless, but you're too constrained in terms of memory. You're too constrained in terms of CPU. You're too constrained in terms of execution duration. You're too constrained in terms of your programming model chasing in and chasing out. You're too constraint in terms of XYZ.

So while people love the attributes of serverless, they do not always like the constraints that come with it. And the attributes are, I think what is dominating these days the conversation in terms of, I don't need to manage infrastructure. I only care about my code artifact. I never pay for idle. I only pay for what I consume versus what I allocate. And those are foundational attributes of serverless that in my world define today what serverless is. And they can be applied to a much broader spectrum of workloads than just what you can handle with functions only.

Jeremy: Right. Yeah. And it's sort of an interesting balance because for me, I really like some of the constraints of serverless. I like that it's event driven, you know what I mean? I like that they don't run for 10 hours or something like that. That you have some of those constraints that almost forced you to think differently about building your application. But on the other hand, I can see why certain customers would say, if we wanted to move to this as a primary compute model that we would have to have different ways that we could overcome some of these sort of artificial limitations.

Michael: Yeah. And sometimes it's a performance thing. If you can get more processes within the same process space, they can do data sharing within that. One of the constraints that has been imposed a lot in the early days of serverless was there was no node to node communication possible. So it was very hard to build up anything that had latency sensitive communication required between the nodes of that distributed deployment. So I think there is lots of goodness in terms of trying to stick as closely as possible to the constraints that were kind of established as part of the almost manifesto of serverless as it was established, but still have the freedom to go beyond that if you want to do that.

Jeremy: Right. Yeah. And I think that's interesting to say, if you need 100 cores, when it spins up on a function for two seconds or something like that, that'd be interesting to have. So let's talk about that. So let's talk about some of these other applications that IBM is looking at trying to expand serverless into. What are some of these other application types?

Michael: Yeah. So what we're seeing a lot these days is like I mentioned before data processing. But not only data processing in a necessarily embarrassingly parallel way, but also data processing that requires more tight coupling between the processing entities. If you just want to do a group buy or join or something that the trust requires more data sharing between those components, that's something we are seeing. And then all sorts of workshops like I said that go up in the high double digits of gigabytes of memory, or that require longer execution times. Lots of stuff in the AI and ML space both in terms of serving, but also in terms of training. So I think everything around data in the broadest possible sense, be it data pre-processing, be it data analytics or be it AI, ML, it's something that we are seeing a big uptake on.

Jeremy: Right. Now what about some of the customers that are using this now? So what are you seeing them doing with sort of the expanded capabilities that IBM has?

Michael: There is some low hanging fruits for people to get into the serverless space. And what I think is important as well is to, from a customer value proposition, I talk a lot to large enterprise customers. From a customer value proposition there is value in running a dozen HDP endpoints on a serverless platform, but data only consumes up so much capacity, right. HDP endpoint is usually not overly huge in terms of its resource footprint. But when you think about data processing or data analytics workloads, they can be really big. And so what I see customers are starting to do is, looking at those workloads ... also batch for example, is a typical case.

Not batch in the necessarily traditional sense where you have a batch run start at midnight and it only processes then, but maybe continuous batch in the sense of, since I'm serverless it doesn't matter whether I spin it up all at midnight or always on demand at a point in time when I need it. And so from that perspective customers are seeing value in taking forward those kinds of workloads to say, "I can get it more real time maybe almost close to interactive for certain use cases." Where in the past I had to wait for a few hours to get something done. Now I can sit in front of the screen and wait for it because I'm getting 1000 cores instantaneously.

Jeremy: Yeah. Well, and that's one of the things that I really like about an ETL task with serverless would be this idea of running things in parallel as well. So there's some cases where you can't just split a job up into parallel and hope it all finishes in 15 minutes with I mean, now you can get 10 gigs with Lambda. But I mean, I certainly see there being a huge benefit to saying, well, maybe a job has to run for 30 minutes. Maybe it needs X number of cores or whatever it is. And being able to sort of parallelize that out. So is that something though that you're seeing the sort of the mindset is, all right, we can use these serverless compute and scale them up really, really large, and be able to do that. But are you still seeing people saying, "Well, if we parallelize them, we can do this much faster?"

Michael: Absolutely. The rethinking of certain ways of doing data processing it's a very interesting one. To give you one example, and it's quite popular specifically in these days. There's one customer, European Molecular Biology Laboratory, quite a mouthful of what the name is. The interesting part is they're doing life sciences research basically. They are dealing with certain data that's taken out of the body, how cells work, how will the internals work? And they have large data sets and they want to be able to process them. And the folks working at EMBL, not all of them are computer scientists.

Many of them have expertise in a different domain, but they use the computer science as a means to an end to do their job. And so what they started doing is, they started to rethink how they analyze these metabolomics, it is called. Those metabolomics and data sets, how to break them down and process them in parallel with cloud functions on every other side, instead of traditional ways. And that gives them both the benefit of it's much cheaper, it's way more cost effective. And they did a very detailed analysis about that. And at the same time, the folks working there get more productive because they get the results faster than they did before.

Jeremy: Yeah. That's amazing. And I love this idea of again, just the parallelization of jobs that would normally take, I can think like Hadoop jobs or things like that, they would just take hours and hours to run and then you'd wait and then say, "Oh, I think I didn't write the query correctly," or something like that. And then you end up with junk data and you have to do it again. So that's amazing. All right. So there's a whole bunch of customer examples I know that you have. And we can talk about some more of those, but I'd love to start talking about IBM Cloud Code Engine and what that is and how that helps people go serverless.

Michael: Sure. So when we made this observation that I described in terms of customers wanting to have the serverless value propositions, but have the ability to go beyond the constraints they were imposed on. We made it a few years ago. We started this effort which is now in the market as IBM Cloud Code Engine. And it's basically our implementation of this mindset of, I call it serverless 2.0 which is preserving the value propositions, but giving people the freedom of high amounts of CPU, memory disk, long duration times and so on. And we're focusing on initially again, HDP serving workloads, of course. We are focusing on batch workloads. And we offer people the ability to run their container whatever it is on our platform.

So with that it becomes a general purpose capability as well. And they don't have to make trade offs. In some other places you have to make a trade off whether you want to have scaled to zero and very fast scaling, or you want to have large amounts of capacity. What we are trying to do with Code Engine is to not force people to make trade offs, but to say, I can have all of those value propositions in a combination I like versus having to choose between one or two or three different offerings.

Jeremy: Right. Yeah. And so now with Code Engine, basically you can write code like you normally would with a serverless function if you wanted to. Right. You could just upload a snippet of code. But you mentioned you could also load a container. So what does that give you, do you get event driven capabilities when you load your own container? Can you schedule them? What are your options there?

Michael: Yeah. So we're making a distinction in terms of what is the artifact you are providing to articulate your business logic versus what can you do with it. And we give people the ability to like I said, to provide their container, package whatever they want to into it, and have that as the artifact that articulates the business logic. It can be a function. It can be an HDP serving app like a Node Express app, or it can be a batch job as well. But irregardless of what they are choosing as the input artifact, in the back end it's all treated in the same way. And all the capabilities in terms of scheduled execution event, driven execution, scale to zero, node to node communication. All of those capabilities are available irregardless of what the input artifact is.

Jeremy: Right. I don't have to set up a Kubernetes cluster or anything like that, right? I just go and it's there for me?

Michael: Yeah. I mean, the Kubernetes and serverless they've been in interesting discussions over the last years and they're very opinionated people on both sides. The way we are approaching it is to say, what do people out there want to have? Is there a market for people who want to have something like a Kubernetes cluster, but not the cost and the pain of operating it? And so what they can do as well is, they can use Code Engine as if it was a Kube cluster, but without having to own the Kube cluster.

They get the Kube API, they can deploy something on there. They can run something on there. They can use CTL, get parts and see their parts and so on. They can do all of that, but they don't have to. So we give them an abstraction where they never have to get in touch with any of that. If they choose to run batch jobs or web apps or functions. But if they want to, if they have the expertise, if they need it for provenance determination purposes or whatever else, they can also drop down into that as well.

Jeremy: Which is crazy. All right. So then in terms of users going in and setting this up. So from a pricing standpoint, this is all sort of on demand. It gives you those serverless qualities that you were talking you about?

Michael: Yeah. It's all on-demand. It's granular pricing like on 100 millisecond granularity basis. The pricing is similar to all the other players in the market as well. So people can again, run batch jobs, HGP applications, functions, containers, but the pricing model is the same for all of them.

Jeremy: Right. So then in terms of some of the use cases, you mentioned HDP and batch jobs and some of that stuff. So what are people using it for now? Are they taking advantage of all those capabilities? Are you seeing them sort of breaking those barriers of what you might be constrained with with typical fast offerings?

Michael: Yeah. The EMBL case for example, I mentioned before is one of those use cases. Then we are working with some large enterprises who have workloads for revenue forecasting and things like that. Which they had run so far in a traditional way, which they are now rethinking in terms of how to run it on a serverless platform. And containers, that's a very horizontally applicable, very general purpose kind of thing. So people are using that often. It's the kind of catch-all for everything they cannot get addressed elsewhere.

And then HDP endpoints are really often used as well as a kind of entry level thing, because also that is very horizontal. And then people are starting to specialize in terms of they say, I have this, like I said, this revenue forecasting application, I use the batch capability. Or I have this embarrassingly parallel executable that I want to have executed 100 times in parallel for that batch is also really super useful, because you just specify your command, like on the Linux command line what you want to have executed, you specify as a perimeter. Do you want to have it executed 100 times or 1000 times, and then you just fire it off and it does all the work for you. So it's that broad spectrum of capabilities they're using.

Jeremy: Right. And you're also, depending on I guess, which artifact you use. So if you just upload a function, you're patching the operating system and you're doing all those upgrades and all of the security things. What about if you're using a container, is that something you as a deployment method, do you have to sort of patch some of those things if you're containing runtime in there? Or how does that work?

Michael: Yeah. Excellent question. And that is often missed. It's a sudden discussion point. If somebody wants to provide a container, that person makes a conscious choice of wanting to have control over what is being deployed. So he can do anything within the container, but with that also comes the obligation of them having to patch it if there's something to be patched. If somebody says, "I don't want to have to patch my node runtime, whenever there is a new CBE coming out. I just want to upload my app artifact." They upload the app artifact, and in that case, we know what the artifact is they want to take care of versus what we have to take care of. And when a new CBE comes out for node runtime time, we will be patching the runtime automatically under the cover for that customer.

Jeremy: Awesome. All right. So what about integration? So IBM Cloud has got a lot of different offerings. There's a lot of different things that it can do. But from a serverless perspective, what does the Code Engine ... does that have native integrations to some of the other services in IBM Cloud?

Michael: Yeah. So we have integrations in terms of events sources, as you would expect the usual ones like object storage, Kafka for data streaming, things like that. But we also have specific optimizations built into it, like setting up a really well heightened performing communication between a compute node or a set of compute nodes and object storage can be painful. You need to know which signature you're using. Are you using V4 or V2 depending on how many threads should you should be using? How should the operating system be configured?

So for those parts, we have basically tested and optimized optimizations built into it. So if somebody wants to interact heavily with object storage, which is a super popular case, they can do that as well. What we also have is for example, some of the AI capabilities. What's an assistant where you can build your chatbots and your interactions, and sometimes you need custom logic. You can articulate the custom logic in Cloud Functions. So there are integrations in both sides, both offerings calling us, but also us being able to call others.

Jeremy: Right. So with the containers you can probably use whatever runtime you want, right. If you want it to run node or something else you could do that. But if you're using one of the predefined runtimes, what are some of the runtimes that are supported if I just want to upload code? Because I just want to upload code.

Michael: Yeah. So it's the usual suspect. It's Node, it's Java, it's Python. Python being really popular, increasing in popularity. And it's all the other runtimes as well in various versions like Java, for example. And yeah.

Jeremy: Awesome. All right. So you're clearly breaking some of the maybe, I don't know if I'd call them the cardinal rules of serverless. But you're sort of pushing past what people are like, "Well, these constraints were supposed to be here." And so you're pushing past that which I think is amazing. And I've said for quite some time that Kubernetes it's an amazing technology. All the things that run with it are great, but it's just too complex for most people to manage. And you're not going to get a small startup with three or four people installing Kubernetes clusters and trying to do that.

And so I always looked at Kubernetes as just the open source that will eventually be productized by all the major Clouds so that you don't have to worry about it. And so in this I think Cloud Code Engine is sort of doing that here. But what's the future like? How far are you going to push this? Because what serverless looks like in five years is anybody's guess. With longer runtimes, more memory, less constraints, we're talking about state now, and how much state can be stored and whether or not you can do, again, node to node communication is something that's still relatively new. So where are we going with this? What's IBM's view? And maybe you as being sort of intimately involved here, what's your view? Where is this going in the future?

Michael: That's one of my most interesting topics I'm working on these days as well. So from my perspective, like we talk about serverless, and if I had to project out like five years from now, we would probably be talking not about serverless, but I just have one server and that server has the capacity I need. And it has in terms of CPU, memory and so on. But I think so far we've been treating the cloud as a place to deploy web service, app service, and database service. If you want to articulate it in very broad terms, right? We've been looking at the cloud as a collection of virtual machines and of the different server technologies. We've not been looking at the cloud as a computer in itself.

What if I treat the cloud as a platform as a single computer, which treats like a simple computer from a CLI perspective, from a programming perspective. And I think that's where serverless will be going. We sometimes use this term, "serverless supercomputer," which basically means is I can define a computer of any size that comes within seconds and goes within seconds, and it's fit for purpose. It's customized to exactly the job I want to hand it now. And if I need now for 15 seconds, 1000 cores to do something, I can get that. It feels like a computer that I get for 15 seconds that has 1000 cores and not a collection of containers or VMs or something like that.

So I think from a user perspective mentality point of view, I think that's where we will be seeing lots of trends towards. In combination with programming frameworks that address this on a higher level, like in the Python space there is interesting stuff happening and in other places as well. But I'm very much behind this start of a serverless supercomputer which lets us treat the cloud as a cloud computer and not as a collection of stuff.

Jeremy: Right. Yeah. No, I love that idea because I think about serverless sort of the first iteration of it which was upload a little snippet of code and it just runs, right. And I mean, then it get to be more confusing or more, not confusing, but more complex where you had to start saying, "Okay. Well, I want this event source. I want that event source." And then, oh, there's some failure modes in here and then oh, wait now I need to run things in parallel and then maybe I need to run multiple jobs simultaneously, or I need to compose functions which has always been I guess, a debate within the serverless community.

How do you compose functions? I know it's always straight functions be calling functions and things like that. And of course, there's been a lot of technology advancements that have made that easier, but what we're doing is we're just stitching together a lot of primitives, right? We're saying, this can do this, this can do that. But I just want to run a snippet of code that can buffer events coming into it, but I have to set up a queue and I have to set up the function. I have to do some of these other things. I have to worry about scaling. I have to worry about maybe downstream throttling so that I don't overwhelm the downstream system.

I mean, there's just so much to think about now when you're building distributed applications. So let's go back to the supercomputer for a second though. So what's the vision? How do you envision that working? As a developer, I just want to use like you said, 1000 cores. What's that experience look like to you?

Michael: So the way I'm thinking about this is, we all know, most of us love the Linux command line. It's a command line. It has often been used as the poster child for good practice of how to develop capabilities, small chunks of functionality you can stitch them together. So today you have a CP command, where you copy data from one machine to the other machine, or from one place on the disc to a different place on the disc. How would that CP command look like in a cloud? That CP command would look like in a cloud in a way where you enter CP, you copy data from A to B, but under the cover, maybe 1000 cores gets spun up, or 150 gigabit interfaces get spun up instantaneously.

And they all transfer data from one object storage to another, or from one device to another. They do the chunking of the data behind the scenes. So it's still just a CP comment that you enter on the command line where you have within, just for a few seconds, maybe if it's only needed enormous amounts of network bandwidth available, because it's handling all of that behind the scenes, you have enormous amounts of compute available if you do not own, you have to copy the data, but you want to have an ffmpeg. There could be a super power ffmpeg version. It could be a superpower of pick your favorite executable or Linux command executable.

And then if you do it intelligently, you can also pipe them together. And that is just what I'm talking about. It's just that the command line part of it. The same thought process can be applied to writing Python applications, or writing Java applications, or running batch jobs, or executing your favorite application. You can rethink all of them in a way of, what if I want to make this as usable as something's running just on my laptop, but with superpowers behind the scenes?

Jeremy: Right. Yeah. No. Again, I love that idea. I mean, one of the things I love about the supercomputer idea is just the cost of what that would be. I mean, you can kind of do it now. And you mentioned that example with the research that you can just sort of spin up all this stuff very quickly, do some big job, and then spin it all down and then not pay for it, because I do think that would be really interesting just in terms of what you could find out. What could somebody do from their garage if they had an idea and they wanted to run some sort of model and figure something out?

Rather than spending millions of dollars on virtual machines and paying for those, or even running serverless functions and having to figure out all that stuff just gets really confusing. So I know there's sort of we're not quite at the super computer yet. Although I do think some people are starting to use that or at least take advantage of that. So do you have other customers that are doing interesting things by just doing this massive parallelization?

Michael: Yeah. So there is one customer, I can't use the name. But it's a large enterprise that does revenue projection. So they always forecast for the next 30 days what is their revenue. And they've been running it so far, I think once a day or so, because they couldn't execute it faster. So they needed that one day to execute it and then they had a weekly run. So they moved this over to serverless and it's now almost operating at interactive speed. And I think that is a big part of serverless. So this is getting us now is in general, I think the next big wave is making basically everything that was not interactive so far because of those users constraints interactive.

Give people the ability to do large things, but not run them overnight, but do them right here right now. And that unleashes another degree of productivity, because if you can turn around things, we all know that from our daily development. If your interloop development is super fast and you do a command S and refresh and you run your code again and again, that unleashes enormous productivity. And if we can apply that to other domains as well, like data analytics for example.

I don't have to worry anymore about subsetting my data to a small chunk of data that I can test something on and then I run it on the big data. I can always run it on big data because then I can be sure that I didn't choose the subset of data that's not representative for the bigger amount of data I want to execute on.

Jeremy: Right. Yeah. Again, I think people are already doing this with CI/CD just to build and deploy projects faster, they're using serverless to deploy ... it's pretty fascinating. All right. So one more thing I want to talk to you about before I let you go. And so Cloud Code sorry, Cloud Code Engine gives you all these capabilities. You can do all of this stuff. How does that compare to just Cloud Functions? Are they the same thing?

Michael: Yeah. It's a good question. So Cloud Functions is really from a technology perspective the most competitive one in terms of cold start times, super rapid scaling from zero to 1000 in the shortest possible amount of time. Then we have lots of optimization built in. Lots of optimization in terms of prewarming machines, keeping them around for a longer period of time, caching them and so on and so on.

So Code Engine implements that capability and addresses that segment of the market with Cloud Functions. And with Code Engine we address this bigger space of the market where you have large capacities and all those constraints unleashed. We want to bring them together, so people can use them in combination. So it's not one versus the other, but rather one complementing the other.

Jeremy: Right. Okay. So if I was starting out and I was just starting to build on IBM, would I go with Code Engine or would I go with Cloud Functions? If I wasn't worried too much about the constraints to start let's say, are Cloud Functions just easier to work with? I mean, I'm sure it's all easy to work with, but what would you suggest for a beginner moving over there?

Michael: For a beginner moving over there I would probably start with Code Engine for the simple reason that it's applicable to a broader spectrum of applications. So there is a very wide spectrum of applications that they can serve with this. If they have requirements that make Cloud Functions really well suited in terms of, they have high dependencies against those codes at times, and things like that, then I would go for the Cloud Functions product. But I would suggest as the entry point because it's much broader in terms of adjustable workloads, I would suggest the Code Engine.

Jeremy: Awesome. All right. Actually, I have another question for you because this I'm just. I know we talked about a lot. I know we talked about a lot of customer examples, but I love hearing customer examples. I love hearing use cases. I love to know what people are doing with serverless because I always hear new ones and it just, it opens my mind. So what's your ... You mentioned a couple of them. You talked a little bit about the ETL ones and obviously the big data ones. What's your favorite use case or customer example of people using serverless and IBM?

Michael: My favorite one? There are so many. I think the most recent one and most favorite one, because it fits into this day and age so well, is what the European Molecular Biology Laboratory is doing. Because they are doing research. They're doing medical research. They are looking for new medicine. How they can cure certain diseases in a better way. And I think it just hits the world we're living in today so well. And we can help with technology to accelerate what they're doing. So that's why I think today I would be picking them because they fit so nicely into the world we're living in.

Jeremy: Yeah. No. That's amazing. I've heard so many stories of people using various serverless products and databases and other things to do research on COVID-19, and all these other things is just solving these problems, which I just don't think would be anywhere near as easy or as quick without this technology.

Michael: Yeah. Exactly.

Jeremy: Awesome. All right. So listen Michael, thank you so much for joining me and for the work that you're doing. Again, continuing to move the ball forward on serverless is not an easy task. So I really appreciate that you're on the front lines of that and moving that forward. And thinking about it differently, right. If everybody thinks about it the same way, we're going to just probably repeat the same patterns that we've done in the past. So thinking about it differently is amazing. So I appreciate that. I know others appreciate that. So if people want to get in touch with you and maybe ask you some questions, or they want to find out more about Code Engine and Cloud Functions, how do they do that?

Michael: So they can get in touch with me via Twitter. I think you have the Twitter handle.

Jeremy: Yes. @Micheal_BEH. And I'll put that in the show notes. So we have it.

Michael: Yeah. Excellent. And so if they want to use Code Engine, they go to cloud.ibm.com/codeengine and they can try it out. It's in beta. Watch this space, it will be evolving quickly. So I'm looking forward to any kind of feedback people are having. And reach out to me on any of those topics we talked about. I'm interested in what people out there are thinking, and we can maybe keep the dialogue going also asynchronously.

Jeremy: Right. Yeah. No, I mean, getting input from people. That feedback is going to be super important as this whole thing grows. So all right. So Twitter, I'll put your LinkedIn in the show notes as well. So I have ibm.biz/codeengine or cloud.ibm.com/functions for Cloud Functions. We'll get that in the show notes. Michael, thanks again.

Michael: Thanks for having me. Really enjoyed it.

This episode is sponsored by IBM Cloud.

View Details

About Tyler McMullen

Tyler McMullen is CTO at Fastly, a global edge cloud platform, where he is responsible for evolving the system architecture and the company’s technology vision. He leads a team of experienced technology innovators focused on internet scale, and working on future-facing, ambitious projects and standards. As part of the founding team at Fastly, Tyler built the first versions of Fastly’s Instant Purging system, API, and Real-time Analytics. Prior to joining Fastly, Tyler worked on large scale web applications, text analysis, and performance. He can be found debating about edge computing, networking, and distributed systems all over the world.

Fastly: fastly.com
Email: Tyler@Fastly.com

Watch this episode on YouTube: https://youtu.be/3F5COSkQlf0

Transcript

Jeremy: Hi, everyone! I'm Jeremy Daly and this is Serverless Chats. Today, I'm chatting with Tyler McMullen. Hey, Tyler. Thanks for joining me.

Tyler: Hey, Jeremy. Nice to see you.

Jeremy: So, you are the CTO at Fastly. I'd love to know a little bit about your background, and what Fastly does.

Tyler: I'll start with what Fastly does. Fastly is an edge cloud platform. What that ends up meaning is that we help people to move their content, as well as their logic, their actual programs, out to run on the edge of the network. The whole goal of that is to make things much faster for your users, better user experience, as well as much more resilient.

It's actually a super exciting place to be, in my opinion. I got into, we founded Fastly, oh, wow. 10 years ago now, maybe more. I can't remember off the top of my head now, but it's been a while. I remember getting into it specifically because Archer, who was our CEO and our primary founder, came to me and he was like, "I have this idea. It's a content delivery network, but it's more like an edge computing network." I was working at a startup at the time. I said, "That sounds extremely exciting." As a distributed systems nerd, that was just, oh, man! It's catnip to me.

Jeremy: Right.

Tyler: So, for the last 10 years it's continued to be exciting. That's how it got started there.

Jeremy: Awesome. What about your background?

Tyler: My background is, I was just a kid who taught myself to program, and got started working when I was about 16 years old, and just never stopped. I skipped the whole college thing and hopped from startup to tech company to startup.

Jeremy: Awesome. So, I'm a huge fan of serverless. Again, I do a serverless podcast, so it's probably quite obvious to people. But one of the things that I am absolutely fascinated with is the idea of serverless computing at the edge, which is one of these things that Fastly is doing. I think that there's a possibility that this could be the future of serverless computing. No more data centers, or things like that, or regions. It's just right at the edge, and as close as possible, that we could get to the user that is actually interacting with this stuff. So obviously, a huge challenge, lots of things that need to be done to make that happen. But I think what would be great for the listeners is if we just take a step back and explain exactly what we mean by compute at the edge.

Tyler: Sure, sure. It's actually a great question, because this is something that keeps coming up. For years, I have been trying to explain exactly what is edge computing. The problem is that everybody has a different opinion as to what exactly it means. I think that the ultimate problem is that depending on who you talk to, that person is familiar with or working on one particular line. One particular edge, effectively, of that network.

So, if you're talking to someone who works at a telecom, they're going to talk about 5G, and how it needs servers inside of cell towers, effectively. Meanwhile, you talk to a traditional ops person, talk to an ops person from the '90s. The way that they think about the edge is actually the edge of their own network. It's kind of the border between their autonomous system and the rest of the network, the rest of the internet. You talk to me, we're going to talk about metro area data centers, as well as even more narrow ones.

Anyway, the list goes on and on. So to me, I think it's actually kind of, the problem is, in my opinion, in the word. The problem is the word "edge," because it implies a line. It implies a specific point within the network, and I don't think that's actually true. Because if you think about all of these different places that we're talking about having computation, they all have really important similarities in their models. The point is that it's not the client. It's not actually the person that you're interacting with. It's also not within your own specific data center. It's not within your core computing.

Everything in between there has a certain set of problems. It means that you don't necessarily have direct access to a database. It means that you probably have to think about doing things in a little bit more of a stablest way. It means that you need to think about doing things at high performance. So, I think that when we talk about edge computing, what we're really talking about is computing in the middle. It's between you and your data center, and your actual client.

Jeremy: Yeah. I think about it a lot. I try to look at it like a CDN. I think of something like a Cloudflare, or even CloudFront with AWS, where they have all these points of presence all over the world. Generally, even Akamai, and some of these other ones that have been around for a really long time, thinking about, you store some sort of static asset somewhere at the edge. It's a .pdf that people can download, or it's an image that loads faster, or what's been really cool happening now is a lot of the stuff with Jamstack, where they're putting HTML, pre-rendered HTML pages on the edge. So, things are just loading insanely fast.

But the idea of finding somewhere to do compute, where actually you can run some sort of business logic. That business logic might be as simple as saying, "Do I route them to the login page, or do I route them to a sign in page?" Or whatever it is, I route them somewhere differently. But the logic could be much more complex, as well. That's what's interesting to me is, if you think about it as a CDN, but with compute, then that unlocks a lot of really powerful use cases.

So, I'm just curious where you see edge computing, maybe a mixture of what we just talked about, some sort of hybrid of the definition, where you see edge computing integrating with what you think of as the traditional CDN.

Tyler: Oh. That's not where I thought you were going with that question. That's really cool. No, this is great. I think it's the mirror of it. You talk about a CDN, you're talking about moving the content. Now we're talking about moving the logic that generates the content. So, the integration there I think is actually going to end up being, for a lot of folks, super tight. It's actually, in my opinion, going to be pretty hard to have a proper, widely used, edge compute network without actually having a CDN attached to it.

I think there's a bunch of different reasons for that. One of them is that almost by its own definition, you're going to end up running the same code repeatedly. If we're talking about an HVP, like a website of some kind, or an API of some kind. You're going to be loading the same things repeatedly. Realistically, that's how the internet works. There tends to be a tail, a spike and a tail for how content is accessed on the internet.

When we're talking about putting servers out at the edges of the network, we're almost certainly talking about a limited resource of some kind. If you're talking about, say big data like machine learning, where you need a large amount of compute power to do it. You're not doing that at the edge of the network. You're not learning, you're not doing training models at the edge of the network.

The reason for that is because it's a lot more expensive to have servers in downtown Tokyo than it is to have them in the middle of the desert in Utah, for instance. So, coming back to it, ultimately, you're going to need to be doing quite a bit of caching. You're going to need to store data so you're not having to repeat the same things over, and over, and over again. I think to me, that's one of the key reasons why the two are almost inseparable, in my opinion.

Jeremy: Right. Yeah. I like the idea of, again, the caching aspect of it, of being able to cache those static assets, whether they're HTML. With compute added to it, there's a lot that you could do to those static files that were cached, where you wouldn't need to make those home runs, and you wouldn't need to do that. You could use things that were local to that particular CDN, or that particular POP.

Anyway, I find that fascinating. But I think there are a lot of different use cases, and I'd be really interested to hear from you. What are some of the use cases that you see people doing with compute at the edge? Maybe what are some of the ones that will eventually open up?

Tyler: Yeah, yeah. I think this is similar to any other new technology that comes out. You're going to have the initial use cases, which we're going to think are really cool. Then eventually, in a couple years, you're going to get the ones that are actually the real killer use cases that we didn't even think of yet.

So, a lot of the initial ones are really simple. They're simple things that make a big difference in end user perception of performance. For instance, instead of having to go all the way back to your centralized data center for every piece of data, what if I actually have 90% of that data, because it's static data, that's already sitting at the edge of the network. Now I just have to go grab that 10%. Or maybe I can feed you some of that while gathering the remaining stuff.

A lot of people think about, I'm thinking about how to put this. A better way to put this is, imagine running a GraphQL server that runs at the edge. You get one request, which actually fans out to multiple different requests. Most of them are already cached, so you're dealing with a much smaller amount of latency, a much smaller amount of variability in latency, I think most particularly.

You also see quite a bit of page rendering at the edge, in my opinion. A lot of that static data is already there, so why send down two different responses? Why send down multiple different responses? Let's just smash it together, right there at the edge, and it's down. Longer term, I think we're going to see all sorts of wild stuff. One of the ones that we worked on internally, just as a prototype, as a little idea, is actually games at the edge. What if you could use an edge compute network to do not only matchmaking of games, but to actually store the state of an ongoing game.

So, one of our little Hack Day projects that we had was doing a multiplayer version of Doom that ran at the edge. It's actually fast. It works. It's really cool to be able to get a bunch of people together to play Doom, and have all of the state actually just sitting there at the edge, ready to go. You can get much closer to a real time type of environment than you could typically, with a traditional game network.

Jeremy: Right. Yeah. I love some of those ideas. One of the things you said about, you maybe can request, 90% of the things you need are local or cache, and then you have to go and get that other 10%. I think about asynchronous processes that you could kick off where you could, say a user goes to a particular page. Then you could say, "The likely place they're going to go next is going to be X page," or something like that. So, now you could preemptively fetch pages, and make sure those are loaded into the cache for things like that.

Now of course, with the GraphQL example, that's an interesting use case because, think about the complexity of that request, to knowing when to fetch it from local cache versus when to fetch it from a home run, and things like that. So that opens up a lot of interesting challenges there.

Tyler: Yeah, yeah. No, fully agree. The other one I wanted to bring up is actually security/compliance/privacy-related things. That's one of the hardest things for us to deal with, and we keep seeing the ramifications of this not being done particularly well in our industry. But imagine being able to have a single layer of your network to be able to do all of that, to be able to say, "Actually, that's a password. I've seen that it's a password. That definitely can't be printed out in plain text within this page."

Or being able to confirm that certain data isn't leaving a particular layer of the network. It gives you a single point, I say single point, but it's actually spread across the world. But it gives you a single deployment point for you to be able to say, "This is our last chance. This is the point before you actually get to the end user." I think it's going to end up being a really powerful security tool for people.

Jeremy: So, speaking about all around the world. This is one of the things that is really interesting about edge computing, or even CDNs. The idea of replicating this to these points of presence, these POPs that are all around the world. The question I have for you, because I've read all of this stuff. I am not an edge expert in any way, shape, or form, but I try to read up on this stuff because I find it fascinating.

One of the things I've seen is companies like Verizon, and even AWS, partnering with other people, doing the 5G thing, putting compute or POPs on the cell phone towers. I guess my question for you is, how close do we actually need to get to the customer? Because that's pretty insane if you can do, again, we'll get to the data piece in a minute. But if you're able to do compute and pull data from the cell phone tower that's a mile down the street versus having to route it to somewhere on the West Coast of the U.S., or North Virginia, or something like that.

Tyler: Sure. Yeah. That is wild. I have actually heard even some even wilder ones, where there was one person pitching me on the idea of putting servers inside of light poles in neighborhoods. I'm like, "Why?" So, there are undoubtedly going to be use cases where that kind of thing is actually useful. The trouble with it, though, is that there's going to be such limited computing power in these place, such limited storage in these places. In order to make it worthwhile for you to use these, whatever you're doing has to be something that is not just for one particular user there. This has got to be something where it is actually so popular and so important, that it is worth it for you to spread this to, say tens of thousands or hundreds of thousands of locations around the world.

My argument is that metro area, city layer, city area, basically being within 15 milliseconds, 10 to 15 milliseconds of users, is plenty for the vast majority of use cases. Now, I could be proven wrong about that 10 years down the line, 15 years down the line, when we come up with some wild new use cases that require you to be within 100 microseconds of where your end user is. But that's not what we're seeing right now. We'll undoubtedly see some specific use cases where this is very valuable. But that's, to me, not the most important form of edge computing, and that's why. Because you can find use cases for something like what we're developing with computed edge for nearly any site. You can use this to make nearly any site faster.

It's going to be a lot harder, in my opinion, for the cell tower layer ones. There's actually a bunch of other reasons why it gets concerning, as well. One of the things that, people trust Fastly quite a lot to be able to handle their private data. We hold TLS keys for a lot of our customers. So, it's really important for us to have incredibly strong security, to be able to keep that sort of thing safe, so that your connections can't be snooped on. It's a lot harder to keep 100,000 locations safe than it is to keep the number of locations that we have safe. You can't make as many strong guarantees about the security of a cell tower as you can about a heavily guarded data center that is nearby in your town.

Jeremy: I guess one of my questions is, when I see people wanting to do those things, like putting them in the cell phone tower, it sounds really cool. There are probably use cases for that.

Tyler: Oh, yeah.

Jeremy: The idea of self-driving cars, for example, that maybe need to ping a network, or something like that. Or the remote surgery, although I don't know how much that would use edge networks. But things where maybe 15 milliseconds isn't enough. Do you see there being use cases where there's some extremely low latency that's needed?

Tyler: Sure. It's certainly possible. I have very mixed feelings about the whole self-driving car pinging the network thing. To me, if it requires, if your car requires the network to move safely, oh, man! We are going to have some real troubles in the future. Again, I think there's definitely going to be some use cases. I don't see them at the moment, though. Maybe that's lack of imagination on my part, but we'll see what happens.

Jeremy: Anyway, I do think that there are probably use cases. Not necessarily to drive, but for traffic updates, or if there's accidents. Things that would potentially, although, again, 15 milliseconds is pretty fast.

Tyler: Exactly. That's exactly where I go back to, as soon as people bring those things up. I'm like, "You really need it in half a millisecond rather than 10?"

Jeremy: Probably very true. All right, awesome. So, the other thing I think that happens with this, and again, you mentioned securing hundreds of data centers, or hundreds of cell phone towers, or these smaller POPs, or whatever. That gets really difficult from a security standpoint, sure. But what about just from a, I guess for building applications. How does the idea of now moving compute to the edge, how does that affect the future of distributed applications?

Tyler: Yeah. This is a great topic, and I think that no one really knows the answer to this yet. We're working on this. I think there's going to be a few stages to this. The first one is kind of where we're at now, where people think of the edge as, it's a proxy of some kind. You think of it the same way as you might think about, "I'm going to put some logic into my engine X server that will run across all of my microservices that are behind it," or something like that, or, "...into my ELB," or something.

I think that over time, what we are going to see, and we're already starting to see this actually, with some of the ideas that are coming out now is, people starting to think about the edge as part of their application. In the same way, and here's why I believe this. In the same way that people now think about the client as part of their application, that's not how we thought about the client a long time ago. I was a developer back in the '90s. I remember how we thought about browsers. The browser was the dumb thing. It was essentially a dumb terminal. You would do all of the rendering, all of everything, back at the server layer. The browser was just there so that the user had something pretty to look at. Over time, that's not how we think about the client anymore. Front end development is real development. It's just as hard and just as serious as back end development these days.

Jeremy: It might be harder.

Tyler: Yeah. No, you could definitely make that argument. You'd probably be right.

So, I think that we're kind of in the early stage of that with edge computing at the moment, where people still think about it as, "I can run little bits of code there." But at a certain point ... let me step back. What I really want people to think about with this is about where is the most efficient, advantageous place for me to run this code. For some things, that's on the client. For some things, that's in a data center somewhere. Some places, it's in a database somewhere. But there's going to be a large swath of things where the edge is actually the correct answer to that. Where if you're not needing to do, in some ways, I actually want to think about it as, being as close to the user as possible should be the default. If you can run something on the client, and it's a powerful enough client, it makes a ton of sense to just run that code there, because it's right next to the user.

So, unless there's a strong reason not to, moving things as close to them as we can, I think is actually going to be a pretty important development over the next few years.

Jeremy: Yeah. No, I totally agree. My concern is, and it's less of a concern, and more of, we don't know yet is, how does this affect how we've learned to build distributed applications over the course of the last five to 10 years? It was always, you start off building monoliths, and then we get into the cloud, and then we start building distributed systems, and we're getting better and better at that. Then all of a sudden, we're hyper distributed systems now because we want to replicate our applications closer to the client.

So I have a whole list of things that this affects and I would love to go through it. Let's start with your existing code base. What does this mean for your existing applications? What do you think a migration to edge computing even looks like?

Tyler: So again, I think this initially is going to look like ... Okay. Think about a traditional application architecture. I came from the Ruby world back in the day. It's been a long time since I've done any Ruby on Rails or anything like that, but that's where I came from. A lot of times, we would have what we referred to as middle wares, and things that would be running between the actual web server itself and the core business logic. So, if you think about it like that and go, "Those would probably be the easiest things for me to try moving out to the edge." If there's just a thing that is completely stateless, that is processing a request, and modifying it, transforming it along the way, that's a really easy thing to move out there.

I think that when we start thinking about the architecture of React apps for instance, there's going to be some really easy wins with that, with service side rendering and things like that, that doesn't actually need to be inside of the core application. It doesn't need direct access to the database, for instance, to be able to do it.

But over time, I think the limiting factor is of course going to be the lack of direct access to your database, or the lack of a really strong, stateful system to work with. I think that what's going to end up happening with that is that we're going to have to modify the way that we think about our data. Suddenly, man, this is something I've been thinking about for a long time. It's a really tough problem. We have good solutions for what to do with problems, with architectures that cannot have a strongly consistent database, that cannot have a strongly consistent distributed system and instead, have to be eventually consistent because of the distribution mechanisms happening.

The problem is that they don't naturally fit the way that we as humans think about problems. They're all these eventually consistent ideas which, you have to imagine multiple different things happening concurrently, and how these things merge together or don't merge together, and things happening in different orders in different places. They're tough to get our heads around, I think. So, I think one of the things that's going to have to happen is that we're going to end up having to develop almost use case specific versions of these things. You can imagine, say, I don't know, a sessions store that exists at the edge of the network. Even that is actually complicated. You think about that as one of the more simple things that a web application can do is just, "Okay, I have some data that goes along with the session. Easy enough."

When we're talking about the edge of the network, we're talking about thousands of servers. We're talking about a user that may be actually in motion.

Jeremy: Like driving. Yes, I was thinking the same exact, or on an airplane.

Tyler: Exactly. So, you could be connecting to one data center, you could be connecting to one server in one data center, even. Then next request, you're somewhere else. So, that data actually has to move along with you, for something like a session store. There's going to be other uses cases where that isn't the case, and we have other constraints that come up. I think that's going to be the trickiest part of this whole thing, though. I spent the last three, four years working on our computed edge product, this specific version of our computed edge product. I had commented to someone recently that I thought I was doing the hard part. Turns out the hard part is actually going to be the state.

Jeremy: Totally.

Tyler: But ultimately, again, coming back to what I was saying before, how we develop applications is going to have to change, I think, in order to take full advantage of this kind of system. I think that it's worth it, though. Again, coming back to the idea of, what if you had the ability to say, "I want to deploy this piece of code to the place in the network where it is most efficient to run, where it has everything that it needs, and nothing it doesn't, and is as close to the user as it can be."

That's a pretty powerful concept to me. I think that, in addition to the state opening things up, if we can find ways to decompose our applications into smaller components, that's going to make a big difference here, as well. Moving an entire application, wholesale from inside of your data center to outside of your data center, that's a big ask. But if I can say, "My application is actually composed of these 16 things or these 32 things," cool. I can pick and choose the ones that actually belong here, and that are communicating amongst each other, and have minimal communications across the wire, that's going to be really cool, if we can get to that stage.

Jeremy: Yeah. I wanted to ask you about the data, because that's one of the things where I can understand and I can wrap my head around building small, reusable components that can be deployed to the edge, because I do a lot of that with Serverless. Where again, you're building small bits of compute. You're separating those things. You're understanding how each one of those things interacts differently. You have to understand how you communicate between functions if that's something you need to do. So, that's something that makes a lot of sense to me. The state aspect of it, though, is just really, really hard to wrap your head around. Because even if you're doing something like a DynamoDB global replication, or something like that, you pick which regions you want it to replicate to, but it replicates all your data.

In the example that you gave of the session store, if I log in, and I'm in Boston, Massachusetts, and then I drive down to Providence, or something like that, and then make my way down to somewhere in Connecticut, I've now passed multiple metro areas that my data has to follow along with me. But under normal circumstances, maybe one, maybe my closest POP is fine. But you certainly don't want to replicate data from a user in Dublin to a user in New York City, if that user's never going to be near that POP. So, understanding what data needs to get replicated, understanding when you purge data, and things like that, even what regions you might want to replicate to, again, compliance, security, all kinds of reasons like that. That's just a really, that is the hard part, I think. I totally agree with you on that.

Tyler: I think so. Yeah, yeah, yeah. There's going to be some use cases where replicating it all over the world is actually the right answer.

Jeremy: Right. For certain things, sure.

Tyler: For certain things, right. I don't think that's going to be the common case, though. For instance, exactly as you were talking about, having a session store for one user that's replicated all over the world makes no sense. It shouldn't actually be replicated anywhere, until, of course, that user moves in some way. So, there's definitely going to have to, if I immediately start thinking about, obviously as an engineer, I immediately start thinking about, "How would I actually do that?"

Jeremy: How would you solve that?

Tyler: Right. It's almost certainly going to require some collaboration with the client, where the client remembers where it was, and can tell wherever it connects to, "By the way, I used to be over here." So, now your local one for wherever you are now, can then go, "Oh, okay, cool. Let me go back and get the state that was associated with this user."

State itself is also going to be, what does state actually mean? Are we talking about a MySQL database? Are we talking about a MongoDB? There are actually multiple ways to think about this. So for instance, one of my favorite systems that has been developed in the last 10 years is something referred to a Microsoft Orleans. Are you familiar with this by any chance?

Jeremy: I'm not, no.

Tyler: Entirely fair. There's no reason you should be. It was the system that ran the matchmaking and users for Halo 2; Halo 2 or Halo 3, one or the other. I can't quite remember now. One of the ideas that they introduced in there was the concept of durable actors. They introduced the concept of durable actors. The whole idea here was that every user, every individual player had a program that was running for them at all times. So, if they're not connected what happens is, that program gets serialized and stored. As soon as they reconnect, we just break that program, in its paused state, out of storage and they're right back where they started.

There's just so many really cool ideas for this. So, you could effectively, if we were going to go down some path like that, you can imagine, essentially you have a program for you on some particular website that has been running, maybe for years at that point. It just keeps getting reconstituted whenever you log back in. There's definitely going to be some interesting ideas that come out of the next few years. We're working on some of them, anyway.

Jeremy: There's going to need to be. Because again, that's one of those things where I'm just like, "How does it work?" I think what that opens up to as well is, the data piece of it is one thing, and the security piece, and compliance, that's another thing.

But what about operations teams? Even in a serverless world, some people talk about no ops. That's not really a thing. You still need to understand your infrastructure. You still need to worry about security and compliance, and do all those things that operations people need to do. How are operations people going to start dealing with thousands of POPs around the world? How does it affect them?

Tyler: Oh, man. No, this is tricky. Again, I think this is one of those ones where we are still in the early days of edge computing, because we don't have, or not many folks have great answers to these questions yet. There's definitely going to be new patterns that come out. There's going to have to be new patterns, because the things that we do now aren't going to work when we're talking about something like this. What does it mean to be able to attach a debugger to a program that is running inside of a server in Mongolia?

Maybe that's actually possible. We have a prototype of something like that working. But is that actually the right way to do it? Or do we just fall back on print app debugging? How does observability work inside of this?

Jeremy: Right. I was going to ask the same thing. Again, you think about that, it's hard enough to observe distributed applications when they're running in one region, in one data center. Spread that out across the world, what does that look like?

Tyler: Right, right. Not only that, but coming back to what I said about being able to break applications into components to be able to run them across multiple different layers of the network. So, if your application is broken into 16 different components, good luck observing that at the moment. So, this is going to require a lot of work from us, and a lot of work from any other Edge Cloud provider that comes onto the scene. But we're going to require, we've already developed some integrations with folks like Datadog, and Honeycomb, and so on, being able to feed data directly back down to them.

But it's also going to be about, if I have multiple components, if I have multiple hops that are happening here, I want to be able to see what's happening between these different places. Where did this request go wrong? Where did it get routed to the wrong place? Where did the data get corrupted, or something like that? I think that's actually going to come back to distributed tracing. It's something that, we all know this. This isn't a new concept. But I think it's going to be so much more important than it was a few years ago. It was a novelty, I think, for a lot of companies, for a lot of people who are working on it. I don't think it's going to be a novelty anymore.

Jeremy: No, it'll just be table stakes for cloud computing.

Tyler: Yeah, exactly.

Jeremy: So, the other thing, again, observability and being able to debug is one thing. But what about the overall developer experience, or just global deployment? I know a lot of these edge networks now, you deploy one place, it automatically replicates, and that makes a lot of sense. With CDNs, it's pretty easy. You just publish to the origin, and then everything picks up from there.

So, those types of global deployment strategies, how are those going to be similar with compute? Then, mix in the data aspect of it and say, "How do I know this node of my compute can access this type of data, or can't?" That seems like that's a pretty hairy problem, as well.

Tyler: Yeah, that's definitely a hairy problem. I think it's not actually that dissimilar, though, to having to do a big deploy, trying to do a big deploy onto a big cluster of machines as it exists today. Now, you may have 1,000 app servers for some companies out there. You already have to deal with the fact that some of them are always going to be out of date. Some of them are going to be broken. Some of them, even just when you're doing a deploy, there's going to be this wave that goes through the whole thing.

So, I don't actually think it's that dissimilar. I think we actually, this is possibly the one place where we do have the tools to be able to do it right now.

Jeremy: But what about Canary deployments, and roll backs, and things like that? That's certainly, I guess you can just roll back by redeploying, essentially.

Tyler: Sure.

Jeremy: But it does seem like there is more tooling, and more thought that still, a little bit of thought that needs to be put into this, probably.

Tyler: Yeah. No, that's definitely true. In some ways, I think we actually have a fun advantage with this. You want to do Canary deploys? We could actually start rolling out your application slowly throughout the entire network. You put it on one, let it run for a minute, put it on two more, and then let it epidemically spread. Sorry to reference epidemics at the moment, but let it spread throughout the network that way.

The developer experience question, though, that you brought up is such an interesting one. This is something that we talk about internally, quite a bit. We had an internal engineering summit a few months back. At the end of my personal talk that I was giving in there, I brought up a couple things that I'm worried about, that I'm like, "We don't necessarily know how to do this thing yet, and I think it's super important."

One of those is the developer experience for it. It's one thing to be able to say, "Great. Three steps, you can have something deployed at the edge." But that's not really the same thing as building an entire application from scratch, or breaking apart an existing application and spreading it onto the network, and spreading it across multiple layers.

I don't think anybody has the answers to this yet. I think it's going to require some new technology to do it. So, that's something that my team inside of Fastly is working on at the moment is, especially in the WebAssembly world. Do we have the tools that we need there, to be able to take multiple different components and have them work together seamlessly, without it feeling like every hop is a new network hop for you, effectively.

Jeremy: So, I do want to talk about WebAssembly, but before we get there, a couple other things on, big questions. These are maybe some of these are theoretical, at this point. But terms like regional compliance. Picking and choosing where or what POPs your applications and your data replicates to. Is that something you see as, a problem that you're solving at Fastly, or something that you will be solving?

Tyler: Right. I don't want to say, yet, whether or not that's something we're actively trying to solve, or will solve. But I do think it's actually a really interesting problem that is likely not going away any time soon. We've been dealing with this for a number of years, in terms of China, and European laws, as well as, we've seen pushes for this inside of the U.S., as well as inside of places like Australia. Regardless of what you specifically think about those laws, they're clearly not going away. I think that edge computing is actually in kind of an amazing place to be able to help developers solve this problem for their users, though. One of the reasons for that is because if we do have locations in all of these different places, it makes it easy for you to say, "Okay, this user's coming from their. Their data can't follow them."

Hearkening back to what we were talking about with that session earlier, maybe your data follows you all the way through the U.S. as you drive across, but then you hop on a plane and head over to Japan, and maybe it doesn't follow you over there. Okay, that's fine. It just means it's going to take you a little bit longer to get it while you're over there. You're going to have to hop back over to the West Coast of the U.S. to get it. That's something that would be really, really hard without edge computing, without something like edge computing coming in. Being able to serve users in all these different countries, it would be nearly impossible. Or at the very least, you're having to do a ton of the work yourself.

So, I think this is something that edge computing is poised to be able to solve for people, but I don't want to say much more than that at the moment.

Jeremy: I think it's interesting because what you potentially get there is a little bit of compute. Even if it's that small piece of compute that says whether or not a particular file can load, or a particular document can load based off of region. It's just much more accurate, or it seems more accurate than trying to guess peoples' location from their IP address, for example. Especially where people are using proxies, and things like that, where some of these other things are harder to fool.

So, I think that's really interesting. Then I guess one of the other questions I have around this, too is, we always get this question of vendor lock in, no matter what you're doing. I'm using AWS, so if I'm on Serverless and AWS, I'm using Lambda, I'm locked into Lambda. To some extent that's true, but you're also locked into MySQL, or Mongo, or some of these other things that you're going to have to do some work to migrate.

But I find it interesting with the idea of edge compute, where if you start spreading around compute to all these different places in the world, what if certain edge networks have better coverage in certain areas, and you want to use multiple edge networks. Intercommunication between them, or interoperability, is this something where you see maybe standards developing around this, so that not everybody's doing something different? Where there could be some way for maybe multiple networks to talk to one another?

Tyler: Oh, yeah. I'm so glad that you bring this up, because this has been the basis of our strategy in this area. We recognize the fact that building out an edge compute network is not something you can just do by yourself. We are one player in the space. I think we're the best player in the space, but we're going to have to be able to work with each other.

Even going back to what I was saying with the different layers of the network, when we're talking about, what if I can move a piece of computation to where it runs best. Whether that's on the client, or on the server, or it's somewhere in between. That is almost certainly going to require some sort of standard way of being able to have a piece of computation, a program, and being able to run it in multiple different locations and expect the same results. So, this is why we have spent so much time on the standards around WebAssembly. I'm undoubtedly going to keep coming back to WebAssembly until we talk about it.

But there's WebAssembly itself. There's WASI, which is the WebAssembly system interface, which is where we are putting a lot of the effort on this standardization thing. There are already multiple different companies that are using that at the moment. My personal favorite one of these is actually Shopify. Shopify has an early product that they have put out where you can run scripts of some kind within, essentially working on your shop, itself. That thing is actually using WebAssembly. It's using some of the software that we wrote, and it's also using that WebAssembly system interface.

So, in theory anyway, I can't say 100% that this is the case at the moment, but it will be soon if it's not. You could have a piece of software that runs on the Shopify platform, that will also run in the Fastly platform, that will also run in your browser, that will also be able to run in the server, as well. So to me, I think that standards are going to be super important for this, and that's why we're putting so much effort into that.

Jeremy: So, let's talk about WebAssembly, then. We can go to some of the other topics later. So, WebAssembly. Again, maybe if people aren't fully aware of what that is, why don't you give them a quick overview of what exactly WebAssembly is.

Tyler: Yeah, sure. So, WebAssembly is something that was developed for browsers, actually. So, it was kind of a response to, if you think back to, some of your listeners might be familiar with Native client that existed in Google Chrome back in the day. The whole idea with this is that it was for running existing C and machine code applications inside your browser. It was used for games and various other things.

Some other folks came out with Asm.js. Asm.js was a way of taking almost a response to that, being able to say, "Okay, that's cool and fast. But what if we could make Java Script really fast? What if we could make it so you've compiled that C application down to a Java Script program, with just a few little tweaks in it, and it would be nearly Native speed?" Then WebAssembly was essentially a response to that, and being able to say, "Okay, that was neat, and so was Native Client, but what if we made a standard way of doing this? What if we made a specific machine code-like language, that we can compile and we can run as fast as near Native speed, and can run in every browser?"

So, that's what WebAssembly was designed for initially. However, it turns out, it's actually great for things outside of the browser, as well. At its core, what it really is, is a super fast, super lightweight, super secure way of, I don't know, cross-platform language. So, if you have multiple different languages that can target this one, and you have a compiler that works for it, suddenly you have a platform that works across multiple languages and multiple different servers.

Jeremy: Right. I think back to, you said you'd been working on web in the '90s, so you're just as old as I am. Remember Java applets?

Tyler: Oh, yeah.

Jeremy: So, WebAssembly is like that, but not terrible. It's very cool. It actually works this time.

Tyler: That's the goal.

Jeremy: I just remember how bad Java applets were, and everybody wanted to do them.

So, WebAssembly is one of those things now where, again, compiling down to Native, runs extremely fast. I've heard a lot of people talk about browsers being those dumb clients, like you mentioned earlier, but you have all this compute power running on your laptop. Why not use some of it?

Tyler: Yeah.

Jeremy: When you have to use something like Java Script in order to do it, you run into all kinds of limitations. But with WebAssembly, it basically opens up that operating system in a way that you can use the full power of it to do a lot of work there. But then that same program, or some variation of it, runs at the edge. It runs in a data center. It can run anywhere. So, that's just fascinating to me.

Going back to this, you recently acquired, or Fastly recently acquired the Mozilla team that created WebAssembly, right?

Tyler: I wouldn't say acquired. We hired them.

Jeremy: Or you hired them, sorry.

Tyler: But yeah, one of the people on that team is Luke Wagner, who is one of the co-creators of WebAssembly, yeah. This was the team that was primarily working on their WebAssembly out of the browser projects. So, they're responsible for Crane Lift, and WASim Time. If you've been working with WebAssembly, those are nearly ubiquitous at this point. You've probably heard of them. You've probably used them.

When we started chatting with them, we were working with this team to create the Bytecode Alliance a couple years back. We've been collaborating with them for a long time. So, when we started talking, we realized that we're actually working toward exactly the same goal. So, when the Mozilla layoffs happened, they were happy to hop over and continue, actually, doing the same work that they were doing before, but now targeted at the edge, instead of at a more central location.

Jeremy: Awesome.

Tyler: Yeah, they're a fantastic team. I think we are super lucky to get them.

Jeremy: So now that you've brought that team in, I'm assuming that WebAssembly is going to be a big part of Fastly moving forward.

Tyler: I think that would be a pretty good bet, yeah. Yeah, yeah, yeah. That's kind of a fun story in itself. This started out as just a couple people working on it over a holiday break a few years back. You hire one person, two people. Now, we have probably one of the largest WebAssembly teams that exists out there, as well as one of the most experienced, probably the most experienced WebAssembly team out there, at this point.

So, it is simultaneously very exciting to me, to be able to really get things done. Fastly, historically, wasn't a language company. Historically, we're not a company that produces compilers and so on. So, now we have turned into a world class place to be able to work on those sorts of problems. But at the same time, I think that there's a lot of responsibility that comes along with that. We've hired up quite a few people who work in the WebAssembly world, and who are responsible for the future of WebAssembly. I have no desire for it to turn into Fastly WebAssembly. WebAssembly needs to exist on its own.

Even for our own benefit, WebAssembly has to be a strong community, has to be a really solid piece of technology. It can't just be for Fastly. That wouldn't make sense for anybody.

Jeremy: Right. I think that's super interesting, and good for Fastly for picking up that project. Again, I see that as being another exciting step. For serverless computing, as well, just the ability, again, to get that speed that we're looking for.

Tyler: Sorry, can I hop in there for a second?

Jeremy: Yeah.

Tyler: To me, you mentioned that for the speed in particular, and I totally agree with that. I think to me, the thing that's actually more exciting than that is the ability, coming back to what we were talking about before, is the ability to be able to take a program and run it wholesale, just bop, bop.

Whether it's on your Serverless platform that is running in a central cloud, whether or not it's in a regional area, or whether or not it's in a cell tower somewhere, or in the browser; that is such an exciting thing to me. Please continue.

Jeremy: No, I totally agree. Right. So, let's jump back to edge computing for a minute here. One of the things that I think is interesting with the amount of compute you can currently do at the edge. It is very, very small. I think 50 milliseconds, depending on which cloud it's on, or which edge provider.

So, what are some of those limitations that you're going to see? I guess edge computing versus your traditional, which is kind of crazy now, but you think about distributing computing, we're talking about hundreds of data centers, probably, anyway. But edge computing verses the data center approach, or the cloud approach. What are some of those limitations that you're going to see? Do you think that it lends itself to this hybrid approach, where some of the compute is done on the edge, but then maybe the more heavy stuff is done in a cloud computing data center somewhere?

Tyler: Yeah, yeah. The limitations for a lot of the edge computing providers, including us right now, are very low. That's just a thing. 50 milliseconds isn't a tremendous amount of time. But it turns out that for a lot of these initial use cases, it's actually enough. When we're talking about doing a GraphQL request, the actual computation time involved with that should be relatively low. That's something that I think we'll see lighten up over time. You'll see those numbers start to go up as people gain operational experience with it, and as the demand, as well, for more computation time increases.

But I think you actually hit the nail on the head, with this actually being about a hybrid approach. The reason for this is that, does it actually make sense to do 30 seconds of hardcore number crunching at the edge, in someone's neighborhood, where you have had to buy up a little piece of real estate in order to put your server? Or does it make sense, in some cases, for that to actually fall back to a centralized data center, where you have those massive computers, where you have thousands of servers all running simultaneously.

I think that's going to be a theme for the next couple years, and I mentioned it earlier. Where is the most efficient place for me to run this? Do you actually really want to do 30 seconds of hardcore computation at the edge? I think that we are eventually going to see a lot of computations, a lot of the shorter computations moving out to the edge, and lot of those heavier weight computations moving back into a centralized data center. That seems like the natural progression to me.

Jeremy: I think that some of those use cases, too, are small bits of compute at the edge, that maybe trigger larger compute jobs asynchronously, that then prepare data, or whatever it is. I think that's a potentially interesting way to think about it.

Tyler: Yeah, I think that's fair. But also, I think the other thing that comes to mind to me is that, if we are 10 milliseconds away from your end user, and you're doing, say 150 milliseconds of computation, let's say half a second. Let's say 500 milliseconds of computation.

So, now your total time do this is 10 milliseconds to get to you, 500 milliseconds to do the computation, and 10 milliseconds back. So, we're talking about 520 milliseconds. Your actual data center might only be 200 milliseconds away from that user. So, is the savings of time worth it in this case? I think this is going to end up being that balance that comes out.

Jeremy: That comes down to intelligent networks, too, being able to understand what's the closest one.

So, speaking about networks, this is a very big market. A lot of people are getting into this. Fastly's been around for quite some time, but you have Cloudflare, and Akamai, and AT&T, the mobile companies getting into it. There's a bunch of them. So, I guess maybe my question here is, who's going to, and maybe you can't answer this, but I like to think of these bigger questions here. Who's going to win? Is it the hyperscalers? Is it going to be the telco providers, like the Verizons and the AT&T? Or is it going to be more neutral providers you think, like a Fastly? I know you hope you win. Or is it just going to be a combination of everybody working together?

Tyler: I am an optimistic type of dude. I'm very Kumbaya, I guess you could say. I think that's it's almost necessarily going to have to be everybody. In order for this thing to work, we are going to have to learn to work together on it. One of the things I say a lot internally is, "I don't want to compete over the basics. I don't want to compete over who is able to run WebAssembly. I don't want to compete over who is able to run an edge computing network at all." I want us all to be able to do that. Once we have all reached this level of, now users can actually use all of these different things, let's compete over the features on top of that.

Competing over the ability to do basic stuff isn't good for users. No one wants that. That's, I think, a lot of the times how we think about standardization work, as well. We want to bring some of our, I don't know if I would say competitors, but some of the folks who are also in this space, we want to bring them along with us and be able to say, "Look, we're not just developing this for us. We're developing this for everybody. Cool. You deployed this, too? Great. Let's compete over what we do with it."

Jeremy: Well, it's funny because I almost look at edge computing almost as another utility. Similar to the internet itself; even though there are companies that own pieces of the network for the internet, for the most part, it's pretty open. So, I could see cloud computing, or edge computing, becoming one of those things where everyone's participating in this, as you said, Kumbaya, creating this thing together, but then building features and applications on top of it, is sort of where the differentiation might be.

Tyler: I think there's some pitfalls in thinking about it strictly as a utility. That direction, I think is reasonable. Again, just to reiterate my point, I really want us to compete over the content, the meat of this problem, not the basics, not over the infrastructure itself.

Jeremy: Right. So then, another question is, talking about competition. We talked about the hybrid piece. Do you see edge compute as being a competitive thing to public cloud providers, where people could eventually host all their applications there? I know you mentioned that you see that, I guess, the hybrid approach.

But I'm just curious because there is this thing called fog computing. I don't know if maybe you can explain it, but where you're using a combination of all these different services; edge, potentially your own data center, public cloud, things like that; and mixing them all together. So, is this something where you think it's just going to be all these things working together? Or is there going to be a niche space for purely edge computing services?

Tyler: Oh, yeah. I think there's definitely going to be a space for just purely edge compute services. A lot of our services inside Fastly work exactly like this. We do everything we can to avoid having a centralized component for almost anything. Obviously, there are quite a few things that we have to. But actually, yeah. I think that the most common case is going to be that hybrid one. To me, edge-only is always the ideal. Honestly, if I set aside my own personal interests from this and say, honestly, client-only is really the idea, because then there's no network involved. You don't have any of the failures associated with that. But of course, that's not realistic for most things.

Edge-only is the next best thing, but I think that realistically, hybrid is going to be where most applications land. I would have a hard time trying to define fog computing, because I remember reading about fog computing in 2002, or something like that. The idea has been around for a long time, and I'm not sure what exactly it's evolved into at this point. But it sounds like, based on what you said, similar to what we're talking about. Where ultimately, I think what developers are going to need to do is, stop thinking about their application as a thing that runs in a place, and start thinking about it as something that is actually spread across the whole network.

Again, I think we have kind of already gotten there to some extent, when people think about their client as just being part of the same application. So, it's a much shorter hop now, I think, to get to edge computing being part of it than it would have been a few years back.

Jeremy: That's a super interesting way to think of it, because that's one of the things is, wrapping your head around that your application does not run on a server somewhere anymore. It runs everywhere. You need to be hyper-aware of what that means for communication, what that means for security, what that means for reliability, for resiliency, for all these other things, observability, like we talked about. That is not an easy thing for people to wrap their head around right now.

Tyler: Yeah, yeah. I think that, oh, man. I think that, honestly, the React developers out there are going to be in one of the stronger positions to understand and really use this. Because React has this concept built in. You have the server-only components. You have the client-only components that go along with it. Why not an edge-only component? Or why not the ability to move some of those between multiple different locations?

I think back to, man, the asp.net days, when I was doing that, in the early 2000s. They had a couple of these concepts in there. They were almost there. Obviously, it stuck around and there are plenty of .net developers now. So, we'll see.

Jeremy: That's an interesting point, though. Again, even early 2000s, even the 2010s, we were still just learning about microservices. Now, we're doing serverless compute, and we're doing containers, and Kubernetes, and all these other things that are just becoming ... the technology is moving so quickly.

Again, serverless, at this point, is probably five, six years old; a solid six years old, if you think back to the beginning of Lambda. But with things moving so fast, I look at edge computing as, you're still day zero, so far to go. Not to use the AWS term, but basically, you're still right at the beginning. It's a very, very new thing. So, what's that path forward? I always ask my guests, "What do you think serverless computing is going to be in five years?" I'd love to ask you, where do you see edge computing in five years?

Tyler: Yeah. So, in some ways, I actually want it to, this is, it's hard to put my finger on exactly what my answer is to this. I simultaneously believe and want it to be, obviously, way more prevalent than it is now. I think that this will be a common piece of folks' infrastructure. That said, I also believe that it won't be obvious that, that's the case. This will just be another layer in your stack. This'll just be part of what you naturally do, where you're not even thinking about it, in the same way that you're not necessarily thinking about the individual server that you spun off on AWS.

Just deploying onto Fastly, or even some other edge cloud network, a rising tide raises all ships in this case, I think. I think will just be such a natural part of developing a high scale application that it won't even be a question. I won't necessarily have to be on podcasts explaining edge computing, because people will know what it is because they're using it.

Jeremy: Right.

Tyler: Again, I think it's also going to change how we develop applications. As we were talking about earlier, it's hard to, the entire concept of a monolith application can't work in a situation like this. We've already seen this broken down, again from the server to client division. It's going to have to be broken down even further, in my opinion. I think we'll see that in the next five years.

Jeremy: That's crazy. Well anyway, I think there's a long way to go, but you're clearly doing some amazing work over at Fastly.

Tyler: Thank you.

Jeremy: So, thank you for that. Thank you for being a guest, and taking the time to talk with me. I know I learned a lot from talking to you.

Tyler: I'm so glad.

Jeremy: Hopefully, the guests found this enlightening. I would say study up on edge computing, because that's the next thing. We'll have to start an Edge Computing Chats at some point.

Jeremy: But anyway, thanks again for being here. If people want to get ahold of you or learn more about what you're doing at Fastly, how do they do that?

Tyler: Check out Fastly.com. You can find me at Tyler@Fastly.com. I don't look at social media anymore, so there you go.

Jeremy: That's not a bad thing these days.

Tyler: I agree.

Jeremy: Well, again, Tyler, thank you so much. I really appreciate it.

Tyler: All right. Thanks for having me, Jeremy. I appreciate it.

View Details

About Tim Suchanek

Tim Suchanek is the lead TypeScript developer at Prisma, which makes advanced data infrastructure developed at large tech companies accessible to all developers around the world. Based in Berlin, Tim has been creating web applications for over 10 years, and enjoys great developer experience and cutting edge technologies.

Twitter: @TimSuchanek
LinkedIn: linkedin.com/in/tim-suchanek-08219346
Prisma: https://www.prisma.io/
Prisma Client: https://v1.prisma.io/docs/1.34/prisma-client/
GitHub: https://github.com/prisma
"Generics, Conditional types and Mapped types": https://www.youtube.com/watch?v=PJjeHzvi_VQ

GitHub: https://github.com/prisma"A Practical Introduction to Database Migrations": https://www.youtube.com/watch?v=xfaps6hgFvI
"How Prisma Solves the N+1 Problem in GraphQL Resolvers": https://www.youtube.com/watch?v=7oMfBGEdwsc

Watch this episode on YouTube: https://youtu.be/Cci65o4IxaU

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today, I'm speaking with Tim Suchanek. Hey, Tim. Thanks for joining me.

Tim: Thanks for having me, Jeremy.

Jeremy: You are a TypeScript lead at Prisma. Why don't you tell the listeners a little bit about your background and what Prisma does.

Tim: Yeah. First of all, thanks for having me. It's an honor to be here. I have listened to several episodes already. Prisma is basically a database tooling company, you could say, where our core is implemented as open source. Everything we build is available for everyone, and everyone can contribute. What we basically focus on right now is a database client, database access, but we're also working on schema migrations. What this database client is doing is mostly giving you type-safe access to your database. We do that in TypeScript.

The way Prisma is architectured, we can also implement clients in different languages like Go or Java. We have the core of the query engine is written in Rust, and TypeScript is basically a layer on top to give you type safety for your database. With type safety, I mean if you for example say, "I want to select a certain field," then you will also in your types, in your code will have the guarantee that this field will be there. I believe that still until today this is the only client out there giving you really this kind of experience. How this works is through code generation.

We employ code generation quite heavily, and you define declaratively your schema. You say, "I have a user. A user has a post and so on or has posts." Based on this schema definition, we then generate the whole client in TypeScript. This is now in GA since 2020. We have been working on this for two years, and in general Prisma already exists for nearly five years. That's what we're doing. Obviously, if you use your database, you oftentimes in 2020 you use serverless, and so we see many users, we see a lot of adoption rising there. While still many users are using this in a containerized fashion, we see a big growth also how this is being used in Lambda and serverless.

Jeremy: What about your background? How did you get into TypeScript?

Tim: It's funny because I started JavaScript I think nine years ago and did not really besides having had type languages in uni like Haskell and Java, the usual stuff we had to look into, besides that I was really into dynamically typed languages, not really a big fan of anything. The type system was my enemy basically. TypeScript came around 2015, I had the first look into it. Actually, Johannes who's the founder of Prisma, he made me aware of it that TypeScript even exists, so I looked into it.

First, I had I would say quite strong resistance to it because once you ... It feels a little bit like it's taking away your freedom. I'm so free in JavaScript, I can just do what I want. If you are looking a bit more into it, if you're reflecting a bit and thinking, "Okay. Where can this really help me?", you are step-by-step understanding that the compiler is not your enemy, but your friend, and just really protects you from your own stupidity basically. We are just humans. It doesn't matter how good of a programmer you are, you will do these mistakes and TypeScript really helps.

Since 2015, I have not used anything else anymore. Here, I really have to site a colleague Ryan from Prisma recently tweeted if he is not writing code in TypeScript anymore, it feels like going outside without clothes. It's really like something is missing. It's like the safety net, the safety layer somehow missing. I am basically someone who converted from JavaScript as my religion to now TypeScript, both in nodes and in the front end and using it four or five years.

Jeremy: Awesome. I've been programming in JavaScript for 23 years.

Tim: Okay.

Jeremy: Yeah. It was a very difficult change for me. It was a huge mind shift for me, but anyways, all right. I don't typically fool my listeners here, but I think we're going to fool them a little bit because even though this is a serverless podcast, we're going to talk a lot more about TypeScript. Of course, we're going to link it back to serverless here, and maybe we'll start that in the beginning, but there are so many really, really cool things that you can do with TypeScript. If you're building serverless applications, it's going to help you. I'd like to start by maybe just getting your thoughts on why TypeScript is going to be an important thing for serverless.

Tim: Yeah. I just checked the stats. New Relic just released a report about which runtime is used in serverless most where the New Relic product is activated, and over 50% is using Node off the Lambda runtimes. It's clear. Node is the majority here. I would now claim if you do Node, you should do TypeScript. It's quite relevant to say most of them are using Node and in Node instead of having it on type in JavaScript, I think it's really useful to use TypeScript. Now is the question really why is that so useful, why should you do that. Maybe a quick introduction again or reminder what is TypeScript. Rather a reminder because I think most of the people know it already or know that it exists. There was the State of JS report 2019 where they asked, "Did you ever hear about TypeScript?" 58% heard about it and want to use it again, and I think only less than 8% haven't heard about it yet.

Jeremy: Right.

Tim: I guess most people have heard about it, but the question is really what is it. I would say you can define TypeScript as a super set of JavaScript. The JavaScript syntax, what you can express with JavaScript is a subset, so that means everything you can do in JavaScript you can also do in TypeScript. That was one of the goals upfront when they designed TypeScript. You need to be able to in hindsight be able to type any program you have written in JavaScript with TypeScript. You need to be able to put the types on top.

Why is this so interesting now? Once you have these types defined, this helps users to explore your API. It enables auto completion, it helps you with documentation, you can add comments. JSDoc and TypeScript types are really working well together. You can do both together. I even recommend doing both so you don't just type things, but you can on top add nice comments in JSDoc. They work well together. After all, if you use VSCode, I guess also many developers are using VSCode these days, the VSCode JavaScript language service is maintained by the TypeScript team. This is one thing because again, it's a subset of it. If you are using JavaScript today, chances are very high that you already experienced auto completion or an enhanced developer experience enabled by TypeScript, although your code is not written in TypeScript.

I think the main point is really developer experience and also giving you safety. If you look into the whole serverless world, then you oftentimes deal with APIs, third party APIs. Let's say we're dealing with AWS SDK. It's huge. The AWS SDK, it's impossible to know everything, and it's also sometimes hard to find exactly what you want in the docs and in examples. In that case, the types are really helpful. I recently looked into Timestream DB, the new serverless time series database from AWS which became GA a couple of weeks ago. Because it is so new, there was not so much content around yet, so I just checked out the types of the AWS SDK. That really helped me to understand some more edge cases details of the API, what is even available. Maybe they didn't document it. I these days even use types oftentimes to explore an API.

Jeremy: Yeah. No. I totally agree with exploring the APIs because the problem with AWS documentation is it's often very complete, but you have to keep digging and digging and digging to find sometimes the right thing that you need, and just to have it pop up and tell you what methods are available or whatever. The other thing is with the JSDoc stuff, that is something that is ... I've actually been working on another library that I'm trying to add all the JSDoc stuff into as well as converting it to TypeScript, completely to TypeScript, which we can talk about in a second. I found that that is a really, really helpful way for you to do ... It almost forces you to do documentation on your own code so you document your own methods and things like that, but then when that is converted to types, it helps just from the user perspective and gives you that really, really good developer experience.

Let's talk about third party modules for a second. The idea of writing something in JavaScript, this is something I did for a long time, I was writing npm packages in JavaScript, and this was early days of TypeScript. TypeScript was not super popular at that point, but then people were coming along like, "Do you have the TypeScript types for this?" I was like, "No, I don't," whatever, and people were then like, "Oh, I'll contribute it." They added these TypeScript types, the type definition files, which were great. But if you look at the type definition file and then you go and you look at the actual code, anytime you make a change to your code, you have to say to yourself, "Wait a minute. Do I have to update the types now?"

One thing that I've started to do is convert some of my projects over to TypeScript directly because one, it's very, very helpful just to have that capability there. Again, a lot of my services interact with AWS, so being able to use the autocomplete for those is good. What's your suggestion for people who have written applications or written modules and packages in JavaScript? How big of a motivation is it for them to convert those over to TypeScript?

Tim: The beauty of TypeScript is that you can adopt it incrementally. By the way, here I want to also point a little bit to Flow. Back in the days, 2015, '16, it was not clear who's the winner, Flow from Facebook and TypeScript from Microsoft. Both wanted to solve the same problem, wanted to make JavaScript a bit more safe, and if you have huge code bases, you are happy to have something like that. In the early days, that was one of the advantages of Flow, the incremental adoption. You could easily have JavaScript files and TypeScript files in the same project. That was not originally so easy in TypeScript.

However, later they added the allowJS effect to the TypeScript config, and with that you can basically start turning file by file into a TypeScript, and also you can make the type checks very loose so to say, you can start with any type, and then you step by step introduce types. I think with that, you can introduce it quite easily. What most developers I guess even in Node.js Already have some kind of compile step. If they don't, yes, you now need to introduce a compile step because you just run the tsc command basically which takes all of your input and it just ... It's a little bit a bundler, not really a bundler, it's rather transpiling.

If you for example use generators, if you use Promise's asynch functions and you want to turn that into ES5, then you may need to transpile that. TypeScript at the same time basically covering Babel, so oftentimes if you have Babel in your project, probably you don't need that anymore if you switch to TypeScript. TypeScript, the compilers for example, also able to parse React like JSX syntax. I think oftentimes, it's getting this compiler into your build pipeline is probably not the biggest thing, especially if you start with file by file, the TypeScript compiler will be very fast in the beginning. It can get a bit slower later if you have really, really big projects, and also depending on how you define the types, how much advanced TypeScript stuff you do because somewhere the compiler needs to run something somewhere, but overall, I suggest really directly starting to introduce the compiler in your code base and file by file convert that over to TypeScript.

Jeremy: Yeah. I found that with TypeScript, trying to incrementally add it, anything works out really great. The problem though that I've had is I always thought a lot of what I was doing in JavaScript with these nice little elegant maps or a lot of ternary operators in there, just things that made the code seem really, really clean to me, and found it very, very difficult to add types to some of those. It almost forces me to rewrite things to if statements, make things for loops, make things a little bit easier and a little bit clearer. I feel like it makes me write more code sometimes, which is fine, but I think it probably makes it clearer and almost enforces some standards that you have to follow.

Tim: Yeah. A typical example where I really have to say that TypeScript improved my code quality is when you do duck typing or you need to check a stat at a specific time. Let's say, I don't know, this type that I get could be an object or a string. In JavaScript APIs you have that oftentimes. What I used to do more is that I just do it in line, maybe a ternary statement, just check is it a string, and then do this otherwise that.

Jeremy: Right.

Tim: With TypeScript, you have a specific keyword built into language that is is. The keyword is called is. With that, you can have let's say a function that returns if that is indeed that type or not. It returns a boolean, but it has a special meaning now, a semantic meaning in the types because you now can do your if statement. Let's say it is my car, and if this if statement is true, TypeScript knows that everything you do in that if block is definitely accessing your car. This way, you really can build, you can make this connection between runtime and compile time because the reality you not always know the type of it. You need to sometimes validate user input or input from other APIs, and this is a very useful way to discriminate the types. Putting that into a separate function alone already is it a blah, blah, blah, type, this kind of patterns is not possible to do that without a function in TypeScript.

TypeScript really forces you. You need to put it into a separate function, that function has this specific return type that is is, and then the type, car or something else, and that way you can clearly separate that. Now, writing these functions is still a bit tricky. Actually, yesterday we did a meetup, new meetup format at Prisma called Advanced TypeScript Trickery because we recently saw more and more crazy stuff popping up at Twitter. We can also later geek a bit out and talk a bit more about the crazy new stuff that TypeScript enables. Someone talked about, he really uses this keyword a lot, and talked about how you in practice really implement that function and found out that what you just mentioned with okay, I have JavaScript code here, I have my declarations, TypeScript declarations, how do I make sure they match, that is still a problem in the is function. That's not yet solved properly.

However, we have great libraries for that. One of them is called IOTS, and there you can basically define a TypeScript type in runtime. You get a programmatic API, you can say tdot and then string, and you can construct and nest your types, and it will give you a valid TypeScript type that will make a valid TypeScript type for the compile time, but also gives you a runtime checker. You basically take the whole TypeScript compile time checking into runtime. If you now get an input and you don't know is it a car, you can use one of these libraries IOTS which are really useful for validation here to validate if things are proper.

If you look at the whole validation space, in general it's really exciting what is changing there. If you look into the libraries that people used to use like Yup for example, some people may be familiar with that, they in hindsight added the TypeScript on top which works reasonably well, but now we have native, TypeScript native libraries coming, Zot is one of them, IOTS, and they give you the full package because obviously it's awesome if you can already validate in compile time if possible. And then, later in your code once you have run this check from the library IOTS, you have to guarantee in your code yes, this object has these properties. You don't need to write any code there anymore.

Jeremy: Right. Which is interesting. Maybe we can move on to the project you and I first got connected on, which is the DynamoDB Toolbox.

Tim: Yeah.

Jeremy: I spend probably, I don't know, 80% of my time coding on writing runtime checks for the data. Make sure that this particular ... That when you're passing options, that this particular option is a valid option, and then making sure that the type and so forth. I had thought about maybe using joi to do that. Is it pronounced joi? J-O-I, whatever.

Tim: Yeah.

Jeremy: But then, I was like, "That just seems like a lot of extra work, and I'm basically doing the same thing, and some of them don't need to be as complex." That is something that would be really, really cool to have is to simply say, "How do I take my TypeScript checks and bring those onto the runtime to enforce them on that side?" Before we get into the DynamoDB Toolbox for a second, maybe we can take a second and talk about something like Deno, some of these runtime for TypeScript things, because those seem really promising. What are your thoughts on those sort of things? I know there's some performance issues.

Tim: Yes. Deno is a really exciting project. It's really awesome, and we are also getting more requests now at Prisma to add Deno support. We don't have it yet because we use the Child Process API quite extensively, and that's just a little bit different like the process model in Deno. We get more and more requests. I still didn't get an answer from anyone because we asked, "Do you use it in production?" Not that many answers yet, but I think that's coming. That is definitely quite expensive to see. If you run the TypeScript compiler, and also a few weeks ago at the TypeScript conference I gave a talk about pushing the ... It was a bit a catchy title. Pushing the compiler to the limit. It was both in terms of complexity, but also in terms of quantity, of how many types we are even having in there.

Performance is an issue if you have really big projects. I think that is also for sure a little bit detrimental in the beginning at least for Deno because they need to do all of these type checks, but I think help is coming. What is really exciting if you see for example the project esbuild by Evan from Figma, so it's basically a let's say Babel alternative written in Go. It's just I think 100 times faster than anything in JavaScript. Now, there's this project called swc, I think. It's written in Rust, and they already are able to compile TypeScript, which makes it already 20 times faster, really much, much faster than the TypeScript compiler, which is self-serving by the way. It's written in TypeScript, which is good. It's a good indicator that it's a proper language. I think Anders Hejlsberg, the founder of TypeScript, is proud of it.

Now, they are really looking to move this to Rust, which is hardcore. If you look into the compiler code, I used to contribute a few little things, it's crazy. You have one file, 20,000 lines, and that is only Anders Hejlsberg who came up with C# and Delphi, he's the only guy who's allowed to touch that file. It's really like you need to have this extreme context in your head to be able to work on that. I even heard from someone who did an internship at Microsoft that they wanted to move the TypeScript compiler. It was an internal experiment at Microsoft. They wanted to move it I think to .net. There is still some artifact on GitHub somewhere of that experiment, but they gave up. They realized this thing is too complex.

Also, there are for sure type systems that we stop, we don't do any generics like Go that kept it simple. TypeScript, they are just going hardcore on the features. It's insane what kind of features they have. They have sometimes features that even Haskell doesn't have, and Haskell is one of the most powerful type systems. It will be a challenge to move that over to Rust. Coming back to Deno, I think Deno is a really exciting project, and also having not to deal with an overhead of adding TypeScript, it's beautiful because you can just write it and you don't need to deal with the extra compile step anymore.

In practice, I also have to say the compile step is not really a problem for me anymore because I'm 99% of the time using TS Node which is just a little CLI utility that injects the ... so in Nodejs you can overwrite the require system, and what they do, if you require a TypeScript file they quickly transpile it and then they run it basically. That's what TS Node is doing, and that way you can directly say TS Node, then the TypeScript file, and you can just run it so you don't really need to care about this extra transpilation step all the time.

Jeremy: Yeah. The way I found you, by the way, I was building this DynamoDB Toolbox library. I've been building this for over a year now, and I had a lot of people saying, "You got to move to TypeScript, you got to move to TypeScript." I said, "You know what? You're right. This is the time we're going to do it." I spent several weeks importing this thing over to TypeScript. The biggest challenge though was that what I wanted to be able to do was I wanted people to be able to define their schema for entities in the DynamoDB table, and then be able to have that autocomplete basically for them when they got returns from whatever operation they did, whether it was a GAT or a query or something like that.

The way that I figured out this could be possible was you could have them define their own type. If they're building in TypeScript, have them define their own type or schema and pass that into the library, and then that would carry that through. Essentially, they would have to do that twice. They would have to define the structure, define the schema for a particular item, and then they would have to write a TypeScript interface or something on top of that that would allow them to basically retype it or add types to that, whatever. I was like, "That just seems like a lot of work."

I went down this rabbit hole and I came across a talk that you did called "Generics, Conditional Types and Mapped Types," which was absolutely fascinating. What I'd love to do is just quickly tell me or tell the audience what it was you talked about that type in terms of what you were building for Prisma.

Tim: Yeah. It's nice to see that you came across that because that was just ... I think we started the TypeScript meetup back then, and then we just said, "Hey, let's put in some talk that let's say talks a little bit more about advanced concepts in TypeScript." It's funny. The definition of advanced also changes. Yesterday again, we had this advanced meetup. It seemed like this, what I am doing is totally basic. I'm seeing what other people are doing is more crazy. This situation that you just described that you don't want users to do a definition in JavaScript and in TypeScript on top or two definitions of their types, of their schema that they need to keep in sync, that's exactly the same problem that we have with Prisma.

Prisma Client, again, you define your schema in a DSL that we came up with at Prisma. We call it the Prisma Definition Language, and it's for the people ... I think it's actually quite similar to a TypeScript definition, or for the people who use GraphQL, it's similar to a GraphQL SDL, schema definition language definition. You basically can just say I have a model user, and then you can point to another model and you can command click it in VSCode and you have this nice experience. That is how you define the whole schema, how you define relations.

And then, another route you have is CLI, that's also written in TypeScript. That generates the whole Prisma Client implementation basically for your particular schema, so all the capabilities it has, query, update, delete. Maybe you have a JSON column in there, and so we give you a specific JSON filter. All of this is generated depending on your schema on your database that you use where you support SQLite, MySQL, Postgres, MariaDB and mSQL. MongoDB is coming, but MongoDB is a completely separate paradigm because it's no SQL. Right now, the SQL database is still more feasible to support them.

What we were looking into is how can we reduce the amount of code generation in order to give people a nicer developer experience, but still have the type safety, and in particular about querying data. Let's say I say I only want the ID. How can I achieve that in a type-safe manner? You could say you for example have ... Let's just build our own database client now. You could have a folder that is called queries, and you could have a file there that you call the user ID query.JSON, and then you just have a JSON definition of that, of what you want, and then you would run a CLI that looks into that and that generates the types or generates what is available. And then, in your code, I have the ID available.

How about you could skip this step of extra generation? How about you can build that into the type system? That is what we did with Prisma Client. Depending on the object which is the query that you write for Prisma Client is just a JSON object. How about we can, depending on the shape of that object, have different types for the output of the query function. In other words, if I'm adding to that object now the name field as well, then suddenly I have this available and the return type of this function, let's say Prisma user query for example, this would be really awesome if that is possible.

We looked into that topic last year January, and we knew there's a mechanism in TypeScript that might make this possible, but we didn't know if that's actually possible. This mechanism is called conditional types. The conditional type really, there's some languages also call it dependent types, there are not many languages first of all which have such a concept. I think C# has it, Haskell has it. The idea is really that you say depending on the input I have, I have a different kind of output. You use generics in that case. Generics, I just refer to it as a variable on the type system, and you can say let's say if the input is a number, the output will be a string, and if the input is a string, the output will be a number. Not sure if that function makes sense, but this kind of combination you can do.

You cannot just say if the input is a string, the output is a string, that you can oftentimes do with generics or you can pack let's say an array around it. You can provide more let's say complex conditions. We looked into that topic, and another important ingredient to be able to implement such an API that maps the input to a specific different kind of output are mapped types. TypeScript has a so-called structural typing, and that means you can define types just by defining a certain structure of the type, and the structure can be let's say an object type for example.

In JavaScript, if you have an object, the object has keys and values. Now, in object type in TypeScript you can say what keys are allowed and what values are allowed. Only these certain three keys, let's say ID, name and email are allowed, and the values have to be a certain type. For example, the email is not allowed to be a number. It has to be a string. With this, you can define object types in other languages, oftentimes abstract. What you can do in TypeScript now with mapped types, you can map other object type and you can turn it into something else.

Some of the build in object types, mapped types into the TypeScript language which are part of the TypeScript standard lib, but they're just implemented with primitives that are accessible to everyone. For example, require, the required type, another option of type. What it does, it goes through your object type and it makes all the properties optional, or it makes all the properties read-only, all the properties required. With this, you can manipulate types and you can take one type and turn it into something else. How this mapped type is working, there's a loop defined in it, and that loop is oftentimes defined with an in statement, and you look through something, through a list for example, and before we dive deeper into mapped types, how you would do that in TypeScript, let's quickly go back into our JavaScript mindset.

If we now want to build, if we want to loop an object, oftentimes we need to do that in JavaScript, we need to dynamically build up an object, we don't just get an object and return it or something like that. Let's say we have a bunch of entries, we have a bunch of key value pairs, and we need to turn that into an object. What we do, we loop through the keys and we take the keys, put them on the object, and then we give them a value. That's how we do it, and after all we end up with an object no matter what we use. For in, we can use a reduce.

The same thing is possible on a type system level in TypeScript. You can loop again through something. What is this something in TypeScript? That's a union type. You can imagine a union type as a little bit like a list or like an array in JavaScript. It's either A, B, C, D, and it can be these four things, and you can now loop through them. You can say the key in and then the list of possible keys of anything, it can be numbers, and then you can give them some value, you can loop that value up somewhere. You can for example say, "I picked this from this other type," or, "I want to wrap it in an array," whatever. This way, you can dynamically in the type system define a loop that is running through the type that comes as an input. This way, you can turn any type into another type. You can do crazy stuff.

What we are doing now, coming back to Prisma Client, that was a quick discourse into more advanced TypeScript types, Prisma Client, why do we need something like that? Also, to the listeners, we don't have any code here that we are showing on the podcast, but I still hope it's kind of understandable. The idea is now in the mapped type, you give Prisma Client an object type that queries let's say ID and name, we loop through these keys ID, name, and we now map that to a different type. In Prisma Client as an input, we always get a boolean. You just say ID true or ID false if you wanted to be part of the payload, of the result, and we can now on a type system level check is it the concrete boolean true or false, and based on that we can calculate the return type. That is basically how all of this comes together.

Again, when we were looking into that back in January, it was not clear at all, January 2019, it was not at all clear if this is even possible because we did not see anyone out there doing it. Actually, I asked a bunch of questions and the issues in the TypeScript repo, and the answers rather sounded like, "No. What you're doing here is stupid. Don't do it." We did it anyway, and actually that's also what I mentioned in the talk at TS Conf, for a long time Prisma Client was not really type-safe. There were edge cases when you were doing some specific definitions of types that you could break the types basically, which we then also later could fix with some more TypeScript trickery.

What I really want to say here, I think TypeScript might be a bit scary. If I'm now talking about all of these crazy types, do I need to understand them to use TypeScript? No, not at all. I rather think this is something that belongs into a library. This is not something I'm doing in my application code. If I am writing an application, I've never used this kind of types. However, also yesterday at the meetup you saw people for example implementing a game engine, game framework in TypeScript. If you are on that level, library level, there you can put a lot of complexity into it because it's like a surface, it's hidden. I believe that with this, having the complexity rather in the library, we can keep our application code more simple.

Jeremy: Yeah. No. First of all, if you didn't understand what Tim just said, go and check out this generics conditional types and mapped types talk. I will put the link in the show notes so you can see this. And then also, you're pushing the compiler to limit. I'm sure you talk about a lot of it in there as well. What you said, I'm going to boil this down quickly. Essentially, what it is is you can in your query, when you write that query in TypeScript, but you write that query just defined as a simple JavaScript object that when that query returns in ... This is at the type checking stage. This isn't even compiled, it actually generates a new type for you that shows you which items are available on the results of that query, which is crazy because I think about just defining an entity type.

Let's say I'm just defining an entity type in a database and I say, "Well, it's got an ID, it's got a name, it's got a date, it's got a client number, something like that." The problem is that if my query doesn't return all that data, then when I'm using TypeScript, it's going to say that it's there even though it's not going to be there. This magic of saying we can dynamically generate types so that even at the coding stage when you're writing the code, that you know whether or not a certain thing is going to be available. To me, that just blows my mind. You're totally right, though, because I think about how I write TypeScript when I'm doing a simple project versus how I write TypeScript when I'm doing a library. Those are two very, very different things.

Tim: Yes.

Jeremy: For me, I jumped into the library writing side of TypeScript very early, so I went down that rabbit hole and those generics and all these mapped types and stuff like that. The other thing I want to point out too is you mentioned these things as you were talking through, but this was another thing about that talk that I really appreciated because it really helped me connect TypeScript with a programming construct that I already understood. That's when you compare JavaScript to TypeScript in terms of saying generics are like variables. You said this. Mapped types in TypeScript are very much so like loops in JavaScript. Union types are like lists, conditional types are like if else statements and keyup was like an Object.keys where you can pull that out. And there's a whole bunch of other things too, even never. Never was always one of those things I kept on seeing. I'm like, "What's the point in never?" You're like, "It's essentially a type there."

Tim: Yeah.

Jeremy: It's not always true. These things don't always hold true, but it is a really good way to wrap your head around it.

Tim: Yeah. Exactly. I think that's really also the challenge in terms of education that we can make this bridge between what it is in runtime, what is it in compile time, but I also see more and more people doing more advanced stuff here. One of the advanced things that I just need to point out now, what came in TypeScript 4.1. Yesterday we had an engineer from Spotify, actually it was two days ago, from Spotify also at the Advanced TypeScript meetup. What he talked about is typed string and template strings. Template strings, probably most people are familiar with that if they're doing JavaScript. Let's say a nicer way to concatenate or have multi line strings and all that kind of stuff. It was a great addition to the language.

Now, what they did in TypeScript, you can now type them. It's a bit insane. In TypeScript, there's a keyword that is called the infer keyword. What you basically do, you can define a crazy type setup, you can say there's a Promise and there's a function around it, and three more functions around it, and you say in the third argument of the fifth function, you put the infer and then you just call it T for example. If the type that you got in has exactly that form or five functions nested, then you can now take that type of that argument and you can do something with it. You can type check it, you can do certain things.

You can do the same with these template string types, so you can say that if there's for example a template string that has a specific structure, let's say it has ... What they even did, they implemented a split, a string.split on a type level. What is a split? You have a left side, you have a right side, and you need to have split in the middle. They just defined that as a type, and you can do that in TypeScript now, and you can even make that dynamic what it's splitting on so you can again have that as a type parameter, and so you can say now, "I have a left side and a right side." What TypeScript is doing it splits as soon as it finds that pattern. It's a little bit like how Regex is working in greedy fashion.

What you can then do, again, you can take what you split, you can take the third part of it, you split into three parts the first one, the middle you may throw away in a normal split implementation, the third part you again call split on it. That's the powerful thing here. You can define a split function on a type level. What people even started with now is implementing a JSON parser in TypeScript on the type system level. What does that mean? I have a string. Also important to understand why does this even work, in TypeScript you not only have the string type, but you can also have a type that is a very specific sting literal, the string with a specific content that can also be a type. You can say only the type ASD is possible.

If you now have a Foo and Bar, if you have a union type between these, you can basically build your own Enum. You can say it can either be Foo or Bar. It can be the string with a concrete content Foo or the concrete content Bar. What you can do now, you can make these even more complex. You can also do union types in the string, template strings, which then distributes them. You can do combinations and that kind of stuff. People are still exploring where this is useful. I think where it gets useful is if for example in a timestamp API, let's say I am allowed to use the ISO 8601 standard for timestamps, and you can obviously do a runtime check, but now that you can actually do a compile time check if that string, that particular string that you pass in is even valid. What that parser is then doing is it's basically people implemented Regex on the type system level. This kind of things, they are not making or breaking it, but we get more and more these nice things to make the developer experience more awesome.

Jeremy: Yeah. That's amazing. All right. A couple more things about TypeScript specifically, and then I want to bring it back to serverless for a minute. With TypeScript, one of the ... Maybe this is just a debate that I see because I'm reading all these different forms on Reddit and all kinds of things, but when you're building your types, when you're adding things like interfaces and types and things like that, how do you structure your documents? I know I have multiple classes and things like that. Do you include interfaces and types? Do you put those in the same file? Do you separate them out into other files? How do you structure your TypeScript project, and where do you put those definition files?

Tim: That's a good question. I have seen very different approaches for that as well. Some people put it in a types folder. In TypeScript, there is something called definitely types organizations, and what it's basically doing, you mentioned earlier JavaScript libraries adding TypeScript support, it's also possible to do that outside. You can have a separate package that is then called convention @types/the module name, the npm package that you want to type, and that way you can put types on top of something that is untyped. This @types, calling a folder @types, that is a pattern that I have seen more and more often. What people do, they have a separate folder with that kind of stuff, and then they just write down all the interfaces, all the types in that folder and import that in their application.

What we are mostly doing at Prisma still is keeping it simple and having the types collocated with a code, so that means the types are just at the head of the file which is useful to understand the requirements that that function has. As soon as you don't have a simple let's say two argument function anymore and you want to really have five, six, seven arguments, it's useful to have an object as an input type. There, especially also in React, you need to do it anyway when you have prop types, you will just define that type at the top of the ... It's quite useful in my opinion to have that at the top of the file so you see the data requirements, what that function, what that class needs in order to run.

I think collocating it is one way I saw that, and also the @types directory, and then also if you're writing TypeScript, oftentimes the types are embedded. I think this location of types we're mostly talking there about interfaces and object types, but oftentimes you have the types embedded in your code if you write it in TypeScript upfront so that you have your return type typed or you have your parameters typed in there. That's I think even more readable if possible.

Jeremy: Awesome. I'm in the collocation camp. I love to collocate my types in the files. I just find it easier. The second you separate them out and you put them somewhere else and then you have to go look them up, even if I have to import a type into a different file, I just import the other file that has the type in it, and it seems to work pretty well for me. I don't know if I'm doing it right, but it's good to hear that you do a similar thing at Prisma.

Tim: Yes. Exactly. Yeah. What I quickly wanted to mention here is that this really helps us and the team as communication. We have ESLint rules activated. For the people who are familiar with ESLint, there was the effort of TSLint done by Palantir, and they merged that into ESLint, and now it's a plugin for ESLint because they just realized they cannot build such an awesome project again, so now it has access to the ESLint AST, and ESLint has the capability to pass a TypeScript. There are some rules that even force you to make the types a bit more explicit. That helps us in the team. In the moment that you are writing the types, especially coming from the front end, or in general I see that a lot with the front end developers, types are in their way or it feels like, "I don't want to write this type. Why do I even need this?"

I have the feeling the acceptance for this in backend in Node.js Is a bit higher still because also to be fair, the React tooling and so on, it's already quite awesome. ESLint, you could nearly call it TypeScript compiler because it understands the code already quite deeply. But anyway, we have a rule in our setup that forces you to explicitly write down the return types. Why is that interesting? For you in that moment, it's trivial. You just hover with your mouse and you see it's whatever, the TypeScript language servers in VSCode will tell you this is a number. Where this is really useful is for the PRs. If you have the PR review in the GitHub UI, you don't know. You don't have your intellisense. In that case, the types are really useful to just understand what's the contract.

After all, a type system is a contract, what gets in, what gets out. Also, then later if you are refactoring and you are not aware what was the type actually and you think you didn't break anything, with having the return type explicit, you just know what is expected of this function. We found that quite useful.

Jeremy: Yeah. The other thing too that's great besides ESLint is Jest. I use Jest for all my unit testing. With the TypeScript there, it just works really, really well. All right. Let's bring it back to serverless, and then I'll let you go and get back to your TypeScripting. The cloud edge, edge computing seems to be something that is getting more and more popular. We've got Fastly and Cloudflare workers and things like that. I know a lot of them run on the V8 engine, but there's also all this stuff with assembly script and WASM. What are your thoughts on that?

Tim: Yeah. That's really interesting to see. If we talk serverless, oftentimes it seems like serverless equals AWS, but I have the feeling that is changing more and more. I think Cloudflare can now be considered a legitimate player on the market in serverless with the Cloudflare workers. It's also really interesting to see how Fastly and Cloudflare are targeting completely different markets. Cloudflare is rather going for the application developers. They make it easy, they give you JavaScript, they made the V8 engine very fast, so they did some tricks to basically give you nearly zero code start. What they did is that while you have a TLS handshake, they already put the function for you. As they have this very minimum V8 runtime, it's much faster than if you have a Lambda function that now needs to start the whole thing. This way, you have much more lightweight functions.

They cannot do that many things. You cannot run native code in there. However, let's say the gate into a native code is WASM WebAssembly. Also, crazy news that Fastly basically hired basically all the WASM people.

Jeremy: Yeah, I saw that.

Tim: Either if you're a Rust developer, it seems like either you're working for Fastly or I don't know. Then, it's long tail, so they have extreme competence now in that team. For Fastly, it seems rather they're going for the advanced use case. That's the whole Fastly CDN is built in that direction. Let's say TikTok is using it for example, or Stripe. You can do these crazy configurations. It's not necessarily built for application developers, however, in order to get into their runtime, we just talked about the Cloudflare runtime which is based on V8, you can just deploy JavaScript and that's it, which then also by the way means that you can deploy TypeScript, and there are great examples out there how you can run that.

In Fastly, you can also theoretically do TypeScript, but rather a subset of it called AssemblyScript. They are pushing it really hard. They are also sponsoring the AssemblyScript project. What is AssemblyScript? The idea is really they selected a subset of TypeScript, it's the same Syntax, you feel familiar with it if you use it, that is able to compile to WebAssembly. The Fastly runtime is a different piece. It's completely different design than Cloudflare. They claim I think they have 34 microseconds boot up for the runtime, so if you give them your WebAssembly file, then they can boot that on the edge in 34 microseconds, which is impressive. That is traditionally only Rust.

Jeremy: Right.

Tim: Also, their first toolchain, what they're supporting there in the edge beta, I had a look at that, that's only Rust. However, now they're adding the AssemblyScript support, which is exciting because this opens up this extreme let's say advanced edge computing, which normally if you are let's say a solo developer, a small team, you would never look into that. Now, you suddenly have access to that and you can write very fast edge functions. The difference between Fastly and Cloudflare, there are also a bit because I looked into that quite a bit because I'm working on a side project related to that. In Cloudflare, the function is not a Lambda function. It's a function it's always triggered. That's basically the outer edge. Even before the caches hit, that function is always called, and then in that function you have access to the cache API and can get something out of the cache and return it, so the cache there is basically powering Cloudflare.

In Fastly, it's different. Fastly has Varnish in front of everything, which is this quite old cache implementation, but very legit and just works well. You can now talk to Varnish, you can communicate, you can configure Varnish on top, and then if Varnish didn't have the cache, then you can write your function behind that, and that can now also be with AssemblyScript, basically TypeScript. I have to be fair here, it's not really yet TypeScript. Most of the packages will not work. I had a look into it. It's not yet there. It will take a while. Yes, it has the same Syntax, but it's still way to go. If I would use WebAssembly at the edge, I would probably rather do Rust today because the tooling there is much more advanced, but it's still really exciting to see that they are pushing forward this. Cloudflare I think is a great addition as well, and I am actually using it as a cheap alternative to API Gateway because you have pricing of 50 cents per million requests. I think API Gateway, you're somewhere at $2.50 or something, $2.50.

Jeremy: I think it's $3.50 for the REST, and a dollar for the HTTP APIs if that's not confusing enough.

Tim: Yeah. Exactly. Now, you are down to 50 cents basically with Cloudflare because what you can do, there are packages out there that you can directly call your Lambda function with AWS API from a Cloudflare function so you can now have that as your gateway, which is really exciting. Again, all of that can be type-safe if you write it in TypeScript.

Jeremy: That's amazing. Yeah. No. I think the whole WebAssembly thing and edge computing, that to me is this next evolution of serverless computing. What's great about the WebAssembly too is that again, running that on a browser, there's so many more things that you can do there and package so much more compute there. The less that the internet is required here to connect to a little bit of data, but again everything powered in that toolchain is just absolutely amazing. Really exciting stuff. Tim, thank you so much for sharing your knowledge here because honestly, I've learned so much from you just through the video that you did and our previous conversation. I am very, very happy that I found you because it completely changed the way I think about TypeScript. I appreciate that. I hope this did something similar for people listening to this. I will put all of your talks in the show notes. If people do want to get ahold of you, what's the best way for them to do that?

Tim: I guess just on Twitter @timsuchanek. Yeah. I guess in the show notes maybe there you can link to it, but it's just T-I-M-S-U-C-A-N-E-K. The name comes from Czech Republic. My heritage is German, but that's where it originally comes from. Anyway.

Jeremy: Awesome. And then, if people want to check out Prisma they go to github.com/prisma, correct?

Tim: Yes. Or Prisma.io. Exactly.

Jeremy: Awesome.

Tim: Prisma Client is out there. They can check it out. It also works for JavaScript, and we have a Go client in beta, so you can also check that one out. Yeah.

Jeremy: Awesome. All right. I will get all that stuff into the show notes. Thanks again, Tim.

Tim: Thanks for having me.

View Details

About Nicole Yip

Nicole Yip is an Engineering Manager at the LEGO Group. She has been working as an Infrastructure and DevOps engineer for over 5 years mostly as a consultant helping teams of all sizes get their services into AWS. Her roles have often become the catch-all for everything non-application-developer but that matches her passions for AWS, Infrastructure as Code, CI/CD, and Security. Nicole is a frequent speaker at various conferences, with past events including Serverless Architecture Conference, DevOps Con, and ServerlessDays Virtual.

  • Twitter: https://twitter.com/Pelicanpie88
  • LinkedIn: https://www.linkedin.com/in/nicole-yip-5b792292/
  • Medium: https://medium.com/lego-engineering
  • LEGO: https://lego.com
  • ServerlessDays Virtual 2020: https://www.youtube.com/watch?v=9oYS_5eL610

Watch this episode on YouTube: https://youtu.be/_09c6maJ3Uc

Transcript

Jeremy: Hi everyone, I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Nicole Yip. Hey Nicole, thanks for joining me.

Nicole: Thanks for having me.

Jeremy: So you are an Engineering Manager at the LEGO Group, so I'd love it if you could tell the listeners a little bit about yourself and your background and what you do at the LEGO Group.

Nicole: Sure, so as you said, I'm an Engineering Manager. I have joined the company about a year and a half ago, as a Senior Infrastructure Engineer and then moved up to Engineering Manager and I've joined in the Direct Shopper Technology Team. So we look after www.lego.com, all of the pages where you're browsing for products, completing your order through checkout, and redeeming VIP vouchers, that's what my team looks after. And specifically, I look after the platform. So I head up the platform squad there where we look after the infrastructure and hosting and developer experience CI/CD, security, and operations of the site. So quite a big remit there and specifically my background as in AWS and managing production workloads and that's really where my interest is and what I'm doing at the LEGO Group.

Jeremy: Awesome, well so I am super excited that you are here because I love the LEGO Group, not just because I loved LEGOs as a kid, but also because I started talking with Sheen Brisals a long, long time ago and he was so super excited about the whole serverless process and just building things with serverless. And so it was really interesting to hear the process that the LEGO Group has gone through. And it's been, I think over a year since I talked to him on the show here. And so I'd love to set the stage here because there is this talk that you gave at ServerlessDays Virtual recently about this audit process that you do at the LEGO Group in order to make sure that you're following best practices with serverless and that you're always kind of upgrading. And I wanna get into that but I think to set the stage for everybody to know just how serverless the LEGO Group is, maybe you could give us just a quick timeline of where it started and where you are now from a serverless perspective in your engineering group.

Nicole: Yeah sure. So the story starts a little bit before I joined back in 2017 when we had, it was kind of an event that was the last straw, the straw that broke the camel's back, so they say. And yeah, so there was a launch event that was highly anticipated, we didn't survive. And that led to us looking at options that weren't on-premise, that weren't hosted on-premise. So we started back in 2018, actually scoping out serverless and AWS and seeing if it would work for us. So we migrated a single user-facing service and a couple of backend services over to the cloud, got them running and they were handling high season traffic by the end of 2018. And so yeah, fully in production and ready to go. And that then led on to us making the entire lego.com site serverless. So the pages that I mentioned that are within our team's remit, we moved all of them into Fargate instances and serverless Lambda functions. And so the only on-premise system we have is our source of truth, our warehousing system and we've wrapped that up in Lambda functions so we only talk to that asynchronously.

Jeremy: Nice, so now where are you now? You've mentioned Fargate, you've grown the number of engineers that are working in serverless, so like where are your just rough numbers, like how many Lambda functions do you have? Things like that.

Nicole: Oh, this is fun. So we had four Lambda functions in 2018 during Black Friday, Cyber Monday. In 2019, which was last year, we had just gone fully serverless. The platform was handling high season levels of traffic and at that point we had four Fargate instances and I think it was 36 serverless services. And a serverless service can be made up of many Lambda functions. I think we had around 150 to 200 Lambda functions. No, I think it was around 150 Lambda functions in production at that time. And now that we've gone through another Black Friday, Cyber Monday high season period, we're now over 260 Lambda functions in production with over 56 serverless services and still the same original four Fargate services. So we're growing pretty quickly.

Jeremy: I would say. So going back to the engineers 'cause this is another thing that fascinates me is how different companies group engineers together to manage different services and to make sure that, you know again, everyone's sharing information across teams and so that everyone's working together. And you mentioned you're on the platform squad so you are broken into squads. 'Cause I find it fascinating but I think the listeners might be interested in, how are those squads set up? What's the makeup of them? Like what are the disciplines that are there? You know just how does that work?

Nicole: Yeah so within the Direct Shopper Technology Team we have seven different squads right now. But back when we went serverless, we had two squads. We had a back end and a front end focus squad. So we largely had Lambda focused and serverless focused engineers in one squad and React, Fargate, Express, and Apollo engineers in the front end focus squad. And then once we launched the platform live we reorganized all of the squads into being product-focused. So they split up into five different squads each focused on different areas of the site. So that was one squad for checkout, one squad for VIP rewards, one squad for just exploration and how you discover products, and so on. And these squads are as full stack as we can get them at the moment. So the application engineers, QA engineers, the product owners, and now we're including stakeholders in there as well.

So people from other teams that are actually setting these requirements, these business requirements. So we've got quite a few people in each of those squads but we still have an essential platform squad. And the reason for that is it's a brand new platform that literally was written and launched like a year and a couple months ago. So we're still in that pattern of having a dedicated operations or platform squad but we're actively trying to move away from that. We're trying to train up each of those application engineers in the product squads to become operations focused, to have that mindset of building things in a way that makes operations easier, because next year, they don't know it yet, they're going to be operating their services there. They're going to be the ones that are on call. They're going to be the ones who have to wake up in the middle of the night. And yeah, that's where we're going.

Jeremy: That's awesome, I mean I love that idea too, of cross-functional teams. It's always been a big thing that, I shouldn't say always, but more recently has been something that I think works really well. And I love that idea of training up the application developers to be responsible for their applications. They're not responsible for the servers right 'cause there are servers but there are no servers. So that's what I think an interesting approach. I'm gonna have to have you back on to tell me how that works out later on 'cause I'll be interested to see that. All right so this idea of training up application developers to understand how to run or to be able to manage their own applications, their own stacks, and just in terms of making sure that what it is that each team is building now, sort of what you're responsible for on the platform team. I mean a big thing here is ensuring best practices, right? I mean to make sure that everything is written in a way that makes sense, that you're not duplicating a lot of code, that anything that can be shared again, I think the security aspect of it.

What else? I mean just in terms of maybe Middleware and things like that. Like the common things that you need to have, there are a lot of unknowns in the serverless world, right? I mean this is something that didn't really exist five years ago and certainly to the level now that people are using managed services. The amount of knowledge you need, the amount of best practices that has just expanded dramatically. So I'd love to talk about the serverless audit that you did at the LEGO Group. Because again, I may be overusing this word, but I find it fascinating, I love this idea of being able to say, okay in our organization we have a set of best practices and we're now adopting a new technology that we can go to conferences and we can see some of these people talking about this stuff, but in order for it to work for us, in order to make sure that we're the most productive and secure and move as quickly as we can, we need to adopt our own standards. So let's start right at the beginning. So why, and maybe I gave it away, but why did you say, "Hey we need to do this audit of all of our serverless applications."

Nicole: Yeah so it's really two reasons why. So the first one is that in order for us to have written so many serverless functions over the last year and a half, we've had to grow the team. We've got a lot more engineers now. So we went from having around 20 engineers back at in like July last year. Now we have around 60. So it's a three times increase in the number of engineers in the squad. And they're at all different levels of their careers. They're juniors, mid-seniors, and not all of them have been exposed to AWS, let alone serverless and serverless best practices. So when we started building out the new platform we had some engineers who knew serverless and it was manageable for them to talk to a single architect and have a shared vision of how we were gonna build our services.

Now with 60 engineers and still one solutions architect that's not scalable. And so the reason to try and the way that we're trying to share out that knowledge and really make sure that we're maintaining that high standard of building functions that are scalable and meet all of those requirements of security and all of that is really just by writing it down and creating a standard. And that's where the audit came from. So we took all of the best practices that we were already doing in our newer services. So every time we write a service, you improve a little bit here or there and we just said, "Okay all of our new services are implementing these things. Let's write that down. Let's have a look at our old services and see what needs to improve." And so really the audit came out of us having a lot of old services that weren't really being maintained or kept up to date with the new standard that we were setting with our new services.

And so then came the fun of auditing these 36 services that we had last year and really calling out all of the areas that we needed to improve them on. And we also took the opportunity to start thinking, "Well we're growing the team so quickly. We can't have a central operations team anymore either." So at the time there was four infrastructure engineers if you could believe it for this global website. And we couldn't handle having from 20 to 60 engineers throwing functions at us. So we took this opportunity to also add an operations related thing. So making them have Canary deployments and basic alerting to implement on their services. And this is really setting the stage for where we want to go next year, of well now you've written these services, now you can own them and operate them. And that's how the standard came together. Both from existing practices and new ones to help operate the services.

Jeremy: Right, now we were talking the other day and you referred to the older services as legacy serverless applications. I think is quite funny 'cause it is amazing. I mean they're only a couple of years old. You said 2017 so just these services that are just a few years old. I mean so much has changed with serverless over the last couple of years. I mean again, we just finished re:Invent or we're in the middle of re:Invent right now. There's so many new things that are happening and I'm sure these will evolve over time as well. So let's go through the focus areas. 'Cause I thought this was a really interesting approach and if you're listening to this take some notes because honestly this is really good. I mean your team did an amazing job of outlining what these focus areas should be and helping standardize that. And so if anybody's listening, this is just a roadmap for you, if you're listening to do this. So let's go through that. So what were these eight focus areas? What were the main things that you said this is what we need to make sure we get right in our serverless applications.

Nicole: So we started out with, because we're an infrastructure team and we had the frustration of having to get this operations burden off of us and into the squads, we started with alerting observability and logging guidelines. So those were the three that are really key for if you want to know what's going on with your service and production. If something goes wrong you need to have observability and traceability through your stack. And you also need the logging to know if particular transactions or batches have been dropped. So those were the three first focus areas. We then started thinking about how do we get engineers to get their own code into production. So we added, we thought about safety and deployment. So deployment mechanism being another focus area that we added in.

So I mentioned before Canary deployments adding in that safety net when you're deploying into production, that's something that we needed to, in order to actually hand off deploying services to the squads. And then the others kind of come out of the best practices of what we were already doing in our newer services. So integration testing, making sure that you've written integration tests that are run in the right parts of the pipeline. So unit tests on Prs, integration tests at that stage as well, and also in QA and acceptance. We also have a standard around secrets management. So we've mandated that we use SSM parameter store as our secrets manager but also we want to use it in a specific way so that it's easier to audit when we come through and say, "So which service does this secret belong to?" So we've got a certain naming convention for that. And then there's another thing for Middleware. So Middleware has made a huge difference to our services. Where one of our developers put it really nicely, "Where it takes all of the bits that you always have to do as a developer that no matter what-

Jeremy: The bootstraps stuff, right.

Nicole: Exactly yeah, yeah. The stuff that you're always writing whenever you start up a new service, Middleware takes care of that and then your handler just becomes business logic. And that's essentially what we're trying to focus on, right. We're trying to get our functions to do something with the business outcome. And so taking all of that out of the handler makes it not only easier to read but it tells you exactly what the service does, 'cause you can read the handler and it's all there. And then things like pausing like validating that the request that comes through is in the expected format. Making sure that responses that go out are consistent, things like that. So Middleware has been amazing. So that was another focus area to get onto all of our services.

And the final one was documentation, which I know it's been around for a while. It's not fun but it's so important when you're growing a team. 'Cause what we found directly from one of our engineers was saying that, "You've got new engineers going into these new squads. I've told you we've changed them up two or three times already. We're gonna keep changing them up." And so when you're inheriting services or when you're onboarding a new person, not only are you pairing with them but you can give them a document that says here's exactly what the service does. Here's how you get it up and running, here's all of the service limitations around it. Here's the architecture diagram. And so that was really key to add into the service audit to make sure that not only can we maintain our services, are they built according to best practices, but we can also maintain them with new engineers. So that's where we got to with our focus areas.

Jeremy: Right, so now you come up with these focus areas and this is all based off of learning from multiple years of implementing this stuff. You're building new services, you're learning from that. So you put together this set of focus areas and I've worked with a lot of engineering teams, I've worked on a lot of engineering teams, I've managed some engineering teams, and I can tell you, and I'm sure you know this, that the best way to get somebody not to do something is say, "Hey, here's something that we came up with that you now need to implement and we know better than you so this is the way that we do it." So that's like, I guess the declaration from on high doesn't always go over well. So how did you take all these squads with all these engineers and get this big of a project or this big of initiative implemented?

Nicole: Yeah so the rollout was really interesting. We actually had to do it in two phases. So the first way we did it was we conducted the audit within the platform squad. We went through the code base, looked at every single service, compared them with each of the checkboxes and the checklist, and raised tickets for each of them. So if we saw that there was no alerting on that service then we raised the ticket, linked it back to our guidelines and said, "Implement these alerts as a basic level." Did the same thing about Canary, same thing for each of the focus areas. So we ended up with a bunch of Jira tickets that then went out to all of the squads that owned each of these services. And then they sat in the backlog.

Jeremy I was gonna say and how did that go?

Nicole: Yeah, it didn't go too well. We went and talked to all of the product owners. We did demos on why we were putting in each of these tickets for each of these services, but it was still in a time when we had just formed these squads. We had just told them that they were owning these services. And now we were throwing tickets at them saying "You own this service now can you please fix it?" And so that sat for a couple months. And then we got onto the... We then tried to like reinvent our approach and go, okay let's focus on one overarching goal. We just want observability. We just need to know exactly what's going on in our platform. And that went really well.

So we just said to each of the squads, "Can you just implement our monitoring and traceability tool? Here's a guide on exactly how to do it." And we did a couple demos on here's all the benefits, here's all the information that gets pulled out of them and that really started to gain traction. I think within the first three months we had a couple services on, within six months we had even more. And now we're, so this was back in March, when we started that roll out. And now in December, we have all of our services into our monitoring tool. And not only that, the ones who adopted it early are in and creating their own dashboards and using these metrics already and of their own initiative. So we haven't given them any guidance or motivation to go and start actually owning and operating their services. We're still trying to focus on the last few to get your logs and your monitoring and the ones who adopted it early are already taking it and running with it which is amazing and exactly where we wanted to go.

Jeremy: It's amazing. All right, so then the other thing is, and this is probably people asking this question, is there's something called the Well-Architected Framework, you know about this and you know about the Serverless Lens for this. So I think this came slightly after you started working on this auditing process. So what did you do ... now AWS is publishing a bunch of standards saying, "Hey here's how we're supposed to do it." Now I think that it's very broad, right. It doesn't get into the weeds. It is very helpful but so how did you reconcile those eight focus areas you had with the Serverless Lens that came out?

Nicole: Yeah so it's quite interesting actually because preparing for our talk, the AWS Well-Architected Framework has been out for about four or five years already and it was across all of the AWS best practices and services and very, very generic. And it was really just a guideline, you came across it in the certifications and things like that. Then the Serverless Lens came along and it was really focused around best practices in the serverless space, but still something you only really encounter when you're getting guidance from architects or in the certification process. Then they launched the Well-Architected Tool, so different from Framework, where they integrated it into the AWS console and I think that was in 2018, I believe. And that was really just a checklist where you could check through and say you've considered these things. And then the bit that came just after we implemented our audit was the Serverless Lens as part of that tool.

So the Serverless Lens was around when we were creating this. It's just that we weren't overly aware of it. But when the blog post came out saying it was part of the tool, then we realized, oh we could use this and see maybe it can highlight some gaps in our audit process because we weren't really seeing the Serverless Lens as an audit. It was more like guidance on how to write a good service. And then we had to think about what our audit was. Oh, it's guidelines of how we build good services. So there was an overlap for sure. And when we did the comparison we realized actually we have two gaps or like two pillars that we haven't really covered in our serverless audit. And so that's giving us somewhere else to really go and define and call out for our engineering team.

And I mean, on reflection, the reason that those gaps are in our audit is because we're already doing them. So you know how I said that the guidance, like the Well-Architected and the Serverless Lens guidance, is built into the core principles of how you build on AWS well. And so because we've had a core AWS architect in the team from the beginning, because anyone who's gone through the certification process is vaguely aware of each of these pillars, you know operational excellency, performance, cost optimization, reliability, security, we're already building really well in those spaces and so we didn't feel the need to bring our older services up to date in those spaces. So that's why we haven't explicitly defined what we do in those spaces. We already have practices and patterns that are best practice for us. And so long as you're following those patterns you're doing okay. But for completeness, it's really good for us to add it to the standard because, I don't know, maybe we won't always have an AWS architect and who knows those guidelines, maybe we switch to a new service, they won't always be the same. So that will be like our next evolution of the audit.

Jeremy: Right and that's one of the things that I actually really liked about how you built your own standard first and then took the Well-Architected Framework and said where does this fit in? And maybe where do we have those gaps? 'Cause I think that is a good way to do it. I mean the Well-Architected Framework and the Serverless Lens specifically, I mean it's really just a bunch of questions, right. It isn't like here's exactly how you do it. It's just have you considered these things and you kind of check them off. So let's just quickly talk about that, where those overlaps were and how you took one of these pillars and said okay we're gonna go a little bit deeper and we're gonna add additional standards to it. So let's start with the operational excellency pillar. So what else did you add, sort of is part of your audit on top of that?

Nicole: Yeah so it's really about what did we define as within this pillar? So within that pillar there were one, two, three, four, five focus areas. So alerting, observability, and logging, so everything around operations, and then also the bits around how to get your service into production. So the deployment part, deployment mechanism and integration testing. So all of those five focus areas were within the operational excellency pillar. And that's because that was always held centrally within the infrastructure team, within the platform squad. And so we really had to define this is how we do it because we were about to get all of the application engineers to do it on their own. So that's why we had so much definition within that pillar.

Jeremy: Right and then the security pillar.

Nicole: Yep, so security pillar would be the secrets management focus area. Really most of the way that we were building was already following some pretty secure patterns. I mean it's an E-commerce shop, we have to be pretty secure from the get-go. And it was really that tidy up of how we manage our secrets.

Jeremy: Okay and then the reliability pillar.

Nicole: Yep so reliability and the next one, the performance pillar, are the ones where we haven't defined guidance around them because as I said before, our services were built to be pretty performant and pretty reliable. E-commerce throwing them through Black Friday, Cyber Monday.

Jeremy: Right, now is that something within that pillar though, I mean thinking about resiliency or chaos engineering, things like that or just even request rates and stuff like that. That's all stuff though that's sort of being considered within your audit, right?

Nicole: So again within our audit explicitly is areas of the services that we needed to bring up to speed for our older services. The ones that are being built now are following our existing patterns and so they're still quite implicit. So long as you're following one of the pattern, architectural patterns that we've set out, that we've proven work for us, then reliability and performance come along with that. And the next evolution of the audit we'll want to define and explore each of those areas. So we do have chaos engineering on the roadmap and that should be part of our exploration into the reliability pillar of defining what do we do now to increase our reliability but what can we do in the future? And so chaos engineering is going to be part of that.

Jeremy: Right, all right so then cost optimization. The cost optimization pillar. That's one of those things where, I mean maybe, I mean I've always thought about costs, right. 'Cause I've always been very early in companies and having to think about that and infrastructure costs. I think as companies get bigger, engineers don't typically think about that. And I had a guest, Eric Peterson, who said, "Every line of code an engineer writes for the cloud, they're making a cost decision," which I think is a brilliant way to think of it. So from a cost standpoint, what did you implement there?

Nicole: So within our audit we only have the Middleware component and specifically this adds to our cost optimization because we use the SSM Middleware. So SSM or the parameter store, there are two tiers; there's a free and a paid tier. And we've had to opt into the paid tier purely because of the rate limiting that is on the free tier. So by introducing Middy, by caching secrets when a container starts up, rather than on each Lambda invocation, saves us quite a lot. And we're doing a lot more in cost optimization on the infrastructure side, on the supporting infrastructure for our Fargate containers rather than our serverless services because they're pretty performant already. And so that's why we don't have too much more in the cost optimization area for our engineers to pick up on just yet.

Jeremy: Right, right and I know your team uses EventBridge quite a bit and I mean there's all kinds of pricing around, so I just think that's an interesting way or an interesting thing for developers to start thinking about is what services I use and how that's gonna affect overall costs when sort of building out and planning that stuff, so very interesting there. Okay, so you started getting people implementing it and it sounds like you're doing pretty well with that which is amazing 'cause that in and of itself getting your engineers to adopt something new is a huge hurdle. But what about some of the challenges people have faced or what are the responses you're getting from these developers? 'Cause I know they're doing it, right, but what has been the impact on performance or morale or just how people are dealing with it?

Nicole: Yeah and as a platform squad, our customers are the engineers. And so their feedback is paramount to how we approach anything really. And that's why we had to change our approach right at the start of instead of putting tickets into their backlog, we had to reframe it as, "Let's just know what's happening in our platform." That's something that everyone can align with, right. I mean if you're gonna put a service out there, you kind of want to, whoever's operating it, whether it's yourself or another team, you want to help them know what's going on. So that was an easy one for them to really align with and get on board with, which was great. And then as I mentioned before, off their own back, they then started using it themselves which is amazing. And also jumping on more of the audit categories and saying, "Okay, well that was really useful. What else is in this audit? What else can I do?" And so like half of the services have alerting on them now even though we haven't made that as our key focus to get rolled out. And so it was really that feedback of silence for the first part. And then they started getting on board and adopting it quicker and quicker. And then not only did they start implementing other parts of the audit, they started adding into the audit. So they added the documentation thing, that wasn't in there from the start.

There's also something about shared code. So when we have a monorepo structure and so shared code tends to be using learner named spacing. We would refer to it as "at namespace slash shared code package." And that's tricky when you look at how we release code because if you make an update to that package, it then triggers off the pipelines for all of those services. So it gets a bit messy checking what's going on. So we've published our shared code as private GitHub packages. And now one of the squads has added that into the audit of use the GitHub packages if you're gonna use any of the shared code. So that's something that as an infrastructure or platform squad, we don't really see that too much. And so it's really the engineers who are adding in those aspects of things that we wouldn't consider but are definitely part of really like best practices when you're writing services. It's a mesh of infrastructure and engineering, right. And so that's kind of the spectrum of reactions of how we've rolled it out.

And then one team has actually said, "Okay, we're gonna own this. We're gonna take all of these tickets that you've given us at the start of the year. And we're gonna put them into a roadmap and a timeline." And they listed out all of their services, all of the audit categories, added in their own, that were really just opinionated as part of their squad. So they decided that they wanted all of their functions to be named. They didn't want any default exports which yeah good thing to have. And they put it into a timeline, they set themselves deadlines. This was not imposed from any other squad let alone the platform squad and they delivered it. And now all of their services are pretty compliant with the audit. And they inherited, I think, three or four old services that they knew nothing about beforehand. And they learned the services, they put in the test, all of the things to bring it up to speed and now they can operate the services.

Jeremy: I'm just curious 'cause you said I think you had 36 services that you were applying the audit to, and I'm sure they were all in different levels of compliance just because as they were being built. But how did you balance and hopefully you can answer this, but how did the teams balance that idea of going back and refactoring all this code as well as continuing to move forward to support new features and things like that? Like how do you get to that sort of operational maturity but still balance this idea of still releasing new features?

Nicole: Yeah, that's something that we've had to tackle with each squad. So each squad has their own PO with their own, I guess idea of what should be prioritized, delivering features or engineering work, and how you strike that balance. And each squad has found their own balance. I don't think any two squads have been the same. So I mentioned that one squad who created their tech investment roadmap. They were building features all the way up until I think it was September, and then they switched and said, okay we're going to bring in, I think it was like 30% tech investment for every single sprint, until they got it done. And so they were fairly structured in their approach. There are some squads who just brought in one or two tickets every sprint. There were some squads who didn't pick it up at all and have kind of made the last push right at the end. So it was we really left it up to each of the squads to figure out how they wanted to strike that balance with the one overriding goal of, we just want to know what's going on in the platform if you can help us there as a minimum, that's great. And all of the squads helped us achieve that, and some of the squads went way beyond.

Jeremy: Yeah, that's awesome. All right. So let's talk about the Middleware again for a second because this is another thing where, I know it was always a big frustration in the past was, you're bootstrapping new services and you're doing the same thing over and over and over again. And the promise of serverless, or at least one of the promises of serverless, was just write your code, focus on your code. And so there's a lot of different ways that you can bring in some of the bootstrapping stuff whether that's with layers or custom images and now they just released containers, so now you could have your own packages or your own containers or base containers that you could use. So you chose the Middleware route so how was that sort of implemented and what are the things that that does?

Nicole: So we chose Middleware before layers were even introduced. So that's another thing that maybe we'll switch over to layers, who knows. So at the moment what Middleware does is it gives us a consistent way to handle errors, to manage logging and also to manage secrets. So those are the three Middlewares that we have customized and have implemented on most of our services and the rest are covered by the service audit. So it's giving us a consistent logger so that when you're looking in our logging system which merges all of the logs for all of those applications together, you can filter it out consistently and performantly as well, because there's a way that you can log a structured logging where you can say, function name is this and logging level is that. And having that consistent across all of the services makes it a lot easier to go and cross-reference and find areas that are affecting multiple services.

Jeremy: Yeah so that's an interesting thing where, when you start standardizing things and again, the audit process I think, again it's clear cut, it makes a lot of sense. Like once people start using it, you said the developers were like, "Oh, well we could do more things." So you mentioned documentation as one of those things where the developers just said, "Hey you know what, this is great. But we need to add some documentation to this process." I mean did the development teams come up with other improvements? Like did they adopt new standards or other things that they kind of pushed back up to the top?

Nicole: Yeah, so I mentioned the thing around shared code and introducing the standard to use published packages. I'm trying to think of if there was more, we added in another one around doing um, no, I'm not sure about that one. Yeah, I think the main one that was added was around how we handle shared code. And it's really exciting to see the engineers are actually going back and feeling that empowerment to go and update or even add to an audit. Some of the focus areas that we had written from a platform perspective they went and rewrote after we had paired and taught them what to do. They wrote it in a way that made way more sense to them because it's different when you've come from an infrastructure background and when you've come from an engineering or development background. The words don't have the same meanings and the way that you phrase things is slightly different. And so rewriting it into their own words, means that it's going to be a lot easier for the next engineer who comes along who has to adopt it. So those have been some of the really great contributions back into the audit.

And we're hoping that it's not something that the platform team have to add to any more at all. We're hoping that it's going to be maintained by the application engineers, who you know, they're the ones who have been building up all of these patterns and practices already. We kind of just put it into our own words and said, "Here we wrote it down," and I'm hoping that as they keep building and improving their services they keep adding it in. I mean one of the squads I know is using Lambda pre-warming on their services. So rather than using provision concurrency, they actually have a package that's going and triggering their Lambda to make sure that at least one or two containers are warm. That's something that if they find that it's really giving them an improvement on their cold start times, maybe they'll add it into the audit for specifically for services that aren't invoked often, but often enough to have like a really big impact on cold start time. Maybe they'll add it into the audit too. So we're hoping it becomes a living thing that the engineering team keep and uses, I guess, a way for them to communicate with each other. 'Cause they're different product squads, right, we need to try and figure out how to share the knowledge that each of the different squads are gaining when they're building new things. So this is one mechanism to try and centralize that.

Jeremy: Right, now have you standardized around like a particular programming language or is it just whatever the squad wants to use?

Nicole: So we started off with one or two squads, right. And so we said everything is written in Node.js and the front end has React and Node.js. So that's the languages that have been selected and we're sticking with them.

Jeremy: Right, it seems to be popular. Although some people do like to jump into, Go or Python or something like that.

Nicole: Yeah, I mean every language is fit for a certain purpose, right? So if we run into a situation where one of the other languages suits us better, because it's balancing that need of being able to have any engineer maintain any part of the platform. So having everything written in a standard language or does this thing really need to be performant, does it really need to do a specific thing? In which case it kind of has to be written in another language, then that's when we'll start making those calls. But at the moment we want every engineer to maintain any part of the platform.

Jeremy: Yeah and I like that. I like the fact that the front end and the back end are both essentially written in JavaScript. It's just that way, especially with cross-functional teams, I mean even if somebody has to go in and make a small change on either side, I mean it's just good to have. So all right that's awesome. So what about sort of advice to others? And then when I say others, I mean other engineers maybe, that are getting in audit or some sort of thing like, "Hey you need to do this." But also to the teams that were like you, that are trying to implement these standards to make the quality of your applications better, to make the standards better, the best practices better, so what's some of that advice that you might be able to give to people.

Nicole: I mean, the core advice, whether you're on the receiving end or the implementation end of this kind of thing is empowerment. So feel that you are empowered to really make a difference and make your services better. And don't feel like anything is ... don't try to make it an imposed standard, make it something that is there to help. Because the main thing around this checklist is if you do these things, you can own and operate your services. You'll be able to find the bugs quickly when you're under pressure. That's what we're trying to empower our engineers to have in their mindset, of you can build good services and here's our tips and tricks on how to do that. If you follow these things, we know that they work. And so it's really about taking that line of we're here to help, we're here to empower you to do good things and do great things. And not have to try and guess that, if I do this, is that going to work? Or if I do this, will that work with the whole stack that we've got? Because no one can know every part of the platform, right.

And so that's really been the consistent message that I would want to get across with this process of, if you're implementing it, make sure it comes through from an empowerment and an enabling perspective of, this is what you should know and here's some starting points on how to do it. And if you're on the receiving end, maybe don't wait to be on the receiving end. Maybe start writing down what your best practices are and say, start sharing it out, start sharing the knowledge because once you start leveling up as a team, then you know your whole platform benefits, right. So yeah and also keep it as a living document, best practices don't stay still. So I mentioned that we implemented Middleware before layers were even introduced. We now have a layer as part of the standard because our monitoring system has written a layer. And so the standard has been evolved to implement this layer, implement Middleware for other things. Maybe the things in Middleware will move into layers as well. So keep these up to date, keep them as the document of, here's what a good service is for your team. And then any new joiners, anyone looking at the legacy serverless applications if you have them, knows what to do.

Jeremy: Love that, great advice. All right, last thing. You said it's a living document or it's an ongoing thing. You mentioned chaos engineering on your roadmap. What are the other things that the LEGO Group is considering adding to this process?

Nicole: So this process started at the start of this year, right? It was a big table in confluence. It's very hard coded. It's very manual. The most automation we have is we've linked Jira tickets, so that they automatically get ticked off when they're done. It's not great. The actual implementation of the process is not great. The outcomes are amazing. So what we want to do with this, is actually figure out how we can show all of our engineers, all of the services they own, what state they're in according to our audit. So you can maybe give each service a score and say you're lagging behind a little bit on operations. So it's written really well but you're gonna find it a bit hard to operate. Or maybe it's written well, it's really easy to operate, but a new person will have no idea what's going on when they join so get that documentation up. We want to try and figure out how to get that feedback back to our squads in an application that's not yet another application, right. So we're trying to figure out where that fits within our ecosystem. And then I've got this secret, it's not really a secret anymore. I talked about it in a conference. I want to put a leaderboard in. So I want to make it competitive. I want to make it a point of pride of, my service is the best in the platform, that kind of thing. And I'm hoping that that will put a bit more fun around an audit process, you know, and really drive our platform to continually be better. And I mean if a new team introduces a new standard, that then tanks all of the other squads' scores, I mean I'm all up for it, you're continually improving, right.

Jeremy: There we go. Awesome. Well, that's amazing and like I said, honestly if you're listening to this and you're thinking about anything even remotely close to this, this is a great sort of roadmap for you. And I think the talk from ServerlessDays Virtual is up. You can get that, I think search serverlessdays.io or virtual.serverlessdays.io. You can find that. So again, Nicole thank you so much for being here, for sharing this, for continuing to do this great work. And again putting this information out there is awesome because again, I think it can help a lot of teams. So if people want to get a hold of you, how do they do that?

Nicole: Yep so the best place to find me is on Twitter @Pelicanpie88. I know it doesn't resemble my name but I picked it several years ago and LinkedIn as well if you want to talk to me on a professional level, but mainly Twitter for AWS and serverless stuff. And I guess the main message here is DevOps as a journey, infrastructure is a journey, where I'm just kind of sharing where we are right now. And I'd love to hear more stories about where everyone else is at the moment in their journey.

Jeremy: Yeah, it's amazing. All right and then don't forget to check out the LEGO Engineering Medium blog, that's at medium.com/lego-engineering and then of course, lego.com. This will probably be after Christmas, you won't have time to put your orders in but maybe some January LEGO purchases for the kids and family. So again, Nicole thank you so much. I'll put all this stuff in the show notes. It was great having you.

Nicole: Thank you for having me, this was a great talk.

View Details

Watch this episode on YouTube: https://youtu.be/QVauc83L8WU

Transcript

Episode #30: What to expect from serverless in 2020 with James Beswick
Yeah, it's really snowballing in terms of popularity and certainly seeing just the sheer number of people from all these different companies. You have startups and enterprises and so many different types of industry all starting to pick up serverless tools. And a lot of things that we talked about just a year ago, that really seem an incredibly long time ago now, the conversations that don't really necessarily matter that much anymore.

There was a discussion about what is serverless and all these sorts of things. And now we're starting to talk about architectural patterns, and starting to talk how it's not just lambda anymore. Serverless is this concept of taking different services from different providers and combining them. So I think, you know, we see people building things where you connect API Gateway, DynamoDB, S3, but also with services like Stripe or with Auth0 and then Lambda is just connecting things in the middle.

Episode #57: Building Serverless Applications using Webiny with Sven Al Hamad
So when you're a small business, it's the cost of infrastructure that really matters to you because it's really efficient. You don't pay if you're not using it. But for the big guys, it's a combination of factors. And sure your bill might be slightly higher in some cases running on serverless, the cost of infrastructure. But the cost of managing infrastructure will go way down. You will have to hire less people, or the people you have will have to spend less hours working there.

But also what that does, it releases a big chunk of the budget or resources or man hours that you can now focus on product iterations. So your product can grow faster. And if your product grows faster, you can out innovate potentially your competitors, which can't afford that same level of innovation. So what I see with enterprises is that they see serverless as a competitive advantage, and that's why they moving to serverless. Although, you see all the blog posts about cost savings and stuff like that. Yes, that's true, but there's that agenda of outpacing my competitor, which serverless actually unlocks. And the moment you migrate to serverless, you can use that potential.Episode #35: Advanced NoSQL Data Modeling in DynamoDB with Rick Houlihan (Part 2)I mean, I get the question of is DynamoDB powerful enough for my app? Well, absolutely. As a matter of fact, it's the most scaled out NoSQL database in the world, nothing does anything like what DynamoDB has delivered. I know single tables delivering over 11 million WCUs. It's absolutely phenomenal and then the other question is is not DynamoDB too much overkill for the application that I'm building?

I think we can have great examples across the CDO of services. Not every one of our services is massively scaled out. Hell, I've got services out there, I've got five gigabytes of data and they're all using DynamoDB and the reason why I used to think that NoSQL was the domain at the large scaled out high performance application, but with cloud native NoSQL, when you look at the consumption based pricing and the pay per use and auto scaling and on demand, I just think you'd be crazy.

If you have an OLTP application, you'd be crazy to deploy on anything else because you're just going to pay a fraction of the cost. I mean, literally, whatever that EC2 instance cost you, I will charge you 10% to run the same workload on DynamoDB.

Episode #44: Data Modeling Strategies from The DynamoDB Book with Alex DeBrie
I introduced the concept of item collections and their importance pretty early on. I think it's in chapter two. And it was actually one of the solutions architects at AWS named Pete Naylor that that turned me on to this and really made me key into its importance.

But the idea behind item collections is you're writing all these items into DynamoDB, records are called items in DynamoDB. And all the items that have the same partition key are going to be grouped together in the same partition into what's called an item collection. And you can do different operations on those item collections, including reading a bunch of those items in a single request.

So as you're handling these different access patterns, what you're doing is you're basically just creating these different item collections that handle your access patterns. And that can be a join like access pattern. If you want to have a parent entity and some related entities in a one to many or many to many relationship, you can model those into an item collection and fetch all those in one request.

You can also handle different filtering mechanisms within an item collection, you can handle specific sorting requirements within an item collection. But you really need to think about, hey, what I'm doing is I'm building these item collections to handle my access patterns specifically.

Episode #79: What to do with your data in a serverless world with Angela Timofte
So the scenario was that we have, so people can sign up, but then they have to activate their account. That’s quite, like, a normal scenario, right? So they have to activate and if they don't have to be in like 30 days then we need to delete the account. And we're doing that in our only Mongo database for where we're keeping all the data for consumers. And of course, we're putting a lot of load, unnecessary load, on our primary database. So we decided to actually take this entire scenario out and we started, okay, of using events when consumers sign up. We will send an event to store some data in a DynamoDB which would say this consumer signed up and then we'll have another event coming from the activate ... like the activation API, saying this consumer activated, so then we'll delete the data in DynamoDB and we had one Dynamodb with all the unactivated accounts. And then from there we could look at, like, when the account was created and we can delete whatever accounts that are not activated in time. So this way we took that whole pipeline to serverless in its own context and, like, its own service and then doing it’s spin there separate from our primary data. And we did it with, like, three events and DynamoDB and then, yeah, another Lambda that was listening to ... was querying this database.

So it was a very simple scenario but we took a lot of load from the main database by not going like every, I think was like every day, queried the database to get like all un-activated accounts. And so, yeah, it was a very simple scenario, but like this just shows how you don't have to, like, refactor your whole database. You can just take parts of it or, like, queries like whatever it … This was just a scenario and we took it out and its own being ... I haven't checked it in, like, a very long time because it's just working, you know? I'm thinking maybe I should go and check it. No, but, like, that's like one example, where, as I said, like, you don't have to refactor the entire thing.
Episode #33: The Frontlines of Serverless with Yan Cui
I don't know about the major breakthrough, but I definitely think more education and more guidance, not just in terms of what these features do, but also when to use them and how to choose between different event triggers. That's a question I get all the time. ""How do I decide when to use API gateway versus AOB? How do I choose between SNS, SQS, Kinesis, DynamoDB Streams, EventBridge, IoT Core. That's just six application integration services off the top of my head. There's just no guidance around any of that stuff and it's really difficult for someone new coming into this space to understand all the ins and outs and trade offs between SNS and SQS and Kinesis and so on.

Having more education around that, having more official guidance from AWS around that, that would be really useful. In terms of technology wise, I think I like the trajectory that AWS has been on. No flashy new things but rather continuously solving those day to day annoyances, the day to day problems that people run into. The whole cold start thing, again, often overplayed, often underplayed it's never as good as some people say, it's never as bad as some other people say. But having some solutions for people with real problems, where with cold starts we speak of various different reasons.

I really like what you've done with provision concurrency, even if I think the implementation is still, I guess it's a version one. So hopefully some of the kinks that they currently have would be solved. Other than that, I'd like to see them do more with some multi account management side of things. A control tower is great, but again, there's a lot of clicking stuff in the console to get anything set up, and it's also very easy to rack up a pretty big bill if you're not careful you can provision a lot.

NAT gateway for example and things like that. One of the companies I've been talking to recently as well, a Dutch bank, they are actually doing some really cool tool themselves to essentially give you infrastructure as codes. Think of it as a CloudFormation extension that allows you to capture your entire org. Imagine I have a resource type that's defines my org and the different accounts and then when they configure CloudTrail set up for multi-cloud to configure security guard and things like that all within my cell template, which looks just like CloudFormation. So some really amazing tool that those guys have built.

But having something like that from AWS would be pretty amazing as well. Because again, we've seen more and more people getting to the point where they have a very complex ecosystem of lots of different enterprise accounts, managing them and setting up the right things. The STPs and things like that. It's not easy and we certainly don't want people to be constantly going to the console and clicking things. And that's another annoyance I constantly have with AWS documentations is, they keep talking about infrastructure as codes, but every single documentation just tell us, go to this console, click this button.

Episode #37: The State of Serverless Education with Dr. Peter Sbarski
I think that's going to be what makes education effective in the future. It's that curated personalized education. You spoke about A Cloud Guru, we have full time training architects, instructors, who what they do every day, right, is they create content. Whenever anything changes, they update it. Right?

So when you go to the platform, you know that what you're getting is the latest version. You're getting that latest best practice. So suddenly, what you are learning in a lecture hall, right, doesn't really match what you could be learning online because, that content is much more up to date. So that's an interesting aspect as well. Yeah, the currency and the quality. Yeah. Because we can continuously iterate on it. I think that universities too have an important function that cannot be done with just an online delivery of education.

That element is really that ... It's going to sound harsh, but it's babysitting, right? Because just after you finish school, right? There's still a little bit of time for a lot of people to mature, right? They need to go through that maturation phase. Going to university, going to college allows people to do that, right? It allows them to build social connections. It allows them to learn how to work in a team, maybe better than they did at school. So it gives them that opportunity to mature before they go into the industry.

As much as I love online education, and I think it is the future, there is that element that still needs to be solved, that social element. But I think we'll figure things out, maybe it'll be some blended learning, where you do get that up to date curated delivery of education online. Then there's an additional element where you go and you socialize with your peers. So yeah, we'll see how all that pans out.
Episode #60: Going Green with Serverless with Paul Johnston (Part 2)
In the end, I know the joke is cloud is just other people's servers and all that kind of stuff. It's always underneath it. There's just servers and there's just servers. But I think that trying to make these servers more efficient, trying to make these data centers more efficient, there is still constant churn. We don't keep things efficient. Two, three years down the line, the server that you were using is not efficient. Six years, seven years, it's old. You don't want to be running stuff on there, you want to be running stuff on something that's efficient and new. Actually, there's an enormous amount of e-waste in terms of the data center industry. It's not straightforward.

The conversations around all of this are not straightforward. I think everyone needs to start thinking about moving to the cloud simply because we need to be reducing our impact. If you're running stuff, I think it's important to be able to go, "Actually, we need to be able to reduce the amount we run." But that means, understanding how that cloud, that you're choosing to work with, is working in terms of its sustainability. You can't just go, "We'll move it to X cloud, or Y cloud or Z cloud, or whoever it is, but we'll trust them to do the right thing."

You've got to still have that relationship. You've got to still be able to go that cloud, "You, Mr. or Mrs. Cloud person, you've got to tell me, are you using green electricity? Are you using renewables? How are you disposing of everything? What is your supply chain?" I think that conversation over the next few years is actually going to become a much more common conversation. It's going to become more important. You are not going to be able to get away with, "We just run efficient data centers." That's not going to be the standard and reasonable response. That's going to be a table stakes. Green data center will be a table stakes conversation, and the best practice will be, "Well, we're actually running 100% renewables and we're putting more into the grid, and we're being as good a partner as we possibly can. And all of that. We haven't got diesel generators, we've got batteries."

It's all of that conversation that I think comes back to. Maybe we will end up not using certain companies because their data centers are not green enough. Maybe that is where we end up, that actually societal pressure actually pushes these companies to do better. But I don't think we're there yet. I think we're probably a couple of years away, two, three, or four maybe, away from that.
Episode #42: Better Serverless Microservices using Domain Driven Design with Susanne Kaiser
...As mentioned that a domain model cannot exist without a boundary, then that's where we come to bounded context and a bounded context provides different types of boundaries for a domain model so it forms, it form a consistency boundary around the domain model and protects its integrity and it could also form a linguistic and semantic boundary so that the language's terms are only consistent inside of its bounded context. So for example, pending in one bounded context could have a different meaning than another bounded context, for example. And it also serves as an ownership boundary, so for example, bounded context could be implemented and evolved and maintained by one team only and a single team can, on the other hand, can also own multiple bounded context but it's really, really relevant that multiple teams are working on the same bounded context, because this enables a ton of teams working at their bounded contexts independently at their own pace and with minimal impact across other teams.

And this also serves as a physical boundary and can be implemented as a separate solution and can be deployed independently as separate artifacts and also enables separate data stores which are not accessible by other bounded contexts and, for example, also the source code could, of each bounded context, can be maintained in separate git repositories with their own CICD pipeline.
Episode #65: Serverless Transformation at AWS with Holly Mesrobian
Yeah. We recommend using a separate account per microservice and then also thinking about an account for each of your environments as well, your pre-production environment and your prod environment. Each one should have its own account as well. What that does for you, if you think about it, a lot of times, two pizza teams own a service or a small set of microservices, and you want to reduce the number of people who can actually access those services and make changes. I mean, it's an operational risk.

It's also a security risk having too many people have their hands on a microservice. You really want to make sure that the people who can access it are knowledgeable and know what they're doing. That will help you have a high availability as well as ensuring security. Of course, availability comes back to not only potential for someone to make a change that is a breaking change, but also things like ensuring that your limits are used and planned for in a way that makes sense for you.
Episode #36: The Cloud Database Landscape with Suphatra Rufo
Yeah. I think this is where things get really interesting. When I was at AWS, I worked almost exclusively on my creations off of Oracle and Azure at AWS. And a database migration is, by and large, the most difficult thing that you can do in cloud computing. It's really hard. You've got to do a lot of data modeling. You've got to do your schema conversions. I mean, it's really just a ton of work and what I have found is that when people are charged with, "All right. We got to migrate our database." We tend to do it in multiple phases and that will take multiple years, so oftentimes they'll first just re-host. Let's say they're on Oracle. They want to get off Oracle, but they don't want to be penalized. So they take their Oracle license and bring it to a different cloud provider. They keep all their data with Oracle still. They're just moving it. That takes six months to a year, then afterwards, they say, "Okay. Well, I think we're now going to replatform." And that's a whole nother workload and that's even more work, and even harder down the line is refactoring, which is where they might actually go from a relational database to a NoSQL database.

It's much more rare that you see people to a database migration where they go from a traditional relational database on one provider to a NoSQL database on an another provider because it's a really difficult piece of work.
Episode #39: Big Data and Serverless with Lynn Langit
Well, it goes to the CAP theorem, which is consistency, availability and partitioning. This is sort of classic database ... what are the capabilities of a database? And it's really kind of common sense. A database can have two of the three but not all three. So you can have basically the ability for transactions which is relational databases or you can have the ability to add partitions is really kind of to simplify it easily. Because if you think about it, when you're adding partitions, you're adding redundancy. It's a trade off. And so are you adding partitions for scalability? And so when adding partitions makes a relational database too slow, then what do you do? So what you then do is you partition the data in the database to SQL and NoSQL.

And again, I did a whole bunch of work back in 2011, 2012, 2013. I worked with MongoDB, I worked with Redis. And one of the sort of key, I don't know, books I guess, would be Seven Databases in Seven Weeks. It's still very valid book even though it's many years old. It tells how you do that progression and really turn the light on for me, because prior to that point it was, oh, just scale out your SQL Server, scale out your Oracle Server, which still would work but these NoSQL databases were providing much more economical alternatives. And of course I'm always trying to provide the best value to my customer. So if it wasn't a great value to buy more licenses for SQL Server or for Oracle, rather you want to get a Mongo Cluster up or a Redis Cluster up, you could partition your data if that was possible because there's cost to partitioning your data and writing your application.

So I just found those trade offs really, really fascinating. And of course during that time, cloud was launched, led by AWS. Microsoft had an offering, but they didn't really understand the market until a little bit later. So Amazon had an offering and they first started, it was really interesting. They started by just lift and shift with RDS at a PaaS level taking SQL Server and actually making it run effectively in the cloud. That was how I got started, because my customers wanted to lift and shift and maybe go to an enterprise edition and run it on cloud scale servers.
Episode #71: Serverless Privacy & Compliance with Mark Nunnikhoven (PART 1)
Yeah. And the tiering system is frustrating as it is for a lot of users, a lot of it does have that. If we use the AWS term, it's about reducing the blast radius. You don't want everyone in support to be able to blow up everything, and if you look at the Twitter hack was actually an interesting example, somebody raised the question and said, "Why didn't the president's account get hacked?", "Why wasn't it used as part of this?" And because it has additional protections around it, because, it's the leader of the free world ostensibly so, you want to make sure that that's not the average, temporary employee on a support contract, being able to adjust that. So the tiering actually is a strong play, but also understanding that the defense in-depth is something we talk about a lot in security. And it gets kind of a bad rap, but essentially it means don't put all your eggs in one basket.

So don't use one control to stop just one thing. So you want to do separation of duties. You want to have multiple controls to make sure that not everybody can do certain things, but you also want to still maintain that good customer service. And I think that's where, again, it comes down to a very pragmatic business decision. If you have two sprints to get something out the door and you go, well, I'm going to build a proper admin tool, or you're just going to write a simple command that your team can run, that will give them the access, you're just going to write a command that does the job. And you know what, in your head, you always say the same thing.

You put it in your ticket notes, you put it in your Jira and you say, we'll come back to this and fix it later. Later never happens, so most admin tools are this hack collection of stuff just to get the job done. And I totally get it from a business perspective. It makes sense. You need to go that route, but from a security and privacy perspective, you need to really think holistically. And I think this is a question I get asked often, actually, somebody just asked me this on my YouTube channel the other day, they said, "I'm looking for a cybersecurity degree, and I can't find one. All I can find is information security. What's the deal?" And I said, well, actually, what you're looking for is information security. In the industry, and especially in the vendor space, we talk cybersecurity because that's typically the system security.

So locking down your laptop, locking down your tablet, locking down your Lambda function, that's cybersecurity, because we're taking some sort of cyber thing and applying security controls to it. Information security is an academic study, as a field of study in general, is looking at the flow of information as it transits through systems. Well, part of those systems are people, are the wetware. Right? Or the fact that people print it out. This is a big challenge with the work from home was, you said, well, your home environment isn't necessarily secure. And you said, well, yeah, it has different risk models. But the fact that I can connect into my corporate system and download a bunch of stuff and then print it, that's information, that's still needs to be protected.

So I think if you think information security, you tend to start to include these people and go, wait a minute, Joe from support, we're paying him 15 bucks an hour, but he's got a mountain of student debt. He's never going to get out of it. That's a vulnerability that we need to address, not from locking it down, but help that person out and make them feel included, make them feel, as part of the team so that they're not a risk when a cyber criminal rolls up with some cash and says, Hey, give me access to the support tools.
Episode #52: The Past, Present, and Future of Serverless with Tim Wagner
I have these two strong reactions to that statement, right? One of them is I would say in some ways the most successful thing Lambda has done is to challenge thinking, right? To get people to say, do you really need a server stood up, turned on taking 20 minutes to fire up with a bazillion libraries on it and then you have to keep that thing alive and in perfect condition for its entire life cycle in order to get something done in terms of a practical enterprise application? And challenging that assumption is one of the most exciting, important and successful things that I think Lambda and other serverless offerings have accomplished in our industry. The flip side to this is to be useful, sometimes you have to be practical. And it's equally true that you can't walk up to an enterprise and say, "All right, step one, let's throw all your stuff away and then step two, you're not going to get past step one."

It's funny, we talk about greenfields, brownfields, it's all brown in the enterprise. Even if you write a net new Lambda function, it's running against existing storage, existing data, existing APIs, whatever that is. Nothing is ever completely de novo. And so I think to be successful and be as adopted as possible in the long run, serverless offerings are going to also have to be, they're going to have to be flexible. And I think you see this with things like provision capacity. I mean, when I was at Lambda still, we had long painful debates about is this the right thing to do? And for understandable reasons, because it is less stateless. It took the ... it's obviously optional. We don't force anyone to use it. But by doing it, it makes Lambda look more like a conventional, well, server container, conventional application approach because there is this piece that is a little bit stateful now.

And I think the arc here is for the serverless offerings to not lose their way, to find this kind of middle ground that is useful enough to the enterprises that still challenges assumptions that gets people to write stuff in a way that is better than what came before and doesn't pander completely to just make it feel like a server. But is also practical and helps enterprises get their job done instead of just telling them that ... because just sermonizing to them is also not the right way to do it.
Episode #78: Statefulness and Serverless with Rodric RabbahI think accessibility of the platform. And I remember when I first met you, right, we had this conversation about, we called it “serverless bubble” at the time, right, and maybe “bubble” isn't the right word because bubbles burst and that's not a good thing. Maybe “echo” chamber is better. But I think … one thing I've learned, and I learned this very early on when I left IBM sort of went to a developer conference at, yeah, there's a thing called serverless, the greatest thing, and it was like what's a micro service? Right? Instead of recognizing that the world hasn't yet caught on. There is part of, you know, the technology community that has sort of, you know, good for that. But recognizing that there are still a large interest in Kubernetes, still a large interest didn't EC2 instances in VMs. There's a massive world out there where building applications for the cloud is still hard. You know, just log onto the Amazon console and look at everything you can get. Where do you get started? Right? So the opportunity for us is making the cloud more accessible.

And so we like to think that from a Nimbella perspective, you can create an account within 60 seconds. You can deploy your first project, you know, not even having to install any tools, right out of GitHub. And hey, I have stood up an entire application. It's got a front end. It's got a dedicated domain. It's served from a CDN. My functions are entirely serverless, they scale. I can have state. I just did that, right. So, it's about really making the cloud accessible for a large class of developers from the enterprise, all the way to the indie developer who just has an idea for a mobile app or a website that they want to build. I think this is where really the opportunity is, you know, whether you're running things in a container or an isolate like Cloudflare does. It comes with implementation detail nobody's going to care about in the future.
Episode #67: The Story of the Serverless Framework with Austen Collins (PART 2)It's the potential that's democratized for everybody, whether you're a large organization, or you're just a solo hacker, like in the basement or something, trying to get something off the ground like this power has democratized everybody. And that going back to our mission, like, yeah, we want to help every single person build more, manage less, leverage higher levels of abstraction, help them focus on outcomes more than ever. We're going to try and rethink developer tools and what that means in order to deliver that experience. And then the last part for us is just we firmly believe serverless is bigger than any one vendor at the end of the day. And we feel very strongly that there needs to be an application framework that provides an open level playing field for serverless cloud infrastructure across any vendors, because yes, we've talked a lot about AWS and the majority of our users are using AWS.

And the majority of the infrastructure is AWS, but not all of it actually. They are still bringing out, our users, our audience are very product focused. And if you want to build the best products, you got to be free to use the best of breed services that are out there. And so we see a lot of people still bringing in Stripe, still bringing in Algolia, still bringing in MongoDB Atlas, Twilio, right? There's so many great things out there. And helping people, developers have this, again, this open framework where it treats all these things as neutral. This level playing field where they could compose serverless infrastructure across any vendor into applications really, really easy. It feels like the destiny of the Serverless Framework to us.
Episode #58: Observing Serverless Observability with Erica Windisch
From a perspective of open source developers though, my biggest issue is the culture. Every one of these open source projects or projects, however small or big that they are. Because I think, I said things like Kubernetes right? Are now multiple projects. You have things like Falco and so forth that are sub projects or adjacent projects or however you want to define them. But you have a community here, that operates a certain way, they have their own culture. And that culture is different, potentially than a culture that you as a company founder or as HR or a manager, or whoever of a company, that has not necessarily the same culture that you want your company to have, or your team to have, that is in the open source. Right? And how do you kind of resolve that difference because, one of the other things is that a lot of people hire from these open source communities.

So if you are building a team that is going to work in open source, and you want to make this a diverse team, for instance. But it's not a diverse project. How does that work? Right? Is the project and the other people in that project, going to discriminate against you, either implicitly or explicitly. It may not be intentional, right? There are implicit biases that exist. And I think it becomes very difficult because, when you have your own closed source application, and you're building things for your own self and your own teams, you have control over what you're building, how you're building and the construction of your team, etc. And I think that you lose a lot of that, when you're working in an open community.

Because if you're only working on open source, it's almost like while you're employed by one company, your co-workers are almost in a sense, a set of people that are not hired by your company. That may not actually hold the same values that you or your company holds. And I don't have a solution for this. But it's something I think about a lot. And it's one of the reasons I no longer really contribute much to open source.
Episode #50: Static First Using Serverless Front-ends with Guillermo RauchI think what's amazing about serverless is that it's exposed the essential complexity of the problem. It stopped developers from sweeping hacks under the rug. The best example that, I think, from this is you can no longer do async computation as a result of invoking a function that easily anymore. In the world of no JS, I would see a lot of customers just put lots of state in a process. When they respond, they continue doing things behind the scenes in that same process.

Functions have altogether made this impossible, but for a great reason, right? They were exposing, "Hey, that side effect that you were computing, you should have not been doing in that same process. You should have used a primitive like a queue to put your side effect, your event there, and then use other functions that respond to that event." Then it's so smart that they also put the developer into this state of success of saying, "Well, if it's a side effect that now can no longer be retried by the client," because the client is executing the function.

The function is responding with 200, and it queued the side effect, so there's no reason for the client to retry. Now, the side effect is loose in the universe of computation. That means that we need a system that can retry it because we want that side effect to run to fruition. Now, it forces you to put that into a queue, and queue can retry and then eventually also fail and go into a dead letter queue. So now just like going through all this in my head, I'm going crazy about the amount of complexity.

But here's the thing, and this is why I love serverless. That was to begin with the essential complexity that had to be managed to begin with. What we were doing before was chaos, was side effects that maybe sometimes run correctly and sometimes not, was unscalable systems and so on and so forth, but it is a complicated world.
Episode #40: HTTP APIs for API Gateway with Eric Johnson and Alan Tan
The most dangerous part of an application that I'll ever build is my code. Right?

So, when I build an application, I want to get that data stored first. That's the thing. I tend to go DynamoDB because that's what I like, that's what I use, but there's different purposes.

I know Jeremy, you and I have had this conversation before, and you're an SQS guy, so that's where you tend to go, and we do this because we look at okay what's the pattern for the retry or the DLQ or different things like that.

For me it's because I'm going to continually write back to Dynamo. On the app, I'm specifically thinking about it. But the idea is if API Gateway can directly integrate with the storage, be it S3, be it DynamoDB, something like that, then I've stored the data and I don't have to go back to the customer if my logic fails, right?

So, in an application I've stored the data, let's say I'm using DynamoDB, I do a stream, it triggers a Lambda, I start processing that data. If somewhere in there, something breaks, and again, it's going to be my code, but let's say something breaks, then I don't have to go back to the customer and say hey guess what, I blew it.

Can you give me your data again? Can you resubmit that? And continue to trust me, because I'm sure I won't lose it again. Instead, I've got that data stored, and I can write in some retry or take advantage of the retry from an SQS or an SNS or something like that.

So, I think it's a really cool pattern for building resilience into our application. Serverless comes with a lot of resilience anyway, that's how AWS has approached this on look as much as we'd like to say nothing ever breaks, let's write as if it does, right?

So, let's degrade gracefully. I think this adds even another layer of that, where I can degrade in my code and know hey I've still got the data. I can write some retry logic. I can use existing retry logic. I think it's a safer pattern.

It does require ... The storage first is the pattern I call it, but it requires asynchronously. What can I do after I've responded to the client and how do I work with them?
Episode #51: Globally Resilient Architectures with Adrian Hornsby
So a soft Time To Live is your requirement in terms of staleness, right? So you say, my Twitter trend lists, I want to refresh it every, let's say, every 30 seconds. So you give it a TTL of 30 seconds, a soft TTL of 30 seconds. So if my service requests the cache and the TTL, the soft TTL is expired, and everything is fine you go and query the service, right? But if my service doesn't answer at that moment, so you are, you've passed the soft TTL. Now, your downstream service doesn't give you the data. What do you do? Do you return a 404, or you actually fall back, and you say, alright, my soft TTL is expired, but I'm still within the hard TTL which is it's one hour, right?

And then you say, okay, your service returns the hard TTL and you say, "Oh, sorry, we just have one hour old data, because we're experiencing issue." So again, it's a possible degradation. And actually quite often cache could be used like this. I think it's all about how you create your cache and things like this and how you define your eviction and policies and things like this.
Episode #48: Serverless Developer Culture with Linda Nichols
I just started thinking about the fact that if I was developing something for the cloud, or just in general, if I start typing a lot, I pause and I go, okay, somebody's already written this. I'm not that clever. Not really, there's a lot of smart people in the world. There are a lot of people that code all the time. This is already done somewhere. It's either in a library or it's a service. And I talk to so many customers and people who they're like, "Oh, here's my great idea of a thing." And almost always I'm like, "No okay, so that is this." And I mean, even like messaging systems like Service Bus on Azure. I mean, there are so many developers that have tried to write messaging systems. And there are so many out there, there's so many people that tried to write Kafka. And they still are.

And sometimes I talk to people that are trying to create something, and they will say, "Okay, well, I'm going to put this in a function, this in a function." I'll say, "No, you don't need functions here. This is already a service, or you can already use something like Logic Apps, like you don't have to write any code." And, you just kind of connect some things together, or there's already built in services and that's still serverless, right? Like serverless is not just fast. Like, I don't have to write 100 Lambdas to be a serverless developer.
Episode #49: Things I Wish I Knew Before Migrating to the Cloud with Jared Short
I think it takes practice, right? You're giving up a lot of fundamental control that I think people are used to having, right? I can't walk into my data center, open a rack and turn off or turn on a server or pull wires or things. That's a huge fundamental shift for a lot of folks. And as we're migrating to people now, these days, that have never even walked into a rack of servers, we're having people that are coming out of college that AWS and going into ec2 and clicking launch instance is their concept of a server.

I think what we're starting to build towards in terms of this cloud native mindset is, we fundamentally can trust these larger providers to provide mostly good experiences, let me be careful there. Mostly good experiences around these cloud primitive services. And we have S3, which has kind of been referred to as one of the seventh or eighth wonder of the world. It's like this modern Marvel, right? That thing holds so much data and performs so well, and it's so scalable.

When it goes down the internet is just basically done. That's incredible that they have this service and we're trusting it. As cloud-natives, we're trusting these providers. I don't care if it's Azure or GCP or anybody, to provide these primitives that we can build on top of it. I think cloud-natives look at those primitives and you have an implied level of trust, and you're willing to build businesses and business value on top of them.

And I think it's control and being able to trust somebody else with giving up that control, so you can accelerate what you're doing and looking to build in terms of business value, is more of a cloud-native mindset than anything else.
Episode #73: Optimizing for Maintainability with Joe Emison
In general, I do choose third-party services for everything. My general view is, prove to me that this third-party service won't work. Now, again, I have a very strong difference though between a third-party service that's serverless and one that isn't. You can find third-party services where they want you to go into the AWS marketplace and run it on a VM. That's not serverless and I'm not interested in that. Or like, “Oh, it's an open-source project. Run it yourself.” Again, I'm not interested in that, but when it's serverless … My short definition of serverless is it's not my uptime. I literally can't influence uptime. Beyond, I could put bad configuration or a bad code in, but if some server fails, it's not on me to bring it back up. I think if you can have a serverless third-party API, I think your default should be to use that unless you can prove that you shouldn't use it.
Episode #80: Revolutionary Serverless at re:Invent with Ajay Nair
Flat-out, I think that is the biggest factor to speed that serverless brings to the table. Like the fact that you can cherry-pick components of your customer or product by relying on other people's expertise, right? So going out there and saying hey, I know Jeremy Daly, you have built this great chat service that has … and I trust you to offer me full lines of availability and a certain performance guarantee, and as long as the user API, I’m good, that the incentive for me to go and rebuild that elsewhere is negligible. Like it doesn't help my business to go and rely on anything else.

And I think what that basically does is your now recruiting an entire collection of experts of really deep domain experts to be part of your operational team, to be part of your development team where they’re continuously improving their portion of that tiny little product and making it better to move faster. The scale is getting better. The performance is getting better. The capabilities are getting better, while you innovate on the part of the stack that you want to. And what's fascinating for me is, you know, that is the true vision that we all had when we went on microservices development as well. Like you can do independent development of different pieces. They're all you know, small pieces loosely joined that talk to each other and they can innovate separately. The only difference is this is not just your organization sitting and doing it, your two-person startup. You have now, you know, 22 person startups and AWS innovating on your behalf, just to make your product better. Right?

Like, you're 1 millisecond example is a great one. Like if you were a start-up who was running on us today and you happen to use Lambda for your backend compute, your bill just got 40% cheaper, which you can now pass on as end-user savings with you doing nothing. Like imagine how much work you would have to do to go and get that kind of behavior over there and just one more thing, Jeremy, since you brought that up. I do believe the true power is going to be connecting all these ourselves together and getting them to interconnect a lot more.
You're starting to see this with some of the bigger ones, right? So Twilio, Workday, Atlassian, they’ve all added this programmable size component to them. They’ve got Lambda based extensions that they showing up, like Twilio Functions and Netlify Functions and others that allows them to add just a little bit of logic to them to then talk to other services via API calls, and kind of build forward over there. So I think just the flexibility and power this enables is really, really cool. And the fact that you can swap out one API for another is quite a testament to the whole dance around “am I really locked into a particular provider or not?” because it's quite easy to change the API call more than anything else.
Episode #76: Building Well-Architected Serverless using CDK Patterns with Matt CoulterYeah. So, it helps that Liberty Mutual as a whole is split up into different business segments, so my segment, GRS we call it, Global Risk Solutions, I’m lucky I remembered that, we’re basically large commercial and specialist insurance. But our CIO made a mandate; he put down what our vision is as a company and where we want to go, and he wrote down that we want to be a serverless first company. So whenever you have buy-in at the executive level, it helps a long way. But the second part of it is I haven’t mandated anything to any engineer who works anywhere because I’ve seen an awful lot of times that it doesn’t matter how good your idea is, if you come in and tell people, “I think I know better than you,” they just say no.

So that’s why I started with CDK patterns external, which is, given I haven’t introduced it yet, an open source collection of serverless architecture patterns and the idea was if I could go external and say, “Here is a thing, here is an actual industry thing, here are all the AWS Heroes that talk about the patterns that are in this, here’s the links to all their blogs posts, here are all their articles, here is me talking about it in the world and then go to them and conduct a well-architected review with their team and then instead of mandating it, just ask them, “Okay, I see you’re trying to build this particular solution, have you considered.” And then at that point because the things already exists, it’s already coded and they can pick up on it, I think you’ve reduced the barrier from the direction you want them to go rather than forcing it.

View Details

About Ajay Nair

Ajay Nair is the Director of Product Management at AWS. Ajay is one of the founding members of the AWS Lambda team, in his current role, drives the serverless product strategy and leads a talented team driving the product roadmap, feature delivery, and business results. Throughout his career, Ajay has focused on building and helping developers build large scale distributed systems, with deep expertise in cloud native application platforms, big data systems, and streamlining development experiences. He is also a co-author of Serverless Architectures on AWS, which teaches you how to design, secure, and manage serverless backend APIs for web and mobile applications on the AWS platform.

  • Twitter: @ajaynairthinks
  • LinkedIn: https://www.linkedin.com/in/ajnair/
  • Serverless Land: https://serverlessland.com
  • Serverless Architectures on AWS
  • Building revolutionary serverless applications: https://virtual.awsevents.com/media/1_wrjleiff

Watch this episode on YouTube: https://youtu.be/QMOLE2-SUjU

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Ajay Nair. Hey, Ajay. Thanks for joining me.

Ajay: Hey, Jeremy! Finally on the show, yay!

Jeremy: Well, I am glad you're here. So, you are the Director of Product Management for AWS Lambda at Amazon Web Services. So I'd love it if you could tell the listeners a little bit about your background, kind of how you ended up at AWS and then what does the Director of Product for AWS Lambda do?

Ajay: Okay, first, thanks for having me on the show. I've been a great follower of both your talks and blogs for a long time, so I'm excited to kind of finish what's been an interesting year by spending time with you. I’ve been at AWS now coming up n about seven years; been with the Lambda theme for pretty much the whole time. Tim Wagner and I were the founding folks for AWS Lambda way back when. I spent a whole bunch of time at Microsoft and some other software companies before that in a combination of development and program/product management roles. I ended up at AWS only just looking for an opportunity to go and build a new product or a new service in the cloud space. I’ve done a whole bunch of things with developers in big data platforms so far and they signed me on this top secret effort which they said was going to be a new way of doing compute. Here we are seven years later with me as the Director of Product Management.

So my role as Director of Product really is to help figure out the why and the what of what we should be building and evolving Lambda for, so everything that's happened to Lambda over the last seven years is in some way my fault. So, yeah, in all seriousness I get to spend time with customers to figure out what the right thing to go and build for them is and help the team figure out, build it, and then help the marketing and sales team sell it. That's kind of what my day job is and it's been a great ride for the last seven years and here I am.

Jeremy: That's awesome. Well, I am super excited to have you here. You know, you said it was an interesting year, that's probably an understatement, but not only an interesting year in terms of everything that's been happening, but also an interesting year for serverless as well. And we just finished, I think it was ... what? ... like week 27 of re:Invent. Oh, no, it's just week 3! But it felt like everything this year's just felt like it dragged on incredibly long. But so there were a lot of really cool things that happened with serverless this year and in your purview is more around Lambda, obviously, you're the Director of Product there, but there's so many services and things that happen at AWS that interact with it.

And I think what would be really great to do, and I want to be respectful of your time and of our listeners’ time, because I'm sure you and I could talk for the next ten hours about this stuff and then have to take a break and talk for another ten hours. But so we'll timebox this a little bit. But I do want to start with just kind of a year-in-review of the things that have happened to, you know ... with serverless with Lambda. What are some of the new capabilities, what use cases do those open up? And so let's start with re:Invent. Let's start with the big ones that happened at re:Invent. We can work, sort of work our way backwards and then hopefully you can kind of put all this stuff together. But so let's start there. Let's start with the big one. At least I think this is a huge one because it opens up a lot of, I think, capabilities for other people to get involved and that has to do with container packaging support. So what's the deal with that?

Ajay: Yes, the idea behind this, as you said, is allowing you to bring Lambda functions packages, container images, and run them on Lambda. You know, this is an evolution of the team we have seen for a while where there's a set of people who say I like to build my code a certain way, but I want to run it the Lambda away and zip just isn't my style. And actually more specifically, I think the interesting aspect that is Lambda is enforced this sort of dynamic packaging structure, right, like where the runtime and layers are bound and execution time was doing something statically and I think something has happened since the beginning of Lambda is this evolution of more consistency across local and online development and trying to push that forward.

And we just saw a great opportunity of saying, you know, the container ecosystem’s done a really nice job on the tooling and developer for front of this, driving consistency across the two brings the best of both worlds over there. Then we try to do some interesting bits over there too, like with the runtime interface client allows you to kind of work with Lambda’s event over execution model while using the container development model. The runtime interface emulator lets you get much more consistency on your sort of local testing than we have had in the past.

I mean, you have great Community Heroes like Michael Hart essentially powering large pieces of that, you know, we have taken some of the burden off his back too by standardizing some of those components and taking it over there. But it just to your point opens up a whole set of new use cases, right? Like if you've previously committed to the container ecosystem as a tooling in a block for you, you now have access to all the goodness that Lambda was bringing for you as well.

Jeremy: All right, and that's one of the things that I thought was sort of Interesting when I first heard about containers on Lambda, I was like, oh no, what's happening here? And I love containers. I know I sort of joke about it. Not a fan of Kubernetes, but that's for different reasons. But the idea of containers running on Lambda, I was thinking that seems like … you know, now we're really confusing things, but it's not really like a container running on it, it’s really the packaging format, right?

Ajay: Yeah. Yeah it is. It's just a packaging format. You know the team I joke that if people thought serverless was awake then wait till they hear about containers! You know you start realizing the one containers used as a packaging format, as an execution model, as a slang for Kubernetes, as a sub for an architectural pattern like microservies, and when you say go and tell people, “Hey, Lambda now supports containers,” they're like, wait all of the about our work on Lambda and so you kind of have to do a little bit of separation, say no it's the packaging format, the execution model stays the same. It's still that, you know, if I will model event open behavior that you get. You get still all the security and isolation that you’re used to, you can just package code in a much more familiar way and get access to a broader ecosystem of tools.

Jeremy: Right, and especially with the idea of the images being able to be 10 gigabytes, I mean, now you have this ability to put all of those libraries and their packages all together and not necessarily have to worry about the Lambda layers and connecting all those things.

Ajay: Yeah. Exactly. The size is one … that's one we actually debated quite a bit. Like I'm a big fan of less is more. I'm not a big fan of writing fat Lambdas and all the other variants that are out there, but I think the 10 gig one is really interesting especially for some of the emerging use cases like, you know, machine learning, and larger images, and dependencies coming in. So yeah, I'm just excited to see what people do with this larger limit and kind of play around with the two.

Jeremy: Now, I know Andy Jassy had announced EKS Anywhere and ECS Anywhere, but really what you kind of get with Lambda now, too, with this packaging format is you kind of get Lambda anywhere, right. You can sort of run this Lambda execution model in different places. Not that you would want to, but if for some reason you did that's certainly an opportunity.

Ajay: Yeah, I got it. I would say you can run Lambda functions anywhere, that's a key distinction. I think, you know, one thing I just want to make sure it is, it's ... this is not about recreating Lambda the entire service also, which is what ECS and EKS really let you do. They let you replicate the service in your own environment, but from a perspective of taking the same code and running it in multiple places, kind of what that container ethos really embraces, very much possible with the runtime available to you. There’s still a lot of work for you to do, you know Lambda does a bunch of work for you underneath the covers, but at least it's possible right now, much, much more than it was before.

Jeremy: Awesome. All right. So then the other big one for me had to do with reducing it down to one millisecond billing. And I’ve talked to a lot of people about this. They said, “Well, you know, my Lambda bill’s only $100 a month anyways, right? So it's really not that big of a deal.” We'll get into a couple of reasons why it is a big deal connecting some other services, but one millisecond billing. What's the thought behind that?

Ajay: You know Lambda has to get faster, better, cheaper. That's been what my driving philosophy was always from the beginning. It's funny, 1 millisecond billing was one of the things that Tim and I discussed immediately after Lambda GA and we were like, well, we're not quite there yet. We just got the service started. Let's figure out what the response is to the product before we go and push there. But realistically what we saw was, you know, you're kind of seeing this breadth of use cases on Lambda where a whole bunch of people are like, look it's a couple of seconds larger workloads, big data kinesis, etc. for these data intensive processes but a lot more interactive workloads, especially when there's a bigger push for performance, you know lightweight one times go and others are running and that sub hundred millisecond bucket.

And we’re, like, look, there's a way for us to save them, you know, 40, 50, 70 percent of their bill by just changing the billing granularity at no effort for themselves. Like, that's a huge rally point and if you can go and tell a customer saying that, hey, you actually making a performance better is going to make that difference right? You're actually literally saving money by making the experience better for your customers. That bag was a really compelling value proposition for us. Like we just felt like there was an entire class of use cases we could make 70 percent cheaper. We found a way to make the money work and we’re, like, let's shoot it.

Jeremy: Right. No, that's awesome. And I actually saw a tweet, somebody showed a graph of what their Lambda bill was before and after 1 millisecond and it was, like, a 40 or 50 percent reduction. I mean, it was huge. And I think there are some of those use cases where if you run enough of them and enough frequency, you know, that that really matters. And optimizing to get under a 100 milliseconds, you know, in not meeting it and running 401 milliseconds and getting built for 200 milliseconds. I mean, there's a huge cost savings there.

So another big thing just in terms of optimizing speed, you know, and this is something great that Alex Casalboni has done with the power tuner, right, now is 10 gigabytes of memory with up to 6 virtual cores. So now you have an opportunity to really tune that even more and get those things just cranking and of course the use cases that opens up.

Ajay: Yeah, exactly. I think one of one of the internal demos we had done was running a 30 millisecond ML inference which used about 4 and a half cores which spun up to about 15,000 concurrent at peak and the bill was, you know, sub 50 dollars at that particular point because it ran for such a short duration of time and I think for us these two features are sort of enabling different points in the spectrum, right. Like 10 gigs and 6 cores enables you to far more data intensive use cases, compute intensive use cases, for running sort of these bigger, beefier workloads not feeling capped or restricted by the amount of compute available to you.

And 1 millisecond was saying, well, you can still do that in this sort of probably microcosm of use case of usage to get that cost efficiency even when you're running these really, really large-scale workloads. And to your point, Jeremy, I think that the combination is really powerful because you can now performance-tune to a point where you get to sub hundred milliseconds and you are incentivized to do so, right. You're incentivised to keep pushing the number even lower with the new 1 millisecond billing.

Jeremy: Yeah, that's amazing. All right. So then another couple of other things that launched at re:Invent had to do with sort of just event sources and controlling event sources and that … this is a really long conversation, so we're going to have to try to keep it somewhat short, but the the big news I think was, you know, again Apache, Kafka, and Amazon MQ opening up as additional, you know, event sources for Lambda, which is important, I think, because you do have a lot of enterprise workloads or existing workloads that run on those types of services. So it's a nice little gateway into serverless to start introducing some more of these serverless things. You can comment on that if you want to; less interesting to me because I moved away from those because I use all serverless things now.

But what I thought was really interesting were two things that launched. One was custom checkpoints for Kinesis batches and for DynamoDB streams, right? So you have ... if people are unfamiliar with this, and you should probably be explaining this, but you have the ability to bisect batches in the past where you could say if the first half of the batch, you know, you keep splitting the batch so you can get rid of that poison pill, but you would reprocess the same event or the same message over and over and over and over again until it was finally in a batch that fully succeeded. So that's changed now with custom checkpoints. Can you explain that a little bit better than I just did?

Ajay: I'm going to try; that was a pretty good one. But that's the philosophy behind it is that this is a more advanced failure management scenario when you're processing complex patches, right. So, Kinesis Dynamo is one of the ... is the oldest event source for Lambda at this point, right? So you're going to see more of the enhancements coming out on that front and what this now enables you to do is get far more granular control when you're passing these larger batches, right?

So one core use case we’ve seen for Kinesis and Dynamo is these sort of analytical and aggregation patterns, API signals or you're bringing in, like, machine operational data in there and doing so, and we kept seeing this pattern of people saying well, it's not just a collection of records that's bad that showed up over the window of time. There’s this one record that has malformed patterns or records. And this just enables them to say that up to now I failed; let's stop right here checkpoint, move on. That was one thing that I really liked with the ... what they do with, sort of, KCL and if they’re self-hosting, but that sounded like an artificial choice. You know, like, I have to either use KCL or self-hosting. We'll go over there.

Jeremy: Yeah, no. And again, that just ... it just opens up a huge level of efficiency and the other thing is maybe people don't think about it this way. And I don't know why I'm always thinking about the billing aspect of it, but that reduces billing because that is less invocations of your Lambda functions in order to process those events, right? Because if you're processing the same batch over and over and over again, you’re reprocessing the same thing multiple times, we can just … another way to kind of bring that down.

The other one, though, is the idea of tumbling windows in Lambda, and this is super exciting for me because I actually … one of the PM's reached out to me many, many, many months ago and I gave some early feedback on the design on this, and one, I think that's a huge testament to, and I'm sure I was just one of hundreds of people who probably commented on this, but it's a huge testament to what AWS does in really getting these things in front of their customers early and trying to figure out what those use cases might be, you know, before they end up so you're not just, you know, sort of building things in a void. But tell us a little bit about tumbling windows.

Ajay: Yeah. No, it's funny that this is but this is actually a really good example of that because when we originally started out this feature, we just thought of it as stateful processing for Kinesis, and over time we were like, look it doesn't make sense for us to just go and say, “Here’s state; good luck.” You have to kind of make it work for a specific aspect that works in this particular scenario. And, again, tying back to that analytical use case, right? We kept seeing this repeated pattern of people running code on their Lambda functions, writing some little bit into DynamoDB, reading one record from DynamoDB, and then processing and moving forward over and over again. And we said, look we can just simplify this entire stack if you just enable Lambda to have a little bit of smarts and how it passes data back into the reprocessing of the Kinesis or DynamoDB stream and that's kind of where we took this tumbling window operation. That's a pretty standard one in most stream analytical behaviors out there and said, we now support that operator.

The actual primitive enabler is what’s more fascinating for me, tumbling windows is just one specific pattern that this enables, right? Like you could go to look forward and say hey, can you do other kinds of aggregations on this? Like, can you now have a formal … you know, this is a primitive reduce, can you do something even more smarter over there? Can you now combine it with something like step functions and start building even more smarts on how these all orchestrations come together. And I think that's what is really cool about this in terms of how It's actually built out. All right. It is one of the top ... it's only been … it's been less than a week since it's come out, but the internal data shows, it's quite popular already.

Jeremy: I can imagine. Yeah. No, it is. It's a huge solution because every other solution you'd build around that with something janky, right? It was like loops and step functions or, like you said, writing data into DynamoDB and then reading it back just to do some simple ... just to pass in the aggregation, or whatever it is, and the aggregation, the state itself, I think can hold up to a megabyte of data, so, I mean it can hold a pretty good chunk of data that gets passed from invocation to invocation. And yeah, so just super interesting there.

You mentioned step functions. And again, I know step functions are a little bit outside of your purview, but super important as an interface into AWS Lambda because they enable you to do orchestration and this is going to go back to why I think the 1 millisecond billing is so important because what we used to see with Lambda functions is ... I'm sorry, with step functions ... is you would use that oftentimes to do function composition, right? You might have a function that does some sort of conversion of your data, another function that maybe writes that data somewhere, another function that then maybe, you know, processes it some other way or generates an event, and then maybe something that returns it, you know, and back to another system.

And that was always one of those things where it's like you're paying for every transition. You're paying a minimum of a 100 seconds for every single step function that runs and, by the way, it's asynchronous, so really it's got to be a background job that runs anyways. Synchronous express workflows, I think, are one of the coolest things I've seen come out of AWS in a very, very long time. Maybe even cooler than Lambda because what this allows you to do is this is the answer to function composition, at least in my mind.

So no more fat Lambdas, no more, you know, Lambda lifts, no more, you know, having to push a bunch of things behind the scenes. This is now a way that you can say, I have a Lambda function that does this very specific thing, generically, by the way, right? It doesn't have to have a whole bunch of specific things that it does, or it does have to be tied to resources. Then I have another Lambda function that does this, another Lambda function that does this, I want to put those all together and I want that to happen in a synchronous loop. That's possible now.

Ajay: Yeah. No, you know, it has been six years since Lambda came out. So it's about time something more cool than Lambda launched for sure. This was actually one of the really exciting launches for me too, so as you can imagine in my role I end up working quite closely with a lot of the broader serverless portfolio at AWSt anyways, and you know, the step functions of Lambda are sort of PB&J for us at this particular point. So this particular pattern … it's funny you bring this up, one of the things we were really excited about was this exact granularity of resourcing and duration that will enable because of the synchronous patterns, right? So we had customers who were doing metadata retrieval, analytical and processing using an ML model, and then a long I/O wait time to write the output into something all-in-one Lambda function, but then they had to run it as you know, it was like I think a 1.8 or 2 gig function because they were like, hey, we have to run this major processing on it.

Now that whole thing splits up where you have like the cheapest Lambda with like actually now you can do some 128 mills. You can have a 64 meg function just doing simple I/O finishing in a couple hundred milliseconds. You're ML model beefing up to 10 gigs and doing the whole thing, and then a simple I/O right there at the end, you know, doing super cheap. Your overall costs were reduced by at least probably 20 to 30 percent in that but your performance behavior looks kind of consistent, right? I mean, this is one of the things we are really excited, even when we put Express out there is it enables you to run a whole bunch of things at scale asynch of the first pattern that it went out with and now with sync use cases you have, like you said, all these new things that are opening up so that that was one of my favorite non Lambda launches inside re:Invent as well

Jeremy: But it ties so closely to Lambda and the reason … and then you mentioned this, this is the pattern, it's that 64 megabyte or the 128 megabyte of RAM in order to do something simple. And I wrote a blog post a while back that was basically, you know, about if you're paying to wait for Lambda, especially if you're paying to wait for something like an API call, then don't, you know, don't run it at a gig of memory. I mean, that doesn't make any sense, right. Run it at 128 because it's not going to be any faster or slower and if it's any slower we're talking milliseconds.

But this is what I love about single-purpose Lambda functions in the first place is the isolation model not just from the idea of, you know, just that code is very simple and it's running it there, but you have the security isolation. You have the concurrency isolation. You have the memory isolation, have all that stuff that's there. And if you think about the ability to say I can run this particular Lambda function to maybe, you know, hit an API and all it’s going to do is bring that data back and I can run that like you said 60, or get a 64 megabytes then pass that response into something that can do the actual processing on it or whatever. So that's super huge. I love that pattern.

And what's another thing that was launched at re:Invent that was announced, was now that you can invoke these synchronous Express work flows directly from HTTP APIs, so now you don't need a Lambda function sitting there saying, okay, I'm going to invoke this, wait for the response, and then spit it back. Now you just take that step out. You go directly into the step function and, again, any things you want to add, of course security and authentication, that's all added at the API level. But anything you want to do within that, you get all that data, you get everything you need to do these really complex synchronous workflows, by the way, in parallel, too, if you want to run parallel executions. There's all kinds of crazy things you can do all within this, you know, single round trip to the server. I'm gushing about this, but to me, this is just fascinating.

Ajay: Yeah, Jeremy. I mean, the stacking was conscious. Right? So HTTP API is on synchronous API as a Lambda function was exactly the pattern you're going for. I think this is one I'm really excited to see how the CDK folks will respond to this because it lends itself really nice to, sort of, these programmatic creation of more complex things using these service primitives in the end, like the API in the workflow and the functions expressed this code and I'm bringing it together. And once even about the isolation model is one pattern we started seeing already is customers splitting up their execution roles for all these individual functions. So that the first function that’s retrieving metadata only talks to the DynamoDB table. The second function only talks to S3. And the third one only talks to the HTTP API, like, that granularity of isolation and you know, even interface isolation, so to speak, and what they speak to is extremely powerful, not to mention the fact that the teams can work on those things separately if they so choose.

Yeah. So yeah, this is one that hopefully next year you and I are having an entire dedicated session just about this. With my Step Functions friends, of course.

Jeremy: And the reuse. I mean, that's the other thing to me that I think is really interesting is the reuse aspect of it. So, you're right. You can have one function that can only talk to DynamoDB while another function only has the secrets available to communicate with the Twilio API. I mean, there's just so much isolation and then the, you know, principle of least privilege there, but then the ability to reuse it. I mean build a generic function that queries Twitter and all you do is pass in what the, you know, what the hashtag is that you want or something like that. I see there's just a lot of really cool stuff you can do with that.

Okay, I want to move on because we've got more to talk about. So, before re:Invent, there were a number of really cool things that came out as well. One of the big ones was EFS integration. I know this happened quite a while back, but this was something, I think, big not for a lot of my workloads, not a lot of use cases that I have, but certainly the naysayers on serverless ML, you know, you can't do machine learning in serverless. This was a big one, I think, to kind of quiet them down a little bit.

Ajay: Yeah, I mean … I will say one thing that is always fascinating to me is how much social media chatter happens about Lambda for using web and API use cases whereas how much of our internal use shows up for these really brutal data intensive use cases, right? Like one of them, you know, one of our big probably customers who talks about it all this time is, like, Fannie Mae running Monte Carlo simulations on Lambda, right, and, like, this whole ML inference is a huge segment that has kind of grown for us even more and you saw this in the NGB launch as well. I think for me EFS is a combination of things: one is, like you said, just knocking off the I can only use 512 megs limit inside Lambda.

You're actually getting a solid persistent store that comes with it that is as performant if not, you know, matching Lambda’s behavior in terms of familiarity and billing and others that go with it, but sort of just enabling sort of these early fast access patterns on Lambda that meets the performance needs the customers have. Like, I think Azurian has a story out there about how they're using Lambda for these ML model storages at this point using EFS for exactly that reason. Like, they're able to serve customer-facing requests on that particular stock running really, really fast and that combination of 10 gigs EFS etc. is kind of pushing you towards this new use case direction of ML inference that I suspect we'll be hearing a lot more of in the future.

Jeremy: Yeah, no, that's ... and again, it's, to me, it's a lot about use cases. I mean even one of the other ones that launched sort of pre:Invent I think it was maybe in October or maybe November was SQS batch Windows. Simple, simple thing. I mean, but again, when you added SQS as a trigger to Lambda functions, I think was in 2018, that opened up a whole bunch of really cool things where now you didn't have to have cron tabs running to poll it, whatever, but then I think what people quickly realized was now, you know, if you've got small batches or things coming in too quickly or not quickly enough, I should say, you've got this issue where you're executing a lot of Lambda functions over and over and over again potentially needlessly. Whereas now you can put them in batches of 10,000 and process some big batch. Now again; no bisecting there, no iterators and things like that on SQS yet. Sure but yeah, but that, I mean, that's one of those things, that's one of those things. I thought that was a really interesting one.

And then Lambda extension. So that was another one that got … that was fairly recent. So this is something where, and maybe this is a pattern that we're starting to see, but the extensibility of Lambda, right, like, making it and, again it's called extensions, but no, this idea of extending Lambda so that you have more control over the runtime. You have more control into the execution model into the lifecycle hooks and things like that. So, what's the thinking? Well, maybe explain what Lambda extensions are and then what the thinking behind that is?

Ajay: Yeah. Now that's a good one for us to get into. So, the idea behind … so Lambda extensions is built of something we launched called the extensions API. So it's a pure to the runtimes API that we launched in late 2018 that allows you to access essentially life cycle events from the Lambda execution context. So, you know when the execution context is spinning up and it's been shut down when something is running inside it so when your function is actually being invoked. We also launched something called the logging API that gets you programmatic access with the logging streams that are being generated from the limited execution environment and the code that's running into it.

Now what this conceptually enabled ... this was very much one for our partner ecosystem. Right? So one thing we realized very quickly is customers like using their own tools, right, like it's good if there are defaults, but they like their own tools and we've had this whole great ecosystem even around Lambda with, you know, people like Epsagon and Thundra and, you know, IOpipe, and Datadog, and others who were trying to make sort of the Lambda debugging and diagnostic experience really powerful, but the fact that they didn't have access to this additional metadata was kind of kneecapping them. So we said okay, let's open that up.

You will notice one of the things we really tried hard to do with extensions is it doesn't change the experience of the person writing the function, right. The person writing the function still just either includes a layer or does something different. It's what … it's these partner ecosystem people who are building the extension who get these additional capabilities like, oh, now I can know when a function started so I can start tracking a specific choice at that point. I know if the execution environment is spinning down so I can flush my log buffer and send it out there, or I can expire my credential because the function is gotten ... is finished.

And so I need to write something back to Vault or CheckPoint. Like it just enables a whole bunch of these patterns around how the function itself is operated that I think is really, really powerful and it's another one of those pots that you're clearing out for Lambda, right, where you’re like, I would love to use Lambda, but I can't use my own operational tools. Well, now you can with identical capabilities really to any other computes to that you have out there.

Jeremy: And I want to get into the partner aspect of stuff there. I want to finish up with, sort of, these launches and then we can jump, we can tie that … sort of, jump back to that. So, a couple of other things, and I'll just mention these quickly, EventBridge archive and replay events again, not necessarily something that you're working on directly, but I feel like most of the tail-end of events end up hitting Lambda function in some way and then you know x-ray integration with S3. Just this ability now if you're using S3 with your Lambda functions, being able to trace that all the way through is super important. There are a number of really cool launches with Amplify and some AppSync things, just giving you different ways to do stuff.

And these patterns and we talked about this a little bit earlier, you know, whether they're DLQs for Lambda functions or, you know, event Lambda destinations, or it's tumbling windows or it's iterator control, or it's more control over how Lambda consumes these events from other things. You know, what is it? They seem, I don't want to say they seem inconsistent, that's not the right way to say it, but some services offered X and some services don't. Is that a general goal? Was that something that we might see where we're seeing some more consistency across the consumption?

Ajay: No. So, I will say inconsistency is not the goal, but neither is consistency, right? So I think for me, event of an architectures is something serverless has brought to the forefront: the idea of composing services together with strong contacts in APS events is the lifeblood of what, you know, people like you and me spend our days obsessing about. So making that pattern more resilient and performant is something you’re going to keep seeing. What you're seeing with these controls is having them show up where they are most definitized, right? So with event and replay what we kept seeing was, sort of, this idea of backlogging and replaying state events, especially the services that our current coming to EventBridge made the most sense that ... so that's, kind of, where it showed up first.

With tumbling windows analytics was the big use case that we saw the control work the most for and that's why you're kind of seeing it show up with Kinesis first. My prediction, and please don't hold me as a roadmap goal on, this is what I would say is you will eventually see that sparse metrics getting more filled up right. Sometimes more is a bet that says hey, this is something we think will be useful for this customer base because we're seeing sort of this cross pattern, but in other cases just because, like you said, the demand will start showing up and I think DLQ is a great example of this. Like SQS started with it way back when, Lambda launched it, EventBridge now has it. And I would predict each of these integrations is going to see that pattern get more and more.

So, so yes, you will see this get consistent over time, driven both by customer demand and where we see opportunities to make life simpler for these event driven patterns built using AWS services.

Jeremy: Yeah. Well, I did a talk that I gave a number of times called “How to Fail with Serverless” that basically was analyzing all the different ways, all the different failure scenarios, and how AWS is built to handle them. A lot of different things, you know, synchronous versus asynchronous, for stream-based versus push, a lot of different ways that these things get handled. So seeing these patterns, you know, the broader so that you can use them on different services is, I think, is going to be a great thing. So awesome stuff there.

I want to mention two other things and then ask you a question about something that wasn't launched. So Aurora Serverless is v2, again outside of the Lambda team, but I just think generally a really good ... a much better way to do MySQL or postgres, you know, at a serverless scale. So just a really good on-ramp for serverless. I mean it just kind of, it lessens the pain of somebody moving into Lambda functions and realizing, “Oh, I need to set up RDS proxy and I'm going to have all these issues with things.” Just the scale, you know ... just the reliability of the databases when you overwhelm them. So when thoughts on Serverless v2. I mean obviously it's a good thing but, you know, just overall impressions on those types of services being built that are really, I guess, complementing the scalability of Lambda.

Ajay: Yeah. No, I think you're going to see this pattern of serverless-driven primitives, right, so databases as a service in the true sense of saying pure usage-based, highly scalable, and burstable like Lambda is really cost efficient on a program the basis of all more and more. You know, the Aurora team’s done a really nice job with serverless. We do ... I've had a chance to play with the early versions of it as well and kind of how it plays at Lambda. I do anticipate that will drive some consistency between that and IDS proxy over time so that you kind of have this continuum of saying your own database, connectivity serverless, you know, serverless database or RDBMs with Serverless v2 and then who knows, right?

Move on to DynamoDB if that's kind of where your flavor stands, but it's more about enabling that continuum for me and kind of making sure that you have good checkpoints along the way of going for it. So, if you are a customer who cares about the IDMS pattern, but is willing to kind of go AWS native on using some of these core capabilities with the cost efficiencies and performance you can get, it's a great choice that works really well with the way most people are running applications right now in this behavior. It's not just Lambda, right. Like even if you’re using containers are easy, too. It's the same behavior that you will see.

Jeremy: And I’d love to see v2 sort of handle the RDS proxy thing for you so that you didn't have to do that yourself, you know, and again data API was a step in that direction. Anyways, very cool. Right. I want to shift to my favorite sort of runtime environment, which is Node, right? So they announced the other day AWS SDK for JavaScript version 3—very, very cool because it allows you to import just individual service packages as opposed to the whole thing. Probably not overly exciting for you Python and Go developers and things like that, but exciting for us Node developers, but the question that I’ve got was why no Node 14 runtime for Lambda?

Ajay: Yet!

Jeremy: Yet!

Ajay: Look, I think we've had a pretty good record of keeping up with Node releases. We could not have done as well as in keeping the window close to 90 days as I would have hoped to, but, you know, it's a question of when, it's not if.

Jeremy: All right, I'm sure that that will make people happy and again, you can always run your own custom runtime if you really need it. All right, so let's move on to the talk that you gave at re:Invent this year and it was called, you know, “Revolutionary Serverless” or something. What was the name of your talk?

Ajay: It was “Building Revolutionary Applications the Serverless Way.” So there were a couple of versions of the title then they kind of changed the launches coming in, but that’s the final one that ended up.

Jeremy: So I watched this talk and I was highly anticipating this talk because, again, you being, you know, involved with all the different teams that are building these features for AWS. And then honestly, you know, connected with all the other teams that are building these ancillary services and other things, I was really looking forward to this talk and I wasn't disappointed. I thought you did a good job of all of the things you can't say, because I know that's a tough thing with AWS. It's like I wish you could say, “Oh, we're building this, we're building this, building this,” but you can't say that and that's fine. But I think what the talk did, for me at least, was it reaffirmed the commitment I think that AWS has made to serverless.

And I know I was disappointed last year that there were very few mentions of serverless, even though there were things that were launched and there were a lot of sessions, I felt like serverless was very much so front-and-center this year, and I think your talk and also Dave Richardson's talk and things like that just kind of went over what has been launched and why you're launching it. What the … I guess the tenets, behind those things are, so I'd love to get into this a little bit and I'm going to put the link to your talk in the show notes. I think you probably have to be registered for re:Invent to watch it. It will probably eventually be on YouTube, but it should be on demand at this point. But I do want to start with the idea of these tenets and part of it is ... and I guess there's maybe three or four of them in here, but let's start with this idea of serving builders. What's that about?

Ajay: Yeah, I know. For me the reason I included that in the talk was to reaffirm the idea that the serverless motion is ultimately about delivering value to your customers, right. The ultimate customer for us is someone who's building software to deliver and customer value. The goal is not to make, you know, infrastructure cheaper. The goal is not to just, you know, drive utilization to Amazon servers. The goal is not to offload workloads of data centers. It’s to help builders go faster.

And that's a tenet that's repeated often within the team just to kind of reaffirm saying the ultimate customer is the developer. They have an entire ecosystem of people helping them out: you have operators and others to kind of go and do that. But there's a developer problem we’re solving. The developer's job is to, you know, deliver value. The challenges they'll face at doing that at scale, at cost efficiency, and others. The things they fret about, the things like security performance and scale, and that's kind of what needs to be our world and how we go and build over there.

And hopefully you'll see this reflected in the whole thing. Like you’ll see our services are designed as application primitives. We talk about applications front-and-center all the time. We talk about application patterns that are enabled. Everyone who's out there talks about how quickly they are able to build and deliver value, and I think that's what resonates for us when we say, okay, this thing has actually got legs, you know. This whole motion is about doing less to do more, as I think you've said quite often, too.

Jeremy: Yeah. I know, and I think the idea of, you know, being more productive, building things that are ... that aren't, you know, this term has been used a million times, but undifferentiated heavy lifting, right? This idea of doing the same things and just enabling people to build better things. I remember, I think it was last re:Invent, we were having breakfast together actually, and you asked, “How do we explain serverless to people?” or something like that. And I said serverless is just the way, at least for me, it is. And then it's funny because now, you know, “this is the way,” is from the Mandalorian. I don't know if you watch that show. Anyways, I think he's talking about serverless stuff.

But there's some design philosophies that have to sort of go into, you know, understanding how it is that you provide people with the tools that they need. And so you've got some design philosophies that are this idea of, you know, “ephemeral” which I love, right? Like this idea of, like ...I've had so many companies that I worked for that have had a server up and running for, like, you know, this server has been running for six years. Don't reboot it because if you reboot it we have no idea what will happen and that's the worst thing you can do.

Also this mixing request, right? So you have things that handle multiple concurrency. I know it sounds good that you can handle a lot of requests for the single server, but that introduces a lot of problems, right? And then you also have this issue where, again back to the idea of not being able to shut that server down, where if you have to babysit something because if something changes or the memory gets wiped or something happens, you know. And this goes back to that sort of “cattle not pets” argument. So talk a little bit about that design philosophy.

Ajay: Yeah, so I spoke about this in the context of what I call “compute for all,” right. So one of the big things the Lambda’s enabling was saying it's the ultimate democratization of compute. Like how can I give you access to the entirety of AWS’ compute power without you having to become an expert in distributor computing, right, like dealing with all the scale problems and others that go over there. And the reason we kind of picked these three as, sort of, are driving tenants was exactly what you called out, right. If you look at cold problems our customers deal with around, you know, security and maintenance and otherwise, part of it is driven by the fact that they assume this thing is alive for a long time.

The longer it stays the more craft it enables, the more complicated it becomes to spread the workload around, right? Like you get into, say, things like affinity and state which gets far more complicated. The entity becomes addressable not just for your attachment purposes, but for security purposes, right, and security is like really top 10, it is the top 10, and for us as we're building through and it's one that we want to pass on to customers on that particular front as well. And for me, you know, people often say, oh you're saying ephemeral; ephemeral means not durable. For me it's more the temporal ephemerality of it. Like it exists, but it only exists for a short period of time when you need it to exist. And, you know, I'm not saying that that means, oh there's a finite … as long as there's a finite bound to existence, that's what matters, and that's that sort of human contribution.

So if I start saying, oh, it's finite bound, but it's bound for a month, that kind of breaks the model a little bit over there, right, but like 15 minutes, 30 minutes, an hour? Like, sure, that's within bounds, like, that you're still within bounds of things being cycled and cleared out and going out over there. And then like you said the same thing with tenancy and isolation, right? Like one one request, one execution environment was driven by the same thing. You need consistency of performance, that's a hard problem to do in distributed systems. You need isolated resources to run and execute the code that you're doing. That's a hard problem to go solve. You can enable multi-tenancy.

But again, like you said, it sounds great in practice and there are a whole bunch of patterns you have learned in the past that do it really for efficiency reasons, right? Like the funny thing I keep and when I talk to people about multiple requests for execution, they're like, well, it's because it's cheaper. I'm like great. So what if made it cost you, you know, $0.01 per billion. Is it okay then? And they’re like like, yeah, oh, yeah, then I don't want to write multi-threaded code. I'm like great. So let me do that then I'll make it cheaper for you than forcing you to write a multi-threaded code on that particular file and it's that combination for it.

And for me the last bit that Lambda does really uniquely even now is saying you deal with resources, you get CPU, you get memory, you get configured. and your code is the thing that's important. That's the addressable thing; not the resources, you're not binding your function to a collection of machines or a pod or anything that's addressable. You're just saying I want this much memory every time it runs and whether I spin up 18,000 quarters in the backend or 1.8 cores, that's not your problem. That shouldn’t be something you worry about where those cores exist, where all that happens. That's going to be the driving philosophy of it. But then again the whole idea is the less things you have to think about from a distributed computing perspective more than what you need to build and deliver value to your customers.

Jeremy: Right. Yeah. And so the other thing, the other piece of this, is that you provide all these these primitives and capabilities and I love this idea of democratizing, you know, compute because that's one of those things where I remember long ago, you know, just having to order racks and racks of servers from Dell and paying thousands of dollars a month in order to put them into a colocation facility somewhere. I mean even just EC2 instances and VMs. I mean that was a huge step forward where I remember I was paying, I don’t know, something like five or six thousand dollars a month just to run a colocation facility. I moved it all into EC2 instances and dropped to $700 a month, right? I mean just that huge shift there, but now we're now we're not even talking about $700 a month. We're talking about, you know, maybe a dollar a month, if you're building a start-up, I mean, that you can get so cheap to do this that now the barrier to entry is incredibly low.

But there's a caveat, there's always a “but” to these things, and that has to do with one of the philosophies you mentioned was personal productivity. And I find this to be one of the most frustrating and the consistently or ... the thing I hear from other people I talk to all the time, is the frustration over developing serverless applications. There's SAM. There's serverless framework. There's Claudia.js. There's Begin, the architect framework. There's all these different ways that try to make it easier for you to build serverless applications, but one of the complaints that I've had, and I know other people have had this complaint, is that you know, AWS isn't always the best with tooling, right? I mean the tooling is somewhat disconnected.

It's not always consistent, but that's something where, I mean ... and again, I'm not even sure there's a question in here, but just more along the idea of saying I get that that's a tenet, and what are the plans? What are you planning on doing, I mean, to bring that to make people more productive? I mean, obviously I think the container aspect is one thing, right, you know, meeting people where they are, giving them more capabilities, but what are the other things maybe that you're trying to do or have done that you think are sort of solving that problem?

Ajay: As you can imagine, Jeremy, like, this is one of the top things that keeps me up at night as well. Like, how do you make ... on one hand we're saying that about making builders more productive on the other hand, you know, the feedback we hear is you're not doing enough. Like, you have to kind of keep pushing the bottle with that and there's a lot of things AWS does well. I think we are really good at building services and growing them, but when it comes to aspects like developer experience, and this is where the personal productivity aspect comes to me, I think the nice thing about the philosophy at least my team follows, and I know AWS overall does this, we can't do this alone. Right? Like, it's not easy to go in and just say, this is the way you're going to do it, take it or leave it and then get out of the way, right, and that's fine. That will work … that will always work for a niche audience. But if you want to go broad with your story, it has to be something that you do in combination with others.

So, I think for me a big push … that that's kind of where the big push between the runtime API container support and even the extensions comes in. Like how do you get the rich ecosystem of partners around AWS to help customers solve that problem and kind of do better on that particular function. I think that's one. I think the, you know, I'll go back even to Serverless Framework. Serverless Framework was actually put out by Austen and Go first even before SAM and others came out, right, like, that was one big enabler for the early drive around it. And that was really nice innovation out there. It kind of started tackling the problem of standardizing in deploying serverless applications that inspired a whole bunch of other tooling pieces that came around as well. So that's one.

I do think serverless has a unique challenge where you cannot have … there's a new conceptual learning that you have to go through. Applications are built composed of services talking to each other through APIs and events where there really is no defined pattern. So you're now starting to create tribes, right? So you have a bunch of people are, like, no this should look and behave exactly like Web Frameworks and I'm trying to build end-to-end stuff, and you have kind of the Claudias and others show up over there. Then others were, like, no, I'm just going to treat this as a general-purpose, slightly better infrastructure-as-code story. And that's where you kind of have the SAMs and the CDKs with their own tabs versus spaces debate that is kind of sparked off over there.

And I would say the real big opportunity, and I actually really like what Begin and architect folks are doing over here, is starting to sort of embrace that service for nature of it and go. How do you make it easy to compose ourselves together and build forward and do more over there. And I would say that's what you will see … the place you will see AWS potentially in when, and I do this pretty cross out of services, is when you see commoditization of these particular patterns, right? So that's kind of what happened with SAM. We saw every single tooling provider going and saying, I need a way to express a serverless application as a combination of services. We’re like, okay, all of you don't solve it ten different ways; we’ll do it. We’ll give you a default standard CLI, but by no means are we saying only use the SAM CLI? It's an easy and default way and we're going to keep trying to make it better but there's always a rich ecosystem of tools that you’re going to go and do it with. The same thing with diagnostics, with extensions, you've gone past CloudWatch. You know, how can you DataDog as much as you do over there.

So like, I gave a non-answer to your non-question. But the whole idea is that I think we're not going to be able to do this alone. This is gonna be an ecosystem story over there. AWS has to get better in offering more vertical solutions in these particular things and I think that's kind of where the space you will see as investing more. I love what the Amplified team is doing. That's my favorite example of a good, you know, simple experience, and then able to go with code defaults into an experience focusing on a single use case.

Jeremy: All right. Yeah, and then and that is always where … you know, and again, my criticism is only because I want it to get better, and I think that you know, the constructive criticism is always good. But AWS is always very good at building these stacks of these stacks of services … like, these services that do a very very good or simple thing. And and I guess what that's like you mentioned in your talk, you know, this idea of application primitives as a service which I think is is a really good, you know sort of way to think of it where you've got your computed data, you’ve got your integrations, you've got your tools, security and admin all baked in.

I think those are great things but you mentioned something that probably is the hardest thing to explain to people sometimes and you said, you know about architect sort of doing a good job of connecting services. Cross service connectivity is extremely hard. It is not an easy thing to, one, do, but also to grok.I mean, just to understand how, you know, service X connects the service Y and then, of course, the observability challenge that's in there. But so just a little bit more on that. Like I know there have been things that launched at re:Invent and this past year, but what is AWS doing to make that easier?

Ajay: Oh, man, I think this is actually something AWS just needs to solve on our own. This is not a tooling or ecosystem problem. We own the services beyond the interfaces between them, we have to make it easy. So as with all things, you know, security comes from so you will see sort of this consistent pattern of resource-based policies between each one, IAM-based roles policies between each one, granular tenets being talked about each one, across the spectrum of services that can talk to each other. I think for me the next big one is around sort of these reliability controls.

So, the DLQs, the checkpointing, and others that enable to kind of go and do over there, because a message … you need to know who you're talking to, you need to be able to talk to them securely, you need to be able to get the message to them quickly, and in a reliable fashion and you sure that gets one to the other. And I think then it opens up this pattern of saying now what are the new sort of line specific use cases that you can enable, right? So this is where your batching, your replays, your aggregations time windows, and all of those show up.

But, you know, we by ... because you own both sides of the equation that's kind of where the power of AWS can really come in. Like, we're really good at doing collaborative common security scale. Like we should make that as much as possible easy for you to do. I think my prediction for you is I would say you're going to start seeing a far more consistent API-driven story for enabling all these controls across all these connections. You're going to see less and less of this required to be solved at a tooling level and more of more of this being solved at the API level within the AWS services.

And even in the broad … for the broader set of services to participate in this I think that's where EventBridge comes in. Like, EventBridge has to encapsulate all these connectivity as a service capabilities, and so then if you have your own self service and you're like, well, in order to talk to an AWS service, I need these, you know security availability, reliability controls, you just plug into EventBridge and that gives you all of that in a box, but for all the other connectivity, it should just be part of the API, right? Like my dream is the events all snapping that we have for Lambda just sort of universally works for any pattern that you see out there. X-Ray flows occur within the service, tracing is enabled by default, logging is enabled by default. But you know, you need an odd start to keep going through it and as with all things AWS, it comes together quickly and over time.

Jeremy: Yeah. Well, and I think, you know … you mentioned, too, the API economy in your talk, and I think this is fascinating and I know one of the developer advocates, I think for the Amplify team, had written an article about sort of a big difference between sort of the Haves and the Have Nots in the the API economy a little bit, and that's probably a longer discussion. But, you know, there are a lot of APIs, people are building service full applications now, that's just the way that it's going, right. So you have this idea of saying okay, if I want authentication I can get that from Cognito, but I can also get it from Auth0, and if I want I want email as a service I can use SES, I can use SendGrid, if I want SMS I can use Twilio or I can use SNS.

So, there's all these different services that are there. But what I think is really interesting in terms of what can be done, and you mentioned this, is that interconnectivity between those services where, you know, everything is in silos right now. So you say SNS is a service and I have to understand the nuances of interacting with that particular service like it is with any API and that I think is a huge opportunity for AWS to say if you just need a cue, you know, you just attach Lambda to it and it just handles all those things that you would want it to handle and you're not writing all that code, right, just reducing the amount of code that you're writing. And again, I don't know if there's a question in there. But just I think that's really interesting as ... especially considering the fact that serverless to me is more service full now, right, that's really a good way to think of it. It's just what are the, you know … maybe just for the benefit of the users or people who are not quite convinced yet. Why is this sort of service full movement such an important thing?

Ajay: Flat-out, I think that is the biggest factor to speed that serverless brings to the table. Like the fact that you can cherry-pick components of your customer or product by relying on other people's expertise, right? So going out there and saying hey, I know Jeremy Daly, you have built this great chat service that has … and I trust you to offer me full lines of availability and a certain performance guarantee, and as long as the user API, I’m good, that the incentive for me to go and rebuild that elsewhere is negligible. Like it doesn't help my business to go and rely on anything else.

And I think what that basically does is your now recruiting an entire collection of experts of really deep domain experts to be part of your operational team, to be part of your development team where they’re continuously improving their portion of that tiny little product and making it better to move faster. The scale is getting better. The performance is getting better. The capabilities are getting better, while you innovate on the part of the stack that you want to. And what's fascinating for me is, you know, that is the true vision that we all had when we went on microservices development as well. Like you can do independent development of different pieces. They're all you know, small pieces loosely joined that talk to each other and they can innovate separately. The only difference is this is not just your organization sitting and doing it, your two-person startup. You have now, you know, 22 person startups and AWS innovating on your behalf, just to make your product better. Right?

Like, you're 1 millisecond example is a great one. Like if you were a start-up who was running on us today and you happen to use Lambda for your backend compute, your bill just got 40% cheaper, which you can now pass on as end-user savings with you doing nothing. Like imagine how much work you would have to do to go and get that kind of behavior over there and just one more thing, Jeremy, since you brought that up. I do believe the true power is going to be connecting all these ourselves together and getting them to interconnect a lot more.

You're starting to see this with some of the bigger ones, right? So Twilio, Workday, Atlassian, they’ve all added this programmable size component to them. They’ve got Lambda based extensions that they showing up, like Twilio Functions and Netlify Functions and others that allows them to add just a little bit of logic to them to then talk to other services via API calls, and kind of build forward over there. So I think just the flexibility and power this enables is really, really cool. And the fact that you can swap out one API for another is quite a testament to the whole dance around “am I really locked into a particular provider or not?” because it's quite easy to change the API call more than anything else.

Jeremy: Right, No, absolutely and the speed of that is, I mean, just the speed and the less maintenance and all that kind of stuff ... I actually saw a tweet the other day that said something like 99% … or your library only handles 99% of my use case, so I built my own. Right? Like it's the same thing with API, it could fly there too. So don't … you know, if it does 80%, just use that, you know what I mean, and work around another way. But yeah, building your own service is crazy.

All right, we're running out of time here. But I do want to just go over at least a couple more things quickly. One of the big things is that at the end of your presentation you had this slide that was, you know, why going serverless is revolutionary, and it was because it's 30x faster development, 60% lower TCO, and just you know, the availability of security, scale, all built in, trillions of invocations per month. The biggest thing for me here is the TCO because I think people miss that the total cost of ownership of having to maintain these other things like yes, maybe a particular Lambda function costs you 30, 40, 50 dollars a month, and maybe maybe thousands of dollars a month depending on what your use case is, but you might say well, it's, you know, it's 20% cheaper if I just run a, you know, an EC2 instance or maybe spin it up on containers or something like that. But there's a lot of operational work you're missing there.

Ajay: Yeah. I think this is the hardest construct for people to understand but also the most powerful one for serverless to internalize. To your point, infrastructure costs are different to compare. I would argue what we've actually seen is in most cases unless you're running a really highly utilized EC2 instance or a container instance, Lambda would look cheaper for you as would most of the other managed services that are out there, but there are cases where you can say I can run this cheaper if I really squeeze it out of my own infrastructure that I want to. But at that point you are the one squeezing that money out, you are the one pushing the efficiency out of it, you are the one management infrastructure. And this is just my personal note, builders are really bad at putting a dollar amount to their own time.

Jeremy: Absolutely.

Ajay: They don’t know how valuable their time is. And you know, even if you just value your time at a hundred dollars an hour, but that quickly adds up. This is one of my favorite discussion points, people saying, well, I can run this on a three-person instance, and I'm like wait, that's good. How many people do you spend on this? It's like, oh, it's one on call for a month. I’m, like, great how much are you paying for on-call and then they're like, well, I don't know, what, five grand, ten grand a month. I'm like, okay. So now how did that compare to that Lambda bill that you just changed. And you kind of see the, you know, the GIF with graphs and figures floating around they had to go through. We have to do a lot more to help customers internalize that and you're going to see more, I think, material and content come out from the AWS team on helping customers understand both their individual costs as well as sort of how you think about the overall TCO.

There's a great paper by a VC out there, you know, it would be a good link to include in your in show notes as well, that people can read to get a really good model on how to think about the overall TCO, too. But, yeah, that that's the big one. Yeah 60% cheaper over a five-year window is big savings whether you are a small company or a big one.

Jeremy: Absolutely. All right. So again, we've been talking for a while. I do want to move on to a couple of other things. One thing that was really exciting to me was during Andy Jassy's keynote, he mentioned that 50 percent of all new services being built on AWS use Lambda, which is just an insane number if you think about that. Which is great from a serverless adoption standpoint. So that's great. And I wonder why this is … I wonder, you know, again, is it just because it's becoming more popular? Is Lambda just now one of those things that is becoming a little bit more mainstream and I'd love to think that, yes, it's just gaining in popularity. But I think part of that has to do with this, you know, we can't use Lambda because X, right, and all those objections that we've seen just getting crossed off the list and the talk that I want to bring up is Adrian Cockcroft had an architectures trend and topics for 2021, and he spent the first part of a talking about serverless.

And in the talk, he basically said serverless is the fastest way to build a modern application. I agree with this, you agree with this, we know this, right. Not everybody agrees with this, and most of those objections have been around things like portability, you know, scalability, you know, cold starts has always been a good one, you know, state handling, run duration, complex configurations, and there's been all these objections. And back over the summer, there was a conference that he gave a talk at where he basically just picked these things off one by one, and he updated that through this talk and he mentioned a few things like portability, new container support, scalability, now 10 gigs of space, complex configurations, AWS proton, which we don't have time to talk about that. But, you know, maybe some other time.

Ajay: Part two, right?

Jeremy: But just your thoughts on this growth of serverless, like, I mean, if you go all the way back to the beginning, I mean, it was a new thing. So, like, what's happened over the course of the last six or seven years that has, just, you know, that has made this thing such a juggernaut?

Ajay: Wow. I will say when we originally wrote the part for Lambda and we launched it, I remember Tim and I sitting up like the hour after it was announced watching the previous sign-up counter go up, right, like will people grok this? Like, will the idea that you can have a managed compute service that does things for you, this whole concept of events and others, really grok or they just work, and it did. Luckily here we are, hundreds of thousands of trillions of invocations, so to speak. I think for me, Jeremy, that the big thing what we’ve seen is, we found a new way of helping people move faster.

The core problem we are solving of saying we have democratized access to big distributed compute in order to build these complex applications at the way that you can't do before. Like that's always been the underlying philosophy behind it, really. You don't have to know how to build a service, just give us code and you're basically getting code as a service that goes over there. I think the journey of the last six, seven years has been enabling new patterns as you called out, by ... I would actually die back to the same talk things, right.

One has been expanding the capabilities the compute can do for you, things like EFS, things like firecracker, we’ve enabled a better isolation model, expand the amount of compute available to you for 10 gigs and otherwise, the second dimension is being sort of expanding your patterns, a big push being, again, event-driven computing connecting services to each other, service full architectures. I think we are at like hundred forty blocks event sources at this particular point in terms of the capabilities you can use with Lambda where we started out with three, across Kinesis, S3, and DynamoDB. That's been a huge growth factor for us. And now bringing that payment services that are non-AWS, so, you know, you call them in cue and self-managed Kafka now that we just announced and including AWS-managed Kafka, that pattern continues to be evolved.

And then sort of this third one around enabling more developer tooling and productivity. And I think that last one honestly internally the big flip for us was when we started becoming standardized with the internal AWS tooling, right, that was kind of when we saw the big inflection point and that was one of our earliest signals where he said, you know, you can have the compute ... be really powerful and enable whole bunch of new application classes and motivated people will jump the graph to do it, but you have to keep smoothing the plank, so to speak, to help customers come on board, which is where sort of this push towards opening up the ecosystem enabling the tools the customers care about is something that you will keep seeing us do a lot more.

I think for me, the biggest fascinating thing about the serverless ecosystem is what Lambda has sparked. Like, we never went in defining Lambda is serverless and serverless is the new way of doing things. So, like, hey, here's an easier way for you to run code, how go that sparked the serverless ecosystem. I think kind of services that we have seen inspired by the idea that you can keep having, you know, spin down to zero, highly scalable, completely apps built by the millisecond, conversations with services out there, which wouldn’t have been the case, you know, six years ago.

Like, we were still talking about billing instances by the hour, and discussing the next ..XXL.5p king thing that came out at that particular time and that conversation’s changed. Like, I was really … the most ironic moment for me was getting into a debate with a customer in June who was basically really upset at us by the fact that we were doing 100 millisecond billing. He's like, that is not acceptable. Like, that's a really low standard for AWS. I’m, like, dude, I'm so happy that you're complaining that I'm billing you for too many milliseconds. Like, that is the dream that we're getting into that argument versus anything else.

So, and, you know, when you look at companies like Steadi or stuff like Joe Emison and Steve are doing over at Branch. Like you're now starting to see this new flavor of single-digit people startups that are just going to become, you know, a hundred, two hundred billion dollar companies and this is very similar to the way of that you saw in the early days of AWS as well. I think there's a new microsized ecosystem that serverless has spun up that's really, really fascinating to me which then also feeds up into how other people are building applications, right? Like, the services themselves are going to be enabling new application patterns.

So, I do feel the big thing that's happened is the vision of saying, like, “builders build, let them do more with less” has been realized, not just because of Lambda. I think the entire ecosystem has evolved around it as people have realized and kind of putting things together. And I think that's what's the biggest excitement about this for me is that now people are building things that they never thought they would, they're launching companies they never thought they would, you've seen this whole wave during Covid when people are building, you know, response sites in distribution systems and others in like, weeks that they couldn't have thought of handling and it's handling like millions of requests as it's coming in. And for me that's really humbling and powerful, right. Like, something that you built is enabling other people to build things really faster and deliver value and I just hope to see more of that coming out.

Jeremy: Yeah, and I love that you use the term “revolution” or “revolutionary” in your talk, because I've heard a lot of people be like, oh, it's an increment. It's an evolution of whatever and I just don't think it is. I think it is a revolution. I think it is a completely different way to think about building applications and it's a revolution in that it's the people that can rise up to start building these things where there's just not that walled garden like you should have talked about in the past.

All right, I want to ask you one more thing and I think this would be a good way to sort of end this conversation, and that's to go back to the idea of the partners. Because I think AWS has been very, very good about, you know, it creating partnership opportunities for people. But you also run this sort of interesting, I don't know, this sort of interesting dichotomy of building the tools to allow people to build things and then trying to build the tooling to solve the other side of the problem, right?

So you think about observability. I know observability is a frustrating thing for a lot of people, you know. CloudWatch had some ... you know, CloudWatch is CloudWatch, you know CloudWatch laws. I mean you add in metrics, there was insights, there was other things that have certainly gotten better over time. But really they don't compare it to the Epsigons and the Lumigos and some of those, like, they just do a better job, you know. And so the question is that ,you know, where is the line? Is AWS the product or is it the platform, right? And where do you see that sort of going in the future? Like what are … I know I'm asking a lot of questions here, but I guess just for me I'm curious what you see as the continued opportunities for builders out there to build tools and services for other builders?

Ajay: I should just say yes and call it done, right? But I think it is going to continue to be both, Jeremy, and I think this goes back to AWS’ philosophy of building backwards from the customer. Right? I think what you're seeing reflected in the way AWS is evolving is also the sheer breadth of customer feedback and signals that you end up getting, right? And I think in my time at AWS what I‘ve seen is, there is a class of people who want AWS native, right. They’re, like I need this to be AWS, otherwise I'm not going to get it approved, I'm not going to get it through, I'm not going to use a, you know, pick your own from the toolbox on the side thing. You have to give me a native solution end to end and it needs to do the basic—that's good enough, right? So there's one school over there.

And there's the other one who’s saying, no, look I won't use what I consider best degree that works for my style, my productivity ,and others over there. And like I said in the beginning, right, like there's no way AWS can do this overall. So I think for AWS, you're always going to see the core investments in what we consider sort of the core aspects of the service, right? So security, performance, scale, capability is on enabling sort of API and service driven innovation. You know, I always think of this as the space being big enough. It's not like if, well, if AWS releases services, it's one and done. Ultimately customers are going to use the best services that they care about and I don't think AWS is the only one who can build the best service.

All the examples you just called out, right. Like, that are great ecosystem stories that are thriving and big over there that continue to go big. I think you see the same thing reflected in the way Lambda’s evolving like we talked about opening up the … the core aspect of the service which is sort of that distributed compute, democratization is going to continue to be something we innovate in. We have opened up an API on the areas we think where other people can do stuff. Like, well, hey, you want to bring your own runtimes and patch it better than we do? Go for it. You think you can manage operation controls better than you do? Here's an API go for it. Have fun with it. We're going to continue to offer an end-to-end vertical solution for those who do care about it. Right? Like so you do want sort of the sensible defaults, I think is a strategy that we're going to go over to there.

One thing I have a lot of partners bring up is how you can enable better crossovers, you know discovery, how do you make sure that it's not like they're able to just pick CloudWatch because they're forced to being picked into CloudWatch etc. And I think that's something we're going to continue to evolve. I actually like what the EventBridge team has done really nicely over here. So if you just go to the Lambda console and try to select an event source from EventBridge, it shows all the different ones who are out there, like all the different size providers show up at the same footing as any other AW service and I think that's a philosophy you will see evolve more.

So, I think for me the … if I kind of bring back to your original question, I think AWS’ core value-add is going to be solving what you call undifferentiated heavy lifting which lends itself to be more of the service tier, so to speak, not necessarily a platform there, and then enabling these experiences on top of it, some which are going to be able to be AWS native and others which are just going to be really enabled by the broader partner ecosystem over there.

Jeremy: Yeah, no. And you said to me earlier, you know, the goal is not to get everything right, it's just to make everything possible.

Ajay: Yes.

Jeremy: Which I think is quite fascinating.

Ajay: Yes. Exactly.

Jeremy: Well, listen Ajay, this was awesome. I love talking to you. Maybe I'll stop recording and we can talk for another 10 hours and not necessarily for the listeners. But this was great. Thank you so much and not only just for being on the show, but for everything that you're doing at AWS the, you know … I know you were there right from the beginning with Tim Wagner and the others, and just you know, making this what it is at this point. I think I've said this to others who have been involved early on, like, this has just changed my life, right? It's been a revolution for me and the way I build applications, the way I think about applications, and I think this has changed the world for a lot of people so “revolutionary” is the word that I would use.

So if people want to find out more about you, find out more about serverless, what AWS is doing there, how do they do that?

Ajay: So first of all, Jeremy, like, you know, thanks to you and the community as well, you know, it's the customers who keep us honest in helping the world of service over here, and I think one of the biggest powerful aspects of serverless has been the community around it and, you know, keep spreading the word, keep telling the builders or listening to your talk about how we can do better. Like I said, the path to the revolution is serverless and the fastest way to build is serverless and we're going to keep taking that through. I hang out on Twitter quite often. You can find me @ajaynairthinks on Twitter. You can find me on LinkedIn, and I'm usually pretty responsive over there. But otherwise you could always go yell at Chris Munns who is our Principal Developer Advocate and he has a way to find me, too.

Jeremy: Right. And a great resource that was launched not too long ago was serverlessland.com, which is really good. So go check that out. Awesome. All right, Ajay, thanks again.

Ajay: Yes. Thanks, Jeremy. Happy Holidays.

This episode was sponsored by Epsagon: https://epsagon.com/serverlesschats

View Details

About Angela Timofte

Angela Timofte is the Tech Lead at Trustpilot, a global review platform that helps businesses collective and leverage customer reviews. Angela has a proven history of transitioning legacy applications to new platforms and product offerings. She is driven to build scalable solutions with the latest technologies while migrating away from monolithic solutions using serverless applications and event-driven architecture. She is a co-organizer of the Copenhagen AWS User Group and a frequent speaker about serverless technologies at AWS Summits, AWS Community Days, ServerlessDays, and more.

  • Trustpilot: Trustpilot.com
  • Twitter: @AngelaTimofte
  • LinkedIn: Angela Timofte

Watch this episode on YouTube: https://youtu.be/jHE0VYfQUaY

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Angela Timofte. Hey, Angela, thanks for joining me.

Angela: Hi, Jeremy. Thanks for having me here.

Jeremy: So, you are the Data Platform Manager at Trustpilot, so I would love it if you could tell the listeners a little bit about yourself and your background and what you do at Trustpilot and what Trustpilot is all about.

Angela: Yes, of course. So as you already mentioned, I work as a Data Platform Manager at Trustpilot and I've been with the company for almost six years. So, quite a long period of time for a company that it's only been for like 11 years on the market. But yeah, I started in the company as a backend developer and then moved to be more of a full stack developer and now the data platform manager because my love for data was always there and I kind of did everything that I could do to move closer to the data. To be honest. no matter where I was in my career it was always data that, like, attracted me the most. Like how do you handle it? And to be honest nowadays, like, data is everything. Like you can't make any decision as a business without data and now everyone is seeing that. So it's really cool the position I am in right now because I can push all these, like, these data meetings that we do so that we take, like, the right decision. And then, yeah, Trustpilot. I mean, hopefully everyone heard about Trustpilot. At least that's what I want to think, but in case you haven't, you should use it. It's an online review platform and I mean our all … our whole mission is to help people to have better experiences when it comes to purchasing online. But of course, at the same time, we want to help businesses to connect with their customers and also improve their offerings and for that we offer all these analytics tools. So they understand better their customers and they know where they need to improve their business.

And I mean if you think about, like, the situation now with Covid, honestly Trustpilot came perfectly. Even for me, like, I'm talking from my perspective now, but, like, I had to order everything online, like, from food to, like, toilet paper. I didn't go to the store to fight for it. I went online and fight for it, and fought for it. So, but it came super handy because I'm based in Copenhagen and we don't have Amazon here and then you have to purchase from, like, small businesses and you don't know about all of them. So I had found myself searching on Trustpilot. Okay, can I trust this business because I've never heard about it, right? So I don't want to throw money out there. And yeah, that's what you can use Trustpilot for if you are a consumer, and especially in these times it's perfect because I can trust that what I find there it's real data and real people and I can get their opinions on all of these. So, yeah.

Jeremy: Awesome. So I am super excited to have you here because as much as I am a huge serverless geek, I love data; like, I'm with you on that. Like everything that I do, I'm always finding better ways to build or abstract data or interact with data. I built all these open source libraries, and I think all my open source libraries have something to do with data. And what we've seen over the last several years is this move to more and more serverless data, right. Like, capabilities that allow us to use databases that are more and more serverless, DynamoDB obviously being one of the big ones; all kinds of crazy stuff happening with relational databases.

So, you're an expert on databases and data and I would love to get some of your, you know, sort of your comments and insight on what are those choices that people have right now for serverless data. And also we're in the middle of re:Invent right now, so just in the beginning of re:Invent there's already a whole bunch of announcements of new things that Amazon has released and more options that are available. So let's start there. What are, you know, what are those serverless options for data that people have?

Angela: Yes, I mean you've already mentioned DynamoDB and for me that's like the first choice when it comes to building serverless applications, to be honest, especially when you think about scaling and that's my first pick and then Aurora and now we have Serverless 2 let's see what we can do with that one, right. Super excited to see more about the version 2 and then yeah you can use S3 Kinesis and they've released some other things now for databases like Babel Fish to, like, export the data up, and yeah. Then you had ... what was it called ... Glue as well …

Jeremy: Oh, yeah, Glue Elastic Views. Yeah.

Angela: Yes, which will actually replace some of the pipelines that I have probably with, like, cold DynamoDB streams going to Lambda then going to Elasticsearch, so probably that will change some of the things that I've done previously. But, yeah, when it comes to databases and serverless, I know like especially before, I don't know if it's still the case, but when people are mentioning serverless was always like Lambda functions, and containers, and things like that, but no one was talking about databases and I think it's a huge mistake because, I mean, no matter how scaleable you make, like, your services, if your database is not scalable, then, like, you're missing the point, right?

Jeremy: Right, right.

Angela: And yeah, and that's why I think it's super important to look at these serverless databases so that you make your entire pipeline scalable from one end to another. And that's where DynamoDB comes super-handy.

Jeremy: Right. So with DynamoDB, I think if people don't know what DynamoDB is, just go and look it up. It is an amazing key-value store document database. You can do some really cool things with that. But it does have limitations in terms ... especially if you're building like OLAP applications, right? Because you can't do all these different queries. You can’t slice and dice the data different ways.

So, beyond DynamoDB being a very good sort of ... I look at it as a perfect application or a perfect database for your frontend users that need to access data quickly where the access patterns are very consistent and you know what those are going to be. But when you have to start exploring data more or you've got relational things that need to be done, what are some of the options there? You mentioned Aurora Serverless. right, so let's dig into that a little bit and then let's talk about V2 because that's kind of exciting.

Angela: Yeah, so on Dynamo the way that we use it is, as you mentioned, it's for our front application, right, so that they are scaling accordingly and also that everything gets like pre-calculated and the way you start, like, you need to be very strict when you start your baseline Dynamo so that you actually take the benefits out of Dynamo otherwise, yeah, don't just throw data out there hope for the best. But, yeah, then, like, you can go to Aurora Serverless. The thing with Aurora Serverless, at least version one, we don't use it as much because of some of the limitations that you have with Aurora Serverless and one big one for us at least with that you can't import from history. And that I know ... I can’t remember at what conference I complained about that, but I know I complain about it when people from AWS were there. I was, like, I need this, people. And also the pricing on Aurora Serverless, it's quite, quite high. But of course, when I mention pricing like even if you go to Dynamo, which is much cheaper, you need to calculate, like, how much time, like development times, it takes to put things in Dynamo for instance or, like, for people to actually understand because it might not be so easy to do things in Dynamo. And then, yeah, with Aurora Serverless, unfortunately, we don't use it as much but that's why I'm super curious now with version 2 which seems that they are investing more into the serverless version of Aurora and hopefully they’ll work more and more so that we can use it in more, like, heavy production, like, workflows because before it felt like it's not really for your heavy workflows, I would say.

Jeremy. Right. Yeah, I mean in the scaling characteristics of version 1 was this doubling of capacity where it actually had to, like, move data between instances and it took like a minute to scale up and it just wasn't one of those things where it was as elastic as you wanted it to be. V2, very, very promising. It's very cool because the instances themselves scale which is kind of crazy and it will just … and I did a bunch of tests the other day, or last night actually, and just I threw as much as I could at the thing and it just laughed at me. I mean, it was, like, no problem. And then the other cool thing is according to the website, it says it's going to have all the features of Aurora. So that should mean S3 imports and global tables and all these other things. They actually said there would be global tables so that could be really, really cool.

Angela: I hope they listen.

Jeremy: Twice the price, though. Yeah. I know, I know. Twice the price though as ... just as the v1. I did some calculations on it though. I mean it still might be cheaper depending on your workloads and they say it's 90% cheaper, but anyways, I think that's a super interesting option. Right. The other thing you mentioned was S3. And I don't think a lot of people think of S3 as a database. But if you store data in S3, you've got a lot of options, right?

Angela: Yeah, that's … it's actually right that people don't think of S3 as a data store and it's also that it's been there for so long and all these other new “cool” have appeared. I feel like people forgot about S3 and how powerful it is and how flexible it is. And I think the problem with S3 it just doesn't come to your mind as like a data store. Like how would you go about? But it's very flexible once you start using it. And it's the same as with Dynamo: it might take a bit of time to, like, actually adjust to how you take your data, but I think it's a very powerful tool that people forget about. So, yeah.

Jeremy: Right. And there's a bunch of interfaces into it as well. I mean you can use S3 Select so on, like, really large files you can select just a portion of them so basically, you can query a file or an object within S3. And then you've got Athena, right? So what are your thoughts on Athena?

Angela: We’re actually not using Athena. Yeah, I know. So I can't really say much on, like, production work because we don't use it. That's my take on it, you know, we don't use it!

Jeremy: That's it! Well, I mean, you know, I think the funny thing is that ... I mean with this large of a footprint that Trustpilot has and all the different services you're using, I mean, again, it's impossible probably to use all of these services, right? So, you just have to pick the ones that actually work for you. But Athena, I mean, what I love about Athena is just the fact that you map over these S3 buckets and you can query it like normal SQL which, I mean, is sort of ... and gives you the performance of something like maybe BigQuery. Like what about things like BigQuery or Azure Cosmos DB or things like that. Have you played around with any of those?

Angela: Yeah. So the reason why we don't use Athena is because we use BigQuery. So, yeah, it's very powerful. And that's what we use to analyze over our data and precompute everything and then we push it back to our data space where we then, like, use it in Dynamo or Aurora or like any other database to actually use it in our applications, but for analytics we use BigQuery.

Jeremy: Awesome. Right. So, we just talked about a bunch of different database serverless that are available to you. And we mentioned some that are cross-cloud right? We're multi-cloud right? They're not just Amazon options. So this is something I think is a really difficult choice for people who are building new applications. You know, how do you ... and not just new applications but refactoring old ones as well. What do you have to think about, you know, when you're moving to either build a new application or refactor an old application? Like, how do you think about, you know, what you choose for a database? And when, and I guess, when is serverless a right choice for you?

Angela: Yeah, I mean we have, like ... we use both AWS and then Google Cloud, but we kind of try to stay in the AWS world when it comes to all of our production data and services and serverless and so on. So first is, like, think which cloud provider is the right for you. I'll go for AWS, but I mean, I might be biased there. But then, yeah, the way that you have to look at the application that you have, so if you start with, I know, Monolith and then you want to split it, I would go with serverless and that's something that we try to do at Trustpilot, to go with, like, serverless first. So we have this principle that whatever you want to build, new or refactor something, you should think of how you can do that in a serverless application because of all the benefits that you get from using serverless applications like scalability and price and so on. So I would start with that: like, think how you can put your application into a serverless infrastructure and then of course if that's not possible because there are still limitations on the serverless choices, then we go to containers. So it’s ECS or EKS, and then if that's not the right choice still, like, the last resort is an EC2 instance and then you just dump your things there and pay for it because you do that.

Jeremy: Right, right.

Angela: So that's kind of the mindset that we have around, like, when we go for something new or, like, refactoring. And to be honest, now it's almost everything it’s serverless when you think about, like, a new ... building something new, because you have so many tools there that you can, yeah, you can get around serverless. But as I said, like, there are still some limitations that, yeah, I find and I'm like, “Oh no, it can't be serverless!” And, yeah, you have to go for something else.

Jeremy: Right. So, you mentioned the sort of this process that you use at Trustpilot and that's super interesting because, again, I always love getting insights into other companies and how they go through these processes, and I know you had mentioned to me in the past your first attempt at moving data into DynamoDB because that's something you’ve really got to think about, right. I mean, and again, DynamoDB it's a NOSQL database. You've got to precompute a lot of data, you’ve got to think about your access patterns, so what ... tell that story because I think that's really interesting, sort of the experience you went through.

Angela: Yeah, absolutely. So with DynamoDB, we started to look at it around 2017 when there weren't that many tutorials about it. So, we started to look at it, we were ... so at that time we had all of our data and MongoDB and then, we’re like, so used with how the way MongoDB was working, that when we started to look at DynamoDB, we're like, what is this all about? Like what's with this key-value store that you're trying to push us to use and, yeah, it wasn't that good for us to start using DynamoDB, to be honest. And in the beginning, we're like, yeah, this is just some silly type of database and we're going to use it for, like, I know some simple scenarios that we have and, yeah, that was the beginning. Like, we really didn't think too much about DynamoDB and now we changed to be our preferred type of data store. So there you have it.

Jeremy: So what's your advice, I guess, out of that? I mean because, again, there are a lot of different options like you said, learning S3, figuring out what you can do with that. You don't even use Athena so that, I mean, so knowing how or what you can do in Athena, like, so, what's the advice for somebody that wants to make that leap to some sort of serverless database or, you know, Dynamo or Aurora Serverless or something like that?

Angela: Yeah. I mean it’s ... definitely try to get all the information that you can about it. And yeah, look up for tutorials. Nowadays, like, yeah, you can find a lot of information and then start using it, practice, because especially when it comes to DynamoDB, like, it's quite, I mean, it's not that easy to see those patterns, to be honest, especially in the beginning because if you come, for instance, from a SQL database and you kind of know how to store your data and how your query will look like, your indexes and so on. But then you go to Dynamo where you have a primary key and sort key and do things with it, you know and, are like, “ So what can I do with this?” So it takes practice. It takes a lot of practice to see the access patterns and yeah, I would say like whenever you have something new that you want to store, like, just give this a try to see how would you do that in Dynamo. And also watch lots of tutorials from people.

So yeah, that would be my advice. And when it comes to companies if you want to push people towards something new, I would say really give them time to adjust because companies are trying to say, like, “Oh, DynamoDB will save a lot of money, Aurora Serverless is saving a lot of money,” because they just went through a presentation where they were saying that, and then it's like, yeah, we need to do this and then they expect that overnight, but it's not like that. It's like if you actually want to get the benefits you need to give people time to actually adjust to whatever new technology you want to adopt.

Jeremy: Right. And that's something I find with DynamoDB, too, is you can't just sketch it out and put it in theory. Like, you have to actually start using it; like, you have to start putting things down, putting data in there, and figure out how you can actually access that and whether it's going to work for you, you know with what you're trying to do, because you will get your modeling wrong a number of times before you finally get it right. So, don't just model something and then throw it into production without going through a number of iterations.

It's funny though because I remember way, way back when, and I'm getting much older than I'd like to admit, but when I first started using SQL and I was writing, you know, basic select queries and insert, like, fine easy, you know, I was using SQL server a lot in MySQL. But you get to a point where you know it seems like third normal form and building all these things it's just impossible to understand … like, I mean, it eventually gets to the point where it's very simple to understand once you get it, but it does seem like a daunting thing. Then you make the move to DynamoDB and I look at how you would structure a relational database and you're like, wow, that is easy, like, that's so simple. Now, I've gotten much better with DynamoDB and understanding the patterns but, yeah, it takes an investment. It takes a huge investment for you to get there.

Angela: Yeah, absolutely. It … And especially now that people have experience with other types of databases I think it's more difficult to make that switch because it is quite different and it takes time to see these patterns. And as you said, like, you'll get it wrong many times probably before you get it right, and, like, yeah, you need to start with your queries, you know, you kind of need to do a lot of planning before …

Jeremy: Right.

Angela: ... with the Dynamo, right? Because you can't just dump your data, you need to do all the planning on, like, okay what I expect, like, everything needs to be planned before you actually do it. But of course, I mean it's easy now to play with the data and moving, like, around. Like, before was like two years to transition from, like, one table to another, you know.

Jeremy: Right, right.

Angela: Yeah, it was painful but now you can play with data much easier. So, in that terms ... yeah, here you can practice with Dynamo and see if you go … if you can store things in there or not. I would say you can but it might not be super easy in the beginning.

Jeremy: Right. Yeah, so I think the takeaway here is if you're a company and you're putting stuff into DynamoDB and you get it wrong the first time you're not alone because we've all done it, right, so just keep on working on it and you'll get there. You mentioned a little bit about limitations, like, you run up against limitations when you're moving things into, you know, some sort of serverless database offering, or just serverless in general. I mean, there are limits there. So how important is it to understand those limits because I find that they’re very high but when you hit them, they’re also very painful.

Angela: Yeah, I mean, when it comes to, for instance, with Lambda, one of the limitations that I'm kind of ... I find myself lately hitting it quite a lot is the concurrency. Like, I want to … and not just limit the concurrency, but when I limit the concurrency, you know, to not throttle everything that is being invoked with and I know, I was, like, I still ... because in my mind was like yeah, we'll just do this. It's simple. It's serverless. This event is triggering this Lambda and then this Lambda will call some third-party API and then we've hit this limitation of which I didn't think about that like the third-party API had its own limitation with like ... or you can only call up 50 times per second. And then I was like, oh, how am I doing this with serverless? And I’m kind of trying to choose to stay in the serverless world, you know, and, like, really like move things around and it got to like a super complex solution. And I was, like, okay now I just need to say no to serverless and move to a container which will be so much easier to explain to people what I'm doing than try to stay in the serverless world. But I spent a lot of time with, like, finding a solution in there and, yeah, it's good to note the limitation so that you don't just spend a lot of time trying to reinvent things just so that you’re staying serverless. But yeah, as you said, like, that’s one thing that you need to do, like, learn the limitations so that you know what you're dealing with. But they are adapting everything and like, yeah, updating all these things that you kind of need to look at the news all the time, I would say, because they are there constantly working on this.

Jeremy: So, yeah, I totally agree. I mean, that's one of those things to where it's like ... and you mentioned this I think in the beginning where you said, you know, you get one part of your application, one piece of your architecture, that'll scale really well, and then you got other pieces of your architecture that won't, so solving the database problem is really great. I mean, things like Aurora Serverless v2, DynamoDB, solve a lot of those. You know, Lambda functions can scale infinitely, you know, whatever they want to do and then, but then you still have those problems with APIs, right? So you still have to figure out how do you manage those quotas and do that. And I think a lot of people get frequency and concurrency confused. Right? So if you or, you know … so if you have a quota of 200 calls per minute, you can't set your Lambda functions 200 concurrencies because what if it only takes, you know, two seconds for that to run then you're running, you know, I mean, thousands of invocations against it. So lots of things to think about there, not necessarily solved yet, but like you said, getting there.

So, alright, so let's go back to Trustpilot for a second because I'm really interested again into kind of digging into your architecture and how you solve some of these problems. So let's start first of all ... I mean, you're obviously a huge fan of serverless, right, which is great. Love serverless enthusiasts, especially serverless data enthusiasts, but what about the rest of your organization? Like how did you bring that in? How did you say, “Hey we need to go to a serverless database, or we need to start thinking serverlessly.” Like, did that change happen within your engineering?

Angela: Yeah, so it definitely took us time and we started quite literally looking at Lambda functions. And I remember it was after they released or announced it at re:Invent and we have a few colleagues that went there, and when they returned from re:Invent, like, every time when someone comes from re:Invent, you’re very excited. And they were like, “We're going for this,” and all of us were like, “What? What's this Lambda function? Leave me alone. I’ll just build my application here the way I know.” So it took us time and I think, like, what helped us was this idea of, like, just trying using serverless, but you're not forced to use serverless. And we were constantly trying to teach people, and, like, constantly talking about the benefits and showing, like, we had, of course, we had people that kind of wanted to use this from the beginning because it was something new and you always have those people in your companies, right, that they are attracted to whatever is new.

So we had people testing out and then, like, sharing with the rest of the company on what they achieved and how it's possible, and then slowly get everyone in the company starting with, like, serverless architectures and then, like, that's how we grew. And the same with the database, the serverless database, which is like some teams we tried to use it and then show the benefits to the rest of the company and that's how we kind of got everyone to be interested in serverless. And they've seen ... so we kind of showed the benefits in our company not from just general presentations, you know, marketing presentations where it's like this is serverless, this is what you get, you know, it's like no, no, let's get real. Like what do you actually get? And, yeah, that's how we got people excited and now everyone kind of wants to go to serverless first.

Jeremy: All right. Now, so what about your ops team where you like your SREs like, what was the impact on them?

Angela: Oh, I think they looked at them the most I would say, because all of the sudden they didn't get requests like, “Oh, can you please give me access to these,” or, “Can you create an EC2 for me?” And you know, those requests that no one really wants to deal with so, yeah, now they can actually focus on building the infrastructure that helps us and building, because before it wasn't about building it was about, like, supporting developers in doing their job more and now as a developer you can do all those things yourself. So the SRE or DevOps team can focus on building the infrastructure that helps us in different ways.

Jeremy: Right. Yeah. And I know I had a colocation facility for many, many years and I would always get the text message at 2:00 am that a blade went down or that, you know, there was a problem with the SanDisk Array or something like that. And you're always driving to the colocation, you have to physically change out hardware. So besides just not having to necessarily do that, I mean, obviously for me, and I say obviously, but maybe this isn’t obvious to people, but since I went serverless, I haven't gotten a single alert at 2:00 am to say that a server went down or something like that. Which made my life a lot better.

Angela: Yeah, that's actually right. Since we moved our data and our services as well to serverless, like, no more alerts at 2:00 am at you, like, “Oh, you need to scale this database,” or, like, “You're having problems with this service, you need to provision more.” So that's a huge benefit. And I know people that are on-call, they appreciate that for sure, like, no more midnight calls to do like, yeah, I need to click these three buttons, and you’re, like, really? Like someone can't do that automatically. You know?

Jermy: Yeah, It's a huge ... I think it's just a huge morale boost. I mean, and I love that these are more and more stories of this where you hear that operations people might be, sort of nervous or I guess intimidated by serverless, and then the ones that implement are like, “This the greatest thing ever, because now I can focus on things that actually matter,” which is something that's important as opposed to, like you said, just, you know, doing the daily toil.

So, let's talk about the Trustpilot architecture and look at sort of what you have now. I know you gave a presentation a while back. I'm sure things have changed, you probably move more stuff there, but just what's the typical overview look like? I mean, are you just using DynamoDB or are you using a sort of a broad range of databases?

Angela: Yeah, I mean one thing I always try to mention when it comes to database use is that you shouldn't put everything in one type because there are so many types of databases and, you know, you need to look at the purpose of that database so that you use the right one for your use case. And yeah, that's what we try to do at Trustpilot. We have data in Dynamo, Aurora, ElastiCache, Elaticsearch. We still have some data in MongoDB as well, in Redis, and like, you name it, different datasource for sure because we have different use cases and that's what you have to look at when you choose the data store.

Jeremy: All right. And now are you continuing to, sort of, evolve your architecture and move things away from MongoDB and some of these other sort of non-serverless options?

Angela: When it comes to MongoDB, I mean, we've tried to move like a lot of our data from MongoDB because of scaling. So we moved a lot in Dynamo, but some of the data that we still have MongoDB, it just makes sense to use MongoDB for the use cases that we have, so that's why we still have some data there. It’s still the right choice for us. But who knows? I mean, now we have Document database from Amazon which supports MongoDB. So that might be the right solution for us. But, yeah, for nowMongoDB works, so we kind of keep the data there, because we've also been in this, like, continuous refactoring for a very long time. So if ot works, we keep it there for now. You know?

Jeremy: And I think that's not uncommon. I mean it's going to take a while I think for people to get everything moved over and it's really hard when you have something that is running okay, and it's working for you and is not a problem for you to say, “Maybe we'll just leave that for now,” and then work on some other things, but you're definitely right. Once you've established, once you have sort of a legacy application, it is hard to think about just spending all that time refactoring it just to, you know, just to get the data piece changed, especially if you're not having the performance issues.

Angela: Yeah. Yeah, exactly. I mean, it is a lot of time that you need to invest to do that and we've done a lot of it, not to say that we have invested a lot of time in refactoring because of all the scaling issues that we were having previously, but now that we've hit like a moment where we kind of we can breathe in that space and we can focus on other things. You know, it's like okay, let's leave this as it is for a second because it works and focus on other things that might be on fire and that's the case we have here with MongoDB like we don't have that much data left. It's your point, quite a lot of data there, but it works for now, and like the complexity that we have in there it doesn't make it so easy to refactor. So that's the reason why we, right now, we are keeping things as they are with the data that we still have in there.

Jeremy: Awesome. All right. So I love to talk theory and a lot of what we talked about I think will help people but what about actual, like, real-world stuff? So let's talk about an example. I know you have a couple of examples here, but what are some of the problems that you were having? Because, again, if it's working, you know, maybe we don't need to invest the money. But when we start to have problems with things, you know, we need to think about refactoring those and maybe taking a serverless approach. At least that's how I think about it. So I know you've done this a few times at Trustpilot. So what’s a practical example? What was that problem that you solved with serverless that made it the right choice?

Angela: Well, it's all about scalability, to be honest. That was one real problem that we were having with scaling the database, the Mongo database. And to be honest, some of the issues were because of the way that we configured things and the way that we stored, And, yeah, exact example for this one was, in the beginning, we were storing all of our data into one cluster in MongoDB. And then, like, whenever you're putting a lot of pressure on one data point the entire class that would say also all the applications that were using that data were down. So we changed that in Mongo, but then still, like, scaling was a real problem for us.

And the company has grown a lot. So, yeah, that's why we knew we had to do something and that's why DynamoDB is the right, definitely the right choice for us, because it's safe. Yeah, it's scalable. And yeah, we know, and I shouldn't say we know, but we definitely hope that the company will grow even more, and that makes it the right choice because I don't see us ... I don't see the need of refactoring again what we have in Dynamo, for instance, because I know it can handle this scale even though it might double, you can still handle it. So that makes it a great choice for us because, as I said, we've been in this refactoring mode for a long period. We've been changing from MySQL database on-premise to MySQL database in the cloud then from MySQL to MongoDB. So, once you factor in that now that I can say, like, we have this in Dynamo, and I don't see the necessity of switching to something else because of scaling issues. It makes it just perfect, I would say.

Jeremy: Right. Yeah. So, you mentioned in that talk that you gave, you know, this idea of the user sign-ups and expirations on, I think was it, like, temporary accounts and things like that. So that was one of the big things. Was that one of the things you moved from Mongo to DynamoDB?

Angela: Yeah. Yeah, that that was one example. So the scenario was that we have, so people can sign up, but then they have to activate their account. That’s quite, like, a normal scenario, right? So they have to activate and if they don't have to be in like 30 days then we need to delete the account. And we're doing that in our only Mongo database for where we're keeping all the data for consumers. And of course, we're putting a lot of load, unnecessary load, on our primary database. So we decided to actually take this entire scenario out and we started, okay, of using events when consumers sign up. We will send an event to store some data in a DynamoDB which would say this consumer signed up and then we'll have another event coming from the activate ... like the activation API, saying this consumer activated, so then we'll delete the data in DynamoDB and we had one Dynamodb with all the unactivated accounts. And then from there we could look at, like, when the account was created and we can delete whatever accounts that are not activated in time. So this way we took that whole pipeline to serverless in its own context and, like, its own service and then doing it’s spin there separate from our primary data. And we did it with, like, three events and DynamoDB and then, yeah, another Lambda that was listening to ... was querying this database.

So it was a very simple scenario but we took a lot of load from the main database by not going like every, I think was like every day, queried the database to get like all un-activated accounts. And so, yeah, it was a very simple scenario, but like this just shows how you don't have to, like, refactor your whole database. You can just take parts of it or, like, queries like whatever it … This was just a scenario and we took it out and its own being ... I haven't checked it in, like, a very long time because it's just working, you know? I'm thinking maybe I should go and check it. No, but, like, that's like one example, where, as I said, like, you don't have to refactor the entire thing. You can just take part of it.

Jeremy: Yeah, and one thing that I love about that example is I think a lot of people who think, oh if I want to tack on or I want to go serverless that I've got to somehow re-engineer my existing application and that's a perfect example. Think if you got flooded with a bot or some sort of attack or something like that that was just adding new accounts and adding new accounts; it’s not even touching your Mongo database, right? Because it's all getting buffered in this DynamoDB table that is going to scale. And then unless it's activated right then, those events don't get passed through. So you're taking off all kinds of load. You're making your primary database that is already scalable but maybe can't handle those transactions or that number of transactions, you’re offloading all of that and that's just a perfect example. I love that and it's a great, you know great ... I think anybody could use that exact same scenario for their sign-ups right now and take a ton of load off that database and, like you said, just the maintenance of having to query through a MongoDB and look for the accounts that were not activated yet. That's just ... it's just a waste. I mean that's a complete waste. And with DynamoDB were you just using, like, the TTL to expire on activated accounts?

Angela: Yeah. Yeah, so that's like ... I remember in the presentation when I was presenting this scenario, I gave like two options because one is to look at the ... just have a cron job that will trigger your Lambda to, like, query the DynamoDB based on the data that you have in there, but the other one is to just turn expiration, like expired because, you know, like, your accounts will expire in 30 days, for instance, and then you can use TTL to expire those items. And then you can use this trim to trigger a Lambda that with that data, for instance. So that's another way of looking at these. So you have multiple choices. That's a good scenario to be in and it depends on what you're doing, right? Because remember on streams he will trigger any kind of activity and in our case, we only cared about expired so I think we just went for cron job in this instance, but if it was possible to filter what kind of events the stream was in, definitely TTL would have been the right approach for us.

Jeremy: Right. And even getting them in batches, I think the TTL approach would be interesting because then you could have that stream every time an account wasn't activated you could forward that off to S3 or something like that, so that you could store a record of the accounts that never were reactivated, have all kinds of data on that but not have it in the operational database which, again, is another really, I guess, good pattern to use in serverless, right, is to say the things that need to be in operational databases, store those and operational databases; the things that can be stored for reporting and for, you know, data like analytics and that kind of stuff put, those in another place, like S3 or something like that.

Angela: Yeah. Yeah. Absolutely. I mean that's the beauty with serverless that you can split things and you should split things like all of these as you say, like keep your main database doing what's important for you, but all these like extra scenarios put them outside and have them on their own so they just work on their little things and they are very good at doing that thing, right? So that's the beauty with serverless and it's something that people should consider and not just put everything in one solution and build like a monster.

Jeremy: Awesome. Well, that's pretty good advice to end with, Angela. Thank you so much for taking the time to chat with me. Super informative stuff. If people want to find out more about you, connect with you, how do they do that?

Angela: Yeah. So, I mean you can find me on LinkedIn by using my name or on Twitter also by using my name. And yeah, I'll be happy to talk more about databases, serverless, or whatever, anything else. So yeah, just use my name.

Jeremy: All right, and so and Trustpilot, Trustpilot.com if you want to check that out and sign up there. So I will get all that in the show notes. Thanks again, Angela.

Angela: Thank you. It was a pleasure talking to you.

View Details

About Rodric Rabbah

Rodric Rabbah is the co-founder and CTO of the serverless computing company called Nimbella. He is also one of the creators and the lead technical contributor to Apache OpenWhisk, an advanced and production-ready serverless computing platform. OpenWhisk is open source and offered as a hosted service from IBM and Adobe. It is also deployed on-prem in several organizations worldwide.

Twitter: @rabbah
Personal website: rabbah.io
Nimbella: nimbella.com
Apache OpenWhisk: openwhisk.org

Watch this episode on YouTube: https://youtu.be/xVZhFHmEuKY

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Rodric Rabbah. Hey Rodric. Thanks for joining me

Rodric: Hey, Jeremy. Thanks for having me. I'm really excited about our discussion today.

Jeremy: Awesome. So you are the co-founder and CTO at Nimbella and I'd love it if you could tell the listeners a little bit about Nimbella, but I'd really love to hear about your background as well.

Rodric: Okay, yeah, thanks for giving me the opportunity. So, I started Nimbella about two years ago, just over two years ago, and it was after a long stint at IBM for 11 years, and IBM Research specifically. And there I did a number of things that touched on programming languages, compilers, hardware synthesis, and FPGAs. And the last project I really did was creating IBM serverless functions offering, which is now called identify functions, but it started as OpenWhisk. And OpenWhisk is now an Apache project that we have donated to the Apache Foundation four years ago and really is what started my serverless journey six years ago. So my background is mixed. It has experience from programming languages, compilers, hardware systems, and I've done a lot of things that I think have a common theme across verticals, or I build verticals that cross lots of different layers of the stack.

Jeremy: Awesome. All right, so I want to start with IBM and OpenWhisk because this fascinates me where, you know, six years ago and just recently, I mean, it was a six-year birthday of AWS Lambda and I think it's, that this sort of kicked off a massive sort of investment and maybe almost like a space race except for the cloud, I guess, a serverless race against all these different vendors. So, you were involved with this very, very early on, right? I mean, like I think it was your project, right? So I'd love to hear how this all came about, like, why did IBM suddenly say, “Okay, we need to build a serverless offering?”

Rodric: Right. So, before I started working on OpenWhisk, I was doing something completely different, where I was debugging hardware and looking at how to generate hardware from software. And then we saw the Amazon Lambda announcement at the time and it was literally right around this time. And as soon as we saw it, it was sort of one of those moments where you just realize, “Oh, my God, this is a dramatic shift in technology that's coming.” And even though it was basically day zero you could see, you know from the right perspective, you could see what the future is. And now I say serverless is inevitable, because you know, whether you're on the bandwagon or not, you will be in the future because that's the only way developers will want to build. So, in the early days, you know, we saw the announcement we were looking at the project we were doing like they gave us the terminology that we were looking for. It was just take hold and just run it in the cloud and we were trying to do something like that with hardware, you know take software and now we can accelerate for you on an FPGA, which is reconfigurable hardware without you having to worry about compilation running.

And so we were doing work in a similar context in a similar area but completely different context, when we saw it. It was like this is it. So we got together as a small team within IBM, six people, and we were having discussions about Lambda and the future. What did it mean for IBM Cloud? And you know from IBM Research our job really was to sort of look at technology that's on the horizon and you know, five, ten years in the future and start thinking about, what does that mean? And after about, you know, a few weeks of just talking about it, I got tired of talking about it and over a weekend I built the first version of what became Apache OpenWhisk.

And it started with, you know, a command line tool which is the programming interface essentially to the cloud, allowed me to create functions, run them, get their logs, and recorded a short video by Sunday night and then sent it in, you know, Monday morning it got in front of the right eyes. You know that whole thing about being at the right place at the right time and it started circulating and from there, we're like, okay this is turning into something, and off we went. So it was sort of the Amazon Lambda landed on the scene. There was recognition that this is something really transformative into the future and then the will to just build something and once you start building something I think good things start to happen, you know, when you're surrounded by good people like we were at IBM. And the project root, I mean, we were three people, we launched the early version of OpenWhisk internally; it was called BlueWhisk, I think, at the time and, you know, I think within one year of when the first commit to the project started to an IBM announced at their big developer conference, it took us basically one year from commit to launch.

And we launched out of IBM Research, which was, again, unheard of. And right around the time we launched, Google Cloud announced Functions, and I think Azure also announced Functions. So, we weren't the only ones, sort of, that saw that shift coming and everybody really started basically saying, “Oh yeah, there is an arms race here or space race.” And I was really excited because it sort of transformed what I've been doing and I think it's been really exciting and rewarding for me.

Jeremy: So, I absolutely love this idea of these sort of ideation, like, these genesis meetings that happen in organizations where you're like, all right, there's this major transformational shift that's going on there. So, to the extent that you can and, again, I don't know if there were executives in there, who were in that meeting, you don't have to disclose that, but if you, to the extent that you can take me inside that meeting, what was the conversation. Was it like, “Oh, we just need to do this because we need to compete with Lambda,” or was it, “we need to compete with AWS,” or was it something where the structure of IBM Cloud realized that this truly needed to be done?

Rodric: Right. So, it was a bit of the latter. And in fact, when we started coding we tried very carefully not to use the word “Lambda” anywhere, that was sort of just IBM bureaucracy that maybe was ingrained in us, but it was a bit of the latter. I mean, I remember the meeting very well. I remember who was in it. I remember where everybody was sitting because it was that kind of transformational meeting, at least for me, and as I saw it. And it was a recognition that, you know, if you wanted to move applications to the cloud or you wanted to transform an organization and become more, you know, what people call cloud native today, basically using the tools and technology in the cloud, you had to do something different and what we had been doing wasn't quite working. This whole shift on lift strategy doesn't quite transform your business and to do the value innovation and sort of pushing up the stack to extract more value at your organization, you had to do something like this.

It wasn't complete buy-in as we started the project, as we built the technology we were fighting against currents that were trying to drag us in different ways: containers and containers service was just getting started. Kubernetes was just, you know, landing on the scene in terms of popularity and we had to basically say, “No, the future is here,” and we built and tried to control our destiny as much as possible. There were senior directors, and as the team grew and we were presenting to more and more people, there were executives, IBM product line managers, and, you know, by the end of first year our calls were fairly big. We were doing to two-week sprints where every week we used to call them shock-and-awe sprint because like what shock-and-awe features can we deliver the next two weeks and that became a theme for our team. And it was really fun to do because as we did this, we were sort of operating the prototype internally more IBMers would sign on and start using it.

And we started using it and so it was really exciting and by the end we had a lot of buy-in because IBM Cloud Functions had to be launching that needed sort of a business line justification and buy-in. But early on I think it was primarily out of research and sort of … but it was senior directors were, sort of, seen participating in that conversation and it was really exciting. Thanks for letting me relive some of that, six years ago.

Jeremy: But no, that's awesome. I mean, again, like I said at that idea of vision, it's hard sometimes. I think it's really hard when you're in technology to sort of pick what's the next big thing. And, again, we've got a lot of serverless haters out there and people who still love containers and Kubernetes and all that other stuff. Not that they necessarily have to compete, but I do love that when you were at that moment and that meeting where you say this is it, this is the next thing, so that's pretty exciting. But now you mentioned in there this idea of lift and shift, right? And this is something where I think most clouds took this strategy very early on to say how can we meet customers where they are and make it very easy for them to just take their on-premise applications and move them into the cloud which is why, you know, we're loaded up with virtual machines and EC2. At least AWS cloud I think is still the biggest moneymaker that they have there. But with this transformation to serverless, I mean, you have a lot of limitations, you know, there's a lot of refactorings, but sometimes you completely re-architect your application. But what about just building this in general. I mean there must have been a lot of technical limitations to to get around, right. I mean, you ... I know AWS and things first started building their stuff on EC2 instances. Is that how OpenWhisk started too, just running on a virtual hardware?

Rodric: Right. There's so much there that I would love to talk about and see how many of these I can peel off. So, yeah, we knew we had ideas of how Lambda was running and executing and then we looked at, “How can we build this?” Obviously we had to run on IBM hardware and the cloud that IBM offered us and that did put some constraints on how we actually architected the system, and some of those features, if you want to call them that today, are still with us. And I think that played to our advantage especially as the Apache project has moved towards more Kubernetes native sort of layer that you add on to give you that serverless experience. But early on we were deploying on VMs and, you know, to auto scale up and down required many minutes, so it wasn't latency that you can sort of just easily hide. And so that meant that we had to rethink or sort of think about the heuristic that we would build to give you that elasticity a serverless solution of I can run a thousand functions and some other user comes along among the thousand functions, “Hey, they just scale and they all run.” So we built a bespoke scheduler and a custom you heuristic for how we do the scheduling and it did influence essentially the architecture because we just couldn't bring up new VMs fast enough when you needed it. So there was a bit of that constraints that played into it.

I think this whole notion of containers versus functions is really still with us and it's for a number of reasons. Some of them you touched on. It's hard and you know to go into serverless, you’re re-architecting applications. So it just doesn't fit the shift and lift model; it’s fundamentally opposed to that in a sense and so a shifts and lift looks attractive, but to really buy into and get the benefits of what serverless promises, whole notion of less operations, more focus on value creation, it's necessary that you sort of be architects. So the key is gentle migration and acceptance that there will be a mix of technology: containers, VMs. And so our solution was to pick containers early on and that allowed us also to run containers as functions. So we started with I think the first run time we had was Node and then we added Python after that.

But very early on people said, “Well, I'd like to run my job application,” or, “I'd like to run some other language that you don't support,” so we were able to say, okay bring your container. So very early on OpenWhisk offered … and we might have been the first that maybe offered this mix of functions as code, basically zip file, and functions as Docker container that you just pull from a Docker registry and just go. So, and I think that helped people, sort of the gentle migration, sort of shift and lift, okay, I buy in and I start to get a taste and then I start refactoring. And I think that's how you still have to do it today. It's sort of this whole gentle migration stuff.

Jeremy: Right. Yeah. I think you make a good point about the hybrid stuff. I mean for a very, very long time especially for any larger business who is already either partially in the cloud or is migrating to the cloud, there's going to be a mix of everything, right. There's going to be VMs, there's going to be containers, and hopefully they start moving things into serverless. So when you built this, though, so you built this bespoke scheduler you talked about, you've got, you know … it's running on VMs, you eventually adopted containers and so forth. But when this launched, like this was ready for primetime, right? Like this wasn't like a little side project thing that just happened to go into production. I mean, this was like enterprise-grade, right?

Rodric: Right. Yeah, so we launched February of 2016, I believe, and we had already been in production for several months at that point; maybe 2015, no, 2016 right, and it wasn't, you know, it wasn't still … the code was quite mature and I think we just recorded another podcast with some of our early partners. Adobe jumped on the project fairly early when we open source, and they were an impetus essentially for joining Apache Foundation. So the code quality was solid in that regard and I used to joke that, “Hey, the system is bug-free.” And I meant, you know, it wouldn't crash for a segmentation fault or things like that and it was true sort of held for a long time we didn't have our first real crash from a segfault for like two or three years and I remember it because somebody Slacked me on an IBM channel, and said, “I thought you said this was bug-free!” So it's sort of stuck. No, it was ready for prime time.

We were already doing many thousands of containers a day sort of churning through containers and our solution was basically you take a function, we didn't create containers per function. We had this notion of stem cell containers which are unspecialized containers ready to inject code into and then once you did that they became specialized for a user and for a function, and that allows us to sort of do things with speculation. We can pre-warm containers, and really allows us to deliver performance that was on par with Lambda. And independent benchmarking today still shows IBM Cloud functions, which is probably the largest deployment of OpenWhisk, maybe Adobe would be second, you know does extremely well against Lambda in terms of latency and throughput. So, yeah, so it was really looking at sort of how do you deliver performance. How do you build this technology. And then how do you meet people in terms of giving them a gentle migration step towards this new paradigm.

Jeremy: Now, I know you've been removed from IBM for a while, but it looks like the project now, the preferred way to run it is on top of Kubernetes, right?

Rodric: Yeah. OpenWhisk, like we said, started on VMs but just like with serverless and Lambda, and you saw it was a future here. It was inevitable. I think Kubernetes quickly started eating up every other container orchestration system on the planet. And so we had to shift the project, the open source project, to support Kubernetes. And IBM also had to do this migration. We were already live in several geographies around the world. And so we started with one region, moved that to Kubernetes, operated that for a while and then, you know, the team got the confidence to roll that out. That was right around the time, actually, I was leaving. I think they had just launched a first on Kubernetes version of cloud functions and off they went. So the project now is Kubernetes native, if you will, or basically you can deploy it with a Helm chart at the risk of a patchy repo.

But we did something different and this is where maybe OpenWhisk stands out against some of the other Kubernetes serverless platforms out there today. We don't delegate the container orchestration to the Kubernetes controller and the Kubernetes container orchestration system because it's too slow. If you're looking at the kinds of workloads that are short running or that are Lambda style where you want to invoke really fast get-responses, I think I’ve seen a number of studies that said the average execution time for Lambda is well under a second and even several milliseconds. Spinning up containers that fast on Kubernetes just doesn't work. It wasn't the time for this. And so unless you solve the problem deep in sort of the Kubernetes scheduler, you have to bypass Kubernetes for container orchestration.

So, OpenWhisk until today really does that for the large enterprise deployments. You can use, you can delegate to Kubernetes, or you can sort of bypass and use a bespoke container orchestration system that really allows you to sort of deliver the best performance. So if you want Lambda anywhere other than AWS, there's really only one answer,might be, and that’s the OpenWhisk project.

Jeremy: All right. Well, we don't have enough time to discuss and solve all Kubernetes’ problems on this podcast. But what I would like to do though is ... so you left IBM and you started, I guess, we’ll set this up for you: you started Nimbella, right? And I want to get into Nimbella because I think this is really fascinating, what you and your team are doing over there. But what was it, you explained it a little bit, but what was it about sort of the current landscape? I mean, you'd already built OpenWhisk and had a tremendous amount of success with IBM, you know with IBM Cloud Functions, it went into Adobe, and then you've got all these people using it but you basically said serverless isn't good enough and you did this other thing. So what was it about the market or the current landscape that made you say, you know, we need to do something different here?

Rodric: Yeah, we needed to do more, I think that's the best … it's like, yeah, this is great, but we could do so much more, and to me, I sort of really viewed it as the introduction of Fortran for the IBM Mainframe. We're at that level of innovation in terms of how early we are on this journey and it was a recognition that, “Hey, there's a lot of problems still unsolved, from how do I debug. How do I look at this whole notion of now I’m breaking up applications that were large monolithic into smaller fractions. There's rich opportunities for system dynamic feedback optimizations where, you know, if I start with a bunch of functions do I fuse them together and run them as a monolith because it's more efficient, o by taking advantage of resources that are specialized like the GPU or a TPU if you're doing a IAM answer flow and things like that.

So sort of looking at it from a pragmatic perspective saying there's a lot of opportunity here to do more and recognizing that from the technology perspective, it was so early use it was hard for developers at large enterprises to really get started and it came from a number of reasons. One was this whole notion of how do you build for the serverless style when you can't quite run locally, you can't quite debug your code in Vivo. And watching some of IBM's early client sort of adopt this technology be successful in the end, but what it took to get there, the questions that they were asking you in some ways influenced my thinking as we started Nimbella. And it wasn't just to compute. I think functions of the serverless touches compute. It ignores the whole data aspect of applications and really this is where I started. I was, like, okay, we've got this model for compute with serverless functions. We've got container service. We can mix the two. What about the data model and looking at how do you marry serverless data model? What does it even look like with compute? That's where the genesis for Nimbella really started. Like, I want to be able to build complete applications. I want to deliver this promise of, don’t worry about the resources being allocated for your data, don't worry about replicating it. Don't worry about some of these synchronization and consistency models because for a lot of those there are good solutions that we've learned over the years of sort of distributed systems, research, and technology that we built.

So, I wanted to bring the two together and that's what really started Nimbella: can we do this, can we do this marriage of stateful and serverless and bring them together so that we can continue to deliver on this promise of, “Hey, as a developer, I just want to build. Everything else should be taken care of for me.”

Jeremy: Yeah, and I think that's a thing ... we should probably just sort of define state or at least give the listeners a little bit of background on state. So with serverless functions, or at least with the traditional serverless functions we've seen over the last several years, there is no shared state in the sense that, you know, when you reload, when you execute a function again that it's going to have all this information there, especially because every single request typically runs in a new container. And even if it runs in an existing container, there's no guarantee that it's going to run in the existing container that you were just in that has the same data there. So I know that AWS has added things like EFS integration and there's more things that are happening there. But really even with something like EFS integration, most of the time when a function triggers if you need data in there that wasn't passed in as part of the event, you need to rehydrate that data. So what are ... maybe we can talk about the kinds of state that you would need in the typical application and which ones really are kind of missing from serverless?

Rodric: Right. Yeah. I think you’ve really framed it exceptionally well and I describe it in terms of locality, right? So when you're running functions in Lambda or really any serverless platform, you don't know where your code is running. You don't know the container, you don't know the resources. Every data that you don't need to touch, you have to move. And you lose data locality out of that. So you're spending time transferring data back and forth. That has both an economic impact and maybe even from a sort of eco-friendly perspective, right, that's wasted power. And so by being data locality aware you can bring back computer computational efficiency that we know from traditional building of systems is important. And so our approach really is about looking at, well, yeah, to be able to scale a function to thousands of instances, the system has to say, hey look, so state is on you, right, because it becomes much easier to spin up a thousand containers without having to worry about consistency, sequentialization, etc. But then you're leaving the burden on programmers. Now, you've given them the supercomputer that's basically a distributed system and said, okay go figure out the rest of the data synchronization model they need to do and that's where, you know, things are lacking. Can we do better?

And at Nimbella I do think we're doing better. One of the approaches we're doing that is with declarative approach. So if you're a function or even a container you can say here's the state I want managed for me. So what that means by being able to declare that it could be for example a file that you need to load because you're doing machine learning inference. So you need to load the machine model that you've pre-trained with your neural network. Every instance of that function doesn't need to load the same file. If you could load the file once, mount it, and share it across multiple functions now, you've saved the cost of maybe a thousand X because you're not doing a thousand times. And moreover, if you're reusing containers, it's already there because it's been hydrated as you said. So that's one example of sort of looking at different kinds of state that can be managed automatically by the system files. And these are things that might be stored in an object store for example, and then mounted as actual files that you can use within their containers. So I like to think of it as state because the system can manage it for you bit more efficiently.

In fact, you can't even do it at user level. I think this was sort of the fundamental primitive for us was, like, If I’m a user and I really want to do this maybe with EFS now we can do some of these as you were touching upon, but I can't … the system doesn't allow me to do it. So unless you have the deep integration within the platform, you don't get that computational efficiency from the copy. Another aspect of state that we also take a declarative approach to managing is sort of transient state. You do things that you would put in ElasticCache or Redis and these could be, for example, session tokens, oAuth flows between Stripe, OCA, or whoever you're using, sort of, identity management with and you need that state, you need to store it somewhere and sometimes you just need a very lightweight database, you know, it's a key-value pair.

So can we just provide that and it's just there for you and these are some of the things we do at Nimbella. So when you write a function in Nimbella and you just want to counter that's persisted and shared across functions, it's just there. You say, you know, for this key bump the counter. You don't know where your residence is located. That's our burden. You don't know how it's backed up. That's our burden. You don't know which geography it's running. The only thing you know that your functions can touch that store with sub-nanosecond, latency sub-millisecond to nanosecond latency, because we've taken on the job of making sure that your compute and your data are co-located and so we're delivering that performance. We're delivering that aspect of state management that is really starting to deliver again on that, sort of, serverless promise, but now it's also stable.

So our approach is really to look at what kinds of things people are doing, and we’ve categorized four of them that we've sort of focused on today. It's static assets when you're building an application that you want to deliver out of the CDN. So that's one. Its files like the machine learning models that we talked about that you might store on object store, but then give you essentially an abstraction that lets you treat them from functions as the file system. And then key-value store. So for transient state, we've left some of the harder problems for our future roadmap, things like databases. Now, I'm talking about RDS and so our goal there is not to build all of these things ourselves. In fact, that's a key aspect of what we're doing. We're not building our own cloud in the sense that we're not managing infrastructure. We're building on top of existing cloud providers that do exceptionally great things, it's just too hard for many developers to penetrate. So we just take the existing clouds of the commodity and build these important layers of abstraction on top of them.

Jeremy: Yeah, and I think it's a it's a good point you make about sort of, I mean, I don't know exactly how you worded this, but this idea of sort of like making it abstract for, or abstracting in a way for developers, right. like making it so it's sort of clear that they don't have to do it. Now, I have concerns about adding state to serverless because I think in some cases people would just use it as a crutch, right, and sometimes when you make things available to people that's when you get serverless WordPress and things like that that maybe you just shouldn't be doing, right. But I do think there are a number of use cases where you do need it but also as sort of, I think, just for me personally, I like to really find a way that I can pass as much information in the event as possible so that you can maintain that statelessness because again the promise of serverless is unlimited scale or at least you know massive scale, right? So now if you're mounting EFS volumes or you're connecting to some file system or something like that, now you've got potentially thousands of concurrent users that are all accessing the same file system and so forth. So there's just a whole bunch of other problems that get introduced there, but I think that's really interesting. So, I don't know, I mean, what are your thoughts on that though? I mean do you adding state is a great thing for a lot of use cases, but at the same time ... I mean, are you still in the camp of you know serverless should be as stateless as possible when possible?

Rodric: Yes, and that's because the history of computing has shown us that's the best way to sort of get computational efficiency and maybe this is important in terms of my background right where we started this call. I've come at this from a programming language and compiler perspective. And so I'm just looking at, “Hey, I can optimize these with a compiler if I had the right abstractions,” and what serverless has allowed us to do is basically be proscriptive, right? When we launched IBM Cloud Functions and when Lambda came out, now Amazon said, write your code like this and people wrote their code like that because the carrot was so big, they just followed the recipe. And, yeah, demanded more features, demanded more capabilities and that came over time but it's given us the opportunity to be proscriptive.

And in some ways because the hardware has come first, the cloud is this massive super computer. It's available. It's allowed us to essentially say, oh we can do distributed programming now with the right program and reliable language abstractions and be proscriptive and people will do it. And so that model that you just touched on has roots for me in sort of actor oriented models where your function is essentially an actor and you can think of the steps of execution, the life cycles, as being broken up into pre-work, so things you do before you actually start running your function, the function itself, and then post work. So opportunities to do things like fetch data from a database or fetch data from a key-value store or file system can be done in the pre-work. And then when I'm done with my functions, hey, serialize all this back out to the right places, can be done in the post work. And what's important about sort of thinking about these three phases of execution is that first the top, the pre, and the post can be completely managed by the system if you can take a declarative approach or other approaches for sure, but at least that's that's how we've come at it.

What that allows you to do from a functions perspective is just have this really clean abstraction that says here's my event, my event contains some state, I don't know where it came from, I don't care where it came from. I can write my code against that event, what's happening before and after is now hands-off and that, you know, ties back to the serverless promise. Just write your code. You can think about your interface, your API, and then everything else is match versus how much of that can we do. I mean, that's how we started, it was like how much of that can we do and this is where you know the genesis for our company which really was. And we found that you know for a number of these kinds of states we can do extremely well and the benefits from the end-user are the abstraction of the function is still pure; you didn't have to break that abstraction. How far can we push it? I mean, this is where it's still early but this is what we're trying to do.

Jeremy: Yeah, so speaking about before and after, another problem in a question that always comes up has to do with function composition, right? And we always think about single-purpose functions. That's the way that we recommend you do things. I mean, this function converts the image, this function processes the record and then it sends it somewhere else, and you can do that with, you know, choreography, right, you can just sort of hope the next system picks up. But state machines have been something that most people embrace when you're trying to connect multiple execution components or if you're composing functions, state machines are really helpful. So what do you have at Nimbella to help with those kinds of workloads?

Rodric: Yeah. So, what we have actually we inherited out of OpenWhisk. Once again, some of the early features we put in OpenWhisk, it wasn't ... we talked about bring your own container, run your container as a function The other was composition. It was built in from the ground up and sort of an intrinsic into the system. What that allows you to do, for example, is take functions and then chain them together so you have a pipeline. And later on, and this was sort of … we did it in a way where the composition itself look like a single function, and what that means is that you can take it and then further compose it so it really goes to, sort of, from a software engineering perspective coming up with libraries of reusable assets that then you can take and treat as malleable code that you can integrate in other components. And later on we went from just composition sort of a sequence to, hey, let me write an arbitrary state machine, a data flow graph, and we have again open source that came out of IBM Research called Composer, which is very similar to Amazon step functions in that it compiled down to a state machine and … but you can code to it against library from Node.js and Python.

And I think Amazon is just starting to do some of that work. We've done it back two years, I think, before. This is one rare area where we had innovated something faster than Amazon. So we're really proud of that work. But I think competition is important because of some of the things you talked about. If you have the ability to focus small pieces of code on specific functionality and then build them into larger applications using the right models, a state machine, you can again bring in this way of sort of saying, “Well, I can re-transform this program, I can recompile it into something that's completely different,” and it's easier to do that when you're starting with small building blocks than taking a large piece of code and then breaking it down into smaller pieces. And when you start small and course it in essentially, you go from fine-grained to monolith … you know, you can still deliver computational efficiency. You can scale things out as much as you can. When you go the other way around you start hampering some of that. Your attack surface also gets bigger, your boot times become longer. There's a number of things that become as hard as parallel programming really still is today. So, I like the model where you start with small fine-grained pieces of code and then course in it, you know, one API for per function. That doesn't mean you have to package the code exactly that way. I mean things like XX solve some of those problems today.

But, so, I like the model starting small, building graphs that essentially state machines. What I think it's fundamental, and Amazon will eventually get there, is, you know, can you take these state machines and further compose them? Right? Can you take ... can you call step function from one step function? Can you call from one workflow can you call another workflow? This touches on something we published out of IBM called “The Serverless Trilemma,” basically can you treat this code as an opaque piece of code that is no different than a serverless function? Code in, event in, event out and what you run inside whether it's a single fucntion or a whole workflow, it is opaque to you as an end user. If you could do that and solve the double billing problem. Basically, you're not waiting for the workflow engine to also run and support just black box code, code that you can't modify as the vendor, then you've satisfied the serverless trilemma, and that's sort of like the Zen of serverless compositions for me. So we build essentially a serverless trilemma satisfying composition with open warrants.

And since then your people have sort of quantified and looked at while step functions does it this way, durable functions from Microsoft do it this other way. I think we're sort of trying to define space around compositions. So what we're doing at Nimbella and sort of inheriting a lot of what we did with OpenWhisk and building on top of that project.

Jeremy: Nice. And that's another point you make about, you know, using these, building these little reusable components, that's another really good argument for statelessness in those components because if you want to reuse a component, but it's always saving to the same database or something like that, it's better to have something that maybe converts that object into whatever format it needs to be, past that then using a state machine to another function that maybe then has the ability to save that state and things like that. So that's really interesting.

So, I want to talk a little bit more about Nimbella because I think an interesting approach that you took was again this whole thing runs on top of Kubernetes as well, right, and you can also run it on-premises so you can basically run it in any cloud or on-prem.

Rodric: That's right. And I think we've done this in two ways. One, as you know, we have a hosted service and if you're an end user who just wants to build against the cloud and you don't care about which cloud you're running on, you don't care about anything but time-to-market, time-to-solution, you know, we're building cloud that's basically very easy to get started and it goes to sort of this notion of building projects, building entire applications that incorporate compute front-end, back-end state and just deploy and it's a repeatable unit of execution. Somebody else could take that code and deploy it and run it. For the on-prem, we're essentially trying to say, “Hey, we can bring that experience for you, wherever you're running your cloud,” and the motivation for us behind doing that is the recognition that a lot of service providers out there are building their own clouds and they need functionality like what serverless functions give them.

They have events, they want to be able to allow their end users to operate on those events. We've seen it with Auth0, Twilio, Salesforce, Zoho now even has a serverless offering, and I think that's just, you know, that repeating pattern that I have events, I have an ecosystem that my developers code against, they want this kind of serverless experience. And so we’re essentially saying we can accelerate your delivery and we can do it in a way where you can run it on any cloud of your choice. It's basically like saying we can bring the Amazon experience for you or whichever cloud you want. That's too big to say because we don't do everything that Amazon does; we do maybe three or four things, but that's the kind of … that's the reason we've sort of approached this whole model of Kubernetes. It allows us to say Kubernetes is almost everywhere. Every organization we talked to has now says, well, we have in-house Kubernetes expertise; we say great, point us at your Kubernetes cluster and, you know, within 30 minutes, we've deployed this entire Nimbella stack for them and they can start coding, building projects, and deploying them and I think that's what's been very powerful for us being able to reach those organizations, help them fill gaps in their portfolio in terms of being able to offer these capabilities. We never expose Kubernetes to the end user; it's an operational aspect. So, it's just a normalizing platform for us. And because it's everywhere, it's allowed us to basically say we can run everywhere.

Jeremy: Awesome. All right. So let's take off your Nimbella hat for a second. Let's put on your analyst hat if you could for me. So, if you look at ... I mean, you said earlier, you know, that sort of Kubernetes is starting to get, you know, sort of become the de facto standard for containers. And I think I agree with you there. There's a lot of surrounding tools that are also maturing and sort of becoming sort of a standard there. One thing we don't really have a standard with, though, is serverless, right, and the way that people are ... and I should take that back and say more serverless functions, right,the way that people are building functions as a service. So, we have Lambda which runs on its own proprietary, it's open source, but Firecracker, you know, sort of to run it as close to the metal as possible. Microsoft Azure, you've got GCP but GCP is also doing not only Google Cloud functions or Google functions, and then they also have their, excuse me, their cloud run and some of these other things, Oracle Fn I think is also like a cloud run type thing. You've got Fargate, you've got edge providers, we've got Cloudflare with their workers, you've got Fastly, you mentioned. All of these other companies like the Salesforces and Adobes and building their own, you know, either running on top of something else or building their own, running their own serverless platforms that are integrated into their system.

So there do not seem to be any standards. There are a lot of different approaches. I know there's the cloud events, you know, Cloud Native ... working whatever it is, the Cloud Native Foundation is trying to do this, like, working group for events and standardize that, which I don't think that has had much movement on it. But just what are your thoughts? I mean, are you concerned with all of these different approaches to serverless?

Rodric: “Concern” isn't the right ... isn't the word I would use, and I think we're in the phase where there is room for a lot of innovation and exploration. They think everybody recognizes there are giant opportunities here. So it's greenfields everywhere. Change the context a little bit and hey, you can go a long way. Taking a pragmatic approach, you know, when we looked at this standards issue, what we said there is a de facto standard, it's Lambda and that's because they process more functions on any given month than any other cloud provider as far as I know. I think the number is in the trillions and I remember, you know, my first conversation with Tim Wagner at a New York City serverless conference five years ago, where he said, I asked him, “How many do you do a day?” He’s like, “2 billion.”

So, it's been exponential growth, you know, over that five years to where they are today. But as you also said earlier, it’s this tiny fraction of all the serverless compute. All the compute that's happening in the cloud today. So we’ve got a long way to go. I think there will be standards or, you know, efforts to standardize will rise and you're sort of seeing it. So Google has this Knative project and as part of that they have been looking at, “Okay, what does the interface look like? Can we standardize it?” And because there's sort of it's got the “K” in the name, right? It's sort of riding on the Kubernetes wave. It has an opportunity to sort of become a standard just like Kubernetes is effectively the de facto standard now for port container orchestration. So I think we need this kind of exploration and I think we're seeing exciting technology being developed because of it and, you know, what's happening at the edge with Fastly and Cloudflare is really exciting. WebAssembly.

And you know the future of isolates where you're running containerless functions, you know from a computational efficiency perspective really excites me. I don't think end-users will eventually care; they'll just care about the interface. So because of that there will be some standardization. As a start-up we can't do that, right; it costs too much and it's prohibitive for us, but it has to come from essentially a consortium of the big players. But everybody has a stake to play today and, you know, … so I don't see it happening any time soon. And if you're a pragmatist you look at who's the biggest whale and it's Lambda and you say okay, they’re a standard and you see it, you know Auth0, the signature is very obvious. Netlify is very obvious. They're all Lambda and so it's winning, you know, without actually being declared a standard. Will that change? Possibly, but I don't think we're going to wait around for it.

Jeremy: Right, right. So, what about you know ... so, with these different approaches to serverless as … I mean, for some people it makes a lot of sense. if I'm an enterprise and maybe I have partial workloads on-prem, I have some things running in the cloud, maybe I want to mix and match, and I've got an operations team that can manage my Kubernetes cluster for me, or can deal with all this stuff. That's a lot different than your small start-up or somebody is just hacking on the side or something like that. So, I mean how much do you think these different approaches to serverless are sort of targeted at maybe the, I guess, the different persona of people who are using it.

Rodric: There's a couple of ways of looking at this one, as you know, from the operations side and the other is from the end user side. Actually, I'll give Knative Project here a shout-out because I think when they came on the scene, they did a really good job of sort of separating the persona dealing with serverless. There's the operator which is managing the infrastructure, the Kubernetes, the VMs, the infrastructure that you’re actually running on, and then there's the end user which is building code and deploying it to this platform. Their concerns are completely different and it's ... I think you have to approach it in different ways from an enterprise perspective. They care about both and you can look at some organizations that have gone all-in on AWS: the LEGO Group, Capital One, there's dozens of them, Vanguard. And they recognize the transformation they can get by essentially just delegating all that infrastructure whirring essentially to AWS.

It's a process and it's a journey so it takes time to get there and if you're a small company, you're a small business, you don't have time for all of that. So your time-to-market is what's most important. You're going to look at which cloud can I build on that will give me the best solution and in some ways the choice that you make really becomes impossible to revert because the more you build, the more you're essentially tying yourself to that platform and if you're extracting value out of it, great. So, for the most part, you know, if we talk to somebody and they say we're on AWS, we love AWS, we’re like, great. We are not the solution for you and we sort of recognize that and we would do the same thing. But there are organizations for various reasons and just because they're running Kubernetes, you know is an indicator that you have to meet them where they are. They have needs where they want to operate their own infrastructure. They want to be able to run the same kind of environment on multiple clouds, data gravity, or other kinds of concerns. You have to be able to give them that serverless experience because I think their developers are going to demand it. So there is a bifurcation there between the operator and the developer that serverless can help serve both in different ways: one on the developer experience and sort of normalizing what you're coding against from an end user perspective. The other is the operator where now you can have smaller teams managing infrastructure because Kubernetes does a lot of the heavy lifting for you.

And you can extract some of that value and now repurpose it, extracting a greater business impact in the long run. And actually if I answered your question was a sort of I took it in a couple of different directions, but I think it's a really interesting area for us and even as a business perspective has implications.

Jeremy: I'm not even sure what my original question was, but let me follow up with this, and I hate to ask you about vendor lock-in because it's just one of those things where, again, when you take a million different approaches to serverless, you pick one and in some cases, that's it. You're sort of stuck with that, you know, at least from a, you know, from a compute and certainly from a managed serverless perspective. I think you know if you pick DynamoDB that's a task to migrate to Mongo or to do something like that, I mean, you pick any database you're pretty much locked-in for it. But I'm just curious, you know, sort of from your perspective, and I know that you know that there’s ... it's more anecdotal but, how important is portability do you think for you know, some of these larger enterprises?

Rodric: I think it's important but maybe not for ... it's important because we've had the conversations with very large organizations and some have said something as simple as this: we can run on any cloud as long as it's this one “name on the list,” right? So their business reasons for why some companies must run on a particular cloud and the sort of lock-in aspect comes about like you said, once you start building it, sometimes it just starts by enterprising developers. One of the early choices I made with OpenWhisk was to use Cloudant which is an IBM CouchDB as a service. And why did I use it? Because it was just there. I needed a database, I didn't want to worry about it. I just used it and I regret that choice to this day and should have used a relational database, but IBM didn’t offer one.

So these choices become really hard to reverse as you grow and takes a lot of investment that essentially then moves. And unless you're a large organization that can afford to spend hundreds of millions, of billions of dollars, that choice for me is almost a non-starter but it's there. People actually are trying to do this and I think whether it's by running Kubernetes clusters on different clouds and then having to normalize it, it's just happening. So I stopped questioning whether, you know, it's valid or not. I just recognize there's a massive opportunity because it’s happening and so if you just accept that and say, “Okay, what are they missing?”

And it's this whole serverless notion because it empowers their developers, it empowers the organization that generates the most value. That's where we focus. Our notion really isn't to look at Kubernetes. It's how far up is back we can go and in some ways because we're small, we're a small company, we're highly focused, we can push up the stack much faster than some of the other companies. And this is what, you know, gives entrepreneurs and, you know, startup founders the opportunity to compete in this space.

Jeremy: Yeah, got that. All right, so I'd love to sort of ask you... and we talked a lot about state and the need for state and maybe there's the need for, you know, better controller mechanisms to spin up or scale up serverless faster, you know, whether that's pre-provisioning or something like, you know, Cloudflare workers are doing where it supposedly at zero millisecond cold starts, right, in there. And of course they can only run I think for, I don't know, 30 seconds or something like ... anyways a very small amount of time. They're 50 milliseconds. That's a very small amount of time those run for. But anyways, so what are some of the unmet promises, let's put it that way, like, besides the state aspect of it. Like what else are we missing from serverless? And what do you think that, you know … there's companies like you have to keep solving for?

Rodric: I think accessibility of the platform. And I remember when I first met you, right, we had this conversation about, we called it “serverless bubble” at the time, right, and maybe “bubble” isn't the right word because bubbles burst and that's not a good thing. Maybe “echo” chamber is better. But I think … one thing I've learned, and I learned this very early on when I left IBM sort of went to a developer conference at, yeah, there's a thing called serverless, the greatest thing, and it was like what's a micro service? Right? Instead of recognizing that the world hasn't yet caught on. There is part of, you know, the technology community that has sort of, you know, good for that. But recognizing that there are still a large interest in Kubernetes, still a large interest didn't EC2 instances in VMs. There's a massive world out there where building applications for the cloud is still hard. You know, just log onto the Amazon console and look at everything you can get. Where do you get started? Right? So the opportunity for us is making the cloud more accessible.

And so we like to think that from a Nimbella perspective, you can create an account within 60 seconds. You can deploy your first project, you know, not even having to install any tools, right out of GitHub. And hey, I have stood up an entire application. It's got a front end. It's got a dedicated domain. It's served from a CDN. My functions are entirely serverless, they scale. I can have state. I just did that, right. So, it's about really making the cloud accessible for a large class of developers from the enterprise, all the way to the indie developer who just has an idea for a mobile app or a website that they want to build. I think this is where really the opportunity is, you know, whether you're running things in a container or an isolate like Cloudflare does. It comes with implementation detail nobody's going to care about in the future. It's what is the programming experience? How fast can you let me create and so at Nimbella, we like to think, you know, create, build, and deploy at the fastest pace of innovation. That's what we really want to try to do. So, and that's what excites me about this like serverless is transformational and even transcendental technology because it can unlock all of that and hopefully you can tell how excited I am just talking about it.

Jeremy: No. No, that's ... I think you make a really good point and I always argue that, you know, serverless when it first sort of came out, right, when we first started building Lambda functions, it was so easy. It was simple. It was a really simple way to think about it and then it just got more and more complex, and more and more complex. And now we're at a point where if you log into the Lambda console on AWS, I mean, it's mind-numbing because where do I even start?

All right. So I think that is a very lofty goal. I totally agree with you. So good luck with all of that stuff. Rodric, thank you for joining me and telling the story of the IBM and OpenWhisk and what you're doing at Nimbella and just giving me that analysis of the serverless sort of market and what the future is because I think it's a really messy place right now and it's got a long way to go. So, the more people we have like you that continue to shape it is great. So if people want to get ahold of you and contact you, how do they do that?

Rodric: So, I'm on Twitter @rabbah, my last name. I'm also easy to find by email: rodric@gmail or rodric@nimbella.com. I think you'll share some of my contact information later on, but you know Twitter is where everything happens today so @rabbah on Twitter and you can find me there.

Jeremy: Awesome. All right, well, again, thank you so much. We'll get all that stuff into the show notes. It was great to have you.

Rodric: Yeah, thanks for having me, and, really, thanks. I really enjoyed it.

View Details

About Ryan Coleman

Ryan Coleman is Vice President of Engineering and Product at Stackery, a serverless platform to design, develop, and deliver modern applications. Ryan is an accomplished product manager and ex-sysadmin who spent the last decade working with enterprise operations teams in the Fortune 100 to automate global infrastructure with Puppet.

  • Stackery: www.stackery.io
  • Twitter: @ryanycoleman

Watch this episode on YouTube: https://youtu.be/tEa2eJLwjZA

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Ryan Coleman. Hey, Ryan, thanks for joining me.

Ryan: Hey, Jeremy, good to see you. I'm looking forward to this chat all week.

Jeremy: So you are the Vice President of Engineering at Stackery so why don’t you take a minute, tell listeners a little bit about your background and what Stackery does.

Ryan: Yeah, so I'm mostly a system-man by trade. I kind of been tinkering with computers most of my life and sort of to pay for my college I started doing IT support that led to more advanced operations roles that led me to some VM automation software called Puppet which led me out here to Portland, Oregon from Pennsylvania to help Puppet grow and be a product manager of professional services and sales and I got to wear a bunch of different hats and more importantly explore enterprise operations that some of the largest organizations in the world and yeah then moved on to Stackery this year to help them with their infrastructure as code platform which focuses on AWS serverless and it’s trying to help people sort of design serverless architectures, express that in AWS SAM infrastructure as code as well as some other languages, and then just provide sort of a workflow for delivering that bringing environments to deliver different changesets over the AWS infrastructure, all that kind of stuff.

Jeremy: Awesome. All right, so there's always debate in the serverless sort of ecosystem or peripheral ecosystems to serverless that talks a lot about this idea of no ops or dramatically reducing your Ops. So I tend to believe that serverless dramatically reduces your ops because there's less things you have to worry about but I don't think that reduces the amount of operations work that can be done. And I think you bring a really interesting perspective because Stackery is a hundred percent focused on building serverless architectures, which is great, but it is for operations teams, right? It's not really for your front end developer. I mean, your front end developer can use it or a developer can use it, but it's very much so focused on the idea of bringing a cohesive operations, I guess, I don't know, sort of like Mantra to a serverless infrastructure.

And with all your experience, especially with Puppet, which again was also like, you know, automating pieces of the infrastructure and like turning ... saying to ops people, “Okay, we don’t need you to install patches anymore. We don't need you to do this because this can be automated.” I think you're going to bring a really interesting perspective. And I'd love to talk to you pretty much about operations for … you know, that serverless is really for operations right, in a sense. I don't know, maybe that makes sense, maybe that doesn't. But maybe we could start by just going deeper into your background at Puppet. Like what were you finding when you were bringing in essentially an automation software to take some of the operational load off of the operations teams?

Ryan: Yeah, that's that's wonderful. I think there's a couple ... there's like two big cultural trends there that I think are worth talking about and I joined Puppet in late 2011. So think like early configuration management movement, early DevOps movement. So everyone was kind of chasing this idea of we can automate configurations on VMs, like it’s very ops-focused, but we can automate everything about the VM provisioning configuration and maintenance process and we're going to reform how we think about these teams. I like to think about how traditional development has this sort of waterfall effect where the business is coming up with why we're doing software, why we're doing IT. Development is getting to decide, well, what are we going to do to solve this business need and operations is that tail of like, well, how do we actually get this in front of customers?

And it usually flowed in that direction and dev ops was a lot of saying well, let's kind of get together into some form of a circle but in classic IT operations, ops was always chasing everything even if they were in the circle. They still had to maintain all the infrastructure over time, they had patch cycles. they had upgrades to do so, they were never really participating in that full loop as much as the business and the developers were and so if ... I came into Puppet as a Professional Services engineer during those two big kind of cultural movements, and I got to go to both public and private trainings, and the public trainings I would be doing 30 people handed off hands-on for three days, right, eight hours a day and it's a mix of lecture and it's a mix of hands-on labs.

And these people generally were operations folks who were told by their business to come and attend. Some were leaning in and interested, others were doing it as they were told and they didn't have a whole lot of exposure to infrastructure as code. They oftentimes were learning version control systems like Git for the first time, right. They're pretty behind development trends on that and they're also really new to this concept of automation. Although the time really everybody was for VM automation. So in the public trainings we kind of had this like mix of characters and in the private trainings I was there to also deploy the software and you would get this sort of room of characters, and that was my favorite time because you had the people who were going to learn Puppet and you had the people around them: the developers, the product owners, like the other representatives of that triad.

And what I found in that was is so often people were hostile towards me, towards the company, towards this idea of automation, and we get this sort of persona who I would see at every one of these trainings, you can find the person who's just sort of back in their seat, arms crossed, really just not thrilled to talk to you and you would start to try to open them up a little bit and they would just be like,” I don't understand why we do this thing. My job is working just fine. This automation is just going to remove my job. Why should I even be here?”

And then they would hear about what Puppet does, and they would see me use it. They would go through a lab of their own and by lunchtime, they were asking questions and by the end of that first day arms weren’t crossed anymore. They're leaning ahead in their chair and they're having conversations with me at the end of the day about like, “Wait a minute. So, all this stuff that I hate about my job, this thing just cranks through it and I'm still the decision-maker about what's going on and I get to control this process through?” Like they thought this thing was just going to be some magical AI that was going to totally eliminate their roles. But in the end, I think what is kind of relevant to your question here, it's about freeing someone up to do the creative work in their job, to make decisions that help the business, that do the work that helps the human brain be its most effective, and all this repetitive work where humans make the most mistakes, where it’s most stressful to cause outages like that stuff is what automation on Puppet was solving for and what I think serverless solves for operations as well.

Jeremy: No, and I think that and I think there's two conversations there as well. I mean you have this idea of automating away certain things that are prone for error, right, so that you know anytime you have to set up a new server and you have to install different, you know, different libraries and you have to make sure the configurations are correct and it's got to connect to these load balancers and things like that using Chef and using Puppet. And using those services to do that and build those out reliably every single time is just something where someone still has to configure it, someone still has to sort of monitor it, someone still has to think about it. But wouldn't it be amazing if you could say is there a better way for us now to maybe, you know, scale now. Maybe can we can work on auto-scaling as opposed to just making sure we launch a new server or maybe we can work on, you know, some other optimization there.

So there's that one piece of it, but then there's the other piece of it which is I think where serverless brings us down further, which is this idea of taking tasks that still require humans, that still require maybe a bit of creativity that management piece of it, but also taking that burden away as well, right? So even like patching and some of those things, not all of that can be automated right? There's still some manual work that might need to be done but essentially outsourcing that to a managed service provider like an AWS, for example, that again reduces other parts of the job that I find those that type of work, that stuff that absolutely has to be done, that can't necessarily be automated, it always gets in the way when you're working on something bigger. Right, like, so let's say I'm working on my CI/CD pipeline and suddenly I realize that there's some vulnerability and we have to go ahead and patch all these servers, we've got to do all these kind of things. That distracts you from working on things that are probably much more important. So I look at this and I say the more that your ops team, you know in quotes, can get rid of the things that aren't adding value to the business, that just gives them so much more time to go and start working on the things that actually do matter.

Ryan: Yeah. Yeah, and I think that's so critical on so many of those conversations I have when that sort of hostile individual started becoming curious and started to become engaged was them talking through all of their responsibilities. They're feeling that pressure from the business to meet a developer need, right, we need this service, maybe it's database cluster, maybe it's, you know, a load-balanced web farm. They then are responsible for the customer experience of that. Do we have enough capacity for the load we're expecting? Are we spending more money than the business should really be spending? How reliable is this service? How do we monitor, trace, observe what's going on when things fail so that the team that’s responsible for that customer experience can respond not blindly to outages, right. Do they understand the architecture? Do they understand how to debug it?

And when we started having conversations about those key roles, “I’m responsible for reliable infrastructure, the customers’ experience that meets the business need and doesn't overwhelm the business in terms of cost.” Those kinds of things aren't covered by automation software, aren't really covered by, like, the core managed service that you’d be consuming in serverless. There are things that those people still need to bring to the business. That's the creative decision making, that's kind of identifying the right tools, connecting these things into a pipeline.

We talked about CI/CD as a stepping stone to saying we're going to give developers a consistent way to deliver software as quickly as they want with the right sort of automated controls to say code has to meet certain criteria before it goes out, we can validate code with automated test suites anytime, and then I don't have to be involved in the software delivery process. I can codify what I know to be successful about that process and give everyone in the team tools to improve it. That's all the same conversation, isn't it? How do we provide software that's going to cover that need that isn't core to the business and free up the humans to spend their time on that core business need and Puppet was doing that like crazy for VM Automation and serverless I think is that next wave of saying how do you bring that extraction even further up instead of paying a vendor for software that automates the VM? What if you just pay the vendor to make the VM go away, right? And that's not applicable for every workload not applicable for every business necessarily, but it's applicable for so many commodity services like say a database cluster where now you don't really need to be managing those anymore. You just need a MySQL interface.

Jeremy: Yeah, and you know, in some of these managed services that you can use ... I mean, one of the biggest things I think that you hand off to a managed service provider is the idea of reliability, right, that uptime, right, you know, the redundancy that has built-in with some of their applications. I mean Lambda runs across multiple availability zones, right? So you never have to worry about a server going down and your Lambda function’s not going to spark up. DynamoDb same idea. A lot of these services do that.

You mentioned billing, right, which I think billing is a hugely important thing when you start building in a public cloud because everything is metered or a lot of things are metered so you need to understand that billing piece of it and that's an interesting place, too, where I think developers and ops people can work together is on the billing side of things because if a developer goes and looks and says, “Okay. Well, this is how I've architected something. This is how I wanted to run, you know using all the serverless infrastructure, whatever, and here's what the costs are going to be. Then, you know, part of the operations team could be to work with them on what that billing is. Look at ways to consolidate things or, you know, optimize them or whatever, but somebody's got to do a tremendous amount of research in order to find out what the best way to use an individual service. And I don't know if that entirely falls on the developer or if a lot of that should fall on the operations team.

Ryan: I think it falls on that triad, right, if you have, say, a product owner. I decide most of my background either in product management or systems administration. And if you're operating a SASS, the product manager is that representative of the business need right? We're offering the software service, we have a margin that we want to care about. They're making decisions about price points. They're making decisions about feature packaging. They should be owning, relatively speaking, here's how much we want to spend to operate this service. Here's how much we want to spend to build the service, right?

A lot of product management is deciding how engineers should spend their time. We want to invest X number of weeks on this feature and all the sort of agile methodology or any other kind of scoring system is really to help say, “Hey do developers have a good sense of how long something will take and does the business want to invest that long for the anticipated return of that feature?” The ops infrastructure should be considered as part of that and if you're consuming managed services that is part of it. And so I think that is less on the developers. I think it's kind of the operations teams and that business owner whoever ... whatever role is being played there to decide whether that cost is enough But I think one of the things that I'm curious for your thoughts on here ... the meter billing and how transparent that billing is. I've seen people have sort of sticker shocks of that and then I start having conversations about well, how much time did you spend kind of building up the service on your own with say EC2 instances or VMs in your own data center and the math starts getting real fuzzy real fast …

Jeremy: Right.

Ryan: And the more you poke into it the more you start to think, well are you even considering how much energy use spent? Your wage, like, good dollars, you're spending to construct the service, let alone the upkeep, are far outweighing the upfront cost of provisioning a service. And when I learned a lot through working with these large Enterprises at Puppet is that it's a real high bar. Like you have to be in a pretty high volume production service before that trade-off starts to become so consequential that you really do want to have a bill versus by versus really optimizing things and that's not to say that cloud is cheap. It's just time isn't cheap either, I think is the point I'm trying to make. I’m curious if that's come up in your experience.

Jeremy: Yeah. No, I mean I think total cost of ownership is a hugely important thing that people need to pay attention to. Right, so it's not just about how long does it take you to build something or, I should take a step back, not how much does it cost you to host something, it's how much does it cost you to host it, but how much did it cost to build it? How much does it cost to maintain it? And then what’s that long-term maintainability look like especially if you have turnover, you know, I mean you need to bring new people in to learn something, learn something that is already there that you wrote custom. It's so much easier to say to somebody, “Oh, hey, we're using XYZ product for this” and be able to find people who are doing that then to say, “Oh, we're using this internal product called, you know Apollo or something like that, something we made up. You know, this is Apollo X and so this is the service that we built and now you gotta come in you gotta learn that and I think that is incredibly expensive.

But going back to the idea about the … also with the sticker shock, this is where I think there's a disconnect maybe between what I think about, and I say “I” but I'm sure there are more people who think this way, what I think about in terms of the reduction of operation costs because you're using a managed service. Now, first of all, you're hiring a world-class team to maintain your DynamoDB database as opposed to hosting your own MongoDB or something like that, if you hand that off now does it cost you more? Yes. It costs more but you don't need people to actually do that. You don't need people to be managing that service for you. So that saves you money from people needing to manage that service.

But the way that I look at it is you don't want to say okay just because we don't need someone to manage this doesn't mean we don't need operations people. You want to take those people who would normally be managing that database or you know, those server clusters or whatever, and use those people to work on some of the other things that matter like we talked about earlier. And I want to get back to something because there are other things that I think we need to figure out what falls on the developer and what falls on the sort of ops person and how much responsibility you want to give to a developer and where there are the opportunities for them to work together. So one of these things I think has to do with observability and we've seen traditionally that Monitoring Solutions was how much CPU is this VM using? What is the memory? How many operations are running at a particular time or whatever?

And the problem with those metrics were those were just to keep the servers up and running and when you don't have to pay attention to those metrics anymore, then what metrics become important and a lot of that comes down to application metrics. But I think that developers have less experience understanding metrics than operations people do, you know, with experience with metrics. So if you can have the shift say what are we looking for not only from an operational standpoint? Because they're still operational metrics, right? We still need to know how many invocations there were, what the latency was, you know, what the error rate is, things like that. But then there's other metrics like how many new signups did we get today? And some of these other things now a lot of that can be built into sort of one, you know observability system that goes back and not only lets you track these business metrics and these application metrics, but also then gives you a window into errors and ways in which you can either debug your applications or speed up or you know, speed up finding where a bug is or something like that. So just what are your thoughts on that where you know ops and dev kind of meat now with this new idea of observability.

Ryan: Yeah, that's such a brilliant point and it is something to me that I found so exciting about the DevOps movement was how these people were by the business being sort of encouraged to work closer together. And I think I found in a lot of those conversations, a lot of enterprises, that everyone got a little confused about what DevOps meant. A lot of people miss the point entirely that it's about culture and about teams collaborating and about how that's all meant to align to a business need and how everyone's playing a role. A lot of people kind of got lost in well, DevOps equals certain tools and Puppet might be one of those tools, you know, my CI system’s one of those tools and that's irrelevant. Similarly, I think the task is kind of irrelevant.

People get focused on well, is monitoring now the purview of the developer because there's less CPU cycles to monitor as you kind of alluded to. I think of this more as what is that human who's kind of fit in that role was developer or operations? What are they really specializing in? What is their sort of core gift to the business and generally speaking. I think that a developer is really skilled at taking gnarly logic problems, thinking in terms of data structures, thinking in terms of how that's going to play out in terms of providing some outcome, whether it's backend and frontend I think is really about taking hard problems of, “I need to make something from whole cloth and I'm going to think about the logic and the data necessary to do that.” Whereas an operations person is thinking more in terms of systems and long-term trends. They're thinking about consequences between, you know, it's almost like a mental Rube Goldberg machine where they're in their head visualizing all the little steps that happen. Oh, well if this ball goes down this ramp that's going to hit this Domino and that's going to cause this other spinwheel to fly and we don't want that to happen too quickly or else the whole chain will break down, right?

There's like this difference in sort of mental models that I've seen so clearly in all of these teams and of course individuals differ, but that to me is the general trend and that's where I go with this sort of observability and monitoring trend is an ops function. It's not solely theirs, right. As you kind of mentioned, the CPU cycle monitoring was really the ops purview because a developer didn't need to care about that, really. They just needed … they should have been having more conversations about what expectations they had about, “Hey, this particular part of the application is going to be way more CPU heavy than memory heavy and then the ops team ideally is having a conversation about, well that may change what sort of EC2 instance profile were applying to this particular part of the application or maybe we're going to split the application between a memory focused VM and a compute focused VM. That is a conversation that should come out of monitoring CPU metrics, but if that no longer exists, the ops team is still thinking about what is the overall portfolio customer traffic that we expect, how much are we willing to spend the kind of overcapacity versus scale? And burst on demand? How does bursting behave on this application architecture? How do I get all the different AWS monitoring options to come to bear on that problem. That should be a collaborative discussion still but I think because of that sort of trend in system view that operations people generally bring it's a great role for them in serverless.

Jeremy: No, I definitely agree and it's funny. I've always looked at developers, sort of the role the developers as again being problem solvers, right? You're solving some sort of problem with data with, you know, with code with logic whatever you're doing and then always looked at operations teams as the people who could then sort of implement and scale the solution to that problem. Right? Like they provide the, you know, the infrastructure for you to do that. Now that line, you know with DevOps, as you know, DevOps really wasn't about changing roles as more about just better communication, which you're right, I think a lot of people still don't understand exactly what we mean by DevOps, but that line is becoming very, very blurry as you get into serverless, right, and people start writing cloud formation maybe but then they start using things like Stackery or they start using SAM or serverless and it's easier and easier for the developers to now create infrastructure and part of creating infrastructure is architecture, right? So now where's that line? How much of that architectured design is on the developer? How much of it is on the operations people and where they are now opportunities for them to collaborate.

Ryan: Yeah. I, this may be a frustrating response, but I don't think the fundamentals changed. Like, let's take throughout example of serverless. Let's say you're building a sort of modified web application that has sort of a front end here and has some back end infrastructure maybe a few APIs and a data layer. A developer and an operations person on the same team-building that stack should be having a conversation about what that architecture looks like. The developer’s going to have some constraints. Maybe they prefer this sort of NoSQL approach to data and they're going to ask for like a Dynamo key-value store, or maybe they need a relational data set. So they want more of a MySQL and Postgres. Now, the ops team is going to be talking about, well, maybe if it's you know, let's say MySQL, should we be building that whole infrastructure from scratch? Well, we're not sure how much business we're going to get on this application. So what if we rent it; what if we're not doing that.

Now the ops team should be thinking about well, what needs to talk to that database service? What are the permissions of that? And this is where like if you're an operations person who came up through say Unix infrastructure, you've gotten you know, or Claw or I guess any infrastructure, Windows or Unix is just changing the sort of flavor of the commands. You are the one thinking about firewall rules. You're the one thinking about file ACLs. You're the one thinking about how, you know, certain operations can happen between devices on the network, right? That's your purview. You have a mind that's primed to think about that and all serverless does is change the shape of the boxes to check, right? Now you're thinking of AWS IAM, you're thinking about the security groups that you're applying, your thinking about how fine with Rain you can be about the database transactions given API server can interact with that database. The developer may be thinking about those things but no, not really, right. They’re thinking more about the database transaction that they need to make on the API and I think the beauty of serverless is the saying because the ops team doesn't have to go away for three months and come up with a new MySQL cluster that meets the business need they can just rent one. It can be an Aurora database cluster from AWS that just you know, how many compute units do I need to reserve for this workload: done.

Now, they're spending their time offering the team a really secure-by-default infrastructure and just to be a little Stackery-biased for a moment, that's been one of our most successful feature sets, is that when you go in and take an Aurora database cluster connected to Secrets Manager or auto-generating the rotating credentials that get stored away in AWS Secrets Manager. You then connect that database service to a Lambda function or a gateway and it's automatically generating the IM role that says only this specific ARM can talk to this specific ARM and then the ops person can come in and go even further and say only these types of transactions. That is not something developers are commonly thinking about. They just want to wire it up and start writing their app and I think that's okay and that's where ops can play a big role especially once they're freed from the sort of big upfront project cost and then the long tail of patching that my SQL cluster

Jeremy: Right, yeah, and actually I just talked to Matt Coulter who created the CDK Patterns site, works at Liberty IT, and his approach to building a lot of these CDK Patterns was to encapsulate a lot of that operational stuff that a developer might not want to think about so, you know, a developer can launch an API gateway with all of the security and all the endpoints secured and you don't have to worry about setting all that up. But someone still has to know what that is and manage those constructs and do some of that stuff which I think is interesting.

So I agree that the developers don't necessarily want to think about some of this stuff. I think a lot of them do, right, and I think that also is dependent upon the size of the organization. I think if you think of any small startup, any good small start-up, you know has that one person, right, who he or she knows how to code, they know how to set up servers, they know how to do all these other things, right, and they can come in and they can do that full thing and that full stack piece of it. I think you see that with serverless as well. But if we go back to the enterprise for a second because this is something where I love the idea of serverless in the sense that it it frees up time to think about one other thing that goes well beyond observability, it goes beyond reliability, it goes beyond billing, it goes beyond all these other things and that's this idea of resiliency.

All right, and I think that is one of those things where when you were building single stack monolithic applications, it was like the server's down the system's down, right? It was everything runs together, you know, if something fails everything fails and we've seen as we went to service-oriented architecture, as we moved into microservices, and things like that this idea of resiliency within distributed systems has become hugely important and I'm sure you're familiar with chaos engineering and all that other stuff that's going on there. So what are your thoughts on that? Where are the opportunities for developers? Because again, you can't just say, oh we're going to just flip a couple of switches and now we're resilient. I mean you have to build things into the code, right, there needs to be parts of your application that can react, that can reroute, that understand things like circuit breakers and some of those other things. So where is some of those opportunities for ops and devs to work together to build more resilient systems?

Ryan: I think that's … to your point, though, first on the sort of unicorn human who does exist in every one of these startups and they do exist in those enterprises. I do think of those people as special but they are generalists really right there. Like they know enough of that full-stack to get going but there's only so much time in a day. So like unless they're bringing so much experience and they've kind of iterated over time and they become really specialized in every one of those categories, there's parts of it that they’re like I've done enough to get it working and I haven't thought through maybe that security angle. Maybe it's like the cost-benefit angle and then you're talking about resiliency. I think that's another space where sure, they can ... you're going to get somebody who's written a front end, they’ve written the API, they've set up the infrastructure layer, but did they also go and write the sort of scale testing suite that runs as part of the CI/CD pipeline to check against regressions when someone's code path changes and suddenly a thousand requests per second breaks the app whereas before the app was responding well.

That ... whether that is the developer/ops, I'm a little less opinionated on which side of that role ‘cause I think we’re starting to talk about are we exercising the code path? Are we exercising infrastructure and how it’s handling that. And serverless I think adds a new wrinkle to that where probably both roles need to get a little more collaborative because are you. You know … let's say you've got a bunch of the patterns that you've got up on your site. One of them's a DLQ pattern, you've got then sort of the data layer pattern and you're causing these requests to kind of flow through this distributed architecture 10 different lambdas, you know a bunch of different APIs, a developer might be really good at helping any of those individual code paths get exercise. They know like yeah, I can exercise this lambda a whole lot. Everything's good. But what if it's the reaction between six different pieces that causes the whole thing to come down and the DLQ gets filled up, like, that is where an operations person I believe brings some perspective like, again, that Rube Goldberg analogy: they’re thinking through more of a distributed architecture consequence and they're not necessarily thinking through, “Oh, hey this library call that you're making in this Node.js function, that's going to cause, you know, a rate limit, right?

And that's where I think both of these teams having a shared conversation about resiliency gets you a fuller picture. And again, you have to codify that, right, you have to put it into either some, you know, one-off test that you're running when you make big changes or ideally part of your delivery pipeline whereas changes go out your stress testing things to meet your expectations or not. And you're doing that in a safe space and then I would just give you one more pitch for serverless … in the VM world so many enterprises using Puppet, they're making such trade-offs for their scale test environment because they only have so many VMs to go around, only so many so much hardware to go around. Or the orchestration of those sort of scale tests on VMs are really complicated. In serverless that's not a problem. It's just dollars. So how much do you care about scale testing? Okay. I'm going to spend those dollars for five minutes as every time I open up a pull request. If I don't care about it that much maybe I do it less frequently. All right, but you’re no longer bottlenecked by where are those VMs for the scale tests going to come from.

Jeremy: Right. No, that's definitely true. So, and I think you're right, by the way, that sort of developers being more responsible for executing the code paths with operations or potentially another hat which could be architects, right? I mean that because again, I think if you think of what happens when that DLQ backs up, there are business rules in place, right? So that partially has to do with the business owner in terms of what do we want to do with backed up DLQ? How much can we do load shedding and some of these other things, what are the ways to do it? There's a lot of roles that need to collaborate, I think, in a modern infrastructure where you need to think through all those things. Definitely take your point on the generalist though. I don't even know sometimes if I started recording these podcasts because I'm too busy doing a million other things. I did record this one though. So that's good.

All right another thing just quickly because you did touch on security. Where does that fall, right, because, again, a lot of this with serverless because of the shared responsibility model and some of these things, a lot of that security falls now to application security. So, you of course have IAM roles and you have permissions and all other things like that that have to happen within the infrastructure and securing the infrastructure and that's definitely I think on the offside of things, but what about applications security and how much does DevOps come into play there or sec DevOps? I guess. You know, where does that responsibility lie?

Ryan: I think that's one of the more interesting questions of this movement, ‘cause the thing that excites me the most about serverless, and this is a little biased because one of the problems Puppet was trying to help solve for its customers as I left was this sort of vulnerability management on VM infrastructure. And that's probably in my mind one of the last miles for sort of operations like VM based operations. If you totally embrace configuration management you no longer have sort of a provisioning configuration problem, no longer an orchestration problem, like you have new projects to apply those sort of techniques to, but it's no longer this sort of thing that's always in your way. And sort of taking care of patch cycles and fixing vulnerabilities on the infrastructure side is still a huge problem in the VM space, right? The vulnerabilities are outpacing patch cycles. The IT sec team is always putting pressure on the IT ops team to get faster. They're staging spreadsheets between each other to say like, well this patch needs to happen on that machine and you ... I don't know if you'd be surprised but I'm still shocked at how much time teams spend exchanging spreadsheets to talk about work that had already been done by an automated process. Right, like that time-spend alone let alone how much more there is to do.

Now, I think one of the things that gets me excited about serverless is saying, well, now we're outsourcing that responsibility and sort of that work to the managed service provider like AWS. Now the ops team clearly has a ton of work they could be doing, should they be taking more of an active role in the sort of application security space. Developers have a ton to do too. So that's where I'm not quite sure where that line goes, but I think an ops team who has been thinking about prioritizing which vulnerabilities to tackle first, they're the ones generally exercising the exploits in conjunction with the IT sec team. The mindset they bring I think is interesting. So maybe that looks more like in the serverless world building up the tooling to help reinforce these things. Maybe they're the ones, you know, installing Sneek or they're kind of building up their own tools and they're making those part of the CI/CD pipeline and the application team is done thinking about their library, their dependencies on their npm making sure that those are constantly cleaned up or maybe the ops team is fitting into that because they no longer have the patch Cycles. I don't think I have sort of a certain answer for you there. But I think it's such an interesting space that as someone with a lot of data on the internet I feel like is a really important question to get answered and certainly both roles have a lot to do there, right?

Jeremy: Right. Yeah, no, it is a challenging space and I think you've got some tools that have been developed as well. I mean, even just putting a WAF in front of API gateways or CloudFront or some of these things to protect against sort of basic application level attacks. And then also again just codifying security into something like Cognito or Lambda Authorizers and using those using those sort of things as a way to lock down endpoints and again, you know, containing the blast radius, all these best practices that you have I think a lot of that falls on the on the operations people and on the architects and then and then as that is given out to those developers, it gives them a little more freedom to make some mistakes not as many as maybe you could in some other places with better, you know network tools.

But anyways, all right, let's go take this little bit further into operations in the serverless world because I kind of said at the beginning, you know, I feel like serverless is more of an operational thing than it is then it is like a development style or anything like that because really what you're doing is you're automating a lot of those operations for you. So what does, you know, what does operations look like in a serverless world I guess.

Ryan: So I this is certainly going to show some of my biases but I think it does come from people who are, you know, coming from a VM automation world and become familiar with infrastructure as code in the power that has to codify the need an operations team has on the infrastructure that meets the development and business needs and then is familiar with sort of the automation tool chain around that, whether it's orchestrating, you know, changes across devices that require rollback strategy and require sort of sequencing of, you know, updates and traffic on the load balancer or it's something more mundane of just provisioning to infrastructure. Those things still apply in serverless and now you're just changing what is the infrastructure as code format, right, instead of Puppet or Chef or Ansible, you're talking AWS SAM or CDK or serverless framework or whatever. And, like, HashiCorps Terraform is big in that space, right? That one's bridging the VM side of the world and the manage side of the world. That role then is still fundamentally, what are the services my development team needs to solve that business problem. Right?

That is writ large, number one the operations role in serverless. I think there's a little bit of a red herring here that we should address that I think is very much the same red herring as we saw in the VM movement, which is developers don't need ops anymore because they can go and self-service stuff, right. As we've been talking about through this conversation, there's so many parts of the responsibility that a developer could go and learn and could spend their time in, but in what trade-off, right. What are they not doing in their development life? So that was the same thing that happened in the VM movement where I would walk into a company that was considering Puppet and they were responding to developers getting purchase cards and going to AWS and renting EC2 instances doing whatever they wanted to those instances, a vulnerability would be exploited in the business panel, right? Classic case, happened that every business I walked into.

Same thing’s happening with serverless. People can just get an AWS account, provision a cloudformation template and they're off to the races. But do they think about the IM roles? Do they think about the scalability and reliability of that infrastructure? Are they maintaining changes and orchestrating those in reliable ways so the customer traffic isn't dropped every time you redeploy that cloudformation template and those are the things that I think ops is still responsible for and there's still opportunities for businesses to say we don't need that role anymore, we’re just going to buy it from Amazon. It's not that simple, right. You're paying Amazon for a lot of that responsibility. But those ops professionals still need to come in and think about the reliability of the service, the cost of that service, and how to secure it and it's just the knobs have changed colors and change sizes. Still the same work.

Jeremy: And I totally agree. I think that idea of, you know, am I automating away my own job, that's the kind of thing where it's like if you are spending time doing the same thing over and over again ... I forget the the calculation of this, this is outside the scope of technology, but essentially it's like when you're training somebody like if it takes you five times as long to train somebody one time as it does for you to do that one job, you should invest that time training them if that's going to be repeatable thing. It's the same thing with automation: that's if you have to do the same thing over and over and over again, even if it takes you five times longer to automate it the one time you know that for that one time, but then every time after that it's going to be taken care of for you. And time is the one thing for any human being that you cannot get more of right unless you want to work 24 hours a day, which I don't think anybody does.

All right. So then I think that's a that makes a ton of sense and I'm totally with you on that stuff. So what about that extra time, like, where can you then put that toward? Like so an ops person in the serverless space? What should you be spending your time on?

Ryan: Well, if you don't mind, I'll take this all the way back to my early career and give you sort of a story to illustrate how I chose to spend that time differently once I started automating. So one of my first sort of rolls was at Penn State, I was an IT administrator for the central IT unit. So, it was a pretty small staff and Penn State's infrastructure is wholly, or at least at the time I worked there, wholly owned by the university so they did facilities all the way up to the software layer, right? So there's a loading dock where Dell boxes came through. I was responsible for wrapping those boxes, giving them power. Now, there's people who are responsible for the actual power and cooling of the facility and these people responsible for the network of the facility, but I had to wrap those machines, connect them, and then provision them, configure them, maintain them over time right that whole full staff operations role and the sort of data center was below the office floor, right?

So we had a hallway and one of those little spiral staircases that went down into the data center and if an infrastructure went down for some physical problem, there's people running down the hall going down the spiral staircase into the data center where they're going to resolve the problem. So that was me showing up in this world, really small team responsible for a statewide infrastructure. So think tens of thousands of faculty, students, and staff, their email, their web infrastructure, their shared file storage, their ability to authenticate across the university, right, so their entire identity and their sort of permission systems. All of that was managed by a really small crew who divided into operating system lines. So there was the IX crew there was the Red Hat Linux crew of which I was a part of and there was a Windows crew, right? It was really a few individuals responsible for these platforms for the whole university.

None of it was automated. And this is a time when universities are stops. They're starting to shift from thinking about IT spend from being this cost center they had to control to be in the thing that was giving them a competitive advantage as a university, right? So universities with a better, you know global learning system or just a better sort of infrastructure. We're attracting better candidates. So there's a lot of pressure on this team as I'm coming in to update our services, to do more for different, you know, different colleges within the university who needed certain things from the central IT staff and then there was this pressure of public cloud, right? All of these two sort of distributed colleges are starting to ask for permission to use public cloud so they could serve their business needs so they can compete nationally. They were dependent on the central IT crew to provide an alternative to going out on their own. So I was asked to kind of maintain a Samba infrastructure. We had the privilege of using IBM's gpfs file system. If you've ever heard of that it's like this fiber channel distributed storage network that goes across the whole university.

So it's like, think of … I’m trying to illustrate here that there's all this plumbing to be responsible for and most of the IT staff really skilled people, really deeply committed to what they were doing for the University all day was firefighting. Service was down, changes needed to happen, patch cycles were behind, every single day, every single week: firefighting. So I started to bring in Puppet to just get out of my mindset of firefighting right now. I'm doing the same thing as everybody else just in my little corner and the more I started automating there, the less I'm firefighting the more I'm getting ahead on some of the service requests. Oh, we need to upgrade the Saba cluster from five, whatever to five this. Okay, I can tackle that now because all of this stuff over here is running stably and if I need to add capacity, I just put another blade into the bladecenter, Puppet provisions it, connects to the load balancer, all good. Right? Right. Like I'm applying those sort of considerations, but now automation is taking care of repeating those considerations.

So we had this person who was responsible for one of the more interesting technologies that I've worked with called Shibboleth, and this is part, it’s like a federated identity system that essentially allows someone from University of Michigan to rent a library book from Penn State, right, like that kind of stuff, federated identity across universities. And the network engineer was responsible for this; the person who's kind of responsible, head of the network and crew, and his process was to kind of … and think of like a really old school system in who does most things by hand, had one of those just gigantic clickety-clack keyboards, really sort of old school, I think it was like one of the early HP operating systems was his platform of choice. He had to take this XML document, ‘cause Shibboleth is shaped in XML, take it from the shared file system and kind of fiddle with it. And so you would see him and he would kind of peck type too, right, so he spent his entire morning peck typing XML structures to fill in this new identity they needed to add and then would get it wrong because XML is hard to do by hand and then just rinse and repeat the whole day. And all he wanted to do is just add this, you know, sort of ARN essentially, that was like identifying this other federated identity that allowed them to go and rent out a library book. And so I, one week spent a project with Puppet where Puppet would take the XML document, put it onto his computer, he would edit, it save it back, Puppet would then make sure it was valid XML, ship it if it was to a new Shibboleth cluster that was nonproduction and give him back a prompt that said like go ahead and do your test validation where he would go and basically try to simulate this federated request.And it took his life.

Like, my life was freed up on this automation. I spent a little bit of time giving him a transformational life where he wasn't spending his whole day editing XML. He was just filling in a thing, saving it, he got to validate it, and then he would get set to say go ahead and ship it and I would send that XML document from the nonproduction cluster to the production cluster. And so that's a bit of a convoluted story for you. But to me, it's like this ripple effect of how everyone firefighting I was managed to free up my space through automation and I could start giving those gains to other people and that to me is the empowerment of automation, whether that's serverless or VM based doesn't matter to me. I was applying operations knowledge to make people's lives better.

Jeremy: Right and I and I love that story because it is so it's like it's like deja vu listening to that. I can think of like thirty other times in my life that that has happened to me or something similar and one of the things that you have when you are so busy firefighting like you said is that you get to the end of the week and you might have a development team that’s like, well, what did you do this week, guy? I don't know. I think we fought fires. I think we fixed this cluster, whatever ... and if you're trying to stay on schedule, and you're trying to release new features, you're trying to, you know, go back and refactor things and add functionality or get things stable again, the more time you spend, you know, just fixing things that are broken without automating them so that they’re long-term fixes is just a complete and utter waste of time. So I think that is brilliant.

All right, before I let you go, I do want to talk to you about the jam stack because I think of, you know, web applications, especially in the public clouds of would … like, I mean, I don't know too many applications now that don't somehow touch the internet in a way, right? Even if it's a private SAS or whatever like there is information flying around the internet. That's how we're building applications. Now the jam stack for people aren't familiar that's you know, static site hosting using JavaScript with APIs and Markup or you know, there's a million different ways to do it. We had an episode that we talked with Guillermo Rauch about about Verselle and what they were doing there and again that was very much so tailored to the front end of it. It's like a front end developer being able to quickly launch something and then have a little API if they need to. But there is a significant amount of depth behind the jam stack that an operations team can take a lot of or can take advantage of so, I'd love to get your perspective. I know Stackery has done some things with the jam stack lately. So what are your thoughts on the jam stack and how that can help with operations?

Ryan: It's well, thanks for bringing that up. I'm pretty excited about about serverless jam stack in particular. And so I'll maybe I'll bridge from that Penn State experience just for a moment. One of the other things I was responsible for there was web hosting for any student faculty staff. Sometimes it will be working, you know for assignments in their class in terms of his own personal portfolios. And that was sort of just basic web infrastructure sitting behind a cash system that was running on an NFS mount for the shared IBM file system that was running University-wide and from the ops perspective, we spent so much time especially because it wasn't yet automated keeping that running and keeping it single patched and just keeping it maintained that we totally missed the boat on that early CMS wave like the Drupals of the world, the early word WordPress movement, and all of these faculty students and staff just started leaving this infrastructure to run their own clusters of this stuff and then they would have these scaling challenges and it would all start kind of coming in this sort of tornado of fire through the University where people were saying like why everyone left your infrastructure for this new stuff. It's not meeting their needs anymore. Will you run it?

And that to me is an example of how operations teams who aren't freed up to have that time miss a business need and it ends up causing way more work. And so I've been thinking about that experience a lot with what Stackery has been doing in the jam stack because when you kind of look into AWS has just one managed service example, it's phenomenally easy to ship the sort of front characteristics, right? They have the cloudfront CDN, which is global, at the edge, in cities, where people are bringing traffic. You just point it at an origin server. Could be your own. Well, let's say it's an S3 bucket. You just kind of deliver your hosted content to this thing and it's everywhere. Right? As someone requested is brought down to the cash, subsequent requests are super fast. There's no there's no need for it.

Now, okay, if you're only serving a static site, maybe it is that sort of quintessential jam stack that has client interactivity through JavaScript, but mostly it's just the static pages. You're great. What happens then in the enterprise and that's where kind of Stackery has found this jam stack sweet spot. There's then, okay, what if I want to run my own data layer? I was listening to a talk to you'll be able to find, maybe I'll drop it for you in the show notes, from Infoworld, where this person at PayPal was talking about their jam stack experience where they were implementing a jam stack architecture for their sort of peer-to-peer payment system. They're not using staff rooms bringing this up as sort of a public example of this. They are delivering the static site, which is essentially the shell for this mobile payment system, right and then dynamically on the client side, they're pulling in the user’s avatar. They're pulling in their balance. They're pulling in, you know, transaction logs and they went from running sort of their own cluster to running the sort of in a more serverless architecture.

And through … especially because that most of the content is those static assets, think of the whole HTML shell plus all the CSS plus all the sort of supporting JavaScript all of that being delivered statically through the CDN means that, like, it pretty much hits instantly on the client then JavaScript is making backend calls to fill in that transaction history. AWS serverless takes care of that infrastructure to make those static assets instantly available, but then you have sort of a build chain problem, which I'll come back to in a moment, and then you also have well, if you're sort of a payment platform like PayPal, you're running a pretty robust data system. You're running many many APIs. You want to be able to express those in some way right? And then you also need to orchestrate the change of those APIs with the change in the frontend for that JavaScript to leverage those new routes or take advantage of new data. And that's where I think Stackery’s been applying this approach where so many of our customers were running the backend and then they were going out to the Verselles and Netlifys of the world to ship the frontend and those pieces were disconnected. And we recently, earlier this year, put out a delivery platform so that any sort of serverless change can go through CI/CD, Stackery’s aware of those stacks, so you open up a pull request will spin up an ephemeral version of that stack that you can run your load balancing scale test against, you can verify and sort of a preview URL to see everything's working out right, and then you could, of course, deliver that, promote that to production where you're you know, updating your cloudfront distribution.

So that merger is pretty interesting as your backends and frontends get more complicated, but I want to take that one step further and then I'll shut up. That CMS thing, right, that pull from the early University, right. If you're building sort of the PayPal eCommerce app, it's one thing to just kind of ship all of that in your bill chain, deliver it to the CDN, and everything's taken care of your backend’s covered too. But what if you wanted to offer a CMS to your teammates, then you're back to managing the ends again. One of the things I've experimented with recently and then we have one of our healthcare customers picking up for their own needs is the ghost CMS platform. It's the sort of open publishing platform that provides sort of a Medium-like experience for editing, but it runs either on your own VMs or runs as a Docker container. Well, how do we make that serverless? The architecture I put together is using Fargate’s PCS or ECS’s Fargate system to run serverless containers and in Stackery’s AWS SAM templates you can specify the definition of that task that gets populated from data about the sort of whole build system around it.

So for instance, the Secrets gets pulled in, the OR database cluster is expressed. So this sort of thick CMS client can be provisioned as part of the same stack that is a Gatsby-driven frontend which is more of what you would think of it as a jam stock and every time you publish new content in ghost or you make a change to your frontend it triggers this build loop where the Gatsby side runs in CodeBuild, another service environment, that pulls data from that ghost infrastructure that then when it's completely built successfully ships off to the S3 bucket, which updates your CloudFront CDN. Right, that whole sort of chain of activities happens and then you can spin down this ghost cluster to the size that you need for your marketing team or other folks who want to work with the CMS and your customer infrastructure is super cheap, super fast, and super secure, right. So to me, it's like it touches on so many things that makes serverless powerful in a way that doesn't sacrifice a complicated backend whether it's something as simple as a ghost CMS or something as advanced as what kind of PayPal has been doing for their payments.

Jeremy: Yeah, no, and I and I'm with you there. I mean, I love the idea of using serverless obviously as the backend for a jam stack site because it gives you so much more flexibility and scalability. Right? But the more you can push to the static side of things obviously the better. I'm still waiting and I know that Verselle was doing this, I think Amplified Console is working on this now, my biggest complaint with generating static site is especially if you're pulling it off like a ghost CMS or something like that, is the fact that you have to rebuild the entire site and if you have very, very large side, so I think if you're running an eCommerce site with thousands and thousands of products and you have to rebuild every single page every time, you know, make a change to one word on the you know on the customer, whatever, customer appreciation page or something like that and everything has to rebuild. So I know some of these systems are working on only rebuilding parts of it, which would be really, really interesting in detecting those changes. I think jam stack’s got a long way to go. I think it's just like right at the beginning of the sort of where we are. But I love this idea of static first. You know what I mean. Static first, serverless second maybe, but I think that's a really, really interesting thing.

So, I think we've covered most of it. Is there anything else that we missed on the jam stack or the operation side of things?

Ryan: I think we mostly covered. I do want to touch on one misnomer about that static app. I think ... to me the important bit there is to say, ship is much statically as you can and especially if you’re wrapping this whole architecture in your build chain, and you're using serverless infrastructure. it's very simple to just say anytime I make a change go and deliver it. You're right that there is one outstanding problem this whole space needs to solve which is that sort of incremental builds instead of rebuilding everything every time. Tying takes care of that to a bit but it is annoying and it's part of a larger you grow the more your lead time for change expands. But people get scared away from the static thing. So I just want to kind of encourage listeners to consider its saying most of the things that would just be generated anyway on demand by a server farm somewhere are instead shipped statically to the CDN, which means that first touch your customer gets on your web application is really, really fast and that matters so much.

There was a talk at Netlify jamstack conf recently a couple of them both from eCommerce vendors and from other folks who were talking about how that first content full paint on a website is really the decider on sales and like people are going to leave you at such a high rate. It really matters to your business. Whether that's a service you're offering or as our eCommerce site where you're selling goods, that matters a lot. And so the static app isn't saying well, you can't have any interactivity on this site. It's just saying take all the bits that are interactive and make sure those are as quick as possible and then fill in the interactive gaps and you have this choice in serverless that I just want to cap on here, which is either obviously JavaScript is running on the client browser or, say, lambda function you interact with on that static HTML shell that really quickly interact with backend infrastructure that return data back to that client real, real fast, right as an alternative to that client side JavaScript. Those kinds of possibilities are really exciting and make it more of an interactive app that happens to be built during the development cycle. And then a lot of it is statically delivered to a CDN.

Jeremy: And you're totally right about that first paint. The first pain is so important. And again, it says bring this all back to resiliency. If you have a eCommerce site and a page loads on an eCommerce site that shows you the products. It shows you the pictures of the product. It shows you the description, maybe even some of the reviews are statically cached and those can be you know, those can be delivered immediately. If for some reason the price doesn't load or maybe the availability maybe that doesn't work because there's a subsystem that's down that that information isn't loading. You've still been able to provide some bit of information to your clients which you know, maybe they say, you know, you give a message. Oh, we can't load the price right now. We can't load the inventory right now. But if they like the product enough because they're able to see it then, you know, there's a chance they might come back and buy it, you know save it to a card something like that, whatever, those are opportunities that you miss if you don't have that resiliency built-in.

Ryan: Yes, spot-on. Awesome.

Jeremy: All right, Ryan. Thank you so much for spending the time with me and all the work that you've been doing over at Stackery. I love what your team is doing over there. I work with Farrah quite a bit for a number of these serverless days things and, whatever, so absolutely awesome work that you are all doing there. So, if people want to get in contact with you, find out more about the stuff you're working on or more about Stackery, how do they do that?

Ryan: You can find me on Twitter. It's @ryanycoleman on Twitter and then of course, we've got stackery.io if you want to see what we're doing for work.

Jeremy: Awesome. All right, and then the blog for Stackery which is called Stacks on Stacks, which I love that name of the blog. And then you have a sample website for serverless jam stack called jamstackery.website, right? That's sort of just a demo site.

Ryan: Yeah. That's right. We put that up after the jam stack conf a few weeks ago as sort of a recap. We enjoyed a lot of these talks and it's a way for me to further exercise this Ghost and Gatsby CMS architecture I was talking about so we are able to edit and write the content in Ghost, but then the site you interact with is a Gatsby generated site that's delivered on Amazon CloudFront, of course delivered by Stackery, but it's a cool architecture that anybody can run in their own AWS accounts and that website just kind of shows it off with some recaps of some cool talks I hope people check out.

Jeremy: Awesome. All right. Well, we will get all that into the show notes. Thanks again, Ryan.

Ryan: Thank you. It's been a treat. Take care.

This episode is sponsored by New Relic and Epsagon.

View Details

About Matt Coulter

Matt Coulter is a Technical Architect at Liberty IT and AWS Community Builder. Matt has a proven history of delivering scalable, serverless solutions on the public cloud, and has crafted CDK Patterns, an open source collection of AWS Serverless architecture patterns built with CDK for developers to use. In addition to his work on CDK Patterns, he shares his passion and knowledge on serverless through events like AWS Community Day Dublin and blog posts on Dev.to.

Website: www.mattcoulter.com
Twitter: twitter.com/NIDeveloper
CDK Day: www.cdkday.com
CDK Patterns: www.cdkpatterns.com

Watch this episode on YouTube: https://youtu.be/wKvaCsvfJ_M

Transcript

Jeremy: Hi, everyone. I’m Jeremy Daly and this is Serverless Chats. Today I’m joined by Matt Coulter. Hey, Matt, thanks for joining me.

Matt: Hey, Jeremy, thanks for having me on today. I’m looking forward to this discussion so much.

Jeremy: Awesome. So you are a technical architect at Liberty IT so why don’t you tell the listeners a little bit about your background and what you do at Liberty IT.

Matt: Sure. So I’m in this account enabling architect role now at Liberty IT. What that really means is Liberty is global, it’s huge, so Liberty IT in Belfast and Dublin has about two hundred engineers in my section, but globally there’s over a thousand engineers. And if you Google Liberty Mutual serverless you can see that we have a mission, we have a mandate, we want to be a serverless first company rapidly delivering value in a well-architected way. So my job is to create the environment where our engineers can do that at a global scale, so not just one thing, but do it in a way where we don’t leave everybody behind and everybody feels bought-in and a part of doing that job.

Jeremy: Right. And if anybody has been paying attention on Twitter or is anywhere near I would say the CDK space they’ve probably come across CDKpatterns.com which is a site that you put together and I want to talk about that because that is super-interesting just in and of itself. But there’s also ... part of the reason you built this, and we’ll get into this, but is because the CDK is so powerful. And I’m going to do a little mea culpa here. At the beginning of 2020, I was looking at the CDK as, like, “I don’t know. I like DSLs better than this idea of imperative code for infrastructures code.” I think I have completely changed my mind on this just because of how powerful CDK is, especially encapsulating functionality for teams. So, I want to talk about the CDK first then we’ll get into CDK patterns. So, let’s start there; let’s start with the CDK and in case people are not familiar with what CDK is, can you explain that and give us some of the vocabulary that new listeners might need to know in order for us to have this conversation so they can follow along.

Matt: Absolutely. So, the first thing is, and I learned this just for CDK Day, CDK itself, which stands for the Cloud Development Kit, is actually not a family of products. So, if you just say “CDK” you’re actually referring to AWS CDK which is the original, the main kit, that is used to deploy resources on the AWS using Python, Java, TypeScript, pretty much they’re working on a lot of languages but there’s also CDK for Terraform which has come up through the community and it’s an officially supported product, CDK for Kubernetes, and I saw CDK for Azure. So, what brings those things together as umbrella products is is a thing called Construct, and Construct is an open source product and that is the magic behind the CDK. That is the thing that allows you to write code in your normal language and it gets converted into DSL that was the original thing that was used in the first place.

So, we’re probably going to spend this conversation talking about AWS CDK and what it does is it converts all of your code into cloud formation and that’s the brilliance of CDK. And pairing developers with the languages they know but at the end of the day you can still apply your rigor and compliance and cloud formation knowledge to the full set-up. And just one more term that might come up later, whenever we talk about Construct, we talk about L1, L2, and L3 constructs. So it’s simple-ish to remember: L1 means cloud formation, everything is L1 starts with c and fn. L2 means AWS built it, so that’s their very light opinion on how to make things easier. And then L3 is stuff that we built, that is typically an aggregation of multiple L2s. I think that’s pretty much the vocabulary you need to understand this discussion anyway.

Jeremy: All right, well, that’s good. It’s a good place for us to start. And I think this is why things have changed my mind because of that L3 category there, right, and the L3 constructs, and it’s because what has happened is rather than you just defining your infrastructure using code, whatever, TypeScript or Java or whatever, rather than you just defining it that way, constructs are multiple pieces of infrastructure that can be wrapped up together especially when you start building these level 3 ones and that allows you to wrap up all your compliance, all of your observability, and all of your metrics, all of your alarms, everything wrapped into one. And so you can do similar things with serverless or with … even with SAM, so why did you choose to do the CDK at Liberty Mutual if it’s kind of possible to do some of these other things, I know you have to copy and paste a lot, but why is the CDK so much more powerful? Why did you choose that at Liberty?

Matt: Yeah, so, I’ll tell you a story. So, a while ago, a few years ago, I had this awesome team. And we were known as a team that built a suite of microservices to support the insurance app, so multi-billion dollar apps were reliant on our services. And that meant we were specialists in Spring Boot and as well as Python for deploying machine learning models and Docker. And so there were five of us on the team, I think, and we were supporting and maintaining roughly thirty microservices. Now this was a high-performing team but I spent way too much of my time talking to our business partners saying we need to do ops. It was a case of, “Okay we need to upgrade this service from Spring Boot 1.x to a higher version or … you know, I was talking to our senior architect at the time who challenged me on a project we called “Deploy with Confidence” and he had said, “I want you to tell me if the system breaks before a customer calls you.” So, it was that point in time I had this brilliant idea of, “You know what, guys? Let’s go serverless.”

And I remember calling the team into the room and I said, “It won’t be that bad; just put in a couple lines of code in the Lambda function and we’re good. We can get rid of all the Spring Boot, we can get rid of all the frameworks, it’ll be easy.” It was not easy. It was challenging to say the least. Because our first piece of code that we put out there was an API gateway and one Lambda function and nothing else. And that Lambda function just made a call to a third-party API. That took us months to get working and that was because whenever you’re deploying code into our cloud in particular, it’s very locked-down. We have these tools in place that if you try and deploy cloud formation that isn’t up to our standards, it just gets deleted, it’s just gone. On top of that you need to know what various AWS components you’re allowed to use. So, as you know, AWS offers options for everything.

Jeremy: Right.

Matt: You need to know which options are right for you. So, we had the use at the time of a private API gateway which I don’t know if very many people use private gateways, but we had these private API gateways with a custom authorizer Lambda and by the time we got that code written there were a thousand lines of cloud formation template. And then the team got in the discussions of, okay, we were a trunk-based development team so if we have four or five developers all working on one cloud formation template, it was just chaos, it was carnage. Because somebody delivers something but it’s not quite ready for production yet. We were not used to that. We had all our Java habits down. So, we started pulling the Lambda functions out of that main template and putting them elsewhere and that’s where we had all these different cloud … we started doing the single function cloud formation template.

Jeremy: Wow, yeah.

Matt: It was at that point we discovered that if you weren’t using aliases you had to redeploy the gateway stack even though the gateway stack was completely separate to the Lambda. It didn’t pick up changes automatically. So, we were going through this evolution of it’s not really cloud formation’s fault, but we were learning all these things about serverless along the way. And by the time we got it all done, we were really proud of what we’d built but as I said Liberty’s huge so the amount of teams who had the same idea as me and thought, “Yes, I want to do that.” And then I wrote this five-part blog series that was I don’t know how many thousands of words but it was long but it was practically a book on how to deploy an API gateway and that was in cloud formation. So, whenever I started looking at the other products like serverless framework, and SAM didn’t exist whenever I was looking at it, but you still had to override parts of the underlying cloud formation to make it compliant in our environment so for me the advantage was that it was easier to stick with the pure cloud formation because I needed to know it anyway so what was the point?

The tipping point for me was whenever I tried CDK I was able to take that same API gateway that caused us so much pain and I made a construct for it that in fourteen lines of code any developer could just literally go “new gateway that’s secured at the endpoint on the function” and that has been deployed thousands of times in the past year alone just because the developer experience was what you would expect. And the beauty of it was the other abstractions for me are they are pure abstractions so it’s quite hard to kick the tires, so to speak, and understand what it’s trying to do. But because CDK is cloud formation I was able to do CDK synth to a cloud formation template itself and because I knew the cloud formation I knew everything it was doing and I knew it was good and I could pipe that into SAM and to pair the two products and start up the API gateway locally. So I was able to create this compelling vision to the engineers to say, “Here’s something like what you had with Spring Boot before, short amount of code, you can start it locally, oh, by the way, you can write infrastructure unit tests for this as well so you can do your CI/CD pipelines.

Jeremy: Yeah.

Matt: That’s why … because it brought all those skills that we already had and it transformed the developer experience into pretty much what we expected in the first place.

Jeremy: I mean, that’s amazing. Just the … when you did that putting all that out there saying we’re going to completely change the way that Liberty IT builds applications. It’s sort of like a major sort of career, like betting on your career in a way, right?

Matt: Yeah. I mean, it’s funny, because as you said yourself at the start of 2020 I went out there loud and proud talking about you know what, I think CDK will work and not only will it work, it will work for serverless. And pretty much everyone looked at me going, what has Matt been drinking today? I don’t think he’s right. But I’ve stuck with it and I think the key has been sticking to open source and talking about this stuff publicly rather than if I’d just done everything in Liberty. We’d be having a conversation where you’d be saying I still don’t know what CDK can do.

Jeremy: Right. And I want to get into the CDK patterns site that you did but this is probably a good place to bring this in and then we can go back to sort of what happened within Liberty. So, you’ve got thousands of engineers, you've hundreds of teams, or hundreds of engineers, it’s a very big company, the Liberty IT piece of it. So, you’ve got all these engineers, you’ve got tons of different teams, so you come up and write these few constructs, you start coming up with these ideas showing people how this can work, but how do you then get … that’s in one or two teams, that’s a small pod within that organization. How do you get an entire company, especially an enterprise to adopt that standard?

Matt: Yeah. So, it helps that Liberty Mutual as a whole is split up into different business segments, so my segment, GRS we call it, Global Risk Solutions, I’m lucky I remembered that, we’re basically large commercial and specialist insurance. But our CIO made a mandate; he put down what our vision is as a company and where we want to go, and he wrote down that we want to be a serverless first company. So whenever you have buy-in at the executive level, it helps a long way. But the second part of it is I haven’t mandated anything to any engineer who works anywhere because I’ve seen an awful lot of times that it doesn’t matter how good your idea is, if you come in and tell people, “I think I know better than you,” they just say no.

So that’s why I started with CDK patterns external, which is, given I haven’t introduced it yet, an open source collection of serverless architecture patterns and the idea was if I could go external and say, “Here is a thing, here is an actual industry thing, here are all the AWS Heroes that talk about the patterns that are in this, here’s the links to all their blogs posts, here are all their articles, here is me talking about it in the world and then go to them and conduct a well-architected review with their team and then instead of mandating it, just ask them, “Okay, I see you’re trying to build this particular solution, have you considered.” And then at that point because the things already exists, it’s already coded and they can pick up on it, I think you’ve reduced the barrier from the direction you want them to go rather than forcing it.

Jeremy: Right. Yeah, and I know, again, brilliant idea to get community support, right, and then you get external pressure kind of pushing down on the organization because that is what the community is using, it’s sort of community accepted. And the other thing that’s great, and why I love open source, is community review. Right? You put something out there and you say, “Here’s this pattern for doing x, y, z,” and you have people coming back saying, “Well, it would be better if you did this this way, or here’s an alternate way to do it” or whatever, then it just makes it so much … it makes everything better. Everyone gets better because of that.

So you mentioned “well-architected” and I know this is a big thing in your organization. But also more broadly in the serverless community where standards, best practices, again, patterns, what are the best patterns to use, what are some of the pitfalls, what are some of the workarounds, that is an ongoing challenge in the serverless world right now. Enforcing those standards and enforcing the compliance or the frameworks, or I guess, the well-architected framework within your organization, why is that such a huge priority for you and how do the CDK patterns help with that?

Matt: Yeah, so well-architected … When I first looked at that … Well, I’ll break it down for people if you haven’t been following. So, there’s the well-architected white paper, which is several pages of just AWS thinks you should be building anything on their cloud. And then there’s the well-architected tool, which is in the console that lets you answer a bunch of questions and it gives you advice on where they think you should go based on your workload. But they’ve still been refining it for others. So after that, AWS released a bunch of specific lenses and one of those lenses is the serverless lens of the well-architected framework and that’s a bunch of serverless-focused questions about your architecture based on the pillars of the well-architected framework. And the reason why it’s been so transformative for us is because, again, we haven’t introduced it as a pass or fail function. It’s a mechanism we use to have a conversation with our engineers.

So, we do these things, I’ll say we do it once every three months with it with a theme for the purpose if sitting down and saying, “This is the spec for what AWS thinks, again, not my opinion, this is not the Matt Coulter opinion of what you should be building. This is Heitor Lessa and the brilliant minds at AWS have said this is the serverless lens, so can you tell me … I can see that your solution will scale, but do you want it to scale that much? You know, are you going to cause a problem elsewhere or how do you know if a piece is broken and I just, I think the fact that that’s there and it’s broken down based on those pillars and it’s not my opinion has helped us massively just have a base understanding across the order that it doesn’t matter what business unit you’re in or even which technology you;re coding in, this is universal. It’s been a massive help in that way.

Jeremy: Yeah. And so that’s one of the things I think that, again, going back to the CDK is why I’ve changed my mind on it so much is that you can use these constructs to encapsulate some of these best practices into it. I mean in terms of the actual pillars, you know like just being able to have observability and some of these other things, security and all that stuff, and then again being such a large enterprise, I’m sure there are a lot of lawyers that work at Liberty IT and Liberty Mutual to make sure that all these things get passed and all these things follow proper compliance and then it’s just every other major compliance thing that’s out there, whether it’s PCI or SOC 2 or whatever those things are, all of that stuff, any bit of it that can be encapsulated into these constructs, is just there. I think that’s amazing. Okay, so let’s move on a little bit and get more specific about how you implement the CDK because clearly you’ve got the CDK patterns on the outside, you know, sort of that pressure coming in, I know you use those patterns internally as well. But what about in a large organization, I think about something like dependency management, right, so how do you handle, I mean, you must have shared components across teams and things like that, so how do you do dependency management in CDK at Liberty IT?

Matt: Yeah, so it’s something I haven’t touched upon yet. So, CDK patterns as I’ve said, is the external facing, open source collection of patterns and you use them internally. Well, that is true. We use them internally in the sense that there’s a Liberty Mutual tailored version of every pattern but we also have a tool called the software accelerator. And what that is is essentially, “click, click, click” new pipeline set up, new code base set up, pattern deployed, so it’s a rapid tool for developers getting up to speed. The reason why that’s relevant to this question is because it means say that gateway that I just mentioned, that is a custom construct that we have in npm but the accelerator it just pulls you in a version of … it just pulls in, say version 3, but it’s still npm and you can update it any time. So we handle our dependency management based on standard practices based on your releases so, we’ve been looking at it today as … there’s a big discussion about this in the community for CDK with should you release a new version of your construct with every new version of CDK or should you build it in a way where you say any version above this point is good with this construct.

There’s pros and cons for both ways. The reason why you release it with every version is because you can guarantee you’ve tested it, it works, you’ve pulled in the latest version of the dependencies. But the only downside is you need to release a new version for every version of CDK which happens maybe twice a week, so if you don’t have that automated it’s a lot of work. The other way is you could say my construct works with anything above version, say 1.30, and that’ll work for the most part until you get a breaking API change and then you have to decide am I going to have to release a new version and communicate to people to work about this version you will need to update and the shorter end of that method of the consumer tells you that the construct’s broken before you know yourself unless you’re testing every version anyway. So that’s why we use locked versions and we do updates with every version of CDK entirely but that works well because, as I said, everybody is working off those base patterns so it’s not like there’s 5,000 different things to update whenever the API gateway construct gets a new version, it’s just a case we have a themes channel, just update the themes channel and everybody knows a new version is out there and they can update.

Jeremy: Now, not to get too deep into the weeds, but that’s changing with v2 of the CDK, right?

Matt: Yeah. So, at the minute, originally whenever CDK was launched, they thought it would be good for all those L2s that AWS made, the opinionated constructs the very light opinionated ones, they released them all independently, but the problem was that they all need to be on the same version of CDK today. So, say you start building your project and it’s 1.70 and then you decide to add a new dependency but you just do npm install and 1.71 has been released since then, well, the new thing you just pulled in would get 1.71 and you would get a weird … it doesn’t actually say your dependencies mismatch, you get a weird TypeScript error and the way it affects it is to make sure that all your dependencies are in the same version. But what they’re doing with v2, which I don’t know the exact release date yet, but we keep saying it will be a few months out, is they’re model CDK, all those AWS L2s are going to be bundled inside CDK’s core module. So it means that for third-party ones you might have that issue slightly but all the AWS ones at least will be on the same version because they’re bundled together.

Jeremy: Awesome. Okay now, so what about testing. So you mentioned that you can do some automated testing and sort of build that into the CI/CD to test your cloud formation. So how do you implement testing on the CDK stuff?

Matt: Yeah, so this is something I particularly like about CDK. It’s better in the TypeScript version than in the other languages, unfortunately. There are plans as far as I’m aware to deport the testing for the other languages, it just hasn’t been done yet. But if you look at all the patterns on the CDKpatterns.com they all have tests. And there’s multiple different levels of tests you can do. So at the most basic level you can do a snapshot which is to say, “This cloud formation should not change and if it does change my pipeline should break; I don’t want this to deploy.” And for me, that’s the kind of test you put in after your application is stable. You’re not adding any new features but you’re upgrading CDK versions and you just want to make sure nothing breaks. On top of that, you can actually do some more complicated granular tests. So you can go in to write a unit test to say “I expect this cloud formation to have a resource like a Lambda function with this particular handler. And I tend to write a suite of tests all based around that so that that way I can at least know without having to say “this whole cloud formation is the same,” the very important bits on a unit test run I can check that they’re all there. I think that’s awesome.

Jeremy: That’s, again, I mean, just writing tests … that’s one of the things that is so tough with you know the DSLs if you just have cloud formation running the tests on just cloud formation is not very easy to do. So, if you have constructs that are generating then you can run that, that’s a very good approach to doing that.

Matt: We’ve all been there with cloud formation where, as I said, there was a team of four or five of us and somebody would change something that they thought was minor and all of a sudden the cloud formation doesn’t deploy and then we’re all sitting around going “what changed?” So, yeah, that’s why I love the unit test because you know the unit tests are on the cloud formation itself so they’re not on the Java code or the TypeScript code, they’re on the cloud formation. So you know if they pass at least you’ve a valid cloud formation template that should theoretically deploy which is nine tenths of the battle.

Jeremy: Right, right, absolutely. All right so let’s move on to the CDK pattern site. First of all, let’s take a step back. You explained it: it’s a collection of patterns and constructs for CDK that basically outline a number of things. What are there, twenty-three patterns now?

Matt: Twenty-three today, yeah.

Jeremy: Twenty-three patterns. So all these really great patterns, everything from webhooks to cloud formation to Alexa skills. All these other things. So, we kind of got the background of why you build it but really what was really the biggest motivation, what was the trigger point that said I’m going to go out there and spend all of my free time here building this amazing site that people can use.

Matt: Yeah, so I got all this CDK, let’s say it was probably last July or August, was whenever I really got into CDK and at that point I did not think I was going to start going to open source for coding. But it was about the start of November I launched that API gateway pattern internally. I was so confident. I was like, “this is it, we’re done—we’ve API gateway and Lambda, we are good.” But I got to go to re:Invent last year, I was so lucky to get one of the tickets. And I realized that most of the problem was actually just the sheer amount of options for developers and what they can do. You know, it doesn’t matter who you follow, if it’s yourself or any of the advocates or who, there’s a million and one different patterns and opinions for how you can build these things so I was sitting there and I thought how can I help this situation, because as I mentioned earlier, you get a pattern, say I took the simple web service from yourself which is just API gateway Lambda DynomoDb but if I try to deploy that internally I personally know all of the extra steps I have to go to deploy that at Liberty Mutual, but how could I help all the other engineers know that.

So that’s why I had this moment of decision where I could either do it internally or I could go externally and go open source. And there’s a tweet out there from I think January where I said, “Sod it. Be the change you want to see in the world.” I bought CDKpatterns.com and let’s get this created. And ever since then, as you say, I’ve had no free time, pretty much every free second just reading and coding. But the reason why I’ve kept doing it is because the first couple of patterns launched and I said I was doing it because I wanted the other engineers to know what it is, how they can deploy things, and what they need to configure and it really internally did hit the mark and massively helped. So never mind the fact that externally people have picked up on it and it has made my life so much easier as an architect being able to use it as the base conversation, even the likes of this where we know we have the same language to have these conversations.

Jeremy: Yeah, that’s amazing. So let’s get into CDK Patterns itself. So, first of all, CDKpatterns.com. Go and check it out. It’s not just a list of patterns, it is organized around the well-architected framework as well, the serverless lens of that. So just explain the organization so people get that.

Matt: Yup. So there’s a few different ways you can find patterns on the site. Originally it was just a GitHub repo, and I’ve sort of been agiling my way to a product here. So the first way is you can just view all the patterns. You can just go in and click “view all” and you can just see all the pictures and scroll through them. Or you can go in and view them through the serverless component used. So you could say I need a pattern for CloudFront , go ahead and pick CloudFront and it will filter them by the patterns that use CloudFront. But the most interesting way I’ve been trying to do it is, I talked to Heitor Lessa and said, “I want to try to introduce well-architected to CDK patterns” and originally I had this really like hit thing, but in talking to him it became a case of saying, if I’m conducting a well-architected review and somebody goes through that and they say question two is your problem, how do they know the answer to that question is the particular pattern. So that’s why if you go to the well-architected section of CDKpatterns.com it’s broken down by each question in the well-architected serverless lens and then each question doesn’t have one answer it has multiple best practices. So underneath each question there’s a best practice and if I have a pattern that helps with that best practice it’s right there, if I don’t it links to all the AWS docs so you should at least be able to find something.

Jeremy: Awesome. Let’s talk about some of these patterns. So like we said, there are twenty-three of them right now. So what are some of your favorite ones?

Matt: Yeah. The one that took the most effort, which I don’t know if people realize looking at it, was the CloudWatch dashboard pattern.

Jeremy: Right. This is a good pattern.

Matt: That took me probably the longest of all the patterns, I’d say three and a half weeks of just reading theory. So it’s not so much the implementation but what to implement was the problem there, and that goes back to it’s taking the simple web service that I mentioned earlier, API gateway Lambda DynomoDb but if you want to build a CloudWatch dashboard for that what are the right graphs, what are the right metrics, what are the right alerts that you’d want the test to tell you. So I had to get into metric max and you know, try to work out what do I want to alert on and how frequently. So we’ve used that pattern now internally and I personally it’s one of my favorites because it seems like it’s low-hanging fruit but there’s actually a lot more behind it once you scratch the surface.

Jeremy: Right. And I mean, there was just a tall the other day at Serverless Days Virtual from the team at Lego and one of their audit processes was figuring out the observability piece and what metrics they wanted to track and what they wanted to alert on and make sure that everything was with that. That’s a really interesting benefit of the CDK is to say, “Look, here’s the pattern that I’m launching, the connection of components that I’m launching,” but then everything else that’s around that. Just the CloudWatch metrics alone are huge. So, again, super-cool pattern for that particular implementation but a really good framework to use if you’re building out your own constructs for your own company in terms of what other patterns you might need there or what other alerts and metrics surrounding that.

Matt: Yeah, I mean, if you jump into the pattern you’ll see the list of external references are quite long in that pattern and that’s just because there’s a lot of opinions about what the right metrics are but I will say you are going to have to write your own metrics just because the ones out of the box they’re not really good enough today to just use the pure metrics. That’s why you have to get into metric max and fully understand why you’re alerting on what it is. So the pattern breaks it down the metrics for the gateway itself, the metrics form the Lambda function and the metrics for DynamoDb. So if you’ve any of those three it’s worth checking it out and seeing if there’s anything that can help you.

Jeremy: Right. Yeah, and you doing all that research for us is very, very helpful. All right, so another pattern on there that’s one of my favorites, and this is newer to the CDK patterns site, is the Lambda circuit breaker and you based this pattern off of something Gunner Grosch just released. He released a little script, a little npn package, to allow you to implement Lambda circuit breakers, which I’ve been talking about for, I don’t know, two and half years now or something like that. So, I love that this now is encapsulated in the CDK.

Matt: Yeah, it’s awesome. Whenever Gunner released that I sent him a DM, I was just “Okay, do you mind if I throw the CDK pattern for this,” because internally, you have no idea how many conversations I’ve had internally, about needing circuit breakers. I mean, it should be core at the point but it’s not so that’s why whenever I saw that library I thought, you know what, this is something that it’s not a huge effort for me to put out there but it’s of huge value to the community to be able to say if you’re going to integrate with something that may or may not be reliable, let’s give you a mechanism to decide what to do whenever it’s not at its reliable phase. It seems basic but it’s not.

Jeremy: Right. Yeah, and I still don’t understand why this is not built in to AWS or any of these cloud providers because if people are not familiar with the circuit breakers essentially what it is is when you are reaching out to third-party APIs or anything that, again, you could overwhelm, or could go down, you’re building in a mechanism here that says once I start getting a certain amount of failures I’m either going to back off or I’m going to take a break for a minute and then I will keep trying to see if it’s working but I’ll be that good netizen, if you want to use that word, and not overwhelm these downstream services because the worst this you can do is for a service that is responding slowly is send more traffic to it. So, yes, these are the things that it’s such a common thing. You’re always calling third-party APIs, you’re either writing data, you’re reading data. It will be interesting to see what they do with EventBridge and SAS integration and some of those things around that, but until then, Lambda circuit breakers, check them out, awesome pattern.

All right, what's another pattern that you like?

Matt: Another one I like, they’re all based around well-architected, but the Saga Step function. The Saga pattern, this has been around a lot longer than serverless, this is just a design pattern where for … it’s about managing these big long distributed processes of being able to say what happens when there’s failure when you’re midway through the process. For the example in this pattern, it’s a holiday booking, it’s a case of booking your flights and your hotel and I think your car maybe, I can’t remember if they added the car or not, or if it was one layer too many. But the idea is for every step that you take you need to have an equal but opposite undo step, so that if anything goes wrong it unravels layer by layer, so the system is in the state it was before you tried anything. It links to an awesome talk from GOTO by Caitie McCaffrey. And the Saga pattern itself it’s … if you want to build systems and you’re working across these multiple bounded contexts as opposed to one single unit, it’s something you really need to consider.

Jeremy: Yeah. Absolutely. And the Saga pattern is something, too, from a serverless standpoint, I know that Yan Cui, the Burning Monk, had put some stuff about that in the past, which is just … it’s just a really, really great pattern. If you’re coordinating multiple things, I mean. I’m a big fan of choreography, like using EventBridge and having other systems react. And in most cases that’s fine if someone is going to get a mass mailer or some sort of, maybe a coupon is going to go out if someone makes a certain number of purchases, if that fails it’s probably not the end of the world if that didn’t work or updated in your Salesforce or Marketo. But if the inventory system wasn’t updated, well then that’s a problem. We need to make sure those things are taken care of, so absolutely great patterns.

All right, one more pattern I want to call out. In the beginning of 202 I was thinking that voice control was going to be a hugely important thing. With all of these, I’m going to say … I’m going to mute … hold on, I’m going to mute my Alexa so it doesn’t respond to me. But with all of these voice controlled systems and all this voice interactivity, I think the ability for you to now do really complex tasks with your voice is pretty cool. I didn’t see as much this year as I thought I would. I thought there would be more of a progression. But maybe that was because everybody was locked at home and maybe there wasn’t enough opportunity for these things to grow. But I do think it’s a very cool thing and I think Alexa, obviously, is one of the pieces that’s driving this. You just recently released an Alexa skill, so tell us about that.

Matt: Yeah, what’s cool about this pattern is that I didn’t write it, so this was actually a community contribution from another member of Liberty Mutual’s staff, an IO community builder, Chris Plankey, but what’s awesome about this is we also have two Heroes at Liberty Mutual, Jillian Armstrong and Jillian McCann, both experts in machine learning, lex, and voice. So, since I’ve started CDK Patterns thinking how can I get voice e.i. into these patterns because it is something, it’s going to come at some point and be mainstream. The problem with lex in particular it’s not exactly in cloud formation today. You have to deploy your own custom resource to get it out there. But Alexa was apparently slightly easier, so that’s where Chris Plankey was able to take a very basic skill and you still have to navigate the fact that there is Amazon and there is Amazon Web Services so you need an Amazon developer to deploy it, but outside of that the barrier to deploying it is get your Amazon account, clone pattern deploy, and you’ve got Alexa skill which what it does is just let you say, “CDK Patterns tell me what patterns you have.” But you can adapt that to be whatever you want in your context.

Jeremy: Right, yeah, and I’ve developed a couple of Alexa skills more as a test, I’ve never put one out there for other people to use, but they are very cool. They do get kind of complicated, and you need to … and this is another thing, CDK Patterns isn’t going to help you understand all of the semantics and how you have to create the different questions and all that kind of stuff and all the slots, but it’s great for getting the basics set up and, like you said, just being able to iterate on that and make some changes.

All right, awesome. So, again, CDKpatterns.com, go check it out. There are twenty-three patterns there that will get you started immediately and probably get you to really like the CDK. So, do that. Speaking of people who really like the CDK, another thing that happened earlier this year which is crazy how fast this all happened was the CDK Day. So tell us about the CDK Day.

Matt: Yeah, so this was one day I just had this idea and I thought, “What if we just had one day, we just took the whole day, and just talked about everything CDK.” So I sent a message to a few of my friends and from there it just snowballed. I sort of thought everyone would say no, this is too much effort, but everybody was so supportive. And it was two weeks or four weeks after that message we actually put out the tweet to say CDK Day was happening and then it was twelve weeks after that that the day happened and in that time we had to build the website, get the CFP out there, pick the talks, find a host, get the screening platform up and running, do all the social media. But the community absolutely came through, I mean the numbers of people that signed up was incredible. At the very start for the keynote there was over a thousand people watching for a conference that was brand new and all of the speakers all nailed their particular talks. I couldn’t have asked for more considering it was just a random idea I had one day that snowballed into a large thing.

Jeremy: That’s amazing. And the videos of the talks are up on the CDKday.com, right, so you can go back and watch the replays?

Matt: Yeah, they’re all up on YouTube and you can find them. There’s a rewind page on CDKday.com where they’re all individually up there because something I’ve learned since I was a 5R, back-to-back-to-back of talks. To make it all easier for people to consume, I’ve sliced them all up and reuploaded them so yeah, you can find them all there broken down by speaker name and title.

Jeremy: Awesome. So what are some of the best talks that people should go check out?
Matt: One that is really awesome that’s still in beta or def preview is “CDK Pipelines.” This is the idea that with CDK not only can you use your language to deploy your infrastructure but you can use it for multi-region, multi-account set-ups, so everything in the same code. So Thorsten Hoeger, who is a Hero, he talks through the actual live demo of how to do that with CDK pipeline. It’s a really good talk if you haven’t seen CDK pipelines before, you’ll learn a lot.

There’s a couple of other good ones as well. So, Nader Dabit actually gives a talk on AppSync with CDK. I’m a big fan of trying to get Amplify and CDK to be friends. I’m trying to get this rapid development closer together, but he does a live demo of AppSync which is really good. And then the last one was Elad, who came up with the idea for CDK. He showed off his new project which is called projen. This is like an opinionated way of creating new projects. So you can just go in be like, projen, use CDK module, and it’s like … it abstracts certain things like all of the different dependency versions I told you, it abstracts them into a yaml for you, and you don’t have to manage that anymore. You can just give it your GitHub token and it will keep all your projects up to date and stuff. So it’s a really cool concept.

Jeremy: Awesome. All right. So listen, CDKday.com, CDKpatterens.com. So, Matt, thank you so much for, one, joining me here today, but also for literally the tons of work you have done. I don’t think you quite realize how much good you have done, not just within Liberty Mutual but in the community itself. I mean, there are … the CDK Day, everything you have been doing is amazing. Keep it up and … yeah, just awesome. If people want to find out more about you and all these other projects you’re working on and maybe what you’re taking so that you don’t need to sleep so you can keep working on all of these things, how do they do that?

Matt: Yeah, thank you for that. If you want to find me I’m @nideveloper on Twitter. To be fair, if you type in nideveloper into Google, you’ll find a ton of resources that I’m at. But if you want to keep up to date on the patterns, there’s a CDK patterns Twitter handle. If you want to stay up to date on CDK Day, in case we throw another one, there’s a CDK Day Twitter handle that you can follow. If you want blog posts, you can go to dev.2/nideveloper. That has your bases covered. If you can find me in those places, you can find the others.

Jeremy: Awesome. Well, I will put all of this stuff in the show notes, LinkedIn, the blog, your Twitter, CDK Patterns, CDK Day. Matt, thanks again, really appreciate it.

Matt: Thank you for having me.

This episode is sponsored by TriggerMesh and Amazon Web Services.

View Details

About Taavi Rehemagi

Taavi Rehemägi is the Co-Founder & CEO of Dashbird, a serverless monitoring and intelligence platform for building and operating complex applications on AWS environment. He has over 13 years of experience as a software developer and 5+ years of advocating for the serverless revolution and building Serverless applications at various organizations himself.

  • Twitter: https://twitter.com/rehemagi
  • Dashbird: https://dashbird.io/

Watch this episode on YouTube: https://youtu.be/xeF19VCuoV0

Transcript

Jeremy: Hi everyone. I'm Jeremy Daily and this is Serverless Chats. Today, I'm chatting with Taavi Rehemägi. Hey Taavi, thanks for joining me.

Taavi: Hey, thank you, Jeremy. Nice to be here.

Jeremy: So you are the CEO and co-founder at Dashbird. So why don't you tell the listeners a little bit about your background and what Dashbird does.

Taavi: Sure. I've been a developer myself for pretty much my entire life. I started coding when I was 14 and since then, before starting Dashbird, I was an employee in two different startups. The last one I was working a lot on serverless. That was in 2016/'17, which led me and some of the team at Dashbird to found this company called Dashbird. We're an operations platform for serverless workloads. We help companies who are building on serverless to achieve excellence with their infrastructures.

Jeremy: Awesome. So we have done a number of shows about observability because observability and serverless seems to be that third-party offshoot that has been missing. There's a lot of things that AWS just didn't really tackle initially with a lot of the observability stuff. Now, they've added quite a few things, but again, it's nowhere near as easy to use as some of these third-party tools like the Dashbird are. So there are obviously constant enhancements.

They just launched, and we can get into this in a little bit more detail, but they just launched not too long ago, this idea of the extensions API for Lambda, which allows tools like the Dashbird or whatever, to have more control over the life cycle, if you wanted to have control over the life cycle of the Lambda function being able to get metrics and telemetry data and things like that. But I think there's still a bunch of stuff missing. I think you would agree with me on this, that there's more we have to do in order to understand and observe our serverless applications. So I'd love to get your input because I think Dashbird has sort of a different outlook or I guess a different roadmap for how you want to address the observability problems, and it's super interesting. So why don't we start there? What's missing in your opinion with observability and serverless?

Taavi: Sure. So I think first off observability is one thing we do, but when it comes to operating the serverless infrastructure, we're talking about high load like ad scale environments, there's a lot going on there that we try to help companies with. As an engineering team, if you're really building something that has hundreds or thousands of functions, for example, and a lot of different Cloud resources, then the one thing that's really difficult obviously is monitoring data and getting an overview of the activity going on across those resources and across your infrastructure.

But there's also, how do you detect failures and how do you get notified quickly and how do you respond to incidents and solve them? There's also keeping up with things like security and Cloud infrastructure for best practices, optimizing for performance and costs. So the monitoring this one part of the puzzle and then having been in this role where we were building a pretty substantial serverless infrastructure, there's a lot going on there. A lot of those things as a team you would have to build yourself and to figure out yourself and to construct strategies around how to improve. So that's really what we're trying to do for our organization. So we're trying to build an abstraction level for operational practices pretty much.

Jeremy: I love that because it's a more sort of holistic approach, I guess, to building a serverless. So building and managing a serverless application, as opposed to just sort of being responsible for, I guess, the monitoring aspect of it. Because again operational-wise ... and this is something, I forget who I was talking about this to, but essentially where it's like serverless or monitoring and observability in serverless is great when you get an alert that says something went wrong. But it's also really good and comforting to know that something went right.

Right? To know that events are flowing through the system and that the SQS queues are processing correctly and knowing that those things are working correctly and give you that level of confidence. I think that's really cool. From the Dashbird perspective, and again, I want to keep this a little bit more general. We don't want to just decide all about the Dashbird, but I really do love this perspective that you have. What is the vision in terms of being able to manage, not just the monitoring piece of it, but also the operational piece and implementing those best practices? How do you look forward or how do you plan a product that does that?

Taavi: When we started working on Dashbird, obviously we didn't come up with this vision in the first iteration. At first, we were just building a tool to monitor Lambda functions pretty much. What that came up early on was hundreds of people or companies who are actually struggling with this. And after all of those conversations, I think we kind of constructed this hypothesis around what this platform should look like for those teams that were the early adopters. So what Dashbird is today and what we're building it to be is this platform, you can look at it in three different pillars and I can go into those pillars if it makes sense?

Jeremy: Yeah, let's do that.

Taavi: Sure. So the first pillar that we have is a data centralization pillar. So what we do is we connect your AWS account without any code instrumentation. We don't use Lambda extensions or layers or instruments to code at all. Instead, what we do is we discover the entire Cloud infrastructure that you have and start ingesting all different types of monitoring data for those resources. So that includes things like log data, metric data, tracing data, configuration data, and really everything that the system is putting out externally. And from that extent of data, we're trying to understand the state of the infrastructure and to make that data available to the engineering teams, to be able to search and query and to interrogate that data in all different ways. So basically the first operating is to get everything in one place to break down the silos between logs and metrics and traces, and to be able to look at services and activity across different services and different resources. So that's the first thing that we do.

Jeremy: Well, let's talk about that for a second. So the idea of instrumentation, so this was something right from the beginning with Lambda that you really couldn't do? Right? I mean you can't install an agent somewhere that just listens to all the activity that happens with a Lambda function. Now we got layers, we got custom runtimes, mow we have extensions API. So there's different ways that within a Lambda function, you could add some type of instrumentation, even just wrapping the entire function in another function was one of the strategies that was used.

I know some companies would read off of your CloudWatch Logs and of course, just recently, we've got the ability now to attach multiple listeners to your CloudWatch Logs, so there's all these things that are evolving. But you don't have the ability to instrument all of the other things that are part of that ecosystem, so your SQS skews and your EventBridge and DynamoDB. They are logs that are there, but that's the thing where just adding some instrumentation to the Lambda function itself that's a very small part, I think, of your overall serverless application. So how do you make sense of all of that log data and connect all of it together?

Taavi: So really early on two things we discovered, the first thing is that Lambda is such a small part of the infrastructure and what really makes up for most of your infrastructure are things like SQS skews, databases, API gateway, there's a large surface area there that's actually as important as functions. The other kind of fundamental realization was that functions are more simple than you would have coding your containers or something like that. There's a singular thing that any one function is doing, usually if you design it according to the best practices. So the complexity is simple enough that it doesn't need code level instrumentation most of the time. And we didn't feel the pull of the market towards providing customers with really low level data.

So that's why we took this approach. For us, how we provide value for DynamoDB table monitoring, for example, our API gateways is that first of all, a lot of coverage to all of the resources. So if your API gateway is timing out or if it's having failures or there's an increase in anything, then that's automatically discovered and continuously checked for. Really what we're trying to do is to bring their meantime to resolution across any one resource down to as small a time as we can.

Jeremy: And so if you're not doing instrumentation within the Lambda functions themselves, so you're not capturing, like you said, low-level metrics. Just the data that pours out of Lambda. I mean, obviously you've got CloudWatch metrics and things like that that are really helpful. It'll give you failure rates and invocation rates and concurrency and things like that, but the log data itself sometimes has valuable information in it. But if you've ever looked, and I'm sure you have, looked at the log data that comes out of a Lambda function, I mean, it's a lot of junk, it's a lot of stuff that's just useless.

And if you think about log shipping solutions that just take those CloudWatch logs and send them all to some other system, whether that's a last research or something like that, that is a lot of data that you're storing that at least in my opinion is useless. There's things you don't need to know. How do you, I mean, I guess obviously I think your tool does this, but extract value from the log files? Because it seems to me like there's just a lot of junk in there that you don't need and you certainly don't need to be saving.

Taavi: Yeah, exactly and that's another topic that we spend a lot of time thinking on. So the thing with managed service logs is that they are high in volumes and low in density of value. For example, one API gateway request makes around 19 or 20 log lines and most of them are completely useless if it's a successful in occasion. And I can say honestly the same with Lambdas as well, there's a lot of noise. So in our case, what we do is we apply prebuilt filters on top of the log streams, so that if there's a code exception, if there's a timeout, there's any type of service specific failure, then we have a filter for that. Then that we automatically detect and aggregate to see how it happens over time and manage that as a failure scenario.

For a lot of the not so important logs we stored them away somewhere in Cloud storage, we don't keep it in the log analytics part of our platform where it's warm and it's expensive to retain. So I think that one of the important things is getting the cost down from processing the sheer amount of log data. The other part is equipping the engineering teams with the right filters and right knowledge to catch those known and unknown failures that can happen. So it takes a lot of time and effort to actually put together for each engineering team, what could possibly go wrong in my logs and what should I be monitoring for? So that's how you approach.

Jeremy: So then speaking about monitoring, that's another one of these pillars that you've mentioned in the past is the idea of actually alerting like. I mean, a good monitoring solution is going to have alerts, but again, you take a little bit of a different approach, I think, to how you alert.

Taavi: Yeah. I think when we really started to acquire customers, it was after we did alerting. So at first Dashbird was just tool that you could roam around in data and look at different things, but when we started sending emails, when there is a timeout or something, then that tripled the usage overnight pretty much. So you have to send meaningful alarms and that starts 70% of all the use cases. In our platform is just when somebody gets a notification. What we provide for our users is this coverage of whenever something goes wrong across your infrastructure that you should know about, we let you know. So we cover all of the API gateway failures, or if the latency increases for API end points, or if there is a delay in the queue, we manage that alert setting and alert handling. So I think that the situation is that if you have hundreds of resources, each of those resources has five or six different potential failure scenarios that can happen, so we try to put that over head.

Jeremy: Well, and I think that's an important piece of this too, is to say, you can set up a CloudWatch alarm that says when my SQS queue has more than a thousand messages in flight or whatever that is then I want to send myself some alert. And then you've got the ability to send another alert when the threshold drops or something or whatever. But I guess my question around this is even if you know what's happening, like even if you are a serverless architect and you've been doing this for several years, I think I could look at something like me personally and say, "I know what alarms, I probably want on this particular resource."

But what I certainly don't want to have to do is set that up on thousands of resources and do that. So automating those alerts is one, I think, cool benefit. And I know a lot of services do this as well, but just bringing your experience to understand what the patterns are and knowing when something is a problem. So again, is there anomaly detection or how do you set that up where or you automatically set up these alerts for these different resources based off of your experience with what the right pattern for failure looks like, I guess.

Taavi: Yeah, for a lot of those complex alarms, we're looking at historical data as well and seeing if that is the pattern. The changes basically over time is also a trigger condition for us. The way we look at this, like setting alarms, is that some things could be more critical than others. So we look at API end points that have error rates or high latency, more over something that's perhaps more on the downstream, we look at those user-facing things more critically, we look at something that causes a high delay or kind of affects the user experience more. We treat that as a more critical event than for example, something that's like just a little bit slower, abandoned, or not being used. So we try to kind of prioritize as well.

Jeremy: Yeah, no, I think that's super important, I mean, because again, things change over time too. So I could very easily set an alarm that says when my SQS queue goes over, whatever, 500 messages then I should be looking at that or I should send myself some alarm, but if that is slowly increasing over time and historically I'm getting more traffic, so now my SQS queue is backed up a little bit more and it's common for it to do that, having a system that can adapt and understand, I think is crazy important. But anyways, so the other thing though about, and I think I mentioned this before, is about understanding patterns and knowing what's the best way maybe to implement an alarm. Beyond just alarms, there's just best practices out there.

And AWS has a very good resource, it's the well-architected framework and specifically for serverless, there's the serverless lens. I love this resource. I suggest everybody go and read this resource if you're building a serverless application so you know what to do, what not to do. There's a tool that they have that actually allows you to track how compliant you are or whether you're following these things. But it's a manual review process, it's a matter of answering questions. So again, what's the way that in the future, we can ensure that these best practices are being followed without having to have a human go and keep looking at these things.

Taavi: So in Dashbird's, not to do much product placement, but when we discovered all the types of data that we have, if you have basically all of the monitoring data and you have alerts set up in your platform and realize that there's this opportunity to actually run a lot of analysis on top of that there's this a whole book or framework around what the serverless application should look like and what are the best practices around security, around operational excellence. And we kind of discovered that actually, a lot of this, we could find out using the data that we already have and build the system that continuously surfaces and pushes the user towards the best practices.

So today we have a collection of rules that we continuously apply and check for, and then get back to you with this list of, "Hey, your API endpoints are not encrypted, or you're not using the right encryption in your databases, or you have something that's unused or abandoned or not tacked." And a lot of those things we can service and push the user towards. We're trying to automate and equip teams with the best ways of following the best practices of the industry, so that's what we're building and having quite a lot of success recently as well.

Jeremy: No, I think that's really cool. Because I mean, that's one of those things where it's so hard. I mean, if you think about static analysis of code and you can catch some things with that. You could look at configurations and you might be able to say, "Oh, the security, you've got a star permission here or something like that." But until the code is actually running and data's flowing through it and you're seeing what happens there, that's, I think it's where the rubber hits the road there and you can see how that stuff works. So that is that's fascinating. Now I do have a question though, I mean, best practices and serverless are really hard and I know that AWS has their serverless lens for the well-architected framework and they make really good suggestions, but there's always a time at least for me, I do it quite a bit is you have to break the rules, sometimes to make something new happen. So how is your system going to deal with breaking those rules when it needs to?

Taavi: Well, that's the constant challenge is to keep the alarms adequate. And when a user looks at this and says, "Hey, I do understand why this is here, but it doesn't apply at all." And I think when we first started alarms were going off all over the place to be honest. Every time we removed the ones that are optional, but it's not easy. I think that the other thing that we can play around with is the critical level. How critical is something? If you're really exposed somewhere, then you should have a high priority alert. It's like, you're not tagging your resources perhaps that's not as important.

Jeremy: Awesome. Go ahead.

Taavi: I just wanted to say that we're never going to get to the perfect 100% with those insights, but yeah.

Jeremy: Yeah, I really do like that approach. And again, there's a lot of observability companies out there, there's a lot of log shipping companies, monitoring companies, all these different things around serverless and they all do a really great job. I mean, for what they're doing. But I do really like this approach where you're saying, "Let's take a step back and solve the observability problem, solve the mining part, but also solve that operational quality and operational excellence problem." I think that's an interesting approach. So good luck with that because I know it's not going to be easy. Like you said, you're dialing that in to get it right. So let's talk about just, I guess, monitoring your Cloud and your operations in general, because this is something where we don't always go deep into this when we are talking about observability.

But I guess a question that comes up is now that we're doing serverless and now that we're using managed services for a bunch of different things, what is it that we're actually monitoring for now? What are those important metrics? Because if you think about it, I don't care about CPU anymore or memory uses. My DynamoDB tables don't tell me how much CPU or memory they're using. I only know how many read units or write units I use and as long as the latency is where I need it to be those seem to be the metrics I care about, but those are different across all these different services. So what are we as an, I guess, as a monitoring community, if that's the right way to put it, what are we looking for as important metrics?

Taavi: I think there's two answers here. The first one is that serverless is essentially like a layer of that abstraction. So it abstracts as a way to underlying compute resources. So what we recommend our customers to monitor are user facing things like how fast are the responses from the backend, what's the downtime and quality of the service. Like how well are the users actually experiencing the system? And that's the first layer that we usually recommend them to cover.

So do I have alerts and API gateways, for example, things like that. But on the other hand it's, what's this microservice costing you for example, or how can you make it quicker a bit or those kinds of things. So really the business and the user impact is what we mainly try to monitor. I think the other challenge is that just making sense of all of that data that's the system is outputting is like if you have tons of monitoring data, then trying to extract the value from that and to identify pieces where you should be really focusing on. So I think that that's the challenges and the approach that should be taken. Another thing.

Jeremy: And I think that's interesting too just from a, I guess, a community or an education standpoint that observability companies like Dashbird almost have the ability to help educate people that are using these different services on what the important metrics are. And then not only what the important metrics are, but also maybe what the baseline for those metrics should be. You know what I mean? I guess the error rate on your API you'd love to be zero, but maybe the latency, for example, if you're connecting to a DynamoDB table through a API or a API gateway to Lambda, to DynamoDB table, like what that average response time should look like and things like that would certainly be helpful.

So then, I guess from a more operational standpoint, I mean, there's a lot of people who equate serverless to no Ops, which it's clearly not no Ops. I mean, you significantly reduce your operations and there's many other things your operations teams could do. They could focus more on security, on automation, some of those other things, but what about the overall responsibility of some of this monitoring? So, again, I like the approach where you don't have to instrument your code, so it just happens behind the scenes. But, I guess, where does the data come in when it comes to whether it's optimizing and following those best practices or optimizing for costs or performance or whatever, or just monitoring the overall health of the application, where does that responsibility fall now? Do you consider this to be a developer tool or do you consider it to be an operations tool or somewhere in the middle?

Taavi: I think that the trend we've seen in the serverless era is that a lot more responsibility actually falls to the hands of the developers. And not just the operations side, but also based on the business side or it brings developers more closer to the customers in a way as well, because the task is less on building undifferentiated value and more on actually solving the problem for the customer. I don't know if sadly or it's a good thing, but a lot of the operational burden also seems to fall on the developers. When we really talk to our users or customer calls, then we usually see engineering leads or developers, architects, not a lot of operations or DevOps people, to be honest, there are obviously, but I would say it's like 20, 25%. So it's more developers I would say. And I think a lot of what we do is around debugging still, and improving the system in general, security wise and things like that and those things are always done by developers we see.

Jeremy: That's a good point about debugging. So in order to debug your code in a again, a Cloud distributed environment, is that something where you need to be using one of these tools to do that? Is that what debugging looks like?

Taavi: So, yes, but I think developers can do it with CloudWatch as well. So when we break it down there's two user stories or two ways of using. First is while you're developing and iterating the application and trying to understand all the bugs and to fix them, then that's something that you need to iterate quickly and do deployments and test it out. And one way to do this is with us and we do have some things light tailing and real time representation of what the activity is looking like, which may be a bit more simpler than CloudWatch is. On the other hand, you can still do a lot of that in CloudWatch, or you try to develop locally as well. We see the bigger value posts in environments where that's already in production, have a lot of users, a lot of load. Then the monitoring part becomes a bigger challenge and tags where we would position us more, but we still see debugging use cases as well. It's just that you can do a lot with CloudWatch as well.

Jeremy: All right, so then what about the monitoring strategy? So you said that again, you see 25% or so of the people that are jumping on your calls are operations people, and that a lot more are the developers. So from a monitoring standpoint, I mean, typically you'd be monitoring to make sure that the CPU and all the servers are running? That's how we used to do it. And you might have an Ops team that does that. So what is the strategy now, so for a serverless team that's developing a serverless application, maybe there's somebody in operations that's helping with VPCs or something like that, but what is the monitoring strategy now? Is it the developers who should be in there getting those alerts, or is it still some hybrid dev ops solution that you're trying to mix and match? I mean, how would you suggest a team use one of these observability tools so that they can make sure that their applications are running smoothly?

Taavi: So if you're going into production with your serverless application, or if you're thinking about monitoring the general, what we push users towards is really set some clear goals for the monitoring solutions basically, or if you're building a monitoring strategy, what are the core things that you should be thinking about then? In our case, what we see being the most important ones first is the ability to quickly understand if there's an issue to quickly get notified and to reduce the time it takes from anything happening to your development team knowing about it and there's a lot of things you can do there. You can map out the failure areas where you have the most risk, and then you can map the end-to-end ways, or monitoring API end points, for example, or things that are really user-facing. Map those out then set alarms for those.

And the first part is really about getting notified as quickly as possible. So the second part that we think it's really important with monitoring strategies, having access to the right data at the right time. So if you're discovering that something is not working, you need to have the infrastructure in place to be able to understand why it's not working and to go through all of that data, to have that data available, and to show you where the problem is. I think those are the two main things for any monitoring strategy. If those are clear, then it's easy to make the lower level of decisions from there, like what types of data you need and things like that.

Jeremy: I think that's important. You said, something about understanding where the risks are in your infrastructure. Because one thing you see, certainly with serverless applications is I have not seen very many 100% or applications that are a 100% serverless, you always have some hybrid in there, you're still accessing a SQL database or a my SQL database or something like. So that's I think something that's interesting is what do you do to protect against the brittle components? Is this something where you add more alerting and more monitoring to that? Or is it something that it's just another piece of your infrastructure that you treat as just like you would anything else?

Taavi: Yeah. I think it's important to be aware of those areas, if those exist. If you have a SQL database or sometimes free service that can be easily troubled, then definitely designing around that is important or that in mind and treating it as a specific failure point, paying more attention I think is necessary and that's what we recommend to do as well.

Jeremy: Awesome. Well, so I guess my last question, just, I love to ask people this question. The future of serverless ... and this has been asked and answered a million times, and everyone seems to have a different answer to it, but just because again, I like how you're thinking about this approach to it. Are we going to see serverless dominating the Cloud world? I mean, is it just going to be the way things are ... or, well, let me take a step back, ask you this question: what's the future of serverless, what's it going to look like five years from now?

Taavi: So the way we see it and the future we're building for is the future of our developers is construct their applications out of Lego pieces and doing very little coding or only focusing on the differentiated value that their organization brings and having at their exposure, a lot of tools that they can just kind of piece together and use. And I think the creativity in Cloud is definitely that serverless is faster to build on, or the pricing is based on the actual usage and it's way more simpler to understand different components as well. So that's what I hope will happen and it won't just be AWS, but it will be this entire Google Cloud, Microsoft Azure, a lot of third-party services will play into it. And there will be a lot of different tools that you can choose from, but they'll all be managed in single purpose.

Jeremy: Yeah. No, I love that. I hope for the same thing and I think you're right. I think the ecosystem will continue to expand and Azure is doing some great things in terms of what they're building out for serverless. So it will be really interesting because having this conversation five years from now could be completely different, but anyways. Well, Taavi thank you so much for joining me and sharing all this knowledge. And I mean, again, the Dashbird, the product direction you have there is really interesting, it's a really cool approach. I love that idea of just trying to make sure that you implement those best practices and give people the tools to do that. So if people want to find out more about you or figure out what you're up to and find out more about the Dashbird, how do they do that?

Taavi: Sure. So Dashbird.IO is where you can contact me or our team as well. And my Twitter is @Rehemägi, so feel free to reach out and happy to chat about serverless anytime.

Jeremy: Awesome. Well, I will put all that in the show notes. Thanks again, Taavi.

Taavi: Thank you.

This episode is sponsored by New Relic and Epsagon.

View Details

About Vadym Kazulkin

Vadym Kazulkin is Head of Technology Strategy at ip.labs GmbH, a 100% subsidiary of the FUJIFLM Group, based in Bonn. ip.labs is the world’s leading white label e-commerce software imaging company. Vadym has been involved with the Java ecosystem for over fifteen years. His focus and interests currently include the design and implementation of highly scalable and available applications, container and orchestration technologies and AWS Cloud. Vadym is the co-organizer of the Java User Group Bonn and Serverless Bonn Meetup, and a frequent speaker on various Meetups and conferences.

  • Twitter: @VKazulkin
  • LinkedIn: https://de.linkedin.com/in/vadymkazulkin
  • Email: v.kazulkin@iplabs.de
  • Presentation: https://www.slideshare.net/VadymKazulkin/measure-and-increase-developer-productivity-with-help-of-severless-by-kazulkin-and-bannes-sla-the-hague-2020-238115659

About Christian Bannes

Christian Bannes works as Lead Developer at ip.labs GmbH and has been working in the professional Java Enterprise environment for over ten years. In recent years, he has been working with AWS Cloud and especially with serverless applications. He is particularly interested in distributed architectures, domain-driven design, and functional programming.

  • Email: c.bannes@iplabs.de
  • Website: www.iplabs.de

Watch this episode on YouTube: https://youtu.be/QoKW1KR7rrM

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I am chatting with Vadym Kazulkin and Christian Bannes. Hey, Vadym and Christian. Thanks for joining me.

Vadym: Hi. Thanks for having me.

Christian: Yeah. Thanks for having us.

Jeremy: Awesome. So you both work at ip.labs in Germany. And so I'd love to talk a little bit about what IP labs does and what you two are about. So let's start with you, Vadym. So you're the head of Technology Strategy. So why don't you tell the listeners a bit about your background and what ip.labs does?

Vadym: Yeah. I'm a Ukrainian native, but I live for 20 years now in Germany. I have been working with Java for 20 years but since three years, I'm involved in the migration stuff and AWS as a cloud provider of our choice. And I'm a part of the serverless community since two and a half years, involving me heavily in all this stuff and presenting ideas and our experiences mainly also with Christian. So this is what I do.

And ip.labs is software provider for designing and purchasing of photo products like prints, calendars, photo books, just where you can print your emotions. So they are part of the Fujifilm group, Europe. So founded 60 years ago, approximately 80 colleagues, 30 developers.

Jeremy: Awesome. And Christian, you are a lead developer there. So why don't you tell the listeners a little bit about your background?

Christian: Yeah, right. So I'm a software developer at ip.labs, also working about 20 years, almost only with Java technologies. So I'm working as scrum team. I think about three years ago we adopted serverless and we switched to TypeScript because it fits our need more than Java. And yeah, we are quite happy with serverless.

Jeremy: Awesome. All right. So I have seen the two of you give a presentation. I know you've given this presentation a few times, about measuring and increasing developer productivity with serverless. And this is always to me a fascinating topic, because you see a lot of claims, right? And a lot of it is very anecdotal. I mean, it's like, "Oh, yeah. We were able to move faster with serverless." Or you hear things like that. But the two of you actually sort of did some research on this, dug into the background of this and really outlined this well, and I think it should be super important, or it's super important to share with listeners so that they understand why serverless is such a powerful productivity booster for software development.

So I'd love to start, like maybe way, way back in the beginning, and just talk about software development in general. So when you're building applications, and you're trying to create whether it's new stuff, and you're trying to build greenfield applications, or you're trying to work on legacy applications, what are the challenges when you are trying to, as a software developer, what are the challenges that you have to face?

Vadym: So I think that the best model to explain this is the cognitive load. And this is the term coined by Matthew Skelton and Manuel Pais in their recent book, Team Topologies. And the cognitive load is the total amount of mental effort being used for accomplishing the task. So accomplishing the software development tasks. And then they differentiate between three of them: intrinsic, extraneous, and germane. And probably intrinsic, is very easy to understand, because it's how you write Java class, TypeScript class, or use some framework of the day. So this is something that you can't offload, you have to learn this. So you have to own this also.

But then you have this extraneous load. And it's especially important to understand in our distributed world that many things are currently distributed. So just how to automate your tests, unit tests, integration, end to end web, mobile tests. How to build package, deploy, and run your application. How to configure, monitoring, alerting, logging, everything. So just operate and maintain infrastructure. So how to build fault tolerance system and resiliency and of course, security is also job number one. So just it's not only application security, but also preparation system, networking, hardware, everything. And just huge bunch and you haven't written even one line of productivity code, but you have to deal with all this stuff, probably. And I see a lot of companies which really struggle to deliver value if they go distributed because all of these challenges and distributed system are hard challenges.

Jeremy: Right. Yeah, then germane. So you've got intrinsic, you've got extraneous and then germane. What's germane?

Vadym: Germane is, this is your business logic. This your workflows, this is your core domain that you implement. So you have to become expert in the things which you are doing. So you have to understand what your core domains, what are your generic domains, like probably commercial system, e-commerce system, something like this. Every everybody needs payment, check-out, but it's probably not your core. So you'll have to reduce also this law to only things which really core and meant for your business. So these are three different cognitive load types. And if you think about this, so just, you want to reduce extraneous and germane load as much as possible to focus on the business stuff that matters.

Jeremy: Right. So the other thing we kind of talking about in here though, is like, again, if you're spending all this time writing or working on this extraneous stuff, obviously, it's taking away from you implementing something, and actually being able to ship some product. And this comes back to productivity. So what exactly do we mean by productivity too, because that's probably one of those things where I think people spin their wheels a lot and you check off a lot of to-do items, but is even figuring out how to implement something like to automate a task or to build and package and deploy your applications. And that's not really being productive, is it?

Vadym: Yeah, being productive means regular shipping your product, which of course used by the customers, but I think that some something very obvious but in our time the productivity and the speed really matters. So you have to offload or you have to try to offload as much as possible to focus on shipping things which really call for your product.

Jeremy: Right.

Vadym: So I think that's the idea for definitely productivity, from the word product, probably just... Could be the word.

Jeremy: Makes sense, right? Yeah. So what about things that are holding us back, though. I mean, so you talk about extraneous things. And obviously, there are a lot of things that a developer has to be thinking about when they're writing their lines of code, or they're implementing something because everything they do, every line of code they write is going to have some impact down the road, like someone's going to have to maintain it or whatever. So what holds developers back from being productive?

Christian: So the problem is when you try to implement all the stuff yourself, and that's a problem that we had at ip.labs. So we try to implement really everything on our own. And so writing it is, would say, can be easy. So you have teams and they can do it really fast but the problem is that you have to maintain it for a long time. So we started our platform about 15 years ago and we implemented like payment and e-commerce systems and so on, which are actually not part of our core domain. And the problem is that all the stuff has to be maintained now for over 15 years, and that's a lot and you get more and more like technical debts, because you want to go on and on. You need new features but of course, you will have a lot that you have to maintain. And I'd say this, this holds you back a lot.

Jeremy: Right. And so let's let's talk about technical debt for a second. Because again, I think that people hear the term "technical debt," and they probably in the back of their mind they understand what it is, but how would you define technical debt so we can really get some concrete examples here?

Vadym: For me, personally, technical debt is everything that can happen with your software or the entire lifecycle of your product. Just if you look at the definition of the Wikipedia, you will see this is some kind of suboptimal decision that you met today, which you have to correct tomorrow. But you can make the perfect decision today and it will be technical debt tomorrow either, because a lot of things happening just rational programming languages will be deprecated or end of life. So just that's happened to your JavaScript framework of the day probably very often, but it also happens with such mature programming languages like Java, they do have some breaking changes from time to time. If you have to switch off, but also other things like security consideration, that forces you to upgrade, just the situation with TLS 1.0 and 1.1 which are becoming deprecated, then you have to switch to TLS 1.2, and even further there is 1.3 standard.

So you probably have to update your web application server which may force you to update the version of your programming language and so on. This is some kind of situation that you are steadily forced to update thing. And if you own too much, then you also have to update too much. So there are also some funny situations which we experienced, we use some encryption algorithm and our payment provider forced us to update the strength of the key. We updated this, but this open source breakthrough threw the exception, and we saw, "Okay, this project can't deal with this key." So we searched for the newer version of the project but the project was discontinued. So we had to take another one and reimplement the whole thing.

So it was the perfect decision several years ago, and now we have to do this stuff and it's not very value-generating currently, but it's some kind of must have. This is security. And this is just a lot of things happen in our industry which forces us to do those things and I think developers like this migrations, like this upgrading, but generally, this is what holds us back from being productive with our product.

Jeremy: Yeah, right. And I love that idea of you know, saying technical debt is not suboptimal decisions, right? It's not like I said, "You know what, I'm going to cut some corners here." And that's going to give me a bunch of technical debt. Now certainly, that will give you technical debt if you make those types of decisions. But you're right, you can make a perfect decision at the time, right, and even thinking forward years down the road, you can say this will most likely be the right choice. And it's certainly the right choice at the time. And then you just end up with like you said, things go out of date. I mean, like they deprecate you know, node six, then node eight, and then eventually, node 10, for supported Lambda runtimes, you know what I mean?

So things are just going to start disappearing. So one of the things though, that I think contributes to technical debt is obviously, like you said, okay, you pick some service, some maybe open source package to do something for you. And then that package eventually goes away. If that package didn't go away, it was just a matter of upgrading it, well, maybe that's not too difficult, maybe the API doesn't change too much, maybe there's not too many technical changes, or breaking changes. But the bigger problem is, is that if it goes away, or if there are significant changes to the API and the interface, then you have to go and change code that you've written. And it seems like in almost every case, technical debt is highly related to the amount of code you have to write.

Vadym: Yeah. That's true, it's related to amount of code and it's also related to amount of dependencies used. Like open source project, programming languages, database drivers, web applications, sort of everything is dependency and everything will be changing, and it will be forcing you to upgrade. So this is some kind of the circle, so the only one solution is to own as much code as possible for you, and this code should have as little as possible dependencies, just enough dependencies, I would say.

Jeremy: All right, so now you're making decisions in you own some of the code, but you still going to make decisions, because again, you want to offload some of that undifferentiated, heavy lifting, as we talked about. So you are going to want to use some open source things or some third-party tools, managed services. So how do you then make a decision today that hopefully reduces the need for maybe re-architecting an entire system. Like how do you build code and build applications, so that you can upgrade them as things change with as little effect on the entire application as possible?

Christian: So one way to organize it is with evolutionary architectures. And the idea comes actually from evolution. So the environment constantly changes, so we have to adapt. And of course, some parameters are constant over time. But it's like the climate in Europe was different 10,000 years ago, and then the Ice Age ended and stood constant for a period of time and now it's changing again. And because it's changing, we have to adapt. And the same with architecture. So you have an environment and the environment is like business requirements and the technical environment, and both are changing. And because they are changing, you need to have the ability to change your architecture.

There's the saying, never touch a running system. And the assumption behind that is, if something stays the same, it won't break, right. But the problem is that even if your system stays constant, the world around you is changing. So for example, if you take a laptop, and you put Windows or Linux on it, and you close your laptop and put it into a cupboard. And say you leave it in the cupboard for two years, or four years. So what happens if you take it out again, and you open it. And what would happen is it would install a lot of stuff and update a lot of stuff. And maybe if you have interfaces to the outside world, maybe some things would even break.

So you didn't touch the system, but your system could break. And the reason is, the outside world is evolving, or is changing. So it's not possible today, to say, "Okay, I'll leave my system constant, and I don't touch it anymore." Because the outside world is changing. So you have to have some ability to change your architecture. And the idea behind evolutionary architecture is to build this ability in to this architecture. And change is easy, if it only affects a small part of the system. So if you need to change something, and you need to change the whole architecture, this is really hard. But if you change just a small piece of it, is possible to just change a small piece of it. It's easy to make a change. And evolutionary architectures actually have three components. And maybe then we will also understand why serverless is a really good fit for evolutionary architectures.

So the main three components for severless architecture is a fitness function. The fitness function is basically business requirements, maybe you can automate it like webpage should respond in like 200 milliseconds. And the architecture should evolve in this direction of your fitness functions. And the second thing is the so called granulum. And the granulum is the smallest deployable unit, or the smallest thing that you can change independently. And in serverless application, the smallest thing that you can deploy is a Lambda function. So it's really small, it's based on the function level. And this makes it easy to change. But there is another thing, that's also very important. And I think this is one of the most hardest thing to do right. And this is appropriate coupling. So now the question is, what's appropriate? And appropriate is things that belong together, should stay together.

So if you change one Lambda function, always when you change one Lambda function, you also have to change the other Lambda function. So does it really make sense to separate them into different deployable units, or maybe into different Git repositories. So when you make a change, you would have to check out maybe multiple Git repositories, you have to declare multiple lambdas. So this would make change really hard. So things that belong together, should stay together. But it's really hard to know what's appropriate. And I think you will make a lot of mistakes and lot of errors in this area. So during your evolution or during your implementation of your application, you will probably recognize that you had some parts that are loosely coupled, but should stay together. And maybe you would also recognize that things that are highly coupled, should better be loosely coupled. So as you get more insights about your architecture, you should always refactor your code to match the appropriate coupling and the granulum. This is a really hard thing, I think. But this is really important to get evolutionary architecture, right.

Jeremy: Right. You know I totally agree. I mean, that's one of those things where understanding where the bounded contexts are, and some of those things can get really difficult ...

Christian: Right.

Jeremy: ... And then also just what we mean by a single purpose function, what that single purpose function does, and how that connects to other things, whether that's using orchestration or choreography. There's a lot to think about there. So I want to go back to that actually, in a little bit. But before we do that, let's talk about serverless in general, because you brought up why serverless is so great for these evolutionary architectures. So what is the value proposition of serverless when it comes to this idea, of not only building evolutionary architectures, but just like being more productive as software developers?

Vadym: Now we have mentioned for the first time the term serverless is it's quite unusual for your podcast after 20 minutes. But generally, this was the some kind of preparation job that we have done. And if we are talking about the value proposition of serverless, it's really huge. It starts with such obvious things like no infrastructure, operation, and maintenance, and this is part of extraneous load. And ask yourself, if the infrastructure management, and operation is core for your business? For AWS is probably yes, but the most of the companies, it's probably no. Even Lyft is spending $300 million per year on AWS. And they can spend less in the data center but this will slow them down.

Jeremy: Right.

Vadym: So just it's a decision. And also auto-scaling and fault tolerance built-in, it's also a huge part of the extraneous load, and it's just part of the platform, and ask yourself, how difficult it is, with all this capacity planning and so on. We at ip.labs we have huge Christmas business for three weeks with a huge spikes because people are making gifts, photo book of the year, and so on. And I really know how difficult it is and doesn't make any sense to own too much infrastructure for only two weeks, just 10x, 20x from just doesn't make any sense for us, even to think about this. Because we are also business to business companies, our forecast so the our partners should bring us forecasts. So they don't know, and how should we know, just it's not possible for us.

And, of course, this is the idea, the next idea to do more with less just in case of greenfield project, you can make prototypes very, very quickly. So just you don't own infrastructure, you pay on demand, it costs you really nothing to prototype. So just you can do things easier. And probably by relying on managed sources, we'll really talk about this. Just you can do more with the same amount of people. And I think that's matter. So you have static, so fixed amount of people currently and can you can really do more with them.

Jeremy: Yes.

Vadym: And it's really powerful if you think, "Okay, I can offload some really nice technical things to the people, to the platform." Which for them it's the core thing, and they can help us to do things quicker. And then we are, yeah, we are talking about the technical debt or to have low technical debt. So we're talking about amount of code. And to minimize this and really like [something Johnson? 00:23:48] was very known in the service community. Whatever code you write today, it's the technical debt of tomorrow.

And it's just true. Just as we have explained this and the best code is no code at all just that I know the codes. But also you have to think that configuration infrastructure as code is also a part of your code. So you can't separate this, as the whole is your code. And of course, think of how much time do you spend maintaining your solution over the whole life cycle. And implementation is sometimes quickly but my experience and what I've read you spend 75% of the time maintaining your solution, and it's huge part. So you have to think about the entire lifecycle of your application, how to reduce your effort at maintaining it because maintenance doesn't have too much failure with just this, you have to do this and you have to free you up.

And then of course there are obvious things. So if you can free your app, then you can focus on your business value and innovation and have faster time to market. And it's probably every company want just this. Everybody's crying, we want to be innovative, we want to be fast. And that's what matters.

Christian: Yeah. And I think this is really a problem that we had at ip.labs. So we really try to implement a lot by ourselves. But actually, what you want is you want to concentrate on the core domain and not on the sub domains. Like sub domains, like payment or authentication, and effort in search, possibly you could do it really fast. But you have to maintain it for all time, and that's the problem. And ideally, what you want to do is concentrate on the core domain and don't implement the, like in domain-driven design you call it generic sub domain and supporting sub domain. And if you are able to use a managed service, what you get is, for example, free bug fixes. So you don't have to fix the bugs, because they will fix the bug, you don't have to do the operation, you don't have to do probably the scaling.

And you will also get new features. So you don't have to implement those new features by yourself. But another team will do it for you. So you can concentrate on your core domain because this is where you have your competitive advantage. You don't get any advantage from your subdomains, they're only there to support the core domain. And that's why it's really important to write less code in subdomains. So you can write more code for your core domain.

Jeremy: Right. And I'm glad you brought up domain-driven design again, because this is one of those things, I had a conversation on another podcast about this. And just this idea of how you implement these microservices using serverless, right. And so I know, at ip.labs you have some experience doing this. So what's it look like in practice when you're building microservices using serverless?

Vadym: Okay, just that it's about our service mindset. And I've seen this, this tweet from Jared Shorpe and about how to proceed and we have adopted this to our realities. But generally, serverless is really an operational construct. So our idea is to be as much serverless as possible. So that's services which are completely serverless, like Lambda, S3, S-Square, EventBridge, and there are services which are a bit less serverless, like Fargate, probably also Kinesis because you have to manage your files and so on. And then you have AWS and so on. And just think we ask ourselves, can we be completely serverless? So as much as serverless as possible, without dealing with capacity planning and so on, even DynamoDB they're not now offer capacity completely serverless mode without calculating read and write capacity units and so on.

And just the first thing to ask yourself, can we implement this completely serverless? Does AWS offer such as service which can help us? And then the answer is yes, then and the feature set is good enough for now, I would say then we will always choose serverless currently. In case it's not possible then we are asking, "Is there any AWS service which offer this service but we have to manage a bit?" And then it's the decision number two, and so on, sometimes even we have to choose some service which is outside of AWS like PagerDuty because we also have parts outside of AWS and they have to build incident management for overall system. So just this is some kind of other options. I think really, one important option is to reconsider your requirements. If they had the situation wonderful conversation for using BPMN as a workflow management system. We had the same experience but we said we just currently, step function didn't have enough features for us. But we could reconsider our requirements and still use step function because it's completely serverless.

But now they provided this feature. So this was the right decision to stay within the serverless ecosystem. Because it grows, it improves steadily and it will be... If it's now not the perfect decision, it will become one in the future. So we are constantly thinking how to embrace this system and stay with this system. But of course, if it's your core domain, you have to write this code and you have to own and also run this code, but it's the last option to choose to write the code by ourself.

Jeremy: Right. And I actually I really like that thought of saying that you build something now you maybe change requirements, but then you know that that service is going to get better. And it's going to have more features, right? So it might be even a better choice in the future, sort of the opposite of technical debt, right? It's going in the other direction, which is sort of a good thing.

But so I get the mindset, right. And I think the mindset is super important when you look at that, but then when you're actually implementing it, when you're following through and building these different services, how are you organizing Lambda functions? How small of units do you break it down into? I know, Christian, you said that it's a really hard problem to figure out you don't want too much coupling, you don't want too loose of coupling either. So how do you build out your projects? How do you kind of think about those for long scale, architectural planning and things like that?

Christian: Yeah. First, I would like to say a few words in general about architecture, because I have people saying that for serverless applications, so actually, you don't need to think that much about architecture, because Lambda functions are so small. And if you do something wrong, you can just throw the Lambda away and write a new one. And of course, this is wrong, so just because you're using serverless, doesn't mean you don't need architecture. Architecture is usually a long term investment. So it doesn't really have a big impact, maybe after months, but after years. But this also means for really small applications, maybe it's true that architecture isn't that important.

But for large scale applications, it's really important. So if you have large scale and really complex domains, then it gets really important. And this is true for a Java application based on Spring Boot and this is also true, of course, for serverless applications. So when you have a large scale application, I think the idea of domain-driven design gets important again, and as I said, this is true for other applications or other platforms. And this is also true for serverless applications. And this is also the way how we try to organize our serverless code. And so we have a large domain and we try first to split this large domain into smaller sub domains. And so our sub domains are still too large so we further try to organize those sub domains into smaller bounded contexts, and inside a bounded context, recreate a domain model.

And so we usually don't deploy Lambdas independently. So we use the serverless framework and we bundle everything that's part of the same bounded context or the same service into one serverless yaml file and deploy it in one unit. So we always try to organize around business capabilities. And I think this is really important.

Jeremy: Yeah. I think that's a really good way to break it down. And I think the strategy that you see a lot of companies that are maturing with serverless are going down that path, right? I mean, you see some companies just building monolithic applications where everything's kind of flying around. But what about like data separation and things like that? How do you break up your data between those different contexts?

Christian: Yeah, I think this is a really interesting topic. So I think one question that arises is should I use one? So we are using DynamoDB. It's also a serverless database. And the question is, should I use one table or should they use multiple tables? So multiple DynamoDB databases. And if you look at the documentation from AWS, or watch some re:Invent videos, they recommend you should use one table per application. And you should understand why this is the recommendation. And the reason why you should use one DynamoDB table instead of multiple tables is, when you come from a relational database background, usually you normalize the data. And that means you create one table per entity. So you have a user table, you have order table, you have order item table, and so on. And if you, for example, wants to have a query, if you want to query all items that use them, you will have to do a join over user over order over all items.

And this is really flexible. So you can do ad hoc queries with SQL. But the problem is, this is not scalable. And this can be a huge problem. And DynamoDB is a scalable database and it scales by avoiding joins completely. So with DynamoDB, you should not do any joins. And internally, DynamoDB is working with partitions, so it has I think it's at max 10 gigabytes per partition. But you can add partitions to the database. And so it can scale almost infinitely. But of course, you still need to model relational data also in DynamoDB. So how would you organize this? And you would do it by denormalizing the data.

So of course, this means that you have to duplicate a lot of data but the idea is that your items are already joined. So if I give you an example, say we have an order and you would have the order ID as partition key and the order item ID as a sort key. And then you can do really one query based on this hash key. And you would get back the order with all the necessary information of the order item of the user and so on. And this is actually the reason why you should use only one database table, because you can now avoid joins.

So what does it mean for microservices? Because when we talk about microservices, so actually every service should own its own data. So you wouldn't do a join between different microservices. So it's actually okay to say each microservice should have its own database, it's totally okay to do it. But of course, you will have some additional operational overhead. And yeah, maybe you will start with five DynamoDB tables. But think about what will happen in maybe five years or 10 years? Will you have 50 tables or 100 tables? Will still be okay for you. Because so there's a little operational overhead, it's much less than with relational database, but you still have to do like monitoring, throttling, and so on. When you have a large number of databases, this might be too much. Can be okay, but it might be too much. And that's why we decided to share a database across our services.

So we have just one DynamoDB table per sub domain but we share it across multiple services. But still, every service is only allowed to access the data of this service. And we make sure by using fine grained access control. So in our policy, you can configure that service only allows to see the data of the service. So there is no possibility to break something. But of course, it might be a disadvantage, because now you have multiple services, and they aren't completely independent. But for us it works really good. So we have one DynamoDB table. It's not dangerous, because every service can only see its own data and we don't have this operational overhead.

Jeremy: Right? Yeah. Now, if you want to talk about cognitive load, thinking about how to structure data in DynamoDB to be shared across multiple services, it's certainly something you have to really think about. And I do like that idea. And every time I see AWS make that recommendation, one DynamoDB table per application, I always think of, "What do you mean by application?" And if you have a system with microservices and you have 100 different microservices, well, each one of those microservice or each microservice might be its own application, if you think about it that way.

Christian: Yeah, that's right.

Jeremy: But it's interesting because I always like to see how different companies implement that, because in some cases, it does just seem like it makes more sense to share a table across multiple services. But like you said, as long as they're in the same domain, and it's using the same sort of descriptive language and the same vocabulary, then it's probably less of a risk than if you were sharing across multiple domains, for example. So that's really interesting. So what about ports and adapters? So I know, hexagonal architecture is something you've mentioned in your presentation, and so this is something that you invest in heavily and just in case people don't know, can you explain what do you mean by ports and adapters?

Christian: Yeah. So ports and adapters is an architectural style. So it's actually a very easy idea. So you have your domain or your domain logic, and you try to separate it from any infrastructure logic. And you do it by implementing ports in your domain logic and different adapters, like DynamoDB adaptor or email adapter can plug in into this port. So in the middle, you have your domain code, and around this domain code, you have your adapters that you can switch. And the whole idea is you should try to separate infrastructure logic from domain logic.

Jeremy: And I always say this is something that I liken very much so to like a data access layer, right? Where you sort of genericize your database, not an ORM, we're not talking about ORMs, but like a database access layer, where you'll have a "get customer," you'll write that function, and then how that interface is actually implemented to the database, that's a completely separate thing. So you can always switch that out, if you ever needed to change how that works in the future.

All right, I want to go back for a second, because we talked in the beginning about this is about increasing productivity. We said productivity was this idea of shipping code and whatever. How do you actually measure this success, though? Because that's something where, I mean, I know there's a really good book called Accelerate that kind of outlines some of these success factors. But how do we measure the success in an organization and how does that tie back to serverless?

Vadym: So you, you mentioned this book, Accelerate by Nicole Forsgren, Jez Humble, Gene Kim, they're probably very known in DevOps community, continuous delivery, continuous deployment, and so on. And what they say, but what they found out is they wanted to investigate what make organizations vary performance. And currently, there are two things, the quality and the speed. And the message is to be successful today, you have to combine both. And they define metrics called four key metrics and two of them are for the speed, this is deployment frequency, how often do you deploy? But of course, it's not that you should deploy every change, you should be ready to deploy this. But in case the business says, I need this life.

And the second metric is lead time for changes. So how quickly can you deploy? So how much time do all your tests take? And so on, the whole pipeline thing. So how quickly can it can I deploy my code into production? It's about automation, and so on will probably also talking about, and two other metrics are about the quality. And these are time to restore the service. So if you see something went wrong, how much time does it take to go to the state that everything is working automatically, or by back fixing and just doesn't matter. And the second one is change failure rate. So how often you deploy things and also then break things because it happens if you are very quickly. But the thing is, you have to restore them quickly. And of course, they divide organizational different performance types and, of course, nobody wants to be low performer, but just they have some guidelines.

What do they mean by becoming high performer and then you see the things that in terms of deployment, it's something mainly times per day, it shouldn't be that like Netflix or Amazon, they deploy 1000 times a day, it's probably not the case. And even if you are in the mobile market, it's not very obvious to deploy and update so quickly. But in terms of lead time for changes, it's really about hours, minutes, or hours today. And the time to restore server it's the same it's for high performance organization. It's far less than one day in and that's really these are the four metrics, which are really important. And they're not tightly ... not very tied to serverless. But as a general recommendation, what outcomes do we want to achieve. And these four key metrics are outcomes, and was best practices and so on belongs to do the things how you want to achieve them.

Christian: Measuring productivity is actually really hard. So I think many agile teams use story points. But story point isn't actually a good metric to measure velocity, because it's actually very easy to optimize this metric. So I could optimize it by say, I have a task, and instead of giving 10 story points to this task, I can give a 20 story points to this task. And I'm twice as fast right? So, of course, I'm not twice as fast, and I just changed the estimation. And so it's not a good metric. So what you actually want to measure is the lead time. So you get a request and how long does it take until it goes to production? So this is actually what you want to measure.

But it's actually also not easy to measure it, because when do you want to start? So do you start when you get the first phone call? Or do you start when the manager accepts the request? And that's why you have this suggestion from the book Accelerate. And they say you could measure the lead time for change. And the lead time for change is basically how long does it take from commits to production. Because if you commit something, you don't add any value, you only have value if something goes to production. Because when something goes in production, then you have an outcome. And you don't have an outcome when you just commit and doesn't go to production.

So this is actually the base metric, the lead time for change. But the problem is, if you only take this metric, you can have a high risk, right when at commit and everything goes immediate to production, you can have a high risk, and maybe you can increase your failure rate. And that's why you also have to look at this other metrics, like deployment frequency, which is actually something to reduce risk, because it's actually proxy metric for batch size. So you try to reduce the batch size. But measuring the batch size is hard and that's why you use this proxy metric deployment frequency, which reduces the risk.

And you also want to look at the failure rate. So you want to reduce the failure rate. But as we know, it's really hard to really make no mistakes. And that's why you can also look at the meantime to recover. So when it's a really, really short time like say, one minute, you have failure, and after one minute fixed again, the impact is really small. So it's not that bad if you make a mistake. And that's how this idea of those four key metrics... that's actually the idea of the four key metrics. So the base metric is leads time for change. And the other metrics are supporting this, this base metric.

Jeremy: Right. And so that's, and I love those metrics because you're right, I mean, I think in the book, they talk about how you can just, you can easily manipulate a lot of these other metrics to make them seem really good. But you can't really fake things like, you know, the deployment frequency, you can't fake how long it takes to fix a problem. And the number of problems that increase, I mean, other than hiding that information, it's really hard to fake.

So how does serverless though just because of the practices that I think have been developing around serverless, how does that help with all of these metrics? Like how does it increase those metrics and or, I guess, decrease some of the metrics, but how does that enhance that experience and increase productivity?

Vadym: So generally speaking, the authors of the book also talking about software delivery and operational excellence. So what best practices do we have to gather to know and to take in place to become productive and they have identified some of them like loosely coupled architecture, which really helps to achieve this and as far as we know, serverless enforces this. Of course, you can have monolithic Lambdas but I don't think that the people doing serverless is because of this. So just loosely coupled architecture is one of the points.

The second one they are talking about code maintainability but I think they mean evolution-ability. So this evolutionary architectures, and also the serverless is also about this because you deploy this smallest unit. And of course, there are other practices that you have to consider like chaos engineering and so on to inject failures to see how your system react on failures. And there are a lot of tools around serverless which can help you.

And of course, time to restore services, you have all this kind of supporting tools for blue and green, and also cannery deployments. So you can do this with API gateway, you can do this with Lambda, with Lambda you have aliases and traffic shaping, with API gateway you have stage variables, and also you can combine this with Lambda. So you just with all these things in place, you can really, you can really decrease this time to restore service. But generally speaking, and I think you mean, how does serverless relate to this? It's probably that Simon Wardley tells us that if something changes, and then the core evolution of practices also occurs, and now we see that the serverless and Lambda execution environment, it becomes commodity.

And that means to be productive with this, you have to apply other practices, which is the same as that you can be productive, if you use NoSQL database and apply all the best practices from the relational database, it simply doesn't work. So and this is also true for every evolution of practices. So you can be successful with method of yesterday. So just in case. So that's a lot of things you have to do on your organizational level. And Christian also talked about the practices which they apply in their teams, and so on. So there's a some kind of mind shaped cultural shift that that happens.

Jeremy: Right. And I think that you mentioned evolution again. We talked about evolutionary architecture, we talked about coevolution of practices. I mean, one thing that is evolving very fast is serverless, right. And what you can do with serverless, what the features are available, what services come out. And if we look at recent launches, I mean, just the other day, we got the extensions API, where now you can tap into the lifecycle of a Lambda function, and you can run almost another function running alongside of it to do different things. We've got EFS integration, we had RDS proxy as an extension of that HTTP API's. And we just have so many of these new things that have been launched. And that's great.

And everything's getting better. Lambdas and VPC's are getting faster, I don't have to worry about all ENI cold startup time, and that sort of stuff. But where do we still need to go? Because it's always great to say that serverless can get us or serverless can get us somewhere really fast. Right? And I know, one of the things you mentioned in your presentation is the last 10% trap, right? I mean, if we're building applications that we can get to get something really fast and get almost all the way there, that last sort of 20% is super hard. And then that last 10%, oftentimes, we just have to give up and go back to something else. So how do we avoid that last 10% trap with serverless, like what else needs to be added so that we can sort of get it to be the default choice for everything that we do?

Vadym: So this 10% trap, the authors of evolutionary architecture book, they wrote the article that they compared all architectural styles, and this was the some kind of statement that serverless often suffer from this 10% trap. My personal opinion, it's probably used to be the case, but currently, it really disappeared. What is easy with serverless, if you can't solve some partition problem with serverless you can easily switch to other architectural styles, like containers and so on, you are not forced to be completely serverless. It's not that difficult to switch. And you have mentioned a lot of changes that has happened even elastic file system for measuring cloning and artificial intelligence that you can now attach one or multiple elastic file system to your Lambda and so on. VPC cold start is reduced. And other things like JS proxy, Aurora serverless, or data API, and so on. Just a lot of things has happened and even extension API, which was released recently and cloud watch Lambda insights in preview, which was released recently.

So there's a lot of things happening which will reduce this gap. But generally speaking, even cloud watch service has improved over the last year with the possibility to search in the multiple log groups in cloud watch log insight, which gives you the language to search in your log and even embedded metrics format to send your logs asynchronously. So a lot of things happening, but probably there are some more steps to go for the platform to be mature, to improve. So I really like elastic file system but I think S3 is really, really superior, because it's really very good integrated with all this events. So if the file is created or updated, you can easily call S3 and it's not possible with the elastic file system. And also, all the compliance services are deeply integrated with S3, like AWS config and so on. And all these services are important. So this is, if the people will use elastic file system within Lambda, they also have to ship also the services.

And probably very sensible topic, but I think that cloud watch, a lot of improvements, as I mentioned, but in terms of observability, and alarms, there are a lot of third-party software, the services, which are really superior, they are. So they offer really a much broader experience but of course, it costs you money. So my desire is really to stay within serverless ecosystem in AWS to use Cloud watch and not to go outside because I have to shift my data outside to think about security and so on. And just generally, I would like to stay there. And that's probably hard to improve. And of course, this situation with Xray support, it's really very important service but sometimes it's other sources, like if you've been breached them, they're missing this possibility. So many services, which are called asynchronously don't have this possibility currently.

I know that AWS ships this the same game was closed for SQL, SSNS, and so on, it takes time. But sometimes you wish that this extra functionality will become available from day one if you want to use the service because otherwise you have gaps in your observability and so on. And of course code commit. I know Christian uses code commit in his team. But this service is very, very basic. So I see that many people only use code deploy and other code, and for commit and build and so on other services. So this is something that code commit is currently not comparable, not nearly comparable to Github and Bitbucket.

You can only do basic things, but sometimes that's enough. But if people have gathered experiences with other services, then they want something more. Yeah, with more functionality.

Christian: So actually, we are using still using big Bitbucket and then a Jenkins pipeline pushes this code to code commits to make it available to code pipeline. That's how we work with code commit. Yeah, I agree with what Vadym. So when you use a framework, it can make it really fast until you reach the edge of your framework. So when something is not possible with a framework, then it can get really hard because then you have to fight your framework. And it's the same with serverless. So when you reach the edge, so something that's not possible with serverless and can get really hard. But I also think that a lot of gaps were closed in the last month and years. So it has less and less restrictions.

Jeremy: Yeah, definitely. And I totally agree on the observability stuff with you. I mean, it is hard to be tracing everything through your system and the third party tools work really well for that. It is going to be interesting to see what people do with that extensions API. And how much more insight that'll give us into the Lambda lifecycle, but again, you still have EventBridge, and these other things that still need that Xray support to kind of trace all the way through.

All right, so we didn't even get a chance to really talk about total cost of ownership. I think we did in some contexts, regarding reducing the amount of employees that you need or developers you need. But I'd like to finish up and just ask each of you to give me like, what's your top recommendation for people building serverless? Like, what's the one thing you would say to them, like "Here's the absolute thing you need to implement in your organization, if you want to go serverless."

Vadym: One thing I would say, through DevOps, so just no separation between the Devs and Ops. There is really a good page also from the authors of the team topologists. This is the DevOps topologists, and they have shown a lot of best practices, but also anti-patterns. And the best thing for the serverless is really to DevOps where the people are really working together as a team. And this is probably the hardest thing depending on how your organization currently works. So to get people there, to get the ops people there because you don't have any service, you can't install any agents. And it's something like cultural shock for them. But there is a really good talk from Tom McLaughlin, "What do we do if the server goes away?" And he explains a lot of challenges, where their Ops people can take charge with things like alerting and monitoring for the whole system.

And also, they have the better feeling about the restrictions of each service. Is SQL, SNS or EventBridge, or some kind of combination is the best solution currently for the next several years, probably. Because they understand those restrictions very well. They look into this, they did this with the storage, they did it with the database, they are really very sensible. And of course, chaos engineering can gain days. So a lot of challenges. So even Ops people don't have to be scary.

But it's a learning curve. But it's also a learning curve for the developers. Because distributed systems Microsoft is also hard for them. So just it can become a win-win situation if both parties learn and learn together, then you have really good chance, though, to embrace serverless correctly.

Jeremy: Love that. Great advice. And Christian, what about you?

Christian: So I'd say one of the most important thing is you really must automate everything. So it's okay to try something out on the AWS console but not in your production environment. In your production environment, you really must automate everything, because you have so many small parts and if you start making manual changes, you will get lost very quickly.

Right, awesome. Yeah, totally agree with that. All right. So gentlemen, thank you so much for joining me, this has been an awesome conversation. So if people want to find out more about what the two of you are doing, what ip.labs is doing, how do they get ahold of you? How do they find that stuff out?

Vadym: You can find me on Twitter and on LinkedIn, with mine first and last name. I think you will put it into the show notes, Jeremy, so because it's very difficult to pronounce my surname. So I'm really active on Twitter, tweeting and retweeting about the experiences and Christian and me we are talking at various conferences about the experience, and we will probably continue doing this because there is just so much to learn and to talk about in this community. So we will definitely be active I think.

Jeremy: Awesome. And Christian, I know you don't have a Twitter account or not a very active Twitter account, right? So I have some email addresses here, cbannes@iplabs.de and yours as well, Vadym. So I will put those in the show notes. I will also put the link to the presentation. There's a SlideShare of this, we'll see if I can find one of the videos of this presentation because it's fascinating. The topic is amazing. And I really love this idea just of where serverless can take you and where it can bring you on that productivity, sort of that productivity spectrum. So again, thank you both for being here. This was amazing.

Christian: Thank you.

Vadym: Yeah, thank you, Jeremy, for inviting us.

View Details

About Joe Emison

Joe is a serial technical co-founder, recently launching his fifth company, Branch, in 2017. His previous ventures have been BuildFax, Spacefu, BluePrince, and EphPod. Additionally, he has consulted with many other companies on software development and cloud migrations, including many in the DMGT portfolio. Joe graduated with degrees in English and Mathematics from Williams College and has a law degree from Yale Law School.

Twitter: twitter.com/JoeEmison
LinkedIn: www.linkedin.com/in/joemastersemison
Blog: emison.org

Branch Insurance: ourbranch.com

Watch this episode on YouTube: https://youtu.be/tGwGxfczaJ4

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Joe Emison. Hey, Joe. Thanks for joining me.

Joe: Hey. Thank you for having me.

Jeremy: So, you are the co-founder and CTO of Branch Insurance. I'd love it if you could tell the audience a little bit about your background and what Branch Insurance does.

Joe: Absolutely. I am a serial technical entrepreneur. Branch is my sixth adventure. Branch is a home and auto insurance company and we also sell renters and umbrella. There are two things that make us completely unique in the United States. One of them is we sell the home and auto as a bundle much more easily than anybody else. So really, anywhere else you would go to buy the home and auto bundle, you'd have to buy one and then the other, and it would take you somewhere between 45 minutes and two hours or maybe a week, and you'd get a lot of fake prices along the way. You'd get like, “Well, we think it's about this.” Okay, I need more information. Then, “It's about this.” We will sell you the home and auto together and for most people, we just need your name and address to give you a real price that you can buy instantly.

Jeremy: Awesome. Well, the thing I think that is probably the most interesting to the listeners is the fact that Branch Insurance is entirely serverless.

Joe: It is. This is actually the third company I've started fully serverlessly. Branch, from the beginning, has been built on Amazon's AppSync service. It uses Lambda, DynamoDB, CloudFront, and a bunch of other third-party services that we love.

Jeremy: Awesome. All right. I have been watching your presentations. I've seen you at Serverless Conf. I've seen you at a couple of other conferences. One of the things that I always thought was fascinating is how you're always recommending some third-party service. You've got a third-party service for everything. Now, obviously some of those are our AWS services, but then you use other things like BigQuery and Stripe, and these other ones, and all these other different things. I was trying to think of something clever. We always talk about stuff like serverless first or static first or whatever. I was thinking that your approach was very much so like third-party first, but then after talking to you about it, it's more about optimizing for maintainability than it is just using some third-party service. What is that that you do, especially at Branch, to optimize for maintainability?

Joe: We think about optimizing for maintainability having two central categories. The first is the less you have to maintain, the easier it is to maintain and so in that bucket goes things like a code is a liability, so less code to maintain is easier to maintain. Or running VMs and containers and having DevOps as a core competency that you need. If you don't need it, you run serverless or you need less of it, then that's also easier to maintain. Then there's another bucket which is how do we not be blockers for all of the other departments in the company? We constantly ask this question of how do we empower other departments to go do what they need to do without talking to tech, without talking to product.

Jeremy: I love that idea of empowering other departments. One of the things that I often see when I'm collecting links from my newsletter is people who write these posts that say they're using Google, for example, I'm sorry, Google Sheets as a backend database for something. Some people criticize it, but for the right application, it makes a lot of sense because then someone can actually go in as somebody who doesn't have the ability to query my SQL database or go into DynamoDB or something like that. Can go in and just manipulate some cells in a database and it gives them the ability to run some business process, which can be really, really effective.

Joe: It is fantastic. One of the companies I started, which I did as a weekend project for a friend, was an angular app that sends a bunch of data to a Google Sheet for all these calculations. Since I built that in 2015, 99% of the business logic of that application has been the main founder on that project making changes in Google Sheets. He onboards a new customer. He sets up their logic in Google Sheets and it does everything that they need to do.

Jeremy: Which is amazing. Then there's so many other applications that you can do that with, that make it very, very simple for … I mean that's when you think of things like maybe Airtable or even integrating things. I know in the past I've integrated into services like Asana. Just tie into the API so that you already have those interfaces built because that's one of the hardest things to do sometimes as a developer, is to build a really good admin and if that's already built-in for you, you're off to the races.

Joe: I think people too often will disregard using a service like Zendesk, which is this amazing Swiss Army knife. They'll disregard it because they'll say, “Well, I can envision a world in which we're going to want to do this thing and it won't do that properly,” but the problem is you're a lot better off getting off the ground, getting things going, having a system that works, and then discovering what things really matter because I agree with you that probably you're going to run into cases where this third-party service like SAS tool with an API doesn't completely solve what you want. The bigger problem is what you think that is at the beginning is definitely not going to be what it actually doesn't do later. We live in a world where we don't have enough people who are constantly thinking.

More than 50% of the things that we think we need to build in the way we're going to build them, actually aren't going to work. We're going to build them. We're going to put them out there and they're not going to be usable. They're not going to do what we need to do. We're wrong about these things. It's so much better to get it out, experience it, even if you know that it's only 60, 70% ideal because that information that you'll get will enable you to make much better decisions. Often, what you'll find is when you weigh the priorities, you'll say, “Well, Zendesk is imperfect, but I'd much prefer to keep using Zendesk and use my development time on this other thing that I now realize is critical for our success.”

Jeremy: That's one of the things too that I love about serverless, just being able to string some things together quickly to get a prototype up and running because you do not know actually, never mind what's going to work, but what's going to be used. I can't tell you how many times I've implemented some process or some system that just doesn't get adopted by the organization because either the workflow is too complicated or whatever, but again, being able to test that quickly is a really powerful thing.

Joe: Absolutely.

Jeremy: One of the problems when you are building new applications and you're trying to, as you said, optimize for maintainability, one of the things you need to maintain or what you need to perform maintenance, I guess, is people, so you need to have engineers or software devs that can go back and they can look at that code. They can continue to upgrade it if they need to. How do you make those choices? I know you have a software development guide at Branch. You put everybody through those paces. What does that guide look like? You had mentioned optimizing for maintainability, but what else is part of that?

Joe: Really, it starts off completely on optimizing for maintainability and understanding that when you have multiple choices to make in the world, you choose the one that's going to be easier to maintain, even if it disagrees with whoever luminary on Twitter says to do. Recently, we had a discussion about whether the right way to convert strings to numbers in JavaScript is I guess the preferred way as a plus sign in front of the variable name, but you can just use the number function which is just much more readable. We were like, I don't care if so-and-so says you should use the plus sign. Number and then in parenthesis is a much easier way to read it. Then the rest of the document, the rest of the software development guide really talks about two things.

One of them is we practice Agile, I think Agile in the way that I have heard Alistair Cockburn, who's one of the signatories of the Agile Manifesto, thinks about which is every week we come in peace to ship software and every other week we do a retro to figure out how and do it better and nothing else is sort of set. You talk to a lot of people about Agile and they'll say, “Oh, well, we're Agile. We do Scrum and every two weeks we do a release, and we do a demo on this day.” Literally, number one is "people over processes." If you just describe your Agile as a bunch of processes that aren't changing with ceremonies at the same time, that's not Agile. This is process over people. Part of the software development guide, after maintainability, really goes into helping understand like, “We're trying to figure out how to do this better.”

We don't know how to do it or we have an idea on how to do it, but we're not saying we're perfect on it and so we need everyone's input. We use … Retro is such a critical thing to do to figure out how to do all these things better. Then the rest that we talk about in the guide is really about how what's most important for you as a developer is to think about how you could be doing things better, how you could be shipping software better, what tools you could be using to be more effective. We don't do enough of this. We actually hire people into jobs like Oracle DBA. If you ever wanted to get rid of Oracle, you're probably going to have to get rid of all the people who have the vendor's name and their title first.

We have this problem, where people ... Ben Keough talks about this, where people identify as experts in the tools they use and that's really bad. You should identify in being able to ship software. It shouldn't really … You should identify as someone who can learn things. This is the most common debate I'm having over Twitter these days, really, where people want to say, “The right language is the language you know.” I'm like, “Okay, but what that says is you can't learn.” The reality is if your job is figuring how to ship software and you're like, “I'm never going to learn another language or another framework. I'm only going to use the stuff I know,” you're going to be bad at it.

I'm sorry. You might be good at a college, but, oh my God, if you're still using the thing you learned in college and you're 35 or you're 40, you're not doing the best thing. It's just true. The rest of what we talk about in this guide is like, “It's important to learn and it's important to treat people very kindly and with respect.” There's a lot of discussion, actually about how to do a proper code review, like, “This is how you review someone's code.” This is how we make sure we have great code, but we don't make people feel bad.

Jeremy: I think that's something too where, again, as people grow as developers and they start using new things, it's very, very easy, especially moving into serverless. There's a lot of information out there, but it's a whole different way to think about software development, so it's very easy for people to make mistakes or to do something wrong. And again, having a team that doesn't criticize, but just helps you, you know what I mean, and helps you ship better software, I think is a hugely important piece of it.

Joe: I think about it a bit as, how do you get to a world where everyone is relentlessly trying to do things better and is dissatisfied or is willing to be dissatisfied with what they just did? How do you do that in a world where everyone respects everyone and treats everyone as humans who are great and wonderful people? How do you separate those two things? That's magic, when you can do that.

Jeremy: Well, certainly not on Hacker News. I can tell you that.

Joe: No.

Jeremy: All right. With this idea of optimizing for maintainability, you mentioned this idea, the code is a liability: the more things you write, the more code has to be maintained, the more that has to be upgraded later, maybe or whatever. Of course, you have potential issues with using third-party tools. They might change something. Maybe there's a breaking change or whatever and you have to go back and change your code, but just in terms of that technical debt that builds up, how do you tell your developers to think about technical debt as they're building something?

Joe: We have no roadmap, essentially no deadlines, and unlimited refactoring. We tend not to do a lot of … We do have some like, “Hey. We should refactor this when we get back to it,” but there's a lot ... It is a very friendly environment in which if you, as a developer, are going to work on something and you say, “Hey. I really think I need to spend a couple of weeks refactoring something before I get into this,” we're never going to say no to that. Now, a lot of times there'll be a discussion about the best way to do that and that doesn't happen with any sort of high frequency, but it's not infrequent that you'll get, “Okay, now that I'm going to tackle this thing, I'm going to go ahead and do this refactor as part of it.”

And I think that's really it. Again, this is maintainability. We say, “We optimize for maintainability.” If we don't refactor things that we think need to be refactored, it's not going to be maintainable. I was just going to touch briefly on the … You mentioned third-party breaking changes. My experience is that that doesn't happen with a high frequency. I tend to find that the third-party services, either APIs or SAS, that we use tend to be way better at change management than we would be internally.

Jeremy: That's right.

Joe: It tends to be. I tend to think of all of them as, what if I had a bunch of super-expensive engineers who were really good, documented things really well, had amazing change management, amazing uptime, and I only paid them by use of the thing in a business context? Wouldn't that be great? Shouldn't I do that for everything? So I tend not to have … It happens from time to time, but I never … Anybody who is telling me, “We would never use Algolia over Elasticsearch because we're worried about breaking changes.” That person is just making that up. That's not a problem.

Jeremy: Well, I think I would worry less about breaking changes, but maybe new features that are added, things like that that you're going to have to go and potentially do that, but I totally agree with you. That's one of the things that's nice about having a SAS company, is they are focused on that one specific problem, that business domain, and that's their job. Whereas when you have a team of developers, especially if you're a small team, just maybe four or five of you or something working on a project, if you're building all this stuff yourself, you are across a lot of different disciplines. With the third parties though, obviously, if I said, “Okay, should I have Algolia or should I use Elasticsearch? Should I build my own search there or should I just use Algolia?" You say, “Okay. That is the easier choice because it is easier to maintain that.” Obviously, third-party services are going to be easier to maintain. What is the deciding factor? Because why wouldn't you just choose third-party services for everything?

Joe: In general, I do choose third-party services for everything. My general view is, prove to me that this third-party service won't work. Now, again, I have a very strong difference though between a third-party service that's serverless and one that isn't. You can find third-party services where they want you to go into the AWS marketplace and run it on a VM. That's not serverless and I'm not interested in that. Or like, “Oh, it's an open-source project. Run it yourself.” Again, I'm not interested in that, but when it's serverless … My short definition of serverless is it's not my uptime. I literally can't influence uptime. Beyond, I could put bad configuration or a bad code in, but if some server fails, it's not on me to bring it back up. I think if you can have a serverless third-party API, I think your default should be to use that unless you can prove that you shouldn't use it.

Jeremy: All right. Then what about proof of maintainability? We can say, “Oh, it's easier to maintain this code because I'm using some third-party service or it's easier to maintain because again, it's an Algolia versus Elasticsearch.” How do you actually prove that, though?

Joe: Well, the easiest way to prove maintainability of your codebase, does it have terrible technical debt or do you have terrible other issues with it or some issues with it, I think there's two things that you can do all the time. One, if every developer has his or her own Amazon account or whatever, provided you're using accounts, you can check out the code and run a command and deploy it into a fresh environment, and it works just the same. Everyone has production. You can onboard a new developer, ideally a junior developer and they can start being productive as a member of the team very quickly. I don't know. Within, say six weeks, they can understand front end, backend, how to commit infrastructure as code. They get all of those things. You give them tickets. They can work in them.

That's a maintainable code. I think part of your question was also like, “How would you know if, let's say you used a really bad third-party service and it was really terrible?” I think it's, again pretty easy to ask what percentage of your downtime or what percentage of bugs do you have that are relating to these things? You'll know. If there's a lot of pain, you'll know. Back in 2012, at a prior company, we used a recurring billing service. At this point in time with third-party services, we would do random testing where we would just hit their development endpoints randomly. This service would return 500 errors about every fifth day out of their dev service. The lead developer was like, “I don't feel comfortable using this service.”

We contacted them and they said, “It's totally normal for that to happen,” and we were like, “We totally disagree.” A 500 server should be a defect. We were a really squeaking wheel about it. They had a couple of outages. We complained about those and we got fired as a customer, for bringing that up. I woke up one September morning to get an email. It was like, “You have 30 days to get off our service because you're an obnoxious twit.” The joke was on them because six months later, due to poor engineering, they lost all of their customers' data.

Their whole business was a recurring billing service that was supposed to store all the credit card information so you didn't have to be PCI compliant. Who knows? I don't know what was causing the 500 server error. They had a poor database replication design where they replicated an error in the master. But I've always relied on that at this point. If you do regular testing and you understand the services you use, you can get a sense of which ones are good or not if you're actually using them. My past is a history of services I will never use again because of bad experiences with them, but we found it. We discovered it.

Jeremy: That's a risk too, though. I mean as soon as you use a third-party service, a lot of that control is out of your hands. How do you mitigate against that? When you're integrating with third-party services, do you say, “How easily can we swap this out if we decide to go in a different direction?”

Joe: Yeah. You can do that or you can load test it and availability test it, and see how long it's been in business, and go look at … All of these services that are any good have status pages with lots of history, have customers who are references and so I don't think it's that hard to vet these, to figure out if they would be okay. I don't know. In the high-risk category, I do think … I'm a Fortune 500 company. I'm about to send a million transactions a month through this service that just got out of Y Combinator last month and it doesn't have a website? I wouldn't do that. I think there are standard vetting things you could do. I think a lot of these services, I would certainly recommend to anyone. Algolia or Cloudinary or Auth0, these services have been around for years and years and years. They're big.

Jeremy: If you did have to swap out a service or … Even not that. Maybe you just say, “Look, we need to upgrade some piece of the system. Something's a little bit out of date.” Maybe there's a new feature that's added. You want to go in and refactor. Or maybe you just have something that's not running efficiently and you want to refactor. You've mentioned refactoring a number of times. Now, I know … I've had a CEO in the past. I've had some interesting CEOs that I've worked for.

One of them who still loves the idea of counting lines of code as if that was some sort of an asset and not a liability. Then I had another one who would say that if we leave engineers to their own devices and we don't give them specific projects, that they would just refactor all day. Obviously, you can't refactor all day. You need to build some new things every once in a while, but from your point of view, refactoring is a really important part of the software development cycle.

Joe: Absolutely. Look, I think there's something that's higher level in this conversation, which is that if you have good technical leadership in your organization, your technical leaders will be able to … I mean what they're trying to is get enough trust from the rest of the organization so that you can do things the way you think is best. All of these things, any sentence that starts like, “If you left the engineers to themselves,” is just an example that there's no leader who's managing the trust in the organization across the organization. At Branch, we do not have a roadmap. We do not have deadlines. I mean people can ask like, “When do you think this thing is going to be done?” There aren't company-wide burn-up and burn-down charts. We do size things, but there isn't what day or is everything that everyone's working on going to be done on?

We don't do that. We don't do that because it takes time. Would you rather have it done sooner or would you rather have estimates that are wrong? That's the reality. That's the choice that you're making. It's the only way that you get the rest of the organization to say, “Okay. I'm not going to ask you for a roadmap.” I ran into this a lot at Branch because we've hired a lot of people who've worked at very large organizations than Branch. They're like, “Well, I want to see the roadmap.” I'm like, “Well, there is no roadmap” and they're like, “What do you mean? You have to have a roadmap.” We sit down and we talk and we say, “Well, if I had a roadmap, if you really needed something done, it's going to go at the end of roadmap.” I'm sure that's what you experienced. You ask for something, how long will it take?

“Well, in my last company, it would take nine months to get it done.” I'm like, “Okay. We'll get it done in two or three weeks, but the trade-off for that is there's a roadmap. Are you okay with the trade-off, where you can get stuff done in two to three weeks instead of nine months, but there's no roadmap? Do you really want a roadmap more than you want things done eight months early?” People are like, "Hhmm." They're not sold on it, like, “I don't know. This is new, ” but if you can do this, if you can gain the trust by delivering software … This is why our Agile is we come every week to deliver software. This is also a lot of the accountability of developers. A lot of what serverless gives you is developers have control about, “We're going to deliver this stuff.” We don't have a DevOps team that blocks things.

We don't have a DevSecOps team that blocks things. We don't have an operation team that blocks things. We build it. We have automated tests. We have some manual testing people who are asynchronous and services that are asynchronous, that will look at things occasionally. We have smoke tests. We put things live right. Everybody knows sometimes there will be bugs. Sometimes you can only catch these bugs in production and we will fix them quickly. We will shift things quickly for you. In that world then, if you live in that world of trust like that, then you get to do what you want to do on the tech side.

I totally agree with you. Developers don't want to refactor all day, but developers also don't want somebody breathing down their neck, trying to make them into a feature factory. That they have no say in what's being done and they're not allowed to clean up their room, basically. I mean this is what refactoring is. It's like, “This is bad." Some new features came out in Reapp, let's say they'll let us use them, whatever they are. This is the crazy thing about it. This is no different from management in general, where you try to empower people to have control over their space so that they can be more effective and happier. That's all it is, really.

Jeremy: You mentioned three things that all tie something together for me. One of those was tests. One of those was trust and the other one was feature factory. I had a consulting client. I went into this consulting client. They had maybe 10% code coverage and the CEO and the COO did not trust the CTO because they had recently had a number of failures where they deployed something and the database broke. There was a database change in there that was in staging, but didn't make it to production. Standard stuff happens if you don't have a good process in place. I go in and I'm looking at this test coverage. I said to the CEO, I said, “You need to give the development team time to add test coverage and add some additional tests in so that you can be confident, when we launch a new update, that we're not going to break something and then people are going to get mad,” but the important thing was they wanted to pump out features.

They had a ton of different clients that needed very specific things. It's like you quickly went from being a software company to being a custom web development company, it almost seemed like, but again, those things, where not giving your team time to build the things they need to make sure that the process is solid, that's going to erode trust. Then when you don't have the trust and then you keep pushing them to deliver feature after feature, after feature, without the time to go back and clean those up, without the time to ensure that everything is automated, that's a huge recipe for disaster. Honestly, I don't know many companies that I've consulted for or worked for, that haven't seen that type of vicious cycle over and over again.

Joe: I agree. I think the vast majority of companies that do software development do it very badly. Another way to say this is, I think the gap between the top 1 to 5% of organizations in software development are dramatically … There's a chasm between them and the average. Absolutely.

Jeremy: As a co-founding CTO or as a CTO, you've done this a number of different times. You mentioned things like no roadmaps, giving users, or giving your developers time to refactor, things like no deadlines, no burn-downs, a couple of notes I have here. What else? What else do you do? If you're running a company, which you are, but if you are running that, if you're advising someone else to manage their developers and their development team, what are those things that we have to change, I think maybe as a culture so that we can, again, build up that trust, get to a point where you're delivering good software, but also not stressing people out.

This idea of developers working 70, 80 hours a week, hey, that was great in my 20s. I did that in my 20s and my early 30s. I don't want to have to do that anymore. I just want to be efficient. I want people to know that they can trust me to work on something and that I'll get it done. What's that advice? What does it need to look like? What does a modern development department need to look like, a modern engineering department?

Joe: I love this question. I feel like this question isn't talked about enough. A huge problem with this question and the answer is that people want to find excuses for how they do things today. That's all they want to do. They also want to … Nobody wants to make judgements in this space. They just want to say, “The way we do it is different and that would never work here.” Now, I will say I've come in and consulted with many organizations and my results, consulting with them, have ranged anywhere from utter failure where I failed to affect any useful change to ones where I think I was impactful, but it took years and years to have really good, effective change. I don't think that it is easy at all to go in and make changes and I certainly don't claim to be that good at it. I think the main thing that I would say to anyone, as a starting point, is that we now have the best book ever written about this, which is Accelerate by Nicole Forsgren.

Jeremy: A great book.

Joe: It lays out four metrics and it gives you good ranges for those metrics. I do think it is a fantastic way to manage organizations if there are good metrics to manage them to metrics. In marketing organizations, sales organizations, finance organizations, managing the metrics works and they're happy to do it. I think and I believe that the four metrics in Accelerate are four really excellent metrics and I think they're a great way to align. If you could get … The last few times I've been asked this, I've said to the CEO, “You get on board with these metrics. Let's the CTO on board with these metrics and let's optimize to the metrics.” Then that will at least allow … That empowers the technical teams to figure out how they're going to get there.

Those metrics are, you need to be releasing multiple times a week. You need to resolve errors quickly. You need to have a low rate of failure and low rate is just less than 50%. They give you these metrics. I do actually think that the most important one is deploying multiple times a week. I think that if you're not deploying multiple times a week, you should be. I also think that you should be trying to get to a point where Friday deploys don't scare you or at least, I'll leave it this way, you should be deploying frequently enough and have good enough testing so that most of your bug reports are not in the thing you just deployed.

In my experience at Branch, most of the bugs we find are bugs that have been in the code for a couple of releases before they're found, even regressions, and so because of that, you don't need to fear the Friday deploy because it's already in the code. The next bug you find is already there. Also, you're deploying such a small thing that it's not that big of a deal. I'm totally sympathetic to, maybe you shouldn't deploy something right before you go to sleep or something. I'm not opposed to that. We don't do automated deploys at Branch. We deploy several times a week, but I think we've deployed two or three times today, for example, totally fine. That's where I would start from. That book is fantastic. Everyone should read it. It's really important if you're trying to make an organization run better.

Jeremy: At Branch, part of your development process, I think it's really interesting because again, you try not to overwhelm your engineers. There's no working nights and weekends. I'm sure if maybe there's something critical, then maybe that would happen, but what about using ticketing systems? How do you use ticketing systems at Branch?

Joe: I really don't like Jira. I've used Jira a lot in my past. There's a really good article by John Evans, I think he's at Rezende, that he wrote in TechCrunch. He wrote about how Jira is terrible. When I first read it, I was like, “He's totally wrong,” and then over time, I've come to think that it's one of the better articles on it. We use GitHub projects and issues. This is a classic case of where we started using it in a certain way and it didn't work that well. We used retro. There was a lot of like, “Well, maybe we should use Jira” and I was like, “Okay. Just prove to me that this won't work. I am down with using …”

We got into the linear, beta, and stuff like that, but I was like, “I really don't want to use Jira, but I will use Jira if guys say to use Jira. Just ... force it.” I actually hired a PM who had never been a PM before because I didn't want to bring … I generally have this thing. I don't like senior developers or senior PMs because in general, I find that all they're going to do is bring whatever process they had instead of looking at what's great about what we do. I know it's not all senior people everywhere, but I've experienced that in my career. Senior people don't look at what you're using. They're just like, “We're going to change over to however I did it the last time.”

So, I brought in a PM that had never been a PM and I was like, "Look, you're going to use retro." We're going to use retro to figure out what we don't like about what we're doing and we're going to try to fix it. We're going to try to use GitHub and GitHub projects and issues and if it doesn't work, we'll know why. We'll know the limitations. We're like iteration, I don't know, six and it works really well. We have a super-functional way of using GitHub projects and issues to do everything we need to do, look at statuses, everything.

Jeremy: You mentioned that you don't do sprints and burn-down charts, and some of those other things. When somebody gets a ticket, how does that work? Do you assign someone a ticket? Do they get a ticket themselves? Do they take multiple tickets? Are they dumped in a batch and say, “Get these done by the end of the week?” Again, you said, “No deadlines,” but how does that work?

Joe: One thing that I'll add is that the thing that I don't like about Accelerate is that it is only concerned with work being done. When the work is done, the development work is done. How does that get live and stable? That's Accelerate. There's an entire other book to be written on ... there is a business problem that needs to be solved with custom software development or something by the team that does software development. How do you get to the code is done, that work is done? It's useful to talk about it. The process that we use for this is we really require the business to state problems, not solutions. You can give solutions as examples. Again, this is an example of we have trust, so we get to define the process. When the trust breaks down, we lose control over the process, but our process is you help us understand the problem you're trying to solve and potential solutions to it, maybe if you want to give them.

That goes to our designers. Our designers are going to come back, and multiple times a week, there's a designer review meeting, and you're going to get to go look at the designs and you can approve them or not. You don't get to bring designs. As a business stakeholder, you can't go say, “I need a button here and it's this color. It needs to say this.” You don't get to say that. You can say, “I want to solve this problem and one way to solve the problem might be to put a button here in this code or whatever,” but eventually, you learn it's not even worth telling them that because they're going to design what they want to design. I'm just going to tell you … I'm going to get really good, which is what you want the business stakeholders to do. You want them to get really good at helping you explain what problem they're trying to solve.

Then the designers deliver designs. Once the design is accepted, it could be the first design. It could be the 85th design. The process is the designers get to propose whatever, but the business stakeholders accept the design. Either they accept them or they don't and sometimes it gets a little tense. It's okay. It works, as long as you learn how to disagree with people, which is a skill everyone needs to have. Work through disagreements. Once the design is set, our designers design and sketch. They use Envision for interactions. They export to Zeplin. Zeplin is amazing because it gives you all of the margins and fonts and things. You don't have to eyeball something and try to figure out what the CSS should be, but when a developer gets a task, that ticket has the designs.

They just work the designs and they've probably been … Actually before they get it, there are weekly, basically sprint planning meetings, except there's no sprints. Every week, we try to go and size and talk about all of the tickets that might be worked by someone before the next weekly meeting in which we will do the same thing again. That's all. It's not that … We've got a method that's working. It's Kanban-E, but we steal the ceremonies we need, but this is Agile. We're just trying to figure out what's the best way to get what we want. Then anybody who, … Then what's great about retro is concerns with things, maybe like in a demo, sometimes the business stakeholders seem like they're shocked by what they got. What happened?

It seems like they should have accepted that design. You'll get those. Then that prompts somebody to think about, how could I revise this process? How could I clean this up? Was that a one-off? Those are the types of things that we have in retro, as opposed to retro being … I think when we started retro and when GitHub projects wasn't fully working, it was a lot like, “I don't know what I'm supposed to be working on. I did this thing. I did it and totally wasn't set.” People were like, “That's not what we want at all.” As a developer, you do a lot of work. You throw it away. We don't actually do that. I would say today, most of the work gets done. It's done. It's right. It goes out the door. It's small in scope. We watch it and then it iterates on it.

Jeremy: That's what I was going to ask you about in terms of scoping. One of the biggest problems I've seen when you use a ticketing system is someone will say something. There'll be a job and it'll say, “Move this button five pixels to the right or rename this button,” something simple, small scope. Then the other one would be like, “Build a new billing system,” a very large scope project. What's your process in terms of down-scoping things to get to the right level of work so you don't have an engineer stuck in a back room for three months?

Joe: There are a bunch of tactics here. One of them is when you have trust, you can convince your stakeholders to down-scope themselves. At Branch, we have a saying. We actually have our "Branch Roots," which are the principles, cultural principles that we have, that we say a lot. One of them is, “What's the V1?” This is like endemic language throughout Branch, but essentially, somebody comes in and we ask, “Well, is this enough for the V1?” Everyone's on board with this because they understand that, one, it'll get out faster and two, when they want to make changes to … We all recognize you're going to want to make changes to it. This is the delightful thing about everyone bought into, “What's the V1?” If you can get that, that helps to down-scope. Of course, sometimes the V1 is large.

I think it's very important. One of the rules that we have is that anything large, you've got to build a release schedule and you got to release at least every two weeks, preferably every week. We will have branches that live for two weeks, let's say. Now, we run a monorepo and we do a lot of … We use GitTown, which is an open-source thing that makes it easy to … I mean it's like any of those Gitflow or whatever, but it allows you to merge trunk to branch very easily. Obviously, we have to resolve some merge complex. We have a couple of different teams. We do our best to make sure that there aren't multiple teams working on the same code at the same time.

Then we have people collaborating. You do some sort of planning so you don't collide into everybody. We use feature flags to release things so that they're out there, but not live. I think there are a bunch of tactics there, but in general, that's the strategy. I think another thing that we usually do is if you are working on something that's going to be a week or two that you put at least two developers on it, they'll pair, they'll split up the work or whatever. That tends to be very good in a lot of different ways, to have that extra number of people if somebody gets sick. Then also to have focus. "We care a lot about this. It's big and so we're devoting a lot of time to it."

Jeremy: It sounds like that provides a lot of transparency where everyone can look. Those business stakeholders, they can go in and they can say, “This project is 50% done.” Or it's in someone's queue or whatever. How do you push back against business stakeholders, though, or the stakeholders when they do push something like, “Oh, we really need this feature? I really need you to prioritize that.” Is that something where …? Again, I know the trust is a big piece there, but as the CTO, I don't know if you have engineering managers, is that something that you really push back on the business side of things?

Joe: No. This is the benefit of not having a roadmap. If you're like, “I really need this,” then sometimes there's a conversation with the other business stakeholders. I always find that transparency solves so many of these problems. I've worked with so many CTOs who like to hide information. I was consulting for this CTO, he was, like, “I really don't like to let my vendors know everything that's going on.” I was like, “You're an idiot. Tell everybody everything and if they're bad people, don't work with them.” It's crazy to hide information. If someone's like, “I really need this,” then I go make sure that the other stakeholders are okay with us moving that ahead. One rule we tend to have is I won't interrupt a developer working on anything, so the best I'm ever going to do for you is put it up next.

It's going to go up next for the appropriate developer or developers, which could be lots of them. In some cases, that might be like, “There's one person who should do this” and so you're just going to have to wait until they're ready. I think very occasionally, for each developer, maybe once every six months or something, we might interrupt them. We might say, “Hey. Can you just put down what you're working on? We really got to fix this thing.” Usually, it would be a bug, but that's very infrequent. Everybody gets that, but in general, everyone's primed. I got to tell you. If you're a business stakeholder and you know that something you really care about can be done in about a week, if it needs designs maybe two weeks, and if you've worked anywhere else at any other organization, you're like, “I'll take that. I don't know what other weird processes you have, but that's fast.”

Jeremy: I don't know. I think that's just amazing, that idea of having business stakeholders that understand what the processes are within the engineering organization and having that leadership. I think that's the most important thing, having the leadership within that engineering organization to really go to battle for your engineers. Build up that trust. Have that history of delivering things. Of course, I think it's a fantasy world in a lot of organizations. I think they can get there, but I think the vast majority of organizations aren't quite at that level yet.

Joe: The one organization that I really helped go from no ticketing system, all tickets on whiteboards, terrible communication, business stakeholders walking into the developers' offices and sitting over their shoulder while they worked on stuff. It was no testing, a million and a half lines of PHP in the application. One page took seven minutes to load on average, stuff like that in a SAS product that lots of people were using. I mean it was not in a good place. Now, three years later, we brought in a consulting firm that was really good, that really collaborates with you, but everybody there was like, “I want help.” I know that the biggest challenge to fixing and making an engineering organization go from, let's say not compliant with the Accelerate metrics to compliant, is most of the people in that organization saying, “I want help. I know it's not good. I'm willing to listen to people who aren't me.”

Until you get to that point, obviously, you're not going to change anything. There's so much around trust here that matters. This thing … It really helps to be the technical co-founder. It really helps to be one of the first two people in the company on day one and so you predate everybody else. That helps a lot, but that doesn't mean … I've held lots of meetings in which I have explained like, “This is how we do things. What questions do you have?” People are going, “I don't get why you do this. Why can't I do this thing?” Explaining like, “Well, you're right.”

If you look at it from that perspective, it seems bad. Here's why. One of the concerns that I get a lot is, “Because you won't let me specify designs, I'm reliant on those designers to come back. They get to make whatever designs they want. I could be waiting forever for something.” I was like, “But have you? Has this happened in practice?” Yes, you're right, it's a theoretical concern. In the past two or three years, I've probably spent fifteen to twenty hours, not a ton, but in meetings, just asking people to vent about what they didn't like about it and just explaining why we did what we did or why we do what we do.

Jeremy: Well, that's why I like junior devs because you can just bend them to your will usually and say, “No. This is happening.”

Joe: It's awesome. It's so awesome to take someone who has not had a lot of experience. They come in and they're like, “Well, everyone's happy here. We only work 40 hours a week. No deadlines. I get to just work. I'm not in a lot of meetings every week. Okay. I'll do it, whatever the hell you tell me to do.”

Jeremy: Exactly. All right. Awesome. I just want to ask you one more question because this is something I am always interested to see and again, because you use so many third-party services and you've been fully serverless, like you said, for a couple of companies now. Where is serverless going? This is a cliché question, but five years from now, is serverless just going to be part of the cloud? I mean are we just going to not even think about it? Is there going to be a distinction between what we consider serverless and what we don't?

All of these major cloud providers are creating these Kubernetes management systems where you're not even going to manage Kubernetes anymore. It's just going to be Kubernetes, but you're not technically going to be managing Kubernetes. You're just going to be using K8 then on top of it or whatever it is. Is it going to go away? Is it going to abstract away? Are we getting close to that, anyways? We're just these public cloud providers that are going to take over the management and serverless is just going to be the way forward?

Joe: No. I don't think so. One, never underestimate an engineer's desire to keep doing things the way they have been doing them or to make iterative improvements. If you were managing VMs with Terraform and your job was to manage VMs with Terraform … This is the problem. I talked about the Oracle DBA. I think Simon Wardley talks about this a bit, but we now have a whole … The new Oracle DBA is probably DevOps, so you just built up a big DevOps team. Well, I don't think you're leaving containers or VMs until that team is not and your prod infrastructure runs on those, so you're in a bad place. If you hired me as a CTO in an organization, you dropped me in and all your prod was running on Kubernetes, and three people understand how it's set up and it's brittle in places, that's a tough place to be in.

You're going to be there for many years. A lot of people are jumping into that right now. I think VMware is out there and has a lot of incentive to drive that and is a very competent organization that will continue to proliferate that. I also think there's a huge debate about what is serverless and what are good serverless architectures. When I go out and talk about serverless, there are a lot of senior devs in the Hacker News world who are like, “Well, you just stitched a bunch of APIs together and I don't want to do that.” I live in a world where all … At Branch, we only hire front end developers. I went out to try to see, can I find a senior developer I wanted?

I advertised for a software development leader and 100% of the resumes … I said, “You need to have a year of experience, not necessarily in your job, but just a year of experience developing a modern JavaScript framework.” No applications that I got had that. I was like, “Okay.” All the front-end developers are selecting out and all of the backend developers are selecting in and they're not qualified. I really believe in this full-stack front-end developer. I really believe in stitching these API APIs together. I believe very strongly in all of those things, but I don't think that there is a lot of acceptance on that at all. You asked me about the future. I think the future is a service that is like Draftbit. I think this is … I don't really believe in no code for interfaces because I think in the end, you're going to want to customize the interface. That is the one thing everyone is going to want to do and that is a place which is going to pay off in every organization, to customize the interface to your specifics.

The idea you can do that through drag and drop, I haven't seen it. Draftbit is this interesting product that lets you almost drag and drop and build a React Native app, but it compiles to React Native, so you can inject it to React Native. You can build it there and when you reach a point where you're like, “I got to customize this more,” learn React Native. Then you keep going. Another way to say this is if you look at the future, the future is just better, higher-level abstractions that compile down to lower-level functional code that you can work with because if you build the version that you can't compile down to code, but it's totally proprietary, which is the Airtables of the world, what you have is fundamentally proprietary. Until you give me the ability to do extensive modification in code, I'm going to have to leave you.

I'm going to outgrow you or you're going to be an internal tool. I think this is the future. I think AppSync built on what Firebase did, it's a great high-level obstruction, but you have full access to customize everything you'd care to customize about it. EventBridge, another version of this that's fantastic. We have this great history of these things. We're seeing them on the front end more and more. That's what I think the future of highly effective development is, but I think the challenge is that 80 plus percent of organizations are going to have a bunch of people working for them who are going to resist this tooth and nail because it means they will have to throw out what they have invested so much in today.

Jeremy: I totally agree. That's amazing. Well, listen, Joe. Thank you so much for chatting with me and sharing your immense knowledge on not just serverless, but also just managing an organization. Honestly, if there's any, I guess technical leader out there right now, read Accelerate. Listen to the advice Joe just gave. Push back on your CEO. It's not about lines of code. It's not about a feature factory. It's about having a healthy, productive, trusting, transparent organization where everyone is working together to, just like you said, make themselves better. If people want to contact you, how do they do that?

Joe: Twitter, @JoeEmison.

Jeremy: All right. Then you have a blog too, emison.org, right?

Joe: Yup.

Jeremy: Then Branch Insurance is just ourbranch.com?

Joe: That's right.

Jeremy: Awesome. All right. Well, I will get all that in the show notes. Thanks again.

Joe: Thank you.

This episode is sponsored by New Relic and Epsagon.

View Details

About Mark Nunnikhoven

Mark Nunnikhoven explores the impact of technology on individuals, organizations, and communities through the lens of privacy and security. Asking the question, "How can we better protect our information?" Mark studies the world of cybercrime to better understand the risks and threats to our digital world. As the Vice President of Cloud Research at Trend Micro, a long time Amazon Web Services Advanced Technology Partner and provider of security tools for the AWS Cloud, Mark uses that knowledge to help organizations around the world modernize their security practices by taking advantage of the power of the AWS Cloud. With a strong focus on automation, he helps bridge the gap between DevOps and traditional security through his writing, speaking, teaching, and by engaging with the AWS community.

Twitter: https://twitter.com/marknca
Personal website: https://markn.ca/
Trend Micro website: https://www.trendmicro.com/
Watch this episode on YouTube: https://youtu.be/QXZT-DQwGk0

Transcript:

Jeremy: Yeah. So you mentioned two separate things. You mentioned compliance and you mentioned sort of legality or the legal aspect of things. So let's start with compliance for a second. So you mentioned PCI, but there are other compliances there's SOC 2 and ISO 9001 and 27001 and things like that. All things that I only know briefly, but they're not really legal standards. Right? They're more of this idea of certifications. And some of them aren't even really certifications. They're more just like saying here, we're saying we follow all these rules. So there's a whole bunch of them. And again, I think what, ISO 27018 is about personal data protection and some of these things, and their rules that they follow. So I think these are really good standards to have and to be in place. So what do we get... Because you said, you have to make sure that your underlying infrastructure, has the compliance that's required. So what types of compliance are we getting with the services from AWS and Google and Azure and that sort of stuff?

Mark: Yeah. So there's two ways to look at compliance... Well, there's three ways. Compliance you can look at as an easy way to go to sleep if you're having troubles, just read any one of those documents in you're out like a light. And then the other two ways to look at it are, a way of verifying the shared responsibility model, and then a way of doing business in certain areas. So we'll tackle the first one because it's easiest. So us as builders, building on GCP or Azure or AWS or any of the clouds, and they all have in their trust centers or in their shared responsibility page, they will show you in their compliance center, all the logos of the compliance frameworks that they adhere to. And what that means is that the compliance organization has said, you need to do the following things. You need to encrypt data at rest or encrypt data in transit.

You need to follow the principle of least privilege. You need to reduce your support infrastructure like here are all the good things you need to do. And what the certifications from the cloud providers mean, is that they've had an audit firm. So one of the big five, Ernst & Young or Deloitte, come in and audit how they run the service. So Azure saying that, Hey, we are PCI compliant for virtual machines means that they are meeting all the requirements that PCI has laid out to properly secure their infrastructure. So that as a builder means that we know they are doing certain things in the background because we're never going to get a tour. We're never going to get the inside scoop of how they do updates and they do patching. And frankly, we shouldn't care. That's the advantage of the cloud. Right?

Is like, it's your problem, not mine, that's what I'm paying you for. So compliance lets us verify that they're holding up their end of the bargain. So that's a huge win for everybody building in the cloud, whether or not you understand the mountain of compliance frameworks, the big ones are basically PCI, 27001 from ISO is basically just general IT security. We don't set our passwords to password, that kind of stuff it's basic hygiene. And then the SOC stuff is around running efficient data centers. Right? So it's like we don't let Joe wander in from the street and pull plugs. We have a process for that kind of stuff so great there. And the others are, if you're in a specific line of business. So if you're in the United States and you're doing business with the government, you need a cloud provider that is FedRAMP certified. Right?

Because that is the government has said, if you want to do business with us, here's the standard you need to meet. Therefore, FedRAMP is this thing that vendors and service providers can adhere to, which means they meet the government's requirements to do that. And most of these are set up like that. So even PCI is a combination of the big credit card processors. They've formed this third party organization that said, anybody who wants to do business with us, so anybody who wants to take credit cards needs to adhere to these rules. If you don't take credit cards, you don't care about rules. So, that's the different way of looking at compliance. So it's very case by case. If we're building a gaming company, if we're taking in-app transactions like, Fortnite through the App Store, that's a huge bonus they get, is that Apple covers the PCI side of it. If they were doing it themselves, they would then have to be compliant. So if we're not falling under those anythings, if we're just making a cool little...

We're not falling under those anythings if we're just making a cool little game where you upload a photo and we give you back a funky version of that photo, we don't have to comply to anything, right? As long as it's just a promise to our users. So that's the general gist of compliance. I don't know why I did wavy jazz hands, but there it is.

Jeremy: Well, no, I think that makes sense. I mean, you need to do something to make compliance exciting because I think for most people you're right, it's a document they could read and easily fall asleep. If you have insomnia, then compliance documents are probably the way to go.

So the other thing you mentioned, though, is that, again, you are always responsible for your data. And I think up until fairly recently, there were no super strict laws on the book that were about privacy in general. And so obviously we get GDPR, right? What does it even stand for? The General Data Protection Regulations, right? Did I get that right?

Mark: Yeah, you did.

Jeremy: So, that is European, it has to do with the European Union and that came out and that was really strict. And what they said was essentially, "Hey, I don't care if you're hosting in the United States, if you're Amazon or Google or wherever you are, if it is a European user's data that you have, then you are subject to these bylaws." And then very recently, I mean the same type of law, I don't know if they were modeled together, but the CCPA, the California Consumer Protection Act that came out for the United States. And again, it was just for California residents users' data, but also extends and applies all these different places.

So these are very strict privacy control rules. I mean, it even gets to the point where you're supposed to have a privacy control officer and some of these other things, depending on the size of your company. If we go back to this idea of where our data is being stored, so think about this, I am writing an application that uses DynamoDB, and my DynamoDB application has to de-normalize data in order to make it faster for it to load some different access pattern. Or, I'm using Redis or I'm using a SQL server that I'm backing up transactions, or I'm running data through Kinesis or through EventBridge. I mean, you've got hundreds of places where this data could go. Maybe it ends up in S3 as part of a backup, so I can run it through Athena and do some of these things. Now somebody comes along and says, "Hey, I have a right through GDPR and CCPA for you to delete my data and for you to forget me." Finding that data in this huge web of other people's services is not particularly easy.

Mark: Correct. So a little additional context around that, so CCPA is relatively new. When it was initially proposed, it was fantastic and then it got lobbied down significantly to the point where it doesn't even apply, unless you make at least $25 million a year. So it's not even...

Jeremy: Welcome to America.

Mark: Yeah, exactly. But it is a first test at scale in the United States as to whether or not legislation will work on that. And the reason it's in California is very specifically, a lot of the tech is based there. It is a good first step. So let's use GDPR as an example, because it's been out for two years now and there was a preview for two years before that, and it was 27 different nations coming together to figure out where they wanted to go. And we've got a lot more examples around it, but the core principles are the same, the United States is moving closer, but it's going to take a long time just because of the cultural differences, the political differences.

So GDPR really boils down for the users to something very simple. As a European citizen, I have the right to know what you know about me and what you're doing with that information. And if there's any issues with that, I have the right to ask you to remove it or to change anything that's incorrect. That's the user side of GDPR. And now there's a whole bunch of stuff behind that from the business side of GDPR, you already laid out one of the biggest challenges is, how the hell do I answer that question? Right? Especially if you're not building it fresh, if you have an existing application that was never designed with this in mind.

Now, the interesting thing for GDPR is that there are two very big sticks associated with it, why as a security and privacy guy, I love it. It's not perfect. But the first stick is that if you do not take reasonable steps to provide security controls with your infrastructure, you can get a fine of up to 4% of your global turnover. So not profit, 4% of your global take. So if you make a billion dollars, you could be fined up to 4% of a billion dollars, whether or not that's profit, paying off debt or whatever. So that's for not doing due diligence of adhering to something like an ISO 27001, or the basic security controls, right? So if I'm setting my passwords to password, I can get a big, big, fine.

The second big stick of GDPR is if I know there's a breach and fail to tell you about it, I can get hit with another 2% of my overall global take for failing to tell you within an appropriate amount of time, and that appropriate amount of time is 30 days or less. The average law in the United States says its best case effort for notification or at most 45 days. So GDPR is a very big stick, lots of reasonability behind there from the user's perspective. But from a builder's perspective, what you just laid out runs counter to most of the things we're looking for, right? We are trying to optimize, we're trying to denormalize data. You mentioned S3, think about Glacier. Glacier, just the costs alone of if I archive your personal data and shove it into Glacier, not only do I have to find it, and then I have to pull it out and either remove it or modify it and then put it back. That is a huge thing.

But again, like we talked about earlier, if you plan for this stuff ahead of time, it's not nearly that bad because it turns out when you look at this kind of data management through your application, there's actually a lot of benefits just to building your application, to being able to trace a piece of data through your system, to know what I know about you, Jeremy, as a user of my application, there are huge benefits because you lose these sort of legacy bugs where it's like, "Oh, you opened your account before 2008? Well you have this check mark instead of this one." That kind of stuff gets solved.

So for new businesses, I think if you understand it, it's a minimal cost just because it's really getting the expertise in to help you do those design work. For existing businesses though, it is a nightmare. Literally people spent two years getting ready for GDPR and then the regulators still gave them another year before they hit anybody with any substantial fines because of the massive undertaking it is to actually build that kind of infrastructure.

Jeremy: Yeah, no, I mean, and that's the other thing, I guess my advice would be if you're designing these systems and you're building these systems, I think people hopefully think about tenancy or multi tenancy when it has to do with building sort of bulkheads around different clients. So especially if you're a SAS company and you have a big client, you might want to separate out their data from other people's data and have those in separate places. You can do that to some extent, even with user data, right? And so knowing what data you're logging, where you're saving data, using an identifier that maybe obscures that user.

I mean, one way that we tried to handle this in the past was only having personally identifiable data associated in one place where that could be removed. So even though there was a unique identifier for that person, as long as you're removed in that one place, right, which again, and backups, but at least you removed in that one place then you would essentially forget. So you'd have the data anonymized, but you essentially forget it. Now, does that go far enough? I don't even know. I've read the GDPR documents before. And I mean, I read summaries of GDPR, what the rules are that I think were longer than the actual rules themselves. Because again, it is kind of confusing to go through that. So I think that's one thing, again, people of think about GDPR, think about CCPA.

The other thing that's been around for quite some time around privacy for children has been COPPA, which I always thought stood for the Child Online Privacy Protection Act. But I think you told me that it actually is the rule.

Mark: Yeah, the A is just made up.

Jeremy: The A is just made up. So thinking about Fortnite and YouTube and TikTok and all these things that children like to use and like to share stuff and are very quick to say, "Oh, I was born in 2007, I'm going to say I was born in 2005 so that now I'm over the age limit." And of course, there's no verification, there's nothing that stops somebody from doing that. So I'd love to talk about this because this is something, I mean, again, I'm a dad, I have a 14 year old and a 12 year old, and I will not say whether or not my 12 year old is using a Fortnite account that has an incorrect birthday on it, but it's possible she is in order to get access to that stuff. So what do we have to do from a privacy perspective and from a legal perspective in terms of protecting ourselves from this? Because this one has a lot of teeth.

Mark: It does. It absolutely does. So if we take it from the builder perspective, not the parental perspective, the builder perspective, there is a lot of, and we can cover the parental in a second because both of mine are under 13, so it's double whammy. But from a builder perspective, this is where you see in the terms of service that, again, nobody reads, it says you can't open an account if you're under 13 and I'm not a lawyer, thank God, I didn't even play one on TV, what that means is they're trying to shift liability to the user and saying, "If you lie to sign up, that's on you not me."

Because the nice thing about COPPA and it's design, it does actually have a reasonable structure to it to try to prevent companies from tracking kids under 13 online. So you see it a lot of impacts in advertising. And so YouTube at the start of the year, there was a huge push where YouTube basically asked anybody uploading any videos, is this made for kids? Do you think kids might be interested in this? Because if so, we're not putting ads against it because we don't have the ability to turn off all the tracking in the backend so it's ads or nothing. And there was a huge uproar around it and they've softened that interpretation, but it's because they got hit with $170 million fine against this rule because they weren't following it.

So from a builder perspective, it's being aware that if you're serving to children, like if you have an application that is... So let's back up for two seconds, ignore the case where people are lying to get on, right? You need to put that reasonable and say for a Fortnite example saying, "Hey, we rated it 13+ so it's not marketed towards children. We've said it's 13+ for maturity level, just like the movies are rated, the games are rated, and we've added in the terms of service that you shouldn't be playing, you shouldn't open an account unless you're 13+." So we're pushing liability to the user. So in that case, you should probably be covered.

But if you're actually making something that covers kids and families, this is a very real thing that you need to adhere to the rules of the act, which essentially say you can't track kid, you can't advertise directly to them. So where this falls, a question I get a lot, is around schools, especially now that kids are back in school or going back to school, even remotely is G Suite for Education versus G Suite are the same thing, software wise, but very different things legally. And so school boards need to understand what services from Google fall under the G Suite for Education license, because that license follows COPPA to the letter and says, we're not tracking kids, we're not moving this, that, and the other thing. So when you're signed in as a child from the school board and then surf YouTube, the normal tracking doesn't happen on YouTube, it actually creates a shadow account that's not associated to your account and tracks that and doesn't link it back to you as a child. Whereas if you're a normal G Suite user to start to surf YouTube, all that activity is linked back to your G Suite account.

So as a builder, if you're designing something that could be targeting kids legitimately, you need to understand that there's a very hard line that you can't do a bunch of tracking. You can't take the same level of PII, if any, at all, you need to provide adult controls. There's a whole bunch of things that are worth consulting an expert on this for to make sure that you don't follow through. On the parenting side, it's a great excuse to say no for social media, if you want no for social media for the young kids, when they're like, "I want Insta," and you're like, "You legally can't have it. You're not 13."

Jeremy: Right, right. Well, I mean, I again, I think about those situations that pop up all the time, where people are building things that affect kids, I mean, in Fortnite is a good example in the sense that yeah, because kids aren't going to play Fortnite. I mean, it's clearly made for kids. And again, you say 13 and older, and that's fine, but think about Netflix profiles, right? You create a profile for your kid and it tells you what your kid was watching. Now so that's good, right? Because you can go and see what your kid watches, but are they using that data for advertising or optimizing what they might show to your kid? If they're using that for recommendations, where does the privacy line kind of sit for those things?

Mark: Yeah. And that's a very good use case because a lot of kids I know, unless they're really young, don't want the Netflix Kids interface. They want the actual Netflix interface, right? Because Netflix Kids, for little ones it's great because it just shows you Dora and Teletubbies, cool, I can just click on the icon I want. For kids, once they pass five, they're like, "I want to search, I know what I want, I want Transformers or Glitch Techs," or whatever.

The interesting thing about COPPA is optimizing your service is almost always the out for a lot of these privacy regulations and there's always outs. And this is where it really comes down is, none of these regulations are perfect. There's the letter of the law and then there's the intention behind the law. And that really depends on company culture. Almost every company follows the letter of the law or what they think they can argue that letter to be because COPPA has very real fines behind it. There was "Silicone Valley" on HBO had an episode where they were freaking out because their chat app was popular with pre-teens and I think it's 48 or 58,000 per user fine, right? So if you have millions of users, it's an insanely high fine. And that's great. We want that as parents for that protection.

But the line, what your question really hits on is even with GDPR, even with all the CCPA, is what is personal information is really the core question. And there's no clear answer because what you think of personal information and what the law thinks is very different because this is where a pet peeve of mine with Facebook in general is when they're arguing in front of Congress and the first time Zuckerberg testified, they directly asked him and said, "Do you sell user data?" And he honestly in his data-like face said, "No, we don't sell user data." Because they don't sell user data because their understanding of user data by their definition is data that you have uploaded to the service. So your status updates, the photos and movies that you upload is user data, the things you type in are user data.

But, what Facebook sells is access to your behavioral and demographic profile that they have created from user data. So what's Facebook sells is not user data, Facebook sells access to data about users. Now that seems like a super fine semantic hairsplitting thing and it is, but that's the fundamental thing that you're talking about with even the Netflix example, what your kids watch is that user data, or is that data about users? Because all this regulation protects user data, not data about the users and there's a multi-trillion dollar economy dealing in data about users and they don't care about user data.

Jeremy: Right, yeah. And I don't think we have enough time to get into all the details of that. But what I will say is I do know that for me personally, I don't like to share a lot of my personal data. Facebook is just a pet peeve, nevermind pet peeves about it. I mean the whole thing, I'm not a fan of it just because I do feel like it's very exploitive and it's something that, again, once our parents got on it, it just ruined the whole thing anyways.

But I do think that there are valid use cases for taking data about a user, maybe not user data, but for optimizing your application. So I do think that that does make a ton of sense. I mean, again, what are the most popular products that are being clicked on? If you couldn't record that and then use that to show products, I mean, that would be pretty bad. But knowing your particular preference for certain things, and then being able to sell that to an ad company that they can combine with something else that then can serve you up targeted ads, good for the ad companies. maybe good for you if you're getting really relevant ads, but at the same time, it's just a lot of that feels dirty and creepy to me when you start getting very specific on profiling individual users.

But hey, all you got to do is read these privacy documents, these terms and conditions and you'll see exactly what they're doing. So if you don't have a problem with it, I mean, it's kind of hard because I think most people would just glaze over them.

So another one though, another law that has a ton of teeth and I think is going to be more important given the fact that we are now living in a COVID-19 world and more people are building telehealth apps, or they're building other apps, like even tracking COVID cases and some of these other things. For a very, very long time, there is a law called HIPAA, right, that is to protect medical data and all that kind of stuff. Where does privacy play in with that, especially now with all of this medical data, a lot of it being shared?

Mark: Yeah. Yeah. And that's a really interesting example and I'm glad you brought it up because of the privacy, because there's lots of opportunity here and if you're building an application to service the medical community, and that medical community expand far beyond doctors, it's all the third-party connections and things, they all have to follow HIPAA when it comes to health information, so now we're talking about PHI so personal health information in addition to personal information PII, right? So you have both of these in this application. And HIPAA dictates what you're allowed to do with that health information. And again, it comes down to a lot of transparency required because us as patients want that data shared. If you give a history to your doctor and your doctor refers you to a specialist, you want your doctor to share that history information with a specialist because you don't want to go to the specialist and take half an hour of that first appointment reiterating what you already told the first person, right?

And when you go get x-rays or an MRI, you want the results of that to be sent back to your doctor. You don't want to walk around with a USB stick and go, "I brought my data, please analyze it," right? So there's a lot of efficiencies to be had. And HIPAA dictates the flow of information between different providers. And that's why there's a lot of consent forms dealt with HIPAA. So when you sign up with your doctor, if you go to a GP for the first time, they're going to get you to sign a data-sharing document that basically says they're allowed to share information with other specialists that they refer you to, with insurers in the States, again, an outlier given how the rest of the world works with insurance.

But the interesting thing, again, is that as a builder, you need to make sure the service you're dealing with is HIPAA compliant otherwise you can not be HIPAA compliant. But specific to COVID HIPAA has an exemption like most privacy acts that says if it's in the interest of public health, all of these controls can be foregone and that data can be shared with centralized health authority. So in the States, that information could be shared with the CDC. So everything you've told your doctor could theoretically be shared with the CDC if it's in interest of going and helping prevent the spread of COVID-19.

Now, there are lawyers on every side of it, that is the one advantage of the United States. While you lack the overall frameworks, you have more than enough lawyers to make up for it. So the originator, your doctor's office is going to push that through a lawyer before they release all the information up to the HMO. The HMO is going to go through their legal team before they send it to the CDC and so forth. But it is an interesting exemption saying essentially, I don't necessarily care about your history of back issues or sports injuries or blah, blah, blah but I want to know every patient that has tested positive or had any test for COVID-19 because we need those stats up, right? And we need to roll that up at the municipal level, at the state level and at the federal level, because it's in the public interest.

And in this case, and it's a common challenge and it's always all shades of gray, is your personal privacy is not more important than the general health of the community or of the state or the nation in certain cases. And I think a global pandemic provides a lot of argument on that front, but we had a case here in Canada where law enforcement made an argument early in the pandemic that they wanted a central database they could query that would tell them if someone they were dealing with had tested positive for COVID. And that went up to our privacy commissioner then to our federal privacy commissioner, because their argument, the first responders, the police argument, was we could be potentially exposed and we want to know, which there's validity to that argument. But the flip side was, well, is it enough to breach this person's personal privacy? And the current result the last time I checked was, no, it wasn't.

Whereas the aggregate stat, if I test positive or negative, that stat is absolutely pushed up with not my name, but with the general area where I live. So in my case, my postal code, which is the same as the zip code, that gets pushed up because that's not a breach of my privacy, but it helps the information, right? I shouldn't say it's not a breach, it's a tiny breach compared to the big benefit that the community gets. So fascinating, all shades of gray, no clear answers, but it is an exemption and those are not uncommon. There's also exemptions and privacy laws for law enforcement requests, right? If a law enforcement officer goes to a judge gets a subpoena or a warrant, all the privacy protections are out the window.

Jeremy: Right. Well, I mean, it's funny. I mean, you mentioned sort of, this idea of the contact tracing application and obviously there's privacy concerns around that if it's about specific people, but from the law enforcement perspective, I mean, obviously I'm sure you've been paying attention to what's been going on in the United States, if somebody has preexisting conditions even using a taser on somebody, which is probably a bad idea in most situations anyways, but if there were underlying health concerns that they could say, "Oh, this person does have, I don't know, asthma, or they have a heart condition," or things like that where using different types of force or using different methods could cause more harm than it does good if it does good in some cases, but that would be interesting data to potentially be shared maybe with a police officer or a law enforcement officer or maybe not, right? So, I mean, that's the problem is that, you're right, there's data that could be shared that could be beneficial to the public good, but then on the other side, there's probably a lot that you don't want to share.

Mark: Yeah. So the more common example with law enforcement is outside of the health information is your cell phone, right? So your cell phone location, the question of, so not only getting access to the phone, but the fact that your phone constantly pings the network in order to get the best cell service, right? So at any given time, you're normally within, if you're in the city, there's five cell towers you could be bouncing off of and one or two of them are going to be better than the rest because they're closer physically.

And the question you see this in the TV all the time where depending on the jurisdiction, law enforcement may or may not need a warrant to get your location information. There are requests from certain cell providers where they can file it and get the location of your SIM card right now, or your identifier for your phone without going through a significant legal process, it's just a simple request either from the investigating officer or from the DA, instead of going through a judge.

And that's an interesting one because that's never been argued out in public. And I'm a big fan of, there's no wrong answer, you need to just have transparency so people understand. And when decisions are made behind the scenes, because you can see the argument either way, right? If a kid is lost and besides sending out an amber alert, which I think you guys have, we have them here where they blast everybody's phone, instead of sending an amber alert, if you could ping that kid's phone and know where they were, you may be able to retrieve that child. And we know when children are missing, every minute counts, right? The outcomes are significantly better the faster you find that kid, regardless of the situation. But on the flip side, if they're tracking you because they suspect you of a crime, but that hasn't been proven, is that a violation of your rights?

So when it comes to privacy, one of the reasons I love diving into it is because it is all nuance and edge cases and specific examples that overriding things. But most of it's done behind the scenes, which is the one thing I do not like, because I think it needs to be out in the open so people understand.

Jeremy: Yeah. And I think the other thing that is sort of interesting about this electronic surveillance debate versus the traditional analog thing, I'll give you this example, I remember there was a big to-do, and I think it was going around on Facebook or one of these things where people are like fo not put your home address into your GPS in your car, because if somebody breaks into your car, then they can just look at your GPS and get your home address, or they could open your glove box and take out your registration that also has your home address on it, right? So that's the kind of thing where that to me was kind of dumb.

Now, the other thing, going back to the police example, is that I read somewhere and again, I'm really hoping this is true because it's so hard to trust information nowadays, but that if a police officer or a detective wanted to wait outside someone's house and follow them, right, that's perfectly legal for them to do. If they have some reasonable suspicion, they can follow somebody, they can tail somebody, whatever they call it, they can do that. But to put some sort of tracking device on their vehicle, that they can't do, although it's sort of the same thing, except one requires a human to be watching and the other one... sort of the same thing, except one requires a human to be watching and the other one doesn't. And again, I'm not a huge fan of surveillance, so I'm definitely on the more restrictive side of these things. But at the same time, those arguments just seem really strange to me, that it's like it's legal in one sense if it's analog, but it's illegal if it's digital.

Mark: Yeah. And so even clearer example, in most states, and again, not a lawyer, but you require a warrant to get the passcode or password for a user's phone, but do not require a warrant to use a biometric unlock. So I can force you to use your thumb or your face to unlock your phone, but I can't force you to give me your passcode. Both unlock the phone, right? The passcode, the biometric, I mean, there's technical differences in the implementation, but the end of the day, you're doing the same thing. But it's the difference between something you are and something you know, and you can't be compelled to incriminate yourself in the United States. So, there's that difference, right? And if you go to the border, all this is moot, because there's an entire zone in the border where all your rights are essentially suspended.

I think it comes down to the transparency, but for all the examples you just mentioned, the law is always about 10 to 15 years behind technology. So this comes back to one of my core experiences as a forensic investigator. Now I've never testified in court, but I'm qualified to. Most of the reports and things that I worked on were at the nation state level and never get to court. But the interesting thing there is, the amount of cases I've reviewed from court findings and having forensics done, they're all over the map, right? And the case of like, "Oh, an IP is an identifier." And I'm like, an IP is not an identifier of a specific device. If it goes into a building that has 150 devices behind a NAT gateway. Which of those devices committed the act that you're in question?" You can't prove with just an IP address, right?

But it's been accepted in a ton of court cases, because again, comes back to privacy as well, the law is way behind the capabilities. And this is the challenge of writing regulations, of writing law, you need to keep it high enough that it's principled and then use examples and precedent for the specific technologies, because they keep changing. Because yeah, some of the stuff, like the tailing something, is such a ridiculous difference between the two, but it is legally a difference. Similar, to tie this back to the main audience of builders building stuff in the cloud, especially around serverless, how we handle passwords is always very frustrating for me. So let me ask you a question. Why does a password field obscure the information you're entering into it? This is not a trick question, I'm not trying to put you on the spot.

Jeremy: If people looking over your shoulder can't see it, or so that you don't remember what password you typed in.

Mark: Both are true, but yes. So the design of that security control is to prevent a shoulder surf, right? To prevent somebody looking over the shoulder and typing in. So the question is, why do we have that for absolutely every password field everywhere, when there are very low likelihoods of certain situations where people are looking over the shoulder. Which is why I love, and amazon.com has this, and a bunch of other people are starting to add the show my password, to allow the users to reduce the number of errors you put in. Because if I'm physically alone in my office, this is a real background, nobody is looking over my shoulder and watching my password. So why I can't see what I'm typing, right? Similarly, when people say, "I'm going to prevent copy and paste into a password box."

Jeremy: Oh, I hate that.

Mark: Absolutely. Prevent copy a hundred percent of the time. But paste is how password managers work. And what's the threat model around pasting into a box, right? So you paste the wrong password, who cares? So it's understanding the control and why you're doing it, is really, really critical. Same thing on the passwords, why people freak out when I tell them, "Well, write it down. If you don't have a password manager, write it down on a piece of paper." They're like, "What? Oh my God, why would I do that?"

Well, writing it down on a piece of paper and putting it under your keyboard in an office is a dumb idea, because that's not your environment. But we're both at home and if I put my password under my keyboard, it's my kids and my partner that are the threat, potentially someone I invite into my home. But if I've already invited them in my home or it's someone in my family, there's a bunch of other stuff they already know anyway, whereas that will help me having it written down. Again, no bad decisions, understanding the implications of your decisions and making them explicitly, covers security, it covers privacy.

Jeremy: Right. And I can think of 10 people off the top of my head that probably could answer every security question that I've ever given to a bank or something like that because they know me.

Mark: Exactly, right?

Jeremy: And I think that's a good point about some of these security controls we put into place that are, I guess, again there's just friction and it gives security a bad name. My payroll interface that I use is all online, and whenever I have to have a deposit for my taxes, it tells me how much money it's going to take out of my account for my taxes. Well, I move money into separate accounts for those certain things. So I like to take that amount, copy it, and paste it into a transfer window on another browser in order to transfer that money, so that it's in the account that will be deducted from. I cannot copy from that site, it won't let me copy information from that site. And I think to myself, "Why? I can print it to a PDF and copy it from the PDF. I can print it, I can do other things." So why do you add that level of friction that potentially creates mistakes? Like you said, which is why that show password thing is so important.

So anyways, I want to go back to the HIPAA thing for a second, because this is something where we may have gotten a little off topic. I think it was all great discussion, this stuff to me is fascinating. But the point that I wanted to get to with the HIPAA is, if I'm sharing your x-rays, okay, I get it, I've got to be HIPAA compliant. But where is the line for these builders that are building these peripheral applications around medical services, medical devices, medical professional buildings, hospitals, whatever. Where's the line? Because I think about an application that says...

You see this all the time, I just started dealing with this. We just got a new dentist, my old dentist retired, he was completely analog, I don't even think they had an email address. Everything was phone calls, I mean, they were excited when they could print out a little paper card and give it to you with your next appointment on it. So I moved to a new dentist. This new dentist has a hundred percent online scheduling, right? It's great, you pick your hygienist, you can say when you want to set your appointment. And I think about this for doctor offices as well, because I know with my doctor's office it's through a larger, I don't know, coalition or whatever it is. And so they have this health center that you can log into. I don't think you can make appointments, but there's some stuff there.

But let's say someone's building a simple application that is just a scheduling app, right? Maybe you're a little doctor's office or a dentist office, whatever, and you want this scheduling capability. So if I go and I allow this scheduling, if I'm booking an appointment for a general physician or whatever, or a general practitioner, okay, probably not that big of a deal. But what if I'm booking for an oncologist? What if I'm booking for an obstetrician? What am I'm booking for Planned Parenthood or something like that, that gets into really specific things about, obviously, my health, or my spouse's health, my kids' health, whatever it is. When you start booking into specific types of doctors, even though you're saying, "Well, we're not sharing any information about your medical." That reveals a lot right. So when does that get triggered? When does HIPAA get triggered?

Mark: Yeah. And you'd have to consult a lawyer to get the actual official answer, because it's case by case. And I always have to say that, most of my conversations start with big disclaimers. The challenge here, so if I take this from a builder perspective, right? So if we're focusing on the audience who's listening, who are watching, they're probably building applications like this or interacting with them. It is easier to take a more strict approach from the builder side, because you're never going to regret having taken more precautions, if you do it early so that you're not introducing friction. So treating everything as a personal health information or personal identifiable information is going to give you a better outcome. Because if you're like, "Oh, I treated that as health information." And it wasn't, the cost is almost minimal to you when you're designing it from day one.

Because even, you said, well, the GP is not that big of deal. Well, not only is the doctor still a big deal, because it means something is of concern, even if it's a checkup. But if you have a notes field where you say, why are you requesting this appointment? Lord knows what people are going to type in there. Right? Because they're going to assume if this is just between me and my doctor, they will be like, "Well, I have a lump on my neck that I want to get checked out." Oh my God, that right there, diagnosis or symptomatic information is health information. Right? Because even if they just said like, "Oh, everything's fine." Well that still can be treated as profile information. Now the problem is, just like most of the privacy legislations, there's this concept of user data and data about the user.

So, HIPAA mainly focuses on user data, which is again, what you're typing in and the specific entry. So the information you just said, now I believe this to be true, but I'll have to double-check, the fact that you're seeing a doctor is not necessarily protected under HIPAA. So the fact that you've booked in with the oncologist, or the pediatric surgeon or whatever the case may be, is not a specific class of data that needs to be protected.

The information you share with them, your diagnostic results, all your blood work, all that kind of stuff absolutely is. But the fact that you haven't, that you just booked an appointment or you spoke to them on the phone, isn't necessarily protected. I believe it should be, because it is very clear that I don't care what type of cancer the target has, I just care that they do. Because that's something I, as a cyber criminal, can manipulate to get what I want, right?

So that's a big problem, is that there's not that line. So similarly, a different example under medical is the genetic testing, right? So 23andMe, ancestry.com, hey, test your genes at home. They all advertise like, "Hey, we keep your data super secure. We protect your health information." And blah, blah, blah. But they aggregate out your genetic code and use that for a whole bunch of stuff in the back end. And they said, "Well, we don't tie it to you." Well, it's easy enough, if somebody gets that piece of information to then tie it to you, if they have access to the backend systems anyway. And that's the challenge we deal with, with all these data and specifically with healthcare, is it's very rarely one piece of information that is the biggest point of concern. It's the multiple pieces of information that I can put together in an aggregate to get a better picture of what I wanted as a malicious actor.

So I don't care that you spoke to this doctor, but I do care that you went from no doctor's appointments to five in a month, right? Because I don't know what's wrong, but I know it's something big, because who sees five doctors in a month, from never seeing a doctor in the last year, right? Something is happening. And so if you put your bad guy hat on, if I'm trying to break into the government and you work for the government and I realized there's something wrong, there's a good chance that I can make a cash offer, that you're looking at a mountain of medical bills, and I could probably compromise that way and say, "Hey, here's a million bucks. I need your access." and I'm in. Even not knowing what was wrong, but just knowing that pattern has changed.

So it's again, a lot of nuance and a lot of challenge. But from a builder perspective, if you treat everything in a health application as PHI, your cost is not going to increase significantly. If you plan early enough, your friction isn't going to increase, but you're definitely going to protect your users to a higher level, which is actually a competitive differentiator as well.

Jeremy: Yeah. Yeah. Well, I think the best advice out of that is, seek professional legal help for those sorts of things. And that's the thing that just makes me a little bit nervous. Whenever you're building something new, you might have a great idea, or you're going down a different path, you're pivoting a little bit, that when these things come up, you do need to have legal answers and solid legal advice around these things to protect yourself.

All right, we've been talking for a very long time and thinking back about what we've talked about, we've probably scared the crap out of people thinking, "Oh my goodness, I'm not doing this anymore." But let's bring it back down, because again, I don't think anything we talked about as hyperbole. I think all of these things are very, very real. The laws are real, the compliance regulations are real. Just this idea of making sure that the data that's being saved is encrypted and that you put these levels of control into place are there. Those are all very, very real things.

But you had a talk back last year at Serverlessconf New York. And I thought it was fascinating, because essentially what you said was, "The sky is not falling." We are not now opened up to all these new attack vectors, that, yes, there are all these kinds of possibilities. So let's rebuild up everyone's confidence now that we've broken it all down and let them know that, again, building in the cloud and building in serverless, that, yes, you have to follow some security protocols and you have to do the right things, but it's nowhere near the gloom and doom that I think you see a lot of people talking about.

Mark: Yeah, for sure. And I think that's absolutely critical. If there's a second takeaway besides find legal advice is, stop worrying so much. And I think that's where I have challenges and interesting discussions with my security contemporaries, because when we're talking amongst ourselves, we talk about really obscure hacks, interesting vulnerability chains, zero day attacks, criminal scams, all this kind of stuff, because we're a unique set of niche experts talking about our field, right?

Whereas, when you're talking about general building and trying to solve problems for customers and things like that, you have to look at likelihood. Because risk is really two things; it's the probable impact of an event and the probability that that event will occur. So security is very good, outside of our security communities, about talking about the probable impact of an event. So we say, "Oh my God, the sky is falling. If this happens, your entire infrastructure is owned." But what we don't talk about is the likelihood of that happening. And be like, "Yeah, the chances are one and 2 trillion." And you're like, "Well, I don't care then."

It's interesting, nerd me is like, "That's interesting and I like that." But the reality is that often, the simple things that cause... So if you look at the S3 buckets you mentioned in one of the early questions. I followed those breaches very, very closely. If you want to follow at home, Chris Vickery from UpGuard, his career over the last couple of years has been focused almost exclusively on S3 bucket breaches, fantastic research from him and his team. In every single case, it has not been a hack that has found the data, it has been simply that it was accidentally exposed.

Given the probability, yes, zero-day vulnerabilities are real, cyber crime and attacks are real. You will see them over the course of your career. Your infrastructure will be attacked, simply because it's connected to the internet, that's just the reality. You actually don't have to worry about that nearly as much as you think, if at all. What you need to focus on is building well, building good resilient systems that are stable, that work reliably, and that do only what you want them to do, is going to fix 95% of the security issues out there, and then you can worry about the other stuff. So the S3 bucket are just mistakes, they're just misconfigurations. Even Capital One, who unfortunately got hit with a $70 million fine because of it, it was a far more complicated mistake, but it was still a mistake.

Basically, they had a WAF that they would custom build on an EC2 instance in front of S3 and that WAF's role had too many permissions. That's it, right? A mistake was made and it cost them, literally and reputationally. So it's the likelihood of something happening is that you are going to mess up. So putting in tools in place, things like Cloud Custodian, the great open source project, little testing around security configurations, AWS Config, using Google's security operations center, all these tools that are at your disposal that are free or almost no cost, to help prevent making mistakes. The idea to keep in your head as a builder is that, you drew it up on PowerPoint, that's great. You drew it on the whiteboard, that's fine. You need something that checks to make sure that production is what you drew, and that's going to cover the vast majority of your security concerns. And if you get that done, then you can worry about the obscure cool stuff.

Jeremy: Right. Yeah. And I think you're totally right. I think that if you cover the bases, and again, just like I said, especially with serverless, it almost all comes back to application security. I did an experiment two years ago at this point, where I was able to upload a file to S3 and the file name itself was SQL server injection attack, basically, that was the name of it. And so when it tried to load the ID or whatever the file name was, into a piece of code that, again, didn't think about SQL injection, because maybe it thought the source was trusted or whatever, that then there's the problem there. How many people are going to even try that one? That's one thing.

And then the other thing is, that again, if you're building solid applications and following just best practices, you should be thinking about SQL injection. That's a very real thing that if people don't build... And of course, there's so many tools now that you just shouldn't be building SQL injection anymore, but people still do. But again, I think there is a sense of a doom and gloom or FUD, you know what I mean? Trying to get people to buy these security applications and things that they do. Because again, I think that if you don't think there's a problem, you're not going to buy a solution to fix it, right?

And for the very few people I think, who do get hacked, they're like, "Oh, I really wish I bought that solution." So I don't know what the right level of advice is. I think you're right, as to say, your likelihood is very low, but I still think people should think about security and put that first in a way that says, yes, maybe I don't have to worry about the obscure attack this way, but I should just make sure that I'm following best practices. And like you said, I like that idea of using open source projects to make sure that your infrastructure, or your proposed infrastructure and your actual infrastructure do match.

Mark: Yeah. So let me say this, coming from a vendor. So Trend Micro was obviously a vendor, or one of the top vendors there. I don't agree with everything we put out from a marketing perspective, because sometimes it does skew negative. We try not to, but it still does, because like you said, if you don't believe there's an issue, you're not going to buy a product, and all the vendors are generally guilty of this.

But I think it's a misunderstanding of what security's role is in the build process and in building applications. And I think if you're sitting at home right now, you're listening to this, thank you for sticking along, it's been a long episode, or broken up into a couple. But I think it's all important and it's interesting, but it's not just academic, like you said. There's real issues here, there are real things that are going on, but the best way to look at security controls is actually from a builder's perspective.

You mentioned SQL injection, there is an open source, fully tested, phenomenal input validation library for every language out there. There is no reason you should ever write your own. You should just import one of these, there are any sort of different licensing available as well, and so you import that and you get that to check your validation. Because I think that is an example of larger security controls.

Security controls can help ensure that what you're writing does what it's supposed to and only that. So input validation is a great example. If I take a lambda function that takes an input that is a first name and a last name, well, I need to verify that I'm taking in a name, that's a valid name, first name and last name and not a SQL injection command and not a picture, somebody trying to upload a data file and things like that. That's a security control.

And if you see it as not trying to stop bad stuff, but from making sure that what you want to happen is the only thing that's happening, you start to adjust how you see these security controls and go, anti-malware the classic, classic security control, you're like, "Well, I'm not going to get attacked by malware." Don't think of it as something to stop attacks, even though it will. Think of it as, I'm an application that takes files from the internet, there's bad things on the internet. I cannot possibly write a piece of software, in addition to building a solution, that is going to scan this file to make sure that only good things are in there.

Well, there's an entire community of security vendors that do that for a living. Pay for one of those tools, not to stop attacks, but to make sure the data you're taking in is clean. So when you adjust that kind of thinking and realize the security controls are just there to help you make sure what you think is happening is actually what's happening, you start to change your perspective and go, "Well, there's great open source stuff. There's great stuff for purchase, but I don't need to buy anything and everything. I don't need all this crazy advanced threat hunting stuff. I just need to make sure what I'm building does what I want and only that."

Jeremy: And I think if we tie this back to the original topic here, is that the privacy of your user's data, is going to be dependent upon the security measures and the rules that you follow to make sure that that data is secure. So Mark, listen, thank you so much for spending all this time with me and just sharing that perspective. I mean, this was an episode I really was excited about doing, because I do think this is these things that just people don't necessarily think about. I know we didn't talk a ton about serverless, but I really feel like it all does tie back to it and I mean, just cloud development in general. So again, thank you so much for being here. If people want to find out more about you, watch your video series, things like that, how do they do that?

Mark: Yeah. And thank you for having me, I've really enjoyed this conversation and hopefully, the audience and we can all continue this conversation online as well. You can hit me up on Twitter and most social networks @marknca. My website is markn.ca and everything's linked up from there, my YouTube channel and all that is there as well. So happy to keep this conversation rolling in the community as well, because yeah, even though we weren't specifically talking about a lot of serverless aspects, I think the principles apply to everybody. The good news is, if you're building in a serverless environment and serverless design, you're already way further ahead than most people, because you've delegated a lot of this to the cloud providers, which is a massive win, which is one of the reasons I'm such a huge fan of the serverless community.

Jeremy: Awesome. All right. Well, we'll get all your contact information into the show notes. Thanks again, Mark.

Mark: Thank you.

View Details

About Mark Nunnikhoven

Mark Nunnikhoven explores the impact of technology on individuals, organizations, and communities through the lens of privacy and security. Asking the question, "How can we better protect our information?" Mark studies the world of cybercrime to better understand the risks and threats to our digital world. As the Vice President of Cloud Research at Trend Micro, a long time Amazon Web Services Advanced Technology Partner and provider of security tools for the AWS Cloud, Mark uses that knowledge to help organizations around the world modernize their security practices by taking advantage of the power of the AWS Cloud. With a strong focus on automation, he helps bridge the gap between DevOps and traditional security through his writing, speaking, teaching, and by engaging with the AWS community.

Twitter: https://twitter.com/marknca
Personal website: https://markn.ca/
Trend Micro website: https://www.trendmicro.com/
Watch this episode on YouTube: https://youtu.be/aPg7WE3Q3SQ

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly, and this is Serverless Chats. Today I am speaking with Mark Nunnikhoven. Hey, Mark. Thanks for joining me.

Mark: Thanks for having me, Jeremy.

Jeremy: So you are the vice president of cloud research at Trend Micro. So why don't you tell listeners a little bit about your background and what Trend Micro is all about?

Mark: Yeah, so Trend Micro is a global cybersecurity provider. We make products for consumers all the way through to massive enterprises. And I focus in our research wing. So we have a really large research component. There's about 1400 researchers in the company, which is a lot of fun, because we get to dive into the minutia of anything and everything related to cybersecurity, so from the latest cybercrime scam to where I focus, which is in the cloud. So a lot more what I'm looking at is how organizations are adapting to the new reality of things like the shared responsibility model, keeping pace with cloud service providers, adjusting to DevOps philosophies, that kind of thing, which is a lot of fun.

And for me, I come from a very traditional security background, if there is such a thing. I've been at Trend for a little over eight years. Before that, I was with the Canadian federal government for a decade, doing all sorts of different security work, a lot of nation state attacks and defense, things like that. And my background in education is actually in forensic investigation, so that nerd in the lab on your favorite crime drama when they come up with the burned-out hard drives and are like, "Fix this," and somehow they do, it's all BS, but that's technically what I do.

Jeremy: Very cool. All right. So I have wanted you on the show for a very long time, because I've been following the stuff that you've been doing. I love the videos that you do, the blogs that you write. You're just out there. And I know you're on the edge of the serverless space, I know you do a lot of stuff in the cloud as well, but you're obviously into serverless as well. And just recently I came across this impact assessment video series that you're doing. I don't know if it's a regular series or whatever, but it was really good. And you were talking about Fortnite and Apple, and I want to get into that. But really what made me think about things a little bit deeper that goes beyond just some of these surface-level billionaires arguing with billionaires is this idea of privacy and how important are online privacy is. And I thought it'd be really interesting to talk about how serverless and privacy, since it's in the cloud, is all the stuff that you're sharing, where that kind of aligns. So let's start. First of all, why is privacy or online privacy so important?

Mark: Yeah. That's a really broad and great question. So yeah, this new video series I'm doing, Impact Assessment, is going to be regular. I was doing a livestream called "Mornings with Mark" for the last few years, did, I think, like 200 episodes where it was mainly talking about cybersecurity issues of the day, and a lot of those are privacy. And where I wanted to go with this new series was just a little broader audience, which is why Apple and Fortnite and Twitter hack and stuff like that are coming up, because I think privacy is a really important aspect, and it mirrors security. You can't have one without the other. And it's directly related to the audience, to people who are building in a serverless space or in any space.

But privacy, a traditional definition of privacy is really your right as a person to not be observed, essentially to be alone and to have control over your data and your well-being. And when you go into the digital world, it's infinitely more complicated than a physical world, right? You can lock yourself away in a room in the real world and be relatively confident that nobody is invading that space, that you have kind of control over that space, so if you want to just sit there and veg out, if you want to read a book, that's an activity just amongst yourself, right? When you come to the digital world, everything we do leaves a trail somewhere. There are tons of exposures potentially. You as a user don't really have a ton of control over your data.

And one of the things that I wanted to do with this video series and with a bunch of my other work was just enlighten people to help sort of expose this so that they're aware, because one of the challenges I get on the security side of what I do, and it directly relates to the privacy side, is that people assume there are correct decisions. And really, the only incorrect decision is one that you are unaware that you're making. So you could make the argument that it's okay that you're tracked everywhere on the internet, and I think the trade-off you get for the free services may be correct, but if you're unaware that that is the trade-off, I think that's the problem. So that's the intention behind this video series, is to look at privacy issues, to look at some security issues, to help people just make a conscious decision instead of just being pulled along for the ride.

Jeremy: Right. Yeah, no, and I think that that's probably something that a lot of people miss, is that people say, "Well, I'll sign up for Facebook, and I will share every photo, every place that I visit, all my friends, all my likes, all my dislikes." And what I think people say is, "Oh, well, whatever. It's free." And they don't realize that they're the product, and most of that is because they are giving up so much of their privacy.

And it's actually funny. This just happened the other day to me, and I didn't even realize. I knew it was coming out, but Chrome just released a new update that blocked third-party cookies if they weren't... I think you had to have like "secure" on and some of these other things. So no user is going to have any idea what that actually means. But what happened for something we were doing is, we were loading a third-party cookie behind the scenes for something, and all of a sudden that stopped working. And so the whole flow of this module or this modal pop-up thing completely broke because of that extra thing of security. And I remember way back in the days of early web development dealing with IE5 and IE6 and the browser wars, like what works on this browser and what works on that browser. Now privacy seems to be the new browser war thing that are conflating those two things.

But anyway, so that's one thing, but let's go to this idea of the Fortnite and Apple thing, because I have two kids, two daughters. They've played Fortnite more this summer than I think... I don't know how anybody could play Fortnite more than that. But they love it. And then I told them the other day, because you and I were talking, I saw your assessment video about them not releasing it on iOS because of the whole Apple Store thing and all this kind of stuff. But why is it a good thing, I guess? And maybe we can talk more about Fortnite. I mean, I'm not really into it. I know you are, but I'm not really into it. But maybe we can talk more about why that review process, why that purchase process through Google Play or through the app store, why is that important to your security and to your privacy?

Mark: Yeah, and I thought this was really interesting. So I got into Fortnite a couple of years ago when I did a piece on it for my regular radio column here in Canada. And I thought it was interesting because it's a microtransaction model game, so it's always taking a lot of money from people, but not to win the game. It's purely cosmetic. And I thought that in general, especially as a parent myself, I thought that was a really positive thing, because it wasn't like a bunch of these games where you need to pay to actually have a realistic chance at winning. The only thing you're paying for in Fortnite is to make things look different. There's no performance differences, right? And since that... Then there was this great Saturday Night Live sketch a couple years back on Fortnite, where this character Adam Driver was playing was solely there to learn how to be better than the stepfather, to show off to the kids. And I always think, "That's me," even though... Just trying to be cool to the kids. But I do play regular.

I thought it was interesting, you know, being pulled up in this drama, because most of the drama between Epic and Apple and Google somewhat right now is related around the business side, because the Apple policy... and Google is the exact same, but we'll just use Apple because it's more prominent right now... the policy basically says, as a condition of being in the app store, you need to follow a whole bunch of these rules. And the rules that Epic is calling out is the one around transactions, and it says basically, if you're taking money through the $, so directly through the app, Apple gets a 30% cut. That's their fee as a middleman for bringing you customers. And as a part of that, Apple will facilitate the transaction. So for Apple users, you're well familiar with this. For Android users, it's similar. But that's why you can use face ID to authorize a transaction through Apple Pay, and you don't actually have to enter a new password. You don't have to give them your credit card information. All of that stuff is handled by Apple as a proxy for those businesses.

And so Epic, they make north of $300 million a month from Fortnite. And they said, "You know what? 30% of the chunk we make from mobile, which is north of $100 million, is too much." So they are contesting that, and they actually have plans, and in their legal filings are saying, "We're not going for the money. We want the right to be our own app store." So there's a really interesting business case there, and they're really petty and low blows, which is fascinating and fun to watch from the outside.

But I did a video in the assessment around, what do we actually get from a security and privacy perspective? Because everybody is saying, "Oh, 30% is a huge amount," even though it's not uncommon in the retail space or in other business transactions. But there's a lot of stuff that goes on behind the scenes, and that's really beneficial to us. So when you submit as a developer, Apple makes sure that there's no obvious malware, though this week there was a case where they actually approved malware, which is one out of eight years of app store, which is not bad. They look for malware. They look for using undocumented APIs, which could create vulnerabilities. They look for your use of personal data, which is what I really dug into, was that they have restrictions around what developers can do with your data, how they can track you, what they have to ask permission for.

And that actually goes to your transactions as well, because a lot of the stuff that happens behind the scenes that we don't even think about is when you go to a store, like a retail store, if you still can in these days, and use your credit card, most of the larger retailers actually track that credit card usage within their physical store. So they will take a hash of your number instead of storing your actual number, and they will look for that reused to create a profile for you if you're not actually signed up for the loyalty rewards thing. Same thing happens online. So not only is the money important, but the more... having someone between you and your customer means you can't track them as much. So from a business perspective, they're saying, "I want the data to be able to track Jeremy and Mark more accurately." But as a user, we want Apple or Google in between us, Apple, definitely more so than Google, given the business models, because they're that blocker. They're preventing us from having our privacy unknowingly breached, in that people are tracking our transactions online. And that's part of the big thing we get through the app store.

Jeremy: Yeah, and I think that having that broker in between is another major thing that dramatically helps with privacy, just from a... Not only privacy, but I guess security as well. And I never use anything but PayPal on most sites that are not amazon.com, because I don't trust some little site.

I mean, actually the funny thing is, I just bought something that I was almost... It was from one of those... What's the store there? My mind is drawing a blank here, but the... Shopify, right? It was a Shopify store. And essentially Shopify says, "Yeah, anybody can build a store." I don't even think they check, and I may be wrong on that, so I apologize if that's wrong, but it seems like it, because there's a lot of stories of Shopify scams. And there was this thing listed, and it was actually a pool for... It was one of those Intex pools, those temporary pool things. We just needed something. You couldn't buy them anywhere unless it was like thousands of dollars, which was crazy. So I saw this deal, and I'm like, "I'm going to buy it, but I know it's a scam. I'm almost 100% sure it's a scam." But I used PayPal, and I knew that the worst case scenario was I'd have to send a few emails back and forth and I'd get my money back. It turned out to be a scam.

But if I hadn't, if I had given that person my credit card number, who even knows if that credit card number would have went to a valid processor, or if it would have been run through some third-party thing, or it would have had thousands of dollars of transactions across the dark web or whatever. So I do think that there is a tremendous amount of added benefit to having that middleman protect your privacy.

Mark: Yeah, and that's an interesting example. And I'm sorry that you got scammed. And I understand, especially in these times, trying to get those items in. Because at the start of the pandemic, it was like basketball nets, trampolines, bikes. You couldn't get this stuff, right? And the nice thing is PayPal as a middleman works. There's some downside when you're the collector from PayPal, for sure. But Visa and MasterCard have the same protections in place. It's very rare that you're going to be financially on the hook. But the difference is, it's a pain in the butt to go back review where you have your normal subscriptions charging to your credit card and things like that, to redo all of that. So even though you're not out money necessarily, you're still out time and frustration.

And that's happened to me pre-pandemic when I was traveling, literally one time when I crossed at the US customs to here in Canada. We cross in the airport itself, and I found out when I tried to buy some food that, oh no, my credit card had been blocked, and so I had to get a new one shipped and all that kind of stuff. So I wasn't out any money. I was just out of frustration.

But there is important aspects, both advantages and disadvantages, to the middleman. But specifically when it comes to that online, a great example there of knowing that there's a good potential for a scam, understanding the risk of, okay, a couple of emails? It's not that big of a impact to you to try. And the upside where, if they did actually ship you the inflatable pool, you're the hero to the kids and happy and cool. So it's finding that balance. And again, like we said in the intro, is really, for me, it's, there's no bad decision. It's just making it explicitly. So you just gave a fantastic example of explicitly understanding that you might get scammed here. There's a high chance of it. But then you used a way to protecting yourself. You had four options to pay, and you picked the one that was going to provide you the most amount of protections, because you were aware of the situation. And I think that's commendable. I think the flip side is, most people are unaware on that scale of what we're doing in the online world, of the types of ramifications of those decisions.

Jeremy: Right. And so speaking about unaware, I mean, one thing that I think people might not understand when they make financial transactions or they share data, they're often giving it to a machine, right? And we think it's super secure if we just slide our credit card in with a little chip on it, or if I enter my information on a website somewhere, or I save my password or something like that and I know it's only saved locally. The problem with people is people, right? And I love people. Don't get me wrong. But once you introduce the human factor into any of these security or privacy issues, or potential privacy issues, it gets exacerbated because people are fallible and people make mistakes.

I think the most important one that happened recently is this Twitter hack. And people are like, "Oh, Twitter got hacked." Well, it depends on what you mean by hacked, because nobody brute-forced into and broke into the system and figured out somebody else's password. They literally scammed people who had access to this stuff. It was a social engineering attack. So how do you prevent against that?

Mark: Yeah, and this is the challenge. And so one of the things for those people who look into sort of the history of my work, I always feel like I'm an outlier, because a popular sort of feeling in the security community is what you just said to the extreme, that the users are a problem. Everything would be great if we didn't have users. Well, we wouldn't have jobs if we didn't have users, so put that aside. But the reality is, people very rarely are trying to do something in their daily work to cause harm. So criminals, obviously that's their daily work. They are trying to cause harm. So this case of Twitter was that the people who were doing the support work were just trying to support users and to get their job done, right? Now, it turns out that Twitter was a little lax and they had about 1500 people with access to the support tools.

But if you step back for a second... Okay, ignore the hack. It totally makes sense if you're running a service that supports 330 million people that you as a business are going to need some tools to be able to reset passwords, to adjust email addresses, to give people access back to their accounts, because someone's going to forget their password and not have access to the email that they signed up with legitimately. They're going to change phone numbers, so they don't have the SMS backup. Stuff happens, especially when you have 300 million plus users. So to build a tool to help you deliver better customer service 100% makes sense. The problem in this case, as you pointed out, is that it was also a vulnerability, because the controls around it, the process around it was a little too lax, these cyber criminals didn't do any crazy hack. And I think if there's one fallacy on the security side of things, it's that, and it's partially because of all the TV and movies, which makes for great TV and movies, but very rarely do big-name hacks actually use anything remotely resembling state-of-the-art hacking. Nine times out of 10 it's a phishing email. Actually, 92% of all malware infections start with a phishing email, because they work. They're super easy to send and to confuse people. I always remember a talk from Adrienne Porter Felt who's at Google. She was in the Chrome team at the time. They'd done a massive, million-plus-person study, and basically the key result was, nobody reads security prompts. So it doesn't matter what you prompt the user, they're just going to click okay. Which is frustrating, because you're trying to educate them and move forward.

So with the Twitter thing, it was just a social engineering attack. They got some extra access by basically just tricking a support employee, which then got them access to the Slack. In the Slack channel, to make the support team's lives easier, they had some credentials posted that said like, "Hey, to get into the big super tool here, here's the login. Here's the URL." Which, I mean, you totally understand, working with a team, that you drop stuff in Slack like that all the time, because the assumption is you're in a private room, right? And in this case, that wasn't it. And thankfully it was a very visible hack, so it got shut down very, very quickly.

But it's these of things that I think are interesting, because my point in that particular video was, most people who use an account, A, assume it's theirs, when you're just actually using it, you're renting it kind of thing. And they aren't aware that there's a support infrastructure behind it that gives people access legitimately, because if that was you who lost your password you'd want access back to your account. You've worked hard to grow your social media following. So it's, again, being aware of those trade-offs

Jeremy: Yeah. And again, there's so many examples of things where people are sharing a lot of information that's getting recorded and they probably aren't even aware that it's being recorded. I mean, every time you talk to Alexa... "Alexa, cancel," because she's just going to come up on me. And then every time you talk to Siri, every time you type on your computer if you have Grammarly installed, all of that information is being sent up somewhere. And so when you introduce... Even if you have the best security protocols in the world, and you're in AWS Cloud or in Google Cloud and you're all locked down, you still have that potential that somebody could simply accidentally share their password to some super tool, like you said, and your information gets shared. I mean, think about S3 buckets, right? Apparently S3 buckets are just... It has been one of the biggest... Or I guess the Capital One breach, right? Is this idea that you just make your things public, or you make it easy for them to be copied, or whatever it is. You don't do it on purpose, but those are human mistakes that are causing those issues.

Mark: Yeah, and there's a lot of trust there. So there's a couple examples that I think are really interesting that you gave there. So the voice assistants are popping up more and more in court cases where the law enforcement are actually requesting access through legal process to the records of what they have heard, because Alexa is a good example, and sorry if I triggered yours or any of the audience's. I have a good voice for that, apparently. But if you go into the app, you'll see actually a history of all your commands, everything you've asked for and whether or not it... Because you can provide feedback like, "Yes, it gave me what I wanted. No, it didn't."

You had mentioned for keyboards on phones and stuff. So Grammarly is a good example. When iOS started allowing keyboards, third-party keyboards, I thought it was really interesting, because one of the prompts that people don't read that pops up says, "You are providing this keyboard with full access to everything you type." So everything you literally are typing, even if you delete it, is being sent to the cloud and back. Is that a bad thing? Not necessarily, but if you don't know that that's happening, you can't make that choice. And that's really the thrust of a lot of what I'm doing, is understanding that work. Because at the end of the day, one of the things I hear often on the privacy side is, "Well, I have nothing to hide. I don't care." And a lot of the time that may be true, but you still need to be aware of those data flows that are going out from you out into the world. And that's where things get more and more complicated the more technology we add.

Jeremy: Right. Yeah, I totally agree. All right, so let's take this into the serverless realm here, because this is a serverless podcast, but I think this is super exciting, because I'd be interested to get your perspective on where serverless and privacy meet. And I think if we take a step back and we look at security first, I think we know, I think this has been demonstrated, that the security of a serverless application, just based on the shared responsibility model, how little you need to do from a maintaining a server standpoint, from even just... There's no direct TCP/IP access into a lambda function, for example, right? Like that all has to be routed through a control plane. So you just have all these levels of security. So the majority of the security concerns from a serverless perspective are going to come down to application-level security. And we have talked about it at length.

And again, people make application security mistakes all the time, right? And the social engineering aspect of it is something where giving someone your password into an admin that you build for your customers... But I want to take it a little bit further and go beyond just this idea of, maybe we make an application mistake, maybe something gets compromised, maybe someone shares a password here. So from a serverless perspective, if I'm building a serverless application, how do I start building bulkheads around my application to protect some of this private user data?

Mark: Yeah, and that's a really good setup. It's a good explanation. I 100% agree by default serverless gives you a better chance at security, because you're pushing almost all the work to the service provider, right? That's a huge advantage, which is why I'm a massive advocate of serverless designs.

So maybe it's easier just to clarify for the user as well, because we've bouncing back and forth, focusing on the privacy, talking a bit about security. I said you can't have one without the other. And really, security is a set of controls that allow you as the builder or even you as the user to dictate who and what has access to that data. And then privacy is just the flip side of that of me going, "This data is about me, and I want to know who I'm entrusting it to." And security is then the controls that you... If I entrust you with my personal information, security is then the controls you're putting on top of that information to enable privacy, right? So they're intertwined. They are linked concepts.

So if you as a builder are creating an application that is handling personal data or handling any type of data, you're fighting this inherent sort of conflict of nature, in that we've been taught as developers for the last few years that the more data we have the better, right? The more data that we're tracking, the more awareness. We can get better fine-tuning on our application. We can increase the performance. We can increase the reliability. We get a better operational view the more data we have. From a privacy and a security point of view, the more data you have, the bigger the liability you also have.

So you need to first go through and make sure you understand what type of data you have. So cold start time on a lambda, total route time for a request, those kinds of things aren't sensitive to specific data. They're sensitive somewhat to your application, but in general, that's not something you need to take... You don't need to lock it in the vault that's encased in concrete, thrown into the ocean so that nobody can ever get to it. If I'm dealing with your social security, that's a far more private piece of information that I need to take further steps to protect. If I'm dealing with your health record, same kind of thing. So it's first step for anybody building any application is just listing the types of data you're actually hosting and processing and then mapping out where in the application they're required.

So for permissions, we have on the security side the principle of least privilege, which is essentially, "I am only going to give you the bare minimum permissions you need to access something," which is the S3 problem at its core. When you create an S3 bucket, only the user or entity that created it has access rights by default, and then everything else has to be granted. And all of these breaches, billions and billions of records over the last few years, have been because somebody made a mistake in granting too many permissions.

So understanding what the data is and where it actually needs to flow and saying, "You know what? This health information isn't going to flow to the standard logs. We're going to keep it in a Dynamo database, and that Dynamo database, that table is going to be encrypted with this KMS key, and it's actually going to break our single-table design, because this information is sensitive enough to merit its own table, because I don't want to take the risk of encrypting column by column, because I think I might mess that up. So I'm going to just separate it completely to make it a logical separation to make it easier." So really, step one is mapping that out and then restricting the breadth of that data or where that data touches, and that does a huge amount of effort, a huge amount of the work to maintain privacy right there.

Jeremy: Right, yeah. And so if you're taking that data, though, and you're... And again, I think this makes complete sense. You're saying, "Look at what it is you're saving. If I'm saving somebody's preference, even if I'm saving somebody's like, whether they like a particular brand or something like that, is that really personally identifiable information? Is that something that I have to lock away and encrypt? Or can I be more lax with that? What about usernames and passwords and things like that?" And I think that all makes sense. Think about it that way. But I think where I'm curious where this goes is, you only have so much control over that data if you are saving it in DynamoDB, right? If you are capturing it through CloudWatch logs-

Capturing it through CloudWatch logs because it's coming in, and maybe it's coming in and it's not encrypted, I mean, even though you are using an SSL or TLS, you come through and the information is encrypted from the user's computer or their browser, into the inner workings of AWS, for example. Then once it gets into that Lambda function, that's all decoded, that's all unencrypted. Right? That's all ready for you to do whatever you need to do. So then you need to take that and put that into a database, or send that off somewhere, or call an API, or any these other things. When you do that and you save that data into, let's just start with DynamoDB, there are backups, those are automatic backups. Right? There's again the CloudWatch logs. So this data is going all different places, so that seems like a lot of effort to make sure that a credit card number, or social security number, or anything that you want to be very careful about, that you have to take a lot of extra steps to make sure that's encrypted.

Mark: Yeah, and I think this is spot-on example. And I think this is the number one failing of the security community over the last 20 years or so. And there's a lot of logical reasons for it, is that right now, the vast majority of security work, so that security work to ensure that privacy of data is done after the fact. Right? So if you think of your DevOps wheel and you've got the development side and the ops side, security exists almost entirely in the ops side. Which means we're taking whatever's already been built and then doing the best thing we can. So we end up with this very traditional castle wall sort of scenario of like, I've made a safe area, drop things into it, and they will be safe from anything outside of that wall. But unsafe from anything inside that wall.

And that's had mixed results, I think is a generous way of saying it. And realistically, if you think of security is really a software quality issue, and we know you're not going to do testing only in production, you're going to do testing early stages, you're going to have unit tests, you're going to have integration tests, you're going to have deployment and functionality tests, You're going to do blue-green deployments to make sure that things are running before they hit prod. There's a whole bunch of testing we do as builders before we get to actually interacting with users. We need the same thing from security, because what you just mentioned is a lot, if you're thinking about it once you've designed the solution, but if you're designing the solution with those questions in your mind, as you're going forward, it's actually not a lot of additional effort to map out these security things as you're sitting there.

If we're starting up a new application, me and you, we're doing Jer and Mark's super cool app. Right? And we go, okay, we're going to start logging in users. Well, we look at that and go, well, the bare minimum, we have a username and a password. So we're going to have to do something with that, we need to know what that flow is. So maybe we're going to loop in something like Cognito. Maybe we go, you know what, Cognito is not quite where we need it to be, so we're going to go to a third party Auth0. So now we're outside of, if we're building it in AWS, now we're outside of our cloud into a third party with a whole different set of permission sets. But if we're designing that from day one, we can map that out and go, okay, we know we get TLS from the browser to Auth0.

We know that TLS doesn't actually guarantee when talking to Auth0, it just guarantees that a communication is secure in transit from A to B. It doesn't tell you who A or B are, which is a mistake a lot of people make. But then we go, okay, we're going to Auth0, fine we've got a secure connection for the user there, we verify who that user is from Auth0, our app will verify Auth0. This is the following method. And then we're going to take that data, and we're going to make sure that we don't actually store it, that we don't actually log the user, because what we've done is we've never taken the password out of Auth0. We've just gotten a token and now we map it there.

So I think if you go after the fact to try to do this, it's really difficult. So even if we just simplify the example down to encryption, the thing I know you always see Vernors shirt, "Dance like nobody's watching, encrypt like everybody is." Love that shirt it's so dirty, it's amazing. But if you take an existing application and say, okay, we're going to go encrypt everything in transit and at rest, that's an annoying massive project that has no direct, visible benefit to the customer. It's really hard to get those things past the product manager, because you're like, Hey, I want to take, four sprints to do this work that will save us potentially if something may happen bad, like if a cyber criminal attacks us, we will be protected. But our customer's not going to see anything for four sprints because we're not doing any feature work. That's a hard sell.

Whereas when you're designing that out of gate one, and you say, I'm just going to add a parameter and it's going to encrypt everything in transit, and I'm going to add a KMS parameter to the Lambda, and everything's going to be encrypted at rest that took five minutes and we're done. Nobody's going to bat an eye and you get the same end result. So it's really about planning ahead, I think.

Jeremy: Yeah. Well, I think security first. I mean, I think that's the first thing just with the cloud and so many of these problems that happen from breaches that are again, not necessarily a vulnerability of the cloud, it's just more of these social engineering things. Then again, thinking about security right off the bat is a huge thing. And I guess here's another thing, and I know that like DynamoDB, for example isn't, you can do encryption at rest. Right? And that things like SQS and SNS, I think those have an encryption in transit as well and... Right? So there's a lot of that security built in, but again, all of those really great tools that the cloud provides and the encryption and whatever else, that goes away the second you build an admin utility that someone can log into and just query that data. Right?

So what do you need to do around that? What should you be thinking in terms of, I mean, are there multiple layers, should we be thinking... You always hear things like Tier one, tier two support, things like that. Are those levels of access that they have to your private data? How would you approach that when you're building a new application?

Mark: Yeah. And the tiering system is frustrating as it is for a lot of users, a lot of it does have that. If we use the AWS term, it's about reducing the blast radius. You don't want everyone in support to be able to blow up everything, and if you look at the Twitter hack was actually an interesting example, somebody raised the question and said, "Why didn't the president's account get hacked?", "Why wasn't it used as part of this?" And because it has additional protections around it, because, it's the leader of the free world ostensibly so, you want to make sure that that's not the average, temporary employee on a support contract, being able to adjust that. So the tiering actually is a strong play, but also understanding that the defense in-depth is something we talk about a lot in security. And it gets kind of a bad rap, but essentially it means don't put all your eggs in one basket.

So don't use one control to stop just one thing. So you want to do separation of duties. You want to have multiple controls to make sure that not everybody can do certain things, but you also want to still maintain that good customer service. And I think that's where, again, it comes down to a very pragmatic business decision. If you have two sprints to get something out the door and you go, well, I'm going to build a proper admin tool, or you're just going to write a simple command that your team can run, that will give them the access, you're just going to write a command that does the job. And you know what, in your head, you always say the same thing.

You put it in your ticket notes, you put it in your Jira and you say, we'll come back to this and fix it later. Later never happens, so most admin tools are this hack collection of stuff just to get the job done. And I totally get it from a business perspective. It makes sense. You need to go that route, but from a security and privacy perspective, you need to really think holistically. And I think this is a question I get asked often, actually, somebody just asked me this on my YouTube channel the other day, they said, "I'm looking for a cybersecurity degree, and I can't find one. All I can find is information security. What's the deal?" And I said, well, actually, what you're looking for is information security. In the industry, and especially in the vendor space, we talk cybersecurity because that's typically the system security.

So locking down your laptop, locking down your tablet, locking down your Lambda function, that's cybersecurity, because we're taking some sort of cyber thing and applying security controls to it. Information security is an academic study, as a field of study in general, is looking at the flow of information as it transits through systems. Well, part of those systems are people, are the wetware. Right? Or the fact that people print it out. This is a big challenge with the work from home was, you said, well, your home environment isn't necessarily secure. And you said, well, yeah, it has different risk models. But the fact that I can connect into my corporate system and download a bunch of stuff and then print it, that's information, that's still needs to be protected.

So I think if you think information security, you tend to start to include these people and go, wait a minute, Joe from support, we're paying him 15 bucks an hour, but he's got a mountain of student debt. He's never going to get out of it. That's a vulnerability that we need to address, not from locking it down, but help that person out and make them feel included, make them feel, as part of the team so that they're not a risk when a cyber criminal rolls up with some cash and says, Hey, give me access to the support tools.

Jeremy: Right. Yeah. No, I mean, and the other thing too, when you're talking about, I guess people having access to things, one is having access to data. Right? And so if you have an admin account that can create other accounts, right? And you get into the admin account or the admin account does everything for example. That's really hard to prevent against if you let that go. Right? But there's another vulnerability with admin accounts, especially when it comes to the cloud, is any time somebody has access to a production environment. Right?

So with AWS, if people are familiar, and I'm sure this is true with all the other cloud providers. You have multiple log-ins that are called roles in AWS and you can grant access to certain things. And the easiest thing to do is when someone's like, Hey, I can't mess with that VPC, or I'm trying to change something here, I'm trying to do that. Okay fine, I'll just give you admin access. Right? So admin access gives you everything except for billing access for some reason, but it gives you everything in the AWS cloud. And I'm not saying, I mean, you need that somebody needs to have admin access at some point. But when you're writing code that could potentially expose data by maybe having a vulnerability in an admin tool or just giving too much control in an admin tool, there needs to be a process that separates out, the development environments, the staging environments, and then that production environment where all that actual, sort of production user data is going to go.

So I always look at this as like... And that maybe people don't think about it this way, but to me, CI/CD having a really good, whether it's Gitflow or something like that, that has control as a place where there are very, very, very few people who have keys to that main production account. Everything else is handled through some sort of workflow, with approval processes and things like that. And I mean, to me, that is the staple of saying you want a secure environment, you have to set up CI/CD

Mark: Yes, a 100%. So my general rule of thumb is nobody should ever touch production, systems should touch production. And so the pushback I get on that a lot, especially for people that are still in virtual machines or instances are like, well, no, I need to get data off of there. You should have a system that pulls data out and logs it centrally so you can analyze, because if you need to make a change, you push it through the pipeline because not only is that better for security, that's better for development as a practice in general. For those of you who are watching this episode, you can see how much gray and white is in my beard. For those of you just listening, think like Santa levels of white, and I've been doing this a long time. And the inevitably, I used to be keyboard jockey doing the Saturday night maintenance windows for Nationwide Networks.

And you're typing the same thing in the system, after system, after system, you had your checklist, you did everything you possibly could to prevent yourself from making a mistake. You still ended up making at least two mistakes per change window, per Saturday night because it's late night, you already worked all week, you're only human. Mistakes happen and enforcing a consistent threes through a CI/CD pipeline. Not only gives you the security benefits, but it gives you the reliability that if a mistake did happen, it's the same mistake consistently across everything, which means you can fix it a lot easier. It's not that there was a different typo in every system, there's the same thing on every system, so you can roll forward. And that's an absolutely critical thing to do because, a lot of the time people see security as this extra step you need to take as this conflicting thing, it's going to slow you down at the end of the day, security is trying to achieve the same thing you are as a builder.

We want stable, reliable systems that only do what you want them to. That only act as intended as opposed to some vulnerability or mistake being made that people could leverage into making that system do something unintended. And that, CI/CD pipeline is absolutely critical. You mentioned roles. There are equivalents in GCP and Azure as well. My big thing is accounts should have no permissions at all, other than the ability to assume a role. So if you can assume a role as an account or as an entity, then for specific tasks, you have a role for every task.

So if I need to roll a new build into the CI/CD pipeline, don't give me a permanent rights to kick off builds. Let me assume a role to kick off a build to kick off the pipeline, because then I don't make a mistake, but also we get an explicit log saying, at this time I assumed this role, it's cryptographically signed, it shows that chain of my system, made that request in the backend, and then after assuming that role, I then kicked off this build and you just get this nice fidelity, this nice tracking for observability on the backend. We're so obsessed on observability and traceability and production. You need it as to who's feeding what into the system. And then that way I don't make a mistake and we get clarity. So it's roles are a massive win if you use them right.

Jeremy: Yeah. And I think things like CloudTrail and some of the other tools that are built into AWS, I'm sure a lot of people aren't looking at them, but they should be. But so the other thing, it's funny, you mentioned this idea of, doing late night support. So I think we've all been there. I mean, if you're as old as us, I have as much gray as you do. I try to hide it a little bit, but I remember doing that as well. And I still have some EC2 instances that I have to deal with every now and then. And one of the most frustrating things about trying to do anything, and I think this is why security people try to find workarounds for it, is because security creates friction. Right? So the more friction you have, you can't access my SQL database in a VPC from outside, unless you set up a VPN or some other tunnel that you can get into. Right?

I think about every time I log into my EC2 instances, first thing I do, sudo su. Right? Because, I just know, I don't want to try to go to the logs directory and not be able to get to a logs directory because security is preventing me from doing that. And, so again, being able to have ways in which... It's almost like people have to build those additional tools. Right? So you mentioned only machines should be touching these things or systems should be interacting with it. But those are systems that somebody has to set up, those systems that somebody has to understand. Right? So again, I totally agree with you. I'm a 100% with you. It's just one of those things where it's like these tools are not quite as prominent or they don't seem quite as prominent as some of these other workflow tools are, again like even CI/CD you could build a whole bunch of security measures into CI/CD but I think people just don't.

Mark: Yeah. I agree and so I'll give you a good example that I can't remember who told me, but after talking in an AWS Summit two years ago, somebody gave me a brilliant example that they had set up that I thought was a really good demonstration of how security should be. Now, it almost never is, but it should be. And it was exactly that problem was that they still had cases where people had to log into EC2 instances. And they were trying to figure it out, they knew they couldn't just say no. So what this team had set up was a very simple little ping little automation loop that as soon as somebody logged in some of the SSH and EC2 instance, CloudWatch logs would pick it up, it would fire off a Lambda and it would send a Slack message to that user.

And it would provide a button. And it would say, Jeremy, was this you logging into EC2 instance ID, blah, blah, blah. Yes or no. And if you hit, yes, it would then provide a second little message, it's just like, Hey, we're trying to cut down on this. Can you let us know what your use case was? What were you missing? Why did you have to dive in? Why did you have to log in? But if you said, no, it would kick off the incident response process, because somebody logged in as you, and it wasn't you. And I thought that was a really good example of saying like, look, we know we want to be here. We can't get there yet. So we're going to take a nice little friction-free response of sending out the standard survey to everybody and be like, how many EC2 instances do you log into and open? Nobody cares.

But catch them in the moment. And I think further to the bigger question of, yes, somebody has to build those tools. Somebody has to develop those. And again, if you try to get that past a product manager, it's not going to happen because there's no direct customer benefit, It's against a theoretical issue down the road. The challenges or the failures on the security side, for the longest time, the security teams have been firefighting nonstop and have developed this rightfully so reputation of being grumpy of saying, no, I'm putting roadblocks in place that preventing people from achieving their goals. So people just ignore them or work around them.

That's not what we as a security community we need to do. We need to work directly with teams, we need to hear a thing like you just sat and said, okay, no problem, we're going to build that for you, we're going to make sure we build some sort of flow that gives you the information you need in a way that we're comfortable with from a security side, so that there's no friction in place. And that is a huge challenge because it's cultural and the security teams continue to firefight and can't kind of get their head above water long enough to go, Oh, we could do this in a way better way, instead of just continually frustrating ourselves and everyone we work with.

Jeremy: Right. Yeah. And I mean, the idea of being proactive versus reactive would be very nice. I know every time you get that thing where you're like, okay, something is not right. You can hear everybody in the IT, all the developers just sigh at once because you're like, ah, this is going to be a long night. We're going to have to figure out what exactly is happening here.

All right, so let's go back to privacy for a second, or maybe for more than a second. Another piece of this that is important is we are saving data into somebody else's systems. Right? We mentioned DynamoDB, we mentioned, SQS, and some of these things are encrypted and that's great. But you've got GCP, you got Tencent, you've got Alibaba, you've got Microsoft Azure, you've got Auth0. Right? So you're saving data, personal data into other people's systems. So I guess sort of where my question is from a privacy standpoint, I can put all these controls in place and I can say, Oh yeah, I've done all my tiering, I have all the security workflows, I've got CI/CD set up, my admins are locked down and I know whatever. But where does, your responsibility as a developer start and end when it comes to privacy, when it's being saved on somebody else's system?

Mark: Yeah. And that's a very good question because there are legal responsibilities and then there's "The right thing," quote, unquote. For various definitions of the right thing. I think most users expectation is that they have a modicum of control over their data. Now, the interesting thing here, as we start to get into a little international differences. So if people have been listening to the episode so far, have probably figured out that I'm Canadian by my accent, and Canada has a very strong privacy regulation. We're not quite as strong as it used...

Jeremy: I did not say tell us about yourself though.

Mark: Which is fair. So the Canadian perspective, we have a legal framework and a different expectation, the European expectation is completely different. The outlier when it comes to privacy is actually the United States. Now, the interesting thing is the United States is also the generator and the creator of the vast majority of the technology that the rest of us use. So when we look at the legal requirements, there're different things. When we look at what you should be doing and what the expectation is, it really comes down to cultural. So what a European citizen would expect to happen with their data, is very different than somebody in the United States, because there is a cultural and a legal expectation in the EU for their data to be treated very, very differently. So the generic answer is when you're building out serverless applications, specifically, you need to make sure that whatever level of data you're dealing with, the service that you're leveraging can support the controls you want around that data.

So if we look at PCI, which is the Payment Card Industry framework, there is a legal requirement. If you're taking credit cards to have certain security controls in place, you need to be PCI certified, which is why a lot of smaller businesses just go to a provider, but bigger businesses it's worthy your time to set yourself up like this. There are legal requirements for the controls around it, which means if you're building in a serverless design, regardless of the cloud you're using, the aspects of your design that are processing payment cards, so processing MasterCard, Visa, Amex, need to be on services that are also PCI certified. So you can't achieve the certification if the service you're building on isn't also certified. So there's that aspect of it in general, that you need to just sort of go with that compliance, but it's really tricky because it comes down to what do you want to do versus what do you need to do?

And that sort of, it's a difficult thing to respond to because sometimes there's very real legal penalties for not complying. But the good news is from the serverless aspect is that, the shared responsibility says that you're always responsible for two things, configuring the services you use. Right? So all the providers give you a little knobs and dials that you can change, so you can encrypt or not encrypt, you can optimize for this or that. You need to understand and configure that service, but you are always responsible for your data, always. There is at no point you cede responsibility for your data.

If you leverage a third party... So if I'm the business and you're my user and you give me personal information or information, you want private, I am on the hook for it, regardless of who I use behind me, it's me. So I need to make sure that the services I'm leveraging in my serverless design, have the controls that I'm comfortable with to make and follow through on that promise to you as a user, and that changes but it's always you, and you need to verify as a builder that you're leveraging services that meet your needs.

Jeremy: Yeah. So you mentioned two separate things. You mentioned compliance and you mentioned sort of legality or the legal aspect of things. So let's start with compliance for a second. So you mentioned PCI, but there are other compliances there's SOC 2 and ISO 9001 and 27001 and things like that. All things that I only know briefly, but they're not really legal standards. Right? They're more of this idea of certifications. And some of them aren't even really certifications. They're more just like saying here, we're saying we follow all these rules. So there's a whole bunch of them. And again, I think what, ISO 27018 is about personal data protection and some of these things, and their rules that they follow. So I think these are really good standards to have and to be in place. So what do we get... Because you said, you have to make sure that your underlying infrastructure, has the compliance that's required. So what types of compliance are we getting with the services from AWS and Google and Azure and that sort of stuff?

View Details

About Xavier Lefèvre

Xavier Lefevre is currently VP of Engineering at Theodo, a web development and product consulting agency. As part of his role, Xavier manages five technical teams and leads the development of the company’s serverless expertise. He believes that serverless is a major breakthrough that will allow the industry to redirect its focus on core business needs, and his specialization centers on serverless and problematic FinOps architectures. Xavier shares his expertise through articles such as with Serverless Transformation on Medium, and various speaking events, including Virtual Serverless London meetup.

  • Twitter: https://twitter.com/xavi_lefevre
  • LinkedIn: https://www.linkedin.com/in/lefevrexavier/
  • Medium Blog: https://medium.com/@xavierlefevre
  • What a typical 100% Serverless Architecture looks like in AWS!
  • Serverless Cost Calculator

Watch this episode on YouTube: https://youtu.be/pKc2f8Q0PQI

Transcript:

Jeremy: Hi everyone, I'm Jeremy Daly and this is Serverless Chats. Today I am chatting with Xavier Lefevre, who I am going to have re-pronounce his name afterwards. Hey Xavier, thanks for joining me.

Xavier: Thank you. Thank you for having me. My name is Xavier Lefevre in French, which is not very easy to say.

Jeremy: So you are the VP of engineering at Theodo. I'd love it if you can tell the listeners a little bit about your background and what you do at Theodo.

Xavier: Yes, so I'm going to start with Theodo. Theodo is a product consulting and development agency, so we work with clients of any sizes, companies of any sizes, to build websites and complete web applications for them for different kinds of use cases. So, it can be a big company, it can be a small company, it can be eCommerce, it can be a big industry, anything. We are in France, UK and U.S.A., London and New York to be exact. And we're doing several different types of web ports, so we do mobile, we do infrastructure, that's something that could be interesting, like Kubernetes a lot, and stuff like that, so that's Theodo.

And for me, so I'm a VP engineering of Theodo in France, at Paris. And I have a fun background. So, I went to business school when I was younger and I did business school, but I always wanted to work in tech and when I got out I started to work in tech, but as a business role. I realized what it meant to work in tech and the different roles. And I finally realized that I preferred to be in tech myself, so that's why I'm here today.

Jeremy: Awesome. So, you are a bit infamous now, you have this article that you wrote called, the Typical Serverless Architecture, which got a lot of praise and also got a lot of criticism from people who don't quite understand serverless architecture. So, I would love it, just to go through ... and we'll start with this, there's other things I want to get to, but let's start with that, let's start with this typical serverless architecture. Take us back, what does that look like?

Xavier: So, from experience, and I don't have ... I have a year of experience in serverless, so I'm still, compared to you, I'm still young. But from experience, I started to of course dig into serverless and understand a little bit everything that's included in the technology. And I wanted to show this big picture and this big idea of what the typical architecture is. So, what can you find inside? You can find ... So, we go through each box. Okay? I can talk about the origin as well, which can be interesting. Which one do you prefer first?

Jeremy: Well, so I have them listed here, so let's start with the front end, what does the front end look like in a serverless application?

Xavier: So let's do that. So, front end itself, your front end is going to be a FGA, for instance, like Drax, it can be Next for instance, with SSR. What you can find there are two things that are specific to servers, the first is AWS Amplify, which does a lot of stuff, but among which you can find many components in there that help you work faster. And you can find STKs that help you communicate more easily and find pre-made features with AWS services, like Cognito for instance. You can authorize your users and handle your users directly from your front end thanks to AWS Amplify, so that's one piece you can find. The other one is ... if I go a little bit further, when you host your front end. So, you have two steps, first the basic one, if it's a static React you just have to host static files with your JS that's going to be loaded on your product and that's going to run, we know how it works.

So, there you're going to use F3, it's going to be exposed by your CloudFront, and that's it. If you want to go further, which happens a lot lately, even more because it's partial, if you want to go into SSR or SSG or a bit both, that you can do with Next for instance. Here you can use ... you can take for instance Lambda at Edge, which are Lambdas that are inside of CloudFront, that run close to the users, that are super, super, super fast and that can take care of generating your pages for itself, for you, for performance purposes or SU purposes. So that's one capacity in terms of front-end servers.

Jeremy: All right. So, now you've got your front end and you've got it hosted either in CloudFront or maybe even Amplify Console, which is different than Amplify, you can host SSGs there as well, and things like that. So, what about for domains and certificates, how do you manage your domain names and your certificates in a typical serverless architecture?

Xavier: Okay. So, there are two services that are quite famous as well in AWS, you're going to have certificate managers that's going to help you deal with your certificates, HTPS, and you're going to have Route 53 that's going to help you deal with your domains. They play directly, easily with the rest of your serverless architecture and you don't have much else to do. And we're talking ... something interesting as well, we're talking a lot about AWS, because it needs those services and from experience we found the serverless experience more, let's say more fit and more complete in AWS, but of course you can do this kind of stuff in other providers.

Jeremy: All right. So, now we've got all this stuff set up, this is our front end, now we actually want to be able to process APIs, so how do we build out our business APIs?

Xavier: Okay. So, there you have two choices, first about the routes you want to expose through the API, you have API Gateway, even more complicated, you have two API gateways, the View 1 NTP2, which are not named the View 1NTP2, not easy, and you have AppSync on the other side. Okay? So, if the API Gateway is about making REST APIs, AppSync is about making GraphQL APIs, mainly, if I have to just recap really fast. Of course, they have a lot of capacities inside, one API Gateway has an option to use API Gateway WebSocket, for instance to do realtime AppSync as embedded GraphQL substations, with pros and cons. So, we have those two, depending on what you want to do, REST or GraphQL. Okay, that's your gateways.

Then behind that, you're going to have your business intelligence. So, you can directly connect those gateways to Lambdas at first, mainly to code your site. So Lambdas, what are they? They are functions, just little pieces of code that you split, that are super microns that you split, you give them ... you put them and deploy them to AWS and they do the rest. They handle the scalability and the uptime of those functions, you don't have to take care of that.

So, you connect your gateway to Lambda's, from there ... What we do, and it's not something new, it's not something specific in serverless and you code it differently, we organize our Lambdas in services, we do follow again the principle of microservices. Behind that, the idea is mainly to develop a pattern to have a well organized architecture and to have an architecture that's going to last long in terms of complexity of thinking. So, we follow the main design at max to have a clear separation of concerns between those microservices.

Jeremy: Right. So, then you have all these microservices that are separated. Now, one thing about microservices in general, and certainly with serverless, is it is a distributed system, right? You're communicating with lots of different moving parts. So, you're not going to connect every piece of business logic or everything that you do, that's not all going to happen synchronously, right? So, we want to send an update to Marketo or to Salesforce or something like that, we're going to do that asynchronously, right? So, now you've got these Lambda functions that get processed from your business APIs, then how do we connect all these other services that you have and do that asynchronously?

Xavier: Okay. So, there are several ways of doing that. So, there is this idea of serverless being driven by design, that's something that gets mentioned a lot. And you can see it in most services, Lambda reacts, and most services react, on events. So, API Gateway connected to Lambda, is an event. But you can also connect it to other services, like SQS, like EventBridge. We do use Eventbridge for us to split and communicate between our services in a, let's say uncoupled way. EventBridge is a serverless Event Bus, it's like a RabbitMQ, but easier, which is extremely comfortable and it scales by design. So one microservice is going to handle, let's say is going to be the article microservice, okay, it's going to communicate with the user authentication microservice connected to Cognito for instance, which is a service made for authentication of users, and they will communicate through this Event Bus.

So, the article, for instance, is going to push the fact that there is an article and it's not going to know ... to the whole, let's say the whole architecture, it's not going to know who is going to know who is going to consume that. And EventBridge is then going to make sure this message is related to the services that want to consume this information and react on it. If you want to go further, then NDTV has a service inside of a service, which is complicated, called DynamoDB streams, you can use the fact that you're going to push changes to data inside of DynamoDB to trigger other events, uncoupled, so even events to make side events, like for instance beginning with a Lambda, it could be something like that and doing something like ...

So, this data change in the database, I want to send an email, but I can decouple and ask for DynamoDB stream to trigger this effect. And the effect is that the first Lambda that changed the data, synchronously ended earlier and to the front end, which showed the user that the request was accepted. Okay?

Jeremy: Okay. And so now you have these back-end processes that are running tasks, right? They're asynchronous, so the user's already connected, they can take longer to run if they need to, they can fail, they can retry, you've got all kinds of things like that. What if after one of these processes finishes, you want to push data back to the user somehow, you do that with WebSocket and AppSync, right?

Xavier: Yes, that's exactly what I would do. I can take a look at the API options you could have if you want to push data, if you want to do realtime, so if you want to in a simpler way push data from the back to the front end. The best options are API Gateway WebSocket or AppSync distributions. So, API Gateway WebSocket, from I understood, and I think you can change that, you have to ping the API per connection to ask to push a message to the front end. The difference with AppSync, where you can send a batch of messages in one go. So, that's one difference in one process of AppSync you can find, but in both scenarios, they're amazing services, really reliable. And that's how you push your messages to the front end, with those two services.

Jeremy: Right. So, then in a serverless architecture we have no servers, right, or at least no servers that we know of, so, if you want to upload a file, which is a common thing that people would want to do, maybe upload a profile picture or something like that into an application, how do you do that in a serverless application?

Xavier: So, of course we use the infamous S3, which is a file storage system. So, one thing you can ... you can once again, in this event-driven manner, is that you can generate file URLs, file upload URLs with S3 at your back-end to generate with a token and upload files that you're going to send back to your front end. The front end is going to directly push the file to S3, and then in this event-driven manner, S3 is going to, if you want it, it's going be able to push an event to trigger another side effect, the fact that this file has been uploaded, you want to potentially to mark some things somewhere in your database. So, you're going to do it in this way, completely decoupled.

Jeremy: Okay. All right. So, now use and authentication, you mentioned Cognito, what else do we do in order to authorize users into our system?

Xavier: So, Cognito is a go-to authentication user management service, than anybody else, so that's where you're going to have your user base, that's where you're going to have your user access and you're going to be able to connect it with other services to make those authorizations. So, for instance in API Gateway for each member you can say that ... you can attach an authorizer, which is directly connected to Cognito, and say that this route is not going to be accessible to this type of person. That's how you're going to do it.

Jeremy: All right. And so now let's say you have a complex workflow in a serverless application, you mentioned Lambda functions, each one maybe does a discreet piece of business logic, but let's say you've got to connect five or six of them together, or more, because you're doing some sort of checkup process and you need that workflow to finish on state machines, how do we build those in serverless applications?

Xavier: Yes, this is interesting, because it's something that gets ... after some point workflows gets complicated in most applications. So, you could do it ... of course you could do it yourself with your own Lambda, but in the end it's going to be complicated and tricky to understand where the data is going and stuff like that. So, AWS has a service for this, which is called Step Functions. Here you describe your flow and the configuration, okay. So, you're going to describe your steps, you're going to describe your states, and how the state is supposed to change step after step. You're going to say that it's away, it needs to come back, it needs to be handling errors, and you can do that with Step Functions.

The nice thing with it is that AWS takes care of this flow for you, in this case once again this flow for you and makes it as well super visible, so you have a nice user interface to see ... to understand what's happening, where are you, when you are debating on where you are developing, where are you in your flow, what data went where, and what you also get out of the box. So that's really super comfortable to handle complex workflows after some point.

Jeremy: Right. Two more to go. Security?

Xavier: So, security, the two things I put forward inside of this article is, one is called IAM, which is extremely famous, it's identity and access management. It's where you're going to define and configure all your authorizations inside of AWS. So, this user has access to this service, has access to this exact action inside of this service. But it's also where you're going to say that ... you're going to define the security of accesses between services. So, this one can have access to DynamoDB and stuff like that out of the box. The idea with IAM and a good practice, is to ... You're going to help me in finding this term again. Is the fact that no one has access to nothing. And then bit by bit you will access it.

Jeremy: Yeah, principle of least of privilege, right?

Xavier: Yeah.

Jeremy: Block everybody by default and then the principle of least privilege open up just what that individual function or user needs.

Xavier: Exactly. So, with IM you do that out of the box, which is a pretty good security principle. So, that's one service. Another service I'd put forward is how to take care of secrets, and your API keys for instance. So there are two services for that, one is called Systems Manager and the other one is called Secrets Manager. In my opinion they have some differences, but you can do both for handling your secrets and your API keys, and not having to version them so much.

Jeremy: Right. And then the last one, and this is important, monitoring, how do we monitor a typical serverless application?

Xavier: So, this one is extremely important, we are in a distributed system and they were in an asynchronous system, so understanding what's happening, what even did trigger what action is indeed complicated. So, out of the box CloudWatch is a de facto solution, it's connected to all AWS services, at least all the ones we mentioned before. And that's where you're going to have all your monitoring, you're going to find all your logs, you're going to be able to customize the logs and the metrics. If you want to push that, you're going to be able to define, not easily, that's still an issue, but it's powerful, you're going to be able to define your dashboard, your alarms. And you can do a lot of stuff in my opinion with CloudWatch. In my opinion it's almost self-sufficient. But one thing you need as well, in terms of scalability, is to understand what's in your whole system and CloudWatch is not enough for that.

CloudWatch is going to be micro, it's going to be a bit too shallow and some very specific pieces I particularly want to understand, and this is actually the entry to this whole flow, I want to understand why it created this error then came back to the flow. For this there is one called X-Ray, which is doing tracing end-to-end between your services, that's supposed to be the go-to. But, and you can correct me if you think it's the case, X-Ray is not supported by all services, and EventBridge which is at the heart of our microservice architecture is not supported by X-Ray, which is a little ...

Jeremy: All right. You just listed or we just went through 11 different sort of categories within a typical serverless architecture, and we probably mentioned 20 services, and there's even a few in there we probably didn't mention. So, the criticism I think that you got from this article that you wrote, and it's probably valid criticism, nothing on you, the article was great, and I think this is exactly what a typical serverless architecture looks like, but the criticism was, "Wow, this looks really complex, right, you've got a lot of different things." So, what do you say about that complexity and also who does that complexity now fall on, right, because this sounds like the developers in this case are going to be doing a lot of these things that maybe were the Ops jobs in the past?

Xavier: Yeah, that's the key question and the key topic when I released this article. And to a person that criticizes architecture, this article, I think I'm going to show them the video of ... What's his name already, the one that sings with the piano with all services of AWS?

Jeremy: Yeah. Forrest Brazeal, yeah. Forrest. Yeah.

Xavier: That video is amazing. If I send them this video, I think they are going to tell me, "You see, we are right, it's true." There is a crazy amount of services. But it's just the way of thinking, the mindset needs to shift. The complexity and the pieces you find in an architecture at the same, let's say level of application, it's the same, we are not inventing something, it's the same, we're just reorganizing a little bit the way you connect those pieces. So, of course the critics say, "I don't understand why there is so many pieces, you take a monolith, you connect it to AWS and it does everything you mentioned there for you and there is no issue." And that's true.

The thing is that behind this monolith you will have potential scalability problems afterwards, you will have it at some point, and you will have to have ... Your company is going to grow and you will have to state it in the services and you will have to be able to make them stay. And you're going to bit by bit explode this architecture. And here serverless is asking you to think about that ahead of time, right from the start, but it's also helping you to do that properly and to not have to think about scalability anymore later, not at all. So yes, indeed as well it changes a little bit the waits on the developers, the developers have to think from the get-go about this organization and how to split it and how to communicate between those different services.

But once again, it's something ... eventually ... And we are talking about applications that's going to need a little bit of scalability or that are going to have some up and down traffic, that will need eventually this kind of stuff. And we're going to get it from the get-go, which is super comfortable. And from experience, it's positive.

Jeremy: Right. And you mentioned a good point about eventually companies are going to have to split up monoliths, and I know that there are a lot of companies that run on these big monoliths, but have you ever worked one when there was a new piece of functionality that needs to be added, it's always like, "What am I going to break if I put in some new service?" And I've always told people ... 10 years ago or so, I would say to people, "When you're building a new application, you just want to get it out there, right, like don't worry about scalability, right.

I mean, if you get a thousand users and your servers are overwhelmed, then great, right, then you now have something and maybe you could start re-architecting and getting it where it needs to be. But you don't want to spend all that time building in that scalability on a typical server app or a server-based application right from the get-go, you want to get the project out there." But that changes dramatically with serverless, because you can build these things very easily, you can get them out there very quickly, and if you just think about a few things with the scalability aspect of it, in terms of what you need to do to make sure that your functions are split independently and some other things, you can scale right up and you don't have to rewrite anything.

Xavier: Yeah. And there is a good practice when you develop, when you just developed, you have your development flow, a good practice is always to cut down your features in very small pieces, it's always a good practice, to be able to have a good, like mind control on what you're doing, and to make sure that you're not going to generate [inaudible], because it's easier to sync through. So, this human thing of splitting things in small pieces to make sense out of them is ... It goes as well inside of the building of an architecture. So, in the time, like in the time, five ... Serverless started to become popular five years ago, or something like that?

Jeremy: Yeah, about five, it's been five years or so. Yeah.

Xavier: So, a little bit before that. If you wanted to be able to do those microservices, it was a little bit more complicated because of the infrastructure you had to put around it, because of all the complexity and all the tools you had to put and that you had to manage yourself to make them work. Now, the super nice part is that AWS took care of this complexity for you and you can split this complexity in several pieces to make sense out of it from the get-go. And it's just a simple idea. It's extremely powerful. And you can see it in a lot of other fields.

Jeremy: Right. Yeah. And I think that some of these individual services that are available to us help, just by their sheer, I guess, design or the way that they're meant to be used, already have the thought of scalability built in. Like DynamoDB for example, right, like thinking through those access patterns, your database is going to scale for quite some time, you don't have to worry about it. You build something in MySQL or something like that, and that's going to work great until you maybe have a hundred thousand records that need to do some sort of cross-join or something like that, with these complex filters on them and then all of a sudden that's going to slow down.

Xavier: Yeah, that's true. We didn't ... I think I forgot about DynamoDB in the business part of my typical architecture. But yeah, DynamoDB is a serverless database we commanded as well from the beginning. And indeed, it's amazing to see as well the shift in paradigm and the fact that you have, you change ... It's always around, mainly around performance and the fact of making sure the system you're building is performing and long-lasting at the same time. And DynamoDB is a really great example about that, because when you did SQL, you didn't ... you thought about a clear and easy to sync relational model. Okay. So, it was not normalized, everything made sense, all those models had a link and a clear unity and you saw the direct connection between them. But the issue is like, it was growing and at some point your SQL queries got bigger and heavier, and there you had to do some tweaks, and here you had some performance issues, and complex performance issues usually.

So, you had to think about reshaping your data or improving your scale of performance ... request performances. And other things, you had to think about scaling your database, but scaling the SQL database is much more complicated, because it's a huge blob of data and you don't exactly know when and how you can access the data, because it's inside of your SQL queries. So, doing horizontal scaling on that is almost impossible, because you're going to get some data here, some data here, some data here, in this horizontal scale, and it's going to be terrible in terms of preference. So, DynamoDB shifts that. So, think about how you can access the data, think about preference and think about how you're going to store this data from the beginning, from the moment you put this data inside of the database. And there it works, after it works.

I'm not saying it's easy, entering data into the DynamoDB world is funny, you have to ... Once again, it's a change of mindset, but it's an exciting one, I really enjoy to think about that, how to start that and how to store data inside DynamoDB. But then it works, after that it works. So, there is also always these critics about, "But do we need that, do we need this change and this, let's say extra complexity from the beginning?" But there is that ... Of course, you have to learn something new. But it's not ... the slope is not that big and the reward after is amazing, is just amazing. You don't have to think about that stuff later. And I hope your business is going to be powerful and I hope your application is going to have a lot of users, and you're not going to need to think about that.

Jeremy: Right. Yeah, and I love DynamoDB. And advice I give too is, "If you're thinking about massive scale, DynamoDB ... You're not going to get the performance out of a SQL database, no matter how much ... with MySQL or PostgreSQL, whatever you do." But with DynamoDB too, if you're using very small ... if you have a small set of data, but you need to query it a couple of difference ways, throw a couple of extra GSIs on there, it's so easy to do that. But anyways, so I'm going to put the link for this article in the show notes, because I do want people to go and check this out. But it is very complex, and it has to be, right? As things get more advanced, they get more complex. So, how can people start building this out? You're working on a boilerplate for this, right, that can just help people maybe?

Xavier: Yeah, exactly. So, in our company we do develop and we do Greenfield projects a lot. And so we have to have something that's strong and that's easy to use and that has all the quality tools and good practice and the best items from the get-go as well. So yeah, we are making a boilerplate for this purpose and it's like how to handle well a mono repository across microservices. Stuff like that. A lot of little steps that have a lot of value and you're really happy they're there, because it takes some time to set them up. And so we're doing that and we're thinking about sharing it with the community, so it's going to go ... it's going to come super, super well along with this article.

And then we have some specific services that we want to package as examples for us that are useful, so we think it can be useful to others, it's not that complicated. But there is some specificities around security, being able to push a file anywhere on S3 or being able to just put an index at the XTML at the beginning just as a hacker, to just make some fun. That's some steps that could happen if you don't know about some good security practices. So, we want to push some blogs, just like an architecture, as examples alongside this boilerplate. The web circuit one for instance as well is something we are thinking about.

Jeremy: Awesome. Well, that will be a huge help for the community and people building it. So, I am looking forward to that being available. All right, so I want to move on to costs, because that's another thing, you look at this typical serverless architecture, you've got a lot of individual Lambda functions running, you've got CloudFront out front that obviously the data, the fees from the data gets expensive there. You've got DynamoDB, which can get expensive if you have a lot of transactions and a lot of things happening there.

So, you do have serverless costs, or serverless still costs you money. But the question is, or I guess the question that gets asked quite a bit, and most of it is anecdotal, is this idea of is serverless cheaper? And you have a serverless calculator that you put together, and I want to get to that, but let's talk about what we have to think about, what are the costs involved when we're building a Cloud application, because it's not just about the hosting costs.

Xavier: No, it's not. It's not. Intuitively when you don't know about serverless, that's what you're going to think, you're going to think about your EC2 bill and then you're going to compare it with the corresponding bill in AWS with all the services you have in a serverless product. But it wouldn't really be fair to compare them in this way. We didn't talk about costs along the ... at the beginning of this whole podcast, because we knew we would arrive there, but that's also one of the big advantages of using this technology. So, we talked about those services, we talked about services that are scalable, we talked about services that already have pre-made, let's say prepackaged features, like Cognito. And if you took a step back, that's when you think and you can use the idea of TCO, total cost of ownership, which encompasses infrastructure costs, development costs and maintenance costs.

Xavier: And that's where you can really see the value of serverless. So, we talked about the infrastructure costs, indeed that's your big ... you can compare, that's easy to figure out, the cost of your infrastructure. And then maintenance costs, so it's going to be the time you spend to patch your systems or to make sure that in the middle of the night something ... to fix something in the middle of the night. And the last one, the development costs, is the time you're going to spend to plug an open source authentication bridge inside of your architecture and to make sure that it's working well with the rest of an architecture that was not out of the box made to go with it.

And the idea with serverless and all the services that went through together, is that they all work on those three aspects of TCO. So, when you look at the serverless deal at the end and all the services, you have to look though and think that they help you on the TCO itself. But then my question was, and it's the same for the first article about the particular architecture, when I first saw it, it was complicated, it's a new world, it's still a new world even though it's five years.

And to understand what's serverless architecture and then is it really cheaper, how do I make sure that it's ... how do I understand that it's really cheaper. There are some examples, there are some calculators, there are some articles talking about TCO, some great ones, there are some studies or use cases. But still I was missing something, so I wanted to work on that. That's why I worked on the calculator.

Jeremy: Right. So, I want to get to the calculator in a second, but I do want to clarify a couple of things too, because I think what you see typically when companies try to build out their own component, right, let's say they want to build their own security or authentication component. Part of the problem with building any system that is your own, something that is not unique to your company, and as we always say undifferentiated heavy lifting, right, it doesn't add business value to you for you to own your own authentication system.

The problem with building something like that on your own is not only the development time that it takes, if I can turn on Cognito or even Off Zero or something like that and I can immediately have secure logins and password reset flows and multi-factor authentication, all this stuff's built in for me, and it might take me, I don't know, maybe a half a day to read the documentation and set that up, maybe it takes a day. But it might take me weeks and weeks and weeks of development time to build that myself. Then here's the problem, let's say the person who built that, the lead architect, that person leaves and he or she goes to some other company and now your other devs have to be the ones to maintain a system that they don't necessarily understand fully.

Which means ... what always happens when they do that? Someone wants to rewrite it, make changes or whatever and then you've got more development time, more maintenance costs, things like that. Whereas if you hire somebody who knows how to use Cognito or knows how to use Off Zero, then that's just done for you. Yeah, you have to learn how it integrates into your system, but most of that is becoming very standard. So, that's a thing that I think a lot of people don't necessarily factor in as well, is that that ongoing maintenance cost. It's also about having the people available that know that, to maintain it. Because you can spend a lot of time having somebody relearn something, trying to figure out what some developer did three years ago, and is now unreachable.

Xavier: Yeah, I couldn't say it better.

Jeremy: All right. So, we talked about TCO, which is great. So, let's get into this cost calculator, into the serverless cost calculator. So, you built this calculator, it's in Google Sheets right now, right, so you can just make a copy of it and you can do some of that stuff. I know you said maybe you'll do it as a web app or something like that, but I think Google Sheets actually is really nice, because you can go in, it's an easy interface. So, tell us about this serverless calculator, what's different between like the AWS calculator, that AWS has on their website?

Xavier: So, the AWS calculator itself is extremely powerful, but you have all the services, the 250 services, I don't know how many. You can find it in the video we were mentioning. So, it's crazy, you don't know where to look at. And the thing I wanted is ... my focus is to evangelize a little bit about the added value of serverless to really help people, that's what I did, simply. And I wanted to create something that could help people that don't have crazy, crazy expectations of serverless, to be able to access the potential cost of a serverless project for their use case. So, I needed to think about ... And the AWS cost calculator is like that, it's complicated, it's difficult, it's micro.

So, I wanted to build something like that, so I thought what can I do for that, what's going to make a difference. There was a lot of articles, there was a lot of videos, but which one do I set, which one do I fix and then I don't think about them anymore. And that's why as well I thought about writing, it was just the first step, I thought about writing what is a typical serverless architecture. Because I wanted to sell a fixed picture and starting to set some viables for the potential calculator coming behind. And which services, how do they communicate with each other, when are we using them and stuff like that? So, build this image of what a typical serverless architecture is.

And then I went further and I made the calculator. So, the calculator it has a table ... and it's a spreadsheet, it's easier indeed to play around with it, it's easier indeed to ... The idea as well with the spreadsheet is to be transparent, so it's easier to be transparent with a spreadsheet. Everybody can see everything, the calculation and the hypotheses we made. So what are you going to find there? You're going to land on it and you're going to see the main data, that's where you're going to see, the price algorithm.

You have some basic variables on the left, you still need some, because you need variables these days, of course, and then the price output on the right. And below you're going to have the services. And then you have several facts, each of the AWS services we were mentioning before. So, the entire calculation based on AWS calculator costs and the variables that are necessary for each of them. So, we used a color code to show that there is some that's really an asset, because it doesn't have a big impact on the cost at all, because there is no reason to really tweak it, it's okay like that, your architecture is going to be fine on that.

And then there are some with a different color code that are going to be put back at the central dashboard, to be able to play with them and see the impact it has on the cost. So, you can use the calculator in two ways, you don't know ... you just want to play around, you have a broad idea of the website or type of website, an eCommerce, a blog, a work tool and with a certain amount of traffic. So, just that, so we have some drop downs and you can change them and directly see ... play around and directly see the impact it has on the customer per month. And if you want to go further, from there you have the capacity of changing some user variables, I call them like that because I think that's ...

That you don't know about serverless architecture, so you are not going to know how many numbers you're going to trigger, you're not going to know how long they're going to be triggered. So, I needed to pre-think and prepare those variables with an extra layer. So, here you're going to be able to play with user variables, like the number of sessions per average user, the percentage of users authenticated, so this is going to have an impact on Cognito for instance. So, average size of ... a product size, stuff like that, which are more product-oriented. So, I don't know, I'm a CTO and it's been a long time I didn't code, it happens, and I want to be able to assess the value of this technology for my company. That's my idea, to help him to be able to do that.

Jeremy: No, I mean, it's amazing. And I love the calculator, because it basically says, right, here's that typical serverless architecture, here are the individual components that are running. And then not only that, but trying to estimate well how many DynamoDB requests are you going to make or how many Lambda functions or what does it cost for the events. All these kinds of things. If you just look at 30 services, or whatever it is, it just gets really, really hard to estimate that.

And so, like you said, if you're a CTO or you're just somebody who's evaluating serverless, going in not only can you use it to actually give you some good numbers, but it's a really good learning tool to go dig in and say, "Oh okay, all right, so these would be the components that would be running. If I change this, I see how that affects the cost." I love the fact that you can bump up the number of Lambda invocations dramatically and the cost changes by like $3. But anyway, I ...

Xavier: And you can see the impact on the whole ... on the rest of the system, that's something I like as well, because those programs ... That's something I like as well. Which is still cheaper. We're going to talk about that after.

Jeremy: No, but it is, it's really great and those different predefined scenarios are also really helpful as well. Because even if it doesn't meet your use case exactly, it's going to be pretty close. So, let's do this, because you give a couple of examples in an article, you wrote a blog post about the calculator itself, which was a very helpful blog post as well.

So, let's go through some of these scenarios in terms of low, medium and high traffic, because this is one of those things where you did a comparison for just the cost, not including engineering time, not including maintenance costs, this was just straight compute or infrastructure costs. So, let's start with the low traffic one, so this was about 50 sessions a day. If it was an eCommerce site, what would it cost you in a typical serverless application?

Xavier: So, I wanted to start slow, because I wanted to compare easy accessible values, so indeed the cost of EC2, the simplest process we can find, there is the architecture we mentioned. So, if you do 50 sessions per day of an eCommerce app, you're going to pay $1.6 per month, so including the feature, but the feature's always free, so it's going to be like that for everybody. And if you do a blog for instance, so a little bit different in terms of user variables, it's going to be a little bit more of course, it's going to be $0.5. Okay.

So, I thought in front of that what ... if we're going to do serverless, what are we going to do, it's something super simple, but you need to have some compute, you don't have a choice. So, AWS for that has this too in EC2, the simplest one, the T3a.nano. And how much is the T3a.nano? It's $3.43. So, just this little picture is amazing in my opinion, you can compare, it's two times more expensive than the eCommerce example we mentioned and it comes naked, you don't have anything in this, you have to put your frame up, you can put your database in there if you want, but I don't think it's such a good idea. So, let's say you're going to use RDS to have an extra database and you will manage it. So, you need to add the cost of RDS, which I didn't even include there. So, it's ...

Jeremy: Right. Yeah. No, that's what I was going to say, is that what I thought was funny about this example, for $3.43, that is just running that T3a.nano for 720 hours a month or whatever it is. But you have no redundancy, so likely what you need to do is you need to add a load balancer, which I think costs like $25 a month, and have at least two of these things running so that they're there.

You also said, again RDS, like you would want to put RDS in there somewhere, even if you choose the cheapest RDS, just one server, I still think that costs you maybe $15, $16 a month, something like that. So, already you're exploding those costs. Now, again, that's pennies, right, we're not talking about a lot of money there. But what about if we move on to something a little bit bigger, something like a medium traffic, so where would we ... so 2,000 sessions per day?

Xavier: So, 2,000 sessions per day, same type of applications, eCommerce app, a serverless one, $70 per month. Okay? And the blog is going to be $20. Of course eCommerce is a bit more intense. Then I continued the same kind of comparison at this layer, it was if I'm capable of showing that EC2 is more expensive than serverless, I don't have to go further and dig into the TCO topic, because you don't even have to. And there I took ... from experience, and I asked colleagues and I took a M6g.large, it's one VM which is corresponding to this size, let's say of application and traffic.

This VM is $62 per month, so I said the AWS serverless app was $70 per month, the blog was $20. So, we're reaching ... the VM is slightly lower in cost compared to the serverless eCommerce app, but it's super close. And once again I said it was enough, because once again the VM cost is nothing. You still have to put your data, as you said, you still have to ideas and some background. So, no questions asked.

Jeremy: All right. So, let's move up to the big one, this is the high traffic, 40,000 sessions a day to maybe up to a million sessions a day?

Xavier: Yeah. So, there I reached a certain limit, high traffic and very high traffic, that's how I called it, and there indeed it's more complicated to compare, those systems are a lot more complex, there is much more blogs in there. And if I had to compare serverless projects, like the one we showed, with the same in Kubernetes, with all the crazy amount the pieces you can find there, it's a good one, and that's something maybe I'm going to do one day, but here I did not have this information. So, here I took the cost of course of a serverless eCommerce cost based on the calculator and it's 1.7 K, it's a $1,700 per month for high traffic. Okay? And a blog is going to be $475. Okay?

So, we're talking about 40,000 sessions per day, which is starting to be nice per day, starting to be a good thing. And then you can go super far and think about a million sessions per day. A million sessions per day, when you're there you're pretty successful normally. Normally you would say it's pretty successful. I don't think ... I don't know if I put a website that's close to this, an example of ... No, I didn't. Do you have an example of a good website that has a million sessions per day? I don't know, it would be nice to be able to compare.

Jeremy: Yeah. It would be good to compare that. But back to the 40,000 sessions per day, $1,700 USD per month for that, and then you had for, very high though, if you did a million sessions per day, you had $49,000. Is that right?

Xavier: Okay. Yeah. So, it's $49,000, so you are giving AWS $50,000 per month, so it's nothing to blow at, but at the same time your company in front of that is big, you have a million sessions per day, so you do have something that brings ... Generally when you're eCommerce, so you know you have set products or you're a blog, so you have advertisements. You need to reassess how that $50K enters inside of the business, what are the costs. And I don't think for the size ... it can be acceptable. But then to be able to really ... So, I did not have to compare this with EC2, there is no way of comparing this with EC2, so I didn't put a comparison in from here.

I needed to take a step back to think about this idea of TCO, total cost of ownership. So, I dug inside of studies, of articles, I read a lot about that, and a lot of those articles say that the amount of energy you put on non-specialized apps, it's going to lower, it's going to run by ... It depends on the articles, but it will really have an impact and it will give birth to less amount of time spent on that. There are a lot of examples of companies that were serverless that don't have Ops, they don't have Ops at all, they don't. But it doesn't mean that they don't need Ops, it just means that they use it, instead of the rest of their tech people.

But let's say, for the sake of the size, that you're not serverless, you have two Ops, okay, you have two Ops inside of your team and you patch servers, so you reduce by two. Let's say you reduce by two the size of the Ops team, because half of the staff are not necessary anymore. If you take, for example if you take the salary for the company, of course it's always ... it costs more for the company than what you see at the end on your payroll, of an engineer, it's a $140K per year. Okay?

So, with serverless you're going to go down to one Ops, so it means that you're going to save $12K per month. Okay? So, this team of two Ops, it was for a high traffic website, for the one that had 40K sessions per day. We were talking about an AWS bill of $1,700 per month. So, it means that you saved $12,000 per month and you're paying $1.7K per month straight away. So, there is a saving of $10K in the middle of that somehow. So, if you think about TCO, of course it's always more complex than that, but the idea is here you think about TCO, even when you go further and you go to extreme traffic, it can have ... Once again ...

Of course, we are not talking about Facebook or Twitter, that's an extreme. And if you reach this, it's amazing, you don't even have to read this article if you reach this anyhow, that's for sure. And there of course for a lot of reasons you're going to have your own team to do that and it's going to be huge, it's going to be amazing. But for the most of people, for us, for the ones that read this article, it's showing that indeed serverless can be cheaper and can be amazing for their usage. That's what I wanted to share.

Jeremy: Yeah. No, and I think you're right. And then there was ... you referenced this paper on TCO by Deloitte, I will put the link in the show notes for this, but I think that that's a really interesting point, right, just the idea of reducing the number of things you have to do that are just maintenance. Again, it adds no value to your company to have somebody installing patches on a server somewhere or worrying or being on call. Now, again, you get to Facebook level or maybe you get to a point where you're at that million sessions per day, maybe it ends up being cheaper just to have Ops people that are constantly watching containers and a Kubernetes cluster and things like that.

But until you get to that point, you can save yourself a lot of money by reducing the number of people that you need working on the Ops side of things. And I want to say this, because this is something I know that a lot of people maybe worry about, is to say, "Well, what happens to the Ops jobs?" So, we just say, "We get rid one of the Ops people." And that may be true, maybe it's better to say, "We just won't hire more Ops people." But what would you do as an Ops person in an architecture like this, right, because there's still more things that can be done?

Xavier: There is still a lot of things that can be done. Of course, you have the change of observability that gets bigger, that gets greater, you have a lot of different moving pieces and you need to understand what's happening there. So, there is some more work to put inside of observability, which is a critical field to dig into. And for me, it's also a Dev Ops task. And then you have a lot of security concerns and there is always going to be some security concerns. And as well you can spend some more time to think it over and improve your security, which is good for everybody, because there is always leaks. Every week we have news of a new leak. And so you spend some more time on those kinds of tasks.

Jeremy: Right. And CICD, right, just thinking of getting things into production?

Xavier: Yeah indeed, that's amazing, thank you for that, I forgot about this one, which is a big one. Yeah, CICD, CICD is still super important and you want your developers to be as fast as possible and to have as less issues as possible during the development and testing and validation phase. Serverless gives you the opportunity, like really easily, to start feature implementing, just with CICD you are capable of ... with a simple computation, to create future environments on the go, like that, for testing new developments in an isolated manner.

And this ... it's amazing, honestly before serverless, I worked on projects where we had that, but it was a hassle and for big projects we have to reach a certain level to be able to ... or to want to spend energy in building them. Now, with serverless, finally you're capable of doing it. And that's amazing. So indeed, your Ops team is going to be able to do this fancy stuff that you dreamed up even more easily.

Jeremy: All right. So, if you're an Ops person, don't worry, just evolve and you can work on things that are much more exciting than just patching servers and worrying about your VPC configurations. All right. So, finally, so next steps on the calculator, right, so this is ... you mentioned at the start that you've got some other things that you want to do with it, so where is this going to go?

Xavier:
So, I'm going to continue using it, I'm pushing it internally, but I want to go further. So, I want to be able to challenge it based on some real use cases. So, I have had inputs, I'm working with some fellow community members, which is super nice, to be able to do that. I will challenge the data we have to make our calculator better. And then we can add some of the services, like AppSync for instance, because AppSync is pretty popular, and getting more popular, and I didn't put it inside of the calculator.

So, give a little bit more options to the table. I need to find the right balance, because the idea of an opinion is to have an opinion. The opinion has value in itself. So, if there's too much options, I'm going to come back to the AWS calculator and ... I need to find the right balance, but some options is good. And then, depending on that, next step could be potentially to build an open source website for it. I think it could be pretty fancy.

Jeremy: So, listen Xavier, thank you for your opinions, because your opinions on this have been great. Excellent tool, the article was great, I think there's a lot of learning, just to put it all out there and having your experience and sharing that with people is amazing. So, I know I appreciate it, I'm sure the community appreciates it. And again, thank for being on the show and sharing all this knowledge here. If listeners want to find out more about you, how do they do that?

Xavier: So, Twitter, of course, first. You're going to give it after, I guess, it's ...

Jeremy: I will put it in the show notes, yeah.

Xavier: Yeah, you're going to put it ... But I'm going to say it out loud, it's @xavier_lefevre, X A V I E R underscore L E F E V R E.

Jeremy: Okay. Right, I spelled your name right, awesome.

Xavier: On Twitter, on LinkedIn, of course on Medium and I'm planning on writing articles, but even more than that, I'm planning on updating my articles, I want those articles to stay up to date, I want this particular architecture to still be actual in six months. So, I'm going to work on ... and I have some ideas already, I'm going to work some on an update of my articles, so you can follow it along on Medium.

Jeremy: Awesome. Well, I will put the links that we mentioned in the show notes, as well as all of your contact information, and the serverless cost calculator. Thanks again, Xavier, it was awesome.

Xavier: Thank you very much for inviting me. So it's me that thanks you.

This episode is sponsored by Amazon Web Services!

View Details

About Alex Casalboni
Alex Casalboni has been building web products and helping other builders learn from his experience since 2011. He’s currently a Senior Technical Evangelist at Amazon Web Services, based in Italy, and as part of his role, often speaks at technical conferences across the world, supports developer communities and helps them build applications in the cloud. Alex has been contributing to open-source projects such as AWS Lambda Power Tuning, and co-organizes the serverless meetup in Milan, as well as ServerlessDays Milan (previously JeffConf). He is particularly interested in serverless architectures, ML, and data analytics.

  • Twitter: https://twitter.com/alex_casalboni
  • LinkedIn: https://www.linkedin.com/in/alexcasalboni/
  • Website: https://alexcasalboni.com/
  • Dev.to: https://dev.to/alexcasalboni
  • Medium: https://medium.com/@alexcasalboni
  • AWS Lambda Power Tuning Project: https://github.com/alexcasalboni/aws-lambda-power-tuning
  • Deep dive: finding the optimal resources allocation for your Lambda functions: https://dev.to/aws/deep-dive-finding-the-optimal-resources-allocation-for-your-lambda-functions-35a6
  • Test Functions/Patterns for Power Tuner: https://gist.github.com/alexcasalboni/9ce2cef56a7d052d4f5e798b37083525

Watch this episode on YouTube: https://youtu.be/m2NB_0J5fms

Transcript

Jeremy: if we go back to the sort of the having to run this on every function, and you maybe make a change to a function, you do something like that, it just becomes very, very tedious, and probably a lot of work to run this on every single function. But as you start to see... You run it on a few functions, maybe different types of workloads, those patterns start to emerge, right?

Alex: Absolutely. There is not an infinite set of patterns. I have identified about six or seven. Usually, you end up in some of these. There are other patterns where the output is a bit randomic meaning there is either some downstream dependency that is not scaling as well as Lambda is. So you might see some noise in the data. But yeah, usually you end up in one of these categories. I think there is a last category that is a bit spatial where you actually are downloading or uploading a lot of data. I've seen this with S3 maybe you need to download 50 or 100 megabytes of data from S3. I wouldn't recommend you. But if you really have to do that the power tuning implications are very interesting, in my opinion, because if you also change one line of code, and I did this experiment with Python. So, it was pretty easy. I think you can do also the same with Node or Java or other SDKs.

So, if you enable the multi threading options in the SDK especially at a high power like above 1.8 gigabytes you get two cores. And so, you just start downloading or uploading using 10 or 20 threads, you actually see a massive difference there. So, you might see 10% improvement for cheaper cost. So, if that's what you're doing with Lambda, you might consider full power. But again, check the numbers.

Jeremy: Right. And the other thing too is, again, knowing... You mentioned Python, knowing your runtime is important because node is single threaded. So, even if you do go over the 1.8, you do not get a second core because Node doesn't work that way. All right. So, you mentioned something really interesting. I think this is another fascinating thing about pay per use services is Lambda has a 100 millisecond billing model. So, if you run something for 99 milliseconds you pay for 100 milliseconds. If you run something for 101 milliseconds you pay for 200 milliseconds. So, I think an important piece of this, if you are trying to optimize for cost, is also understanding that billing rounding thing, right?

Alex: Yeah, that's true. I've been talking to some development teams. And it's very common that you develop a service application, and you end up as you were saying with 10, or 50, or 100 functions. And one day the manager wakes up and wants to optimize for cost or for performance. And you're like, sure, but where do I start? I have 100 functions. But I think it's also important to know what your functions are doing to detect the right pattern and to know where it makes sense to optimize. There are cases where your team or yourself or your manager may want to only optimize for cost. It's a cost optimization project, whatsoever.

And you might end up optimizing some functions where there is no way that you can actually shave off enough milliseconds to go down one level, 100 millisecond level. So, maybe you're just optimizing for the user experience, which is great. Or maybe it's not a customer facing app, so it doesn't matter. But I think it makes sense to understand how cost and performance are related in serverless. Because sometimes they are aligned to each other, meaning you can optimize for both just in one shot. But you still want to be aware, especially if it's about prioritizing between a large set of functions. Actually, I got that feedback a lot. I think if I see a direction of Lambda Power Tuning evolving into something that would help a development team handle multiple functions I'll build something like a prioritizer error or something that helps you detect those kind of functions more easily or to help you with a batch of functions, for example.

Jeremy: Right. Yeah, no, I think that'd be another very cool project. I think you asked the right question there. And then at least, this is something for me is sort of what are you optimizing for? Are we trying to make the user experience better, so we have lower latencies? Are we trying to get the cost down on the backend, maybe for running ETL tasks and things like that? Those are certainly things where I think this comes in. Really, this is an important thing to consider is to say, "Are we trying to save 10 bucks a month from our front end just so that we're, again, saving $10 a month, but maybe it takes 120 milliseconds for our API to respond. Or are we trying to save potentially thousands of dollars on the backend if we're running these complex ETL tasks. And that brings me to something... So, Joe Emison who runs Branch Insurance, every month he usually post a screenshot of his bill. And running an entire online insurance agency his Lambda bill was $22.65. So is optimizing for cost in that situation something that should even cross your mind?

Alex: Yeah, that's a fair question. It probably isn't. I would still suggest you run Lambda Power Tuning because you might be willing to pay $26, and get a 30% performance improvement. So, it's not like one or the other depending on your scale, depending on the use case, depending on the customer needs you might decide to invest on performance. And with Lambda it's pretty simple. You can visualize it. You have one knob, and it's fairly simple. So, as I was saying, it's almost free to run this power tuning process. I've seen customers who actually run it at every deployment basically multiple times a day, and it's still less than $1. So why not?

Jeremy: Yeah, yeah. And I totally agree with you. I think the performance aspect of it is the biggest thing to optimize for. And when you see tweets or blog posts that criticize serverless performance, it's often because they don't have a knowledge of what it is that is possible to tweak. But again, again, going back to the idea of being able to measure that is an important component. So, let's get into more specifically some of these things you can do because these things that you do to these Lambda functions then running...because again, if you just run Lambda Power Tuning, and you see, okay, this cost, or I can get better performance if I turn the memory up. That is not the only way to get better performance.

That's not the only way to optimize your cost. There are so many other things that you could potentially do that would bring those things down either lower the execution time, or some of that. So, let's get into those. And I know you do a lot of presentations, all virtual now, unfortunately. I mean, again, I wish we could all get back doing presentations again, but I know that you like to break these things down into a couple different categories of optimization. So obviously, there's sort of our general optimizations that we can do, but then we have things that are very specific for cold starts. And then other things that would require you maybe to re-architect your application. So, maybe you can explain how you approach those categories?

Alex: Yeah. Sure. So, it's what you're concerned about is cold start, we talked about it at the very beginning. There are a few things you can do there. It's not likely to be the majority of your executions. But if it's customer facing you do want to optimize for a cold start. It's about monolith, avoiding monolithic function, you can optimize your dependencies, or rather minimize your dependencies. In some languages, you can minify or uglify your code. You can try to initialize some objects in a lazy fashion depending on what libraries you're using. You can optimize how you import the SDK components in our individual clients.

These are all things that allow you to shave off maybe 10, 50, even 100 milliseconds of cold start execution time. So, definitely worth having a look. Although I always recommend people, but don't stop there. That's probably 5 or 6% of your overall executions. You also want to optimize all the others. So, to optimize all the others, usually you either have to re-architect everything or rethink some components or some parts of your architecture. And that's great. If you can do that sometimes. Unfortunately, you cannot do that or you might decide it's not worth it. Many reasons why you may or may not want to do it.

Alex: Likely there are some low hanging fruits that you can actually take without re-architecting anything, basically, or refactoring your code, basically. We have already talked about one. One is memory optimization, resource allocation, zero code changes, zero architectural changes. We have talked about what you might expect depending on the pattern. There are a couple more. I think, if I remember correctly one is the keep alive option in the SDK. Often, you don't want to re-initiate a connection every time you want to talk to DynamoDB or Cognito or some other AWS service and you can just keep that connection alive just with a one configuration parameter or one environment variable. So, that's a pretty easy low hanging fruit you can see massive impact.

Not in the cold starts because if it's a cold start you will actually see the impact of creating the connection. But you will see a large impact in the remaining large percentage of your executions. There are a few more very specific too. Some runtimes are very specific to some use cases. But that's the way I like to think about it. The rest of the optimization strategies I'm aware of, unfortunately, kind of require you to rethink of some part of the architecture. We can talk about some of these if you want.

Jeremy: Yeah, so I mean, so one of the things I think that is sort of really interesting about optimizing once you get past the cold start thing is every time you have to make a network call, every time you have to do something that requires some sort of synchronous call, you are not only paying for that execution time, right? But you're also adding extra libraries into your code in order to make that happen. So, one of the cool optimizations, or I guess, maybe people might not think of it as optimization or as an optimization. I certainly do, though, are Lambda destinations. Because to me, it's like if you have an output of your function that needs to go somewhere having to put all that extra code and wait for the call to EventBridge or wait for the call to SQS or SNS, that if you don't have to do that. And you can use Lambda destinations to do that for you. I think that's a big optimization right there. I mean, maybe not a big optimization, but certainly interesting.

Alex: But there are many cases where your average execution time is slightly above 100 millisecond interval, 105 milliseconds, 110 milliseconds. And many developers ask me, "How do I shave off those five milliseconds? I can't touch my code further." And so, there are cases where, yes, you can just delegate to the Lambda service, the invocation of S&S or EventBridge or the destination that you want to invoke at the end of your execution. And not many people think of it as a cost optimization or a performance optimization. Overall, it's not like Lambda can do it faster than your code would do. So, it's not really a performance optimization. But because you're not paying for the execution time of that API call you might be able to shave off those five milliseconds. So, in some cases it might show some benefits for sure.

Jeremy: Right. Yeah. And so, another thing too that can really cut down time, especially when you are doing warm invocations is reusing any sort of global variable or connection or thing. If you mentioned the HTTP keep alive, that's great. But you're not really maintaining a connection, in the same way you would maintain a connection to say, an RDS cluster or something like that. So, the ability for you to... And you also mentioned lazy loading in there, which I think is another interesting thing, where you don't necessarily have to connect to the MySQL server when there's a cold start. And maybe that function doesn't need to connect to it until it actually needs the connection. But once you have the connection, having that global reuse of those variables I think is another way that again you don't have to keep reaching somewhere to rehydrate state.

Alex: Yeah, there are other cases too. Like if your runtime configuration parameters that you are fetching from Parameter Store or Secrets Manager. So, you don't really have to fetch those at every single invocation. You can just cache them locally. And as long as you are fairly sure that the value of those parameters did not change, unless it's a new deployment, or situations like that you don't really even need to check or to have an expiration time for that caching mechanisms. There are cases there where maybe it's a database password, and when you rotate it, the next query is going to fail. Because that value is not valid anymore. So, you may want to have some kind of retry mechanism to be able to detect an online invalid password error, and then just go and refetch the new password because it rotated if you're using Secrets Manager. And then just go on doing what you were trying to do. So, there are some more interesting cases there, too.

Jeremy: Yeah, and I mentioned RDS proxy too, that... Or I don't think I mentioned. I mentioned RDS. I didn't mention RDS proxy. That is actually another thing where you might not necessarily think of it as an optimization. But if you do not have to keep retrying connections, and you can get that connection pooling on the backend, so that you're minimizing the amount of stress on the database because you're using connection pooling. Those are all additional optimizations that could actually make query results come back faster.

Alex: Actually, RDS proxy is going to fetch the secrets from Secrets Manager for you. So that's...

Jeremy: Also that.

Alex: ...less code to write and also less execution time of Lambda itself to pay for. So, yeah, the RDS proxy is very powerful. Also, if you think about the resiliency if Node goes down you don't have to re-initiate another connection. It will just migrate a connection to another instance. So, it's pretty powerful.

Jeremy: All right. So, let's talk about a couple of best practices, right? So we talked about some ways to sort of tune or to optimize different things. But there are other sort of, I guess, best practices that can also optimize performance that can save on costs. Things that maybe aren't so much tweaking things. Just sort of general, I guess, concepts. And the first thing would be orchestration, right? If you're doing some sort of external orchestration, why do we use Step Functions as opposed to trying to write a Lambda function to do that for us?

Alex: Yeah, that's a that's a fair question. Actually, as a developer five years ago I would have told you I love to do orchestration in my code. It's simple, I don't have to pull in other services. I would probably do everything inside my Node.js application or Java application. But there are benefits to it, especially if you consider cost. As you were saying before, every time you are invoking an API or idling. In Lambda, idling means you're paying for nothing, right? You're paying for waiting. And one of the best things I love are Step Functions that we have mentioned slightly because Lambda Power Tuning is based on Step Functions.

One of the best features of the function is that you have... Actually, two best features. One is the wait state that allows you to wait up to I think a year without paying for idle. So, that's great if you have asynchronous stuff, or if you need to wait for human interaction and stuff like that. But also you have the ability to coordinate concurrent tasks that will converge into a final decision step maybe. And that's usually why you need to wait. Maybe you invoke three APIs, but one takes longer than the other. And you need some kind of coordination there. So, Step Function, you can do it. It's just a built in feature. You don't have to wait. You don't have to pay for idle in either of those three concurrent branches. So, pretty cool.

Jeremy: Right. Yeah. And waiting, the wait state thing is probably the best. And what I would love to see is, especially with longer running transactions, or longer running API calls. I know I have an API call that sometimes runs up to 25 seconds to do natural language processing. What would be great is if I could send my payload disconnect, and then wait for a webhook response when it was finished processing, and then avoid even more of that wait state in there, but that's maybe a different topic.

Another one, though, that I think is important, and this has to do with architecture and how people think about moving data around. And this is something I think Chris Munns said years and years and years ago. I don't know if it's from him. But he says, "You want to transform not transport data with Lambda function." So what does he mean by that?

Alex: So, it doesn't apply to every use case possible. I think there are cases where you need to fetch data from somewhere. But it could be at a relational database. It could be S3. It could be some service that has some nice filtering functionalities. So, usually, you want to fetch the least amount of data into Lambda because that means less by in the network, less idle time, less I/O time, basically. So you want to use Lambda to modify data, to manipulate data.

Ideally, you probably want to get the data directly in the event, instead of having to go and fetch it. So, if there is a way to do that. If there is a native trigger that will give you the input. Or if there is a way to get the data you need instead of go and fetch it, do it. But also, sometimes there is a better way to do what you're doing. For example, there are many situations where you need to fetch data from S3. But you don't really need to fetch the whole object. It's like if you are doing a select all from your RDS database, instead of using a where clause, right? You don't want to fetch the whole database, and then do the filtering in your application code, you want to do the database do the heavy work of fetching exactly what you need, so that you can minimize the network device over the network. And you can do the same in many situations, for example, with S3 Select. So, it's kind of like a database. You specify an SQL query, and you can fetch data out of a large S3 object without downloading the whole thing. So, if you are in a case like that where you're downloading a lot of data, it's very likely there is a way to only fetch the data you need and delegate that computation to another service.

Jeremy: All right. And you mentioned getting events, having sources or certain systems that will send events into your Lambda functions. And one of the things that you see a lot, especially with certain S3 events, and whatever is that some of those events are going to be uninteresting to your Lambda function. And every time your Lambda function responds to an event, you're going to pay for that processing. Even if you discard it. Even if you say, "Oh, no, I don't care about that event." You still pay for the invocation. You still paid for the hundred milliseconds at a minimum for it to just say, "I don't want this event." There are better ways to do that.

Alex: Yeah, absolutely. Usually, in the train and native trigger of Lambda you can probably add some kind of filtering. Whether it's S3 or SNS, or other kind of custom events in the AWS platform. If you can filter those out in the trigger configuration, it means you're not going to pay for that. And this is typically not a huge issue. But there are cases where, for example, you want to allow all your customers to upload files, but you only want to process images. Well, there are all things client side you can do to avoid them upload PDFs files, but sometimes they will do it anyway. So, you really don't want them to reach your Lambda functions and do a denial of wallet kind of thing to your architecture. So, you really want to discard those events as soon as possible, usually in the trigger config.

Jeremy: Right. And then another way that you can potentially save some money, or you can optimize, I guess, how often your functions run is by implementing things like throttling. Whether through an API gateway or maybe adjusting your concurrency, so how do you manage that? What are ways to kind of figure out what the right concurrency is or how much you should be throttling data to your application?

Alex: Yeah, I wish I had an answer for that, that we could discuss in a minute. It really depends on multiple things. It could be your business model. It could be your SLA. It could be a lot of things that doesn't allow your customers to invoke you 1000 times per minute, or per cycle, or per hour. You might have a freemium business model where free accounts can only invoke you once an hour, or once a minute. And so, those decisions are not really about, we want to make this thing as cheap as possible. It's more about you want to avoid misuse of the service or you want to avoid abuse as well of the service.

So, there are other things that you may not want to do for other downstream reasons. Like you may not want to delete more than 10 records from the database per second, stuff like that, just to avoid race conditions, or to avoid more problems somewhere else. Usually, it's not too much about saving money or those things like if you're scared of a DDoS attack you probably want to use WAF or AWS Shield or something that protects you at the edge, and not too much on the API layer or Lambda layer. But you can do that.

So, if there are good reasons to set a maximum concurrency for a single Lambda function, or for an API endpoint or a single route you can specify that at the API gateway level, or even at the individual Lambda function level. Always remind you though that you have a limited regional concurrency for all your functions. So, the sum of the concurrent executions is bound to that limit. So, there are the situations where you want to allocate a given concurrency to one Lambda function, so that the concurrency of the others is not going to affect the availability of that function. But again, it's not really much about cost. It's more about resiliency and availability.

Jeremy: Right. And I mean, if you're thinking about tenancy and multi tenancy for maybe your freemium, but then also you want to split up the tenancy for large paying clients and things like that, then putting them into different accounts and stuff like that you can control. That's another way that you could potentially optimize that. All right. So, another thing I think is super important is just as a really good best practices. Let's say you go through the power tuning exercise, you want to make sure that you take that information and you bake that into repeatable deployments. So obviously, infrastructure is code, huge best practice.

Alex: Absolutely. Yeah, I think nowadays you cannot do a lot quickly and reliably and securely with our infrastructure as code. I still meet a lot of developers that do not do it. If there is something you want to invest on in the next six months as a developer learn, and infrastructure is code framework. For serverless you have a lot of options. I meet more and more people that are in love with Terraform, or are in love with LPSN, or the serverless framework or the CDK. To me, it doesn't matter which one you choose as long as you choose one, you learn it, and you make it your default in your organization. I've met organizations that use multiple infrastructure as code tools. It's okay. I've used many in my career as well. They're all different and all equal in a way. Some are more vendor neutral, some are more community focused, some are more provided by the vendor. So, pick your religion here and see which one works better for you.

Jeremy: Right. All right. So, last one, and I think this is an important topic, is the idea of observability in your application. So, AWS' X-Ray, there's a ton of other observability tools. But why is having something like X-Ray such an important component in your serverless applications?

Alex: I think it can give you a hint into what the hell is going on when something goes wrong. That's the simple definition I can give you. There are many cases... I was actually talking to a customer a few weeks back there was using just going back to Lambda Power Tuning for a second. And they were seeing completely random results like, "Hey, this thing doesn't work. Every time I run it, it's different." So what I told them is, "Hey, turn on X-Ray, and see what's going on." So they had a legacy system downstream based on RDS single instance, single AZ, and they were testing like 2000 concurrent executions. So, that's never going to work. But somehow they didn't have that architectural diagram in their mind, and they didn't know what was going on.

So, that's a typical situation where a new person comes up or someone who is responsible for optimizing for cost, and they have no idea what are your downstream services or what might go wrong in the overall architecture. So, having visibility into that is the only way to fix the problem sometimes. And if we were talking in 2015, it was a hard problem in the serverless space. I think now you have a lot of options out there, a lot of even community heroes, and community leaders from AWS, and from other vendors as well. So, I wouldn't say it's a solved problem. But you can also pick your religion or your platform of choice.

Jeremy: Right. Yeah, well, and I think the most important thing, especially looking at instrumentation or looking at observability when you're using something like power tools, limited power tuning, and you're trying to figure out what is the most... How can you optimize it? If you keep saying, "Well, if I keep turning up the memory, keep turning up the memory, and it's not having an effect." It's important to be able to go and look and see, well, how long do these API calls take? You mentioned in the third party in the downstream stuff. So, if I'm calling some external API, there may be no way for me to shave time off of that. So, if you know that it doesn't matter if you have three gigs or you have 128 megs. It's still going to take 1.6 seconds or whatever it is for that API to respond to you on average, then there's really nothing you can do to optimize that. But that's important information to have just as a sort of holistic picture of how to do all this optimization.

Alex: Yeah. Well, if you are in a situation where each function is only doing one thing, maybe talking to only one downstream service, that's pretty easy. You might live without observability. The thing is, it's quite common that you are either reading from Dynamo, and then putting something into SNS or into Kinesis or whatever other service. If you're doing two or three things, and something's slowing down, and there is an error and what went wrong, maybe you're doing three things in series, and you can visualize it in the X-Ray trace visually. And you say, "Well, there is no reason why I shouldn't do it in parallel," and you can compress the execution time and do it much faster. So, it gives you visibility into what's going on and for different reasons, it might be very useful. For troubleshooting, for optimization, even just to have a nice picture to post on Twitter.

Jeremy: That's always a good reason to be able to share your X-Ray waterfall screens. All right, did we miss anything? I feel like we covered a lot of information here. I mean, we mentioned the stateless functions thing right from the beginning. Oh, actually, I don't think we mentioned stateless functions. We should talk about that for a second. That's another optimization. And you did mention not trying to hydrate stuff every single time. But again, serverless is sort of meant to be stateless, right? Why are stateless functions a good optimization?

Alex: Well, it's not like you can take a stateful function and magically convert it to stateless. It's still about where is the data coming from? Why am I depending on state during the execution? Or why am I not reading the state from somewhere else? So, if you have a stateful application that, for example, if you have three EC2 machines, and they rely on some sticky session mechanism, so you're talking to the same customer, and the session is stored on the EC2 machine instead of a Redis or a Dynamo, you might have that problem.

I think if you are developing with Lambda, it's less likely that you encounter such a situation where you're relying on sticky sessions or other stateful mechanisms. Usually, you don't really have a lot of storage or a lot of memory to store your state long term anyway. So there are still interesting things you can do at the design time. So, instead of using Redis, or Dynamo, or MongoDB to store state, you might say I will inject state into the execution because data is coming from external service that is invoking my function.

So, that's a way to make your functions completely stateless. And you only add the business logic. You don't even have to know where the data is coming from or where it's going to go next. So it's much easier. And I think here optimizing for cost is not only about execution costs. It's also about the cost of re-architecting an architecture. The cost of extending something in the next six to 12 months. Especially, if you're using orchestration with Step Functions it's much easier to inject state instead of fetching the state from inside the function. So, if you put all these things together, I think it will come natural to design a function that is stateless. Does that make sense?

Jeremy: Yeah, no, it does. Because I'm a huge fan of doing that where you're injecting the state along with your payload. I love JSON Web Tokens now if you're doing something at an API level because you've got that signed bit of information where even if it's just a user ID or something like that, that's passed in, it's verified. You know that, that token is valid. And you can use that ID as a way to save data or whatever. It's a much more optimized way of doing that if you don't have to make that separate I/O call. Awesome. All right. So, again, I'll ask again, did we miss anything? Again, just so much information here, but I think we covered pretty much all of it.

Alex: Yeah, I think we are good. New things might come up in the future. No spoilers, of course. And if we missed something maybe thing goes on Twitter or LinkedIn, I'm happy to learn from the community as well.

Jeremy: Right. So, if something comes up, and people want to get a hold of you, or they want to learn more about power tuning, Lambda Power Tuning and stuff like that, how do they do that?

Alex: So, you can find me on Twitter, Alex Casalboni, we'll add a few links, and also LinkedIn. Those are the two platforms I use the most. And while the project is on GitHub, we are going to add a link as well, I think. And I might actually go and write down a blog post with actual images about all of these, especially all the different patterns and the different visualization scenarios. So wait for it.

Jeremy: Awesome. All right. And then you've got your website, alexcasalboni.com. You write on DEV.to. You write on Medium. There's a test function/pattern thing. I'll include that in the show notes as well. Alex, thank you so much. Awesome information as always. Hopefully, I will get to see you in person again at some point. Probably not this year, but maybe next year once we have 2020 in our rear view. But thanks again, Alex, I really appreciate you being on.

Alex: Thank you very much, Jeremy. And thanks for all you're doing for the serverless community.

View Details

About Alex Casalboni
Alex Casalboni has been building web products and helping other builders learn from his experience since 2011. He’s currently a Senior Technical Evangelist at Amazon Web Services, based in Italy, and as part of his role, often speaks at technical conferences across the world, supports developer communities and helps them build applications in the cloud. Alex has been contributing to open-source projects such as AWS Lambda Power Tuning, and co-organizes the serverless meetup in Milan, as well as ServerlessDays Milan (previously JeffConf). He is particularly interested in serverless architectures, ML, and data analytics.

  • Twitter: https://twitter.com/alex_casalboni
  • LinkedIn: https://www.linkedin.com/in/alexcasalboni/
  • Website: https://alexcasalboni.com/
  • Dev.to: https://dev.to/alexcasalboni
  • Medium: https://medium.com/@alexcasalboni
  • AWS Lambda Power Tuning Project: https://github.com/alexcasalboni/aws-lambda-power-tuning
  • Deep dive: finding the optimal resources allocation for your Lambda functions: https://dev.to/aws/deep-dive-finding-the-optimal-resources-allocation-for-your-lambda-functions-35a6
  • Test Functions/Patterns for Power Tuner: https://gist.github.com/alexcasalboni/9ce2cef56a7d052d4f5e798b37083525

Watch this episode on YouTube: https://youtu.be/31LHFQ1lT78

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today, I'm chatting with Alex Casalboni. Hey, Alex, thanks for joining me.

Alex: Hi, Jeremy. Thanks for having me.

Jeremy: So, you are a senior developer advocate at AWS. So, why don't you tell the listeners a little bit about your background, and what you do as a senior developer advocate?

Alex: Sure. So, I come from the web development, software engineering, and also startup world. And I combine that and I try to help customers using AWS, discovering all the different services and I used to travel the world. Now I do a lot of virtual conferences.

Jeremy: So, as a DA, I know you've been working a lot with serverless. And one of the things that you publish a lot about, you got an open source project that we'll get into about this. But you do a lot with optimizing and tuning the performance of Lambda functions. And I think a lot of people sort of assume that all of this stuff is done for you maybe that it's just a matter of putting your code up there, and it just automatically does what you need it to do. Now, if you look at some of the surveys and look at some of the research data that shows that people just typically use the defaults, I think that people do assume that quite a bit. But there are ways to optimize your Lambda functions. So, what are some of the sort of main things that we have control over when it comes to optimizing Lambda functions?

Alex: Sure. So, there is a lot that the Lambda team, and actually the AWS team is doing to optimize the service itself for performance, to make it faster, to make it more reliable, to make it cheaper, eventually. But of course, it's a service and you can configure it. And with every configuration, you can fine-tune it for your specific use case, right? So, whether you are developing a RESTful API, or an asynchronous service for ETL. Whatever you're building, you might have very different needs in terms of performance or you may, you need to bring the cost to as minimal as possible. So, there are many things you can do. And maybe we'll talk about some of those. I got very passionate about this topic because I keep meeting a lot of developers that are just mind-blown when I tell them about the power-tuning side of things, and how they can actually get a lot more performance, and sometimes even a lot less cost. They can make their functions cheaper just by tuning the memory of their Lambda functions. So, that's what I'm really passionate about.

Jeremy: Yeah, so I mean, and if you think about building any type of application, I mean, obviously, there's a couple of major components to it. I mean, you have to think about latency. You need to think about throughput, especially if you're transferring large files back and forth from S3. And you have to think about cost, right? Cost is always sort of an important factor. So, what are some of the things, maybe we start there. What are some of those things? Do you have control over those three factors? Are there ways besides just the memory manipulation that you can really focus in on those?

Alex: Well, when it comes to the speed itself, the execution time itself, you actually do have control, and there are many things you can do to speed up the average execution. We can talk about cold starts or the large majority of your execution as well when it comes to Lambda. There is also a lot you can do to optimize for throughput. Usually, it's something you do at the architectural level. It's not just like a configuration parameter where you increase throughput. There are services like that. Like, I don't know, Kinesis Streams where you have some configuration that allows you to have more throughput on a megabyte per second, maybe.

But for Lambda itself, you don't really have a parameter to increase throughput. All that is managed by the service. What you can do is at the architectural level, maybe if we have a 10 gigabytes of data to analyze instead of doing it in series 10 megabytes at a time, you can do it in parallel. So, that's an architectural change, that pattern change, but it's not really a configuration of Lambda itself. It's more in the way you use all these services together, in my opinion.

When it comes to cost you also have control over that. There are many guardrails you can take to control the span, to control the parallelism of how many Lambda instances you can add. You can set it to zero if you want to stop everything. So, there are many ways you can control costs. What happens usually in the wild, that's what I see is that most developers I meet are more interested in optimizing for performance just because cost itself of Lambda is a relatively small percentage of their bill. So, they're not so much concerned about, "I want to make a Lambda as cheap as possible." Usually, they want to make the overall architecture cheaper. And they are happy to pay a little bit more for Lambda, if they can make it faster.

Jeremy: Right. Yeah. All right. So, let's talk about cold starts because this is the thing that always comes up. I think there's a lot of confusion around how often you get cold starts, how much of an impact they actually have, when they happen and things like that. So, let's start there. How much of an impact do cold starts actually have on your application?

Alex: It kind of depends on the type. I would say there's no typical answer, I know, I'm sorry. But it depends how often you involve your function. It depends how often you're scaling out to additional run times of your Lambda functions. So, depending on the use case it might be 1% or five or 10% of your executions. If your function is a daily cron job, it's probably 100% of your execution.

So, it really depends. In my experience, unless you really have like one execution per day or one execution per hour you can typically think of it is a relatively small percentage of your executions could be one or 5%. That means typically, it doesn't matter that much. And I understand why many developers are concerned about it because when you go and test it in the console, that's all you "ah, it's low." But at scale, and scale here means if we have more than 10, or 20 invocations per minute, or per second, or per hour it's not like every customer that is using your API is going to experience that.

So, you may want to work there really hard and to try to optimize it. And there are many ways to optimize for cold start times if your application is latency sensitive and customer facing. So, if you want to optimize the user experience. And there are many ways to do that, depending on the language, and the runtime, and the framework. Sometimes there is no way to avoid it. I can tell you the first time I used Lambda in my life it was a 2016, early 2016. And I was deploying a service based on the SciPy and NumPy. These are pretty heavy Python libraries, and the coaster was about two seconds. So, now it's a lot better already. But hey, there was no way to get rid of it. I really needed that library. There are many other cases where even just working on the libraries that you are using can have an impact on that. So, that's a long discussion.

Jeremy: Right. Yeah, no, but I think that you can boil it down to, like you said, if your application is latency sensitive, and I think the latency sensitive ones might be applications that are responding to API events or something like that, that are the synchronous things. I find that a vast majority of the functions that I build now are like asynchronous processing functions. And so, the cold start time, if that takes 300 milliseconds or something like that to start up, that that has a very, very minimal impact. But what are some of those optimizations that you can do just to maybe the code base itself to try to minimize... Or what are those optimizations you can do to minimize those cold starts?

Alex: Yeah, sure. So, one of the best suggestions I can give you. It's not like you do it overnight. But basically, it's about avoiding those monolithic functions that are 10, or 20 megabytes of code and library that they do 99% of what your application is doing. If you end up with a very large deployment package what happens under the hood is that we have to move a lot of bytes over the network. And that means you need more time to spin up your runtime, and then to execute your code. So, that's the simplest, most common answer I can give you.

There are many other examples of very fine-tuned optimizations like if you're using the AWS SDK, for JavaScript, for example, but only the Dynamo client, you can go and just import that single and initialize only that single client not the whole SDK. And that might give you, I don't know, 100 milliseconds back, which is impressive if you think about it. But there are many other techniques. Usually, you will have to either take some tricks in your code like lazy initialization of libraries or moving things around.

For example, I like to have many functions per file in my code. So I can have a full picture of maybe a domain of my application. And very often you end up with a lot of initialization code up there, which you don't always need in every single function. So you're just maximizing the cold start of all this function in that file. So there are ways to optimize it. Either you use one function per file, or you do some lazy initialization for the most heavy objects or libraries that you need to use.

Jeremy: Right. So, I want to get into those in detail. But before we get into that I want to kind of talk about some other optimizations. And maybe they're not optimizations, but they're sort of services and service improvements that were made that help with cold start. So one, we know, especially where you saw a lot of cold start problems or I think a high latency for cold starts was when it came to VPC. That has been, I mean, dramatically reduced, which is great. But then the other thing that came up was provisioned concurrency. So, how does that fit in to optimizing cold starts? And also how does it have an effect on the scaling impact of your Lambda functions?

Alex: Absolutely. So yeah, the VPC impact now is basically zero compared to a year ago. And that's been like that for the last 9 or 10 months since the last reinvent. Since November last year, more or less. Provisional concurrency, that's also been there for, wow, almost nine, 10 months as well. And that's basically a managed solution to the cost of our problem. What developers were doing before is that they were maybe using some cloud watch events to keep their functions warm. But let's say that, that only works at a very small scale. You cannot pre-provision 200 Lambda runtimes if you know that you're going to have a peak in half an hour, and you want to add 200 warm environments.

So, that's exactly what provision concurrency allows you to do. So, typical use case is, I know that, I don't know, my streaming platform is going to handle a live soccer match or whatever live events, and you know you're going to have a peak on your website, which results into a peak on your backend. So, there's a lot of things you can do there. You can do caching. You can try to reduce the load. But at some point, they will hit your backend. And you can try to estimate that, and you can try to ramp up enough warm runtimes so that those customers, or those users, or those consumers are not going to experience a cold start.

The way you do it is you put a number in your Lambda console, or in your Cloudformation template, or Terraform template or Serverless framework template. It takes some time. So, if you need 10,000, it might take a few minutes. I think the ramp-up time is still around 500 runtimes per minute. So, for most use cases, if you know you're going to have a peak well in advance at least a few minutes or hours or days in advance you can plan for it, and ramp up to that provision concurrency. And then you can keep it up there as long as you need it. You can also auto-scale it. So, it integrates with AWS auto scaling. So you don't really have to over-provision. One of the things that some developers told me is that, "Wait, do I get throttled above that concurrency?" No, you never get throttled. You are just going to pay for the regular Lambda implications above that provision threshold.

Jeremy: Yeah. And there's certainly a balance of cost and performance there. Because there is a cost to it. Even if you're just running one function warm... I think one concurrent location assuming you have a function for the whole month it's like $14 or something like that in USD. But for the ramp-up stuff. So, I do think if you're just worried about a cold start every now and again that is something that I would be... I would not worry about too much. But if you do have to really ramp up thousands and thousands of warm instances so that you can handle that peak, then that's definitely something to look at. But anyways, all right.

So, let's go back to just this idea of sort of general tuning of the Lambda function. So, another thing that I think I hear is a complaint sometimes is that Lambda doesn't have a lot of knobs, right? There's not a lot of things you can turn on there. I mean, you have very few things. Really, the main one is just the amount of memory that you allocate to a Lambda function. So what is the relationship though? Because I think people don't always understand this. What's the relationship between the memory setting and CPU and network throughput?

Alex: That's a good question. And that's the, I would say, the most counterintuitive side of serverless. And I keep meeting people that as developers or as operation people will just go and say, "Hey, we have a lot of Lambda functions that we are maybe over-provisioning for memory. Because we have this one gigabyte default configuration, maybe." And then they look at the logs, and the function is only using maybe 50 megabytes of RAM, and the like. We don't need all that RAM. But in reality, what you are provisioning is power. That's why I always think of it as power. Don't think of it as memory. If you allocate more memory, you obviously get more CPU, more IO throughput, you get more power. You get a chance to run faster in simple words.

So, you can go from 128 megabytes up to three gigabytes today, which is usually pretty good range. And one thing that many people do not know is that you actually get 64 different values you can choose from. It's not just 128, 256, 512. You can actually go all the way every 64 megabytes. So they give you a lot of choice. And even going one gig by default there are two potential problems. One is you might be actually over-provisioning if you don't need either speed or you might be saving cost if your function is asynchronous, and you don't care about speed. The second problem is that you might actually go even faster if you allocate more power, depending on what type of workload you are implementing.

Jeremy: Right. And so the biggest problem, I think, with setting that knob, setting that value is obviously understanding what impact that has because like you said, "If you turn the power up, it could actually run faster. That execution time could be faster, and it would actually be cheaper. Or it would just get it done faster and give you more performance." And we can certainly get into cost versus performance. But measuring that is really difficult. You mentioned logs, but that again is what am I going to go in there and go up each 64 megabyte and then try it again and try it again. And then is that an accurate sample, and so forth. So you built the Lambda Power Tuning project, which is an open source project on GitHub that helps you measure this stuff and run these experiments. So, can you tell us what this project is and how it works?

Alex: Sure. So, it is open source project. You can find it on GitHub. Maybe we'll share the link later.

Jeremy: Absolutely.

Alex: And it allows you to deploy into your account a Step Function, a state machine that will take your function as input and run it enough times to get enough data looking at the logs, looking at the implications. And then as the output, it will give you the optimal memory configuration for your specific function with that specific payload. So, the idea was that you could actually automate this process of tuning your functions eventually integrated in some kind of CICD environment. And what I really wanted to solve in 2017, I was an AWS customer. I was really just spending my day back then it was March 2017 just tuning and testing, and tuning and testing. And of course, it's a very manual process. There are 64 different values you can test. And you might actually have very different payloads in production that will have caused some performance implications in production. So, you just don't want to do it manual.

Jeremy: And the tool itself has evolved, right? I mean, you've mentioned it's open source. I know you've gotten a lot of contributions to it. And so, it has a lot of great features, right? I mean, obviously, you're able to send in the ARN of the function, right? It's pretty simple to set up. I mean, you put in the ARN of the function, which values you want to test. Which power values you want to test, the number of samples you want to run for each one because I think that's another important factor where you can't just test one sample. You got to run it 10, 15, 20, 100 times or whatever at each configuration level to understand to get a good average and see... Because especially if you're connecting to third party components like those could vary greatly.

So, being able to see that. Like you said, the payload was great, the parallelization factor is pretty cool. So you don't have to wait and run these serially and takes you hours to run it. You can do it within minutes. But you have some other things in here that I'd love to talk about because I think these are more advanced ways to think about it. And one of them is the strategy input that lets you choose between speed, is it speed and cost? What are the options for that strategy input?

Alex: So, typically, you want either your function to go as fast as possible, so that speed, or you want your function to be as cheap as possible, so that cost. But in reality, what I figured is that you want somewhere in between. You don't always want the cheapest or the fastest because the fastest might be the most expensive, and the cheapest might be the slowest. So, you probably want somewhere in between. Like the sweet spot between these two dimensions. And that's why you also have a third option, which is balanced. And if you choose balanced, you can actually decide what's more important for you. Is it speed? So, you can provide a weight, which is like zero, so you really care about speed. And that's the same as selecting speed as your strategy or you can give it a value of one. By default, this value is 0.5. So, you're giving equal importance to speed and cost.

So, I really recommend the balanced option. It's a formula deviced by an Italian mathematician that I know personally. We just talked about it for a whole day, and we came up with a formula that makes sense. But ultimately, what I really recommend is to run the tool, look at the numbers, and look at the visualization, actually. And it's very, very easy for our brain to find a sweet spot without actually looking at the precise numbers. The formula is useful. The balanced optimization strategy is useful if you want to automate the whole thing,

Jeremy: The auto optimize function, what does that do for you?

Alex: So, that was intended for use cases where you wanted to run the state machine, and leave the function in a state where it was already optimized. So, at the end of the execution it configures the power to the optimal value for you. Initially, I thought of it as a way to automate it in your CICD, maybe. In reality, you probably want to take back that optimal value, put it into your cloud formation or Terraform template, and then redeploy it, right? So the auto optimize option, I wouldn't use it in production. It might be useful in our development environment where you really want to automate it and basically auto optimize at every deployment. But really, it was more of an experiment back then. If you find a very useful way to use it, let me know. I think there are more advanced CICD strategies there.

Jeremy: Right. So, what about the output, though? So, you mentioned the visualization, but what do we actually get out of the state machine when it's done running?

Alex: Right. So, by default, you get a few parameters. You get what's the optimal power value. What's the average cost and the average execution time for that optimal power value. So, that's what you need to... First of all, it's in a JSON format. So, if you really want to automate it, you can just parse it and do something else with it. And that's really just about the optimal configuration. But then I also decided to provide some metadata. So, there is a constant, I wouldn't call it problem, maybe challenge, with AWS services, which is, as a developer, I say that I invoke a service, and I don't know how much I pay for it.

Alex: There are a few services like Amazon Polly that would tell you, "Hey, I've computed 20 characters for this text to speech operation." And so, you have a sense of how much you might have spent for that single API call. So what I decided to do is to end that in the output, the cost of that state machine execution, force that function, and for Lambda. So, you also find that in case you want to keep track of it or log it somewhere or decide it's too much. Usually, it's not too much. Usually, it's less than one cent, one 10th of a cent of a $1, a lot of zeros. So, I consider it free for most use cases. And the last thing you find in the output is a URL. This is a URL that you don't really have to use. But if you decide to open and visit that URL, you will find the same data for all the power values that you have tested, visualized in a chart that is interactive, and you can play with.

Jeremy: Yeah, and I think that, that is the thing that... Again, if you look at a bunch of numbers on a spreadsheet or in a JSON structure or document it's very hard to say, "Yeah, that's better than the other one," and kind of do that. So, the visualization piece, and that's something I think that goes to show the whole power of open source, right? That was something... You didn't build that. That was brought by the community?

Alex: Yeah, that was the idea. My initial idea was that you would just pick and trust me, and you'd just pick the output of my state machine and trust it. What actually I realized is that you want to find the optimal sweet spot visually. It's really impossible to do it with raw numbers on a table or in any other format. So, what happened is, I got actually a university student, the same mathematician I was talking about. Hi, Matteo, if you're watching. So, he decided to help me build this. And actually the visualization is a static website that you also find on GitHub. There is no database. It's not a service or a platform. It's as simple as a client side, static website that is just visualizing some data.

So, when you click or when you visit that URL, what happens is that all the data about your power values, the cost for each power value, and the execution time for each power values is serialized into the hash part of the URL. Meaning, it's not actually sent to the server. It's all available on the client side, so that the JavaScript code in the browser can visualize it for you. And it's completely anonymous and completely GDPR compliant, and there is no data privacy concern there. But if you still do have privacy concerns, if you still want to maybe customize it and build a better one, you can actually provide a custom data visualization URL that you built yourself, and you can provide it at deploy time. So feel free to build your own.

Jeremy: So, just with this topic of open source. So, how did you find working with open source? Because I know I do a lot of open source projects, including one that keeps Lambda functions warm, but I don't really use it that much anymore. But did you learn anything because I just I'm always fascinated by people working in open source because it is such a... It's sometimes thankless. It's also sometimes really rewarding.

Alex: I think it is. I totally agree with you. If you're not used to it, you could find it frustrating. I'm a developer. I like to sit down and code and develop new features. And sometimes you just have to open an issue there and wait for people to provide their opinion, which is what open source is all about. You want to get opinions, ideas, improvements. So, I actually learned a lot via that process. And now I don't just sit down and implement something I really want to implement. I sit down, maybe I'll be surprised in two weeks, there will be an open source contribution that will implement it. And so, in the process, you meet people, you build relationships, you build trust, and you might learn something you wouldn't learn otherwise.

So, I'm really happy about the learnings I've done myself. And some of the best features like the visualization and others have come from other people's idea. So, I could never have done that by myself. Last but not least, sometimes you figure that people are using it in very creative ways. I've met a person, a developer that was using Lambda Power Tuning as a kind of stress testing tool. They would use a very large number of execution just to see what happens to the downstream services. Like are we ready for production? So you do two tasks in one basically.

Jeremy: Interesting. Yeah, so speaking about learning, right? Because this is another thing. I think people are going to start asking this question. And we haven't gotten into too much detail about tweaking knobs and that kind of stuff yet. But this is a lot of work, right? I mean, even if you... Let's say you have 10 functions, and you have to run this tool on 10 functions. That's probably not that big of a deal. You could handle that. But what if you have 1000 functions, and you have to run this and get the optimization.

And so, I think one thing that I've noticed from this tool is you sort of quickly get to understand the type of workloads that benefit from different types of optimizations. What benefits for more power versus less power? And I know you have some examples, and I'd love it, I think you have a GitHub repository that you can actually run these examples and see the different ways to do it. And I'll get that into the show notes. But I'd like to talk about this a little bit. So what type of performance benefit do you see when you're at 120 megabytes versus three gigs if you're just running a standard no op type service? So, you're just doing something quick. Nothing CPU intensive, nothing network intensive, what type of performance difference are you going to see when you turn that power all the way up?

Alex: Well, so for the no op kind of function meaning no API goals, no long running compute task or anything, what typically happens is that you cannot make this kind of function cheaper because you're probably running in five or 20, or 30 milliseconds, right? And that's as cheap as you can get. So, if you're given more power, what happens is that the cost is going to go up. But what might happen, and you may not expect it is that the execution time might actually go down. So, if it's running in five milliseconds, okay, it might not be worth it because what better you can do? Maybe two milliseconds.

There are cases though where you can run for 20 or 30. And maybe you are part of a microservices orchestration tool whatsoever. And if you can shave 10 or 15 milliseconds off each step of your workflow it might be actually noticeable for the end customer. So, even for these kinds of functions, if you can shave off a few milliseconds. Usually, if you go from 128 to 256, that's usually the most you can do in my experience. There might be edge cases. You might see 5, 10, 15, millisecond execution time gain. So, that's sometimes very interesting. Sometimes it's not at all, and you do not want to pay 10% more to gain five milliseconds, I get it. I think that the important point is, "Hey, you have the data. You can take an informed decision about it." I think that's the most important achievement of the tool.

Jeremy: All right. So, what about tasks that are heavy CPU like the NumPy and those sort of things that need to use a lot of the CPU? What do you see memory settings? How does that affect that?

Alex: Well, CPU, functions are in my opinion the most interesting for memory and power optimization because the CPU effect on cost and performance is the most noticeable. So you might have something like a function that runs in 10 seconds, or 20 seconds, or 30 seconds, even more. And it's very likely, unless you're just invoking APIs and waiting. If you're just crunching data or transcoding videos or doing something CPU intensive, it's quite likely that with more power maybe you go from 10 seconds to five seconds. And so, in a way, it's all proportional. So, if you double the power, and you have the execution time the cost is flat. The cost is constant, basically, right? So, why shouldn't you give in more power? You have the same cost, twice better performance just by tuning one parameter. And I think this is the kind of workload that can benefit the most.

I've seen use cases where you run for 20 or 30 milliseconds with maybe one and a half gigabytes, or even two or even three gigabytes you can run in less than a second. So, we are talking 20, 25X performance improvement, right? And usually, the cost is flat meaning that you're either paying the same. You might be paying 10, 15% more. In some cases, you're paying 10, 15% less. And that's like a no-brainer situation, right? You can be 10 times faster and pay 15% less. There's no reason why you shouldn't do it.

Jeremy: Right. Another use case are ones that are network intensive. Because I've seen this quite a bit. I think I wrote an article about this was that when you're just making a network call, especially if you have to wait because you're waiting on the network, that very, very low memory settings really do not affect how fast any of that stuff works. Is that what you've seen in your experience too?

Alex: Yeah, that's the case, especially if you're talking about third party API external to AWS, so that your power cannot in any way affect their performance, right? There are cases where you are transmitting a lot of data. So you might see some kind of improvement with slightly more power. But usually, the cheapest option would be 128 megabytes. I've seen cases where you want to go 256, or 512 just to gain maybe 20% performance. But you need to be ready to pay for something more. Usually, you cannot really make these kind of functions cheaper, which is to be expected. It's their performance not your performance. You're calling someone else. And unless you can control that other API and optimize it too, you can't do much about it.

It's quite different though, if we're talking about internal APIs because you're not leaving a data center. You might not even... There are cases where you invoke a third party API, and you are not leaving the data center anyway because you're invoking, I don't know, a third party API that is also hosted on AWS, but not always, right? If you're invoking DynamoDB though or if you're invoking some other internal API it's pretty common I see that you invoke Dynamo a couple times during a Lambda execution. And you can actually get a lot better performance, usually, for the same cost. Maybe I've seen an example where you were running three DynamoDB queries in series, so there is no way you can paralyze it. You have to run them in series.

So, I think that total was about 250 milliseconds, something like that. And so if you double the power a couple times, you start seeing that the execution time is going down one level, and then down another level to about 90 milliseconds. So that means you can basically get the same cost for two or three X the performance. So if that's what you're doing, invoking internal AWS APIs, it's also worth power tuning. I encourage you to have a look, at least. You might not change everything, but it's common that you can get a considerable performance gain for same cost or very, slightly more cost.

Jeremy: Right. Yeah, no, and I think that part of it too is and if we go back to the sort of the having to run this on every function, and you maybe make a change to a function, you do something like that, it just becomes very, very tedious, and probably a lot of work to run this on every single function. But as you start to see... You run it on a few functions, maybe different types of workloads, those patterns start to emerge, right?

Alex: Absolutely. There is not an infinite set of patterns.

This episode is sponsored by New Relic and Amazon Web Services.

View Details

About Austen CollinsAusten Collins is the founder and CEO of Serverless, Inc. Austin is an entrepreneur and software engineer located in Oakland, CA. His specific focus is on building cheap, scalable Node.js applications while minimizing DevOps requirements as much as possible. An enthusiastic AWS Lambda user from day one, Austen founded the Serverless Framework (formerly JAWS), an open source project and module ecosystem to help everyone build applications exclusively on Lambda, without the hassle and costs required by servers.

  • Serverless, Inc.: https://www.serverless.com/
  • Twitter: https://twitter.com/austencollins

Transcript

Jeremy: Yeah. And I mean, I think that was one of the things that I thought was amazing about how you took the approach to the community, because it was very much so a community driven project. And I think that helped to have AWS who is also very much so like what do they call it? The leisure paths or whatever it is when you basically create paths in a college campus by letting people walk and then you pave over the dirt or whatever it is. I can't remember the name of it, but anyways, that idea of letting people sort of decide where it's going to go and how it's going to develop. And I always remember there being a delay, because there had to be between the framework supporting something new that AWS just came out with.

And whether that be and again, I'll use a bad example, but like maybe it's a new event that it supports. And I remember every time that came out that there was always it seemed like there was a lot of debate, a lot of pull requests and conversations about how to abstract that piece of it and fit it into the existing framework. Right. And so you developed this entire plugins community around the framework too, which was really great because that was the other thing. And that has been one of the things that I've had challenges when I work with SAM is that you don't have the ability to build plugins. So you essentially have to write a separate Lambda function that just does a module or whatever they call it there that allows you to do something, a custom resource.

But with the framework, you just had that ability to write that, to do things that you to do, to extend it, to fill gaps that you may have had during a certain amount of time until the framework supported it. But I always appreciated that more so after the fact, sometimes I was a little frustrated that there wasn't support right away. But after the fact to say, I liked the fact that you took the time to think it through because abstractions are the hardest thing to build and when you do it wrong and then you're married to it forever you know what I mean? It's really hard to change that. So that was one of the things I always liked was that it was the framework doesn't support this yet, but build a plugin and it can, and then you can always deprecate the plugin and then accept whatever the valid way is or I guess the established way that the framework does it.

So I just thought that was a really good approach. And I think really helped with all these people contributing because then also you saw plugins that came up and you're like, we should support that in the framework.

Austen: Yeah, absolutely. I'm glad that you brought that up. That was something that I think on one end the hallmark of like a great developer tool often is how extensible it is. Right. But a lot of this just came out of the fact that I started as one person and I couldn't take on this massive project, just being one person in the early days. And those people who've had, like things go viral on Hacker News. I think they have probably been through a similar experience where one day this project is just blowing up, but the next day, the hangover sets in, when you look at the issues in GitHub. And people say this project does not support my use case, this does not support my workflow.

This does not support my organization's policies. And then on top of that, the ambitious goal of trying to abstract almost essentially it's just so many AWS services. It was like, okay, how is this ever going to work? And that's where plugins came from largely, and you're right, you honed right into it. People can actually just overwrite any part of the framework, extend it, and it's up to them to figure out what the right patterns are. And once we see those emerge, then we'll just merge that into the framework. So it was a lot of outsourcing product development to some extent, but all goes back to just the fact that we didn't have the resources to do all that in the early days and it's great. That's where creativity comes from a lot of the time. It's just limitations.

Jeremy: Actually. It's funny you mentioned limitations because that's the other thing... This has been an ongoing theme with serverless too, is that you do have to work around some of these limitations, but sometimes that's not a bad thing, right? Sometimes when you have constraints, you can build something better because you can't just go and do something crazy, like say, well, we're going to set up a K8 cluster because we want to do some type of processing over there. If you ask yourself, how can I do it within these constraints? Not only is it usually faster, it's a lot cheaper. And oftentimes I think you'd get a better product out of it.

Austen: Yeah, absolutely. I totally agree with that. Constraints are super important. And also, I got to say the amount of stuff, use cases, that people were trying to use serverless for, especially in the early days when it really wasn't meant for it, was just so amazing to me. People would try and duct tape together crazy serverless architectures to accommodate a use case it clearly wasn't ready for. There's so much creativity in that area. And it seemed pretty crazy, but it's just people love those serverless qualities. Auto-scale, only pay for use and massive scalability. People want that so bad. The fact that so many people were doing this and they still do it a lot today because there's still not even great serverless database options, right.

The stuff that people are doing to get around that, the DynamoDB movement right now, the popularity and all that I think is a lot of it is due to the pent up demand that is just waiting for the cloud to become more and more serverless and offer more and more serverless options for all use cases. And that's really I saw that early on and I'm just thinking like, everybody wants this, this is going to be table stakes for the cloud. Again, this is not a fad. This is like the next... This is cloud 2.0, that's cliche to say, I hate to saying that, but this is like the evolution of where-

Jeremy: Cloud 3.0 now I think we're already getting towards 3.0. So 2016, 2017, 2018 you're working on building a community which you did an amazing job of. And I know you hired a few people. I know Alex DeBrie, for example, a ton of growth hacking, which is just brilliant by the way. So if you just take your serverless hat off for a second and go back and look at how you built this company and how you built this community, it is like a masterclass in how that worked. Now, I know I'm sure you made a lot of mistakes along the way, like you said, for every one success, there's 100 failures behind you. But again, it was the persistence. It was the approach to it. It was hiring a bunch of great people. And like you said, having some really, really great people working for you and working on the project that I think got it to the success. But now we get into 2018, maybe towards the end of 2018, 2019, you've got to turn this into a company at some point, like a real company.

Austen: Yeah, absolutely. And we had built out such a kind of groundswell of a user base. It was time to start kind of looking at how to do that. And we had tried a few kind of minor experiments on like, okay, what does commercialization look like for us? We even went into infrastructure as a service at one point. We had a project called The Event Gateway which I still am a huge fan of that project. I think that the serverless era, era of serverless compute, a venture in compute, were again, it's never been easier to write code that reacts to events. I think that there's room for a better type of event gateway or event bus that really takes a lot of API gateway concepts and kind of merges it into a new type of infrastructure there. And that's what the event gateway was seeking to solve. But we were still-

Jeremy: But it was hard to market an event gateway that ran on servers from a company called Serverless.

Austen: Yeah. A lot of interesting conversations were had about all of that certainly. And it ultimately just kind of became a distraction from like building the framework and stuff. We were just doing like too much stuff at the time and organizations efforts thrive or fail depending on where you set the focus.

And so we decided to continue to focus on the framework and just build out the other later stage life cycle application life cycle management features that you need. So the challenge that kicked off the framework in the early days was this distributed system challenge, right, where there's just all these pieces you have to compose together, and that's not just that only exist on the cloud too, so distributed system where you don't own most of the parts.

That's not just a development challenge, right? That's a monitoring challenge. That's a kind of a developer, a team workflow challenge. That's a secrets management challenge. It really presents a lot of new problems for every single phase of building and managing your app.

So that's where we came out with Serverless Framework Pro and our goal is just, hey, if you're using Serverless Framework we know you want to focus on product. We know you want to focus on the outcome. Like we're going to try and set up monitoring for you. We're going to try and set up secrets management for you. We're going to try and set up CICD for you. And we're going to use the knowledge of the application that you've created with our framework to try and do all that for you without any configuration, without you having to figure that stuff out.

Austen: And so that when you deploy with the framework, you're ready to go into production from day one. And that teams don't have to go figure this stuff out. All teams should have to do is just say, use Serverless Framework. And that in itself will put all the guard rails, all the observability you need in place. So I think, again, you don't have to think about that stuff because that's what serverless is all about, right. Not having to mess around with all this, and then of course maintain it, which is the harder part, easier to build some stuff, some DIY thing, but then maintenance is always the hardest part.

Jeremy: Right. Well, and just stitching everything together too, and understanding how X flows to Y through Z and all this kind of stuff, it just gets really, really complicated. So yeah, the Serverless Framework Pro is great. I've been playing around with it quite a bit. Again, the CICD stuff is awesome. And if anyone has ever set up CICD for serverless and I have a sort of a running joke in my newsletter every week where every time I see someone's set up for CICD and serverless, I always make a joke, here's like the 10000th way to do CICD in serverless.

And again, model repos versus multi repos and all that kind of stuff. And the framework supports both those, as well as I like the model repo approach where it can grab just one modified serverless.yml file and deploy that section of the framework and some of those things. The monitoring is great. The alerts are great.

Jeremy: So definitely if you are using the Serverless Framework, I am a big fan of what you've done from the commercialization piece. And I think again being a little bit selfish, I want to see the framework keep going. And you can't run a company if you don't have revenue, right. You can't just always be building community. So I think it's great that people are using that and they're benefiting from that and they want those extra features. It's a great way to not only take care of the things you really need, but also to continue to support your company and the framework that I think is what really made serverless very accessible to most people.

Austen: Yeah. And that's been very successful for us. It's been great. We're shipping features every single day to make that easier. We've got a lot of next generation tools that have been in the works for a while, that we've been releasing lately. A lot of cool stuff that just extends on that theme of don't focus on the infrastructure, of course, focus on your product. And there's a lot of interesting stuff happening in the cloud right now. I mean, it's just such a fascinating space. I think personal predictions might be way off here, but this is where it seems to be going. And it's kind of been my personal conviction since 2015, since Lambda came out. It does seem like serverless qualities are just going to be expected for all cloud infrastructure.

Right. And when you log into AWS, Google Cloud, Azure, you're going to just you see all those services have more and more serverless qualities. And so it almost feels like the serverless is just merging with the cloud and that'll just be cloud... I don't even know if there will be a serverless term anymore. It'll just be what we think of-

Jeremy: Exactly.

Austen: What the young'uns think of the cloud in the future, right. They'll be like, "What is this serverless thing?" And so that's really interesting to watch because that's important for us because we're trying to build the tools for that. And if all of cloud is going to have these qualities and everything's going to become serverless, it's an interesting opportunity. The other stuff that I see is really interesting is I think how code is thinking a second seat in this architecture, they're kind of like, it's not the primary thing.

And I think it was Paul Johnson who wrote this at one of his many awesome serverless articles. And he wrote this and I totally agree. I see this all the time where the people working on these serverless architectures are doing more configuration these days than just code. And then you've got people coming out in the serverless community saying of course code is a liability, all that stuff. And that's really interesting because so much of our development tools are designed around code first and code is the most important thing. Meanwhile, here's this new architecture saying like, well, you put in code where you need it. Right. But then you want to try and lean on stuff that you don't have to maintain over time and rework.

So I think this fundamentally disrupts the workflow and the tools at the end of the day. And that's been a very interesting space for our company to think about. And we've got a lot of innovation that we're just starting to ship out and a lot more kind of later this year on kind of those themes. And then of course, with serverless just becoming the cloud and everything kind of being serverless at some point you get a lot of complexity that comes out of that.

Jeremy: I was just going to jump in there and talk about complexity, because that is one of the things where, right from the beginning I remember it being so easy. And some of those Hacker News comments, by the way, very early on like, "Oh, we already have this, they're called cgi-bins." Which of course it's not a good comparison, but you know what was great about cgi-bins? You upload a little bit of code and it was available on the web and it ran and it was amazing. And that was how serverless started very easily was sort of like, you just upload a snippet of code. You don't have to worry about setting up servers, all this other stuff, it runs something for you, does something and it's done, and that's it. You don't have to think about it. You have to set up the servers, you don't have to any of that stuff.

Then we added API gateway and then we added RDS or VPC, so you could do RDS and ElastiCache. And then you started having all these issues with connection management, because the connection management or the connection model between those were broken. And so you had to start playing games and then cold starts got worse. And then they're like, "Hey, let's support Java." And you're like, "Okay, well now the cold starts are really bad." And so you had all of these things that just started adding more complexity, SQS Queues, SNS, EventBridge. You start adding all this stuff in and you get to the point where you look at... And I remember there was an article that was like your typical serverless infrastructure. And it was like 40 different services all tied together. And people were like, "This doesn't seem simple, blah, blah." It's like, "Well, it's not simple, right?" It has become very, very complex. And you've got all these patterns that have come out of this in terms of how you handle different use cases.

And so I know you agree with me on this. This complexity of serverless has taken some of the shine off of it. And we need to find better ways that we can start focusing less on the actual code and the infrastructure and more on what we're trying to achieve or the outcome of it.

Austen: Completely agree. And that's where the Serverless Framework, that's where we've always been thinking, right. Going back to that simple story of functions and events. When this thing happens run this logic, right. It should be that simple and it's certainly gotten pretty complex, and this is kind of the side effect of the cloud becoming more serverless. And it seems like there's just a lot of stuff too, that we just didn't see coming at the end of the day. Layers like, provision concurrency, like database proxy, like RDS proxy or something.

Jeremy: RDS proxy, yeah.

Austen: It's like, well, I understand servicefull, but there's a lot more services than ever anticipated here. Right. Meanwhile, the biggest companies still are having trouble securing their S3 buckets.

Right. So yeah, what does it all mean? How do we create simplicity here? And again, going back to just like, how do we bring order to this and make it accessible? Because on the other hand, this is highly efficient, very powerful, next generation cloud infrastructure. And myself, our team, we believe that these are the greatest building blocks of all time. Right. Just to have all these things shelves filled with every single type of like API to do almost anything. Right. I think it could usher in a new golden era of software development. I mean, it feels like we've already been in one for a while, but this stuff is just so darn powerful. But again, how do we make that simple for people? So for us, it's really just continuing our story of abstraction.

Right. And again, we started out with a simple story of functions and events. Don't think about the infrastructure, just think about kind of your outcome first. Now, we've got more. Now, there are more serverless infrastructure possibilities, configuration options, all that stuff. But it's very clear that specific use cases are more popular than others and specific use cases for the most part, somethings just should be done serverless almost all the time. Depends on your organization. Depends on some people still need a greater degree of control over their service environment and the infrastructure, all that, but there's just very clear use cases where that this is a great fit and you have to have a pretty good argument for not just spinning up a serverless API right now, because it's just so darn efficient.

So we see those and we know what the best practices are because people we've been building these for a while. And to some extent, we're kind of asking ourselves, well, just like you shouldn't be managing servers, should you really try and figure out how to build a rest API on service infrastructure from scratch or where it's just like, they're ready to go, rest API opinion that has the best practices built in, the best combination of infrastructure from a scale performance and cost perspective. So developers don't have to think about that. They just think about, I need to deploy an API. I need to deploy like an express app, for example. And so that abstraction story is what we're continuing to focus on with our new effort, which is Serverless Framework components, which is our by far fastest growing effort right now where we've basically just honed in on those most popular use cases.

And we're building out specific developer experiences for each one of those use cases. And our goal here we want to do like a 10X better developer experience. So each one of these, like we're trying to get them to deploy in three seconds or less, right? Because we think that when you're working you're on these cloud services you should be developing on those cloud services. You shouldn't have to emulate this stuff locally.

And a lot of people just deployment is too darn slow, right? No one wants to wait 30 seconds to a couple of minutes for cloud formation to do a deployment in order to see their single line of code change in the cloud, that's crazy. So most people go back to emulation and then unfortunately that's easy to start, but then a lot of teams, your team grows, the project grows, the types of services grows. And I've seen so many companies trying to maintain some crazy kind of emulated version of AWS that they do all their development on.

Jeremy: Which is impossible.

Austen: Yeah. And I'm still personally trying to keep score. I'm like, are we being more productive with this architecture or not? What are the issues here? If people are trying to maintain a local version of AWS, then it feels like we're losing a lot of that creativity and productivity that was supposed to be freed up by outsourcing the infrastructure to this cloud provider in the first place.

So anyway, we've looked at these popular use cases and we're trying to create one Serverless Framework component for each use case with other developer experience features like Fast Deployment. So Serverless Express is a good example. If you're a developer, you just want to deploy Express in a way that's auto-scaling, in a way that can scale massively out of the box and charges you 0.00003 cents per request.

You just want to take your Express and run it like that. Then the Express component's perfect for you, just put in your Express app and package it on serverless infrastructure for you. Three-second deployments and there's other cool features. It streams your logs and errors directly into your console. You've got like a dev mode where it just watches your code every time you hit save. It does a pass deployment, streams, log statements, errors, and a lot of what we're doing to be frank with all this new development experience stuff is trying to recreate the experience that developers had before serverless cloud infrastructure. Right? So we're just trying to make it as if you're running this Express app on your machine again, which you could still do, but we want to make working on the cloud just as fast as if you're working locally, because we think developers for serverless to use all these new infrastructure features you got just work on those.

Be able to develop on those, not do some fake thing that you then push into production and encounter all these limits and issues because you weren't working in kind of the reality of the infrastructure at the end of the day. So components is a huge effort here. We've got one for websites who want to deploy a serverless website, highly efficient, Express, scheduled Lambdas. We've got one in the works for Next.js. Also if you want to deploy React applications, we've got event-driven ones in the works, different web hook handlers, and then components for running on other vendors as well.

Jeremy: Right. Yeah. And I tell you, what I love about the idea of components is again, it's almost as simple as saying like, "Okay, I just need this to happen when something else happens." And there may be multiple connecting pieces, but here's my piece of business logic that I need to write. And there are a lot of companies now that are doing these sort of low code workflow systems, right. And I think Paragon, Pipedream, like there's a whole bunch of them. And they're very cool because they say, look, you want to send something to Twitter or you want to read something from a Google Sheet or you want to query a database or do something. Those are all pre-written and done for you, but then it's like, well, but I need to randomize it and add some prefix to it or something like that.

Well, that's a little piece of code that I have to write. And I like that level of abstraction, but I don't necessarily like how it's owned in someone else's infrastructure in the sense where I don't quite have full control over it.

So what I like about the components is that it gives you the ability to do that and own it, but still have a lot of that extra instrumentation sort of have done for you. And the thing that I think is overlooked maybe from this idea of components and any of these sorts of things is it's about repeatability, not just for something small, like, "Oh, I need to just publish a new website, or I need to publish a new API." It's something like we are an enterprise and we have all of these security requirements.

We have all of these bootstrap stuff that needs to be done. All of this instrumentation that has to be written in order to do observability and monitoring and making sure we're following PCI and all these other things, right. You've got this whole long list of things that has to be done every single time. You give somebody a blank template and just do serverless init, right? I mean, that's not going to give you enough in order for you... I mean, do you have to go through quite a few iterations before you get something that is going to pass muster with the lawyers and compliance and all that kind of stuff. So being able to encapsulate that, right, which is why I've rethought the CDK a little bit, because I like that idea of writing these reusable components that you can just plug in together and use these constructs.

And the serverless components feels very much the same way to me. And I like that because I think when you start talking about new organizations needing to build out complex applications that are compliant, that have all the security that they need in there, that have passed all the security checks and all of the sign off from everybody involved. That this is the way, this is the level of abstraction that will allow you to do that, be really productive. And again, go serverless, which I can't imagine anybody now sitting down thinking about building a new application and saying we shouldn't at least look at the serverless piece first.

Austen: Yeah, yeah. Absolutely. And then there's more because what we're doing with components is we've been in the land of infrastructure as code, infrastructure provisioning for a long time. And every single problem that you raised absolutely is a real problem that needs to be addressed, that can be addressed by really powerful templates that you can run on your own infrastructure that your team can use as opinionated pattern. So they don't have to go figure out how to put all this stuff together because everybody's going to do that wrong the first time. I don't think I've seen a team that's put together-

Jeremy: And the second time and the third time and the fourth time, and then eventually they might figure that right.

Austen: And then you get it right and then re:Invent comes around and says, now there's a new way to do this. Now, you've got to do this. Right. In my opinion you don't want your team doing a lot of that stuff. Right. Doing all the low level stuff.

With abstraction, as we all know, it's this battle between convenience and control, right. And not every abstraction is going to work for every company. There's a lot of different types of users out there and use cases, it takes all kinds to make a world. But components were coming out with a ton of different flavors at different levels of abstraction. And we want to go after those darn serverless architecture diagrams, right. That we see going around on Twitter which they're just showing all of that stuff.

And all we're trying to do is say, yeah, we're going to take all 20 of those things and we're just going to put them into one piece. Right. And then you'll have the other 20% will still be some low level pieces. But for the most part, you don't need to go and use every single low level thing and configure it yourself anymore, unless you just love doing it. And then there's nothing wrong with it. Right.

But for me I personally have a ton of APIs I deploy. Right. There's a ton of them. And I don't want to think about low level API gateway stuff anymore. I mean, I've got hundreds of end points now for doing all types of things. And it's just crazy to want to think about that low level stuff.

But again, going back to kind of, there's more here. We've been thinking a lot about infrastructure as code for a while. And in our opinion, there's so much room to innovate here. And it dawned on me when I heard, I think it was Adam from Chef. He was giving, I think it was a presentation on Habitat. And he had this quote here where he said, you should try and figure out how to package more automation with your application. More use case and automation should just kind of come out of the box. And it's really inspired me personally to start thinking about like, okay, how do we package more automation with these components? Right. Because infrastructure provisioning, as we think about it, it's just kind of deploy, remove, roll back, very simple.

But there should be more automation capabilities kind of built in. And the one thing that's interesting about components is that we try and focus on the outcome of the use case first, which is so enlightening when you're building a developer tool. Because you know the goal of the user and we know what the goal of the user is, you could automate way more for them. So an example of this is when you deploy a component, it's going to have metrics instrumented for it in our dashboard now. Like if you deploy Serverless Express and those metrics are going to be the metrics that matter most for that use case. It's not just going to be low level, here's in AWS Lambda function in vocation, right. It's going to be first, we're showing you your API requests, because if you're building an Express app, you want to see your API performance overall, because that's what the customer is facing. Right?

So each one of these are starting to shift with their own custom metrics that are focused on their use cases, ranging from infrastructure and then soon product, right? Like what are the product metrics that really matter for this? These things are starting, there's a few components that have built in testing functionality. Why should you have to rethink how... You learn how to test an API at the end of the day. Components should be able to generate documentation for you, all this stuff. How do we really think less about infrastructure provisioning and more about application automation and pack and use this knowledge of the use case, the end goal to really automate so much more of that for the developer.

So they don't have to think about that stuff. They just have simple action. Each component has simple actions, deploy, remove, test, monitor all that stuff. So we've just started that journey. Like components is still new. We went GA in April, but you'll see us announce a lot more on this front later this year. But this is all going back to the theme of build more, manage less, right. Where we think serverless needs to go ultimately.

Jeremy: Well, I think that idea of the use case or the outcome is really cool. And I love the idea of having the metrics built in. Indulge me in this story for a second. I thought it was hilarious because I was at a company and we had just set up dashboards or the monitors in our office so that we could have these big dashboards up. And it was because the CEO said, hey, we want to numbers up on those. Right. So especially like when investors come in, they can see whatever. And I said, "All right, well, what do you want to put up there?" He was like, "Well, I want to put up our KPIs." Of course, key performance indicators. And I said, "All right, what are those?" He's like, "I have no idea."

So I think there are a lot of people who are going to be in that boat like all right now I'm building a serverless application. And if you abstracted away enough where I'm just entering a little bit of code, just doing my little business logic here, then the question becomes, what do I even care about? Right. Because I don't need to monitor CPU on my Lambda function anymore. Like that is just done by AWS for me. But what do I really want to know? Do I want to know the latency of my routes? Do I want to know the frequency of my routes? Do I want to know my error rates? Some of these other things that would show me something that's actually actionable, but how am I supposed to know those things unless you have a lot experience. And any of that experience that you can just bring and sort of fold in there automatically, if anything it just gives you something to build off of.

Austen: Yeah, absolutely. And it kind of goes back to the question that serverless raises that is, when you don't think about the infrastructure, what do you think about? What matters most? Well, of course the outcome, the product, the customer experience, the business problem you're trying to solve. That is the stuff you need to think about. So it's still early for us here, but we want to answer that to a greater extent built into our tools and this is all just kind of changing how we think of developer tools and what that really means.

And then ultimately like who is a developer at the end of the day too? Because you can see and I've listened to so many of your interviews already with a lot of stuff with people just in different parts of this space, you've got like low code, you've got front ends and then you've got serverless infrastructure.

There's a convergence here kind of happening. There's an interesting kind of movement here happening where it's going to end up, I don't know. My personal theory though, is that the cloud, everything is going to have to focus more on outcomes, not on infrastructure. And that's what you'll get when you go to AWS, just serverless services, API as a service focused on outcomes where we've got business solutions, instant business solutions that auto-scale and never charge you until you call their APIs. And how that's going to change the tools, how that's going to change how we define a developer. I'm not sure what that looks like yet, but we've got a lot of interesting ideas on what it means and we'll continue to try building out some cool products to put some solutions out there.

Jeremy: That's awesome. I mean, I remember sitting down at breakfast at re:Invent with Ajay Nair and he asked, how do you describe serverless? Or like something to that effect. And I said, serverless is just the way, that's just where we're going. Right. Like it's just going to be cloud 2.0 or 3.0 or whatever revision we're on. But I totally agree. I mean, this is one of those things where it does take a lot of convincing sometimes when people look at containers and they're like, "Oh, I have all kinds of control over it."

And it really goes down to this level of abstraction and we get I think distracted by things like cold starts and vendor lock-in and all these other things. And yeah, maybe not all of those things are perfect.

But if you think about something as simple of a use case as catching a form from a website to serving up an entire API, to processing millions and millions of records coming in via Kinesis or something like that. Or converting files, performing OCR on medical records. I mean, all of these use cases now are there and they're supported and they're not all perfect, but it's getting to a point where I just can't imagine my life right now without serverless. And again, I don't know, like you, I just get really excited about it because I think about the possibilities and I think about where this is going and if we get it right, if we continue to build those right levels of abstraction, and continue to convince people that again, it's the way, I think the future of this is pretty amazing.

Austen: Totally agree. It's the potential that's democratized for everybody, whether you're a large organization, or you're just a solo hacker, like in the basement or something, trying to get something off the ground like this power has democratized everybody. And that going back to our mission, like, yeah, we want to help every single person build more, manage less, leverage higher levels of abstraction, help them focus on outcomes more than ever. We're going to try and rethink developer tools and what that means in order to deliver that experience. And then the last part for us is just we firmly believe serverless is bigger than any one vendor at the end of the day. And we feel very strongly that there needs to be an application framework that provides an open level playing field for serverless cloud infrastructure across any vendors, because yes, we've talked a lot about AWS and the majority of our users are using AWS.

And the majority of the infrastructure is AWS, but not all of it actually. They are still bringing out, our users, our audience are very product focused. And if you want to build the best products, you got to be free to use the best of breed services that are out there. And so we see a lot of people still bringing in Stripe, still bringing in Algolia, still bringing in MongoDB Atlas, Twilio, right? There's so many great things out there. And helping people, developers have this, again, this open framework where it treats all these things as neutral. This level playing field where they could compose serverless infrastructure across any vendor into applications really, really easy. It feels like the destiny of the Serverless Framework to us.

Jeremy: Yeah. Well, that's pretty awesome. And what else was awesome is you being here and sharing all this information. Again, I love that history. I think that's just amazing. I think again, like I said if you're building a service now or you're trying to build a company just how you've done it, how you've gone about it, and how that has grown. And also just the advice of listen, you're going to fail hundreds of times before you succeed. It's just the way that it goes. So again, thank you for being here and sharing all this information. And if people want to find out more about components and the Serverless Framework and where all that stuff's going, or just contact you how do they do that?

Austen: Serverless.com. Pretty easy to remember, right? The infamous domain, but that's where you could find us. I'm on Twitter, just Austen Collins, A-U-S-T-E-N though. And you can include that in the show notes or something like that but-

Jeremy: I will put it all in there. You have a great blog too: serverless.com/blog. Lots of great information in there. So thanks again, Austen. Really appreciate it.

Austen: All right. Thank you Jeremy. Take care.

View Details

About Austen CollinsAusten Collins is the founder and CEO of Serverless, Inc. Austin is an entrepreneur and software engineer located in Oakland, CA. His specific focus is on building cheap, scalable Node.js applications while minimizing DevOps requirements as much as possible. An enthusiastic AWS Lambda user from day one, Austen founded the Serverless Framework (formerly JAWS), an open source project and module ecosystem to help everyone build applications exclusively on Lambda, without the hassle and costs required by servers.

  • Serverless, Inc.: https://www.serverless.com/
  • Twitter: https://twitter.com/austencollins

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Austen Collins. Hey Austen, thanks for joining me.

Austen: Thanks for having me, Jeremy.

Jeremy: So you are the CEO and founder of Serverless Inc. The creators of the Serverless Framework. So I'd love it if you could give me a little bit of your background and just in case somebody doesn't know what Serverless Inc is all about.

Austen: Yeah, sure. Quick background on us, we make the Serverless Framework which is an application framework that makes it really easy to build applications on serverless cloud infrastructure. That is infrastructure that's auto-scaling. You never have to pay when it's idle and scales pretty massively. And the goal is to help developers deliver software that has radically low overhead and all of these serverless qualities at the application level as a whole.

So that's our goal. We make Serverless Framework, that's what we kicked off with. And I was excited to chat with you because I was thinking it might be interesting not to do just such a technical conversation, which I'm sure you've done a handful of already. But maybe talk about the history of Serverless Framework a little bit, because the project is now five years old. I think this is my fifth year of serverless development, which is crazy to think about, because it feels like we're so early in this journey in general. But I was thinking you might be interesting to talk about kind of the history of like how things got started, how we got started and our perspective of just like kicking off the serverless movement and kind of, we know what that looked like in the early days and the crazy days where we didn't... I don't think anyone knew how big of a deal this was potentially going to be, or that this would have become a big category for cloud or maybe even the cloud itself.

And when you reached out to me to do this podcast, I thought this might be a great opportunity just to kind of tell that story, at least from our perspective, my perspective. Because I think it's a fascinating one, not just for technical people, but for makers and entrepreneurs, anyone who's trying to get something off the ground. I think there's just a lot of interesting lessons learned along the way.

Jeremy: Yeah. No, I think that is something that is really, really important for anybody in the serverless space. And I think anybody who's developing cloud applications today is to look back and see, I mean, where we were five years ago. Because it has dramatically changed in terms of the technology that we have available to us, the building blocks that we have available to us. And also I think, JAWS as it was originally, and we'll get into some of that, the original Serverless Framework, what that was able to do compared to what it can do now, but also compared to what's available now and just the massive explosion of development tools and observability tools and everything else that has kicked off. Open source projects beyond the framework that have kicked off that really have built this amazing community. So let's start there. Let's go way back to the beginning, like this we're talking 2014, right. Lambda is still in preview and then what happened?

Austen: Yeah, Okay. Going way back, a lot of the credit first goes out to the Lambda team, the visionaries over there who kind of basically disrupted how they do compute over at AWS. Which I've heard a lot about that story. I listened recently to your podcast with Tim Wagner, kind of talking about the early days of that, but really like a lot of what we're doing here with Serverless Framework and building out our developer tool suite is just kind of standing on the shoulders of the effort that those people did, which I'm sure was hard figuring out what that looks like inside of an organization as big as Amazon. And so our story really I'd say they did a lot of the hard stuff and a lot of the really meaningful stuff and kind of my story starts right out when they did that announcement at re:Invent 2014. Yeah. When it was in preview.

Austen: And I was looking around at everybody else and there was definitely some excitement and I was just personally so enthusiastic. There's something, it just hit a note in me that still is driving me to this day and inspiring me to this day to build great developer tools and really capitalize on what the potential I first saw when it came out and still until today. And that is like... This for me, it was, I guess, I don't know about you, but I actually never got into this to be a developer. That was not my personal goal when I was just getting started.

Again, I've always felt more like a creative type to be frank and more of an entrepreneur. And the programming, the development was kind of a means to an end, but also felt like potentially the greatest skill set to have at a time where the cloud programming gives you the ability to make anything, to solve any problem almost.

And I think a lot of my story and all the time that's gone into the Serverless Framework the same goes for the community members and whatnot. I think there's something similar where the people who are attracted to this are very much product focused. They really care about making things, they care a lot about the customer experience. And a lot of the technology is cool, but to some extent, we kind of want it to go the way so we could focus on the customer facing experience.

And that's kind of always been a strong theme for me personally. I see that in the serverless community everywhere, we've talked about it a lot and it was the thing I felt when Lambda first came out. I got so excited. It felt like for the first time there was really a technology where I could just put any logic out there and it would run for me, auto scale and charged me unless it was running.

And so that felt like amazing power. And I was looking around, there's definitely some excitement. It was so early in those days. I think AWS, maybe you remember this better than I do, but they were pitching it as like event-driven code kind of glue code. And there was no serverless category. There was no serverless buzzword or anything like that. It was kind of just a stitch together, kind of shuttle some data from one place to another largely from S3 and had very limited use cases at the time.

Jeremy: Right. If you remember all the way in 2015, even once it became generally available, there was no API gateway to connect. So you weren't able to do those web use cases, which again are probably one of the most prevalent things that serverless is doing now.

Austen: Yeah. Very limited. And I don't think all the pieces were there and enough of the pieces were there for people to get kind of the overall vision, maybe how meaningful it was, or at least how meaningful I felt it was. So I went away from that, re:Invent trying to chat up my network, my colleagues, my friends, and say, and see if they're excited about this new compute services as I was. And I was thinking like, what if we take this new compute service and pair it with other infrastructure that has the same auto scaling kind of never pay per use qualities. We could deliver software that as a whole has a really low operational cost, right. Almost like you set it and forget it architectures. And that again, kind of touched on a theme that I really like, and that's just the ability to build more and manage less.

Right. And so I was going out of my network trying to raise a lot of excitement about this and just see if people were equally excited. And there was some. Again, early days, some people were into it, but not a lot of people cared then. And so I kind of took it upon myself to see what would happen if you could actually take Lambda and build an entire applications on it. And see what that process looked like and see if it was possible. And right at that point in those like early personal experiments, I ran into the problems of kind of the serverless architecture that we're still dealing with today. And that is that this is a distributed system. And the other day you were actually working with a lot of cloud services, perhaps more cloud services than any other architecture out there.

Because if you really want to build a application like a web app, you're going to need Lambda. You're going to need to work with IAM. You're going to need to work with API gateway. You might need to work with DynamoDB, CloudFront, maybe Route 53, a lot more. There's just a lot of stuff in there. And so the value prop was like deliver software that's super-efficient, but how do you actually take all these pieces and easily put them together to form a nice application experience that developers would really enjoy? And that was the initial problem that sent me off on my journey. And I think it's the problem that we still are trying to figure out today. It's like, how do we streamline serverless development and make it as zero friction as possible?

Yeah. So I started working on the framework and it's just a personal nights and weekends project. And at the time I just moved up to the Bay Area. I was still pretty new. I'm originally from Los Angeles. And I think I was working on another startup at the time and also building out a backend inventory management system for a larger company. And it was just spec and this is a personal experiment. And I think that somehow I had published something to GitHub and Jeff Barr found it. So Jeff Barr is amazing as we all know. I don't know how that guy does it, but he's looking at everything at all times. And somehow when I was just kind of publishing some stuff to GitHub he found it and I still don't know how he does it if he has like an army of interns or something like that, but he found it and he sent me an email.

He's like, "This looks pretty cool. Would you be interested in like talking about this at an AWS event or us promoting it on the blog or something?" And I was pretty honored because I've been a big AWS fan for a long time, user for a long time. And the fact that Jeff Barr was reaching out based on this thing I was working on over the weekends was so cool. And I told him, I said, well, let me work on it a little bit more and finish it before you kind of push it out to the Jeff Barr audience. Right, because you know how massive that is. And I told him, I was like, "Okay, give me like a few weeks to really flesh this thing out."

Jeremy: So where were you though when you were starting to work on this? Right. Because API gateway came, I think it was an early 2015 or mid-2015. I think it wasn't until 2016 that there was even VPC support. I mean there was just all kinds of things that have gotten... I mean, it's really crazy to think back at how limited it was back then to where it is now. And when I first started using JAWS the original framework, I know that it had API gateway support in there. So that must have been around version 0.4 maybe or something like that. But what was available at that time when you first started building it?

Austen: Not much. Looking around there's still a lot of some definite challenges there in the architecture. But way back then, it was so much harder. It was crazy hard. And I had done a few versions of those even before APA gateway came out. And I think there was one other project, I can't remember the name of it. And it was kind of really, really strange also trying to work around the complexity of using this new infrastructure that was just so raw and early. But there were a few versions I worked on before API gateway came out. And then when API gateway came out, obviously that was the missing piece for one of the major use cases. Now that it was still the backend, APIs, microservices, all that. So worked on it a lot. And then API gateway came out, I think July, 2015, it was July?

Jeremy: Something like that, yeah, that sounds right.

Austen: And then I did a whole new version of it. And started, I think it was that time where I really kind of honed in on the essence of the framework and that was, okay, there's a lot of cloud infrastructure that you have to work with. Developers have to know about. Now unfortunately, developers don't really like working with a lot of cloud infrastructure. They like kind of building apps and getting stuff out there as fast as possible, and the least amount of distraction in their flow as they're working on something, creating something. And so my goal was to kind of hide the infrastructure complexity first and foremost. And by the time API gateway came out, I think the opinion that was honed in on was that serverless applications are a simple story of functions and events.

And it's like that is the essence of a serverless application. And that whole idea is how the framework is going to be designed. And so that's what you get in a serverless.yml file, you get that functions property, and there you could list out all your Lambda functions and just add in your business logic, your code. And then there's an events property where you could hook up anything to trigger that function.

And while it's a bit strange to define an API as like an HTTP event triggering a function it just brought order to a pretty kind of chaotic, especially back in those days type of architecture where the developer could just quickly look at it, understand that story. And with a lot of kind of abstract configuration syntax, the framework, also something that it did, that's a bit different than a lot of other application frameworks is it helps you kind of structure your code, but it also provisions the cloud infrastructure, right? So it's this kind of weird hybrid thing. And the goal was don't make developers have to know a lot about the infrastructure. Come out with a nice kind of application model that allows you to focus on logic and what triggers that logic to run.

Jeremy: Yeah. And I think it's important to note too, that that idea of event driven applications. I mean, that's not particularly new. There were other... When we were doing even microservices or we were doing what's the other thing that I'm thinking of there, service oriented architecture, things like that, that you had a lot of events flying around. But building those types of applications were ridiculously complex because you had to have a deep knowledge of this idea of some sort of event bus that was running there. You had to have the individual microservices set up as different components. You had to know all the different ways in which these could communicate. And one thing, and I don't know if the Serverless Framework deserves credit for this, but I'm going to give you credit for this, was that paradigm shift of saying, I'm a developer, I want to write code.

And I just want to think of it in of all how does my code get triggered? Right. What's the thing that triggers my code, because if you look at how ClouFormation was structured to do and provision API gateway, I mean, obviously the SAM, which is the serverless application model, AWS came out with, they essentially used what the Serverless Framework had come up with that sort of idea of functions and then having the triggers against the function.

So I think that's a really, really good way to think of it because you're right most developers... I mean, I work with a lot of developers and most of them had no idea what was happening behind the scenes. All they were like is here's my code make it run somewhere. And that was throwing it over the wall to a completely different team. And that has, for the most part, especially with full-fledged or full stack serverless applications, that's gone away.

Austen: Yeah. Yep. Absolutely. Yeah, it just seemed right at the time. It just seemed like there was just a lot of chaos and it needed some simple stories, simple way to think about it, to bring order to that. And looking back even at the time it was just so weird to... Yeah. Again, define your API like that and stuff. But also I think we were just so excited about Lambda in general and just its event driven qualities and whatnot. I mean, it's a really amazing thing maybe this is kind of far out there, but it has never been easier to write code that reacts to events. Right. You could just go put it in a Lambda function, tonight I'll deploy 1000 functions, right. 200,000 different things, just sitting waiting for something to happen.

And in some ways, I don't know, maybe that's how humans work. Maybe that's how businesses work. They're just kind of logic sitting there kind of waiting for events to happen and then it runs. And this felt like just such a nice natural model that could really scale. And yeah, so it felt natural. Looking back certainly fast forward ahead, and we can certainly dig in this later, serverless has grown a lot, the use cases, the types of infrastructure, all that. Is that still the right model? I don't know, but I will say that developers still love it. They get it so easily. And most of them, you know the interesting thing for us was when we first surveyed our audience we realized that 30% of our users had never even used AWS before.

Jeremy: Oh wow.

Austen: They came to the framework because I think it just surfaced the few things that you really need to know about on AWS and gave you a simple model to deploy serverless architectures on Amazon. So yeah, that was the theory back then. And again, just me working nights and weekends on how do we bring order to this complex new architecture that has these great values, but unless we bring order to the architecture, no one's going to be able to realize that value, that potential. And so that was the solution that was designed. Now, there's another part of this though, it's kind of interesting to talk about and that's the marketing. And marketing I don't think is something that developers think a lot about or the cloud industry. I mean especially maybe more so back then now it's increasingly important and whatnot.

But my background, I actually grew up outside of Hollywood and I was always around growing up. I'm a self-taught programmer, but I was actually hanging around a lot of screenwriters and directors and people in the film industry. And I had learned so much from them around designing product for emotional impact. That is taking something and wrapping it in a story. Because if you can get a great technical solution and wrap it in the emotional charge of a narrative or a story or an idea, then I think you could really deliver a more profound effect on like the end user more profound message and experience overall. That's your personal product philosophy is kind of this weird... I had this weird hybrid of being around that culture while also being an engineer.

And now I think a lot about that in terms of how you bring products to market and whatnot. But looking back, I think that was equal equally important, if not more important piece of this whole thing. And because I spent probably just as much time trying to come up with this framework and this technical solution as I did designing some marketing and whatnot. So hence for those who don't know, the serverless framework was originally called JAWS and it had this cool shark mascot icon and the big, bold, all caps JAWS texts. And I had recycled that branding from another project. Also my background is in design motion graphics. And I had recycled that branding from another project, but I guess I was kind of trying to find a way to embody how big I felt this architecture was and this framework was, right.

I was trying to personify it in a way that felt like a blockbuster. And JAWS was kind of like one of the original blockbuster films way back in the day. And I wanted to feel like a big deal. And so I take some of those mascot, the branding, all that stuff and wrapped it in this cool experience. But the last piece of that, probably the most important piece of that was the term serverless, right? Because these days the word was not really around at all. Lambda was always event driven code, glue code. But I had read while I told Jeff Barr, I was like, "Okay, give me a few weeks or something. And let me try and figure this out."

I had read a blog post on the Amazon Compute Blog from Tim Wagner and it was the first time I ever saw the serverless buzzword, the serverless word in general. And he had wrote in there like buried in the middle of the blog post, there was one sentence. And that was, you could use AWS Lambda to build entirely serverless applications. And as soon as I saw that word, I thought that's a great word. I love that. I'm not sure what it means. Right?

Jeremy: Right. I don't think we know what it means now, but sure.

Austen: That's a whole other podcast debating that. Right.

Jeremy: Right. Exactly.

Austen: But the developer in me, the maker, the person who just wants to hopefully build cool things one day. I just loved that word because it meant to me like the technology that kind of gets out of your way and less management, all that stuff. So I loved it. And I started putting it with the JAWS branding all over the project. So I totally admit a lot of over the top propaganda.

Right. So it was JAWS, the monstrously scalable serverless application framework was kind of the tagline with the shark icon. And he's like coming out of the water, trying to take a big bite of the text. And then on the GitHub README there were some badges, 100% server free, no servers guaranteed. And because we're kind of this badge-oriented society, right. Where we're looking at GitHub repos like, hey, does this have all the badges I need? And you're at the market looking at all the badges on the product? And it was just over the top. It was fun. I wasn't really thinking about anything kind of later down the road. And so before I had a chance to even talk to Jeff Barr, I did the typical Hacker News post.

Right. So that was like on a Tuesday, I was about to go out to lunch. And before I did that, I thought before I send this over to Jeff, I'll just put it on Hacker News real quick, because it'd be great to get some feedback before Jeff Barr's audience and the AWS audience hear about this.

And I posted it and I just walked away, had a sandwich someplace and I came back and front page immediately tons of upvotes, a ton of like enthusiasm, people were going and they're pretty excited about it. And I had no idea this would happened or anything like that. It wasn't orchestrated, it was just such a casual thing. And literally overnight just it caught on and it just kept picking up momentum like nothing I've ever seen before.

And I know you're a fellow entrepreneur, developer, and I'd built so things before this, right. So many projects, right. Behind every successful project there's 100 skeletons of projects.

Jeremy: Absolutely, right.

Austen: And sometimes I think there's the notion of product market fit and sometimes maybe you get lucky and I think luck has a lot to do with it, timing, all that. You just hit the nail on the head and it works. And all of a sudden you know it when you see it, it really takes off and starts to form a life of his own and it kind of almost becomes beyond your control. And it was just like that. And it was a totally, totally phenomenal experience.

Jeremy: And so I have two questions for you. One, did JAWS stand for something?

Austen: Yes. So one of the projects I had recycled the branding from was just the JavaScript AWS framework was the goal. So I'm a big JavaScript fan and there was never at the time, there just wasn't a good application framework for AWS. Something that could just help developers be productive on AWS with, again, not knowing a lot about the cloud infrastructure, it's kind of a big theme of mine and especially for the JavaScript community. So I had kind of designed a JavaScript application primer, but it was really early. I think I just worked on the branding and kind of got caught up with that before the project was there even. So, yeah, it was the JavaScript application framework, but then when Lambda came out, recycled the branding all that, and I kind of ditched the acronym because it wasn't as important.

Jeremy: Right. So then my other question is when did you buy serverless.com?

Austen: Okay. Yeah, so-

Jeremy: Or am I jumping ahead too far?

Austen: No, no, no, this is a great question because it kind of starts going from JAWS to Serverless Framework and establishing the company of course. So I think after that initial Hacker News post, things picked up a lot and I was working hard just to promote it and get it out there. And people were really excited to hear about it. So I was doing a lot of events. I was in San Francisco, AWS has one of the AWS Lofts here. And I was kind of demoing the framework almost every week. And I remember like the first time I met Tim Wagner in person, he was there doing a presentation on Lambda. And I just approached him like right before he was going on stage.

I said, "Look, I built this great application framework. I'd love to just show it off. Will you give me some time on stage to just show the audience here." And Tim's so cool. He's just like, "Sure." And I was pretty excited, you could probably tell the enthusiasm. And he probably had a hard time saying no to me at that point because I was just so pumped up at everything about this project. And so he let me on and I presented there a lot and also re:Invent too, the Lambda team, Tim, Jay, they told me the last minute you want to go present at re:Invent at 2015. Yeah. And I was like, "Yeah, absolutely." And I had even done one other cool kind of maybe growth hacking thing.

And or maybe just something that came from my unconventional background, but I had put together a JAWS Serverless Framework, movie trailer. And I don't know if you've ever seen it.

Jeremy: I haven't. But send it to me and we'll put it in the show notes because it will great.

Austen: Yeah. So I took clips from the movie Jaws and I was going to show it on this, but if you put this on YouTube or something we'll probably get in trouble and there's like music in the background and stuff. And it was like, I just took clips from the movie Jaws and I put these titles, like the Serverless Framework is coming for your infrastructure. Intercut with the clips of like the s,hark from the movie, in the water you can see people swimming, you never see the shark, which is the early brilliance of Spielberg when he was making that movie.

But you just see the innocent victim, shark, POV and stuff. And I had posted that before re:Invent and I was so excited. Werner the CTO over at Amazon had retweeted it. And we were like, this is mesmerizing or something like that. And anyway, I've been such an AWS fan for so long. It was so cool to kind of go through that whole experience personally. And so, however, at the same time when this thing was picking up some momentum, that was when, I don't know if you remember, but there was also equal amounts of kind of shade and doubt and being thrown at this whole kind of burgeoning kind of serverless movement thing. And I totally think, well first off, the first time I posted the framework on Hacker News, I think like the first comment or something was this is a horrible idea. And I think that's how they say hello on Hacker News, but-

Jeremy: Right, exactly. If you don't get criticized on Hacker News, I mean, what's the point.

Austen: Exactly. Some skepticism about the product and the project whether it would actually work for real world use cases, which is par for the course, if you're building out something new. And also early days back then cold starts were more of an issue. It took the whole architecture, kind of get out of that cold start kind of fear, uncertainty, and doubt territory. And now we hear about that less and less AWS has done such a great job. All the other infrastructure providers have done such a great job to reduce that cold start. I even remember when I was first chatting with the Lambda team, someone who will go unnamed just like, "After you deploy the application, could you just ping their Lambda real quick to warm it up behind the scenes."

And I thought like, "I don't know. I don't know if I feel right about that. I'm sure it'll get better. Just keep doing the great work that you're doing." And then there was, we were talking a lot about NoOps back then, right. That was a big... I think everybody was kind of starting to get excited about Lambda and stuff and then NoOps term came out, which I don't think it's right. And for this architecture, I think it kind of alienated some people out there. So there was a lot of kind of pushback on the whole idea. And meanwhile I was having some challenges with the logo and the branding. Right. Which I should have known, I mean, it sounds like I should have known, right.

But the trouble was not actually from Spielberg or universal or something like that, because trademarks are usually registered respective to a specific type of service or good. And there was another project out there, a screen reader application called JAWS, which is essential to a lot of people to gain access to computers and start writing code and stuff. And so I remember I put like when JAWS was kind of picking up some momentum, I decided to try, I'll put a sponsor link on here and see if someone might send a donation and the only donations I ever got were from people who kept sending me $1 donations and they always put the hashtag like, change your name, the visually impaired community or something. And I got a bunch of these and then I started to get more and more threatening emails which, totally understandable.

I mean, it's like you're born into this world on someone else's property and you've got to figure out how to kind of carve out your own space. So I knew like almost immediately after it took off, it was going to be kind of a challenge. And while this brand and stuff I thought it was so beloved it seemed by some of the users of the project and everything. I knew we had to change immediately. And so also at the same time I was wanting to build a company around this. I mean, it very much felt to me even back in the early days and more so today that serverless is not this fad, it's the natural evolution of cloud. Right. And that is like the cloud and serverless almost seem like they're on a path where they're going to merge soon and severless will just be the cloud.

Right. And that's just what you should expect from your cloud infrastructure. And I wanted to build a company around this because I felt like there was a huge opportunity to build a next generation set of developer tools to help developers capitalize on this great super powerful cloud infrastructure. And so I was kind of going around and doing some meetings with the VC community around the Valley. Also super funny, it turns out when we raised a lot of our investors are actually Dockers investors, and yeah-

Jeremy: Interesting.

Austen: And I think they are very much looking at us in the early days almost as like a hedge to some extent. Right. And I remember in my pitch deck I had kind of JAWS the branding and everything and a lot of our community members in the early days were coming from the Docker community.

Because I think a lot of those people kind of realized that they didn't want to think about containers. They just wanted to think about product. They just want to think about building apps and getting it to market as fast as possible with the least amount of maintenance. And so we were getting an influx of that crowd and I had put out one of my pitch decks like the JAWS branding. And there was just one slide where there was the Docker whale and it was upside down in the ocean and it had a huge bite taken out of it and this kind of cross over its eye. So and then like the slide after that was just JAWS, that was the introduction to the pitch deck. And I remember I was just looking at the VC's portfolio again, like right before the meeting and I saw, oh, there are big investors that Docker, I took out that slide immediately.

But there was potentially some cool branding we were going to do around that. But anyway, so I was raising capital trying to turn this into a company, this kind of like the scrappy effort into a real company selling great dev tools to people who want to take advantage of all this. And so I had to change the name, I missed all of this. And again, it wasn't quite clear still that serverless was going to be the term, the category that it is today. All I knew is that serverless is the word that generated the emotional reaction in the user base. Right. That was the thing that developers were responding to emotionally that got them to lean forward in their seat and say like, "Oh, this is something different." Right. And so I had like 10 names and I was kind of running them by even some of the potential investors.

And I remember the feedback I got was serverless. Like, what does that even mean? That sounds really weird. Right. Yeah. And so I set my eyes on it and I thought, well, I got to get the domain and see if that's available. I've got to have a great domain because in .com we trust. Right. So I started the hunt for that and it took honestly two months of cold calling because I could not find two on the domain or anything like that. Two months of calling around trying to find this person and then like another month kind of negotiating because the person who owned it kind of sensed maybe that there was some opportunities, I was expressing interest of course.

So that took a long time and it was still just me working on this project, trying to build out the product, the marketing, the promotion, do the name change, do some meetings around Silicon Valley.

And I had somehow been fortunate enough to kind of be able to get the domain name and then eventually like close around the Capitol. And that was like, I think right at the end of 2015, and we became... JAWS transitioned over to Serverless Framework. And we became Serverless Inc. Which was, again, seemingly we didn't know it was going to be so big, but of course now we're in this awkward position where we're named the same thing as the category. Right. So now I hear all types of creative ways of people how they explain when they introduced the company or something, how they differentiate it from the architectural pattern, like the movement and the company.

Jeremy: Yeah. I know my youngest daughter, who's 12 now for the longest time because again, I've been using the Serverless Framework for so long, but then building things servelessly and whatever. So she always was getting confused like, "Wait, so serverless is a company, but also what you do... So do you work for serve...?" I'm like, "Okay, let me try and explain it to you." But it is funny cause I'm the same way. I'm like, ah, serverless capital S or the Serverless Framework, so it is... It is a bit of a challenge, but I think good for you, right. If you default to be sort of the namesake there.

Austen: Yeah. It's good. And it also really confusing, the experience your daughter went through is the experience that almost everybody had to go through at some point. A lot of people had to go through. I've got to take the time in the beginning of a lot of meetings and presentations to try and provide some clarity now. So it's very interesting, but yeah that was always the word to me. And it was just based on, I think something to do with my background. Thinking about how you design products for emotional impact and always looking at what really generates excitement in end users at the end of the day. And it was that word. And I never forgot the first time I read it, the credit goes to Tim Wagner. That's how I felt the first time I saw it. I don't know when you saw the serverless word for the first time. Do you remember?

Jeremy: It's funny. I honestly don't remember. I do remember in early 2016. So that must have been just as you were making that transition because I started using JAWS. I don't remember when I found out about JAWS. Must've been late because I had already started using Lambda. I'd already started playing around with Lambda as soon as it became GA. I wasn't paying attention to the re:Invent stuff. And so I didn't know about it in 2014, but then once it became GA I started using it. And then it was later on that year that I discovered JAWS. And then I continued to use JAWS though up and through the beginning of 2016. It wasn't until I think you switched over to the version one that you changed it to Serverless.

The history is just, again, so much has happened in five years. It's hard to keep track of all that stuff, but I don't know when the term serverless made it into my mind. It's sort of like, but I think you're right though. It was exciting. It seemed like something different. It was a different category of things. It wasn't event driven. It wasn't SOA or whatever, it was something that was like its own category.

Austen: Yeah, it's amazing. It's been five years and all this stuff that's happened since. But yeah that's the story behind the domain and the name and all the things that kind of led up into just transitioning from JAWS over to that. Looking back now, I mean, it just seems now serverless is like such a massive category, a lot of vendors in this space every major public cloud provider has a serverless compute offering. That part of it was just a wild surreal experience.

Jeremy: All right. So now we're into like 2016 at this point. I think it was still JAWS at that point, or whatever, but you started building a company around it, you started getting some other people involved. You made the massive faux pas community projects where you changed it to non-backwards compatible when you went from 0.5 to 1.0, or something like that. I remember being so angry at the time like, why are they doing this? I've written so much now in the old and the old format. And then as soon as the new format came out, I absolutely fell in love with that. So no problems there, but so what happened there? How did it start as a company? When did you start bringing people on? And when did you start growing this thing?

Austen: It's so funny that you brought that up because we still get grief for that. Like after it's been years, that was 2016. Now, it's 2020, certainly a lot of other stuff going on in the world, but I'll still hear from someone back and I can't believe you did a breaking change. Right. Especially at re:Invent. When we go to re:Invent people bring that up and they're like finally I got to meet the people who made that breaking change and tell them about it.

Yeah. I started as a solo founder which is a whole different type of experience for building a company. And I think if you have a co-founders I don't know what the right way is, but I do know that solo founders have to do a lot.

And so it took me a while to kind of build out the team and do everything because it was just one person trying to still build out the software, all that. But the best part of all of this is the community. It's the people who helped along the way and made a huge impact. So going back all the way to Tim and Jay on the Lambda team, all the great folks there who were pioneering. And I remember the early conversations we had back then, not even they really knew what they had yet it seemed, right. Compared to how they talk about it today. And everything just seemed so clear and obvious now, but back then the conversations were just all over the place. People were trying to describe it. There's a weird thing that you see sometimes when stuff takes off, like something takes off before people can describe it.

Right. And that's just an interesting phenomenon. So they did a great job, I think, early on back then, 2015, 2016, there was a lot of AWS community who were kind of helpful along the way. Jeremy Edberg, if you know Jeremy, Peter Sankauskas, I can never pronounce his name and his last name correctly. Mitch, like these were the original at least in back in 2015, the cool kind of AWS luminaries. Right. And I'm sure you see this. And but it's like the people who are in kind of the AWS user community, like there's different waves of people, but back then it was like Jeremy, Peter, and Mitch. And they were certainly helpful. I spent a lot of time chatting with them about this.

We had another friend of mine, Ryan Pendergast, who I haven't spoken with for a while, but he was instrumental in really helping design the first version, or then the first few versions of the framework after we got launched and shaping where it would go. He was great. And back then also A Cloud Guru, I mean, credit to them for being some of the early visionaries in this space. Sam and Ryan. Ant working with them. They were doing the original serverless conferences. So I don't know where you at the first one in Brooklyn? I can't remember when that was. I think it was 2016.

Jeremy: I wasn't, no I wasn't.

Austen: It was such a funky fun conference that was amazing. It was like this cool little kind of almost dive-y conference venue in Brooklyn. And it was like the middle of the summer in New York.

And like everybody's sweating. There was like no air conditioning and there's only, I don't know how many people were there. I don't even know if we hit 100 or anything, but A Cloud Guru got in front of this and really helped promote the movement and all that. And so they should get a lot of credit for that. Ryan Scott Brown, Jared Short, as you know he was big contributors early on. Jared especially loved the JAWS stuff. I remember he loved sporting the JAWS hoodie. And Herique, Anna from Red Badger, Marcia who's at AWS now. Alex Casalboni. I mean, there's just been so many brave people. Rob over at Nordstrom. Eric was at Nordstrom. It was one thing to be a solo founder.

But it's another thing to be a solo founder with all these great people helping out via open source, working on this project, trying to define we all knew, hey, there's a great new architecture here that could really enable more people than ever, but we got to make it easier. You got to make it accessible. Otherwise people will not be able to realize all this.

So I think we only scaled to like a few people in 2016. And for the first couple of years we just focused on community development. That was it. We didn't do much of anything else. And that a lot of that credit goes to Dan Skolnick who was kind of one of the original investors in the company. He was over at Trinity Ventures, now he's at another firm and he was just like, I think first money in the Docker or at least kind of one of the initial investors in the Dockers.

He's got great instincts. And I remember we had a conversation where he said just focus on building the community right now. So just make that community as big as it can be and just take time to focus on that. And that was really instrumental for the company, I'd say. To just to hear that, especially coming from an investor, you know what, don't worry about anything else. Just like build out community and all that. So 2016, 2017, all the way to even like 2018, it was just a lot of like working on the framework, which as I'm sure you know is no small feat because that thing has grown so much as an architecture in terms of like the surface area it has to cover because it's in a tough position of trying to abstract, create a simpler experience over a growing amount of AWS infrastructure and configuration options and composition options.

It takes a lot of calories to keep that project going. So without the open source community, without like the blessing of our investors. I don't know how big the framework would still be today, but that was what we focused on back then. And I would say just instrumental to kind of the growth and the scale that the framework has right now.

View Details

About Holly Mesrobian

Holly Mesrobian is a Board Member at Cascade Public and the Director of Engineering for AWS Lambda. Holly has 25 years of experience in designing, building, and managing globally distributed teams in software development, and more than 15 years as a leader of leaders. She has in-depth experience with building services for builders, and for wireless and broadband carriers; online services for direct to consumer offerings; and commercial shrink-wrapped software. With a double Master’s Degree in Computer and Science and Software Engineering, Holly began her career as a developer before holding leadership positions at companies, like Amazon and RealNetworks, and startup Cozi.

  • LinkedIn: https://www.linkedin.com/in/holly-mesrobian-a1b710/
  • Under the Hood of AWS Lambda 2019: https://www.youtube.com/watch?v=xmacMfbrG28
  • Under the Hood of AWS Lambda 2018: https://www.youtube.com/watch?v=QdzV04T_kec

Watch this episode on YouTube: https://youtu.be/nBYUh7CVUiQ

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today, I'm chatting with Holly Mesorbian. Hey Holly. Thanks for joining me.

Holly: Hi, thank you for inviting me.

Jeremy: You are the director of engineering for AWS Lambda at Amazon Web Services. Why don't you tell the listeners a bit about your background and what the director of engineering for AWS Lambda does?

Holly: Absolutely. Engineering leaders in Amazon are very technical and I think I fit in that class of leader. I've been in engineering for more than 25 years. The first decade that I was in, I was actually an engineer, and then the last 15 years or so, I've been leading large-scale engineering organizations that are also responsible for 24/7 operations. You think about those, they're called DevOps organizations. That's what I've been doing for quite a while now. The Lambda engineering organization is just like that. In terms of my background and how did I get here?

I have two graduate degrees, computer science and software engineering, and as I referenced lots of time, designing and building systems. One of the things that's really great about AWS and leading the teams here, I reference that DevOps culture. My teams, they build it, they run it and they have really great best practices around engineering excellence and operational efficiency. If we have an issue in one of our production environments, my teams are on it, and we have great processes around how we do that. We have a really well established COE.

Anytime there's a customer-impacting issue that happens in one of our production environments, my team's right. COE, which it means correction of errors, we review it as an engineering team every week. I sit down with my teams, we go through operational dashboards, we inspect our metrics. We look at how we're doing across the availability latency scale. We have ongoing scaling targets and scale testing. We're constantly inspecting how are we running the service? Not just how we're building it and how we're building new features, but how we're running it.

We run game days as well, so that we try to break our systems and then see that my team, all my on-calls can recover those systems. One of the things that I really like is we put new people in the team on those game days, because where better for them to learn than we've intentionally broken the system. Get in there and figure out if you can fix it before it's actually fixing something in production. That's really great about Amazon.

Then I would say the other great thing about Amazon and Amazon engineering and the teams that I have, I just love what a high caliber they are and how invested the members of the team are, and how hard they will work to try to make the best service for our customers.

Jeremy: Awesome. Well, listen, I am a huge fan of AWS Lambda and I love what you do. I love what your team is doing. Everything that Amazon is doing for serverless is just amazing. One of the things though that I'd love to talk to you about today, and we could get into the specifics of Lambda itself and how everything works, but you have a couple of great talks. You and Marc Brooker did these talks at re:Invent in 2018 and 2019, getting into the details of Lambda, Lambda Under the Hood, right? Great talks`.

If anybody wants to know exactly how Lambda functions work and how all that stuff works under the hood, definitely go check those out. I will put those in the show notes. What I'd really like to talk to you about today is just this idea of serverless adoption or serverless transformation, because I know AWS talks a lot about how all their internal tools are going serverless, right? Which is pretty cool. Of course there's the external stuff too. There's a lot of adoption from enterprises and small businesses and medium-sized businesses and things like that.

I would love to know the mindset internally. How does AWS take serverless or look at serverless and look at Lambda and use that to build their internal processes? What's the learning on that? How do you keep learning and just keep building with serverless?

Holly: Yeah. This is a really fun topic for me to talk about, and as you might imagine, customers find value in the agility and the operational load or the lighter load on operations that serverless brings. My teams are no different nor our AWS teams or Amazon teams. What we have seen over time is teams across AWS adopt and use serverless. Then my own teams over time have also adopted the serverless architecture and they actually want to use it.

Over time, more and more of the Lambda service, in particular on the control plane, because you don't want circular dependencies in your architecture. So we're really careful about making sure that in early design, when we're saying, "Hey, my team wants to use Lambda, is it okay to use Lambda and serverless?" Because it's building serverless underneath serverless and you have to be careful that you're not doing a bad thing. We're really good about inspecting that in the early design phases.

I've seen more and more of my teams picking up and building control planes on Lambda. In particular, they're using the feature that we launched last year at re:Invent called Provisioned Concurrency. What that does for really high-scale, low-latency services is it gets rid of what people have typically talked about, which is cold starts. Of course, we've done a ton of work over the years to reduce cold starts, but they're still not zero.

We're going to continue to do work on cold starts, but for customers who are super latency-sensitive and need that scale and know that they have low latency all the time, Provisioned Concurrency is a great solution. We have used it within our own services as well.

Jeremy: Right. Is that something now where all AWS teams, when they're thinking about building a new service, that they're going to build that on top of Lambda and do that serverlessly?

Holly: Yeah. One of the ways that people look at it is it's that operational model and where are you sitting in that? Of course Lambda's pretty high. We do a lot more of the shared responsibility on behalf of customers, and so teams like that and they say, "Oh, well, this is going to be easier to operate. We're going to get more agility out of it, so let's go there as a first stop." It's only when they say, "Well, maybe this isn't going to work for us."

That they go to the next potential option or an option after that. Like I was talking about earlier, I see my own teams doing that same evaluation and we're increasingly using Lambda to build portions of our service.

Jeremy: Right. Right. Yeah. Another thing ... And I think that maybe we don't always think about, or maybe people don't always connect these dots, but you can't run serverless in a vacuum, right? You can't say, "Hey, I'm just going to build everything on Lambda, or I'm just going to build everything on DynamoDB." You have to talk and interact with a number of different services in order to make that happen. You think about some of the recent launches, so V2N, Provisioned Concurrency, EFS for Lambda.

These are services that Lambda has to use in order to handle some of these use cases, and because Lambda really pushes the boundaries of these services, you end up making these services better, right?

Holly: Absolutely. To your point, Lambda stitches together a ton of AWS services. I think about it as Lambda is a lot of the glue between AWS services. In a number of the features that we've built, and you referenced V2N and EFS, both of those services, we worked very closely with the teams. You can think about them as joint projects, like at a project team level where we're in the room with the leadership from both of those teams every week, talking about any issues that we're seeing or how the project is progressing.

In the process, we make those products better and the products become better products. One great example is on EFS. Because of Lambda's unique performance characteristics, the scale, the instantaneous burst, we drove EFS to deliver higher burst capability from 1K to 25K, and so the product becomes better for everyone, not just Lambda customers or serverless customers, but all customers in the process of doing that.

We also do lots of joint collaboration and work to make sure that those services are operating at that level as well. We spend a lot of time in the development cycle, ensuring that the products worked really well together.

Jeremy: Right. That's one thing that I'm curious about in terms of like, what is the serverless vision for AWS, right? Or at least from your perspective from the Lambda team, with all of these new launches, all of these things that have come out, I mean, this has solved a bunch of new use cases or have opened up a bunch of new use cases, right? I mean, with EFS, you get the ability to do maybe ML, for example. You've got the Sam CLI that just went GA. You got RDS Proxy that just went GA to help solve the connection pooling issue.

You've got Savings Plans, you've got Provisioned Concurrency, all of these things we mentioned. Is that something where you're pushing or Lambda is pushing the other teams to help you solve these use cases so that more internal teams as well as customers can start using serverless?

Holly: Absolutely. In Lambda, as you referenced, we have continued to work to drive increased adoption and remove barriers for specific use cases. Rolling back all the way to a couple of years ago, we launched Firecracker, which made our service faster and helped reduce cold starts. We then launched V2N which brought that capability and lower latency to networking or VPC networking. Then we launched Provisioned Concurrency because we were hearing from customers that they needed that low latency all the time.

You roll forward. We just launched in June, EFS. EFS is really designed to help make sure that customers who haven't been able to bring their really big data-intensive workloads that they've been wanting to bring to the simplicity of Lambda to Lambda. If you think about it EFS the workloads, like bring a model and run something on that model, or bring big data and do big data transformations that you can do this really simply with Lambda. The data's there and you can do it when you want to do it.

You're not holding a bunch of capacity to do this highly scaled, highly parallel data processing. Lambda is great for that and that's really why we built EFS.

Jeremy: Yeah. No. I love the idea of EFS because it does, it opens up so many more use cases. That was the common complaint with serverless was always like, "Well, you can't do machine learning with it," or something like that. This is just one of those things where it gets us much closer to being able to do those sort of thing. All right. Another thing that is a launch or a feature that I think is absolutely amazing and I think it was last year at re:Invent, was Lambda destinations.

I love, love, love, love this feature because I have so many workloads where you have some background processing happening and when the process has finished, you want to tell somebody or tell something that that process has completed. Rather than putting all of that code in there and having to call a separate service and do all that other work, it's just so much easier for you to say, "Oh, when this is done and this completes successfully, fire something off to EventBridge or to SNS or put something in a queue."

Or, if there's an error, you get all that extra context information. I really, really love this service. I would also love if maybe there was a synchronous version of it as well so I didn't have to write code maybe at the end of a function. I could send some data off somewhere else too. Maybe that's a different discussion. I think what this opens up is this idea of ... And maybe I should say, maybe confuses people a little bit, is this idea of function composition.

This is when we want to have one function end and send information to another function and so forth. Obviously there's two different ways to do this. We can do choreography. We could use something like EventBridge and coordinate them or SNS and coordinate the results of functions. We could also use orchestration and use a state machine like Step Functions. I love that you now have all these different options and I love what you can do with Lambda destinations.

What are the use cases for that? Are we talking about just small workflows and then more complex workflows use something like Step Functions? Or what was the intent of building Lambda destinations?

Holly: Yeah. That's a great question. When we built destinations, it's really designed so that you ... You used to just hit an asynchronous event and you would fire and forget it, and you wouldn't know really how it continued or be able to pick it up. We built the destinations in order to allow people to do those completions, those continuations. In terms of, if you get to those really complex use cases of, "Hey, do this then that," and all this logic and branching and things like that, then Step Functions is a great way to go, because it's really designed more for those complex workflow situations, and is probably going to be an easier use case for that.

Jeremy: Right. Yeah. No. I love Step Functions. I mean, I think that the way that you can do parallelization and that you can fan out and you can fan stuff back in, and you've got all the wait timers and things like that. I mean, it just a very, very good solution. Back to the destination thing though, so one of the things that's really great about them is again, the reduction of my code, because every time I write code I'm introducing some liability into the system.

If I can just finish something and the output of that automatically gets sent and guaranteed to go to some of these services, that's really great from an asynchronous standpoint. From a synchronous standpoint, I would love to have that too. I don't know if that's something you guys are thinking about or is something that you would potentially put in, but I really do like the synchronous use case.

Holly: We haven't seen a lot of requests for that use case.

Jeremy: All right. Well, I would like you to add it. That's on my AWS wish list.

Holly: Okay. I've got it in the intake.

Jeremy: Awesome. All right. All right. I want to move on to serverless architecture in general and maybe just application architecture in general, not only how AWS does it, but how other people should be doing it. We know that there's a lot of isolation with Lambda functions and with Firecracker. I mean, so you're pretty good from a blast radius when you build single-purpose functions. Of course, there's the microservice pattern or microservice designs where you put a couple of Lambda functions together into a single cloud formation stack and things like that.

I'm curious, how does AWS add additional security or build bulkheads? Is that something you do in a single account, or do you have multiple accounts or a separate account for each microservice?

Holly: Yeah. We recommend using a separate account per microservice and then also thinking about an account for each of your environments as well, your pre-production environment and your prod environment. Each one should have its own account as well. What that does for you, if you think about it, a lot of times, two pizza teams own a service or a small set of microservices, and you want to reduce the number of people who can actually access those services and make changes. I mean, it's an operational risk.

It's also a security risk having too many people have their hands on a microservice. You really want to make sure that the people who can access it are knowledgeable and know what they're doing. That will help you have a high availability as well as ensuring security. Of course, availability comes back to not only potential for someone to make a change that is a breaking change, but also things like ensuring that your limits are used and planned for in a way that makes sense for you.

Jeremy: Right. Yeah. No. I love that and I'm glad that we've cleared that up. You've heard it here, AWS separate microservice or separate account per microservice. I do love the idea of microservices and I love that you have all of that isolation, even another level of isolation, and you have the ability to set the limits. Your concurrency limits and all that stuff can all be set per a microservice.

Holly: Yeah. I really like it too. Hopefully, everyone who has a business, that business is growing, you're scaling your service. I know we're scaling our service all the time. You build new features. You break apart services into smaller units to help scale with your teams. As you do that, if you thought about it as microservices with account permissions, then it also makes it easier for you to transition service ownership over time and have a new team pick it up with just that group having access. Kind of [inaudible 00:18:16] for growth as well.

Jeremy: Right. Yeah. No. I love that. I absolutely love that idea of breaking things up because you also have all this extra control over things like concurrency. You can control all those different things. Those different limits are controlled per group. Then the other thing that is great about that too, is each individual account is going to have its own roll-up of billing, so you can actually see what the cost is per microservice, which is pretty interesting.

Holly: Right. Your ownership as well, right?

Jeremy: Right.

Holly: Like when you want to go to your teams and say, "Why is this being billed to me this much?" It's really easy to go and tease that apart and talk to the right team member and get the right answers rather than charging around a very large organization try to figure it out.

Jeremy: Right. Right. All right. Awesome. The other thing that I'd love to talk about is this idea of what are the next workloads at that Lambda is going to be able to handle? If we think about machine learning, EFS handles some of that, but there's a lot further to go in order to handle that type of use case. Maybe support for other legacy databases, maybe even just larger memory, for example. What are some of these new use cases that we're hoping to unlock?

Holly: Yeah. We're looking and continue to look at different types of compute as well as larger memory sizes. Of course, larger memory, because the way we do, the memory and CPU go hand in hand, so larger memory also implies more cores. Again, it's back to that big data, more compute-intensive workloads that we know can be unlocked by bringing more memory and more CPUs to customers' specific use cases. Then the one thing that I know we get asked for, it's increased duration, but one of the things ... And 15 minutes, is it right? Could it go further?

It could probably go further, but one of the secondary considerations is we want to cycle our customers' execution environments. The reason why we want to cycle those execution environments is because that helps with security. You don't want something really that runs indefinitely. You want it to be bounded in time, because then you know that you're getting that cycling of the environment, which then you know that you have a clean environment and that your workloads are safe and secure.

Jeremy: Right. Right.

Holly: Yeah. Those things come hand in hand. Do you want to cycle? How often do you want to cycle? The longer you go you don't want it to be indefinite.

Jeremy: Yeah. No. Right. No. I definitely agree. I mean, I think 15 minutes is maybe a little bit arbitrary, but it's a good number. I mean, maybe 30 minutes would be better for some workloads or whatever. Certainly, that's one of the things I love so much about Lambda is the statelessness of it. The longer that container runs, the more potential there are for everything. Memory leaks, security issues, or whatever, even if you've got variable saved in your global scope and things like that.

Yeah. No. I mean, but maybe bumping it up a little bit. I don't know. That might work. What about GPUs? Have you heard anybody asking for GPUs?

Holly: Yeah. I think that was when I ... I didn't talk as much about the potential for GPU, but we certainly hear from customers an interest and we are evaluating that as well.

Jeremy: Awesome. All right. Here's something that ties back to the vision and maybe this is the AWS vision, maybe this is your vision, but are we ever going to get to a point where like a hundred percent of our compute is serverless? Where we have just no need for containers, sub-millisecond cold starts. Maybe we have some coordinated parallel compute where Lambda function can talk to one another. Maybe just from the Lambda perspective, is that your goal? Is it to, I guess, take over the compute world with serverless compute?

Holly: From a Lambda perspective, we're going to continue to work to remove limits and allow customers to bring more and more workloads to Lambda and to serverless, because we think it has such value. In terms of, will we get to a hundred percent? I think no one knows and only time will tell. From a Lambda perspective, we're going to do everything we can to continue to make it a great platform for customers and to remove things that get in the way of that for them.

Jeremy: Right. Yeah. No. I mean, for me personally, I would love it if I never had to manage another container or a server ever again, and every use case was solved. I think one of the big ones that comes up is the idea of cold starts. I mean, you look at a couple of other platforms that exist and again, they might be limited in terms of their language support and things like that, but some of them maybe they run on like the VA platform or something, have very, very, very, very low cold starts. This is obviously a continued complaint.

I mean, I know we've got Provisioned Concurrency, but is that something where AWS is going to continue to keep pushing and pushing and get that cold start down to where it practically doesn't exist?

Holly: We're going to continue to drive down our steady state case of cold starts. We absolutely continue to work on that. We put a lot of focus into both our warm and cold latency to make our services as fast as possible. We put Provisioned Concurrency there to address customers' immediate need, because we know it takes a little bit more time and some real heavy engineering lifting to address it without that, but over time, those will get closer and closer to parity.

Jeremy: Awesome. All right. I want to talk about this idea of the Lambda supercomputer. I'm sure you're familiar with Tim Wagner, obviously. You worked with him. In terms of this idea of Lambda functions that can run in parallel and can talk to one another. I think the way that that Tim was doing it with his test project was this idea of doing NAT punching and having them being able to coordinate with one another. This could open up a lot of use cases, especially, for big data, for genomics or anything big like that.

Are you thinking about making a way for Lambda functions to potentially mesh together and do this supercomputer use case?

Holly: Yeah. It's certainly an interesting use case, and it's something that I think Lambda is well-situated for, especially if you think about it from the standpoint of all the concurrency and bursts that you can spin up. You can spin up a lot of different nodes and then just based on routing the messages in the right way, end up with this large scale compute environment. I certainly think it is a possibility. It's certainly something that Lambda could do.

Jeremy: Right. Yeah. No. I mean, and Lambda, obviously you can use fan-out patterns and some of these other things, even Step Functions to coordinate and do parallel compute, but I do think it would be really interesting if there was a way for Lambda functions to directly talk to one another. I think that would open up some really interesting use cases. All right. Another thing I want to talk about is I guess, this idea of complexity in serverless. There are a lot of building blocks. Lambda is just one small piece of it.

If you're just building a small application and maybe you're just using Lambda and API gateway and maybe DynamoDB, and that's relatively simple. You can put it all into one cloud formation template, or using SAM or something like that. That's relatively easy. Then you start integrating EFS and then you need an SQS queue, or maybe you're reading off of a Kinesis stream or you're using EventBridge, or you've now got 15 different microservices, all separated into different accounts. It gets really difficult to wrap your head around the complexity that's there.

I'm wondering ... And I know that there's open source things for Terraform and there's the CDK and obviously SAM and some of these things. Is there something where maybe the Lambda team or AWS in general is looking at another level of abstraction? I know you've got SAR and some other ways that you can package up some use cases, but is there something on the roadmap for what that next level of abstraction is to make it easier for companies to come in and adopt best practices and things like that? What's the vision around that?

Holly: Yeah. SAM which you referenced is intended to be that next higher level abstraction for building serverless applications. We launch every new feature with SAM support. For instance, EFS just launched and we wouldn't launch it without having that support. We also are big believers in the broader tooling ecosystem. The reason why we believe in that is we don't want customers to have to learn yet another toolchain, if they have a toolchain that they love. We support meeting customers where they are with the toolchains that they find most comfortable with.

That's a dual strategy. We build SAM as the top level of abstraction for serverless, and then we support a variety of third-party tools as well, so that customers can use those.

Jeremy: Yeah. No. I think the support for third-party tools is great. I know with observability, there's a whole bunch of tools that are there that can help you. I mean, just from an adoption standpoint, as you add complexity ... And again, it's going to get more complex over time. That's just how these things work. Is that something that potentially is going to hurt adoption if it just becomes harder and harder to integrate these services into your existing toolchains or into your existing workflows?

I mean, even like testing, for example, it's very, very hard to test locally. You have to test in the cloud. What's the vision there to just bring it closer to developers?

Holly: Yeah. To that point, we are continuing to invest and have a real focus on the developer tooling and the developer experience. We know that that's an important element of serverless. It's not just having it run great on your data plane. It's also how are you interacting? What's the tooling? What's the customer experience? Then, how do you operate it in and out as an engineer? It's nice being an engineer. We're working with a lot of engineers and then going back to the ... and adopting it ourselves, we see where we can improve as well, even firsthand.

Jeremy: Yeah. No. I think that's super important from an adoption standpoint. Just, I know a lot of developers have a hard time trying to do the testing and wrap their head around all these different changes and stuff or just the different way that some of this stuff works. All right. I'd love to ask you this question too. I tend to ask a lot of my guests, where do you see serverless going in five years? Or where do you think serverless will be in five years? You actually have a lot of control over this. Where would you like to see things go?

Holly: Well, where I would like to see things go, and going back to our earlier conversation on why not serverless? I would love to see the industry be running on serverless, just because I think it brings such a great experience for engineers. Going back to my experience and you heard 25 years in the industry or 25 plus, I've seen all the phases. I've seen the phases of a technology adoption. I've seen what we've asked of our engineers over time. Back when I started, it was you learned a language and you learned it really well, and you programmed on it.

Then you ended up with polygon and you ended up with you're no longer on a box. You're driving a whole bunch of microservices and coordinating them together. Then testing, you used to have a test team who had test and now then you became the test team as well. All this stuff is good, but we've asked engineers to do more and more and more and more. I like that with serverless we are actually asking them to do less and to focus on the stuff that's really value-added.

I think that's a positive outcome for engineers. So when I think about it as a long-term engineering leader driving the most agility out of my teams, I think that serverless ... And I hope to see where serverless it's a why not.

Jeremy: Right. Yeah. Totally agree. Awesome. All right. Well, Holly, thank you so much for being here. I know that you've successfully managed to avoid social media, which is amazing. You're not on Twitter, but if people want to get a hold of you, how do they do that?

Holly: Yeah. They can connect with me on LinkedIn. I'm really easy to find. There are not that many Holly Mesrobians in the world.

Jeremy: Awesome. All right. Then also the two Under the Hood of AWS Lambda from re:Invent 2018/2019, I will put those in the show notes. Thank you so much for being here, Holly.

Holly: Great. Thank you.

View Details

About Raymond CamdenRaymond Camden is Lead Developer Evangelist at HERE Technologies. He’s an expert in web standards, Node, and serverless with a passion for teaching others. He has written about, and presented on, technologies for the past fifteen years, and enjoys helping others become passionate about the web as well. Ray is a prolific writer, as evident from his blog content as well as in industry publications. He has authored (and contributed to) multiple books over the years, such as his book Developing Serverless Applications, and speaks at conferences around the world.

  • Twitter: twitter.com/raymondcamden
  • Blog: www.raymondcamden.com/
  • HERE Technologies: www.here.com/

Watch this episode on YouTube: https://youtu.be/2I3nNsU6HSs

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today, I'm chatting with Raymond Camden. Hey, Ray. Thanks for joining me.

Ray: Thank you for having me.

Jeremy: So, you are a Lead Developer Evangelist at HERE Technologies. Why don't you tell the listeners a little bit about your background and what HERE Technologies does?

Ray: Sure. I'll start with HERE. So, we do everything involving location. So, we have a mapping platform. We have APIs for routing, geocoding, reverse geocoding, anything that a developer may need in regards to location or mapping we have technologies and tools for. People are free to reach out to me later to talk to me about it.

Jeremy: Awesome. And your background?

Ray: Let's see. I've been in web development since '93 or so. So, I've been around for a while. I spent a long time doing back end work. Last decade or so, more front end work. I've been involved in developer relations unofficially for most of that time, because I like to share, I like to give presentations and stuff like that. Officially in DevRel for six or seven years or so.

Jeremy: Awesome. All right. So, if people don't know you, if anybody was ever involved in the ColdFusion community, you are a legend in the ColdFusion community. I was looking through my old books trying to find some of the old books that you had written, I think ColdFusion MX 7. You and Ben Forta had so much material out there. I was a huge fan and always fall into stuff that you guys were doing back then. So, I'm super excited to have you on the show.

What was kind of funny is I was a ColdFusion guy way, way back in the day, as well. I had a big customer, it was a college that was doing everything ColdFusion. So, I learned ColdFusion just for that customer and I actually fell in love with ColdFusion. I thought it was a great language, it was super easy to use, all kinds of features, but I eventually got away from that and I really wasn't following what you were doing anymore.

Then, all of a sudden, I come around maybe two years ago and I see you're working on serverless stuff. So, I would love to get that perspective of how you went from ColdFusion to serverless.

Ray: Absolutely. So, I began actually with Perl CGI scripts back in the old days where there were no real defined roles. I would do everything HTML, Photoshop work, and back end work. I discovered pretty quickly that I don't make things pretty. I enjoyed the back end work because back then, having a web page be dynamic was a big deal. I can remember at college finding some random website that would pick three tarot cards and it was random. I was amazed by that. It would take five minutes to load the images. I would sit there and just reload. So, I kind of migrated there.

I got into ColdFusion because I had been doing Perl CGIs for a while, but we had a client who had a SQL Server system. I knew that Perl could talk to it, but also knew that it wouldn't be fun to write that code and just randomly discovered that ColdFusion supposedly made working with SQL Server and Access, et cetera, it made that easy. That was right. I kind of fell in love and spent probably a good 15 years doing just ColdFusion stuff.

Jeremy: Yeah. No. I mean, it's funny because I have, I think, pretty much the same exact background. I started in college with Perl and CGI. Then, I went to PHP and MySQL and sort of that stuff, but then moved over to ColdFusion. I spent many, many, many years with ColdFusion until I moved back to PHP.

So, when you got to the end of that ColdFusion era, what was the next step after you sort of were done with ColdFusion?

Ray: So, I remember when JavaScript came in, but it quickly became hard to use. You know the whole browser war type thing?

Jeremy: Yeah.

Ray: So, not only with me not being able to design well, the other reason I didn't like doing the front end work because it was such a pain to make everything compatible with IE4 and Netscape 4, et cetera, so working on the back end, somebody would just hand me HTML and I'd make it dynamic.

What got me kind of looking back is I remember I was bored at work. Ajax had been around. Gmail, I think, had just come out, so the whole Web 2.0 thing was fresh in people's minds. I had only done a tiny bit of JavaScript over the years. I thought, "Let me look at it some more," now that supposedly things are a little bit better. And I found out that, yeah, things were a bit better. Browser dev tools were around.

And so, I just started doing more work on the client side and I began to realize that having the app server there, primarily because browsers were so horrible, wasn't necessarily the case anymore. There was a lot more that I could do on the front end and without an app server.

Jeremy: Yeah. No. I mean, I think that was one of the things that was the biggest pain for all of us that were ... I mean, I owned a web development company for 12 years. So, battling with IE6 and all the different browser compatibilities was definitely a huge problem, which is why one of the things that I think we did or we did well was build back end applications that would completely render the pages.

And as you said, as we move towards this more of an Ajax, sort of filling things out, and of course now, we're into a whole new era with single page apps and React and Vue and things like that, but what was it, though, that once ... So, once you started building some of these more interactive JavaScript side of things, were you still using ColdFusion or did you move to something else for your back ends?

Ray: Well, so it's interesting. Yeah. I was still using ColdFusion, but what I began to notice is that ColdFusion began to get more and more dumb in terms of what it was doing, whereas before, I was rendering entirely the entire site. It began to be just a JSON provider.

One of the things that kind of got me off ColdFusion is that, at the time, it did JSON very badly. It would have this wonderful feature where, because it was typeless, it had to make guesses in terms of what your data was. If you would try to take someone's name, let's say possibly an Asian person who's last name was No, Dr. No, maybe. It would, in JSON, convert that to false. You had no control over that at all. So, both seeing how small I needed ColdFusion and seeing it kind of fail in terms of generating JSON, that began my migration to more Node.js development.

Jeremy: Yeah. And so, when you started using Node.js, were you still setting up servers and doing all that fun stuff?

Ray: Yeah. And so, in fact, what really got me into Node was Express, because I would go to intro to Node presentations at conferences. It was 40 minutes of setting up a web server in Node. That was cool, but having used Apache for a decade, I'm like, "I don't want to rebuild Apache." It may be slimmer and quicker and I'm doing it by hand. I just didn't want to do that, but seeing Express and seeing it kind of handle some of that boring stuff behind the scenes and really just kind of get you quicker to building routes and actually building applications, that's what really kind of nailed it for me.

Jeremy: Yeah. And so once you discovered serverless and I know, we'll talk about the JAMstack a little bit more later on, but once you kind of came to all of that, I mean, because I know, for me, building everything on ColdFusion or PHP or any other language and having to deal with the servers and the routes and all that other stuff that you have to build in, a lot of that goes away. So, did you see just an efficiency gain when you started moving to JAMstack and serverless?

Ray: Absolutely. So, I mean, Express is pretty slim. It's pretty quick. I can get a server up and running in four lines of code. I could define a route and then, on that route, say, "Do this logic." The first time I did a serverless function where the route was handled for me, URL informed handling was just passed into me. I literally, I only did the logic. I was like, "Oh, god. I'm never writing a server again."

Jeremy: Right. Exactly. So, you go from ColdFusion, which was, especially in the beginning was relatively simple. If people aren't familiar with ColdFusion, I think it still exists. I think I just looked and there's ColdFusion 2018, maybe, which is now under Adobe. It was J.J. and Jeremy Allaire who started it, I think, '99 or something like that. Then, it was eventually bought by Macromedia and then Macromedia was bought by Adobe. So, it's changed some hands.

So, you're building things on ColdFusion. It's this relatively simple thing where it's very much like PHP. You can write a lot of stuff inline, if you wanted to. You could do some application components and things like that, but mostly you're writing code inline. So, it's a really sort of simple, I build a page, I add some dynamic content to it, maybe I query a database, whatever I'm doing. Very simple, very straightforward, render the page, session management, all that stuff built in.

Along comes the JAMstack and along comes serverless functions that are, for the most part, stateless. There's this huge paradigm shift. You have to think about it in sort of a different way in which we build applications.

So, I'd love your perspective. From going from something like ColdFusion and building more traditional applications that were rendered to moving to single page apps and static sites and stateless functions, how did you sort of, I guess, calculate that or deal with that complexity in your head?

Ray: Hard. It was hard at first. On one hand, it was kind of mind opening to say, "Yes, I could have a serverless function that will do whatever logic and it's just a file and it's just deployed. It's good to go." But starting to wrap my head around, I used to build a site this way, now I can build it that way. It was definitely a process and still one I'm thinking through. I'm blogging right now about migrating a node site to serverless and JAMstack.

But, it was a bunch of things happening at once over a couple of years, although I guess, "At once at a couple of years," doesn't make sense, but when I'm this old, things kind of compress. Again, seeing browsers just get really, really good. We were both there back in the old days, so now they're amazing. So, being able to rely so much more on the clients to be able to do stuff and also realizing that many of the sites I built with ColdFusion, they were dynamic but the data was changing once a month or once every couple of weeks. I was still doing extensive calls to a database just to fetch something that never changed.

So, recognizing that there was data that was more static, even though it was in a database, so it was still essentially static and that accounted for a vast majority of what I was doing, just kind of made me realize I did not need an application server. I didn't even need Node.js. That was overkill in a lot of cases and just to start the think of what is dynamic? What could be fetched on the client, what could be generated at bill time with the static site generator? Just kind of think more about my data, I guess.

Jeremy: Yeah. No. I think that's a really interesting way to start thinking about it. I actually had Guillermo Rauch on several weeks ago now. We were talking about Vercel and Next.js and this idea of static first versus serverless first, which I think a lot of people are kind of praising this idea of serverless first, but the more you think of it and, of course, if you're looking at it from a web development perspective, static first is a very, very compelling thing because it reduces a lot of complexity, like with a blog post. What do you care about on a blog post? It's the content of the blog post. We want to see that as soon as possible. If you've got comments and maybe a little offer that loads on the side or something like that, that can all be done after the first paint, but you get the benefits of the SEO, you get the benefits of it running really quickly, you get the benefits of not having to call a database every time like you do with something like WordPress.

So, yes, so I think that's great. So, what are some of those other things so that you then did sort of to make a JAMstack site or a static site have extra dynamic capabilities in it?

Ray: So, I've been working with JAMstack for three or four years. Back when I started, it was the static site generators. Pretty quickly services began to spring up. So, I remember Formspree was one of the first ones I used where literally you would just point your form at it and then it would handle sending an email for form submission. Now, we have hundreds of stuff like that.

So, not having to do form processing, not having to send mail, we have options for database storage. There's literally nothing left for me to put on my server except for the particular logic, so if I pick something on a dropdown, the form content gets mailed to Mary, whereas, if I pick something else, it gets mailed to Sue, but everything else, the routing, the mailing, that's not handled by me at all, so I write three lines of code to handle that if clause.

Jeremy: Yeah, yeah. Now, I saw that you were using Eleventy for some stuff which I just started using it as well, so why Eleventy over something like Hugo, for example?

Ray: So, I think all of these JAMstack options, they have a different philosophy and I think some engines are a bit more forgiving, a bit more loose in terms of what they allow. I could see philosophically why you may or may not like that. You may like something that is more strict in terms of what it allows. I found Hugo at the time I was using it to be very strict. I remember I was trying to build a JSON file. A couple of years ago, that was very, very hard to do. In fact, I just gave up. Ended up outputting plain text instead, whereas I feel like Eleventy allows you to do anything you want to, whether that makes sense or not, Eleventy's not going to judge you on that. I've yet to find a place where it stopped me from doing what I want.

Now, have I pushed some very ugly code? Absolutely. But one of the awesomest things about JAMstack is I feel like I could do ugly, inefficient code during the build process because once it's live, it's a flat file and that's awesome.

Jeremy: Yeah. No. I totally agree with that and that's something, too, where I found myself, I want to say, "hacking" the Eleventy stuff, a little bit to make some things work, but then I find that it works really, really well. Even the data model that flows down. If you have data stored somewhere else but then you want to reference that data like maybe within some loop that's happening there. Of course, the fact that you can do, what is it? Support Nunjucks and all these different template formats. You can do things in markdown. It uses Frontmatter in order to do the, you can put titles and all kinds of information up at the top. You can control sorting. You can do all kind of things like that. So, I do agree, it is a little hacky, but it gets the job done. I think the most important this is what you're loading once you push that to production.

Ray: Yeah. Yeah.

Jeremy: So, great. So, all right. So, what else around JAMstack have you been doing, because I know you've used Vue quite a bit and you also have written a little bit about Netlify. So, where does Netlify fit in with what you're doing with some of the JAMstack stuff?

Ray: So, they are essentially a JAMstack provider, so they allow you to host static websites, but then they provide all kinds of really good features on top of that, so they have a form processor built in. They have analytics built in. They have serverless functions built in, so you don't have to go anywhere else to do any of those things.

So, to me, they are the gold standard in terms of where to host static sites. Not perfect and you know Vercel is also pretty darn good. You could use S3 as well, so you definitely have options, but I tend to compare everybody to Netlify as the gold standard.

Jeremy: Yeah. And are you still using Vue or you building static sites that are using server-side rendering or something like that or how does Vue fit in there with what you're building?

Ray: So, I tend to not use Vue to build sites. So, I don't use Nuxt, for example. I'm not opposed to it, but just my mental model, when I'm generating a JAMstack, if something like Eleventy or Jekyll or Hugo. What I think about the interactivity that I'm building on the client side, then, yes, I use Vue.js, but I tend to not worry about them together, if that makes sense.

So, I'm building a website. It's going to have a list of films, for example. I know I'm generating that dynamically with Eleventy. On the film page, I want this interactivity, so click for a video preview or whatever. I use Vue for that. I know that Nuxt can do that entire stack, but for me, that doesn't click right. I'm not saying it's wrong. It just doesn't click for me.

Jeremy: Yeah. I know. I'm still on the fence about static site, like where to use Vue within some of that static, because I love the reactive or I love the Virtual DOM and being able to easily manipulate stuff. If you've ever used, whether it's Redux or something like that, in order to use the state management, I forget what it's called, in Vue. But, it's very cool stuff that you can do if you maintain that state across, you can do page transitions and stuff like that.

Anyways, interesting but, I actually want to go back a second and go back to this idea of embracing sort of this new paradigm. So, again, if people know you from the ColdFusion world, you wrote a lot of books. You were an educator. You continue to educate. You even have a book on serverless. We can talk about that in a minute, but in terms of teaching developers, I mean, because again, I think teaching developers ColdFusion was, in a way, similar to teaching people a little bit of ... Because it was a little bit of maybe a mind shift there, but do you find any parallels between sort of what you've been teaching with ColdFusion and how you've been trying to teach people about serverless and JAMstack?

Ray: Maybe. I could not imagine how I would do a course now that was just welcome to the web. I wouldn't know how to do it day one like if you knew nothing. I feel very comfortable talking about the JAMstack and JavaScript, Vue.js, et cetera, but if you were beginning in web development, that's something that's actually been in my mind lately, like how would I start my children on that, but I would not know how to do that.

I can say that, for example, I think some technologies are more welcoming to people of different backgrounds. I found ColdFusion to be very welcoming to people from nontraditional backgrounds. I find Vue.js to be a bit friendlier than React, which is not to say that it's more powerful than React, but if you came to me and said, "I'm an English degree person, no ComSci background," I would absolutely lean more towards Vue.js than React.

So, I myself, I tend to lean on things that are a bit more approachable and a bit more simpler. I don't find myself to be the smartest person in the world, so that helps. Some of the easier stuff works for me as well.

Jeremy: Right. Well, what about the level of abstractions when it comes to serverless? So, if we look at something like AWS, we know we've got all these building blocks. I know that you haven't done much with Lambda and some of those other things, but if you think about SQS queues and DynamoDB and just the Lambda functions themselves and whether you're using Netlify to do that or you're using something else, you've got all these different building blocks. So, is that too complicated, you think, for some of these people getting into web development and that just this idea of using Netlify or JAMstack, does that create the right level of abstraction for people?

Ray: I think if you start with just HTML, then you're doing fine, but that's how we started in the old days. I really like Duran Duran, so I'm going to build a fan site on it. So, you go to Geocities or MySpace and it would have a small amount of HTML that they could build.

So, I think starting with the JAMstack means mostly starting with the HTML. Typically, you have a bit of interactivity with the template language involved. So, I think that would be great because if you could even do JAMstack without any dynamic aspect at all, just to kind of get your feet wet.

Jeremy: But is that something, too, where, I mean, how many people getting into web development because I know, I used to interview a lot of developers not too long ago. So, many of them didn't even know what the cloud was, nevermind getting into any of the details. Is that starting point there is, there's just so much to learn.

So, I think you're right. What is that starting point? If you were to say, "Welcome to the web 3.0," or wherever we are now, that includes cloud, that includes JAMstack, it includes serverless. I mean, again, maybe you can't answer this question but what advice do you give to somebody who's just starting out?

Ray: I think it's less about teaching the cloud and more about looking at a problem solution type aspect. So, I have a web page and I want to put the current time on it. That's a problem. And then you use that as a way to talk about JavaScript, as a way to add interactivity to a web page. When it comes to the cloud, maybe the introduction there is, "I want to put my website out there for people to see. How do I get my website on the internet?" And you talk about Netlify and other providers like that that allow you to push the HTML live.

So, to me, it's more focused on as you're learning HTML and building a website, what are the problems you encounter and then how do you solve those problems?

Jeremy: Yeah. And what about defining serverless? I know you have interesting thoughts on maybe what serverless means, but just in general, if you start talking to people about serverless, where does that put them?

Ray: So, first off, I'm kind of happy that I don't think people are complaining about the name anymore because, for a long time, that's all you heard was, "Oh, you know there are servers, still," and that was a voice that every person said it with. So, yeah, I don't necessarily think that you have to worry about that anymore. I kind of lost my train of thought, too. What was the original question?

Jeremy: No. Just, in terms of introducing people to serverless. Again, I mean, we talked about people, do you need to know the cloud? Do you need to know ... Obviously, you need to know basic HTML design principals or things like that, but just where does serverless fit in there? Is it something where we're getting to a point where it's just that the term itself might just become synonymous with the cloud?

Ray: Yeah. I think, again, if you take a problem and solution type aspect, so there's something that I want to do in this web page where JavaScript won't work for me for whatever reason. So, I can look at serverless as a way to kind of solve that problem. Netlify has done a great job and also Vercel of just making it in the same package, so you're not going to Lambda, even though they're wrapping Lambda, but you don't have to worry about that.

So, I have my folder. I put my HTML pages in there so I use the same folder to write a function. So, I would use that as an opportunity to basically introduce function as a service. I need to write some logic that runs on a server. Here's how I do it. Here's how the input comes in and here's how I return output. You teach that for Netlify. They can transfer that knowledge to Vercel. They can transfer that knowledge to other services as well, because then it's just input/output and what you name your file, essentially. But then, the idea of I couldn't do it in HTML, I couldn't do it with CSS, I couldn't do it in JavaScript on the browser. I needed a server, and this was my gateway to doing that.

Jeremy: Yeah. I think that that's actually good advice is, again, if you go down the static-first route is to think about what you can do without making something dynamic, how often something has to change, and then, like you said, sort of slowly increment and say, "Okay. Now, I need to add this bit of dynamicism or whatever and start adding those pieces in."

All right. So, I want to talk about OpenWhisk, because I know you did some work with this. I know you wrote a book about OpenWhisk. You don't use it anymore, but I'd just like to kind of get your perspective. What was that about and why aren't you using it anymore?

Ray: So, I think I first looked at Lambda and I was a bit intimidated by it and I remember it was a bit hard to use from the command line. I doubt it's hard anymore. I was working at IBM and I knew that we were participating in this open source project called OpenWhisk. I knew it was serverless, so I took a look at it. It just clicked with me. Via the command line, it was very easy to use so it's very easy to deploy stuff. I tend to gravitate to things that are easy to use and quick to play with.

I can make a new file, I can deploy it in one second. At that time, Lambda felt like there was a lot more overhead to get something up and running and to be able to test it in my browser. So, if I could play with something and I had the same thing with Vue is that I could drop a script tag on a page and play the Vue whereas, both React and Angular felt like it needed more of a build process, so I couldn't play as much but OpenWhisk, I was able to play and build a crap ton of really stupid demos, but that enabled me to learn it a lot faster.

Jeremy: Right. Yeah. No. And so, with OpenWhisk, though, you stopped using it and you moved on. So, are you still involved with that community at all?

Ray: No. I mean, I try to check in on it every now and then. I know, for my company at some point, I'm doing a blog post on using our APIs with OpenWhisk. I don't hear anybody talking about it, which I think is unfortunate, because I do feel like it had a very easy-to-use approach. I would like to see it get more attention, but I don't work at IBM anymore, so that's not my job to help promote it.

When I left IBM, I went to Auth0 where they had Webtask, another very easy to use one. That product is now canceled. I think I was about to start looking at Lambda again when Netlify added their thing. I was, "Oh, great! I don't have to worry about Lambda again."

Jeremy: Well, Lambda is great so you should definitely check out Lambda, but actually this is something that maybe draw another parallel between ColdFusion and serverless is ColdFusion was always proprietary, right from the beginning. So, you had to buy a license to it. There was no just download it and run it like you can with Node or with pretty much any other language that's out there now, and so it was very proprietary. There's been this battle in the serverless community about the idea of being or having vendor lock-in, like this idea of saying, "Oh, well, AWS owns my Lambda functions, so I'm locked into using those."

I'm wondering. I mean, I personally think one of the reasons why ColdFusion wasn't more successful was because it was proprietary and it wasn't something that just anybody could use. It wasn't something where you could get the developer edition, but there was still quite a bit of work there.

So, do you see a parallel between that? Do you think that there's going to be an adoption problem because of this idea of proprietary versus open source?

Ray: So, a couple things. There is an amazing open source engine for ColdFusion called Lucee. It's been around for a while now, has a great community. Is it as big or as popular as the commercial release? Maybe not, but it's absolutely good to use. There's been other community members and other companies like Ortus Solutions. They have CLI products where I can literally go into a folder and I can say, "Add the open source ColdFusion engine here," and then just runs.

Jeremy: Oh, wow!

Ray: So, I'm not doing that kind of old school thing of double-clicking the installer and spending a half hour to install ColdFusion. So, the open source version has done really, really well in terms of making it easier to use. Honestly, I don't think the community is growing that much. I do think a lot of people have moved on, but there is definitely an open source flavor of ColdFusion or CFML, if you will.

In terms of lock-in, I don't do a lot of orchestration. I know we're going to talk Pipedream a bit later, but this idea of I have 20 serverless functions and I need them to flow in a pattern. If I was doing that, then I'd be worried about lock-in with Lambda or somebody else. If I'm doing a couple of serverless functions, then I'm not worried. I know I can convert a Netlify serverless function to Vercel in five minutes. I look at their docs. I see, "Oh! They passed this argument instead of that argument." I rewrite a couple lines of code and then I'm good to go.

So, personally, I'm not involved with that, but, again, I tend to build a lot of small demos and a lot of toys. I'm not building an enterprise application again with those 20 functions being called in a chain.

Jeremy: Yeah. So speaking about that orchestration stuff, tell me more about Pipedream.

Ray: So, I found Pipedream six months or so ago. One of the people from the company had emailed me and asked me to check it out. I didn't get a chance to look for months and months and months. Then, when I did look, I was just blown away. I loved it.

So, Pipedream, it's built around this idea of building workflows. So, I want to read data from a Google Sheet. I want to take that data. I want to find a particular row. And then, I want to send an email for that data. Then, once the email is sent, I want to do a tweet.

So, all of those particular things, those are unique steps. They're unique bits of logic. What they have done is I've built a system where they have prebuilt all of these steps for you. You can literally, like Lego bricks, pick a couple, maybe have one step in the middle that's your custom Node.js logic that literally does like an if condition. It does real minor work, but a real example, I have a Google Sheet that has 10 rows of URLs for pictures of the moon. So, it has an image URL, it has a credit and that's it.

So, I built a Pipedream that says, "Read from the Google Sheet." So, Pipedream, they built that logic. I didn't do anything. All I did was plug in what my Google Sheet ID is. I didn't have to write any of that Node.js code.

The next step, the select random, I wrote that. That was three lines of code to pick a random index from an array. Then, when I'm done, I have a random URL and I have that random credit. I then wanted to tweet that. Well, again, Pipedream, they wrote that logic. I dropped that Lego brick in. I say, "Here's the input from the last step. This is what I want you to tweet." And I'm done. And it all started with a cron-based trigger, that, again, they wrote where I say, "Run once an hour." So, I have a Twitter bot where my coding work was literally the "pick a random row" and that's it.

Jeremy: Right. Right. Yeah. And I've seen some other companies like that. I think Paragon was one. A couple of these other companies coming up with these sort of low code solutions that are using serverless behind the scenes, because, again, it's so cheap to use a serverless function here and there and even some of that logic that you're talking about like the crons and things like that, probably using CloudWatch scheduler or something that was already written, too. So, they're just building on top of it.

But that's one of the things with these abstractions that I think we're seeing a lot of is that people are building these new layers of abstractions on top of serverless infrastructure and you've got very, very low level abstractions like the serverless application model and the serverless framework or Claudia.js or some of these other ones or even the AWS CDK that are building these, but things like Pipedream and this idea of saying these repeatable things, these things that you need to do over and over and over and over again, why would you write that more than once, right?

Ray: Yeah.

Jeremy: I mean, again, when I owned my own web development company, I had to write so many forms to process, you know what I mean?

Ray: Yeah.

Jeremy: And you're always writing these form processing things over and over and over and over again and it would just drive you nuts. I think that this is really interesting where we're finding this really powerful paradigm under the hood but then finding companies like Pipedream that are taking advantage of it and writing those powerful pieces of business logic that let you just stitch them together.

Ray: Yeah. So, I think so much of my work back in the old days was boilerplate, was setting up browsers and stuff like that. My work is getting much, much, much more smaller, but much more enjoyable as well.

Jeremy: That's a good point. So, I'd love to ask you, then, if you see a lot. I mean, you're using Pipedream. You're trying to stay away from Lambda, because, again, you're getting what you need from Netlify and things like that. So, I guess my question is and I think if you're building full-on applications, then you're going to need Lambda and you're going to need SQS and you're going to need some of that, but certainly, when you think about the future of serverless from a web development perspective, where do you see that going?

Ray: I think it'll basically be what the browser can't do. It's something where I can't trust a browser to handle this for me. I need to rely on a server that's where serverless is going to plug that hole. So, JAMstack is rendering my HTML, which we still need. People still read stuff. We have JavaScript in the browser to handle 90% of the interactivity. Then, there's certain things that need an extra layer of security. Can't be in view source and that's the last mile, I guess, the last bit will be serverless. I'll only have to go there as a last resort. I'm fine with that.

Jeremy: Yeah. So, I mean, in terms of building out some of these things, I mean, it's funny because now it's like when I think about building a new site or I think about building something, I spend most of my time planning and a lot less time actually writing code, which, like you said, I think makes things a little bit more enjoyable as opposed to just coding randomly and hoping things work out. You can actually do much more planning and structure that way.

So, all right. So, one last thing I'd kind of like to do just because when you get two old people talking, it's always and again, I am including myself in that because I was back there in the early 90s with you. The idea of setting up a site back in the day versus what we're doing now. You mentioned earlier in the podcast Perl and CGI. Well, in order for you to get Perl and CGI up and running, if you either call the hosting provider or signed up for a hosting provider that had it set up for you, but if you were setting up your own service even 10 years ago, let's just talk about that. What was that like?

Ray: So, 10 years ago was a heck of a lot easier than it was 20 years ago.

Jeremy: That's also true. Right. So maybe start 20 years ago and then we can work our way through.

Ray: So, I know, I live in a small city. There was maybe two ISPs, so we were on good terms with them as a web development shop, but you would call the ISP up. You would say, "I need a machine," and they'd get one piece of hardware just for you. Then, they would tell you the FTP address. Then you would FTP to it and you would remote desktop in, if you were lucky, to install software. This whole set-up process was before you got the machine in maybe a day or two. Now, I think it was a bit quicker for us because, again, that we knew people, but if you didn't, it was at least a couple of days. Then, you installed all the software and then you deployed your site. So, it was at least a week or so to get stuff going.

10 years ago, it was a heck of a lot easier because you could go to Amazon and get an EC2 machine and have that immediately. But, then, install your software again, but I was so much happier about that because it was instantaneous. I know at some point, they allowed you to define what was already installed and you could just clone that. I wasn't really around or I didn't really play with that much but yeah.

So, now, I tend to use Netlify for my real sites and Vercel for my quick demos, but at the command line, I typed something and 60 seconds later, it's live on the web. That's nothing compared to how it was.

Jeremy: Right. Yeah. No, I compare setting up an EC2 instance or I had a colocation facility, so I actually or I rented space in a colocation facility, so I was buying servers from Dell, racking them, installing software, doing all that kind of stuff as well. So, that's my equivalent of walking to school uphill both ways in the snow. But, anyways.

Well, listen, Ray, thank you so much for taking the time to talk to me. Seriously, I was huge fan of yours. Still a huge fan of yours. I appreciate everything that you've done, not only in the serverless community lately, but way back in the day in the ColdFusion community, because I think you're right. There were a lot of parallels between what this non-traditional background and making it easier for people to kind of get in and I think, as you've called it, democratize the web, sort of get to that point where more people can start participating and sharing ideas. I love that ColdFusion and serverless are very similar in that regards or ColdFusion and serverless/JAMstack are in there.

So, again, if people want to get in touch with you, find out more about what you're working on, how do they do that?

Ray: So, I have a blog at RaymondCamden.com and I'm on Twitter, @RaymondCamden, and my DMs are open and I like to get questions. So, absolutely reach out if you have a technical question or any question's fine. I'll be happy to try to help you.

Jeremy: Awesome. Well, thanks again. We'll get all of that into the show notes.

Ray: Yeah. Thank you for having me.

View Details

About Paul Chin, Jr.

Paul Chin Jr. works in Developer Relations at Begin, a tool that helps everyone build apps fast with cloud native services and open source tools. He previously worked as a Cloud Solutions Architect at Cloudreach, and is passionate about open source, serverless architecture, and making technology more accessible for everyone. Paul’s a lifetime resident of Virginia Beach, VA, and has a special place in his heart for Nicolas Cage, who tends to make an appearance in his presentations and projects. Feel free to talk to him about anything related to tech, food, or business.

  • Twitter: twitter.com/paulchinjr
  • 757 Dev Group: meetup.com/757dev/
  • Begin: begin.com
  • Begin Training Resources: learn.begin.com
  • Architect framework: arc.codes

Watch this episode on YouTube: https://youtu.be/L8DkEqgVEBY

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Paul Chin Jr. Hey, Paul, thanks for joining me.

Paul: Hey, Jeremy. Thank you so much. I'm really excited.

Jeremy: So you are, or you do Developer Relations at Begin. So I'd love it if you could tell the listeners a little bit about your background and what Begin does.

Paul: Sure. Well, my name is Paul Chin Jr. At Begin, we are a service to make deploying serverless applications just super simple. Really straightforward. We give the developers everything they need from local development all the way through the CI/CD pipeline, and then we put their app onto live AWS Infra. For Developer Relations, my job is to help developers understand this technology, onboard them and get good feedback from them so we can make the service even better.

Jeremy: Now, I noticed you have a sloth hiding behind you. Is that-

Paul: Oh, yeah.

Jeremy: ... something that basically says, "This is how it used to be. And now serverless is so much faster."?

Paul: Yeah. I feel like sometimes we've got to give ourselves permission to slow down a little bit. It's a reminder that says, "It's going to be okay." And even if I step away, it's going to be okay.

Jeremy: Perfect.

Paul: Yeah.

Jeremy: Well, so I'd love to talk to you about Begin. And I want to get into a whole bunch of things with that. But maybe we can begin, I'm sorry, that was a really bad joke. But maybe we can begin by just talking about your path into technology. Because I know it's a little bit non traditional. And this is one of those themes I think, that we see a lot with serverless, is a lot of people from non traditional backgrounds go into serverless. So what was your path like and then how did that lead you into serverless?

Paul: Oh, man. This question is always great and I love sharing the story because it's a good example of just where technology has taken us now. I don't have a CS degree. And I was never doing anything remotely technical. I was just always a big business nerd. And I've had several businesses in the past. And then one day I decided, "I should learn how the software stuff works. I use the internet a lot. So I should figure out how it works. Because it's going to help me, no matter what I do."

And so I decided to undertake, like, "I'm going to learn how to program in JavaScript and make web applications and do all that stuff." And so I started that and I had a wonderful local community of developers that took me in and mentored me. And from there, I just wanted to find the fastest, quickest, easiest way to build a form or to build a CRUD app. And that always led me to cloud native services. I didn't know it at the time, but not setting up my tooling and not setting up a server, all these things that I would read about, I was like," I don't want to do any of this." And now I'd find a new service. And it would just do it for me. And I was like, "Great, I'll just use this." And they always had a free tier and they were always easy to sign up, you just give them an email or a GitHub or something.

And then that's how I continued on this path of, quote, now "cloud native development" to use all the buzzwords. And it was just a very organic thing for me to continue to learn and build and put stuff out. Because none of the infrastructure was in my way. None of the arcana of what it is to be a web developer was in my way.

Jeremy: Yeah, no, I think-

Paul: So that's how I got to it.

Jeremy: Yeah, no, I think that's a common story, this idea of, you look at everything you would need to do in order to build out a full on web application and you're just like, "Yeah, you know what? Maybe I'll take up gardening," or something like that, because it just seems like an easier path. But I agree, I mean, I don't have as traditional as a background or as traditional of background as a CS degree. And some of those things I actually did a split degree and random stuff. But I ran a web development company for years and years and years. And mostly, I taught myself how to program just by reading books and using the internet and things like that. And I went down the deep path of learning how servers work and how networking works and all that kind of stuff. Which I'm glad I did. But for somebody new it just seems like if unless you're getting specifically into that business, if you just want to learn to build applications and develop in the cloud, then that traditional experience I think, is not necessary anymore.

Paul: Yeah. And it changes so fast.

Jeremy: Right.

Paul: I got started right when the MEAN stack was a thing. Everybody wanted to build the Angular, Node, Express thing. And I recently, as part of my job at Begin is to take these example apps and make them better and make them more accessible. And so I recently just took one of the first sites I ever made. And went to the Free Code Camp resource that I had used five years ago, when I first built it. And I go, "Man, none of this is actually deploying anywhere." And I was trying to remember, "Well, how did I actually deploy this thing? I must have done it on Heroku or something." But there was no instructions for it.

Jeremy: Right.

Paul: And I remember, that was a friction point. I was like, "Okay, I've gone through this tutorial. I've built this thing. It works locally, on localhost only. Now what?" And what Begin does is say every time you push to GitHub, your stuff is live. It's just automatic. It's just a given. We think that it's just table stakes now like, while you're building and prototyping, and developing, you should be able to see it on live resources. And so I've updated the Express app. I even threw the entire Express server inside of a single Lambda, and it works just as a proof of concept. But splitting them off and doing more event architecture with it is the next level. And it's just as straightforward. Because if you're pushing to GitHub, you can push to the cloud. That's it.

Jeremy: Right. Yeah. No, that's awesome. So let's talk about Begin. Because I think that's one of those things where if you think about serverless development, and you're thinking of SAM, the serverless framework, and or the Serverless Framework. So the Serverless Application Model, the Serverless Framework, obviously Architect is the backbone of Begin, or is part of that and we'll talk about that in a minute. You've got Claudia.js, you've gotten Breath, you've got all these other things that are ways to deploy. But getting that is part of a process, is part of a CI/CD process that is always repeatable, that you're not deploying from somebody's laptop every time you want to push something to production, which was very common with those tutorials that you mentioned before.

So let's talk about what Begin actually does and what it doesn't do. Because I think that's important for people to understand where there may be limitations. Because Begin is all about being serverless. So I think that could help too, just to say, "Hey, here's what you can do with Begin. And here's where that correlates to what is possible with serverless." So at a very high level, I guess, let's start there. What can you do with Begin?

Paul: So Begin uses OpenJS Architect, which is a serverless deployment framework that can package and bundle up all of your Lambdas. All of the Begin applications are comprised of small serverless functions like that. That's it. And so what we do is we associate your project inside of GitHub with a full CI/CD pipeline straight to AWS. And every time you push to your default branch and commit there, it will kick off a build in the cloud. And it will do it as many times as you push to it, and we give you the full AWS Infrastructure, we don't actually manage any of the infrastructure personally ourselves. We just make it easier for you to launch onto AWS without an AWS account.

I think that setting up AWS with keys and IAM roles and all that stuff is necessary and we have hidden it with a really nice UI, and made it so that people can push their Lambdas without having to go through a bunch of stuff and we do that as a combination of the CI/CD platform and OpenJS Architect.

Jeremy: Right. So you're building your applications with OpenJS Architect, the Architect framework, which is a really, really terse framer. I mean, it is amazing how little you have to write to make that work, which is quite amazing. And I love that about that. I spoke with Brian LeRoux, way back I think it was Episode 17 or something like that. And we were talking about the framework and Begin was just beginning at that point. But it's really interesting that it's super simple to just configure what these applications are going to be and then drop in your business logic. And then there's a whole bunch of other things that go around that. So CI/CD is one of the main pieces. But what about developing locally?

Paul: Oh, yeah. So for local development, I was a big user of the serverless trademark framework.

Jeremy: Right.

Paul: The one that is called serverless. And local development was always invoke the regular function. But now with Architect, we've emulated a good chunk of how all the AWS services work. So you can spin up what we call Sandbox, which is a local development server. And it emulates API Gateway, Lambda functions, events, queues, and DynamoDB for folks that want to use some persistent data, which we always need to get from time to time. It's not all stateless all the time, even though I wish it were.

Jeremy: Right, right. And so the other piece that I think is really interesting, and this was funny, I had a laugh with Brian about this, where I had said that the Architect framework seemed to be very opinionated in terms of how it made some decisions in terms of how to do certain things. And he thought it wasn't, but he gets how maybe I got to that conclusion. But what I think is interesting about the way Begin works, and this is not something that is so much opinionated, more so as it's, I think, just instructive and guiding of best practices. So when you use Begin out of the box, it automatically sets up a staging and a production environment for you. You limit the size of the functions to five megabytes. What's the reasoning behind all that?

Paul: So initially, when I first came to Architect, I wasn't even working for Begin, or I was just using Architect because it had an awesome local developer experience. And it was doing things with infrastructure as code in a completely different way. And I think a lot of the opinions are there just to make it smoother. And the other thing is Architect is really focused on web application. Its main job is to make sure that you can build a REST API really quickly, or you can build a static site really quickly, and get it deployed really quickly.

And I think at that point, we want to make sure that first, default happy path is going to be good enough for production level stuff. Because if you build it for production from the beginning, then there's no upscaling of your prototype. You're going to have staging, you're going to have production. You play in staging as long as you want, and when you want to go to production, click a button or release a git tag. And it's just, that's it. It's already built in. A lot of people will try out Begin for side projects and whatnot, but always the true goal is to build something that lasts and build something that is meaningful, and it's going to need certain features that will ensure that it is repeatable and stable, and has all of the checks and balances that you want when 50 people or more are working on your project. And so that's where a lot of the opinion might come in, is because of the experience of knowing where these projects should end up and baking that into the happy path from the beginning.

Jeremy: Right. Yeah, no, I love that. And that's the thing too, where with a lot of these frameworks or even some tutorials, the direction and the best practices, I don't want to say they're not there. They are baked in, but they're not quite as explicit. So oftentimes you start asking yourself, "Well, should I do this?" Or, "How should I do that?" So I do love that Begin enforces, not so enforces, but definitely encourages you to do some of those things. So-

Paul: What's funny-

Jeremy: Oh, sorry, go ahead.

Paul: Oh, yeah, I was going to say, the things that we have opinions about are really just for production support. And the thing that we don't have opinions about are the things that everybody wants to have an opinion about. I just think that it's interesting, that we don't care what front-end framework you use. And we don't really care if you want to bring another application framework on the back-end either. The only thing that we do really emphasize is that you keep your functions small so that they can load performantly. And we do care that you want your functions separated and decoupled, so that they can have independent deployments. Stuff like that. The parts that we're trying to teach people about that are important for a serverless architecture.

Jeremy: All right. So now what about the runtimes? Because I know now we can only do Node.js and I believe Deno as well. You can write it in Deno. What about, if I want to write something in Java, I don't know why I would want to, but let's say I did, or I wanted to do something in Python or something like that. Is it possible to do that within Begin or will it be possible?

Paul: It will be possible at some point. Right now we're still working on enabling extra layers. So for folks that are not aware of what a layer is, it's a way to load external dependencies to be available for your Lambda functions. Currently it is Node and we did add Deno at the beginning of the year because it's just super exciting to try to do. So we've allowed people to build functions in TypeScript, run them on Deno on this serverless architecture. And the really cool thing about serverless is we can swap out runtimes within the same application. So not available yet through Begin, but if you use Architect, you could totally mix and match your runtimes per function. So you could have an application built out with 60% Node functions, 20% Ruby functions and 20% Python functions and it'll all work together in the same framework. We just load them either directly inside the Lambda or load it as a layer to be used.

Jeremy: Awesome. All right. So now, what can you actually build with Begin? You said it was very much focused on web applications or Architect was very much focused on web applications. So what are the standard things that we can build with Begin?

Paul: Sure. Well, anything you can build with Lambda functions. So if folks are familiar with how they use Lambda functions, then that's what they can do with Begin. What I like to build are just REST APIs. Lots of just separate endpoints. I believe that Begin is the fastest way to really just get an endpoint up and running, if you want to have, I don't know, my go-to example is just a quote machine. You want an endpoint where you can just hit and it will bring back a random quote. You want an endpoint that's going to reach and do some stuff with some data inside of DynamoDB. You want to build a full application with lots of different endpoints, that's a perfect thing for Begin.

Jeremy: Right. And you can also use it for... I mean, of course API is a or REST API is often an over encompassing, I guess description. Because with those, you can build Slack apps, and you could build-

Paul: You're right, yes.

Jeremy: ... Alexa Skills and all those kind of things as well. Right.

Paul: Yep.

Jeremy: Yeah.

Paul: Anything that has that web hooky feel. You need a service to push some data to some endpoint, you can build that. Yeah. And I think that it covers a huge amount of use cases for working developers today.

Jeremy: Yeah, it's pretty cool. All right. So what about data in sessions? Because that's another thing, I don't want to say a complaint. I mean, people who don't understand, I think how to build serverless applications often think of stateful applications, and yes, we need to rehydrate state and we need state in there. What's cool about Begin, or I guess, it's part of Begin, part of Architect or how it's mixed in there, is this idea of having sessions that actually can track people from Lambda function call to Lambda function call. And I know that it uses your back-end database, or it uses DynamoDB, or the I believe, what's it called? The Begin Data service in order to do that. But can you just explain what's Begin Data, how does this session thing work? Because it's really interesting to me.

Paul: Yeah. So early on we knew that people were going to need persistent data. And I guess, one of the things that serverless does not do so well is handle socket connections to databases. And there are ways to get around it and there are ways to make it work. But we feel that DynamoDB has got everything that serverless needs for most cases. And so we integrate that pretty heavily and we expose a DynamoDB client through another open source library called Begin Data. And what Begin Data allows is for people to use Dynamo with a simplified API. So there's just get, put, delete, batch and increment decrement. Some of the basics that you're going to need. And it's new, and it's different for people coming from even something like Mongo. It's still different enough to where, I really need to do a better job of explaining it and putting out different examples.

But for the sessions, that's a basic web concept that needs to be there. And persisting it from function to function, we have middleware that can help. And it'll write the session either to Begin Data, which is a DynamoDB table, or wrap it in a JWT and then send it back and forth and capture it inside that request object and shuffle it down through all the different requests that it may need to touch. So you can do things like authenticated login, get the session, and hold that session until they're done and they log out. And so you can pass certain session IDs back and forth between the different Lambdas and it'll be fine.

Jeremy: Yeah. No, I think that is actually, it's a great feature. Because I mean, the ability to send in a auth token with every API request and then have API Gateway authenticate that and do all that work for you, that's all possible. But what I love about Begin is it just bakes that in. And again, going back to this non traditional path into technology or even people who are in technology, maybe you're a PHP developer, and you want to move over to using serverless, with Node or something like that. This idea of having that baked in session management in there, which, by the way, is just an abstraction on top of auth tokens and things like that. So it's not like it's doing anything -it's doing things that are supported, it just abstracts that away for you. So I really like that.

And I also like the Begin Data thing, because it feels very much so like Redis. You know what I mean? Where it's, it doesn't have all the same features, obviously, because Redis is more powerful from a in memory standpoint than Dynamo is. But in terms of what it gives you, just that ability to say, hey, I just need to save this a little bit of information for this user or I need to get this or I need to increment some value and having that all be powered by a serverless back-end infrastructure is pretty cool. So the fact that it abstracts all that away, I think is, I don't know, I think it's awesome. I know, and I can't find the word for it.

Paul: Yeah. Yeah. And that's why I really want people to just give it a shot and try it out and know that it is a little different. And then a big idea at Begin is this idea of a beginner's mind. It's called Begin for a really good reason. We want people who are going to try out something new, but also see what is possible when they rethink what they can build.

Jeremy: Right.

Paul: And really bring that beginner's mindset to it.

Jeremy: Yeah. So speaking of opening your mind to things, what about Deno? So I'm not sure if people are familiar with what this is. It's been around for a little while, I think it just hit 1.0 or something like that. But what is Deno? And why is Begin looking at it?

Paul: Yeah. So Deno is a new JavaScript runtime from the creators of Node. They wanted to essentially redo Node and create another execution environment that is going to be more secure, more stable, and more web native. They wanted to build an execution environment that's really going to be like the browser. So they've built in more browser APIs like Fetch Native to Deno. They have the ability to do TypeScript if you do that kind of thing. And what we find really interesting about it is again, because it's serverless, we've built a plumbing to just say, like, "Oh, I want Deno instead of Node." You just go, great, just write Deno in the configuration. And now you're executing in Deno and not executing in Node.

Jeremy: Yeah. And what are the advantages though? I mean, because I've been looking at Deno a lot and I've seen the TypeScript execution, for example, isn't quite as fast. Type checking isn't quite as fast as some of the other things. I mean, so that's maybe some of the downsides. But what do you think are some of the upsides besides just maybe the security and stuff?

Paul: I think the biggest upside for me is how it loads modules. It's just a URL. And I get that-

Jeremy: Oh, yeah.

Paul: Yeah. Like require statements, you could technically do it to be like pointing to a GitHub. You could make it a URL from end-to-end in Node. But it's not built that way. But when you are writing in Deno, it feels like you're writing for the browser more than for Node. And I think that that is a benefit, because ultimately a lot of what we write ends up in the browser and it needs to be executed there. So if we get better at executing and writing JavaScript for browsers, and it also is the same style and APIs and data modeling that we have, on the server side runtime, then it just feels a lot nicer. Being able to write a module and say, "Import this dependency from this endpoint." Well, you can control that endpoint. I could have one Begin project that is my dependencies, and I have my own mini NPM. That's just part of it.

Jeremy: Okay.

Paul: And I get to control that from the beginning. And there are trade offs, obviously, to how much you want to control your dependency tree. But you can do it. And it's not insane. You don't need a whole lot of engineering talent to build your own repository now. And you just go ahead and you do it because the underlying infrastructure is AWS, it's already globally distributed, you can make it as fault tolerant as you'd like. And just go ahead and include it in your application workflow. And I think that is super powerful. I think that not relying on NPM for your builds, is something that is going to make a big difference in a few years, as things get faster and faster and faster and more demanding.

Jeremy: Yeah. No I think from the security standpoint too, just around your NPM packages that you're using, or any package manager you use, as things change, you get that potential security issue if something comes up and somebody injects a bug somewhere or something gets in installed incorrectly. So I think that's interesting being able to control your own destiny and your own dependencies on that side. All right. So let's go back to maybe Architect a little bit and talk about this idea of Cloud Function Based Applications or CFAs.

Now this is something that I saw on the Begin site, and I started searching around for it, there's a few mentions of it, but it seems like this is something that Begin has really embraced this term, Cloud Function Based Applications or CFAs. So what is it about those that that makes them so much better, at least in your opinion, and maybe Begin's opinion, than building something on traditional servers like VMs, and instances and things like that?

Paul: Yeah. So I think I would start by saying Begin itself, is a full serverless application. So when folks are interacting with Begin, they are 100% interacting with just Lambda functions and DynamoDB. So for folks that think it's just like Git calls, you can do a whole lot more behind the scenes because functions can also give you queues, they can also give you messaging channels, they can also give you timers, with scheduled jobs. And so this is part of the architecture and infrastructure and your code base all at the same time. And that sounds really impressive. And all it means is that if you need to build an application that has a queue, you just go ahead and write the function for it. Or if you need an application that say, just needs a series of functions, go ahead and put them in a scheduled function.

And at the end of the day you look at your application and you're able to read through what it does by the functions that are present. There's really nothing else to it. You look at what we call it as the arc file the app.arc file, it's a manifest at the root of your project that defines your routes, and defines all the functions that your application may have. And it sounds like, "Oh, man, I'm going to have thousands of functions," but you really don't. When you boil your application down, you have, I think all of Begin is under 60 or something like that. It's not too many where I can just read through, I know how the applications work.

And so I think that's a really exciting architectural point for serverless applications. And it's just less for me to deal with at the end of the day. I just know how my application works because I can look at the arc file and say like, "Oh, there's nine routes here. Good to go."

Jeremy: Right. And then of course, you have all the serverless benefits in general to that, which is you're not paying when you're not using-

Paul: Oh, right.

Jeremy: ... things. And never require a patching, those sort of things as well.

Paul: It's funny, I've been doing serverless now, for so many years I forget what's the basics. But the basics blow people's minds and I'm just like, "Oh, yeah," I don't do VMs anymore. I don't have to specify anything, I just have it available.

Jeremy: Right.

Paul: And Brian and I, when we pair program, it's hilarious, because he'll just open up a GitHub repo in the browser. And then just edit the code in the GitHub browser, and then scroll to the bottom and click the green submit button and it'll kick off a new build and we'll just go look at it on staging.

Jeremy: Yeah.

Paul: That part is super magical. The speed of iterating through your prototype is just astronomical.

Jeremy: Yeah. No you make a good point about actually the things that I think blow people's mind about serverless, those being lost on people like us that have been doing it for so long. I just take for granted not having to set up a VM now. I mean, if anybody said, "Hey, I need a web hook that does this," I'd be like, "Okay, yep, I'm just going to set up a Lambda function API Gateway." Or it's like, I need some other service or whatever, but yeah, no problems write it in the Lambda function. It is amazing that we don't necessarily think about that anymore. But there are though, some things about these idea of CFAs that are going to be a little bit limiting. I don't want to say limiting, but just things that you have to think about. What are some of those things that you are limited or need to maybe rethink the way you build your applications when you're using the Cloud Function based approach?

Paul: Sure, sure. So I feel like all of us will make trade offs in our decisions, in architecture. There's always going to be, do I really need the new shiny? Or do I just use what I know? And for that huge use case, each developer has to decide if they want to put in the time to learn a new platform. Now we have tools and examples and stuff to make that easier. But the things that Cloud Functions can't do so well are long lived things. Things like obviously, the function times out after 15 minutes, if you want it to, or at least up to 15 minutes. But a lot of these limitations are going to be chipped away at over time. And I feel like AWS is going to continue making them into more and more traditional looking execution environments.

But other things that we can't do are connect super easily to, say a relational database, something that's socket driven. We don't have a good happy path for that right away, not like we do with Begin Data. We don't have any way to build container based workflows. We just don't have them. We're not going to pull from ECS. We're not going to pull container info down. And it's less about limitations of Begin and frameworks as it is more about getting back to what is going to be the most web native, really.

Jeremy: Right.

Paul: And that means that we're going to make sure that the things that the browser and 99% of web applications can do, we want to be able to support those use cases.

Jeremy: Right. Right. Yeah. And I think also, most of what people are building now, other than maybe some component that runs in the background that does something fancy, most of what people are trying to build and trying to Architect now, are these are CRUD apps. With a little bit of queuing in there, and some web hooks and some maybe calling Amazon Comprehend and getting some sentiment analysis or things like that. But for the most part, it's a lot about compute and data. Those seem to be the two biggest things.

So this is something interesting that you just said, and you mentioned this idea about trade offs. But I think that is important because serverless is a new paradigm if we want to use that word to describe it. Well, you are certainly thinking about things differently when you build an application. And if you say, "Hey, I don't have state anymore, and I need to rehydrate that state." And there's maybe a cost to rehydrating certain amounts of state in every Lambda function. Now we're talking milliseconds in order to do that. But it's still something you have to think about. The long lived executions, I can only run my Lambda function for up to 15 minutes. If I'm doing some ETL task or something like that, I have to break it up into multiple sections.

If I want to queue something, I'm no longer just saying, "Oh, I'm going to use RabbitMQ," or something like that, "I'm going to use SQS queues." And if I want to connect my Lambda function to an SQS queue, and again, I know that Begin and Architect help abstract some of this away, but I have to think about what the concurrency is going to be. How fast do I want to be processing those things. What are the redrive policies? What do I do if something fails, and it goes into a dead letter queue, then what do I do? And what's the replay? And some of those other things?

So I am of the opinion, and I would like to think others agree with me on this, but that serverless is getting more and more complex. That we took this idea of simply saying, "Hey, here's a function responding to an event." And we created an entire ecosystem around that, that made it much more complex. I like what Architect does and I like what Begin is doing with Architect in that it tries to simplify that. I like the idea of serverless components where you try to simplify the connection between the building blocks, same thing with SAR, the Serverless Application Repository. So I'd just love to get your opinion on this, do you think serverless is becoming too complex? Are we potentially going to shut out new people and new adopters if we continue to go down this road where we get very, very granular with these building blocks and just make it harder and harder to stitch all these things together?

Paul: I think that it has no choice but to become more complex over time. That's just what's going to happen. But a whole new set of tools and services like Begin are going to come out of it in order to keep that overhead low. There will be opportunities for people to create tools and processes that push away and abstract a lot of the undifferentiated heavy lifting, over and over and over again. So as complicated as AWS has become, I always make a joke, "AWS, they've got a service for everything."

Jeremy: Right.

Paul: Including satellites. If you want a satellite, you can go to AWS. I don't need satellites. I just need an endpoint behind an HTTP server.

Jeremy: Right.

Paul: And that's what Begin does. And so we've found that somebody who's making a web application doesn't need to log in to the console, and click through nine different menus to see their logs. Yes, that's complicated, because when AWS first started off, they had less links on their page.

Jeremy: Right.

Paul: Now there's 300 links of places to go to and services you probably don't need. And AWS doesn't have a strong need to simplify. And so that's why Begin exists and we're going to simplify the parts that web developers care about, and really focus that down. And if you need something outside of the AWS services that we provide you, you can try to write your other integrations around that. But really, for most use cases, you just need the eight.

Jeremy: Right.

Paul: And so it can get as more complicated as you want. But still you just need the eight.

Jeremy: Yeah. No, and I think that's a good point. I mean, I think part of the problem is, is you go to the AWS console, and you're right, and you see 250 services, or how many services it has. And you say, "All right, I want to build the CRUD app. I don't think I need the satellite." But I mean, is it something where we keep on trying to replace one level of abstraction with another level of abstraction?

Paul: Yeah. I think that's totally valid. I would say that right now the way I'm building applications is definitely simpler. It's definitely more, I guess in step with doing a thing, seeing a thing, doing a thing, seeing a thing. And in that part, I want to keep that feedback loop tight and going as best as we can, but I just don't know if there's a great way to say, "Yep, we're too complicated now. It's time to blow it away and start over again."

Jeremy: Right.

Paul: I just think we just continue focusing on the core services that you need, and making those faster and better.

Jeremy: Yeah. I wonder too, if the variety of abstractions also adds to the complexity. So even though-

Paul: Oh, yeah.

Jeremy: ... even though one thing makes it-

Paul: Oh, yeah.

Jeremy: ... easier, now you've got AWS CDK, you've got CDK TF, you've got cdk8s, you've got your Serverless Framework. You've got all these different ways that you can provision and then you just have the complexity of Cloud Formation underneath that, that it's funny where it's well, which one do I choose? Because at the end of it all, they all do the same exact thing.

Paul: Yeah, push to CloudFormation. So even Architect pushes to CloudFormation. That's how we make it repeatable, because on the AWS side, not on any of the other cloud providers, but on the AWS side, CloudFormation is the one way to have repeatable infrastructure. And so even Architect will build down to a SAM template, and then the SAM template will build down to CloudFormation. So part of the really cool foresight that Brian had when he was building this was, if you build your application with Architect, you can still deploy directly to AWS with SAM.

Jeremy: Yeah.

Paul: Because it's a hunt like SAM is downstream of us.

Jeremy: Yeah.

Paul: And so if somebody was using Begin, and then they wanted to really expand and get that satellite, if somebody was using Begin, and they're like, "Man, I need that satellite." So great. You can take your same project, send it directly to AWS and then add your satellite.

Jeremy: Right. Right.

Paul: It's fine. We let you do that. And that's a big part of how we run.

Jeremy: Yeah. Well, and the other thing I think, too, is, especially from a serverless adoption standpoint, I mean, when you're trying to pick what framework you want to use or whatever, it's a lot easier if you're building a greenfield application. So if you say, "Hey, I'm just building some new service that does XYZ," I can really choose any one of these frameworks and I can just hit the ground running with it. Whereas if you already have 20, 30, whatever applications or hundreds or thousands of applications, if you're an enterprise, then you have to try to say, "Okay, how do I maybe fit this new paradigm into the tools that I'm already using?"

Paul: Yeah. So that's always, the migration path is where the rubber meets the road. All of the tutorials out there for serverless are greenfield, like you said. And I'm trying to create some better examples of say wrapping your existing applications with serverless components. There's a very famous architecture style, what's it? It's a strangler pattern. Are you familiar with it?

Jeremy: I'm very familiar with it. Yeah.

Paul: I like to call it... For anybody else who knows about it, I want us all to start thinking about it as the hugging pattern instead. I think strangling is just too harsh for these applications that have worked for us for so hard for so long. We don't need to strangle them. What we need to do is hug them. And so I want people to embrace the hugging pattern. And if you're watching a video, I'm like hugging myself right now. Where you take your old application and you lovingly wrap it in serverless functions. And you take that application and you give it to people who are systems engineers, and who are really going to take care of it. You're going to make sure that it's stable, and that no one else is going to mess with it and that it gets to live out its life making HTTP requests until it can't anymore. Meanwhile, you've built a support network around it to handle all of this incoming messy HTTP traffic.

And one example I did of that, was I had another one of my early websites that I had built. And it was still running a PHP LAMP Stack VM somewhere on DigitalOcean. And all it was doing was catching a single form. It was a contact form. And I didn't even know how to write PHP at the time. Okay? I just copied and pasted PHP.email, and it's like 40 lines. I cannot read it. I don't know what it is. It's a lot of brackets, and all it did was send me an email when someone filled out the form. And I was like, "Man, this has been five bucks a year for the last five years. Why am I doing this?"

Jeremy: Right, right.

Paul: And so I just rewrote it. The front-end's all the same. It's jQuery. It's really, really advanced blazing jQuery. And I threw it in an S3 bucket, I put it on Begin and I built a single function that now sends me a Slack message when someone fills it out.

Jeremy: Right.

Paul: And that's better than email. Because that email account, it's the only one email account... You can combine and bring new life and really hug your old applications to give them the life that they deserve. They're working for you, non stop without complaining. You should give them a hug, not strangle them. But anyway, that's my big point for the day. It's hug your servers man. I don't find serverless as a like, "I hate my server." I find serverless as, "I love my server." I love my server so much, I want to let it go.

Jeremy: A corollary between like sending your grandparents or your parents to a nursing home. It's like that. Give them the support they need, but eventually-

Paul: Exactly.

Jeremy: ... you need other people to take on those responsibilities.

Paul: Exactly.

Jeremy: The thing you said about moving that one server that was costing you $5 a year, whatever it was, and just to capture a form, for example, that's the kind of thing where I think if you explain it to someone like that saying, like, "Hey, I had a whole computer setup or a whole server setup, just to catch a form, because I needed to catch a form. And now all that's gone away to a point where I never have to worry if, one, the page that serves up the form itself is still there, or two, whether or not the infrastructure is still going to be running to support capturing that form on the back-end." That is something that I think, personally, is quite magical. It's a really interesting thing. And I say that because you had a presentation you gave at Serverlessconf New York, and you said that serverless developers are magic makers. So I'd love it if you would just give me a your opinion, why is serverless so magical?

Paul: So in that presentation, I did a real time demo of Serverless WebSockets.

Jeremy: Right.

Paul: And that is the most magical experience I've had in technology is when I was creating persistent connections between clients and users, without having to write or stand up a specific WebSocket server. So I didn't have to write any of that logic to handle the connections at a network level. I still handled connections in my applications. I am fully in the application level responsible for everything. But at a network level, all that has been advocated. And so what is magical is the experience that developers can create for their users. I think that is magical. I think it's magical when we get past all of the technicals. And I think it's magical when humans decide, "Oh, I'm going to make this application. And it's going to help someone find their perfect 'whatever.'" And that ability and that power, that is magical for us as developers to say, "Yeah, I can create that. I can create this next thing, and it'll make somebody's life a little easier." And that is some power. And that is the magic that we should respect, for sure.

Jeremy: Totally agree. Totally agree. So another thing that I would have to say is quite magical, and I'm sure you would agree with this is when you watch Nicolas Cage in a movie, I mean, you're just transported. You've just transported.

Paul: Yeah.

Jeremy: You have a, I would say unhealthy obsession with Nicolas Cage and maybe that's just my opinion. I mean, just you incorporate him into your presentations, which I think is great because again, it gives people something to connect on. And when you see boring presentations that are just going through the bullet point after bullet point, I know I give a lot of presentations like that, I'm like, "I can't believe people watch these." But the interest in Nicolas Cage, what's that all about?

Paul: I have been so blessed and fortunate that I can continue in this career, while also serving and worshiping Nicolas Cage. So as we mentioned earlier, I'm a self taught developer. And I also don't learn things super easy. So I have to come up with good excuses for why I need to learn something, to myself. Whenever you want to learn something, I'm sure something drives you. The thing that drives me to learn is using Nicholas Cage, honestly and truly. When I want to learn something new, I think, "Okay, how can I make this ridiculous by adding Nicolas Cage?"

And so when I'm testing out, say a static website builder and I got to build a blog, I could read through documentation. Yeah, I do that. I could do their tutorial. Yeah, I could do that. Or I could build a new Nicolas Cage blog and then spend an hour or so looking for interesting articles to put into my blog. And now all of a sudden, I've got not just a blog, I've got a blog with content and it looks really cool and I learned something about Nicolas Cage, and I built it all from scratch and well, now I've done the thing. And I do that for every single thing I've done.

My very first Nicolas Cage project was a React.Component when I was like, "Man, I need to learn React. It's probably going to be a thing." And so I made one component inside of a CodePen, where every time you clicked it, it went and got a random gif and made it bigger and bigger and bigger and bigger until it took over the whole view port, and then it shrunk down again. And then it got bigger and bigger and bigger. And it was like it just did this over and over again. I was like, "Wow, look at this. It doesn't refresh. This is just updating state on its own. It's manipulating its virtual Dom thing. This is really cool." And I never built a to-do app with it. I just built these silly things.

And Nicolas Cage is there because there's nothing else on the internet that has as much source material as Nicolas Cage. I promise you, go and find another bit of source material in popular culture media that has as much wealth of information as Nicolas Cage. And the only thing I can think of were the Simpsons, and that's about it. And that might be it. The only other longest running television show on the planet.

Jeremy: Maybe Kevin Bacon. Kevin Bacon's got a lot of stuff around also.

Paul: So Kevin Bacon, but it petered out. What was the last thing you ever heard Kevin Bacon do? I'm not sure. It's been awhile.

Jeremy: That's a good point. Right. That's a good point.

Paul: But Nicolas Cage is continuously evolving. He's continuously doing bigger and crazier and badder things.

Jeremy: Right. Well, hopefully we didn't lose too many people on the Nicolas Cage thing. But I know you're a huge fan, so just quickly, favorite Nicolas Cage movies?

Paul: There's so many good movies. The current favorite is obviously National Treasure. Can't beat National Treasure at all. It's just a wonderful, wonderful movie. For the deep cut, I like Vampire's Kiss because that's where the memes start. And that's where we get a lot of his craziness. I think Ghost Rider and-

Jeremy: Oh, yeah.

Paul: ... Con Air. And his big-

Jeremy: Oh, Con Air was a good movie.

Paul: Yeah, yeah. And The Rock. Common on.

Jeremy: Oh, yeah, The Rock. Wow.

Paul: Yeah, yeah.

Jeremy: Yeah, you know what? I've got to rethink this whole Nicolas Cage thing because there are-

Paul: That's right.

Jeremy: ... some good movies in that catalog there.

Paul: That's right. That's right. Yeah. He has 100 movies. 100.

Jeremy: That sounds like a lot.

Paul: Yeah. And part of his working philosophy is that you wouldn't ask a baseball player to play less games. So why should an actor not be in as many movies as they can.

Jeremy: That's a good point.

Paul: And I was like, "Man that is amazing." Because we always think that if an actor does too many movies, it's because they need the money. But what if they just like acting?

Jeremy: Or trying to act? I'm just kidding I'm totally kidding.

Paul: Ah, there you go. There you go. You had to get it in a little bit. But that's fine, I forgive you. As a preacher of Nicolas Cage-

Jeremy: Right. Right.

Paul: ... I have it within me to forgive you for your transgressions.

Jeremy: I appreciate that. And you know what else I appreciate, Paul? I appreciate you being here and sharing your knowledge of not only Nicolas Cage but also about Begin and just serverless in general. I love this community, I love that people like you are in it and sharing these crazy things. So if people want to get in touch with you, find out more about Begin and other things you're working on, what's the best way for them to do that?

Paul: First go get and try out Begin at begin.com. Once you sign up for an account, you can build an app in 30 seconds or less by clicking two buttons. All you need is GitHub. Seriously, go try it. Second, I'm always on Twitter, Paul Chin Jr. @paulchinjr. Please, any questions about Nicolas Cage or serverless or Begin or food. Yeah, I'm a big food person too. So just send me a message. I love talking to people. And then final shout out to 757dev. They are my local developer community here in Virginia Beach in Norfolk, in Virginia. They've been incredibly, incredibly helpful for everything that I've been able to do. So I want to give a shout out to them.

Jeremy: That's awesome. All right. And I know that there's learn.begin.com as another-

Paul: Ah, yes.

Jeremy: ... training resource.

Paul: Thank you. Thank you. Yes, learn.begin.com has tutorials for building a lot of these serverless application concepts that we've been talking about. And not only will the tutorial teach you how to do the app, it'll deploy it. It'll show you where you deploy to. So your app will actually be live and not just sitting on localhost.

Jeremy: Awesome. And check out arc.codes as well if you want to see the Architect framework. All right. We will get all that into the show notes. And I think we might have to add a new segment to the show where we ask people what their favorite Nicolas Cage movies are.

Paul: Please, please do it. I think it could be a good thing. I would love it for Cage to become a patron saint of serverless.

Jeremy: All right. Well, thanks again, Paul.

Paul: Thank you, Jeremy. I really appreciate it.

View Details

About Nica Fee

Nica Fee is a Serverless Developer Advocate for New Relic. She's worked with and written about serverless for the last two and a half years. She recently spoke at Deserted Island DevOps, which you might know as the tech conference that happened in Animal Crossing. She writes regularly for The New Stack.

  • Twitter: twitter.com/ServerlessMom
  • Twitch: twitch.tv/noctnica
  • TikTok: tiktok.com/@serverless_mom

Watch this video on YouTube: https://youtu.be/yM4q0NSFz0M

About New RelicNew Relic One is an observability platform built to help engineers create more perfect software. From monoliths to serverless, you can instrument everything, then analyze, troubleshoot, and optimize your entire software stack. All from one place.

  • Website: newrelic.com/
  • LinkedIn: linkedin.com/company/new-relic-inc-
  • Twitter: twitter.com/NewRelic
  • Telemetry Data: newrelic.com/platform/telemetry-data-platform
  • Blog: blog.newrelic.com/

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today, I'm speaking with Nica Fee. Hey, Nica. Thanks for joining me.

Nica: Hey. Thanks so much for having me, Jeremy. Longtime fan, so it's great to be on.

Jeremy: Well, thank you. So you are a developer advocate at New Relic, and I'd love it if you could tell the listeners a bit about your background and what led you to New Relic and sort of what's new with New Relic.

Nica: Sure, yeah. That's great. I was actually at New Relic back in the day. I was at New Relic as a support engineer until about 2015 I believe, and left to go and become a full-time developer and full-time coder. And my path took me back sort of... As I was sort of coding full-time and just clearing queues and writing features and fixing bugs, I really started to miss some of the community building that I'd done previously. Especially actually when I was at New Relic back in the day, I was one of the people who was starting meetups and doing that kind of community building. And so I started trying to pursue that as a job which is how I got into dev advocacy. Dev advocacy, you get to tinker and you get to play and build stuff, and you also get to try to get other people excited about it and try to show it to people.

So I was doing that for Stackery, which is a serverless deployment tool, for two years, and had some success there and built some skills and really enjoyed it, and that's where I got kind of very into AWS and cloud engineering. So yeah. Now I'm back at New Relic, and it's such an interesting time to be at New Relic, and be looking at how we can go and talk to developers. Something that is interesting about being here is that everybody is talking about the time I was last here. Everybody's talking about, "Hey. There was a time when New Relic was something that lone engineers would install on one server and something would go down." They'd be like, "Well, I can see right here what the problem is." And then some exec was saying, "Hey. Let's go use this tool. It sounds great."

And right now, the question is, can we get back there? Can we get back to the place where it's a tool that developers love and that they're the ones saying, "Hey. We got to use this," rather than... As are so many developer tools being something where most people know it from the CTO coming in and being like, "We're using this tool. I met this guy on the golf course. He's told me great things about it. It's got a great spec sheet. We're using it. Everybody's going to use it now," right? So the question is, can we get back there to being in that space? So that's sort of what I'm doing New Relic because I'm trying to go talk to actual engineers about what it does and how it can help them.

Jeremy: Awesome. Well, I mean, one thing about New Relic is that they just released the New Relic One platform. I want to say the new, New Relic One platform, but it seems kind of hard to say new twice. But first of all-

Nica: We actually do that every year. We should do New Relic but then a starburst at the side. It's new this time.

Jeremy: It's even newer. Well, first of all, I want to thank New Relic because they are sponsoring this episode which is amazing, which again shows their incredible amount of support towards the community as well. So I do think that this is a great opportunity.

Nica: Can I give a quick shout out on that one, actually?

Jeremy: Absolutely.

Nica: As a dev advocate, I am actually really actively looking for stuff that is exciting in the community that we can help support. And so, obviously, you were very high on my list. I said, "Hey. We got to do this." But I don't see everyone. I don't know everything. So if you're listening to this and you have either an open source project on observability or you're doing community events or running a podcast that maybe is a little bit less famous than this guy, get in touch with me. Hit me up on Twitter. Show me your stuff. I would love to hear about it. My situation right now is I don't know enough people to support, not that I can't do that. So yeah. I want to hear from you all, if possible.

Jeremy: Oh, that's awesome. That's a great offer, and anyone listening, please take up that offer because I think it could be quite amazing. So anyway. So-

Nica: ...keep expensing GitHub sponsorships. But for the moment, that's just one that's just one like, "Well, just do it." And I'll just fight with AR after I get it done. And I'm sorry. Go ahead.

Jeremy: No, no. Not at all. I appreciate that. All right, so let's get back to New Relic for a second and this New Relic One platform. I do want to go through this because it is actually pretty cool. I mean, the entire thing has been completely rewritten. It's all new, right? Not all rewritten. I shouldn't say it that way, but-

Nica: Yeah. I had hopped into... So I came on five months ago now, and I got into New Relic. And I had kind of... I was excited. There was obviously tons of new stuff since I'd last been there five years, but I was a little confused. There was some stuff that looked the same as what I'd seen five years ago. There was other stuff that was like, "Oh, this is kind of a nice little interface." And there were things in this new interface which was sort of part of the site at the time. You could do stuff like every chart, you could go see what is the actual query that's building this chart. You could go and edit it. You could go and facet it, and you can make it more sophisticated, save it out. Oh, it was really neat. But that was only kind of part of the site. And it was like, "Hey. This doesn't feel 100% cohesive." And I'm like, "Maybe it's just me. Maybe I'm not trained or something."

But what's been happening and what's been released in the last few weeks is the whole site is the same very clean, cohesive experience now so that you can do stuff like if you're monitoring AWS Lambda and you're maybe monitoring some other service, maybe something that's even self-hosted, but their performance is implicitly connected, you can tie them together very easily. You can even just rewrite the query to connect them both directly. Or I was just writing something that's just trying to do your own kind of basic cost estimation that just applies its own rate of, "Hey. We know for Lambda how much does it cost per request," but maybe for your self-hosted stuff, that has a certain cost per usage and so, it times them together and give you a nice price dashboard just kind of out of thin air. That's pretty nice, yeah. So yeah. That's the New Relic One platform which... We're working on dark mode, but right now it has a lot of quality of life improvements for developers. So we're enjoying that.

Jeremy: Yeah. And I think that if people don't know, I mean, New Relic... I always remember New Relic as being APMs and monitoring and that sort of stuff. Obviously, the buzzword of observability is the newer sort of thing. So maybe we take a step back and just in case people don't know, what is observability? How do you define observability?

Nica: Yeah. This is a really good one. It's that someone saying, "Well, you don't want monitoring. You want observability." Or you say, "APM, application performance monitoring," and now you want observability. What's the difference? And I think about how very often, it's very useful for our dashboards or for whatever else to kind of look at one metric that covers a lot of stuff. For example, we kind of want to combine how fast we're loading, how many requests are we serving, and is anybody seeing errors, and you want that in one number. Observability is kind of an attempt to do that as a organizational goal. Observability asks, "Hey. How fast can someone looking at a problem or a question come up with an explanation and a next step?"

So the classic of course is the service is failing or flopping or we have a bunch complaints from users. Everybody's reporting a problem. We know something's wrong. If we take from that time to the time that we understand what the problem is, that's our measure of observability. So everything has, right, it has a certain amount of observability, right? I would say the only thing where you maybe have an observability score... Not a real score, but your observability is very, very poor is when stuff goes wrong and it keeps going wrong and you end up just resetting the service and then it works and you don't know why, right? You might have a very poor observability situation even if your resolution was relatively quick, right, because you have a black box inside your system and you don't understand how it works.

Now, New Relic can help with part of the observability picture, and obviously monitoring is part of the observability picture, right? How well do you go in, see how your code is performing, and send that metric back to some kind of data warehouse to say, "Hey, show me how well we're performing." That's part of the picture. Other stuff like a new interface actually effects observability, right? Because if you're struggling to click through and see what's really going on, right... If you're clicking through thousands of lines of CloudWatch logs or sitting there and trying to write a regex to sort of try and maybe find a pattern in these logs, you may have all the monitoring in the world, right? You could add a log line to every single line of your Lambda code, right?

So you have all the monitoring in the world. It's all there, but your observability is very poor because the time it takes you to actually figure out the problem is quite long. And maybe you when you find the answer, it's in great detail. It's very interesting, right? But the time it takes it high. So that's why when you pursue observability, you have to think about everything from how data is being collected, obviously how it's stored, and how available it is, but then also just how it's displayed in a way that makes sense and can operate quickly.

Jeremy: Right. I do remember when I worked a help desk, a support desk, very, very, very long time ago, my favorite solution to everything was just reboot your machine. That always works really, really well. Unfortunately, we can't do that in the cloud quite as easily, especially with really large applications.

Nica: I was doing a practice for one of the AWS certs, and I noticed one of the questions involves how you might automate that on an EC2 instance. Well, I could answer that, but I hope that's not what's happening. I hope you're not just saying like, "Well, how would you automate rebooting every 24 hours to keep this, who knows what, from affecting you?"

Jeremy: Right, right. Yeah. No, and I think its interesting too, you mentioned about making customers happy because that's one of those things for me too where it's great if you can get some alarms. And I mean, you can set up CloudWatch alarms. You can even do some interesting tracing with things like X-ray and like you said, CloudWatch insights. There's all kinds of things that can give you data, but your ability to kind of pinpoint what's wrong and act on that data quickly, that's a big thing because even if your reliability is really high and you have a lot of sort of nines up there, I think if you don't have that resiliency built in where you can keep customers happy, that's a big thing.

Nica: Yeah. This is the Charity Majors thing. Charity Majors puts it as nines don't matter if your customers aren't happy. There's a couple ways to look at that. One is right, you may have great observability. You may have great metrics for performance, but you're not really seeing errors that are happening. Another one that I see pretty frequently, and I do get to talk about another feature I genuinely like which is awesome, which is way back in the days, one of my first development jobs, I was working for an online classroom system. And a lot of our users where home schoolers or people with just two or three kids in their little online classroom. And for them, the service always performed great. And they represented, in all their usage, they represented none of the operating cost of the company.

The operating cost of the company was brought in by these people who often had 30, 40 kids in a room, but then often hundreds of kids in the same virtual classroom. And for them, the interface performed very poorly. They were a much tinier percentage of daily page loads, so our average page loads were great, but what happened... Well, the system that we had therefore created was, once you really started to love our tool and give us a lot of money, we just gave you garbage. That was when we started to treat you very badly, and that's just... All the data that that was happening was there, but until you could facet that data by for example, customer ID, or even, God forbid, by sales size, by how much money that... You have page load, sure. How much money does each page load represent in revenue, right?

And so, that's something where again, you can have great instrumentation and great metrics, and I mean, that's in... For example, that's in your CloudWatch metrics somewhere, right? You have all those parameters on every transaction in CloudWatch, but the question is how do you get to that point? The other side of course, and this is the classic New Relic thing, and my title is actually, I'm a dev advocate just for serverless, right, so I'm very focused on the serverless stuff, is beyond, okay, yes. Here's how you're performing overall for whatever customer, right? Let's facet it or let's look at various metrics, is okay, how is this code inside of that actually performing?

That is an area where again, whatever great AWS or Google Cloud or whatever built-in metrics, they're just going to tell you how long that virtual code execution environment ran, right? When did it start? When did it stop? It matters for billing. It matters for everything, but the why of it, right, why are we hitting a problem here? Why are we not performing for some users? Or maybe you're doing something really complex and you're omitting a user response halfway through your runtime, and so climbing runtime maybe isn't a problem because the second part is you're just cleaning up data or gosh, no.

So for that, that's this thing. APM, it may be a classic, but it's a classic for a reason, right, where you say, "Hey. I want to know how much time is sent..." So you'll know how much time is sent inside the express library, and how much is my own code, right? And these are questions that you want to have real instrumentation, right, like code level instrumentation, and ideally you want to not have to sit there and add a bunch of timing points and call points. Adding observability, you really hope that it's not that your whole team can't ship features for two or three weeks while they go and add a bunch of code points, right?

So yeah. New Relic of course, got famous for doing this right out of the box, and New Relic Serverless offers a very similar performance where it will go in. It'll tell you, "Hey, this thing's running really long." "Okay, why?" "It's this function call," right? "That's the one that's taking so long," or, "No. Hey, all the Lambda code's running really fast, but it's sitting here waiting for the DB to come back for quite a while." And you can see that very easily.

Jeremy: Right. So I mean, in terms of adding observability to your application. I mean, I remember back in the day as you said earlier, where you could just install something on the server and then it would just start doing all that stuff for you, right? Yeah.

Nica: Yeah. This is a really interesting area because there's some stuff when you set out to do this that just doesn't make a ton of sense, right? You're just like, "Okay." When something happens, it's maybe... Okay. You load some kind of wrapper and you wrap a serverless code, just suppose you can wrap any other code to say, "Okay. Kind of watch the function. Maybe you look for function calls, and then announce when something took a long time." Okay, announce where, right? Because you're on the server environment, right, so you don't have an agent, right? There are things that just don't translate over like a common call that I would write and explain to people a thousand times a week when I was in support was, here's how you increment a metric, right? You say, make a call to the New Relic agent and you say, "Okay. Increase whatever metric by one."

And sure, maybe there were a few instances of the agent running on different servers but they would... We could work that out, right? We could add that up. But that's not meaningful at all. There's no observer, right, to get all those little requests on the serverless environment, and if you're doing something... Someone told me recently that I use the term naively in a way that's... I don't mean it pejoratively but it's that... I just mean you're just trying something out as a prototype, right? And prototype, say, instrumentation for Lambda might be, "Okay. When you get done running, don't just end," right? You've returned whatever you returned. Go and report your data somewhere and then shutdown, right? The problem is, I mean, the gift and the curse of serverless, right, is that it charges you by the second that that thing is running, or the millisecond that that thing is running.

So if you just... Oh, I just need to make a quick little call, well, that could very easily... Well, Lambdas run for a very short period of time, so that could easily double or triple the runtime up. So then your bill for Lambda has just shot up just to get observability, and that's not a great situation. You don't generally want to see your actual service cost shoot up to do observability. So what the New Relic agent does is it creates a wrapper which is the same code instrumentation that you see with our APM style. So if you're a running a Node Lambda, you'll get the same level of code instrumentation, but instead of trying to phone home every time it runs, it writes it out to CloudWatch and then uses another agent to just snarf that up from every single one of your Lambdas. And so it's a very clean install experience and also has very, very tiny overheads so that's quite nice.

Jeremy: Right. And then the performance of other things, I mean, you clearly know this with the way serverless works is that everything is so distributed, right? So you've got SQS queues and you've got EventBridge and you've got DynamoDB and you've got all these different things happening. Somebody uploads something into an S3 bucket and that kicks off. So what type of observability do you get with New Relic around those other components?

Nica: Yeah. I mean, before we even talk about New Relic, this is such an important thing. And I'm guilty of it. I'm sure if you run the tape back, you'll hear me do it here where I say, "Oh yeah, serverless." And then I say, "Oh, yeah. AWS Lambda or maybe Google Cloud functions." And I'm like, "Those are..." I almost said those are synonyms like those are the same thing. In fact, right, serverless... AWS Lambda is not particularly new but it is still kind of new, but the oldest and best thing from AWS, AWS S3, that's also serverless, right? You don't initialize your storage server there, right? You just give it the object and expect it to figure it out, right? And so, first of all, even the term serverless applies to a whole bunch of cloud services, and then also no... I say this all the time. I say, "Who can tell me how to find your serverless functions IP address?" Okay, okay. Right. So that was unfair. "How do you find a TRL, right?" You just want to go and get it, some of it. It's like, "Well, okay. It doesn't have those things, right?"

To even do that, you need to create an API gateway to have a connection even to your Lambda. So no Lambda exists in isolation, right? And of course, there's going to be at least, even for anything beyond Hello World, even for the to-do list app, right, you're going to need a gateway, probably some kind of file storage, and then you're going to need some kind of database probably, right? So one of the other things I hear all the time when I talk to developers who are working with serverless is that maybe they understand how their code is working, but when they go and hit their API gateway and get a 500 back, they don't know where the problem is, right? Is it permissions between components? Is it the Lambda code? Is it the databases that's returning something wrong? It's just not obvious, right? And so AWS is sort of... Their in-house solution to that is X-ray which is an effort to say, "Hey. Let's see what one request did all the way end to end, right, to give you that insight into saying, "Hey, let's see how this started and ended and what services it hit between.""

So you might even have some surprises there that you say, "Hey. I didn't realize this is relying on this other maybe queuing system or this Lambda always calls this other one." Well, there's a problem with what I just said always, right, where X-ray is just a little sampling of, "Hey. This is what one of these did." And very often, you'll been in the situation with X-ray where you'll say, "Hey. All my X-ray traces are pointing to the same problem but it's tracing only when something's going wrong." Or it's tracing in only certain situations so other stuff you're not seeing.

At New Relic, first of all, we do instrumentation on every single invocation. It's not sampled. You actually see every single invocation and what those code spans where inside of each invocation. And then we also integrate with X-ray data. So we pull X-ray data into our distributed tracing system to help you look at it in a unified place, and obviously, those are going to be sampled. We're not going to send you every single span for every single item because that would a very wild amount of data, but it is going to give you a really even sampling that shows you really broadly across your stack what's going on inside of those stacks. And then you can sit there and see something we were talking about like, "Hey. This much DB time," or, "These are the services that were called."

Jeremy: Yeah. And I think that's interesting.

Nica: So yeah. Some... Sorry, go ahead.

Jeremy: No. I was just going to say, I think that's interesting too where you're sampling everything and I don't know if you can say sampling everything. You're just recording everything and part of the idea behind that, and this is something I talk with actually Erica Windisch about a couple of weeks ago on this show. We were talking about how it's sort of really interesting where it's almost like you want to be able to see when your application is behaving correctly. That's some of what you want to be able to see because then you can do things like performance tweaking and you can say, "Okay. The service is running just fine, but this Lambda function's taking 600 milliseconds to execute. Why is that?" And you can dig down and you can do some optimizations.

Nica: Some of the best leadership I ever got was I had an engineering manager. I'm blanking on her name. Embarrassing. But anyway, I'll remember it and shout it in the middle of the... later I promise, but she said, "Were hunting down these errors but let's just sit down and look at our total number of requests and the number of errors here." And if we're erroring at one kind of request every time then yeah, we have a systemic problem. But I think maybe it's a time out or something else. And if it's one half of 1%, maybe we should be looking at when things go right and seeing how we can improve that experience.

Jeremy: Right.

Nica: In that case, we have some little user response section where they're supposed to put in a percentage, and everybody was taking a minute or something to get through that. And so we really had a user experience problem, right, for all users that we wanted to look at that was much more important than once in a while when people put unescaped SQL into their username that it would error. Okay, great. But that really wasn't... That wasn't the problem that most users were having. So yeah, you want to capture data when things are going right. That's a very smart thing to sort of keep in mind.

Jeremy: Awesome.

Nica: And obviously that's all doable in the serverless world. Though there are these situations, right, and this is the thing where distributed tracing becomes a big issue where in some serverless tools, right, you can write some code and you can go in and get real insight. And then in Lambdas, you can even use layers to grab big, large code packages and say, "I want to use this sort of outside my code." In others, you can do some configuration to say, "Hey. Please log this over here. I'm trying to watch those logs." But in others like queuing services you can't do any of that, right?

So if that queuing service experienced a scenario like... Right? Where does that go? Or especially, hey, when things are going right, how long does it take you to de-queue certain stuff? Well, there's no endpoint to say, "Hey. Queuing service, I want you to tell me about this." So that's why stuff like X-ray integration is super key because you have to figure out... You have to get insight into those things that by design, don't allow you to do any kind of that. There's no custom code you can run around the simple queuing service.

Jeremy: Right. Yeah. I mean, I like the idea too of trying to connect things automatically with tracing headers and correlation IDs and some of that stuff that you do not want to have to try to deal with yourself. And I know it's not perfect yet. I know it's getting there, but-

Nica: Yeah. We're actually, we're still sweating the AWS people to be like, "We want to implement some more open standards for these kind of tracing headers because we're trying to get to the point where it's a little bit easier." It's such common request and of course I understand it to say, "I need to connect all this stuff. I don't just want to be sitting here looking at here's how all my DB calls performed and then over here is how all this queuing stuff performed. I want to be able to see these together."

Jeremy: Yeah. Yeah, that's awesome. All right, so let's talk about the New Relic One platform for a minute because I do think there are some really interesting things in here and I'll read off the website right now. But essentially, it's one platform, three products. And I think that's interesting because... And we're going to get into a little more about this, but maybe we can just talk about each one of these things. So the first one is the telemetry data platform. So what's that all about?

Nica: Yeah. So we haven't gotten to talking much about open source which is probably a focus for me, and is certainly a focus for the entire company. And one of the things that we're seeing, especially since 2015 is that, you've seen a lot of great open tools to do some of the instrumentation that you need. Now, I'll be honest, in my experience, right, if you install a New Relic APM agent, you're going to get really detailed information about function calls, their names, their designations, right? But there are some open tools that can do similar or close to that performance, but also there's open instrumentation for stuff that we of course, never got to writing instrumentation for, right? So open instrumentation is a huge, huge component. There was a tool initially called OpenCensus for PHP but it's now I believe called OpenTelemetry. And you get great results with that.

Now, where does that data go, right? It's great to have open tools for instrumentation, but if you're then saying, "Okay. Well, now we've got to stand up a database and now we've got to standup a data front end to show people what's in that database, right?" You said, "Oh, we used open tools, but we just put ourselves in a difficult situation." So the Telemetry data platform is an attempt to be this omnivore for that performance data, and have a place where those open source tools have a home where they can send data and display it in a really clean and useful way, so that your maybe sales enablement people or other people who have coding skills and they want to write a SQL query to show you the data, but they don't want to sit there and configure your database themselves. They don't want to handle database permissions. They just want to write a few lines of SQL and get a cool chart, right? So the Telemetry data platform is a place that you can send that data and pull it out in a really effective way. And hopefully that helps your observability.

Jeremy: I would hope so because I don't think anybody wants to be setting up databases just to store telemetry data, especially-

Nica: Yeah. I mean, that's the thing is you have a hard enough time running your own databases for customer data, right? You sort of get to this meta point where you're like, "Yeah. I don't want to be setting up services to observe the services that observe the services," right?

Jeremy: Right, yeah. Well, that's the other thing, right? Now, you're going to observe your database platform in order to make sure that that's still up and running which... I mean, if you think about just the promise of serverless or the idea of serverless in general, it's hand off that undifferentiated heavy lifting. Collecting telemetry data is probably something you don't want to try to manage yourself.

Nica: Yeah, exactly. Yeah. And that is actually something I love about... I landed at working at Stackery and I loved seeing it of course as a value of New Relic which is this whole serverless ethos, right, is you're supposed to be focusing on business problems, right? And if you're sitting there and saying, "I got to learn this config value because one of my Kubernetes clusters failed. I got to learn this because again..." It's like, "Well, okay. How did this help the customer?" It's like, "Well, the service is back up so I suppose that's good," right? But the idea is you're supposed to be saying, "Hey, I don't think we're going to differentiate on becoming a platform company," right?

And I think again, saying, "Hey. We're the best at measuring our own performance, our own service performance. We have people here who are great at engineering an observability platform," it's unlikely that that's what's going to differentiate you if you want to be selling shoes online, right? So New Relic can handle a lot of that heavy lifting, right, and present an incredibly clean and incredibly cheap, in my opinion... I'm not a sales person and I'm not deep on these sales numbers but we can present something very, very inexpensive to store and retrieve that data.

Then the second piece is full stack observability, and that's very much like... That's the stuff that... It's what you sort of know and love about New Relic, but very often, I will... The interviewer will say, "Hey. I'm dev advocate in serverless for New Relic," and people will sort of be like, "What do you mean? Doesn't New Relic just do APM?" And it's like, "Well, we still do and we're the best at that, but also, yeah. We'll observe your serverless stack super good." So this is the stuff that we're very familiar with, right, is that you get this really deep insight into what you're doing. It's kind of what we've been talking about. So maybe there's less to say about that piece.

But then the last piece is AI stuff. And when I try to explain internal to New Relic why this is important or why we should take the time to document this or that or talk about it, I say, "I've been doing..." So when did I come on at New Relic? It was like 2012. People used to say to me in 2012, they said, "You have all our data. Why do I have to set up the alerts? When I have something that sees 10,000 requests a minute and yesterday it saw 6 all day, why can't you just email me?" And no. And I heard about it in 2012. I'm old. I heard about it the year after, the year after, and then on a call with a customer, who was a very advanced customer of ours using a lot of data features, I heard mention or say it again. They said, "You have all our data. Why can't you see, "Hey..." Why couldn't you maybe even message us when errors are normally at 10% for this service because maybe user behavior creates an error, and now there's suddenly 0%. Why couldn't you email us about that because it's just so unusual."

And another engineer at that company on the call said, "Yeah. We actually have that. Let's go look at the Slack channel." And the Slack channel was just, "Hey, this error rate dropped. I mean, this throughput dropped unusually low today." And you click through and you can go see a New Relic chart. That's pretty cool, right? That has real promise. And there are of course, as with any ML system or any linear algebra system, there are many times when it presents you things that maybe you don't care about, but just like with any well engineered system, you can go back and say, "Hey. I want to see less of these. I want to see more of that." Obviously there's a huge place for manual alerting. It's something I talk about all the time. But yeah, it can be very, very powerful.

Jeremy: Yeah. And I think the idea of even simple anomaly detection, right? When you have data that's collected over time and you can see your average error rate or your average throughput or whatever. And then also not sort of the cheap anomaly detection where you say, "Oh well, it averages this." Well, averages are great, but only for certain periods of time. Maybe in the morning it's higher. Maybe in the afternoon it's lower. Maybe we got a spike at lunchtime or whatever it is. Or-

Nica: If you're selling delivery food and you're getting just a sort of simple average getting three times a day, right, that says, "Oh my God," or at least twice a day. I don't know. Some people get delivery breakfast.

Jeremy: I don't know, maybe.

Nica: Let's not talk about that now, but yeah. At least lunch or dinner. It can't be a simple thing, right? It really needs to be at least, a second order system that can say, "Yeah, you..." For example, hopefully right, "Hey. You normally see a spike at lunchtime." And maybe you can go and say, "Oh well. It's Christmas day, so okay." It's Thanksgiving so it makes sense, not that people are going to order a pizza right now. But, yeah. You want that kind of at least a second order system that says, "Hey. It's not just the average. It's not just that you're breaking the average, but that something does seem off here."

Jeremy: Yeah. And I think the promise of AI and machine learning and all this kind of stuff, it's kind of funny because I think we're finally starting to see people implementing real AI/ML use cases. I think when you... In 2012 because I am also old, I remember every pitch deck having, "Oh, we do ML and AI," with no idea what that even meant. But I think if we go back to the conversation we were having earlier about observing your application when it is working, that this is the kind of thing where AI can really help because if you're not getting any errors but you are just seeing a huge slowdown for your lunchtime order spike, then there is a good reason to potentially go and look at that. There could be a reason why that's slowing down, right? I mean, especially what if all of a sudden your traffic dropped off but everything seems to be working correctly, that gives you insights where you can go and start investigating those other things. So it's not just about errors. It's also about just fluctuations in the normal operation of things.

Nica: Yeah, and that's actually... It's something I talk about a lot that I would argue... This is maybe a little extreme to actually implement but that everything you're setting a high alert for, you want to think about, would it be meaningful to set a low alert for, right? Now, it might be... And I think that some of the things that seem would be obvious knows like total response time, I think that might make sense. Maybe it's a very, very low threshold, but you'll say, "Hey. We're reporting that your total runtime for your Lambdas is 0.01 milliseconds." Something is wrong at that point, right? You know that something is wrong. So obviously high and low throughput are classics, but another one that covers a lot of these is actually low cost. If your cost just suddenly drops by 30, 40%, something's probably... Unless you really did just push out a big release, something's probably up. And so that's something that I think is really interesting as far as you really can...

When you're thinking about a crisis that is something that is not what you have predicted, something like, yeah... We talked about Charity Majors before but something I really thought about a lot that's stuck with me is how very often we create dashboards for problems that we've agreed we're not going to fix. And that's something like you have a huge system. It leaks memory sometimes, and you really just need to watch memory usage and reset the thing. And I don't think anyone would disagree with that. I said, yeah, this is a dashboard to monitor a problem that we're not going to fix, or whatever.

You've got a steam locomotive. It gets hot, right? You're not sitting there trying to make a cold steam locomotive. You're just saying, "Hey, yeah. It gets hot so we need to keep an eye on that," right? But then those are all the problems that you know about, right? So hopefully anomaly detection and some other observability tools can get you to a point where you get at least a clue, right? It's not that you're getting a text that says, "Hey. Steve just deployed some code and it used an incorrect sorting algorithm and that's not what you're going to get on your phone, right? You're just going to get an email that says, "Hey. I actually need to start looking into this and see that there might be a problem here."

Jeremy: Yeah. And you know what the other funny thing is too is that we've been talking about these metrics and we said how long a function runs for and things like maybe, I don't know, errors and things like that. And we're talking a lot about application sort of level things, right? I mean, there are certainly infrastructure components underneath that, but that's another great thing about serverless too is you can start focusing on a different set of metrics which are not, is my server running? It's where are my performance issues. And I think that's just another thing that's really great about what you can do with observability.

Nica: Yeah. Something I love diving into is I'll just take people through... Maybe they've done some of the New Relic instrumentation on parts of their cloud stack and I'll say, "Hey. Let's look at some of the parameters that you're gathering for every single indication that right now we're not doing anything with." They're the event parameters that are just available within AWS, and we'll step through them and there will be an awful lot there, right? Obviously there's stuff like the event source, where did this come from? Did the API gateway call this thing? Was it an event from some place else, right? And you can see stuff like stuff that would make my heart stop where it's sometimes this function is called by API gateway. Sometimes that's just being triggered by DynamoDB...

Jeremy: Right. That's not good.

Nica: You see like cans of soup and balls of cotton on the conveyor belt, and you're like, "I don't think... This doesn't seem right." But you can really... Often there's so much available on each of those events. They're very rich data objects that you can start looking into actual business logic where you can say, "Hey. Let's look at how one organization or customer or one sort of use type." Like, "Hey. This is a person making some kind of update request," right? And a lot of that stuff is available on the front end often, right, in front end monitoring, but I like to see that coming in more and more on the back end. So instead of just looking at kind of... I often sort of when I'm thinking, I think of it as these engine metaphors where you're sort of seeing, ah, the engines hot or the engines cold. How much memory are we using? Physically, how hot is the CPU, right?

It gets you to the point of saying, "No, this service is very critical for people updating their accounts, and see how that's taking longer or that's performing differently. And let's look at what that might mean," right? Something I think about is looking at the weight of DB information that's coming back, right, because one of my side things is looking at GraphQL and trying to encourage people to do these fully parametrized queries, right? It's, "Hey. Look at how this kind of request we're always sending half a megabyte back every time someone tries to do this one thing." That can be very, very insightful and again, that's much more in the business logic world than it is the world of sort of yeah, as you say, looking how the server is doing. Is the server up or down?

Jeremy: Right. Exactly, exactly. So you mentioned cost in there too, and of course, when you're using on demand or pay-per-use services, especially in the serverless world... I mean, even in cloud in general, cost is one of those sort of first class metrics. And we can talk more-

Nica: Yeah. It's kind of fascinating. Sorry, but I'm trying to get a cert right now and looking at... There's whole classes of AWS stuff that exists because some of the tools you're using only do host based pricing, and you're like, "Oh, you have..." I mean, not to the point of well, I just can't do that, right? You can't go on this cores. But with EC2, people are like, "Oh, I need to own certain cores because that's exactly how I pay. I pay by core ID." It's like, "Wow. That sounds pretty old," right? If I'm going to do... If I'm paying for my Lambdas by how many requests I get and the billing is scaling smoothly, shouldn't that happen for everything that works with it, right? So it's been nice to see that. Again, I'll connect you with great sales enablement people who will tell you all about the exact cost structure, but it is nice to see, "Hey, we're doing usage based pricing," which is very helpful.

For me, the part that I'm super passionate about is I love going and talking to bootcampers and meeting people. I do a Twitch stream that's just for people who are totally new with AWS. It's like, "Hey. Let's get your first web app on AWS. Let's get you to Hello world." And there's just such a weird thing about observability. Observability is this buzzword, very much like test driven development was or object-oriented programming or anything that's like, "Hey. This is good. You do this, it's good." But if someone was trying to pitch test driven development and they said, "Well, what's step one?" Well, step one, you sign an $800 a month contract, right? Step one needs at least $20,000 in sales you have to sign up for. And most of the tools for observability, they're not cheap.

And so this thing, it's called the perpetual free tier, and again, I'm not going to break it down into gigabytes and MIPS and MOPS, but I will say, if you're running a little hobbyist app or for me, I'll be setting up eCommerce apps for people and I just kind of want to set it and not really think about it, that perpetual free tier, that will just carry you. That will gather plenty of data, plenty of usages, plenty of requests. If you have a few hundred or a thousand users, you can use that free tier forever. Hop into New Relic and see how this service is performing. So that's so nice for me because when I started 5 months ago, it was like, "Well, I could sign a bunch of bootcampers up to the free trial on that and in two and a half weeks, they're just going to be out of luck." And so, that's been really nice.

I think there's some real power in that in the idea that just like testing, it's not that necessarily every single person's going to do it, but it's much more about, if you take the time to do it, there is a service available that is affordable, right? Or when I started within web development, a lot of people were... They'd gone pretty far, but they were doing all their hosting on their laptop because they could not go out and just buy hosting, right? So that's what New Relic is trying to do with observability is say, "It either costs you nothing or it costs something that's just very, very negligible on top of launching your business or launching your web work."

Jeremy: Right, yeah. And I mean, and the free tier is... I mean, I don't want to get into the numbers, but you're right. The free tier is very generous. There's quite a bit that you can do in that free tier, and you couple the free tier with AWS and you could probably run a good size application for quite some time before you start getting hit with charges.

Nica: Yeah. It was something... It was in the millions of transactions you can be measuring and you're still on the free tier. I was-

Jeremy: Yeah. 100 million app transactions per month and 100 gigabytes of data transfer per month which is pretty big.

Nica: Yeah. Yeah. So you're right. Of course, I love talking to small teams or agencies and stuff, and agencies is a big one where I think about where it's you would like to set up some observability tools on there so that when the client calls you six months later and says, "Hey. I'm having some problem." You're not having to say, "Well, we got to start billing to even try and figure out what's up," right? Now you just got to click through to your dashboard, but then I also don't want to be bugging them about a $25 bill that has to be paid so that we can keep up the observability. So yeah, that's pretty powerful. That I think is... It really opens up who I get to talk to which is fun because that means I can go and talk about Arduino stuff or talk about goofy CLI stuff instead of having to have these enterprise conversations.

Jeremy: Right. And trying to sell people on that stuff too. Yeah. So I mean, I think what's really great is again, not only do you get the free tier, after that again, it is usage based pricing. And one of the things I love about that because just, I remember I started a startup back in 2010 at one point, and we were building facial recognition as a part of what we did. And I had to go and buy a software that then I had to write a PHP shared object for, the shared object module, and write that so that we could tap into that with a PHP call in order to run facial recognition.

I think it was $5,000 just to buy the software for that, plus all the engineering time, things like that. This is what I love about serverless and this idea of usage based pricing where you just say, "Hey. I need to run facial recognition. I can hit AWS recognition servers or something like that. I need to translate a document or I need to do that." I just hit this one thing and it costs me a few pennies here and there. Extending that idea to something like observability I think is amazing. It's very useful for those small teams.

Nica: And it's so much this thing. Often when people ask me to define serverless, I say it's a goal. It's a goal like Agile, right? You don't buy Agile in a box, right? You don't say, "Oh well, because we're all clicking this Kanban board 20 times a day, now we're agile, right?" And actually, one of the ways that it really is related is here, right? The ability to say... Slack's the classic example. It's like, "Hey, we have something here and we think it could really be big," right? Well, that's great but before you get that huge interest and have those huge sales and have that huge growth, how can you make something that still performs well, but does something sophisticated, that lets you just say, "Okay. This one was successful and these 12 were not," right? And let's you just scale with the success of that product, right? And serverless is so much about that.

And I tell people all the time to say, "Hey. Just write this microservice serverlessly," because very often, you're an engineer. You have a good idea. You don't want to start with having a conversation about how you need to pay this extra AWS bill or you need to do this extra thing. And it's quite wild. You can see people who make... They're taking home 12 grand a month and they're having a conversation about an $80 a month bill that's taking 15 emails to justify why. And very often I see teams where the real message they take back after that is don't experiment, right? Me asking all these questions, right, you know not to experiment.

Now you can say, "Hey. You start this out. It's free or it's very inexpensive," and then you say, "Oh hey. We got a big bill we got to pay because we're taking off," right? Tons of use. People love this area of the site, right? Maybe you want to add facial recognition. Image recognition is a good one, or say, you want to add image editing. You want to add video uploading, and you just don't know how big it's going to be, right? S3 and Lambdas a very powerful way to do that, right, using something like using serverless based video pre-processing and then storing in S3. If nobody uses it, you don't pay very much for that hosting.

Jeremy: Right. Right. Yeah. No, I always suggest that too where I say, "Look. Start with serverless especially from the prototyping phase and if something gets so amazingly big that for some reason, you can't optimize it anymore with serverless and you have to go down, as you said, the Kubernetes cluster, a path or something like that, then that's great. But you don't need to do that when you've got 10 users. You need to do that maybe when you have 10,000 users or more."

Nica: And a big part of that is what expertise are you building on your team?

Jeremy: Exactly.

Nica: When you're using any observability tool... And much like deployment tools, I often say like, "Hey. Go use an observability tool," right? "Go use Gravada," right? That's a great open source toolkit, right? Just do something because what you want to build in your team is expertise at building your product and observability should give you insight into your product. You really shouldn't be building expertise in other stuff, right? You could say, "Hey. I'm becoming an expert in running a metric server in doing my CICD config." Well, some of that stuff is maybe necessary in certain use cases, but ideally, right, you're becoming an expert in your actual product that you're giving your users what they want, right?

Jeremy: Yeah, exactly.

Nica: I think about how every time I sort of struggle with the time zones and time spans, I think how, oh boy, the ladies at Airbnb must be so good at this by this point. They've probably got a team that's like, "Yep. How many fortnights between Memorial Day and Labor Day weekend?" or whatever. They got that. They got all that down.

Jeremy: Totally. Totally agree. All right. So let's talk about open source for a second. So you mentioned open source a little bit in the beginning, things like OpenTelemetry and some of those other services. But New Relic has gone all open source on all their agents.

Nica: Yeah. This has been very exciting. Yeah. Yeah, so this is something that I think was maybe overdue, it's probably overdue everywhere, is that if you're using this tool to get insight into your own code, it seems nice if you could actually look into what its doing. And then also of course, any instrumentation package, even New Relics great instrumentation packages for all these different language web apps, you're going to want to extend it, right? And most of the conversations I have when I talk to customers is about extending it in some way. And so through open sourcing our agents, we've opened a lot of that up to say, "You can take a look at this logic. You can look at where it can be extended as is appropriate for your tool set."

And then a big piece of that is we're doing a ton of contributions to the OpenTelemetry project, previously OpenCensus, which I can remember if I slipped in calling it OpenCensus at the start, but yeah. Tools are very important. Again, New Relic does great out-of-the-box instrumentation but of course, the open source community is going to build instrumentation for stuff that isn't even on our radar, right? There's a new web framework every week, right? So if someone's going to sit down and write some great Deno instrumentation, it would be great if they weren't doing that for just their own shop, right, if that was something that was shared everywhere, right? So yeah. So that's a big push and I want to plug again, if you have an open source project and it contributes to observability and has a code of conduct and it's repo, get in touch with me because I would love for us to backing that and helping build that.

So the other big piece of that, and I mentioned a little bit when we talked about Telemetry Data platform which is the name of the product is if you're using open source tool kit to do observation, we are going to be able to consume and display that data in a very, very powerful way. That part is not just this week. That part has been going on for a few months is, or sorry a couple years, is we've had really great end points to take that data in, and again, on this free tier, we can take and display a ton of that data without it really costing you anything to do that. And even once you grow beyond that, it's not expensive. So that's a very powerful set of tools to say, "Hey. Maybe there is an open project that gets you a lot of the data you need. Let us display that for you right in with the other stuff that we're instrumenting."

Jeremy: Right, yeah. And so besides just making the agents open source, which I think you're right, I think is really exciting, you're also contributing quite a bit to open source. I think you're the third biggest contributor to OpenTelemetry I think, right?

Nica: Yeah. I saw that at the chat the other day. That's really neat. There's some engineers who have gone so deep on how to truly instrument certain behaviors and do full instrumentation on some very big and complex applications that it's really great to see those contributions becoming more open source and seeing that stuff happen. That's been really cool. And there's also, there's some neat stuff happening with what we call programmability. This isn't my baby but there's a really smart guy, Jemiah on the team who... He runs New Relic Nerd Days which is coming up in October that we're all excited about, but where people can also create whole modules inside of New Relic.

So let's say you want to show our error rate or something but you want to do it in a fun way. You want to show it was a ring toss game, or you want to see an elephant that grows bigger and bigger for how many cars you sell or what have you. I'm just thinking of visual ones but whatever. But we have an open source tool kit to do that, so that you can actually build data components. So that's after all your data is already in New Relic and you're sort of in the New Relic architecture. So it's not a great first project but it is just fun to think about having them in your future.

Jeremy: Awesome. All right, so we've been talking a lot about the New Relic One platform and we've talked quite a bit about serverless too which is I think most interesting to the people listening to this podcast. So what are some of those features and some of the things that you can do with New Relic when you plug it into your serverless applications?

Nica: Oh, yeah. Yeah. That's great to talk about. That's where I start. So the first pieces again, is you're going to do a no code commit deploy to do your instrumentations. So you're not going to have to edit your code. You're not going to have to add... I say this with some instrumentation... Oh, just add a few lines of code at the top. Nope, not necessary. We'll do a tool layer and we're even working on better tools for that in the near future to deploy it in an even smoother way. But then what you're going to get is again, you're going to get that code level instrumentation on every single invocation.

So you're going to be able to see, at least for every single invocation like for example, how much time was spent in library code, how much time was spent in your own code, how much time was spent waiting for a database or another server to come back? So that's significant. And then with our distributed tracing, you're going to be able to zoom in to a large number of your transactions and see exactly which functions were taking so long and what were they waiting for, right? Did you have maybe API calls going out from that? You get this nice, I don't know what they call... It's called a waterfall chart or something?

Jeremy: Yeah, something like that.

Nica: Where you can see, "Hey. Was this maybe happening synchronously, so it was really... It was made asynchronously, but that you were holding up other requests to wait on it." And so, you can see that real detail. There's other stuff too which is just kind of quality of life stuff but it really matters is you can see cold starts and you can see memory usage and the memory cap on your Lambdas. So very often, that's critical. Sometimes it will reveal a problem. I actually just was talking to somebody who sure enough, they had tons of cold starts and so they really did have to think about how they were going to handle that.

But also it helps you eliminate that as cause. Hey, you're seeing this request time climb up. Do you need to dig into the tracing and the logging or can you just say, "Look. It's cold starts," right? So it's nice to be able to eliminate that. Using 50% of the memory, you're probably okay for CPU and IO as well. So now we can move on to the code performance. And then the last piece is because we're using this kind of CloudWatch step, you can actually have a Lambda sitting in your service that's grabbing that out of CloudWatch grabbing that logging and sending it up to New Relic, which means we can actually grab more logs if you want. You can define a pattern and grab really extensive logs and send them up. And so, New Relic logs is another way to connect those traces over and again, see really detailed performance information.

Jeremy: All right. And you can actually add additional things. If you wanted to capture specific business KPIs and things like that, you can alter your code and add some of that stuff in there, right?

Nica: Yeah. Yeah. So you can absolutely add stuff as a custom value. You also, again because you get all the event parameters that AWS is sending around, often that stuff is already in there and we have this really clean data explorer where... That's a big stumbling block I've noticed for myself as well. I'll say, "Oh, well. Let me just log out this whole event." And maybe it'll... Okay. It'll log out but it wasn't a complete object so there'll be a few ones. I'm like, "Okay. I'm sure there's others. That's fine. I've got enough detail." But just having a little explorer.

We have this thing called the Data Explorer where you can click through and be like, "What parameters were available on this event and are they on every other event," right? Does every event have a customer ID or is it only some? You can just see that, and that just makes it so much easier to figure out what you might be charting or what you might need to do a code change to see. I can see that saves so many people so much time, and so it's always very nice to show off. And we can do that because we're doing that level of instrumentation on every single transaction, you'll see it for every single invocation in Lambda, so you'll get a really nice consistent smooth data graph on that.

Jeremy: Right. And then you get the benefit of the entire New Relic One platform. So you get the AI monitoring there and the ability to do those alerts, but you can also set custom alerts if you wanted to as well.

Nica: Mm-hmm (affirmative). So you could see for example... Because you can do this right at the query level, so we use a SQL syntax to make all these charts, you can write a query that says, "Hey. Is my container running out of memory or are my Lambdas running out of memory?" And combine those together and get a unified alert for that. Now, that one didn't make a ton of sense, but let's say maybe you're using some kind of EC2 instance to handle some requests and you're using a Lambda to handle other requests, but they're both checkout cart actions, right? So you really want to unify an alert on that. You don't want to see it from one or the other. You want to alert everybody.

And so because you're monitoring all this in one place, you can write an alert that covers both of those together which is pretty nice. And obviously, even if you're not doing that, combining alerts, if you get an alert, you can very quickly click through and say, "Hey. How's the Lambda site doing? How's our monolithic self-hosted application doing?" It's very consistently for me... Anybody whose really successful, they have all those levels of abstraction exists, right? Maybe they still own some bare metal some place. They definitely have virtual machines. They have EC2 instances and they have the Lambda and so being able to click around and see that stuff all together, ugh. Such a quality of life improvement.

Jeremy: Right, yeah. And I think you just actually made a really good point. I mean, this idea of having Lambda functions and EC2 functions running side by side, I don't know many companies that are 100% serverless. I mean, I know a lot are going that way and they want to go that way. I know I would love to be 100% serverless, but even some of my applications still have things that aren't serverless, and being able to put all of those into one platform is really powerful.

Nica: Yeah. And especially as you see ML tool kits get bigger. I mean, there's ways to implement that stuff serverlessly, but I mean, it's not straightforward. So right. That's going to be a really good example where it's like, "Okay. We're doing all this basic CRUD action serverlessly. That's great." But then when we need to... Image recognition and put a fun mask over everybody's face because it's St. Swithin's Day, yeah. That's going to happen inside an EC2 instance. There's going to be those exceptions, right? When we talk about step functions or other stateful ways to do serverless, it's like, "Should I just do this from the start because I like using states?" So it's like, "No." But sometimes you get stuck. Sometimes you have to use those tools, right? Yeah. Trying to get a picture of that whole map, right, that whole system quickly, hopefully that's where New Relic comes in.

Jeremy: Yeah. Awesome. All right. Well, Nica, listen. Thank you so much for taking the time-

Nica: Yeah. This has been a good one. I really enjoyed it.

Jeremy: ...to talk not only about serverless... I mean, or not only about New Relic, but obviously just your insight to serverless is really exciting too and really interesting. And if people wanted to go and learn more about what you're doing with serverless, maybe some of your side projects, but also in New Relic, how do they find out more about that?

Nica: So I'm still pretty bought in on Twitter. You can go follow me on TikTok too. If you go search Nica Fee, you'll see me over there. But most of my stuff is going to go up on Twitter. I am on Twitch twice a week. I'm on on Tuesday and Fridays in the afternoon if you're a US Pacific time person, but I'll announce there. I do a lot of hands-on demos there. But yeah, those are two great places. You also see me on the New Relic blog and various New Relic stuff. I was just quoted in Forbes this week. That was fun.

Jeremy: Oh, wow.

Nica: Yeah. That was neat.

Jeremy: That's awesome.

Nica: It was about how you need to be able to throw stuff away in serverless. You need to be able to just say like, "This service is no longer running. You have to go through and delete stuff."

Jeremy: Right, right. That is a hard thing for some people to do.

Nica: Yeah. Hey, I have games that I've written in C# code where I have reams of lines that are commented out. It's just like, "I might need these some day."

Jeremy: Some day. I know. I always do that too. That's another bad habit of mine, but-

Nica: They're like my old style Apple white lightning connectors where I'm like, "I just-"

Jeremy: You never know when that old iPod mini's going to come-

Nica: Come back.

Jeremy: ... and you're going to need to charge again. So anyways. All right, Nica, thank you again. I will get all of this information to the show notes. I really appreciate it.

Nica: Ah, thank you so much. Thanks everybody for listening.

Jeremy: Awesome.

View Details

About Heitor Lessa

Heitor Lessa is a Principal Specialist Solutions Architect at Amazon Web Services. He has spent the last 10 years in a number of roles, focusing on networking, infrastructure, and development. Since joining AWS in 2013, he’s been helping organizations of all sizes and segments across EMEA to design cloud native applications as well as software development best practices.

  • Twitter: twitter.com/heitor_lessa
  • Email: lessa@amazon.com
  • Serverless Application Lens Whitepaper: d1.awsstatic.com/whitepapers/architecture/AWS-Serverless-Applications-Lens.pdf

Watch this episode on YouTube: https://youtu.be/bFjT3TrpbZg

Transcript
Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Heitor Lessa. Hey Heitor, thanks for joining me.

Heitor: Thanks for inviting me. It is a pleasure to be here.

Jeremy: I'm super excited to have you here. You are a principal specialist solutions architect at Amazon Web Services. Why don't you tell the listeners what you do at Amazon Web Services and sort of what a principal specialist solutions architect does.

Heitor: I know it's a long title. I guess we can just say I'm a solutions architect at AWS. My day to day is basically working with customers and enable developer teams to find the best solutions on how to either build something on AWS or migrate, let's say a microservices monolith to a microservices or optimize something that they have.

But more recently, I'm also working with customers to help them build developer communities' inside. Similar to what we have at Amazon which bring in pros and stuff like that.

Jeremy: Very cool. Now I know you're doing a million different things. I don't know how you're not running AWS yet. I think you're next in line, I think. But you're doing a million things there. And one of the things though that you've been working on in the past and I know you're still involved with it, is the Well-Architected Serverless Lens. And I want to get into this because this is one of those things where if you're trying to find best practices and you're reading all these blog posts and you're looking at anti-patterns and good patterns and all this kind of stuff, I think it gets really, really confusing.

And your team and a bunch of other people at AWS and a bunch of the community heroes and all kinds of people got together and put together this Serverless Lens. If people are familiar with the Well-Architected Framework, which is talked about quite a bit, there's also this thing called the Well-Architected Serverless Lens. What's the difference between those two things?

Heitor: Sure. So the Well-Architected started way back in 2016, even before that to be fairly honest, where customers were looking to use AWS but we have roughly 50 services back then. Compared to today we have a lot more. And basically those customers were asking, "How do I use X service versus the other service? How do I go to production with this critical application? How do I model from my on-premises applications to something more cloud native? How do I migrate?" And things like this. Or even specific questions like, "How do I set up a multi-account? How do I better protect my accounts from a security perspective or billing?"

Well-Architected brings all those best practices that are agnostic from a workload perspective that typically applies to many of them, whether using serverless or containers, so usually what would work really well. But the challenge of Well-Architected as the platform evolved, we started to have more high level services like serverless or some service like AI/ML, which you have to treat them slightly different. The best practices still apply on how you set up AWS accounts, how do you do backups, how do you think about relational databases versus NoSQL databases?

But when you get to things like Multi-AZ and EBS volumes for serverless, they don't quite make sense. The Lens was a project to say, what are the customers using that Well-Architected actually helped them but they still lack a lot of good practices that are very specific to the technology they chose? So serverless was one of them. IoT was also another one. And more recently, last month we also announced Analytics Lens. If you're interested in big data, AI, those pieces, Analytics covered that pretty well. That's the difference.

The Lens is a... It doesn't replace, Well-Architected, it's more as an add on to the all these best practices we've been sharing for the past few years.

Jeremy: Yeah, because as an add on, it makes sense. The original serverless or sorry, the original Well-Architected Framework has, I think, 47 questions or so that asked you about specific areas and there's the five pillars and we'll get into some of that because I do think it's interesting to think of it that way. But the Serverless Lens just has more questions. What's the reason for all those extra questions?

Heitor: Sure. Well-Architected, when we started the Lens, if I'm not mistaken, again, there was the 47 questions but now we had just a recent update where some of those questions might change now. But the Serverless Lens, I think, if I'm not mistaken, we started with 31 questions because we were trying to get every single detail of servers and every best practice. But that was primarily a academic paper. So Lens started as a let's set up a document where you can go and find out when do I use serverless, is serverless as a good thing for me? How do I choose between all these services? How do I know the operational best practices for serverless?

As we started digging into those best practices, we felt we needed a lot more questions to dive into, okay, what type of metrics do you need? What type of alarm do you need? When do you use containers versus Lambda functions? When do you use orchestration versus synchronous calls? We only started in 2017. We had all these questions that customers were asking us. We put together into a document and we started writing. That took us roughly six to 10 months to put together into 50 pages when we announced Serverless Lens.

Jeremy: Right. And then that was a white paper, like you said. That was just a document. But now that's been moved into the Well-Architected tool, which is pretty cool. If anyone's used that or hasn't, I suggest you go and try it out. But that just takes you through and asks you all those questions and you can kind of keep track of your progress. How did you go from the white paper to the tool?

Heitor: Yeah. In 2017, when we announced, Werner went on the stage and talked about this idea of getting those best practices for serverless. And in 2018, we got an immense amount of downloads for Lens. If I'm not mistaken, it was over 20,000 downloads in less than six months, specifically for serverless best practices.

But then those questions started to ask more. How about Alexa? How about X, Y, Z? So what we found was trying to keep writing those pieces into the document was pretty difficult to keep up with the serverless space as well and how much it changes. What we found was instead of keep adding more questions and more pages of documents, we came together and thought, what if we evolved the lens project into the console? The customer would go to the console and say, "I want to review my architecture and I'm also doing serverless. I'm also doing analytics. I'm doing IoT or I'm doing something specific for FSI," Financial Services for instance.

So we thought we would experiment first with serverless and that's exactly what we did. But the challenge of migrating to the console as you probably have seen the console, you go and review your architecture, you have a very specific question and a few best practices you typically are doing or you're not doing yet.

That didn't map well with a white paper academic because you had to read what was the question, what was the best practice, how will you implement it? We went from those 30 plus questions down to nine questions with much more specific best practices. That we announced in February just a few months ago.

Jeremy: Yeah. Yeah. It's a great tool. One of the things, you're talking about best practices. Yesterday's serverless best practice could be today's serverless anti-pattern, because it does. It changes very rapidly and things are always sort of changing. How do you actually figure out what the best practice is?

Because there's a lot of posts out there about best practices and anti-patterns and so forth. Especially whether Lambda should be calling Lambdas and stuff like that. How do you actually go about deciding whether or not something is a best practice or not? And how does it make it into the lens?

Heitor: Sure. I think that's a great question. That was probably the fear number one when trying to think about serverless best practices in 2017. Because you remember API gateway was back there, it was very early days. SQS wasn't even EventSource back in the days. There were so many questions we were kind of unsure whether that would work or not. One of the two things that happens in this Well-Architected is that we have the concept of the pillars, as you mentioned, the five pillars; operations, security, performance, reliability and cost. That helps a lot in breaking that down of those best practices. And think does that fit into here or is that more of an opinion as of now.

And then the second is that we have a rule of 80% to 20%. If it's something that's been working for 80% of our customers and our technical field solutions architects, technical account managers, evangelist, developer advocates. All these people in those communities internally know that these are things that what's working for customers in production, then this fits into the 80%.

Things that are the 20% are what we call the edge or leading practices, are something that we know it might work for certain customers who have a certain expertise already with AWS, but eventually might become a common practice, like the Lambda-Lambda communication type of thing, that back in the day, we weren't even discussing that time or something like EventBridge that until this year, it wasn't something widely discussed. So something we definitely would get there.

Jeremy: Awesome. Yeah. Well, I was actually going to say something like EventBridge, which is probably my favorite service now that exists, maybe after Lambda. That's not in the lens right now. And you've had a bunch of new services that have come out like Lambda layers, Provisioned Concurrency, RDS Proxy. These are not in the Lens yet. Is it just because they're so new or there's certain things about those services? Have they not matured enough yet where they're considered to be best practices?

Heitor: It's a bit of two. I think it's always like... One of the pieces that we love at AWS is we like to launch those features or those services early so we can iterate on those features and those services with customers as we hear from them. It's quite similar to the process of deriving best practices for Well-Architected. When we announced something... Like EFS was actually just announced, immediately you would think, well EFS for Lambda makes a lot of sense for AI/ML use cases, for some shared state if you will, but it's something that we have to observe how that's working out for customers.

When you look at the Lens, the white paper per se, we have the scenario spaces that we not only share with you, these are the common architecture diagram for that specific use case, but we share what we call configuration notes, which is what are the common gotchas and caveats that might be different depending on use case and depending how you use.

One type of use case might be true to the vast majority, but if you have high throughput or high concurrency, the whole best practice landscape changes completely to you. That's one of them. For those new features, we are definitely listening, having our ears to the ground and hearing from customers how you're using, how is that working out for you?

Layers is one example. It's something that we've seen customers using for custom builds like FFmpeg or something very specific like a chromeless browserless if you will. It's working out really, really great. Everyone loves it and it works. But when you're trying to use layers in a very large enterprise, you have a couple of caveats. For instance, when you're trying to share dependencies that change very frequently, you have all those Lambda functions now being redeployed and that causes cold starts.

There are some caveats that we need to make sure we know exactly how to deal with, so when we write, we tell you, "This is why this is a good practice. These are the caveats and if you are in the caveat space, this is how you handle it."

Jeremy: Yeah. Yeah. And I know with layers too. I mean, I'm not a huge fan of the way the versioning works, where it's just an incremental version; version one, version two, version three. And I know you can use like SAR, for example, the Serverless Application Repository and you can wrap a layer up in a SAR app and give it semantic versioning and things like that.

There's just a lot of steps to go through where you're working out what is the best way to do it would be really interesting. I do want to go back to EventBridge for a second though, because this is one of those services where when it was announced and it was announced in the middle of the summer last year, I think it was. So it's about a year old now. And actually, it's just over a year old or it was just announced about a year ago, which is kind of crazy to think about.

There's not been a lot of literature on EventBridge. And I know that you see some people using it. You've got some dev advocates that are really pushing it. I love it. I think it's great. I think the things that have been added to it. But are there reasons why that still hasn't made it into the Lens? I know that you don't have things like DLQ's. Is that still one of the reasons why I think that's not there yet?

Heitor: Yeah, there are a couple of reasons. One, it's primarily time for me as well to distill all the feedback we get from customers and figure out what is a good practice versus what people learned by trial and error at some point, which is not something we want to tell everyone, "Just go and to use it." And there's also the fact of what you just mentioned about DLQs and some of those pieces. In the Well-Architected, when we're about to suggest a service or suggest a way of doing, we always have to keep in mind the five pillars.

For instance, I know EventBridge, we know EventBridge doesn't have DLQ, but nothing stops you from having a SQS queue as a target first, and then a Lambda function which will handle that. When you're looking from a large enterprise that's going to get easily 300, 400, hundreds of queues which then makes it more difficult for you to manage that piece at scale. Some of those pieces would come into play. But DLQs in some of those space, I think I would definitely add them, they're caveats. There are ways for you to work in production and works really well. But these are more caveats. So it's something we're looking at right now to introduce a scenario and not under question specific.

The pieces that we're not entirely sure yet, from a Well-Architected perspective on EventBridge is some of the event modeling, some of those best practices, inter-service versus intra-service, multi-account approach. It's very easy for the EventBridge, which by the way, it's one of my favorite services too, although I shouldn't be saying that. It's great that EventBridge can give you that visibility of how this async communications are going, their schemas that load code bindings. This is fantastic. But there's also some of the tracing capabilities that you want to know where these events went, who filtered, who routed you where.

We're trying to figure out those patterns and how can we tell customers to use EventBridge that will work for just 80% of customers that may not have this extensive background on event modeling, DDD and all those good tools and good practices. Customers were deep in doing event microservices or events at some extent, it's kind of a no-brainer. EventBridge, it just fits the bill and it works like that.

But on our side from the Well-Architected, we also have to be mindful of customers who may not have used the cloud or are trying to use for the first time and they are looking for those good practices to jump straight in. Going from monolith straight to something like events maybe a bit too much for them. We're treading carefully on that.

Jeremy: Yeah. Well, I don't work for AWS. I will tell people to go use EventBridge because I think it's amazing. One other thing though maybe about this and not EventBridge specifically, but you mentioned this idea of leading practices or edge practices. How much of a risk is it? Because every service that goes out there from AWS is... Everything has its problems here and there. There's little caveats here and there, but for the most part those services are solid and you can use those services or at least I've always felt I can use those services in production with quite a bit of confidence.

Is there some sort of rule that I should follow as a developer or as an organization where I say, "Okay, I'm following the Serverless Lens 95% but I do want to introduce EventBridge or I do want to introduce the RDS Proxy or something like that." Is that okay for me to do? I'm sure you'll say it is, but I mean, is that... Where do I draw that line? I don't want to be all leading edge but at the same time, I want to be able to take advantage of some of these new services.

Heitor: Yeah, I think as you mentioned, it's hard to say to everyone, "Don't use these services, don't use these features," because if the service is out there, there is a need for it. As we've seen that Lego going out and about explaining how to use EventBridge is a marvelous thing. It's something that I find amazing how customers are using EventBridge and many other services as well. And so it is with layers as well as an example.

The other pieces that I think I would divide that into two buckets that are ones that we're not entirely sure it's going to work in production. For instance, before RDS Proxy, many, many customers that I worked with for the past four or five years on serverless, we are all basically implementing SQL proxy clusters on multiple availability zones to deal with the issue of the connection pooling or using other practices as well. But we also know that we didn't have a reference architecture that would go and show about them how to do that piece.

From that reason, we refrained from referring to RDS Proxy because it was in preview. So if it's in preview, I wouldn't recommend production. Easy one, clear cut. The other pieces like we just announced... We recently announced Provisioned Concurrency. It's amazing for specifically Java applications on serverless or others that require some predictable latency.

Those are the pieces that because it's Lambda, but it's an additional feature, it's about trying and figuring out if your KPIs, if your requirements still work as you use that feature. It's not so much about AWS telling do not use that or perhaps use this instead, that this 20%, 80% is more of our role to make sense of this plethora of announcements that we also have that we need to make sense it works for the vast majority of customers.

If we don't hear from customers using it, it's difficult for us to prove it's working and we recommend that. One thing that in the Layers because I think it's a good point to bring that up now is what we call general design principles, which in fact, was the hardest thing to write. It took me, I don't know, months, with other people who are still figuring out how to do it.

The general design principles is what I use when I'm trying to use another service or trying to recommend something to someone that might not be into this best practices arena but I know it might work for them. The general design principles are seven principles that usually helps you understand whether serverless is going to be good for you on the use case you have or maybe we need to update those principles, please let us know.

But also if a service is going to help you out aimed at towards that direction, not only from the five pillars, but also from the principal perspective. I remember reading lots of articles from you about using step functions everywhere. And also now you discovered Lambda Destinations, which is a new feature. And there's this discussion about when do I use one versus the other?

In the design principles, we have one specific that we say orchestrate your application with state machines, not functions. Which back then, it made a lot of sense. But now with Lambda Destinations, this might be not entirely true anymore. That's kind of a, "Wow!" As long as I can orchestrate that with that feature, I'm still orchestrating that. It's not inside my code that's handling all of that pieces. Design principles help us to, I guess, navigate toward this fine line, as we call edge or leading practices versus best practices.

Jeremy: Yeah. Actually, from a choreography versus an orchestration standpoint, I actually am a huge fan of choreography. I think that's a better way to make microservices talk to one another, trying to orchestrate them. I love step functions and for workflows that absolutely need to follow every step and you might have to roll things back and so forth. I think step function is the way to go.

But once EventBridge was introduced, and again, having some more visibility into what things do, for me, I always set up one route that catches everything and just dumps it all into a Kinesis Data Firehose and then that goes into S3 and then you can query it with Athena. So you sort of have that... It's not a DLQ, but at least it's a record you can see every event that came into the system.

But once EventBridge was introduced, the idea of choreography just becomes so much simpler to kind of work around as opposed to having to use something like SNS or something else that has to be sort of set up and could be disconnected between services. I always used recommend the pattern of setting up just a microservice that was only an SNS topic that was your EventBus essentially. That was one of the... Because again, I'm a huge fan of that style.

All right. Let me ask you this question because this is something probably... You see a lot of people that just start building and I think if you just start building that's great. You want to build a monolithic Lambda function? If that's how you want to get started, great. Then you start realizing some of the benefits of breaking it up and some of that sort of stuff, whether it's scaling independently or it's the principle of least privilege and things like that, where you can really get very specific about individual routes and things like that.

But anyways, I do recommend people to start building. But at what point do you need to seriously look at the Serverless Lens and say, "Okay, we can't launch a production application until this is live." I guess the question maybe is what's in it for me as an organization, as a developer to follow this really strictly?

Heitor: Sure. That's a very great question and I'm happy you brought this up because there's a misconception in the fields that you use Well-Architected once you've already finished your application and now, you're thinking about go to production and what else do you need to fix or implement. This would be a costly way of doing business for one particular reason. When you're creating your sprints and how you're basically designing your backlog and what features you're going to prioritize and things like this, by the time you get to the Well-Architected and you get this, lots of questions and over 100 best practices, the reaction that most customers had based on my own experience for the past seven years working in AWS is, "Oh no, I won't have the time. I need to go live next week." And, "I will fix it later."

But actually that later may never come. And we know that. Other things get in the way and that happens and it's just a natural thing. I like to recommend people to use Well-Architected when you are thinking, when you are researching, when you are thinking about which service should I use, what are the common patterns should I use? Or what are the common things should I watch out for?

In the console, it might not be that obvious at first because we do that for a reason too. When you say review my architecture and you start answering those questions, even if you answer a single question or two questions, in the report or under the status of your application, it shows how many high risks you have and how many medium risks you have. I would only recommend you go to production if you have no high risks. Medium risk are something that for instance, if you have no Cloudwatch alarms in your serverless application or no tracing or no structure logging or centralized logging at all, then it might be difficult for you to go to production.

But if you don't have, let's say Canary deployments as an example, or if you don't have high itempotency in certain parts of your application, it's something you can definitely go in production and improve as you see along. Looking at high risks first, if you get that nailed, absolutely go forward. And medium with something, there's always room for improvement and we know that.

Jeremy: Yeah. I like that. I like the strategy of serverless anyways in applying that to the Well-Architected piece is it is very iterative. And actually, James Beswick just had a post that showed like, "Hey, I'm going to start by capturing this one thing, but then I'm going to send it to EventBridge or to SNS and then I'm going to process some secondary component and then I'm going to do something else." I love that idea of building incrementally, but I agree. If you're running npm and you see that you've got 900 high risk issues, you don't want to deploy that. So it's the same thing I would say with the Serverless Lens.

All right, so let's get into some of the details of the Serverless Lens itself. Because we've been talking about best practices and some of the services you can use. So let's actually get into those. We don't have to spend a lot of time on it. But I think it'd just be interesting to kind of review those so people know which services are available to them and what are the sort of current best practices. Like you said, I'm sure that's going to evolve over time. But anyways, so here's... Let's start with the compute layer. This is sort of in the white paper, there's a definition of all these.

And you should definitely go and read the white paper, by the way. If you haven't done that, that's just a really good resource to have all that data and you can kind of get that beforehand, maybe even before you start building your application so that when you start going through the tool, that those checkboxes are there for you. So let's start with the compute layer. From a serverless perspective, what are the compute options available to us from AWS?

Heitor: Specifically on the lens, we have Lambda for the doing the compute side of things for you, we also have API gateway for doing some of the REST spaces and we have step functions for doing some of the orchestration of that state. Ideally, AppSync would be there too. But we're trying to figure out whether that's going to be in the next updates for that.

But these are the primary ones, based on the best practices we have. Compute layer for us is any service within the Lens that we selected that process your external requests, do some sort of computing, do some sort of a controlling access to those requests before they get you your business logic for instance.

Jeremy: Right. Yeah. And I think it's actually interesting that API gateway is in that compute section because it does actually do quite a bit. You can do throttling and you can do transformations. API gateway is a very powerful service that has some really cool features around it. All right. What about the data layer?

Heitor: The data we have... Well, you basically are working with persistent storage, we're not dealing with the cache specifically yet. This, we were looking at DynamoDB as one of the clear winners. DynamoDB is definitely being used by quite a lot of customers, specifically on the serverless. And when we call on not only Dynamo, but we call out specific pieces of Dynamo like DynamoDB Streams, and more recently, DAX that we've added too. And we also cover other pieces like S3, we cover Elasticsearch. And then that's where we cover AppSync as well.

The reason why we're undecided between AppSync being the data layer and also on the compute layer is that a lot of customers are using AppSync on designing what we call a schema first. So they're dealing exactly how your model application which looks similar to a database modeling if you will. I wouldn't say kind of but we had to make a decision. And so AppSync that in that case, actually, it could fit in both criterias. On the compute, because there's a lot of authorization, a lot of logic, VTL, like API gateway. But there's a lot more of data aspects in AppSync and we tend to say that specifically in the Lens, if you're building data-driven applications, when you're trying to model things around your data, then AppSync from a GraphQL perspective, if it's something new, then it makes a lot more sense.

Jeremy: All right. Now, question for you. How did you let Elasticsearch creep into a Serverless Lens?

Heitor: One of the things that happened in the definition was we had all these questions first, do we add containers in there, specifically Fargate? Do we add something like Elasticsearch because even though it has servers, it's something that we know customers are... Specifically in 2017 was the most common solution for analyzing your logs and there was kind of a best practice we needed to include.

I think it has to updated this year to rehash some of those. But the definition for us is a way to introduce all of those services that we're going to talk to out the lens. And we created categories to basically introduce what exactly that service does within the architecture you choose. For an Elasticsearch, Elasticsearch was specifically you wanted to... You have a scenario for mobile applications, so full-text search for those. Elasticsearch is actually the only option right now. Well, you could do that in Lambda in somewhat different ways. But that's kind of off the point now.

Jeremy: Yeah. But it's funny with Elasticsearch, because that is... It's been my go to for, I want to say since 2009, maybe 2008, well before Amazon even had the Elasticsearch service in place. That is one thing. That is definitely one of the missing pieces of serverless is to have a full-text search capability around that.

Also, I guess, caching as well for just a generic caching layer. I know some other providers have other things. What about Aurora Serverless, though? I know that's not in there now. Aurora Serverless still has problems, the same problems with exhausting connections and zombie connections just like you would if you're connecting to RDS. But actually it doesn't work with RDS Proxy. Which is interesting.

And if anybody wants to know why that is my guess, and maybe you can answer this question, my guess is because you need to exceed a number of connections and CPU usage in order for the autoscaling to trigger. And if you don't, if you had RDS Proxy in front of there, then that might not trigger your server to scale, which is actually part of the problem with my serverless MySQL package is that if you put that in front of Aurora Serverless, if you don't set it right, it doesn't scale up, which can be more of a problem. Anyways, where is Aurora Serverless on that list? How much of a risk is it to use that?

Heitor: I wouldn't call it as a risk because we know customers aren't using that. So the reason is not in the Lens yet. It's exactly for the reason you just mentioned about that connection pooling because we didn't have specific guidance on what are the best ways to tackle that. Now we have RDS Proxy. Now it's something that we could not only bring RDS, ElastiCache and a bunch of other services that are VPC specific, which previously had a lot of latency.

Now we can bring all these services in the new updates. And then we can add some call outs, if you are going to use Aurora Serverless for instance, here are some of the caveats that we know customers are using successfully in production. For now, it's not because it's from 2019 from the last re:Invent. But in upcoming updates, we do plan to have that.

Jeremy: And there's the data API too, which is very cool for, I would say, for asynchronous stuff. I don't know if it's ready for synchronous processes because it does have a higher sort of startup latency. But yeah, again, so many... I don't even know how you're going to fit all these things in a single Serverless Lens. There's just too many things to add. All right, so what about messaging? The messaging and the streaming layer? So we talked about EventBridge not being there yet. What do you have available to do that?

Heitor: Yeah, so before EventBridge, which is something again and we do plan to add in upcoming update, especially now that so many customers are using and James Beswick, Developer Advocate, has been doing a great job evangelizing some of the possible use cases. Before that, the classic ones you just mentioned about; SNS, SQS. I didn't have SQS specifically call out there. But SNS is basically like the go-to for low latency asynchronous communication between services. And that's still the case today. If customers are using EventBridge and SNS, and you can use SNS for very low latency.

Although, it's not the same feature set. It's a very different service. And then streaming well, Kinesis kind of clear winner. And then we also add the Kinesis Firehose. The only difference there is that we worked out a lot more, I guess specifics about Kinesis and streaming. But we didn't add because well, now it's public, we have a specific Lens about analytics that dives into much greater length about Kinesis Firehose and some of the configuration pieces that you might want. Messaging, we basically kept it short for SNS but we do plan to add now with SQS as an event source, and EventBridge.

Jeremy: All right, so now even though you have to pay per shard for Kinesis, do you consider Kinesis to be serverless?

Heitor: I think we had this conversation when we wrote in the 2017. I can't tell you how much of a debate. Just to give an idea, when the first Serverless Lens, I had roughly 70 plus revisions before we got out. And even before we went out publicly, I had over 700 revisions, 700 edits on things like, "Is this serverless? Is this serverless? Is this..." What I basically had to do is... The agreement we made was, there are things that are not serverless. Elasticsearch is definitely not one of them. Kinesis on the other hand, you definitely have this knob that you have to tweak about the shard counts.

But at the same time, it's something that empowers or is basically the backbone of many serverless applications doing streaming. For that reason, we decided to say, "Is this something that is the backbone of a serverless application successfully running in production?" If it's it, then what can we do to make it more easier to manage and easier to operate following those best practices? So Kinesis falls into that bucket.

Jeremy: All right. You didn't answer my question. I wanted you to email Chris Munns and tell him that yes, it was serverless. All right. How about the user management and identity layer?

Heitor: Sure. Well, in this case, we only have Cognito. Cognito helps us and that, I would say serverless. Though I implicitly answered the previous one. So Cognito helps us to do the old offload mechanisms or more recently, a lot of customers are now using for custom authentication mechanisms like passwordless, signatureslack or many other communication tools nowadays.

So Cognito falls into that bucket. There's not much to say there. We didn't spend too much time explaining too much about identity pools versus user pools. We briefly talk on the security pillar about ensuring that you're using identity metadata like Scopes in OAuth flows and many other mechanisms to do something more secure. But beyond that, it's plain simple Cognito integration and using Federation if you can too.

Jeremy: Yeah. And what about like JSON Web Tokens? I know that the new HTTP API's, which we haven't talked about, but those are primarily just used... I think that's all they uses is the J-Web Tokens. Is that something where the Lens might eventually get to a point where it says, it's okay to use OAuth? That might be fine for the identity layer? Because you're already integrating with those. Or is it something with a lens is going to say very specific to just AWS services?

Heitor: No, not really. If you go to the security pieces, when you go to the very first question actually, when we ask about some of the security identity or throttling, if you will, we do a rundown in the paper specifically, of when to use IAM authentication-64, when to use API keys, when to use custom authorizers, when to use something like OAuth like JWT. Previously with a REST API gateway, it was very contrived example of just using validating the data that the token was valid. That works for simple use cases but it wouldn't work for something more enterprise where you need to verify a lot more logic on JWT specifically.

We don't make that distinction about do not use JWT or use this instead. What we call out in the Lens is here are all these possible ways of you to do authorization, and specifically, this is authentication, this is authorization and these are the pros and cons of which. That's kind of the line we go.

Jeremy: Right. All right, so what about the edge layer because we I think edge computing is getting extremely popular. I don't know if you've been following along, but CloudFlare just did this huge thing where now it's like nanosecond cold starts and expanding the workloads they can do, adding more languages. So I think the edge is going to be really interesting, especially from a compliance standpoint in terms of where you're processing data and handling workloads and things like that.

Where are we now with the edge in terms of the Serverless Lens, but where do you think... I'd actually asked you a question beyond the Serverless Lens, but where do you think AWS is going to go with that edge computing stuff?

Heitor: One thing I was... Actually first I would definitely agree. I think it's something that's becoming more and more popular. One thing that I was very surprised to see the uptick of customers using it and without naming customers specifically yet because they're not public, is the amount of customers going from single page applications which we've been seeing have been popular over the years to something like going back I guess, if you will, into the server side rendering and now more specifically something like Gatsby which is something super popular and super handy.

But that Lambda app edge thing or doing the compute at the edge is becoming hugely popular in the streaming. And more recently, customers are using edge to do not only click streaming of those analytics pieces, but also doing data ingestion in multiple regions which is something that has been quite popular now. For the edge layer, it's something I want to add in the new update as well. Specifically cover the server side rendering.

There are customers doing hundreds of thousands of requests per second on server side rendering that it's something that we want to detail a bit more what we mean by server side rendering to do that at the edge, how do you do cache, so you reduce your cost, but you equally get the performance and SEO of that.

At the moment from what I've been following on this on the edge pieces, customers are now more comfortable with server side rendering and now more recently, incrementally static generation with something that Next.js just did. We have being following on that pieces. The applications that live completely in the edge, I haven't seen that much yet. But as edge makes more progress, lifts some of those limitations for timeouts and RPC calls, I think we might be seeing that shortly.

Jeremy: Yeah. No, I agree. I think that SPAs are great and they have their use and I think if you're doing server side rendering and then that rendered page then the first paint happens very quickly, then you can start interacting with it and so forth. I think it is a really popular way that we're seeing a lot especially like you said, with Next.js and like what Vercel doing and some of these other companies. I think that's really, really interesting.

That'll be cool to see where that's going to fit into an overall serverless strategy, because certainly if your front end is going to be rendering web pages, then I think the edge is going to play a major part in that. And if AWS has a good strategy around that, that'll be very interesting to see. Okay, so system monitoring and deployments.

Heitor: In that case, it's a very simple one. We see Cloudwatch as being like the backbone of all those metrics and logs and KPIs that customers use. X-Ray, which is our official mechanism for doing distributed tracing. And then SAM is when we call out our official way of doing deployments as well. But we don't rule out specifically choose these framework over the other framework, what we recommend instead or from being timeless, because as you mentioned, it changes a lot, the landscape, is do use a serverless framework.

In this case, we're basically outlining Sam as the official version from AWS. But we also recommend many other frameworks as well. It's all about what services to use for metrics, KPIs and logs and tracing and you use should make sense of all this little Lambda functions that tend to grow organically. How do you handle those? That K is a framework.

Jeremy: Right. And actually, I think that that category is probably one of the largest categories that's been or the category that's been affected most by third party services. So you have all of those monitoring tools that have launched. You have a number of different deployment engines and frameworks that help with that.

In the system now or in the current version of the whitepaper in the Lens, you're recommending SAM as you said. What do you think about the CDK though and also maybe SAR. Where are those going to fit in do you think in the future of the lens?

Heitor: The SAR, we have some references already as links but not specifically as a service because it wasn't something that, I think it acts more... Aids your deployments as opposed this is how you deploy things as in SAM or Serverless framework, if you will. And the CDK is different though. CDK, it wasn't GA until recently and some of the constructs... I think, if I'm not mistaken, API gateway still is not GA, it's actually either preview or beta at the moment.

There are certain things that I wouldn't be able to recommend in the Lens. As in, we know customers could use L1 constructs like lower level constructs and make their way up because it's basically CloudFormation either way, but in this case, we're basically recommending SAM because we know it's working in production and we know they can just use it. CDK, it's more of a discussion. We need to think how we can frame this in the Serverless Lens.

I think what CDK enables today, it's something we couldn't do easily before. Like the likes of Liberty Mutual and many others. Alma Media, is one of the examples that I was blown away by how they use a CDK so effectively. When you're doing multiple best practices at a larger organization in multiple teams, CDK makes it so much easier to onboard those practices, internalize those blueprints than if you were to do SAM or Serverless Framework. While Serverless Framework has the components if I'm not mistaken, which can now do something similar, but doing it in an imperative way, it makes it easier.

But at the same time, I have recently have found customers actually tripping up with the amount of abstractions. Developers got abstract. It's just the case. If you have the power of doing it, why are you not doing?

Jeremy: Exactly.

Heitor: But then at the same time, you get the problems with, I now have no idea what this line, build the best application possible serverlessly on AWS constructor and then you have to dig in and it's kind of a complicater. And this is not a new problem. This is something that we've been seeing.

Like 2015 when I started doing microservices in production, when customers would have 7 deployment tools because one found a better way to abstract things. I think CDK has its place on those customers looking to use programming language to easily deploy those applications but in the serverless, SAM is definitely predominant. CDK on containers on the other side is definitely made life so much easier to deploy containers on AWS.

Jeremy: Yeah. And I was sort of against the CDK initially, because I was like, "Oh, there's another layer of abstraction on top of every layer of abstraction on top of a layer of abstraction." But what I do like about what people are doing with the CDK and then you mentioned this about sort of baking in some of those best practices and same thing with serverless components is sort of, if you're a team and you say, all right, here's the bootstrap for a serverless microservice. And really all we want to do... All of my X-Ray and my logging and whatever my security best practices are an all that stuff, any layers that I need, if you can encapsulate that all into one construct and then be able to just add services or add routes or whatever it is that you're doing on top of that, I think that's really powerful level of abstraction.

Because I think that could give us the ability to just say, all right, I don't have to have a 600 file bootstrap template that I use to start every new serverless project, that a lot of that stuff could just be baked in. And again, same thing with serverless components. Even Pulumi, and some of these other ones that are doing sort of similar stuff. I do like that idea of potentially being able to sort of encapsulate that. But anyway, so deployment approaches. You mentioned Canary deployments earlier. What are the best practices now for those deployment approaches?

Heitor: Deployment approaches haven't changed much. It's still the same as before. We still have all at once which to basically deploy a thing specifically on Dev. You're deploying something, you're iterating fast, and you want to make sure whatever you deploy, it's working or it's not working. When you're going to production, you still have the mix between should I do Blue-Green? Should I do canaries? Canaries is an easier one to think about, to reason about because you have to have a lot of traffic to be able to shift a percentage of your traffic to a newer version.

And that's kind of where most people get tripped and when you're doing server side rendering, which is why I want to have a dedicated piece of server side rendering at the edge before trying to use Canary deployments at server side rendering when in fact more than 70% of the traffic was being cached.

You wouldn't be able to see any of that and then you would go and then you'd break. Blue-Green is kind of a classic one. You keep both of them and then you switch or using a DNS or using some other pieces. It hasn't changed much. It's not specific to serverless per se. But in the Lens, we cover how you could do that using SAM or using any other frameworks. I think there's a table. I'm just looking at the lens paper now myself. There is a table that we basically tell you, "These are the differences between all these three and when to use each."

It's more common for people to use linear deployments. So you're shifting a percentage of traffic over a period of time. And then you use KPIs to revert if something went wrong.

Jeremy: Yeah. That's interesting, because I think what you... The best thing you can take away from that is just do not do SAM deploy right into production from your laptop, for example. You should have some strategies in order to especially see ICD and some of those other things that I think make a lot of sense. Read the paper, figure that stuff out. I do want to talk quickly though, before I let you go about some of the use cases and that's one of the things that the paper does is it outlines a few scenarios.

This is what I think more people need to see because like you said, your best practices and the things that make it into these papers are based off of whether or not customers and technical specialists and evangelists and so forth are using these successfully. Let's just go through these quickly. Just give me an overview of it. And maybe you can outline some of the best practices for these. But like RESTful microservices, for example.

Heitor: The RESTful microservices... Well, actually one of the classic ones, once API gateway came out, most customers were using this as the go-to use cases. The microservices, instead of creating a bigger picture of what the microservice might look like, it wouldn't fit into your diagram, we chose to do something more conservative, which once you update now, in the next one to show a bit of caching, a bit of other pieces that also come through. Now, a VPC that enables you to do more interesting things.

In the RESTful API, what we cover is how do you have an API that your client will interact with with a contract and then how your backend, in this case, using Lambda functions can interact with your persistent storage. We chose something very simple and then you basically just store something into DynamoDB.

However, into the caveats or configuration notes, we actually covered some of those pieces that are more specific. Like we talked about data like geographically being close. How do you work with API gateway access logs. Back in the days, people were just enabling logs for API gateway and all of a sudden you have your incoming requests and your responses all in plain text in your logs. Customer were more sensitive to security, they would be like, "Oh, no, there's got to be a better way."

In the configuration notes, we would basically talk you through some of those things. How do you do logging for the REST API gateway the better way and how do you basically model some of those other pieces to do full-text search on logging operations. It's a very contrived example. It's more to show. This is how simple our RESTful micro service could be in serverless. But it's something that we want to update to include now DAX which wasn't easily done before a DynamoDB with Lambda or ElastiCache now and things like this.

Jeremy: Yeah. And I love just the idea of RESTful microservices with serverless. Because again, it just... You don't need web servers anymore. It's just amazing what you can do with these APIs. And I know there are a lot of blog posts out there that say, "Oh, the cold starts and so forth. It's not ready for primetime." It is ready for primetime.

So there are lots of your customers using this. I know I've been using this for, I don't know, probably 25-30 projects at this point that are out there. Very good scenario and example and use case for that. All right, so another one that I know that Aleksander Simovic would like is the Alexa skills.

Heitor: Yeah. The Alexa skills, it was a partnership with the Alexa team. We know many customers have been using Alexa for Lambda. There are also other use cases as well. But Alexa was one of them that not only Alex as well used a lot, and advocates a lot, he did so many things as well. But one thing that we saw was missing from the Alexa was we were always explaining to customers how to use Alexa with serverless into here's how you can choose a random number from one to 15. Or here's the Hello World example.

There wasn't anything about good practices or good design decisions, because this is also a very different way. You also need to think about UX of your audio and your transcript and how you're interacting with the customer. It basically gave a little bit more room for Alexa scenario to tell you when you are designing a skill, what are the things you have to keep in mind and what a good experience or what delightful experience actually means.

Some of those kind of things should keep in mind, and we also go into more detail about how a proper Alexa's queue might look like. It's not going to be something like an Alexa talking to a Lambda that talks to a Dynamo. There are other things as well. That we cover things like DynamoDB outscaling, what if you're using IoT with Alexa homescale? How does that fit together? That Alexa skill covers a common Alexa's skill how to design it, and when you expand and Alexis skill to do more things, how does that look like as a whole?

Jeremy: Yeah. All right. What about Mobile backends?

Heitor: Mobile backends is a... You probably have seen the server line example. It's becoming one of my favorite ones nowadays. The Mobile backend, it's covering AppSync or GraphQL specifically on how customers are specifically building mobile applications these days. There's been some changes with data store and Amplify changed a lot recently we have to update. But it covers things like when you are dealing with SMS or multiple-factor authentications or user registration or assets or dealing with a single graph as we typically call in GraphQL and handling multiple types of data sources like Elasticsearch for full-text, some part of your mobile application, parts of your data that could be into DynamoDB or NoSQL, parts of your data that could be in a relational database and some other third party communications that you want to use Lambda with.

The mobile covers all of these aspects on how we use all these different services to hydrate data that's in DynamoDB that now goes into Elasticsearch or how does your user use a single API that can talk to different data sources based on what your customer wants?

Jeremy: Yeah. I love the approach to Mobile backends especially with GraphQL and being able to avoid that overfetching and underfetching problem. That's all great stuff there. All right, what about stream processing?

Heitor: The stream processing is one that it got a new update not in the Serverless Lens, but specifically on the Analytics Lens. It covers things like the stream processing and how you handle batch processing specifically. It doesn't go into a very detailed like Analytics Lens as of now, but we do cover things like best practices about using a single shard but when you have to use a new parallelization factor or how do you design a good streaming solution to your payload and stuff like that. How do you handle high throughputs in DynamoDB with streaming or partition keys and stuff like that.

For Lambda doesn't have a specific library for handling like KPL or Kinesis Producer Library or Kinesis Consumer Library like you normally have in a EC2 or container. So it gives you some of the workarounds on how you can handle that, some other libraries that you could do or dealing with duplicate records or idempotency, and things that evolved as well.

Like we used to say, because of the way Stream Processing work, if your Lambda function fails, it will block the stream and it would keep sending the same records. On queue, you'll probably have some data loss. Recently, we announced async controls that give you more flexibility. We added that recently too.

Jeremy: Awesome. All right. And then the final one here is the web application, which is sort of goes beyond just, I guess, the RESTful microservice and adds S3 and some of the other stuff, edge computing. I know you said you want to update that with some SSR and some of that. But what do we have currently for that scenario?

Heitor: At the moment it's very similar to the mobile application. It learns from the mobile where you have the static assets into your S3 and you have a CDN on top. So you separate the two. You still use Cognito exactly the same way for using user authentication, user management. But you're now dealing with your API gateway, handling the authorization of the JWT token that you got from Cognito and then landing in DynamoDB, which is very similar to the REST API.

The difference is that you're now using some more sophisticated authorization mechanism that you probably would do in a REST API, because it could be service to service communication where IAM would be a lot simpler. Or you could also use custom API keys when you're doing things like throttling, but not only throttling, but also tiers. Your application is in freemium, in premium or business or enterprise. It got a little bit more details on if you're going to go down that route. These are some of the good practices you have to follow.

Jeremy: Right, perfect. All right. I wanted to talk to you about the five pillars, but I think we're running out of time. Maybe what we could do is just... So you mentioned them earlier. You mentioned operational excellence, security, reliability, performance efficiency and cost optimization. But I think what a really interesting aspect of serverless on this has to do with cost.

There's a whole bunch of things around reliability and security that are already baked in. But I'd like to talk to you just, use some time wisely here and talk to you about the cost aspect of it. Cost optimization in serverless applications. And I guess, in the Well-Architected Framework in general, what are your thoughts on that? Why is that so powerful, you think?

Heitor: There are many aspects that we could tackle. I think when customers think about cost, they typically would think about if I have the server running, how much it would cost versus having a serverless approach? They will try to do apples to apples when in fact it isn't. I think we have this discussion many times written in blogs like yourselves, or Yan Cui or Ben or even Lego as well on growing serverless teams, when in fact, one of the most costly aspect is actually developer hours.

Actually, those developers, one of the things that... I think the main reason I was so passionate about serverless was when I used to work with customers where we had to have a platform team, we have to have a SRE, if you like and many other people. Basically maintain a basic platform to run those services that they needed. And serverless, when I was working with British Gas, specifically or Centrica, as they went on to re:Invent three years ago, all we had was you know what, let's start with four developers, you add one architect to help us and you have someone from security as well and someone with ops so we can basically have a team that we can have people on and off. But this developer should be able to do.

Basically three months, less than three months to be honest, they had no AWS experience and they got off the ground and got something production as well with those practices. That changed the cost perspective because there was one of the applications that they had over 100 people to maintain. When you're thinking about a server or a... We don't even have to go too deeply on the load balancer versus Fargate, versus in all these nitty details, by people specifically, instead of saying, "Oh, we don't need all these people anymore." It's actually quite the opposite.

You could train these people now to do a different role. And you now have this army of talented people already in your organization, you could be retrain to add more features, and you can ditch the competition in a way. I think that cost is something that you've seen the Serverless Lens as well. Which was also the hardest. How do you ask questions about costs when serverless is mostly cheap, if you will, inexpensive if you will.

Jeremy: Exactly. I think that to me is the biggest... The biggest cost factor is not how much does it cost to run a Lambda function or what does it cost for API gateway? There are certain services where you can start running up some bills, but for the most part, it's just how much does it cost to have those SREs like you said or to have those DevOps people or to have all these other ops people that have to constantly monitor servers and make sure things are up and running and then just the wasted processing time that you're... All that idle time that you're probably getting over provisioning, under provisioning, trying to set up autoscaling.

There's just so much that you can save by going serverless. Awesome. All right, I have one listener question for you and if anybody wants to ask questions to the guests here on Serverless Chats, go to serverlesschats.com/insiders, sign up to be an Insider, and you can ask questions to our wonderful guests, like Heitor here.

I have a question from Michael. And he... I'm not 100% sure I understand the question exactly, but maybe we can break it down. He said, "I'd like to learn about best practices on sharing models between services that use the same table." So talking about I'm assuming single table design in DynamoDB. "So in my case, I have an API service and an ETL service which share the same table, and I haven't found an approach that I'm happy with yet."

I don't know if there's talking about entity models or API gateway models, but I don't know. What are your thoughts on that question?

Heitor: Yeah. I was going to ask you the same. I'm not sure if he means about API gateway models where you basically define your contract and how your client or your consumer is going to work with or if it's about the data entity model when you're trying to design a database or a single table, if you will.

In the case of API gateway for models, I think there's a great example of... We have an open source and example code serverless e-commerce platform, where I think is the most... It is the most comprehensive example that shows how you deal with API gateway models, especially event schemas for EventBridge as well and the tooling around it. And then it shows you some of the design decisions and why we made that what we made.

Have a look at that one to give you an idea about the tooling and how to share some of those models across services, when it makes sense because there are some parts of the contract that you can share. From a database perspective, I think he goes into where we think, is that multiple services accessing the same database or should we... Maybe we need to think about that? Can you explain a little more? Or is it something just like an ETL, like a service airline. Is another service adding new flights into the service that already handles those flights.

Then in that case, it's more about making sure you don't get into the situation where the ETL uses more of your database than your service can use at the time because I've had many incidents in production like that as well, even in serverless. When the ETL function basically took over all the concurrency of the whole account. So you need to be mindful of both things. But from the models perspective, I don't know if that person means single table but it goes... As long as they have a way to protect your ETL or not over consuming or over utilizing in a way that impacts your customer experience, I don't see that much of an issue.

The issues I saw with models is mostly managing with frameworks like service framework or SAM, how do I store this things that are stored in a file? How do I make sure it's easy to change? But also, especially in that single table context, it's quite complex to get it right. But once you get it right, it looks like a dream. But then you also need to think about when you make changes, how do you make those changes? I don't think I have an answer if I understood that correctly. I think we're going in circles, I guess.

Jeremy: No. Yeah. Well, I'm wondering too if maybe and from the ETL perspective and maybe this doesn't answer the question. But if you have an API service that's accessing a table and loading certain types of data, then you have an ETL task that runs at night, it depends on what that ETL task is doing. If that's converting or if it's doing aggregations maybe. So it's aggregating counts across logs or something like that, I wouldn't want that ETL task touching my production table directly.

I think I would want that ETL task, if it's doing some sort of roll up, that should operate maybe in its own table and then just send those aggregations through an API gateway maybe or through an event. Maybe EventBridge. And so that the API service could accept that event or those updates, but through a contract, maybe through an API or through something like an event schema. I don't know.

And maybe this is not answering the question. But I think it's interesting debate too, because that's part of the problem with microservices in serverless is where's the boundary of the microservice and does every function get its own table? No. But I mean, that's the kind of thing. Like how many functions should be interacting with it and should functions ever cross between bounded contexts? I don't think they should. But I think there are still a lot of people trying to figure that out.

Heitor: In fact, I was looking at the analytics lens. I was just searching for ETL and they actually go into a great length about how to choose your ETL. Whether you need like a nightly batch or if you're doing a string batch or on demand batch, or high frequency ETL. Have a look at that. If hopefully that answers your question on the design principles of analytics lens, there are a bunch of scenarios specifically on how to use ETL. And separate your query like you're just mentioning, Jeremy, when instead of an API, you have a data lake, if you will and you use Athena to search for that specific aggregation.

If that's what you're after, it's quite difficult to know the question but Analytics Lens covers a lot specifically about ETL. Ping us on Twitter. We're happy to help, have that debate publicly as well.

Jeremy: Awesome. All right. Well, listen Heitor, thank you so much for spending the time with me sharing all this knowledge. If people want to find out more about the Serverless Lens or they want to contact you, how do they do that?

Heitor: So the Serverless Lens, if you just search for Well-Architected Serverless Lens, you will find the whitepaper. But you also go to the console and if you search for Well-Architected, you will find a Well-Architected tool right in the console. When you're searching... When you're creating a new application, actually, you can basically select Serverless Lens and you get all the questions and the best practices.

That's the best way to find the best practices. If you don't want to read the 80 pages upfront, because it will give you more summarized version of what is the best practice, why is it important for you to do and how exactly do you do step by step, how do we evolve that.

And if you want to find me, I'm also on Twitter, @heitor_lessa and if you need to message me as well, if you're doing something on the Serverless Lens and you want to give some updates or if you are have specific feedback, reach out to me on email as well. Use my last name, Lessa, L-E-S-S-A@amazon.com.

Jeremy: All right. Well, we will get all that into the show notes. Thanks again Heitor.

Heitor: Big pleasure. Thanks for having me Jeremy.

View Details

About Paul Johnston:

Paul Johnston is an interim CTO, CTO and strategist who has particular interests in serverless, cloud, startups and climate change. Formerly, Paul served as a Senior Developer Advocate at AWS for Serverless and CTO of multiple startups, including one of the world’s first serverless startups. Paul is also a co-founder of ServerlessDays.

  • Twitter: twitter.com/PaulDJohnston
  • Medium: medium.com/@PaulDJohnston
  • Project Drawdown: drawdown.org/
  • Roundabout Labs: roundaboutlabs.com/
  • Leading Edge Forum: leadingedgeforum.com/
  • White Paper: The State of Data Center Energy Use in 2018
  • IPCC Special Report: Global Warming of 1.5 ºC
  • Blog post: To fix Climate Change, stop being a techie and start being a human

Watch this episode on YouTube: https://youtu.be/DTpP7RGXV6g

Transcript

Jeremy: All right. We talked about the big three a little bit and compared them in terms of their green and stuff like that, but in that paper that you wrote, you have this cloud league table in there where you compare them. I'd love to know more, what about Alibaba and Oracle and IBM and some of these other things, where do they all stack up against one another?

Paul: They aren't as big. Let's just be clear on that one. They aren't as big, and their green credentials are less clear. For example, Alibaba is very big in China for obvious reasons, it's a Chinese business and yes, they have a footprint outside of China, but they're a primarily Chinese business. When we looked and researched and were trying to find out about all of their green credentials, we found very little information whatsoever. It was almost non-existent. We found a little bit about efficiency in data centers and putting things in cold regions of China. You're like, "Well, that doesn't actually change anything if you're growing at a massive rate." It changes the conversation a little bit.

Jeremy: Can you do that though? Can you pack your servers in snow? Does that help with the cooling bill?

Paul: It depends on the server, I suppose. But I'd say you end up with this, not every conversation is equal. This is shown across the political spectrum as well. The conversation in China is very different to the conversation outside, in terms of industrial nations. Industrialized nations such as the U.S., UK, Australia, and Europe as well. You have very different social context. But anyway, coming back to Alibaba, we just found very little information. IBM and Oracle, actually, both had lots of information. IBM has had a commitment to renewables and sustainability for a very long time. I think since the '70s if I remember right. They have good credentials, but they don't offset their renewables in data centers or regions.

Paul: I think I remember, from the top of my head, because I can't remember everything, they are renewable in the UK. I think Oracle are renewable in the UK anyway. But some of these regions, they are renewable, but not all of them. But it's not clear unless you dig into the paper and unless you dig into their information. Nobody, as far as I can tell, has a little green dot against a region that says, "This one's renewable." It would be so much easier if they did. These other organizations, none of them, apart from Google and Microsoft have really made a play for being the green advocates in the space. Amazon does, but that's a whole other conversation which I'm not going to go into at this point. But the three lower down, I think, are struggling to be able to play in the same conversation. I think they would like to be seen as green, but I don't think they are really pushing the agenda because they don't see it as a point of differentiation.

Jeremy: All right. Then what does the tech industry have to do as a whole? I know you had some recommendations in your paper about this, but just what are maybe the top two or three things that the tech industry as a whole could do to address the climate crisis?

Paul: The climate crisis as a whole.

Jeremy: Or I guess their impact on it anyways. Let's start with that, we could build from there.

Paul: I don't know anymore. I think I've gone backwards and forwards on-

Jeremy: Don't give up Paul, don't give up.

Paul: It's not that. It's just there are so many things. I think my biggest thing is the tech industry needs to find its activist voice. I think that would be my point. I think sitting there and going, "Oh, everything's going to get fixed by technology," is entirely the wrong approach. I think my personal view, as much as I like Tesla's technology, I don't think Tesla is going to save the world. I'm not an Elon Musk fan, I find him very difficult in a number of different ways. That's as much as I'm going to say, but I find-

Jeremy: I don't think he listens to this podcast, so don't worry about it.

Paul: That's fine. But I find a lot of people within tech look at techno utopianism, and let's call it that because I think that's a pretty simple way of it. That technology will save us. The more that I look and the more that I look at what is happening in the world and the speed the technology is evolving and what we need to do in the speed that we need to do it in terms of climate change, I don't think we're going to get anywhere near fast enough technology evolution. We have to do something else. We can't just expect technology to catch up, fix it, and just for us to carry on.

I think technology needs to learn to find its activist voice. I think we need to be activists against those organizations who are not doing enough, who say they're green and are not, and in this, Amazon, I will call this out. Much as I love their serverless technologies and the people who work there are brilliant, the wider organization I think is not doing a good enough job in terms of its green credentials. That hasn't been good enough from my point of view. I want them to do better. Because actually, I think a green Amazon would be a great thing for the world and I think they can do it. I have seen Amazon and what it does when it's amazing, and I think if they turn themselves around and actually did the green thing properly, then I think that the hope for the future would be significantly higher.

That is why I want Amazon to change is because I think they have the power to be a force for good, and I think they're not doing that at the moment enough. That's one. I think find the activist, find the place that you within tech want to change and go and change it. Because I don't think we have anywhere near as much time as people think. I think your career in tech in 10 years time will not look like the career in tech that you think it is now. We are in a completely changing environment. Depending on what happens in the U.S. in the next few years, the world is changing around us. We are in an inflection point that I don't think many people are aware of. This is just slightly rounded conversation, but in terms of people in tech, it is very much, don't just let it happen around you. If you want to make a change, go and do something about it.

To that is, green your data centers. It is go and find an organization to go and get involved with. It is find your political voice, whatever that is. It is inform yourself. It is get involved with all of the things like all of the black lives matter protests and all of that. I think it is important to find all of those areas and get involved, because once you start getting involved, you will find other areas of intersectionalities that will then help you look at it and see all of the wider issues and all of the systemic problems. And then you will start to go, "This is huge, and we've got to do something about it." It just becomes a bigger problem.

I think the tech industry needs to, and I know this is a massive rant, but speaking as someone from the UK, I look at the U.S. and I see the U.S. tech industry, and I see the money in the U.S. tech industry, and I actually think that I don't think we need the technology, I think we need the money and the brains to go into other areas. I think we need them to stop thinking about how to make technology, I think we need to start them thinking in how to start being humans again. There's a very big difference in that. I wrote a blog post about it, it was quite fun.

Jeremy: Well, I don't even know where to start to respond to everything that you just said, other than I agree with you. Brilliant. That was great. The funny thing you mentioned about Amazon just not doing as good a job as it could, I think that is again, partially an infrastructure problem, in a sense, partially a priority problem, partially this thing where I think it's FedEx decided to stop delivering their packages so they had to speed up their own fleet for delivery. Why not electric delivery vans? The technology's there. Maybe it would've cost a little bit more, maybe it would have taken a little longer to roll out, but those little things, I say little things, huge things like that that they could have done that would have had a massive impact.

Again, singling out Amazon is probably not fair. I think every company in the world, the vast majority of them are under that same thing. Any incremental change is good, but I think you're right, we just need massive systemic change at this point. Otherwise, that clock is ticking and we're going to run out pretty fast. I guess another question though about that is big companies like Amazon and Microsoft and IBM and even Oracle and some of these other ones, they have resources. Big tech companies, Facebook, Google, the Twitters and things like that, they're building their own data centers, or of course they're building data centers that are shared, that other people can use. Are they more likely to become green or have more efficient data centers and be more up to date because they're not...

I think about when I rented a rack in a co-location facility when I had my web development company. I had one rack with the power coming in, whatever, I had servers in there that were six or seven years old. Because you're like, "Some of these old websites running on that thing, I'm just going to keep it running, and it was probably terribly inefficient." It's a massive investment for companies to continue to upgrade their servers and continue to make their data centers or their on-premises locations green. Is that something where maybe there would be a nice point where moving to the cloud actually would be the smart move from a green efficiency standpoint?

Paul: It's a massive question.

Jeremy: Sorry. Well, you had a massive rant before that. I'm just getting you back.

Paul: In the end, I know the joke is cloud is just other people's servers and all that kind of stuff. It's always underneath it. There's just servers and there's just servers. But I think that trying to make these servers more efficient, trying to make these data centers more efficient, there is still constant churn. We don't keep things efficient. Two, three years down the line, the server that you were using is not efficient. Six years, seven years, it's old. You don't want to be running stuff on there, you want to be running stuff on something that's efficient and new. Actually, there's an enormous amount of e-waste in terms of the data center industry. It's not straightforward.

The conversations around all of this are not straightforward. I think everyone needs to start thinking about moving to the cloud simply because we need to be reducing our impact. If you're running stuff, I think it's important to be able to go, "Actually, we need to be able to reduce the amount we run." But that means, understanding how that cloud, that you're choosing to work with, is working in terms of its sustainability. You can't just go, "We'll move it to X cloud, or Y cloud or Z cloud, or whoever it is, but we'll trust them to do the right thing."

You've got to still have that relationship. You've got to still be able to go that cloud, "You, Mr. or Mrs. Cloud person, you've got to tell me, are you using green electricity? Are you using renewables? How are you disposing of everything? What is your supply chain?" I think that conversation over the next few years is actually going to become a much more common conversation. It's going to become more important. You are not going to be able to get away with, "We just run efficient data centers." That's not going to be the standard and reasonable response. That's going to be a table stakes. Green data center will be a table stakes conversation, and the best practice will be, "Well, we're actually running 100% renewables and we're putting more into the grid, and we're being as good a partner as we possibly can. And all of that. We haven't got diesel generators, we've got batteries."

It's all of that conversation that I think comes back to. Maybe we will end up not using certain companies because their data centers are not green enough. Maybe that is where we end up, that actually societal pressure actually pushes these companies to do better. But I don't think we're there yet. I think we're probably a couple of years away, two, three, or four maybe, away from that.

Jeremy: Right. I think that that conversation about e-waste is probably really important too. Because now I'm wondering, where did my Pentium 166 megahertz computer go from 25 years ago, and my 32 megabyte RAM module? Is that a landfill somewhere? Do they melt it down? Is that in my new MacBook? I have no idea. I think that's interesting what you said, that again, you want to make sure that the cloud infrastructure is efficient and it's green and you want to do that.

But I think there's a bigger conversation around this as to say if I was to buy, I don't know, a million dollars worth of Dell servers, brand new, highly efficient servers. I throw them into my data center, I'm running my own power, I have to buy my own batteries, I have to have my own backup, I have to have all of the waste that's involved with that. Even if I make it incredibly green, two years from now, there's faster chips, there's new servers, maybe I can swap out the chips in the servers, I don't have to get rid of the actual metal, the casings and things like that, but I have to make another massive investment in order for me to then upgrade.

Whereas if I put my stuff in the cloud, even if the cloud is not as green as I want it to be today, maybe tomorrow they're a little more green. And then the day after that they're a little more green. And then five years from now they're 100% green and it didn't cost me any more money other than what it cost me to run my workload. I go back to this, and I don't want to offend anybody here, but climate change deniers, people who don't believe it's happening, and I know we had to change it from global warming to climate change because they're like, "Oh, it's actually getting colder in certain areas." You're missing the point.

But my thought here is to say, is there anybody who disputes the fact that if you dump oil in a pond, that that's a bad thing for the environment? You don't want to drink that water, you don't want to swim in that water. Can we agree that pollution is a bad thing. Whether it's making the earth hotter or whatever, can we agree it's a bad thing. If we can agree on that, then that's a good thing. That gets us where we can think of maybe the moral part of this. But let's take it back to a more selfish level. If you are a business and you can implement green things that are going to cost you less money over time, I use serverless. I spend, I don't know, a 10th of the time writing code than I did before. I have teams that are smaller so I don't have to have 20 people to write an application. Now I can have five.

The impact of that on the planet is huge, but also, the impact of that on my wallet is huge. I'd love to get your perspective on that because I think that there is a snowball effect, that as you make one small change to increase efficiency, to reduce your carbon footprint, to do these things, save yourself some money, but then the impact of that is huge.

Paul: I think that you make a very good point. Going back to 2015 when I was CTO of a startup and it started off with me building serverless stuff with AWS Lambda and DynamoDB, we were doing half a million monthly active users with a backend team of three. We weren't building complex servers and putting servers in a data center and having to run that and having to think about automation and dev ops and all of this kind of stuff, and having to think about end of life of whatever and all of that kind of stuff. We were literally able to change with whatever we thought we needed to worry about, and we could have done it in seconds if we actually needed to. We could have changed the way.

That was the architectural choices that we made in terms of the technology and moving to the cloud. I think that makes a huge difference. We were able to do things like be all remote from the very beginning. Being remote means that we are not traveling to an office, and when we're not traveling to an office there aren't the carbon emissions from traveling to an office. Yes, we're using electricity, but we can all use green electricity at home. That's a heck of a lot easier than having to think about whether or not our train or our car, or everyone buying electric car is ridiculous because that's quite expensive.

All of those knock-on effects of having a smaller team, having a remote team, thinking about smaller workloads, which takes less electricity, all of those additional elements, it's not just making a technology choice. As you say, it is that continual change. That comes back to the original point that I was making, I don't know, an hour ago was it? A while ago anyway. Which was that-

Jeremy: It's been awhile.

Paul: It's been awhile. But that constraint, that you put that constraint on yourself of how can I, in making this application, in making these decisions, reduce my impact as much as possible? Things like going remote, it was 2015, it was a no brainer. I didn't want to be traveling to wherever I wanted. Why do I want to go to an office? There's only a few of us, we can all sit and talk over Slack. It was very straightforward. It was irrelevant. It was like, we don't need to be in an office, it's unnecessary.

You reduce the emissions in that way, and then you've reduce emissions in another way, but we'll just use the cloud and then we'll reduced emissions by, well, we'll stick constraints around the amount of code we use. You just start to put those constraints in place and then the snowball effect is that actually we had a tiny, tiny, tiny amount of money that we were spending with AWS, but that we also had a huge amount of people that we could reach. If people stopped using it, the budget, we were spending almost nothing. From a cost perspective, my CEO basically never bothered talking to me about the budget. It just was irrelevant. It was like, "It just costs us what it costs us."

I remember a conversation and he said, "Well, it costs us," I think he said it was $1.00 to acquire people, and he goes, "How much did it cost you to run this platform?" I told him the number and he goes, "Well, that's less than the $1.00 a person then isn't it?" I went, "Yes," and he goes, "Well, that works then doesn't it?" He's sitting there going-

Jeremy: Right. The numbers work out.

Paul: ... "Okay, that's fine." It was that ridiculous. When you started to work it out, it was that ridiculous the way that we were running our company, and it was so positive for the company that he just didn't bother asking us how we were doing it. It was like, "You do your thing, just go on with it, you're fine." I think that's where the conversation gets lost in terms of technologists. When they start talking about data centers, especially when you start talking about containers. You want to sit there and you want to go, "I just couldn't care less." But I think one thing that is worth saying is that being serverless doesn't mean you never use a server. I think it's one of the conversations that a lot of people go, "Well, it means you're always using functions. You're always using the smallest... " It's not that.

It's about using the most appropriate technology, and using it so that it simplifies your application so that if nobody's using it, then you get the scales to zero if possible. But there are times when that's not possible. You have to use something, so you use the most appropriate technology for the right reasons. Sometimes that will mean you need to use something like an EC2 instance. It just might mean that that's what you have to do. It's just that, you just have to understand the constraints, understand why, and then understand that if you do that, you then have other concerns that come along with that.

That's what being serverless means. It's not about no servers, it's about understanding the constraints, understanding the platform, and then building your application. Understanding that you will then need to manage that server in a different way, and understand the application life cycle and retire that server at some point or change it. It's a whole lot more holistic than just, "Oh, we just-

It's a whole lot more holistic than just, "Oh, we just use FaaS." So I think there's a whole conversation there, but it does, for me, come all the way back down to, how can we reduce our impact? Because one of the other things is if you can reduce the impact in terms of operations, and if you can reduce the impact in terms of maintenance. No maintenance over time, significant carbon impact, because you're not having to go back in and change stuff. You're not having to make more code. You're not having to test stuff, all of that. The knock-on impact is huge.

Jeremy: Right now, if I can summarize what you just said, what I got out of that was that Kubernetes is bad for the environment. So I'm just going to use that as the takeaway-

Paul: And go with that.

Jeremy: And go with that. So the other thing that you mentioned though, this idea of working remotely, maybe a benefit of COVID, which I don't think there are any benefits of it, but the experiment of, can most people work from home? Now, I fully understand, now again, get back to the Black Lives Matter and under-represented communities, things like that. They are just not possible for people to work from home. Eventually kids are going to have to go back to school and teachers are going to have to be in the classrooms and you're still going to have to have your janitors, you're going to have your baristas, you're going to have your doctors.

You're going to have that whole range of the economy still needs to exist. But a lot of people in tech are very privileged and the ability to work from home is something that, I mean, I know a lot of people say, Oh, I want the socialization, and that's fine. But think about it this way, I leave my home, I don't turn my thermostat down to zero degrees and let my house freeze when I'm away from home and at the office, I still have to heat my house.

So I'm heating an empty house, even if it's at a lower temperature while my employer is heating a 10,000 square foot office. And so we're heating two places, we're wasting a lot of energy and I spend all that time traveling. I spend, well, not time-travel, but I spend all that time traveling to work, commuting to work.

And then, I wish I could time-travel, that would be nice. It could probably make a few things better, but that's the Back to the Future reference. But if I could, the amount of gas that I put in my car or the train that I ride or whatever it is, those efficiencies there by having to power two places when you don't necessarily need to. And I think what we're going to see, and I'd be curious, your thoughts on this, but I think what we're going to see is companies starting to build much smaller in-person offices, right?

They could be regional, but I mean, having an office building that can fit a thousand employees just might not be necessary anymore. It might be that having an office building or an office space that has room for 50 employees and a few conference rooms and things like that where certain people can go in on certain days.

But for the most part reducing that overall footprint, I think that's going to have a huge impact. And I think companies are going to need to start thinking about how they reduce their footprint because I'm pretty sure that at some point the hammer's going to come down and these companies are going to have to rethink their entire global supply chain.

Paul: Yep. And I agree with that completely. I think we are going to end up and I think tech is probably going to be the lead for an awful lot of this. We are an incredibly privileged group, that we are essentially all working on computers. We all work... Basically need a computer and an internet connection and we can do our jobs. We don't even need to be at home. We can be at a coffee shop, which is not even fair for most people. It's like, well, I am working from the coffee shop all day. Well, stuff you, everybody else.

Jeremy: While you are fully garbed in PPE trying to treat patients in the emergency room, exactly.

Paul: We will happily sit here with our latte and a Danish, it's like, well, thanks.

Jeremy: And we gave you TikToK and Facebook and Twitter, so stop complaining. Yeah, totally agree.

Paul: We should recognize that privilege. And I think we should be at the forefront of trying to make society better. And I think we should understand that. That includes recognizing that moving away from offices is going to have a significant impact on the economies around those offices. So there are going to be shops and coffee shops and things like that, that are going to have to change and grow and move. And so I think we do need to recognize that impact is going to be huge.

But I also think that we need to think about tools and using serverless to build those tools is very good. But I think we need to think about tools around how we make teams better, how we make organizations work better. And it's not so much just the video. The video is lovely, but it's a brief conversation I had with someone yesterday was like, are we doing the storming and the norming and all of those kinds of things anymore?

So, are we actually creating teams? Normally when you get together and you're basically all together in an office, you have the fights and then you get all together, you then get back together and you figure out how to... I'm not sure that you do that over a video call anymore.

So, there's all these other things that we're going to have to start learning how to do differently. And I think technology and technology companies are going to need to be at the forefront of all of that. And so I think we're going to need to learn how to do all of those things a bit better.

But that will have an impact on carbon. And I think we'll be doing things like smart homes better than just having a nest controller or whatever it is in your home.

I think we'll be doing things like repurposing office buildings into something else, possibly homes, possibly something else. Because I think that there are going to be a decent amount of empty buildings and I think our city centers may change. This might take 10 years, but I think we're going to see some real changes in the way that cities work.

But this is all a much bigger, broader conversation. Coming back to the serverless side of things, I think we are going to have to... I think we could well see a change in the way that companies approach technology as well. I think we're going to see that we can't just throw money at tech. I think companies are going to want to see returns. They're going to want to be able to manage and understand how technology is used and where it fits. And that I think serverless is probably better placed than it realizes for that.

Jeremy: Right. Yeah. I totally agree. All right. Well, we've been talking for a very, very long time, and I appreciate the time that you have, if you have a few more minutes, though, I do have some questions for my Serverless Chats Insiders list. And again, if you want to ask questions to guests like Paul, go to serverlesschats.com/insider, sign up for the list, and you can do that.

All right, so first question is from Eduardo. And he asked the question, does serverless contribute to climate change in a good or a bad way? And as cloud providers are investing in more green infrastructures, what does that mean to users?

Paul: So, the difficult thing is that it is probably a net good in terms of serverless, it is probably a net good, because I suspect... It is impossible.

The reason it is hard to answer this question is, and I have asked this multiple times over years now, it's actually very difficult to build an equivalent application, not serverless and serverless, and then work out what the actual electricity usage is. And then, in terms of how do you basically calculate load? How do you work out how much electricity is being used and for how long and over what period, and it's actually incredibly difficult to do.

And also, take the pet shop example, which is the one we all grew up learning, what the heck it was. And, it's like, how do you build pet shop in serverless? It's like how everyone would do it differently. And so what is the reference implementation? There isn't one.

So we don't know, but if you take the idea of function as a service, and if you take the idea of just using... Your functions only do one thing, single responsibility principle. If you take the idea of not running relational databases, but running something like DynamoDB, some completely different idea.

If you take that, I suspect that you are probably net positive. It's a net good. Do your own research, have a look at yourself. But that's what I'm going on. Because I think that the only real good proxy I have is the cost and the majority of serverless implementations are reduced cost. And it's a pretty good proxy for the amount of electricity, somewhere along the line.

Jeremy: Right, you would think that if the cloud provider is charging you less than they're paying less, which means you're probably... And energy and electricity is probably one of those things. I mean, the other thing I would say about that too, is that, obviously, you have a very large server that is running all these multi-tenants, that is running firecracker or whatever it's running under the hood there. The hypervisor in order to spin up these little containers that are your Lambda functions or whatever your FaaS of choice is.

Obviously that server's running, the whole server's running. Can't shut down parts of the server. There's probably some efficiencies obviously to the CPU that if the workload is low, then it's not burning as much electricity, but you still have to have a lot of servers turned on to be able to handle the spike in load.

You can't just have somebody saying, Oh, wait, now we've got every server's used. Let's turn on another one. Capacity planning is still a thing that needs to be done in these data centers. So even as efficient as serverless might be from your implementation, I think at least what you're saying is even though those servers are running full-time, I am only using energy when my part of it runs. So I'm at least reducing it a bit. So I agree with you. It's a hard question to answer, but I-

Paul: It's an impossibly difficult one.

Jeremy: I'd like to feel better-

Paul: What was the second part?

Jeremy: The second part, was there one? Oh, what does it mean for end users? But I think you've sort of answered that.

Paul: I did yeah, I did answer that.

Jeremy: All right. So, Mark asked another question, and this is, again, going back to probably putting AWS in the spotlight, but AWS has the worst credentials for powering their data centers of the big three. And this is what Mark says. They also have the most comprehensive and power efficient serverless offering, presumably even more so once we have Lambda on Graviton2 which lets us use less energy overall to run our apps. So, how should we look at the trade offs? So Amazon, maybe not as green for the data centers, but a much more efficient serverless offering. How do you make that trade-off?

Paul: Yes. And this comes back to the conversation around... It's a difficult one. And again, this trade-off is complicated. I think if it's complex to build an application it's going to be complex to maintain and complex to manage at the end of the day.

So, if you are finding that it's difficult to build, difficult to take forwards, you need a bigger team, whatever it is. I think the amount of emissions you create doing that is probably, in terms of the people, in terms of the amount of maybe get meetings, the conversations that you have around it, the longterm effects are probably going to be higher than if you take the more efficient offering and then choose to do your own offsetting, take an approximation of an offset in some way, shape or form.

I think that's the way I approach it anyway, which is to essentially say, yes, I know they're not perfect because no company ever is, but yes, I know AWS is the worst of the three and I've looked at all the offerings very carefully, but I look at it from the point of view of, but I know how this serverless offering works. I know how efficient it is. I know that it will give me the best offering for my users.

So the likelihood is that I will spend less time in support. I will spend less time trying to work around the things that I know aren't particularly good in all the things that aren't particularly working in the way that I want them to work in other platforms. Not to say that you couldn't work around the constraints and the other ways, but I can see that I know how much less effort it is in AWS than it is in the other two clouds to produce a good application and to manage it and maintain it.

And I've tried to build in all of them. So, I see that efficiency in terms of time as being a carbon impact as well. And that's the trade-off that I see. So, if I didn't see that, I would be moving over, does that make sense? So, that's how I look at that trade-off.

Jeremy: Yeah. Makes sense to me. All right. So, one more question. This one is anonymous and I'm assuming, because it's such an easy question that they didn't want to be named when they ask us an easy question. It says Paul's articles on serverless have been thought provoking in a good way about how he sees and has seen service evolving with the industry adopting serverless more and more. I wonder what he sees in the future of serverless and what he sees as the answer to what is serverless 2025? So simple question, see if you can...

Paul: Thank you very much. Where do I see serverless evolving and what do I see serverless in 2025. Well, thanks for that, whoever sent that in. Oh, that's so, okay. I see tools, I don't think we have the tool sets yet to properly build serverless applications. I see we have an awful lot of people building web applications on top of serverless. I don't see we have a lot of applications building the ability to simply connect.

So at the moment, I think we have things like SAM and CloudFormation and all those bits that allow us to automate, but they're complex and they're not straightforward. And then we've got other tools on top of those, which your serverless frameworks and your others, I can't think of the top of my head.

And you've got all these, but I think they are not yet evolved to where we need everybody else to get to. And I think until we have that next evolution of tools, and I think I haven't seen it yet. I haven't seen that tool yet that allows someone to come along and go, I just want to do X. And that it's like three or four lines of code. And I think Begin and I think Architects and all of that kind of stuff and all of that, I think has some of those elements, but it's not quite there yet. I think there's a little bit missing a few pieces, but I think it for what it does, it's brilliant. So I think we're just missing some of those tools.

I don't think we'll see a multi cloud tool. I think the main three clouds in terms of serverless are so different in the way that they've approached function as a service, data and all of those things. I think we will see that they will diverge. I don't think compute is the key anymore.

I think we will also see serverless as being more about data. I think we will see it less about compute, and more about data. And I think we will see the divergence in the clouds being about how people store data, how people use data, how people build pipelines around their data and how people move data at speed, get queries out to their data. I think that's where the innovations will come. So basically building the building blocks, I think we'll get more of that. And then the keys around the data.

So, we will see less relational databases. Yes, I really can't stand them. I think we will see some really, really fascinating data tools appearing. That's where I think 2025 is going to be. So when someone asks what serverless is going to look like, I think we probably need to be looking at the... I think what people are going to be really scared of, is it's going to move away from that three layer, which we still think about the three layers, your client, your server and your data and all that kind of stuff. We still think in that world.

I think it's going to be blown away. The data layer is going to all of a sudden learn to compute and it's going to learn to do all of these clever things that we didn't realise it could do.

And then all of a sudden these data scientists going to walk in and essentially become very, very much more key for everything we build and that's where serverless is going. So when we get to that level, when we start to go down that road, and I think some people are there, but not everybody. I think that's when we start to see what, away from microservices, that I think is not where serverless is going, that I think is where we're heading. So anyway, that's my broad thinking on where 2025 is going to be. How's that?

Jeremy: Yeah, no, I think that's great. And I agree, I think that one of the most difficult things about serverless right now, despite the lack of tooling in many cases, is the fact that everything is a primitive, right? And you've got a Lambda function that needs to connect to SQS and it needs to connect to Dynamo or whatever. You have this whole, all these little tiny primitives.

And what we really need is, the building block needs to be more abstract than just those tiny things. And as those patterns emerge and we are already seeing a ton of them, and as we start to encapsulate those into things like the CDK and into serverless components is something that is an interesting thing that they're doing. Like you said, Architect, and some of these things that abstract away some of the more complex individual building blocks and make them much larger pieces that you can put together, that's interesting.

And then the data stuff you're already seeing this, you're seeing Adobe, you're seeing Salesforce and some of these other ones, I think Twilio is doing it now too, where you can run serverless functions on their platform in reaction to data or in reaction to the events that data changes in your system.

And that's going to be really interesting because then that's hands off the compute and then the data handling piece of it gets smarter. But, I mean, anything else? I mean, we've covered quite a bit here, but any last words on serverless and going green?

Paul: No, just more serverless and go green. I think it's just very much a case of get on with it. Become an activist. I think it's what we all need to be doing.

Jeremy: Awesome. All right. Well, Paul, thank you so much for this conversation. This was excellent. I had a great time. I hope you enjoyed it. Hope the listeners enjoyed this. I think this is a little different than what we normally do, but honestly, this is stuff that I'm passionate about. I know you're passionate about it. So anyways, if people want to learn more about what you're doing and learn more about green energy and how they can make an impact, what's the best way for them to contact you or find out more?

Paul: Just Twitter. I'm on Twitter @pauldjohnston. And just also find me on roundaboutlabs.com. Just my details are on there as well. So, either way is absolutely fine.

Jeremy: Awesome. And then your blog is medium.com/pauldjohnston as well.

Paul: Yeah, that's correct.

Jeremy: Awesome. Thank you again. I will get all this information to the show notes. So if people want to go check out the show notes, they can find the information there. Thanks again, Paul.

Paul: Thank you for having me.

This episode is sponsored by New Relic and Homeschool

View Details

About Paul Johnston

Paul Johnston is an interim CTO, CTO and strategist who has particular interests in serverless, cloud, startups and climate change. Formerly, Paul served as a Senior Developer Advocate at AWS for Serverless and CTO of multiple startups, including one of the world’s first serverless startups. Paul is also a co-founder of ServerlessDays.

  • Twitter: twitter.com/PaulDJohnston
  • Medium: medium.com/@PaulDJohnston
  • Project Drawdown: drawdown.org/
  • Roundabout Labs: roundaboutlabs.com/
  • Leading Edge Forum: leadingedgeforum.com/
  • White Paper: The State of Data Center Energy Use in 2018
  • IPCC Special Report: Global Warming of 1.5 ºC
  • Blog post: To fix Climate Change, stop being a techie and start being a human

Watch this episode on YouTube: https://youtu.be/SI2-WU_0zgs

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Paul Johnston. Hey Paul, thanks for joining me.

Paul: Thank you very much for having me.

Jeremy: So you are a consultant through Roundabout Labs and a research associate at the Leading Edge Forum. So why don't you tell the listeners a bit about your background and what you have been up to lately?

Paul: So, yeah, background's a bit confusing. It's always a little bit strange. I don't have this whole 14 years at any one big company or anything like that. I spent many years in tech, 20 years, working with various different startups from my own business, that kind of thing. Then I worked in a startup in 2015 that effectively started using AWS Lambda.

So this is where the serverless comes in. And I was one of the first companies to start using Lambda in any kind of scale in a startup as a kind of first principle. And then I went from there to using it in that startup in 20 countries. Went a bit mad, a tiny, tiny budget from AWS. And I was like, "Well, this kind of worked so I'm going to keep doing it," and started telling people. AWS took notice, gave me a job, that was quite fun. Was a senior developer advocate for serverless at AWS for a while.

And then didn't stay there all that long, but it was really enjoyable while I was there. And moved away from there to go and do some consulting, which I've done since 2018. 2018? 2018. And then from there...

Jeremy: What year is it again?

Paul: Honestly, this year has gone on for a very long time.

Jeremy: Right, right.

Paul: And since then I have done some consulting in various different projects, tech projects. But one of the things I've done is worked on working out how tech and climate work and how they intersect. And one of the projects I've been working on is a research project for the Leading Edge Forum, which if any of you know Simon Wardley, that's the organization that he works for.

And I've been working on a project to look at how climate change is going to affect business over the next 10 years from a tech angle very much, so from a data and a tech angle. And just trying to see what lessons we can learn and what things are going to be coming up in the future. So kind of many and varied, shall we say?

Jeremy: All right. Well, listen, I have been wanting to get you on the show for a very long time, because I think this whole climate change thing is hugely important. I have two young daughters. I think about their future. I think about the junk we pour onto the earth, the pollution, the amount of carbon dioxide we're creating.

And one of the things that I think really attracted me to serverless in the beginning was not just, you know, obviously not having to manage servers, which is great. But this idea that maybe by sharing tenancy on a big server and only using the compute that we needed to, I was thinking in the back of my head, I'm like, "Well, maybe that reduces the amount of energy we use."

And so I know you have dug into this tremendously, and I mean, you're an expert on this stuff. And so I'd love to go through all these things with you, just get your insights, get some thoughts on this stuff. We can talk about serverless and some details of serverless as well. I'm sure that's what the listeners want to hear, but I love this idea of going green with serverless. Because I think it's hugely important. I think it's a step in the right direction.

But maybe we could start and just, or start by saying, how does serverless technology compare to traditional technology or traditional servers when it comes to green computing?

Paul: So it's a very, very good question. It's almost impossible to answer in some ways, but it's really, really easy to answer in others. So one of the things you want to look at, first place you want to start is, well, effectively you want a definition of what serverless computing is.

Paul: So let's just kind of take function as a service is kind of the base enabling technology, shall we say, for most serverless computing. Because I think most people will kind of see serverless and they'll go, "Right. What does that mean?" And so you want to drop it down to something that is kind of tangible.

So you want to talk about function as a service really as being the base enabler, because serverless for me is about business value and getting as much as possible out of your technology in terms of applications and all of those elements.

And so I think when you talk about function as a service, what you get is you get a pay for what you use. So you know that you are using as little electricity for your application as you possibly can. And so what you're trying to do is go, "Well, I want to be as green as possible." So being as green as possible means actually reducing as much of your usage as possible.

That's essentially what we mean by being green is actually reducing and actually using as little, as close to zero, in terms of compute as we possibly can. So what does that mean? How do you do that? Well, in terms of building an application, don't build the application to start. Just don't build the application at all if you possibly can. If you can build it on the basis of lots and lots of caching or not running any servers at all, then great, do that.

If you can do it on the basis of only running compute when you absolutely have to, then great. If your application can scale down to zero and it literally can, nothing is running if nobody's using it, that again is... That's the kind of thinking that goes into being green and being serverless, which is why serverless is something that for me works really well alongside an environmental conversation. Because it's not just about what does the techno- how does the technology work?

It's actually, well, this approach allows me to say, well, it gives me business value. It gives me environmental value. And actually when you come down to it, it just works out as better common... It's more common sense when you actually try. And when you build the application, you come out to the other end, it's usually a better application and easier to build going forward as well. So you've got all of these things working positively. It just seems to... It seems to work out better.

Jeremy: Right, right. Well, I love that idea of deciding whether you even need to build the application. That's a really good way to just cut your carbon footprint is say, "Don't even build the application."

That's super easy. But unfortunately there are applications that still need to be built and we still need power. I'm sitting here with my 16 inch MacBook Pro drawing 100 watts of power with the fan spinning a million miles an hour because my CPU is going nuts and just generating a ton of heat.

Obviously in the data centers, you have millions of computers that are generating a lot of heat, that are drawing a lot of power. And like you said, every time there's some execution that needs to be done, that CPU has to spin up, which means we need more power, caching, SSD drives, things like that - very low, low, I guess, wattage types of equipment like that obviously draw less and there's ways for efficiencies there. But there's only so much of that we can control. We're still going to need to do compute.

So what about data centers themselves? What type of impact do those have? Because I think it's more than most people think.

Paul: Yeah. So data centers are, let's put it this way. They're a huge impact. And they're a significantly greater impact than most people realize. And I wrote a white paper in 2018 with a friend of mine, Anne Curry, and she and I did an awful lot of digging.

Now when we did the digging, we found out an awful lot around people saying that it was going to be, there was going to be a tsunami of data, which meant that data centers were going to grow at a huge rate. And we found that it was going to be five times more data was going to be going through data centers and using electricity.

That that may or may not come to pass. That's a prediction. So there are some who say that is, and some who say that isn't going to happen. But it's a huge amount of electricity, and it's a huge amount in terms of carbon footprint. And some estimates go between about 1% and some say up to four to 5% of electricity is...

Jeremy: In the world?

Paul: In the world.

Jeremy: In the world, right.

Paul: Yeah. And if you take that in terms of carbon footprint, we're talking in the magnitude of somewhere between something like a quarter to two to 3%. It's a huge amount of actual carbon emissions that go into data centers.

Now there's an interesting conversation about what constitutes a data center at this point, because some people go, "Well, it's only hyper scale computing." And you're sitting there going, "Well, actually that limits what you mean." And it's a whole other conversation.

But I tend to put that it's actually around 2%, probably a little bit less, which puts it somewhere in the same region as aviation pre-2020.

And so you end up with this... Even if the numbers aren't right, they're in the same kind of ballpark. And it's really actually very difficult to know how much electricity data centers actually use, because most of the providers don't tell us. And a lot of the people who do the digging around the numbers are doing things like going and counting the number of diesel generators outside to work out what the backup power is, and therefore how much backup power you need for the size of... For a type of data center and what that would...

So they're not, nobody really knows, is the actual answer. But it's actually significantly more than most people realize. And while there are some that go, "It's really, really terrible and it's huge," and there are some that go, "It's nowhere near as bad as all that and it's coming down," and it's probably somewhere in the middle. It's really not good, however we skin this cat.

And so we end up with data centers being a problem, and they are an issue we need to solve. And actually, if you look at someone like Google, and if you look at someone like Microsoft, they're both trying to do an awful lot in this area.

Jeremy: Yeah. Well, that's what I was going to say. So I mean, you do see press about this and you do see companies talking about this. What are they doing to try to become more efficient?

Paul: So if you have a look at, let's take Google for example, because I think they're probably the best example of good practice in this space. They are trying to look at their electricity usage. They're trying to offset on a, I think it's an hourly basis. It might be an half hourly basis. They're trying to offset their electricity usage usage with purchases of renewables in the same grid. So in an electricity grid, they'll buy. If they using X, they'll try and buy X amount of electricity from renewable sources for the same hour in the same grid. So there are some places in the world where that's not possible. So some of the, I think it's the Chinese data centers. I think it might be Taiwan or something, they can't do that because there aren't the renewable sources available. So they have to do it in another way.

But a lot of their US data centers, they're definitely able to do that for the majority. So that's a really, really good practice, but they're still not able to do it completely. So they're still offsetting in other ways. And offsetting here is, we use... You use the electricity in the grid, which may not be completely 100% renewable. And then you buy an amount of renewable energy from somewhere else to offset the fact that you've used bad energy, carbon emitting energy. That's all an offset is.

Microsoft again, they've got interesting practices here. I love this topic, by the way. They use renewable energy for something. I can't remember the exact numbers. It's something like 40% of their data center usage. So that's pretty good. And then they offset the rest, but they also have an internal carbon price. So if they emit, then they have to pay internally. I think the last time I looked, it was 12 or $20 or something like that. It may have gone up since then. So for all the carbon they emit, they actually price it internally, which means that if they move to renewables, it's actually cheaper for them.

So it's actually really, really interesting things going on at Microsoft, and I think that's an inte resting model to look at. AWS...

Jeremy: It's a good way to play with... A good way to play with your balance sheet is to do that.

Paul: Very, very good. Really good practice. AWS, they have four green regions. So there's Frankfurt, Ireland, Montreal, and Oregon. Really good, except for the fact that they don't tell us how green they are or how they're offsetting or what they're doing there, but they just tell us they're green. So it's a little bit unhelpful. And they don't tell us how green any of the other regions are. So you can't actually offset your AWS bill, which doesn't really help if you're... I'm sure if you're really big, you can go to them and say, "How much electricity are we using at these different points, and can you tell us how to offset?" And they'd probably do it, but you have to be kind of... Hundreds of thousands.

And if you're serverless, the likelihood of you doing that is very, very low. So I think you see that there are interesting conversations going on in that space around how best to start thinking about the carbon footprint of data centers.

Jeremy: Right. So it's not just the carbon footprint and the offsets, and some of that stuff is... I think those are great, right? I mean, you're still using dirty energy, I guess if you want to call it that, but you're subsidizing essentially green energy on the other side.

But is there a difference or a distinction between sort of the idea of just offsetting your carbon footprint saying, "Hey, listen, we're going to run up the meter, but we're going to pay for it in another way." Is there a distinction between that overall environmental impact and the efficiency that can be created in a data center?

Paul: So I think efficiency is a really difficult area for a lot of people. Because a lot of people like to go, "Oh, we've bought better servers and they're more efficient, therefore it's better for the environment."

Jeremy: We installed LED light bulbs.

Paul: Exactly. We've done our bit. There's the thing called Jevons paradox. It's a very old economist who's long dead who basically worked out that if you made things more efficient, made energy production more efficient effectively, then people would use more energy.

Jeremy: Right, exactly.

Paul: So, you basically make it cheaper to do something, people will spend more money on it and therefore they use more in the end because it's like, "Oh, it's fine. This is easier."

Jeremy: And not to not to interrupt you, but that's actually the same argument people make about serverless, where they say serverless is faster and cheaper. That just means you'll build more with it. Right? So you're not actually spending less. You're just doing more, which is still great. But back to your point.

Paul: Yeah. And it's exactly the same with cloud. When we all had to buy servers and stick them in a data center ourselves, that was hard work and it was difficult. And then along came the ability to buy a virtual server. And then we literally just bought virtual servers like they were water. And then it was like, "Oh, it doesn't matter. It's tens of dollars and it doesn't matter."

And now we just go like, "Right, we can build whatever we like." That's efficiency. And the fact that a data center is more efficient does not make it more environmentally friendly. It just means that everybody wants to use more of it. And that I think is a... I think it's a fallacy that if we make things more efficient, it's more environmentally friendly. Yes, we need to make things more efficient, but we need to make things more efficient and we need to look at reduction of our carbon footprint overall.

And I think that reduction comes first, and then efficiency leads to reduction. Not efficiency is just the overall goal, because that doesn't lead. And that I think is a personal concern as well. Because most people talk about, "Oh, it's fine. I'm becoming more efficient. I'm personally, I'm doing all these wonderful things."

And then you sit there and go, "Well, that's fine. I I'm doing my bit. We're more efficient. We don't do this. We don't do that." And then you're like, "Well..." But then you use more heat because your house is... You just do and you just use more because you've made yourself more efficient, or you...

A good example, I think is, most people would understand right now is in the middle of having spent three months in our homes, which is when we're recording this, you all of a sudden realize that you can have packages delivered to your house.

Jeremy: Right, it's so easy.

Paul: It's so straightforward. The efficiency of packages being delivered to your house. Stuff, you just go, "Well, I don't need to go anywhere. I can just order it. I can just order whatever. Well, I just need more of this. I will order it." And it's just the efficiency of stuff. Doesn't mean you use it less. It just means that you use... You just find it more and more. You will use it more. And I bet everybody who's listening to this is going, "Oh yeah, I've ordered far more stuff in the last few months than I have in the previous two years." And that's how it works.

Jeremy: Well, I also love the fact that you get that Amazon box delivered and it's a box within a box sometimes within a box, right? Like, I mean, it's efficient for you, but it is a lot worse for everything else in terms of how many boxes need to be made. Which by the way, I had this idea, I don't know if anybody's on board with this, but when you get a package from Amazon, now that Amazon is doing most of their own delivery or at least in some places, you should just leave your broken down boxes out on your front porch or wherever it is. And then the Amazon people should take those back and then reuse those boxes as opposed to trying to throw them into recycling, which I think most towns say they have recycling and then it ends up in a landfill somewhere. But anyways, separate idea.

Paul: Whole other conversation.

Jeremy: Right, exactly. So you mentioned though, this idea of again, efficiency probably creates more demand in the sense, because it's easy. Right? And so if it's easier to do, then again, you can consume more. And I think that is certainly true with things like EC2 instances. Like you said, I can just spin up as many EC2 instances I want to, and it's going to cost me incrementally, but for the most part, it's pretty cheap. Whereas when you build with serverless, you kind of build with some constraints. So does that tie into green energy or, I guess, maybe fighting against Jevons paradox there?

Paul: Yeah, I think it does actually. I think there is more in there because your constraint is you're trying to reduce, you're trying to use as little as possible when you're trying to build serverless, or at least in the way that I built with serverless. You're trying to say, "How can I do as little as possible work in terms of compute and get as much as possible out of it?"

And I think that is the most efficient thing you can do. And it's like your end point is, with all of these efficiencies, is to get to the point of only doing as much as you have to, to achieve the goal that you're trying to get to. I think if you're using something that is inefficient, then you always have efficiencies to bring in.

I think if you're trying to build something that is as efficient as possible and there is pretty much no... There's nothing to make more efficient in that process, then you're essentially building as green as you possibly can, whatever the goal of the solution at the other end at the end of the day is.

I mean, just as a quick example, one of the principles that I had when I was building the startup in 2015, '16, was that we didn't put any... We were using Node and we were using Python as our libraries within AWS Lambda as our languages. And we basically said as a rule that we use no libraries. We use no libraries, we use no frameworks, we use nothing. It was all pure language in each of the functions. And each function should only do one thing.

And it was like, the principles meant that each function, I think we only had like three or four functions that were over 150 lines ish, maybe 200 long. So if you went into a function, every single person in the company who was able to code could go in and probably fix the function if it went wrong. And it also meant every single one was lean to within an inch of its life, do you know what I mean?

It was that kind of, it was so efficient and it was so clean that it was getting ridiculous. And it was easy to do because it was just the principles that we set in place. It was the constraints and the principles and everything else. But if you don't start with those principles, if you don't start with the understanding that that is the constraint and you just go, "Right, we've got all of this, all of this. We can do what we like within the context of the world. Oh, isn't this lovely?"

If you don't start with those understandings and constraints, then you end up in, "Well, how do we make this more efficient? Can we make this more efficient?" If you start with, "How can we be as efficient as possible," then it's a different conversation going forward. And I think that's the serverless...

Jeremy: Yeah, no, I think you're totally right. And I mean, that's one of those things where, like you said, the single purpose function thing is... Not only is it great from an efficiency standpoint, you're using as little, you're spending as little money as possible because it loads fast. You don't get the cold starts as badly and things just run really quickly.

It does one very specific thing and it can scale independently, right? So if you have millions of people hitting against that one thing, that one action scales, whereas something else like maybe your delete AP, your delete path or whatever doesn't have to run. That doesn't have to scale.

And it's not like a microservice that's running on a container where that container is running and have to scale up everything just to process one small bit of code. So I think that's really interesting. And I love this argument about thinking about efficiency right from the beginning. And you and I have talked before about where that efficiency comes in later on and what that means for global impact or for climate impact later on.

But I want to talk quickly about Project Drown, sorry, Project Draw Down because you talk about this a lot. And I think this is a good way for developers who are unaware of their impact to start thinking about this. Can you tell people a little bit about that?

Paul: So I think there are a lot of people think that to fix climate change and to kind of hit all of these goals, the obvious thing is we've just got to sort out the electricity system. One of the obvious things we've got to do is go and just make everything renewable electricity. And they don't know things like that 33% of global emissions come from agriculture, or that cement is a significant emitter of carbon emissions or that the oil industry itself is like 10% of emissions because of the fact that actually taking stuff out of the ground that is going to be burnt and emit emissions actually produces emissions itself. It's a massive, complex world. It's not just about this we need to make more batteries and the world will be a better place. The complexity of the argument about how we fix the problems that we have are huge and systemic. And most people haven't actually looked at this problem to know about it. And I think Project Drawdown is brilliant because one of the things that we almost certainly will have to do to hit the kinds of targets we need to hit to make a livable world in the future is that we're probably going to have to take carbon out of the atmosphere in some way, or at least reduce, significantly, the amount of emissions in certain areas.

So one of the organizations that looks at this is something called Project Drawdown. And they actually produced in April I think it was an updates to a previous piece of work, which was looking at the biggest impacts in what kind of things will reduce carbon emissions significantly. So what will effectively draw down emissions over time as quickly as possible. Previously, it was things like air conditioning units, because actually the stuff in air conditioning units is impossible to recycle. But actually air conditioning units, when the planet's getting hotter, we're going to need air conditioning units. So if we produce more air conditioning units and the planet... It's just not good. So we need to do something about air conditioning.

Jeremy: It's a vicious cycle.

Paul: Brilliant. Yeah. And, but previously they were also talking about things like educating women and actually because educating women is important for understanding how to, across the world, this isn't in various of place, but across the world in terms of understanding, because they see that the more educated that women are in terms of in areas where they aren't educated, there seems to be greater population. So this just seems that this is a good idea. So creating organizations to go out and educate women is a climate. It's a good thing for the climate. You're like, "Okay, this is a positive." So you see that it's not just about going and building massive wind farms, and it's not just about building batteries. And it's not just about getting everyone to drive electric cars, although those are also good. And it's also things like agriculture, and it's also things like lab-grown meat. And it's also things like all of these elements actually are a part of the solution and it's understanding what are good things and bad things to do and understanding how they all fit together. And I think it's worth just Google Project Drawdown, it's worth having a look at their website, reading all the stuff around it. I think it gives people a more rounded understanding of what's needed to be done rather than just going, "It's just this."

Jeremy: Right. No, actually I was taking a look at it after you and I talked before, and just the fact that like family planning, just, that has a huge impact on carbon emissions. You know what I mean? And just everything from the medical services that are needed to actually having another child and what that means and just being able to understand those impacts. So there are obviously those big things that you can do. These are global things like, yeah, let's get rid of fossil fuels and some of that stuff, but what personal actions actually matter when it comes to... So we'll move away from technology here for a second. Because this is an interesting subject for me. My wife and I talk about this all the time.

But what personal actions actually matter. And because the other day, and this is something silly, but we recycle or we try to. We put the stuff in the recycling bin, hopefully it gets recycled. We try to buy aluminum or glass instead of plastics and things like that. Recently, we just switched because I was getting sick of throwing away all these paper towels and napkins. We just switched to cloth napkins. Now, again, does that have a huge impact? Probably not. Does it make me feel better about myself, a little bit. But what are your thoughts, because I know you're a vegetarian, like what do these even mean? Does this help?

Paul: Yes, I think it does. I think there are two things that's worth saying. I think all these things are worth doing. I think they are worth doing from the point of view of changing who you are and changing the way that you live. I think those things that are going to be important, I think over the next 10 years, the world is going to change. I think we are going to see a radical shift into something looking like a low-carbon economy. And I think a lot of people are going to struggle to shift over to that. And so I think the more that we can do to understand that and understand that we are probably going to be eating less meat, that we're probably going to be reusing an awful lot more of our items, that we are probably going to have less access to some of the things that we see as convenience items.

We probably will still have cars, but I suspect that fuel will be significantly more expensive because it will be taxed, not because the fuel will be unable to be gotten. It's just that I can see an awful lot of things. So I think we need to think about our own personal lifestyles in the context of the world is going to be changing. And I think it's probably a good thing to start considering how your life needs to change around all of those things. So I think it's important to think about those things and understand that actually we do live an incredibly privileged life. We talk about switching lanes a little bit, we talk a little bit about how we have... Talking about Black Lives Matter.

Being white, I literally don't know, well, that I have privilege simply because of the fact that I am white. It's just, it happens. And I don't know. And I've learned an awful lot about all the little things that I've never had to worry about. And I think if you switch lanes back, we live in the industrialized nations. We have had the oil and it's a similar kind of thing. We have had the money and the wealth and the oil that. And that other nations who haven't had the money in the wealth of new oil haven't had. And I think we need to realize that that isn't going to last forever. That tap is going to get turned off. And I think the wealthiest nations are going to struggle the most because they are going to have to change the way that they think.

And I think that lifestyle, the idea that the green economy is going to change all of that. And they're just going to be able to carry on and keep staying rich, I think isn't going to keep going. So I think in terms of going back to the personal because that's where it all started. I've just gone on a bit of a rant. I think you and I, and all of us, I think we all do need to start thinking about the personal. I think we need to change those things. Recognize that they aren't in and of themselves going to change anything in a big way, but getting involved in a movement is. So in the U.S. it'll be writing to whoever your representative is or ringing them up or whatever it is.

I see all these things from the U.K. and think, "Well, we don't do that. We just write to our MP," and then that's about all we can do or go on a march. But it's those kinds of things. Finding the groups that actually talk about these things and understanding the intersectionality of climate with other areas of justice, including race. And I think these things are quite important to understand that there is a connection. For example, it's something that I think is quite important to point out is that is an awful lot of non-white. And I think it's worth just saying outside of white people, non-white people live in more polluted areas simply because they are poorer. So there is a racial element to climate justice. And I think that you can't just separate these two things out. So I think you can't just go, "I'm just going to buy better things." I think there are other elements to being part of the conversation that just look into it and see where it fits.

Jeremy: Yeah. And I think you're right. I think everything is so intertwined and complex. And I think that goes back to something that I thought about for a very long time where I feel like it's the infrastructure that is part of the problem. Like you said, we've got oil and we've got combustion engines that we use to power our cars. We have highways that are terribly inefficient in terms of backing up traffic and cars idling and things like that. In the United States, I've been to Europe a number of times and in a lot of the cities in Europe, they have very efficient public transportation systems. In the United States, we don't. I live about 45 minutes outside of Boston, Massachusetts. If I drive my car in during rush hour, it takes me an hour and a half.

If I take a train in, which is just an old clunky commuter train or whatever, that takes me an hour and 15 minutes. So the trade off of the flexibility and so forth. And again, I'm probably not saving much in emissions either, but just going back to the infrastructure thing, this is something where if I'm sitting at my house right now, like you said, I'm privileged, I live in this nice house. I've got central AC, I've got heat that I can just turn up to as high as I want to and as be as comfortable as I want to burning oil. Because that's what was built in this house. That's what I have. I have oil heat. The question is as a consumer, and this is super selfish, but why do I have to be cold? Because the government and because corporations, and because everybody made these decisions for me to say, "This is how you get warm and you get warm by burning oil," as opposed to saying, "Hey, we've covered the Texas panhandle with solar panels and we're collecting all this energy and we're going to have electric heat it's going to be a hundred percent efficient and you can turn the heat up as much as you want to and that's not going to affect anything."

Now again, I know that's a little bit selfish, but I think about these things just from an infrastructure standpoint to say, "Why do I have to make those trade offs?" And I think a lot of people would ask this question is like, "Why does my house, why are modern houses built with one waste pipe? Why isn't there a gray water and a waste pipe? Why aren't those separate? Why can't I capture that gray water and use that to water a garden or something else? Why aren't houses built that way?" And so I think that's just this thing that bothers me so much is that asking people to do a lot of these personal things, and I have no problems doing these things, I certainly don't have problems doing these things, but I feel like the infrastructure is fighting against us.

Paul: Yep. And I would agree with you. The very simple answer that I think, but yeah, the slightly longer answer is that oil companies, it has been shown in various different ways. And I think it's worth pointing to other resources on this, have spent an awful lot of time, money, energy resources in making sure that they are the ones who make the rules. And they are the ones who have the power to do things. And it's like there were electric scooters in the 1920s. So people think you get an app and then you go and you scan the QR code, you pick one up and you scoot off and isn't that wonderful. We had electric scooters back in... we were doing stuff like this. The first cars. There was a discussion whether it was going to be electric back in the day. We have an oil industry that is incredibly powerful across the world.

And oil is seen as incredibly cheap, and it's incredibly efficient in terms of you burn something and it produces an awful lot of power and energy. And so there are some really positive things about that, but the downside is that it produces the emissions and those emissions have a secondary effect that is incredibly negative. So we need to move away from that. It is definitely a systemic problem. It is definitely a systemic issue. And it is definitely something that you can't fix on your own. So part of this whole conversation about people going off grid, and I've seen an awful lot of people in tech, I've watched them, they're going, "Yeah. I'm doing my bit, I've got loads of solar panels. I've got a heat pump in my garden. I'm doing this, I'm doing that, blah, blah." And I'm sitting there going, "Well done. You're going to be fine while the rest of us are absolutely stuffed." And it's just like we have to kind of think about all of us. And all of us is politicians. It's companies. It's understanding that a pledge to do something by 2030 by a company is not the same as we've done this this year. And it's just thinking about those kinds of things is far more important. And putting pressure on companies, I think is also quite important.

Jeremy: Right, now, are you a fan of the Back to the Future movies? Have you ever seen those?

Paul: Yes.

Jeremy: All right, so, because I'm just thinking that if electric cars would have been a thing, then the entire plot of Back to the Future III would have been unnecessary and they wouldn't have had to make that terrible movie. So, but anyways.

Paul: It would have been different.

Jeremy: It would have been different. So one of the things though that I guess... now that you and I, we've solved the world's problems here because we've discussed them on a podcast, obviously there's a lot that needs to be done. But practically, what can people in tech actually do to try to have an impact on this?

Paul: That's a massively broad question. So right. I think the biggest thing is make yourself aware, start to do the research. A really good place to start is, it's something that a lot of people have heard of, but probably not read, is just to read the IPCC, Intergovernmental Panel on Climate Change, Special Report of warming of 1.5 degrees, the special report at 1.5 degrees, it's fine. See IPCC. It came out in October 2018, but it was the thing that triggered most people into understanding that we had not a huge amount of time left to really try and keep warming below 1.5 degrees. And that was where the idea of a carbon budget came from. The amount of carbon we had left to burn.

The amount of carbon we had to put into the atmosphere before we basically breached 1.5 degrees as a threshold, or two degrees. And so that was where the idea of 12 years came from. And so all of these little things came from that paper and it's well worth reading. So I think that's worth kind of just pointing as a good reference document for all of this. But in terms of tech, simple, start looking at your data centers, where you store stuff. Are your data centers run on green electricity? If they're not, go and ring your data center provider and say, "Move yourself over to green electricity, please. Thank you very much." It's really not hard and actually half the time it's cheaper. If you're in the cloud, move your workloads over to green regions.

So if you're already in Google or Microsoft, you're probably already in a green region. If you're an AWS, move them over to Oregon, Montreal, Frankfurt, or Ireland. If you can't move them over, because you're basically stuck, which is often what happens, then move your workloads that can be moved. So your things like your test environments or your machine-learning workloads or things that don't matter which region that they're running, move those over to green regions. There is always a way of moving more stuff over to green regions. Turn stuff off is also something else. Don't leave servers running overnight. It's all of these little bits and pieces that people are like, "I don't understand." Why do people just go, "It's fine. It's just it doesn't matter. It's over there. I don't need to think about it." Turn computers off overnight. Your own work unit. These are habits. Once you understand that, it's just simple things to do.

That's kind of important. There was another story I read, was it last week in Wired, or it might have been this week, about a guy who has a WordPress plugin, and the plugin, he looked at it and he realized that if he just cut down some of the code, and like 2 million people use it, but just cut down some of the code, he could reduce what the download speed was, the amount of download. And he's worked out it's around, I think he said 57,000 kilos of carbon emissions saved just by doing that. I think that's off the top of my head, but it's like it doesn't take a lot to start to think about how can I reduce the impact of what I'm doing. And even to the point of going, "Can I just reduce the amount..." Taking libraries out of code. Taking frameworks that you're not using out of the code. Actually just reducing the amount of code you write is just a simple thing to do.

And it's not that it's going to save the planet if you take one line of code out. Well done. It's just the principle. It's just following through on a principle. And it's just taking these principles and saying, "This is what I'm going to do ongoing." But that's all it is. Is it's a set of principles that if you can stick to them, then after a while you find that you just keep doing them and keep doing it. And you are finding that it gets better and better over time.

Jeremy: All right. It's the straw that broke the camel's back, but in a good way, going the other way. Like every little bit helps. Yeah, and I think that's something that is a powerful motivator for startups is saving money. So trying to be efficient in those ways. And I know I've worked for a lot of bootstrap startups or a lot of startups that were like, "Let's keep the cost down." So I think, again, working with serverless, having those constraints, having that mindset of we need to be more efficient, having the mindset of we need to save some more money. These are all benefits that really should benefit your company in the long run anyways. And I guess that's another thing too. What about organizations? So what can you do? You're just a staff engineer at a company, what do you do to try to affect change at your organization?

Paul: So this kind of goes back to the work that I've been doing with the Leading Edge forum up until, well, I'm still in the middle of doing it. It's kind of working with them on looking at organizations and how they change and grow and look at climate and look at sustainability as a wider goal. One of the things that is clear is that actually organizations don't know things like where all their emissions come from. They don't know things like understanding whether their supply chain is full of good or bad emissions. They don't know these things. It's actually quite difficult to understand. So it's actually a simple thing to go and talk to your head of sustainability and go, "Well, actually, what is the situation?"

And I actually, I know this, I've done this at a number of companies. If you go and talk to these people, they love talking to technical people because then they can go, "Can you give us a hand with this complex technical problem?" And you're sitting there going, "Brilliant. I've now got some things that I can..." It's like, "They want something." And over the next few years, I think we're going to see a connection between data and sustainability that is going to become a quite important key thing. So I would suggest that if you're just a staff engineer at a small company or something like that, I think putting yourself into a conversation with head of sustainability or asking who is leading on this, if you're in a smaller company, I think it's actually quite important.

I've talked to a number of smaller companies who, they've asked me, "Well, how can we reduce our carbon footprint?" And actually, that's a very difficult question if you're a smaller company. It's like, well, this is pre-2020, but they were doing things like flying to conferences and understanding what can we do then? It's like, well, either you reduce the amount of your flying to conferences or at the very least you offset and offset more than you would normally offset, but you need to understand reducing is actually quite important. So think about what you're doing there. So maybe think about doing remote conferences and remote videos and being a leader in front of that and talking about why you're doing that. And a couple of companies have taken that on board and started to do that anyway, before all of this.

And I think we might see more virtual conferences anyway after this whole change in the way that we do things, I think people are becoming more used to doing virtual work and virtual conferences and seeing the value in all of that. So I think it's just starting to become aware of, again, all those little things. It's like, "Well, we use servers, we do this, we do that." Talking to the senior people and actually starting to ask questions, "What are we doing in this company? How are we doing? How can I help?" It's just a really easy question. And you'll find that actually not many people are actually asking that question. Most people don't. A lot of people who are, even in large companies, even in very, very large companies, they don't necessarily ask that question. They want to do something, but they don't necessarily go up to someone and go, "Well, what can I do?" And I think that's a really straightforward and simple thing to do. And people will respond.

Jeremy: Well, you made a good point about the travel for conferences and things like that. And I know you were one of the founders of the ServerlessDays Conference or JeffConf originally, which I think is an amazing thing. I really like that idea of regional conferences, in the sense where people can take public transportation or they don't have to fly to get to these, and then you bring in some speakers that do fly, but it's better to have 10 people fly than to say, have 65,000 people fly to Vegas for re:Invent or Google Next or some of these other, really, really large conferences. I like those small intimate regional conferences because again, those are more efficient, I guess, than making a lot more people travel to one place as opposed to having a few people travel, and it's all easy to get to.

All right. We talked about the big three a little bit and compared them in terms of their green and stuff like that, but in that paper that you wrote, you have this cloud league table in there where you compare them. I'd love to know more, what about Alibaba and Oracle and IBM and some of these other things, where do they all stack up against one another?

Join us next week for Part 2.

This episode is sponsored by Amazon Web Services and Epsagon

View Details

About Erica Windisch

Erica Windisch was the co-founder and CTO of IOpipe where she helped organizations maximize the benefits of serverless applications by minimizing debug time and providing correlation between business intelligence and application metrics. She is now a Principal Engineer at New Relic.

As an advocate and pioneer of cloud computing since 2001, Erica is always pushing forward as technology and the industry adapt. She was an early contributor to OpenStack and maintainer of the Docker project where she worked on hardening Linux containers and establishing corporate and community security policies.

Erica is a champion of AWS Lambda and serverless technologies, and she speaks frequently at conferences about AWS Lambda and other AWS solutions. She's passionate about systems architecture, security, and the future she sees for machine-automated, low-code development.

  • Twitter: twitter.com/ewindisch
  • New Relic: newrelic.com
  • Personal email: erica@windisch.us
  • Professional email: ewindisch@newrelic.com

Watch this episode on YouTube: https://youtu.be/T1t_P_zqOiE

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and this Serverless Chats. Today I'm joined by Erica Windisch. Hey, Erica, thanks for joining me.

Erica: Hello. Hi. Thanks for having me. Or thank you for having me.

Jeremy: So you are a principal engineer and architect at New Relic. And you're also an AWS Serverless Hero. So I'd love it if you could tell the listeners a bit about your background and what you've been doing at New Relic.

Erica: Oh, gosh. Okay, well, my background is pretty deep. So, I'm at New Relic now. Before New Relic, I was the founder and CTO of IOpipe, which was an observability product for serverless applications. Now, I am working as an architect and principal engineer for New Relic. And if we're going to rewind history a little bit. I previously was a security engineer working at Docker, where I founded their security team and their security processes. I was involved in OpenStack from very early, since its founding. And then before that I had actually had my first company and we had built a cloud. We actually had our own cloud services. We were building from 2003 actually, building out horizontally scalable cloud services. And I said, "Well." We bought really early into that pets versus cattle idea.

Jeremy: Nice, nice. Well, so obviously you're doing a lot with observability. And you're doing that in New Relic, that's sort of what New Relic does. IOpipe was all about that. I know a lot of the team has gone over from IOpipe to New Relic, to continue to work and expand their services. And I'd love to talk to you about that today. We've done a number of shows where we've talked about observability. But that was probably almost a year ago at this point. And I'd love to get a sense of sort of, where things have gone, where things are going. You know maybe what the future is going to look like. I got a bunch of other things I want to talk to you about. But maybe you could just start, just in case listeners don't know, what do we mean by observability?

Erica: Oh, gosh. The way I see it is, being able to really see what's happening in your applications, in your infrastructure and doing that... I would say early monitoring. Things like Nagios was, I would not consider observability that was monitoring. It was very much, very reactive. There was zero product... It was not productive at all using something like Nagios. Logging products give you some ability to start getting into, being able to be proactive. And I think that observability kind of ties in some of the concepts from logging, and ties it in with your metrics and ties in being highly correlated. And also deeper into your application, having traces in your application, having contacts for your applications. For instance, just having a trace and knowing that, say an API gateway triggered a Lambda is one piece of information that you can have, but knowing say, the resource path, the HTTP method, things like that. That's a deeper set of insight that I think is necessary. And definitely fits within an observability picture that is very much say different and distinct from something like Nagios. Or even just plain text logs.

Jeremy: Right. Yeah. And we've talked about on the show the three pillars, right? You've got monitoring, tracing and logging. And so monitoring, like you said, is that sort of general like just something goes wrong, maybe you get an alert, something like that. The logging bit is obviously logging data. But let's get more into tracing a little bit. What do we mean by tracing?

Erica: Sure. The way I look at tracing as being able to see the relationship between various components, and not just the components. And I think this also where maybe tracing generally... in our industry historically, has been this service talks to this or that service. And that service talks to another service, etc. I think of it as this function communicates to this other function. And that is true, even outside of serverless, where functions are the primitive. Serverless was a really great place for us to start because it's already segmented into functions. But if you're looking at a microservice, there's no reason that you can't think about your code, about say this functions or this component or this resource path is communicating to this other function and also, contextually. So, for instance, maybe this service only calls DynamoDB when it's inserting data. Or when the API gateway, there's a put request, right? That triggers a put into DynamoDB.

You don't get a put into DynamoDB when you do a delete on a API gateway. So that's the kind of context that I think is really interesting for things like tracing. That is a little bit more I think, beyond what traditional tracing solutions have been doing.

Jeremy: Right. Because I mean, it's a lot different in these distributed systems. I mean, even if you're just talking to one microservice, it's usually you talk to one microservice and maybe you want to see that continuity there, or service X called service Y. But now with serverless specifically, we have you function X calling service Y, which generates an event that then gets picked up by EventBridge, and then another service picks it up and so forth. So we can get into why it's important in serverless applications, but is there anything else, where observability is different in the sense of monitoring? I mean, you mentioned the idea of being a little bit proactive. What do you mean by being more proactive?

Erica: Well, by being proactive I guess I'm referring to the fact that things again, rewinding history a little bit and going to something very distinctly not observably like Nagios. Again saying, "This very reactive." Something went down and now we then asked it, "If it's up?" And I said, "It's..." And we couldn't reach it, so then we determined that it's down. I think that kind of step two of that kind of journey towards observability, would say, "Okay. Well, we have logs, we have a logging product. And the logs told me that when, I don't know. This service tried running to the Syslog Server, they got an error." Well, when I get that error, I know that at least this system cannot talk to the Syslog Server. In fact, maybe I know that a hundred systems cannot talk to the Syslog Server. And I think two things come out of this. One, is that it eliminates the kind of the, maybe the falsehood of a binary status for uptime, right?

Because maybe that Syslog Server is up from the perspective of say Nagios. Or from the perspective of machines on a segment or on a subnet. But maybe there's machines in another AZ or another subnet, and those are the machines that cannot talk to it. And that's contextual information that is really critically important. I guess you can argue that it is still somewhat reactive, because you're still basing it off of say something like logs. But you're not polling for that kind of information directly, necessarily in order to have basic fundamental understanding of things that your applications should already be knowing about themselves.

Jeremy: Yeah. And Sheen Brisals wrote an article the other day that I thought was really interesting, where, with all the observability in place on the serverless applications that they have at LEGO, they basically said the system reports when everything's healthy. Or the alerts are, "Hey, everything's working, right?" I mean, you can see that everything is going through that pathway. I think that's really interesting too, because it's not only about maybe this service not being available. It's very much so about this service not being available if you try to put data this way, right? So you can see that with tracing and you get a much better understanding of, "Well yeah, the system's not down." If I was just getting alert saying, "System's fine. Systems fine." But then you're seeing a consistent pattern of certain messages failing, then it's really great to have that tracing and the ability to go in and then dive into it and say, Oh okay. When it's shaped this way or when it comes from this component, then there's the error."

Erica: Yeah, exactly. I think that being able to know when and where, is a vital component; Like Nagios for instance, would tell you what, but it wouldn't tell you... it would tell you when, but it wouldn't tell you really where necessarily right? Which applications are having problems communicating? And I think context is really the important key for me here. And being able to facet that data and tell you exactly, where it's happening and for who, tells you a lot about why it's happening. Because going back to the subnet example, if you can easily look in your observability tool, and see that all of the services that are in this subnet are having a problem communicating, you start to really flesh out the why of a problem, much more quickly than if you just know that, that service cannot be communicated to, and you don't have any other additional contexts.

Jeremy: Right, right. Yeah. No, context, I think is super important. All right. So why is observability so much, or I don't say so much more but extremely important in serverless?

Erica: Well, I think one of the things about serverless is the fact that, it is broken up by default into these many pieces, right? So, you by default you have much more sprawl, you have many more services. From a perspective of, instead of a monolith which contains many functions or many endpoints or resource paths or whatever. You get, maybe potentially many functions that serve that application. And I think that, kind of two things come out of this. One, is the capability to pinpoint things more accurately because that context is kind of baked into serverless. Because when you know there's a problem with that function, and that function has a very narrow scope, right? That gives you a really strong context into what is happening, versus it's this application. Right? Because no one gets a function, you have that built in context. So I think that serverless actually enables you, more so than the fact that it just needs it.

And, yes, I think that's for me, the biggest thing. In terms of other reasons why maybe it's important for serverless, is just because there are maybe a lack of other traditional tools. So, you wouldn't run, maybe some of the more traditional tools in the traditional way, with a serverless application. So, you're not necessarily getting that broad picture that isn't clearly defined. So, in some ways where serverless kind of forces you into this deeper contextual awareness of your application. It also kind of requires deeper contextual observability for this applications. They kind of go hand in hand.

Jeremy: Right. Yeah. No, and I actually find with a lot of the tools we use in serverless, whether it's SQS or EventBridge, or some of these other ones where, you don't really see what happens. It's very black box for some of these transport mechanisms that are in there. So, being able to connect that stuff together. You can't go and look at your RabbitMQ logs, for example and see what happened if messages got lost or if they didn't get delivered or something like that. Whereas, you put it into SQS that's just not available to you, right? You have to see whether or not it actually happened. And without recording that and being able to trace that all the way through, there's obviously a lot of data that's missing there, if you don't have the right observability tool in place. So what about the challenges that the developers see when they're trying to, monitor these applications and troubleshoot those applications?

Erica: Sure. I think a really good point with the SQS, and I think this also exists for services like Kinesis, you don't have traditional logging for these services necessarily. Sometimes they can act like a black box. So knowing the context for which your application is consuming from those services. What kind of messages are getting the rate? How it's partitioned? A lot of that information is contextual provided to the Lambda. So I think that observability of the Lambda itself for instance, can give you some insight into those services that you don't otherwise get. I think another challenge is again, relating back to the sprawl. There could be many components of a serverless application. And I think that, first of all these are distributed applications. And not everybody's familiar with and comfortable with shipping and observing and operating a distributed application, in the way they are with maybe monolithic, non-scalable applications.

And I think that a lot of users do really need tools, to help them bridge that experience gap as well. And also just, even with the experience, it can be a really valuable tool to help you visualize what is happening in your application.

Jeremy: Right. Now, what about the fact that it's so ephemeral. I mean, we see containers being very ephemeral now as well, but as the fact that functions disappear after a few seconds. Is that creating problems?

Erica: I think it's a different way of writing applications. I think something that happened, especially... well something I saw a lot early on when we were doing IOpipe, was that users wouldn't necessarily always account appropriately for how the Lambda environment worked. So they would assume that things could be long living where they couldn't. Certain libraries which made that model very difficult. Some of the database libraries in particular, were a really frustrating challenge for a lot of users. AWS has made some progress on building proxies and data APIs and things like that, to kind of bridge that gap. Because some of those libraries are kind of fundamentally, maybe not incompatible but less compatible with the serverless, with the AWS Lambda model, at least.

Jeremy: Yeah. So when you say incompatible though, I think you'd mentioned to me before that there was a W3C standard, that is sort of standard now but not necessarily standard in X-ray?

Erica: Yeah. Well, there's a W3C standard for... well, trace context. WC3 trace context. And New Relic was actually involved in creating that. We have some engineers, I think Justin Foot and Erica Arnold, in particular were involved in that and maybe some others. And, with that the idea is that HTTP... It defines it for HTTP headers in particular. Although the actual encapsulate data could theoretically be put over other protocols, but it does the over HTTP, W3C I'm sure. Right? But the idea is that, this a standard set of headers that can be passed along vendor agnostic, throughout services. So, if all of your applications support W3C trace context, even if they're using different libraries by different vendors, as long as they all support W3C trace context, you can actually have complete traces through all these applications. Now, the AWS services do not currently, as of the time I'm speaking, support this trace context. They do support X-ray headers, so they can pass those along.

Jeremy: All right. So what about the advancements that have been made over the last year or so? Because again, a year ago it was pretty cool, right? There was a lot of great tracing software, there's been more vendors jumping into the space. So has there been some maturity you think with these tools over the last year?

Erica: I think so. I think one of the big things we had out of New Relic, is the recent launch of Infinite Tracing on New Relic Edge. I was actually involved in creating the Edge element of this which is, primarily in the first cut a provisioning solution for various services that New Relic will provide on the Edge. The first application for that is Infinite Tracing. And Infinite Tracing allows you to throw millions and millions of spans, at a service that lives on the edge. So it lives, say if you're in AWS, it lives in your AWS region, it receives those traces at high data rates right? So we can ingest at tens of gigabits per second, per trace observer. And then once we consume those, we can then apply machine learning and other filtering mechanisms, to help you sample appropriately. So rather than, traditionally with tracing what would happen is, your agent whether it's open tracing or it's a New Relic agent or one of our competitor's agents. It has to make sampling decisions in a vacuum right? In a fairly stateless way.

What we're doing here instead is we're receiving all the traces, but then batching them together and able to filter them out on the back end. Right? So that we're only storing so many, but then we're actually able to do back end filtering, of larger batches. So there's much more context as to which traces are important and which ones we should keep, and which ones are unimportant and we should throw away. So that I think that's a really big change in how tracing works at New Relic, and for the industry, potentially. And something that we're also doing is we're releasing, I don't know the final product name. Because we're doing this call a little bit in advance of the launch. But it will be an X-ray integration where we're able to, ingest X-ray traces and correlate that with data that we have in New Relic. So when you have a Lambda, you'll be able to see not just all the traces that are within the application and the traces out, and context for that trace, but also be able to see the AWS services and see through those services.

So, for instance, if you're triggered by an API gateway, now you're going to have context for those traces, in the same way you would get from X-ray, but you have that now pulled into New Relic. So, in those places where we don't have deep observability because there are components that we cannot instrument, because of third party services or because these are third party tracing products like X-ray, we can actually pull those and they tell you a more complete story.

Jeremy: That's awesome. Yeah. Let me ask you some questions about, how some of that third party stuff works or the X-ray stuff. Because again, I know AWS has added some capabilities where they pass trace headers through SQS and things like that. But that's not available on all their products, obviously. And I think, probably I think of it this way, because I mostly build web applications. I'm mostly thinking about HTTP. Right? But there's a lot of other things that are happening. So where are we with that kind of stuff? With sort of the non-HTTP, messages being passed around?

Erica: I think it's interesting, because it's something I've been thinking a lot about recently. In particular, because AWS does have that for SQS and I think SMS. And that's not a place, where I think a lot of vendors are necessarily looking for trace headers. It's a place where W3C does not define standards, for how to pass along data in these non HTTP ways. But I do think that W3C trace context, like trace pair headers, the values of those headers could be passed along in places that are not HTTP. And I think it's going to be really compelling, once all these different services are able to support these. When we actually look at, what does it mean to have... for instance, do we get to a future when you write data into DynamoDB? That you can actually pass along a trace header? And then the trigger that comes out of that, right? The Lambda trigger off the DynamoDB stream, can actually have context for some of those traces? I don't know. It'd be really interesting to see that feature. I think we're just kind of at the beginning of that a little bit.

Jeremy: Right. Yeah. Because I mean, if you're doing something with DynamoDB now and you want to read off the streams, or even I think Kinesis, right? You still need to pass in your own correlation IDs, in order to trace those back, right?

Erica: Yes. And there's lots of questions about how that would work. I think in the case of Kinesis, which I know a little bit better than Dynamo, to be honest. You have individual records. So I think that in this case, it would be a trace for where that record came from, not necessarily... because you don't want it from the batch, right? Because if you have it from a batch, it's not really very useful. You want to have it down to the individual record.

Jeremy: Interesting.

Erica: But yeah, you're right, you can kind of encapsulate that in yourself right now. But none of that is kind of built in by default. And there's no way that a New Relic could just say, modify people's Kinesis records because that's arbitrary base64 data.

Jeremy: Right. Yeah, exactly. Exactly. No, I mean that's what I'm thinking. It's like it'd be really great to have that extra stuff. Now does X-ray have any of that data, that you're you're able to import now?

Erica: Oh, gosh.

Jeremy: Right. Because you can't trace DynamoDB all the way through with X-ray. I don't think you can.

Erica: Yeah. So the SQS example, I'm sorry. The SQS example is something where we could potentially do that. We're ingesting that X-ray data. So if X-ray has that data, it's passing along those values and X-ray collects it, then we can ingest it and we can give you that context. Our agents are not directly getting that data, which just means that it's going to be harder for us to correlate it. But honestly, until, say SQS and AWS have W3C trace context support, we're probably going to be a little bit of a gap period before, we get the kind of correlation that we want between native New Relic traces, whatever that is, and native X-ray traces. Because once you have W3C trace context, you don't really have this concept of native necessarily anymore. There's something called a trace state which is a vendor specific field. And for the large part, we might ignore those. We might decide to support some of those. Well, for the large part, we're currently working with the trace parent, which is a highly vendor agnostic field.

Jeremy: Right. All right. Well, if AWS if you're listening, let's get moving on that stuff because it would be nice to have-

Erica: Yeah. And I definitely give them my feedback.

Jeremy: All right. So take off your New Relic hat for a second. And I'd love to just get your insights, just into the overall landscape that, where we are with serverless observability. So obviously, the landscape, the number of vendors that are getting into this space. When we have IOpipe and now New Relic, Thundra, Datadog, Epsagon Dashboard, Lumigo, Honeycomb and then AWS recently launched their ServiceLens. So we just got all these different tools. So my question is less about, which is the best one to choose. And it's more about this idea, I think they're all doing something slightly differently, or slightly different they're trying to add a little bit of, I guess what's the right word for it? Distinction between them. But I guess it's a good thing, right? That we're getting all this competition. What does that mean for serverless adoption? Does it mean anything for serverless adoption?

Erica: Oh gosh. I think that, the success and failures of serverless observability over the last, couple years. And the larger social economic, landscape or economic landscape rather, especially around COVID and everything else. We have a world right now where, I'm not really sure. I think that, I definitely would have preferred that IOpipe could have stayed independent for longer, to be quite 100% honest. And, gosh. I think there is a distinction between these products. One of the things that we determined in IOpipe towards the end, was that we wanted to start going broader. We had gone very deep on serverless, we wanted to start going broader and not because... well, for a couple of reasons. One was because we found that, almost nobody is running just serverless applications, right? They're running serverless applications as part of, a bunch of applications that they're running right. They have business needs and those business needs are not entirely serverless business needs.

IOpipe was a company that we were running everything on serverless but that was not the case with the vast majority of our customers. So we wanted something that was broader. And I think that New Relic was a really great way for us to look at saying, "Here's a way that we can go broader without having to, become a full competitor to New Relic. Without having to build out everything for every service." Because that's really important. Users need to have their whole application observability and IOpipe was doing serverless observability. I see some of the other competitors now also going a little bit wider, a little bit broader with their missions. But I think it's challenging because, there's a lot of pieces there to build, and you have to decide which ones you're going to build first Are you going to build out, really fantastic tracing? Are you going to build out fantastic logs? Are you going to build out fantastic monitoring? Right?

How much of these pieces are you individually building out? How're you connecting them? Because you want to make sure it's actual observability, and it's not just a piece of it. And worse, I think that IOpipe was observability. I believe it was. I believe that it encompassed all of these things, but it did it very narrowly just for serverless. And that was an intentional thing that we did because we couldn't build the entire world. There was only so much we can do, with so much money and so much time. So we focus very narrowly. To try and do those, a broad set of things for a narrow market segment, is easier than doing a broad set of features for a broad market segment.

Jeremy: Right.

Erica: But that's what we're doing now at New Relic right? Is going for the broader market segment, of not just the serverless part. Yes, we're doing serverless but it's not just serverless because realistically, you have things that are not just serverless. And it's very hard I think, for really any vendor to do all the things and to do all of them right. I know you told me to take off the New Relic hat but this a question that was really hard, to take that hat off with. Because I do think that we're doing a lot of those pieces and we're doing a lot of those pieces right. I think that it's very possible. I think that companies like Honeycomb do really fantastically with, doing their market segment very, very well. And maybe better than we do at that particular market segment. But that is, a segment of the market.

And broadly speaking, we have customers who have mobile apps. We have browser apps. I want to get to the future where I can look into my dev tools, and I can authenticate it and logged in of course, that if I am trying to debug my application, and I'm having a failure in my browser, I want to be able to click into dev tools and then jump straight into the line on GitHub, that it's giving me that problem, for the back end service that generated... the problem that went all the way to the front end. That's the future I want to get to. And I don't think that's possible, just doing narrowly focused market segments.

Jeremy: Right. Yeah. And I think that, what I like about companies that are established, like a New Relic and like a Datadog and like these bigger... these companies that are covering, this wider swath of these broader market segments. What I like about them getting into the serverless piece of it is, I think for a lot of people having a good service observability tool, it's an absolute necessity. You cannot have one of those. And if you are trying to build an application and you're all on containers, maybe you still have some EC2 instances running. Maybe you still have some on prem, but you want to get into serverless. If you have to go out and buy a different tool, and try to integrate that into what you're doing, I mean, that just becomes a really hard problem and a really hard sell.

And if you've got these bigger, more established companies that can do all these different things. And you start mixing all of those things together, then I actually think the observer, or the adoption of serverless becomes easier, because now you have those standard tools in place that are just a natural extension of your cloud infrastructure.

Erica: I think that's 100% true. And I think that was one of the biggest challenges we had IOpipe was like, "Great." But users didn't want two tools and the fact that the serverless tool was separate, made it very hard to migrate the users. Right? Because they had to migrate then, not just the applications and the way that they built their applications. But those developers had to also learn a new tool and use two tools. It is significantly better now that we have, a unified platform.

Jeremy: Totally agree. Totally agree. All right. So speaking of hybrid applications, because I think that's what we're talking about. Some people there may be running their main workloads on containers, maybe they're using Kubernetes, or something like that. And then they've got these peripheral things that they might be doing with serverless. Maybe their ETL, maybe their DevOps tasks, whatever. But clearly, you do have a lot of hybrid apps and that's great. That's fine. Do what you need for your workload. But one of the things that I thought was interesting sort of, this relatively new with Fargate and with Cloud Run. Is this idea of trying to take containers and make them more serverless. So how do you feel about serverless containers?

Erica: Oh, gosh. There's so many things here I cannot talk about.

Jeremy: Do your best.

Erica: So I think that it's interesting. I think that Cloud Run in particular is pretty interesting. I think that it's important to meet users where they are. And building out serverless container runtimes, is a really fantastic way of meeting users where they are. That said, I think there are reasons why serverless... So artificial constraints, I think are one of the most powerful tools that we have as builders of infrastructure products, right? I come from a history of building infrastructure products, things like Docker and OpenStack. And one of the things I wanted to do with Docker, and I advocated strongly for, was actually fewer features. I wanted Docker containers to be able to do less. And that was because of a number of reasons. I wanted to have more immutability for the services. I wanted to have more immutability for the logs. One of the things that I found with Docker, was that if you restarted a container, you stopped it and you restarted it, you would get a new set of logs. So if you did Docker logs, you didn't have any of the logs from the previous run.

Why would we throw those away? Those should be immutable. And I was like this should be immutable record. Logs should never be erased. And I lost that battle. I lost the battle of saying that we should not be able to ping out of containers. Because the ping out of a container requires net raw. And if you have net raw, you can do things like spoof the IP addresses of other containers on the same host. So these are the kinds of things that you can do in Docker, that I thought that you shouldn't be able to do in Docker. I thought that we would enable users, by taking away features because the problem is, that in enabling users to ping also enables them to compromise adjacent containers on that host, right? And those are things that we don't want to enable our users to do. We don't want to enable users to lose their log files, right? We want to enable them to have immutable logs. And I think that's serverless Lambda at least, right? Because I don't think you should say it's a serverless thing. I think it's a Lambda thing.

Lambda has done a really good job of having really tight constraints on the workloads. And allowing arbitrary containers, arbitrary Docker containers, for instance or OCI images to run, would mean that your applications can do a lot of things that they really probably shouldn't be doing ever. You should never have an application that can write to arbitrary... if your application was to escalate to a root user, that root user should never be able to write to the Etsy password file or the Etsy shadow file. That should be impossible. Your root user should not be able to do those things. You should have an environment where you are contained in a way, where you cannot escalate in those in that fashion. And I think that, enabling containers, right? Arbitrary containers does take a step back from that.

On the other hand, we do want to meet users where they are, and enable them to build applications in a way that, actually accelerates their development. I'm thinking back to the CGI days. We have so many users that I mean, not just users, like tutorials and blog posts, so that even when... like web operation. So this one of the things when I had a web operation in 2002, 2003, 2004. I mean, it was really the whole 2000s. But that was when for me, it was those were when we had users were shipping CGI applications and PGP applications or PHP applications. And then we started moving, we tried to force users to go into virtual machines and containers, in the mid 2000s because we wanted to stop having users doing bad things on our infrastructure. They kept doing bad things but they did it inside their own sandboxes.

Jeremy: Right. Exactly.

Erica: This is something else that we learned, right? Was that enabling users to do things in a secure way, did not actually get them to start doing things in a secure way. It just isolated them from the other users, so they did it-

Jeremy: From the rest of the system. Yeah.

Erica: Right. But like you, you would find blog posts where they tell you to make your directories node 777.

Jeremy: Oh yes. Yes.

Erica: Right? And that's something that users should have never done. But they did it because a lot of providers didn't have the right security isolation. But when you did provide the right security isolation, and you had your PHP application running as your own dedicated user, and your own container in your own VM, which we did. Users still set their directories to chmod 777. It was completely not necessary. I think it's the same struggle. Users, you give them Docker containers. Almost every developer is going to do it wrong. And forcing them to do it, the best practices. You didn't have another choice. Don't give them a choice to do it on.

Jeremy: Right. And I totally agree with you on the artificial constraints thing. And that's one of the things I love so much about serverless where, it was like there was no state, right? So you had to just think about things differently. And there were circumstances where you're like, "Wow. It'd be really nice to have access to state, to do this one specific thing." But it was a best practice, or it was a good idea to use state to do this. But under normal circumstances like let's say massively horizontal scaling, using state was a terrible idea, right? Because you just wouldn't get the performance. So then we get EFS integration with Lambda. And that changes quite a bit. Now, I think there are a lot of really, really good use cases for that. And Lambda would be perfect workloads. But back to your point, I think people can do some really bad things with this.

Erica: Oh gosh. I mean, true. I think a lot of users can do bad things with it. I've actually been thinking about some really awesome/awful things I can do with it. So I set the EFS and I have my own VPN from my house, into my AWS environment where I have an EFS, and I can locally in my house, on my home computers, mount those NFS folders, which is amazing. And then I can run Lambda jobs against the data, that I basically throw onto my NASS. So I can keep my photo libraries on EFS like out of Aperture or out of Lightroom right? So my Lightroom can now store on EFS and then I can have Lambda process my images, that I've been, or my photos. That's a really powerful thing. But also, is that a way that we should be working with our things? I think there is value in the fact that we are enabling use cases and workloads.

Another application I've been working on has been email. And I did a whole talk on how I kind of failed at building out an email system. And one of the things I did not talk about was how EFS would make this better. Because EFS wasn't announced yet.

Jeremy: Wasn't an option, right. Yeah.

Erica: Yeah. But one of the things was, you can have SES right into S3. Have S3 trigger a Lambda to write those email messages into EFS. So now I'm using Lambda to actually write in EFS. My applications that are reading from EFS are actually container applications running on Fargate. And that's because, an IMAP server cannot run behind API gateway. It can't run really anywhere serverlessly. If you want to run an IMAP server, you need to run it basically in Fargate or EC2. And so that was the model I picked. So now it's like I have, an IMAP server, Dovecot running on Fargate reading off EFS, and the files in EFS are written to it from SES. And the only way to do that is with Lambda. My alternative would be to write I guess, an SCS consumer that would pull from it. Or put it in a SQS and then write an SQS consumer that runs on Fargate.

And here, I could just write the Lambda, which is, a lot easier, a lot more powerful, a lot less to maintain and then I only have to have the Fargate for the IMAP. And the other thing is, the IMAP doesn't have to scale as much too, right? Because the IMAP only has to scale for the number of people who are reading say, a mailbox.

Jeremy: Mm-hmm (affirmative).

Erica: Writing the email messages from SES. I mean, that's the many to one problem, right? The IMAP is a futile one.

Jeremy: Right. Right. Now, what I'm concerned about is that someone's going to be like, "Oh well, now I have a file system that's shared. And I can just connect Lambda to it. So now I can just use Lambda as a web server. Right? And just load files off of that." Because, you just know someone's going to do that. Right? I mean, we've already talked about in the past like serverless.

Erica: I will definitely do that. I will definitely do that.

Jeremy:
Just for fun. But I think that like you said, meeting consumers where they are. I mean, I wonder if EFS though, especially when you get down to things like machine learning and some of these other things, you get to load really large notebooks or you've got a lot of data that needs to be loaded in and streaming that from S3 is just one, expensive and slow. Do you think that maybe EFS might dissuade or open up new possibilities where Docker containers might not be as needed?

Erica: Well, I think in my IMAP example, right? That's an example where I would have had to build that application entirely on top of Docker or EC2 previously. Now, with EFS, I could build a hybrid application, that is partially built on top of Lambda.

Jeremy: Yeah.

Erica: It doesn't get me all the way. And I guess there's an argument... well, containers on Lambda wouldn't solve that problem for me either though, right? Unless they were able to give me arbitrary ports, which would be amazing. But until AWS gives me like arbitrary TCP/IP, I'm going to be stuck having to least run the non-HTTP services on Fargate. ECS definitely did enable me to take that particular application, and not run half of it on containers.

Jeremy: Yeah, right. Right. All right. Well, so let's move on to talking about maybe the future of the cloud, because I know you did a lot of work on OpenStack. We've got Kubernetes. Right? Do you see, and maybe we bring this back to open source, right? So you've got all these big open source orchestration systems and cloud orchestration. Is that what we think it's moving to do? We think we're going to see, the OpenStacks and the Kubernetes being just the dominant players in terms of how people are building cloud applications?

Erica: Oh gosh. I don't know. I've become a very skeptical of open source over the years.

Jeremy: Okay.

Erica: I think there's a lot of traction for the fully hosted services. I mean, Lambda, excuse me. Lambda is interesting because it's completely closed source. I guess you can say Firecracker. I have hiccups. Firecracker is kind of a partial open sourcing of Lambda. But it's open sourcing of the pieces that are very much not serverless. Right?

Jeremy: Right, right.

Erica: It's open sourcing of something that looks a lot more like traditional architecture. But we also have Kubernetes and like EKS, alas the Google version, Google Container Service. It's interesting because they are shipping open source solutions, but part of this is like now AWS is charging you 10 cents per hour I think, to run an EKS cluster. And I mean, I think it comes out to $70 or something a month, just to run an EKS cluster. To run no applications on it. I want to tell you as an individual developer, I am not going to do that. It's not important to me as an individual. Now, of course, as a business, building business applications, it can very much make sense to pay $70 a month, to manage your application. As an individual, I mean, I can go buy a Raspberry Pi and put it in my garage. And I think that is, kind of important because even though that may not be the market that like AWS is looking for, it does mean that you have fewer developers experimenting and learning, with these technologies because, from a learner's perspective, they're more difficult to access.

And I do think that open source in theory, provides a lot of opportunity for learning. But all of these solutions are way too complicated. I think Docker was a really great example of a successful open source project, in that it was very easy for developers to use it and learn it. Kubernetes is way too complicated for the majority of developers to pick up and run in, their house on a Raspberry Pi or on a small server or a small VM, for them to experiment and play with. It's too much, it's too difficult, it's too expensive. And honestly, I don't think that it's a fundamental problem with building orchestration solutions. I think it's a fundamental problem of these being corporate solutions. These are solutions that are being built by enterprises for enterprises. They're not being built for learners. They're not being built for developers who need to enter this industry or to get their next job. In fact if anything, they actually create more barriers than they create solutions in some ways.

Jeremy: And I wonder, you think they may be victims of their own success, right? It gets popular, and then you start getting a whole ecosystem around it and then they get more complex and more complicated. And then, then you get things like EKS, where Amazon says, "This is too tough for any normal person to manage. So we're just going to build a service that abstracts that away." Is that something you think about?

Erica: It is. And I think that EKS can be strongly contrasted against ECS. Right? Where, you have a service that is fully managed for you. And for me to get started on ECS, I had to spin up an EC2 image, which honestly, I have to say it was a little harder than I thought. You could theoretically use Fargate although, I've had a lot of trouble with Fargate. I'll get like just error saying... I'll set everything up in the way that it's supposed to work. And then it just says, "Oh no. Fargate can't actually run this workload." And it just says, No. It doesn't say why, it just says no. And I'm like, "Okay, well, I'm just going to spin up an EC2 image and run traditional ECS." But even then, it's what? An EC2 image and a Docker container, it's not, "Here's my cluster. Here's my configuration for that cluster." It's like EKS is, you have AWS managing a service, and then there's still the service that you have to still kind of manage yourself in there as well. And it's significantly more complicated and more costly.

And I don't necessarily want to run, a cluster at all. I want to have the ECS experience for Kubernetes. Or maybe just no Kubernetes at all, as somebody who doesn't necessarily need it. I just want to run my applications. That's what I want to do. And I want to pay as little for them as possible. I want them to be as easy to set up and run, easy to shut down. A lot of the reasons I like Lambda. Because, I don't have to worry about it or think about any of it. And for the majority of my applications, that is fine. That said, I also have a programmable open source networking switch in my basement, that I have built my own operating system for. So, I can go a little bit both ways with this. But the thing is, that's a choice. I want it to build an operating system for my networking switch and run that. And I don't want to do that with Kubernetes. I just don't.

Jeremy: Right. Well I mean, you also have, I mean, it's fine for your own personal stuff. But if you've got enterprises that are relying on this stuff, then obviously it needs to be fully tested, and it needs to have lots of developers contributing to it. Which is another thing I think is interesting, or I guess an interesting trend is that, companies enforcing is, forcing is probably the wrong word. But they're having some of their employees just work on open stuff or open source stuff, right? And I like the idea of companies dedicating some time and resources to help keep some of these open source projects up and running. But just what are your thoughts on companies that have open source teams that are doing a lot of contributions?

Erica: I mean, I've been on one of those teams. I have been an employee who was just working on open source things. When I was working at OpenStack, I mean it was the majority of the work I was doing, was in the open. I also did a lot of building of like Chef recipes, and integrating those components together and making them work. And I think this one of the things that was kind of touched on a little bit in the question, was the fact that, maybe this is less true with Kubernetes, although maybe not completely untrue. It was definitely extremely true of OpenStack, which was that you had these loosely coupled components, that as an operator you had to figure out how to put them all together and make them work. And everything was tested, but nothing was really integrated. And you needed to have companies that integrate these things for you. That's why you have companies like FTO and VMware and everything for Kubernetes as well. So I think that was an issue.

From a perspective of open source developers though, my biggest issue is the culture. Every one of these open source projects or projects, however small or big that they are. Because I think, I said things like Kubernetes right? Are now multiple projects. You have things like Falco and so forth that are sub projects or adjacent projects or however you want to define them. But you have a community here, that operates a certain way, they have their own culture. And that culture is different, potentially than a culture that you as a company founder or as HR or a manager, or whoever of a company, that has not necessarily the same culture that you want your company to have, or your team to have, that is in the open source. Right? And how do you kind of resolve that difference because, one of the other things is that a lot of people hire from these open source communities.

So if you are building a team that is going to work in open source, and you want to make this a diverse team, for instance. But it's not a diverse project. How does that work? Right? Is the project and the other people in that project, going to discriminate against you, either implicitly or explicitly. It may not be intentional, right? There are implicit biases that exist. And I think it becomes very difficult because, when you have your own closed source application, and you're building things for your own self and your own teams, you have control over what you're building, how you're building and the construction of your team, etc. And I think that you lose a lot of that, when you're working in an open community.

Because if you're only working on open source, it's almost like while you're employed by one company, your co-workers are almost in a sense, a set of people that are not hired by your company. That may not actually hold the same values that you or your company holds. And I don't have a solution for this. But it's something I think about a lot. And it's one of the reasons I no longer really contribute much to open source.

Jeremy: Yeah. Well, I mean and that is, where the problem is. That you get brilliant developers and engineers like yourself, and then because of the culture that just exists in tech, which is many cases pretty bad. That if it's discriminatory, or it's just like you said, maybe they don't accept that PR request because, "Oh, it's from you." Or whatever it is. And you don't know who those maintainers are sometimes or how they feel. And then there's no accountability. Right? That's the other thing that, I think is a challenge in open source. But that's too bad. I wish you would contribute more rather than just writing your own operating system for your network switches. But anyways, so I have one more question to ask you just about this idea of, open source versus proprietary systems. So I love Lambda. Right? I think Lambda is a great product. It's got so many awesome features.

Yes, it doesn't do everything perfectly. Yes, there are constraints, that are some good, some bad. But then you've got Knative or, what's the other one there? OpenFaas and some of these other things. What do you think about that? I mean, I do love open source projects. And I do love what you can do with that open source stuff. But on the other side of the coin, I mean there's, having somebody making a profit off of it, and constantly monitoring it and improving it and listening to customer feedback. I think that's important too. So where do you stand on those types of products?

Erica: Yeah. This is, again, the challenge of open source versus corporate engagement. Going back the reasons why I'm hesitant on open source. But, it's just, it is important I think, to understand where your users are coming from and what your users need. And I think that, a lot of those corporate interests are really good at, having product driven decisions. I don't think that it's necessarily a requirement. But, on the other hand, you have a lot of things that, some of the more successful open projects that do not have corporate sponsorship, do also tend to be things that are more straightforward. The use case is really well in known. I think that some of the video game emulators for example, are very great projects, that maybe don't have as much corporate sponsorship as other projects in open source. But also it's really obvious what you need to do, right? You need to make the thing work technically, to a defined standard. Whether that standard is written down or it's a black box, you're replicating it.

What was interesting for me with Knative and OpenWhisk and some of these others, was that they didn't necessarily actually go to Lambda and look at like, "How are we going to kind of emulate this service?" They kind of went and did it their own way, with their own product, Oryx. And they didn't necessarily learn the lessons, that the other products had learned or this other teams have learned. So yeah, I'm a little bit... I don't know I guess I'm a little conflicted on this, because I don't necessarily see corporate engagement always actually delivering the right product. Because I'm not actually sure that Knative is the right product. I don't think it's picked up the way that, a lot of people hope that it will pick up.

Jeremy: Interesting. Yeah. I just wonder too about, whether there's that question of lock-in. I feel like that lock-in question is just... so many people still ask it or still, I think in a way that is part of their decision making process. But I just think of something as simple as compute as Lambda. And yes, it's got all these other great features, things that can connect to all the eventing that's built in. And then you look at something like Knative or OpenFaas or OpenWhisk or any of these things that are, open source implementations of these. And they have their sets of limits and their features and other things that they do as well. I mean, do you think lock-in and I even hate to ask this question. But I mean, it's lock-in a factor there? Or is it one of those things where, moving a compute service is probably not as challenging as trying to, design for the lowest common denominator?

Erica: I think that for open source, lock-in has a few factors. One is, what is the velocity of that project and its uptake? Because a lot of companies do not want to be the first ones to adopt something like Knative. And they don't want to be locked into it, if it turns out that the project ultimately fails, right? Because now they are locked into something that is abandoned. And nobody wants to be locked into something that's abandoned. But increasingly, kind of going back to, the culture thing a little bit. I don't think most people think about this, but it's something I personally think about a lot. Is also locking yourself into that culture. Because if I, let's say use Linux, I am locking myself into the Linux kernel community to a certain degree. And if that's not a culture that I want to be associated with, or if I don't feel comfortable with that culture, I'm not locked into an operating system, that I don't feel that I can contribute to. That, is Linux open source, if it's not accessible to me as a developer?

If I did not feel that I can contribute to that project successfully, for various factors. Is it actually open source? And is my ability to engage with that project and work with it, really any better? Or is it actually worse than working with something like macOS? Where I might, or maybe in... or even Windows where, I can maybe build a business relationship with Microsoft or Apple, that is non-discriminatory? These are really interesting questions that I've been asking myself recently. And I think it relates a lot to the lock-in, because as soon as I choose a technology, I'm choosing the people that build it.

Jeremy: Yep. Yeah. No, I think that's a great point. I remember seeing not that long ago, someone who posted on Twitter that they were, and I think it was some SQL group or something like that where knowingly refused to address them by their proper pronoun, even though they knew what it was and knew that that's what that person preferred and just ignore that fact. And I think it's little things like that, that push people away, and again it's hurtful, it's hateful, it's disgusting. And those things, that bothers me too. So I mean, keep fighting the good fight for that. I'll do whatever I can, to be as open and welcoming to these communities as possible. It's just one of the things I like about the serverless community, I feel like it has been very, very open and welcoming. And it's just, hopefully a safe space to be. So hopefully you can make more of these tech communities those things.

So anyways, so Erica, thanks again for joining me and giving me all of this insight into observability, as well as into this open source stuff. It is a lot of things that we need to be thinking about in 2020, that I think people have ignored for too long. So I appreciate your voice on this. So if people want to get a hold of you or find out more about what you're doing at New Relic, how do they do that?

Erica: Well, I have a Twitter. It's not only technology though, and I guess you could email me if you want to, personal as erica@windisch.us or professionally I have ewindisch@newrelic.com. If you want to reach out directly, find me on Twitter. Yeah, I guess those are the main places.

Jeremy: All right. And then newrelic.com just if you want to check out all the stuff they're doing with serverless there. Right?

Erica: Yeah.

Jeremy: Awesome. All right. Well, I will put all that into the show notes. Thanks again, Erica. Appreciate you being here.

Erica: Great. Thank you.

This episode is sponsored by Amazon Web Services: Check out the How to Use Objects in Amazon S3 to Trigger Automated Workflows Using AWS Lambda Learning Series.

View Details

About Sven Al Hamad

Sven Al Hamad is co-founder and CEO of Webiny Serverless CMS. Sven has worked with the largest media and ecommerce customers in Europe as their trusted advisor on the topics of web performance and architecture, and has a proven track record of successful delivery of several multi-million dollar projects for large enterprises. Sven is also an experienced entrepreneur who has acted as a CTO in 4 different startups.

  • Twitter: twitter.com/svenalhamad
  • Email: sven@webiny.com
  • Webiny: webiny.com
  • Webiny Twitter: twitter.com/WebinyPlatform
  • Github: github.com/webiny

Watch this episode on YouTube: https://youtu.be/9TSmOcLBr0k

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Sven Al Hamad. Hey Sven. Thanks for joining me.

Sven: Hey, Jeremy. Thanks for having me.

Jeremy: So you are the CEO and co-founder of Webiny. So why don't you tell the listeners a little bit about your background and what Webiny is?

Sven: Yeah. So well, in terms of my background, I'm a developer. I started maybe 20 years ago, even more coding websites and many other stuff along the way, but also worked in the enterprise world for several years and decided to start Webiny about a year and a half ago, maybe two years back. But now I'm more focused on the business side. And what Webiny is, it's essentially an open source framework for building full stack applications that deploy to serverless infrastructure like AWS Lambda and similar. So it's all about creating serverless solutions.

Jeremy: Awesome. All right. So that's what I want to talk to you about obviously the Webiny platform, what you can do with it. And I'd like to start by kind of going through more of the details, right? Because I think we get confused with maybe what a CMS is versus what a application platform is versus what a cloud provider is. I mean, so I think it can get very confusing if you don't dive deep into the docks and even I was looking at the Webiny site and I was like, "All right, it's a CMS, but it also has a framework for building things." And so I'd love to go through that. And then we can talk about a couple of other things just to get your insight on serverless, but let's start at the beginning. So why did you build Webiny? What was the thing that triggered that?

Sven: So I started researching the serverless market in general. That was late 2017 and the more I dived deeper into the potential of serverless and serverless infrastructure, I kind of understood that serverless has such a big potential that actually it can become the standard, how we are building all applications in the future. Essentially, if you want to build an application five years from now, you're going to build it on serverless infrastructure. Serverless is going to become that standard.

And at the same time, I looked at the market, in terms of the solutions that are available today to help you do that. Well, there was nothing that I could use out of the box to help me build an application. There are tools to help you monitor and deploy serverless applications, but there are no tools to help you build actually a full stack serverless application. So I saw a big opportunity there, but also at the same time, I had a web design development agency many years back where we were all testing different CMSs and different solutions on building things. So from that learning and that experience, I decided, "Okay, let's take that. Let's take the serverless market, which is new and has great potential. Let's build a solution for that market that also is open source at the same time, so it benefits the whole ecosystem and the community as a whole."

Jeremy: Yeah. And who among us, hasn't owned a web development company in the past? I know I did for about 12 years and honestly it's funny because we built a CMS. I mean, CMSs were, this was before WordPress, right? So I mean, WordPress comes along and it changes a bunch of things. But one thing that was always a pain for me, and I know we can get into this as part of what the tool offers. But it was things like building forms. If I had to build one more HTML form in my web development company, I mean, I was ready to just, I don't know, go stick my head in a closet or something like that because it was just, it's so tedious, it's so repetitive. And as one of the things I think we've done really well or we've done a good job of is we've abstracted away a lot of these things and tried to make it so that where we make the undifferentiated heavy lifting, much easier.

But a big part of that, and this is something that always scared me, every time you build a new project, it's less about the interface, it's less about, maybe the backend, it's a lot about the data, right? We want to make sure that our data is secure, that our data is backed up. So I'd love to just talk about the data model that you have built into Webiny, because again, it supports a bunch of different things. But could you explain that?

Sven: Yeah, so well just handling data in a serverless environment comes with its own challenges. We found that really early on. So we decided to go with a MongoDB, particularly MongoDB Atlas to store the data, but how you actually talk to a MongoDB database from a Lambda function it's not the best today. If you don't use some specialized solutions, but you also find a similar problem with MySQL. That's why you created the MySQL library to help you do that. But essentially, what we did with Webiny, we have this notion of multicloud in our minds and we didn't want to get locked into specific databases. And we built a data library, a library that handles how we talk to databases and how we model the data models actually inside Webiny.

And that library is also open source, it's called Komodo and through Komodo, we built pretty much all the data models that you see in Webiny today in all our applications. And the beauty is that with Komodo, on the other side, you have these adapters for MongoDB. We also have an adapter for MySQL that's not published there, but we're planning on also building an adapter for DynamoDB and things like that. So we had to kind of also put some constraints on the data model. We didn't want to make it fully no SQL because we also want to support SQL in the future. So we put some constraints there but Komodo is kind of lifting off all the complexities there for us as a user. And on the other side, what we found is another challenge with handling pretty much TCP connections to the database. So we built, if you look at our architecture, we built something called the database proxy, essentially a Lambda function to which all our Lambda functions, talk to.

And only that Lambda function has the actual connections to the data, but it's like a funnel. So you can kind of limit how many connections you send to the actual database, reducing the number of zombie connections and so on. Some database have really, really smart abilities. So you can programmatically tell it, "Kill this connection or open another connection." But MongoDB doesn't have that. So it's pretty much kind of a, just having that funnel really low and then queuing up all the requests there. So but yeah, the data model has being a challenging topic for us there was a lot of iteration, a lot of work we did in Webiny, but now we've got it into a place where we're certain you can build really big business applications with it and it's going to handle any type of a use case.

Jeremy: Right. And then the other piece of it is the access to the data. So you mentioned that proxy that you've built, but in order to get that data back down to the website components or to interact with it in the admin, that's all built with GraphQL?

Sven: Yes. So everything's done via GraphQL and there's a bunch of scopes. You can pretty much control who can access what, which service can talk to what. Also when you create users, there's hundreds of different scopes and settings you can control in terms of securing your data and what you want to expose and so on. I mean Apollo server helps greatly there, but you still need to put quite a lot of logic in front of it. And also if you look at our architecture, it's all micro services with the central Apollo server through which everything goes to. So there was some also engineering there. How do you connect everything? How do you deploy everything? When you deploy something, how do you update the Apollo server? But again, this is something that Webiny just handles for you under the hood. And if you're just putting a new microservice, the Apollo server will know about it and it will just update your schema and the new scopes will automatically be visible to you in the admin section, under the security module.

Jeremy: Right. And so another choice that was made was to use React and not just on the back end, but also the ability to deploy these as components on the front end. And I want to get into that deeper, but just what are some of the basic reasons for choosing React?

Sven: Well, the whole story, when we decided actually, just to take a step back, when we decided to build Webiny, it was like, "What stack do we want to use?" Right? So we then looked at the market, looked at the market surveys. What are the trending technologies people use, trending stacks and libraries, right? Because we wanted to kind of tap into an existing ecosystem of developers, but also make it easier for the majority of the market to adopt Webiny without having a really steep learning curve. And we found React being one of those libraries that a lot of people in companies and organizations were adopting. So we just figured, "Okay, that's kind of a direction we want to go with and when we did adopt React, when we started to learning it, it was a really pleasant experience, what I want to say. But also some people asked me, "Hey, can I use Vue or Angular or other bits with Webiny?" Even like, "Hey, can I build an interfacing Python for webinar?"

And the answer is, yes, you can. Because how we architected Webiny is we've got these notion of stacks. We have an API stack that holds all your backend logic, done in Node, fine. But then the app stack is done in React and you can replace our app stack with yours and then you can code the interface in any language or stack you want and hook in to the existing GraphQL, API endpoint, essentially.

Jeremy: Awesome. All right, well, let's get into the CMS part of this because I think when you originally built this, this seemed like that was the original intent. I know it's grown beyond that, which is awesome. But let's talk about the CMS. So there are a couple of different components let's start with the admin.

Sven: Yeah. So we started with the CMS because that's the area we knew the best and the admin is kind of the core of it. So when you install Webiny, you get this what we call an admin app. It's a whole user interface where, where you can expand it. It has a menu. It comes with bunch of ready-made components. And everything that you see on the screen in the admin side is expandable. You can change the logo, you can brand it, white label it, do anything you want, but essentially a starting point, if you want to build a new app, you would just add a new plugin to register your new app with the menu in the admin side. And then admin, when you click it, admin will do the routing and everything, but it's essentially the core through which you build in all other apps and then hook them in.

Jeremy: Right. And then you have the standard serverless use case of taking images and converting them and that sort of stuff. And that's all through the file manager?

Sven: Yes. So file manager comes, as a built in module together with the admin app. And handling files, we thought it's always, you just place an S3 bucket and then it works. So it doesn't ... So if you look at the architecture of just the file service, it's even more complex than the API and the front end together. So essentially we had to build a whole solution on the backend and the front end side. And that solution is called the File Manager. But essentially when you upload files into Webiny, we do this post data token exchange with S3 so that you don't upload files through a Lambda function, but they'll upload them directly to an S3 bucket and that bucket is private. Only that token allows you but also when you request the token, it goes through our security module, ensuring that you have the right to request a token.

So you can upload the files. When you upload a file, we detect if it's an image, if it is, then we have separate Lambda functions that create thumbnails out of it, or any other dimensions you want. There's also some DDoS protection behind that in place. And also when you delete a file, there's a trigger that invokes other Lambdas to clean up all the stuff there. So, yes, there's a lot of stuff behind it, as you can hear. But the beauty is that we built a solution and if you upload one image works great. If you upload 10,000 images, it still works great or even a hundred thousand. So it scales and that's the beauty that serverless brings. So ...

Jeremy: Yeah. And that's all stuff by the way, you do not want to build yourself. I've built little file upload managers and things not for commercial use or not for to share with other people, but just for my own internal projects and things I was working on. And that process of connecting Lambda functions and doing all of that post-processing. It's straightforward, but it is not easy. So it is definitely something that you can get wrong. So that's great that that's built. All right. So then another thing you have is a page builder. What can you do with that?

Sven: So a page builder actually, it's a simple drag and drop page builder for building landing pages. But the cool thing is that the stuff that you drag and drop are not static HTML widgets, but actually full blown React components. And it's all pluginable. So technically you can build a React component that holds your business logic and you can drag it and drop it, and it's going to get rendered. And that business logic is going to work. But also when you publish pages with the page builder they go through a service side render, a snapshot gets created, that snapshot gets automatically pre-populated on to the CDN. And when a user visits that page, it doesn't hit a Lambda function, but actually get served all for CDN. But in the background, what we then do is we rehydrate the application, meaning we asynchronously load all the JavaScript and then do the dirty checks if the page has been changed or not via an API call that goes through a Lambda function, but that's all in the background.

It doesn't slow the first paint time. It doesn't slow the user experience and so on. So the page builder might seem like, "Hey, you can build landing pages with it." It is a really powerful tool. You can build dev dashboards, business logic with it and so on. So yeah, essentially it's a piece of really powerful technology there. So ...

Jeremy: Right. And that does server-side rendering, right?

Sven: Yes. So every time you publish a page, we do the server-side rendering for you. That's a separate Lambda function, separate user flow, it happens all in the background. It's not server-side rendered on demand, but when you publish a page, you use the server-side render to create that snapshot because if you will do it on demand, then you might have bigger cold starts because the Lambda function that does server-side rendering has a lot of stuff in it. So you never want to do it in the main thread, in the main user flow, you want to kind of have snapshots stored on S3 or in the database in case of Webiny and just pull it really fast there and populate it onto the CDN. So ...

Jeremy: Yeah, love that. That's awesome. All right. So then the other thing is this form builder, and I love this because, I mean, even when ... What was the name of that form company? There was a couple of those places that just gave you simple drag and drop forms, things like that. But it was always kind of a pain because again, you're plugging it in and you're trying to do form validation and some of these other things. But you've got a very cool form builder and I know every time I have to build a form, now I just go find something else to do because I don't want to build a form, but this makes it super easy.

Sven: Yeah. So like you said, earlier, building forms is something nobody wants to do. So we wanted to make sure this is kind of the ultimate solution and we never have to build another form again but essentially yes, with a form builder, you drag and drop fields. You can select the field types and stuff like that. But then we also got really powerful validation features, set patterns, pre-made patterns, custom patterns and stuff like that. But also at the same time, how you building, dragging and dropping fields, you also can control rows and columns. So you have also how the layout is going to look like it's all fully mobile responsive that what you get at the end. And then you also provide you with a set of triggers, web hooks. There's a plugin for recapture, there's a plugin for accepting terms of service and stuff like that.

So we really wanted to kind of build a solution that works well. And then finally, there's also GraphQL API through which you can pull the data off Webiny from all submissions. So you don't need to kind of export them manually. Also, you can export the manually, but you can have a programmatic access to who submitted the form and then do additional processing if need be.

Jeremy: Yeah. That saves you so much time. It's crazy. All right. So then the other thing you just launched was a headless CMS, which is pretty cool.

Sven: Yes. So that's the last product that we just launched. As you can see, we have a lot of products and they're all open source, which is a cool thing and that they can be combined together. You can take the headless CMS and build a Gatsby site, right? But then you can also take the form builder and render the form inside Gatsby. Right? They can be combined together. But the headless CMS is one that first of all, it allows you to model your content like any other headless CMS, but it's also has multi language support, support for multiple environments. It supports aliases. So if you switch from production or version one to version two, you don't need to change anything with the client. The switch has done on the backend in Webiny and if you need to roll back, it's instantaneous and things like that.

And it's all done again on top of serverless infrastructure. So it scales really, really well. And if you're not building anything and you're not sending any requests through the headless CMS, well, your cost is zero. There's nothing that you need to pay for but it's got all the features that you can expect from your typical headless CMS.

Jeremy: Right. So another thing that is super important to everybody or should be super important to everyone is security. And that's another thing that this platform just has completely baked in.

Sven: Yes. So security comes in many flavors in Webiny, but also what we found is, especially when it comes to enterprise clients, they want to bring part of the security that they already have. For example, a lot of enterprises, they are locked in, in their user identity providers like Okta, Active Directory, things like that. So when they take on a new solution, right? That solution has to work with their existing security user pool providers and so on. So when we designed our security layer, we had that in mind. So out of the box when you install Webiny, Webiny creates a user pool using AWS Cognito, that's the default behavior. If you have an existing Cognito pool, you can use that as well, but you can also bring in something like Okta or Active Directory, as I mentioned. It's just a matter of fact of writing one or two small plugins and say, "Hey, my users are actually located there."

And then Webiny will automatically talk to that process, to that API, see if the user with the same email exists. If it does, it's going to do the token exchange, do the JWT token as well, and then pool all the roles and permissions from within Webiny. But if that user leaves your company, you deactivate the user in that user pool and Webiny will automatically deny him access. So that's kind of your starting point from the security. And then from that point, it's our standard ACL with all the scopes access to the API and everything that is kind of ...

Jeremy: Right. And you can use those identity pools and you can mix and match for doing things like backend access to the admin. If people want to do page management, or if you build other types of interfaces there, but then you can also use it on the front end. If you wanted to do an intranet or maybe a permission-based or, SAS that people can log into or something like that as well, right?

Sven: Exactly that. So we just had a big enterprise that adopted Webiny for their enterprise intranet solution. And of course they want to protect the public site from the actual internet, only that employees can access it from within. And they are looking to do that with Okta, actually putting Okta in front of the ... They're using the page builder. So in front of the page builder pages, and they can do that. So if a user accesses ... It doesn't have a single sign-on token, it will deny him access there. So everything you see in Webiny is either a plugin or a module, and you can, as you said, mix and match and combine things to work in any way you want.

Jeremy: Right. So speaking of plugins, I mean, that's an important piece of projects like this or products like this, is the ability for it to be extensible, right? And so, I mean, you can plug in all these different security providers and you have some of these internal, or these other things like the form builder and some of these other plugins. But what about more of a general plugin system? I mean, something like a WordPress plugin system, is that something that is happening with Webiny?

Sven: Yes. So it's a good question. It's definitely happening. It's in plan to have a plugin repository that other people can contribute fully open and free and we want to build essentially a community around the whole product and having a plugin repository is really important for us. We just haven't gotten to that stage that we have resources to invest in that area, but it's coming soon. And at the moment, if you want to expand Webiny with plugins, we already have a documentation for that. So later on, once you build those plugins, we've built the plugin repository, you can publish them there. And we also want to give exposure to people. If you've contributed to a plugin, when you click on it, it's going to go to your site, to you so that we also help our contributors get extra visibility there.

Jeremy: Right. And you already have a couple of plugins. I know you have the Google Tag Manager one, I think Facebook integrations, things like that.

Sven: Yes. So we've got couple of plugins already. The Google Tag Manager that you mentioned, so you can embed third party scripts into Webiny. But we also have MailChimp subscribe form. In the page builder, you can just drag and drop your MailChimp subscribe form. We're going to ask you for your MailChimp API key, then you're going to select the form and it's going to be rendered there. We also have the recapture plugin for the form builder and things like that. And if you look at the source codes, which is all public on our GitHub for those plugins, it's the identical way, how you would build any plugin that you need for your site.

Jeremy: Awesome. All right. So moving beyond the CNS piece of this, because again, I think whenever I see Webiny I just see Webiny CMS. I guess what I think of it's a serverless CMS, which is awesome, but you have the serverless web development framework. Explain what that is.

Sven: So you can call it a bad copy on our side in marketing Webiny as too much of a serverless CMS, but essentially yeah, I know. We were worried just to explain the backstory here. We were worried if having a serverless development framework would be too abstract for people. Having a saying a serverless CMS is something that it's much easier to wrap your head around and already picture what you can do with it, but actually where Webiny as a product is, it is that framework for building applications and actually all page builder, headless CMS they're all just example apps, what you can do with our serverless web development framework. And that is the service web development framework is actually our core product. So that is essentially the bit that you take as a foundation for building any type of a serverless web application.

Sven: And that foundation solves many of the serverless challenges and pitfalls like managing files, security, handling database connections. So that is what is in that framework. But also that framework gives you a really good structure for your project, regardless if you have a small project or really big project, it works both ways. And that framework also handles the deployments, the creation of serverless resources, the state files, it has a CLI to bootstrap projects. So that is essentially the core of Webiny. You install the framework, you can pretty much start your project right away, focusing on the business logic. You don't need to kind of do many configuration and bootstrap and stuff around it.

Jeremy: Right. And that's all the stuff that you likely do not want to do anyways. It's just going to slow you down-

Sven: Exactly.

Jeremy: Every time I start a new project, I just am like, "All right, I know I have some, sort of bootstrap templates here and I can and do some of this, but yeah, that's amazing. So that is again, similar to the CMS portion of this, it uses React UI components. You can build your own. You've got that login component, right? It's just drag and drop and now you've got that login that ties into all those other things. Yeah, I mean, that is just that, it's just a very cool tool.

Sven: Thank you. It took us a lot of time to actually build it, but also when people approach this, they think, "Oh, I can do this by myself." Sure, you can, but do not underestimate the challenges and pitfalls that come with serverless. It is so, so, so hard to get some of those components right. It took us over a year and a half to get some of those things right, right? They might seem trivial, but they are not. But also-

Jeremy: They're not.

Sven: ... what people think is, especially developers within organizations, they build a serverless solution on their own, right? But suddenly now the organization wants to do many more serverless projects. So how do you scale the knowledge? Right? How do you kind of do many different teams with many different projects that all run on serverless? Unless you have a solution that has documentation, that is proven that it can scale, it's not going to work. So yeah. I mean, that's kind of the pain we want to solve.

Jeremy: Yeah. And I totally agree. I mean, that's one of those things that if you can certainly build this out yourself and you can go through that process and spend all that time and learn all that stuff but if you don't document that learning and understand why X works better than Y or whatever it is, if you don't do that, then you're right. Passing that knowledge on to somebody else is really, really, really hard. So having that all encapsulated in one project where again, you can divert from that too. If you have to kind of eject and go a different way or whatever, there's still a lot of capabilities in there. But yeah, no, I think that's hugely important. And the other thing I'll just say is proof of concepts or a proof of concept in serverless is typically very easy to create.

And in most cases, depending on the complexity, that'll just scale, right? Which is one of the nice things about serverless but it might not be the most efficient way to do it. It might not be the smartest way to do it. The most secure way to do it or whatever. So again, having those best practices baked in and sort of done for you is amazing.

Sven: You touched on a really good point. I think that easiness of serverless is one of it's greatest virtues, but also one of its greatest pitfalls, right? So you don't get a problem right away. You don't see it now, but the moment there's more pieces to it, suddenly now I have to refactor a microservices application that's humongous. It's you don't want-

Jeremy: Yeah, exactly.

Sven: ... to get tangled in that, yeah.

Jeremy: Right. All right. So I want to move on to multicloud because that is something that the website for Webiny talks about and the ability to deploy to multiple clouds. So what is the multiple or multicloud support like in Webiny right now?

Sven: Yeah. So today, Webiny, you can only deploy to AWS and that's intentional. But when we started building Webiny, we had this requirement of multicloud from the beginning. But we didn't want to kind of make it available to all cloud providers at the same time. But what we did, we built these abstraction layers towards cloud providers but behind those abstraction layers for each cloud provider, you would need to kind of build a driver for it. Let's say it like that, but we only build the drivers for AWS now, but if we want to move to GCP or Azure or somebody else, we can just add additional drivers. The actual core code doesn't need to change. So the multicloud support is built in, but without the deployment drivers for other clouds at the moment, purely because having those deployment drivers takes a lot of effort and we figured, "Okay, let's go with AWS. It holds 75% of the market today." So that's where the biggest user base is. Let's get it running there. And then if we see a lot of requests or ask for, I don't know, Azure, for example, then we're going to focus on it. But until we see that need, we will kind of just let the community to decide on it. And we will waste our effort to where the community tells us they need us now."

So it's just kind of that's the train of thought, but also why we were approaching multicloud is it's likely, from a different angle, it's not about resiliency or technology or things like that. I've worked in an enterprise and I've learned that at the end of the day, who you're going to choose is all about money, right? But often how, for example, you start using AWS today and your bill is a hundred dollars, but if your business grows, you might pay a hundred thousand dollars to AWS and suddenly you want to move away, but you can't. You super sticky because you didn't have those abstraction layers, right. But the other thing you can do is try to negotiate your pricing with AWS, which unless you can't move away, they won't budge, right? But if you have that possibility to move away, suddenly your business has a leverage to negotiate and some businesses pay millions if not billions, to AWS. And having that opportunity to do multi-cloud is really, really important.

Jeremy: Yeah. No, I love to believe that the negotiation and the leverage of having a huge AWS bill would be helpful, but there was just that recent article that was sort of about, you get companies like Dropbox moving off of AWS because the bill's too high or whatever. And maybe AWS loses a hundred million dollars from a giant customer for a $4 billion run rate or whatever it is. It's a small fraction of that. I do agree with you though, that I think that having the flexibility if you're not compromising the underlying services is hugely important, because you don't want that lowest common denominator ... So if you did want though to do let's say I am on Azure and I say, "You know what? I really want to use this Webiny CMS, can I ... I mean, there's hooks in there, right? Aren't there hooks? I could build that in. I mean, it'd be a lot of work, but I could build that in and be able to connect to those other things because of the way that this is architected.

Sven: Exactly that. So you can literally hook in into any part of Webiny, but also if you look at just how the AWS components for deployment are done, it's pretty much just using the native SDK. So if you want to not use AWS Lambda, but to deploy it to the Azure functions, you would just write a small adapter that creates those resources the same way. But also Webiny doesn't kind of lock you in, even if you do use AWS, if you have certain services from Azure, you can combine stuff, right? You have access to the code, you have access to the deployment hooks. You can pretty much customize it to fit any needs. So ...

Jeremy: Yeah. So another question about multicloud because I think if you're paying a hundred dollars a month to AWS that if you're thinking to yourself, "Whoa, this could get expensive. I mean, I really should be thinking about trying to be cloud agnostic or some way in which I can move things around." Your thoughts: should small businesses, startups, things like that, should they really be thinking about multicloud at this point?

Sven: Not at all. So that is a concern for what I would say for enterprises and larger organizations. If you're small, I mean start with being sticky, that's fine because your priorities are about building your business, not controlling the cloud cost, I would say. But be wary that controlling the cloud costs will come at one point. But don't let it constrain you in the beginning, especially potentially just constrain your iteration cycles, making them slower because you're suddenly working on multicloud while you don't need it, right? So at this scale, at the small scale, don't waste your time there. But have it in mind if you plan to go big at one point. So, but only invest those resources once you have those resources and that is a priority.

Jeremy: Right. And I would say if you're building a serverless application and it gets to the point where that gets to be really, really expensive, you're probably doing something right there. So hiring a few extra engineers, if you need to start diversifying clouds probably wouldn't be that difficult at that point. All right, so I want to talk to you though about this idea of just why serverless, right? So let's say I'm a regular everyday person and I'm trying to figure out how to build something for a client, or I'm trying to build something for myself, why not just WordPress on WP engine or something like that? What's the underlying benefit to building my site or my service on serverless, even if I'm not expecting a ton of traffic, maybe.

Sven: I mean, how I see it. First of all, building on serverless might seem slightly complex at one point, comparing to VP engine because it's two mouse clicks right? But the thing is with serverless, first of all, your cost is way more efficient. If your service is not having any traffic, why should you pay for it? With serverless, you will essentially stop paying for stuff you don't use. It's an important thing for especially small businesses, being efficient at how they spend their money, but also imagine your site suddenly gets on the front page of Hacker News. By the way, we were, with Webiny on front page of Hacker News.

And that sends a ton of traffic. And it's a really big spike. And that is when you need your servers or whatever you're running on the most. You do not want to go down when you have 50,000 eyeballs on you. That's when you want to deliver. And that's where VP engine and others will fail short on you. They won't scale to those demands while serverless doesn't care about it. It just works. Of course be prepared for a bill shock, but hopefully it's going to be worth it for you. So ...

Jeremy: Right. Yeah. I mean, and I think that if you think about something like DigitalOcean, for example, I mean, it's very easy now. They've got these engines in place where, you just go on and say, "Oh, I want to create a WordPress site and you click a couple of buttons and it's magically there, but Webiny is getting pretty close to that, right?

Sven: Yeah. So at this stage, Webiny requires you to kind of put in API credentials for AWS, connect to database, in one command deploys and creates all the resources. But we're also preparing another offering, which is going to be one of our commercial offerings where essentially you're going to get those two buttons and you're going to have a Webiny site deployed to your AWS cloud. So we won't be hosting anything, but we will be providing a user interface to make it really easy for you to deploy Webiny sites, have multiple environments, multiple projects everything would just kind of mouse clicks. Make it super, super easy.

Because no matter how advanced a certain technology is, if it doesn't get to that level, that it's really easy for even non-developers at one point, because a lot of people ask us, "Hey, I really love your page builder, but I'm not a developer. With WordPress is just two mouse clicks." Serverless is not there today but Webiny we're maybe one step behind that. We are going to be there pretty, pretty soon and hopefully that the serverless market would kind of benefit from that and the adoption will also grow.

Jeremy: Right. So the other thing that I'd love for you, sort of get your input on, or your perspective is I often see arguments saying, "Oh, serverless is great for spiky workloads, but if you don't have spiky workloads and you have steady traffic, it's just easier to put it on a VM or something like that. But I always find that serverless has benefits, whether you're a small business or you're a huge enterprise. What are your thoughts on that?

Sven: So what I'm seeing is, so when you're a small business, it's the cost of infrastructure that really matters to you because it's really efficient. You don't pay if you're not using it. But for the big guys, it's a combination of factors. And sure your bill might be slightly higher in some cases running on serverless, the cost of infrastructure. But the cost of managing infrastructure will go way down. You will have to hire less people, or the people you have will have to spend less hours working there.

But also what that does, it releases a big chunk of the budget or resources or man hours that you can now focus on product iterations. So your product can grow faster. And if your product grows faster, you can out innovate potentially your competitors, which can't afford that same level of innovation. So what I see with enterprises is that they see serverless as a competitive advantage, and that's why they moving to serverless. Although, you see all the blog posts about cost savings and stuff like that. Yes, that's true, but there's that agenda of outpacing my competitor, which serverless actually unlocks. And the moment you migrate to serverless, you can use that potential.

Jeremy: All right. And then what about the big business? So what are the benefits for those larger enterprises?

Sven: Essentially that. It's just having to spend less on managing that infrastructure. Sure, the infrastructure costs might be higher, but the cost of managing will be drastically lower because there's way less things to manage than if you had a fleet of hundreds of thousands of servers or containers, that all requires man hours, and you as an organization, you should strive towards autonomy, putting the resources where they really matter and managing servers well matters if you have to have them up and running, but if they suddenly go away you can focus on other bits.

Jeremy: Right, yeah. And I mean, I think it's crazy to have all of these people who specialize in running data centers and building software for the data centers and managing that for you and then say, "You know what? We're going to not let you do that and we're going to have other people work on managing that and installing patches and trying to figure that stuff out." I mean just the amount of effort and time and energy that goes into that, just backing up databases, just having a DBA that does your database backups and optimizes the database and migrates data. It's just all taken care of for you, if you choose the right tools.

Sven: Exactly that. And especially, I believe the younger businesses that are starting today and building their technology have a tremendous advantage because they have such amount of brilliant tools in front of them that they just need to kind of piece together and starting maybe five, 10 years back, it's just networking load balancers. For me, mentioning that, I just get depressed-

Jeremy: Gives you flashbacks. Yeah, I totally agree.

Sven: But not positive ones so ...
Jeremy: Right, yeah. You can only wrestle with security groups to an RDS and since so many times before you were just like, "You know what? I don't want to do this anymore." All right. So let's talk about this idea of who you're targeting because this is something for me that was I think the first question I asked you when you ... I think you and I started chatting maybe a year ago or two years ago, whatever it was, and I was like, "Well, who are you targeting with this?" Because if you ask the average developer, they don't even know what serverless is. I mean, even the average cloud developer, they're like, "What's serverless? And again, it's probably not the best term. We know that. You ask somebody off the street, maybe some small business owner or somebody who's like, "Oh, I'm going to start a little side business."

And they want to create a website. I mean, the first thing they're going to do is they're going to go for WordPress. So who are you trying to target with this? And is this, I guess maybe the question is, is this trying to be a competitor to WordPress or is this taking a different angle?

Sven: So there's two answers to this question. One is that is about the short term mission and vision. And one is about the longterm. Short term, yes, if I ask many of the developers out there, what serverless is, they won't know about it. And those users are not our target users. We are letting the big cloud providers educate the market about serverless and let them spend millions and millions, if not, billions, on educating the market. But once the market is educated, that's when those users will start looking for ready-made serverless solutions. And that's where they're going to find Webiny.

At this stage, we're targeting developers that understand what serverless is because educating them just takes too much resources for us, but the moment an engineer or developer and knows what serverless is, and you show him Webiny, he just gets super excited because he knows serverless comes with certain challenges that with Webiny they just go out of the window. And suddenly I unlocked the potential, just writing business logic and all gets deployed into microservices. And I don't care about how it all works under the hood, but I still utilize all the power of serverless.

So those developers are people that we're talking to today. But in the future, how I see the future in five or 10 years as I started, as I said earlier in the session is I see serverless as being the standard of how all web applications are built. But don't get me wrong. The traditional servers won't go away, but they will be used for specialized purposes, machine learning and things like that but for anything that's event driven that runs on the web, serverless shines there. So I see a clear future where serverless is that future, where it's the standard of how we build the web.

But that standard it's not enough to solve it on the infrastructure orchestration layer. You have to solve it also on the application layer because people are after, when they say, I want to build a serverless solution, they're after. They're thinking in their heads, the application layer. And that's where we with Webiny come in. We're providing the tools on the application layer to help them build their business logic. And de facto, a CMS is one of those tools out of the box, but they can innovate other tools. But if you think about WordPress for a second, WordPress powers 34% of the internet, maybe even more at this stage.

And there's two reasons for that. One is they made it super easy to install and run and integrate with many other systems. But secondly, WordPress, although like a lot of people perceive it as a blogging platform. If you open their documentation, it's going to be clear. It's not a blogging platform. It is a foundation for building applications, but those applications run only on virtual machines and those traditional architectures. With Webiny, we want to be something like that, that foundation, but for the serverless world, for when the serverless is that standard. So will we replace WordPress? I would like to do that, but I think that's a mission for the next 30, 40, 50 years because Matt did an amazing job with WordPress but the mark is going to change because if you think about it, the Nokia phone, everybody thought that that's never going to go away. And then-

Jeremy: I mean, playing snake on your Nokia phone was amazing.

Sven: Yeah, but-

Jeremy: I mean, the graphics were intense.
Sven: It looked so real. That is the perception of the virtual servers and containers today. And then serverless is that smartphone. It's going to have a slightly longer adoption curve, but that's how I see the market. That's how I see this evolution coming.

Jeremy: Yeah. And I think we'll see even sites Wix and GoDaddy, I think they use, maybe they use WordPress, but WordPress and other platforms like that starting to migrate to more serverless technologies on the backend anyway. So I think that's great because if the Webiny can help influence that and set those standards and help sort of drive that ecosystem, I think that is absolutely amazing. So all right, I have two questions for you from our Serverless Chats Insider. So this is a new thing we're doing. So I'm going to throw these at you. They're pretty easy. They're softball questions for you, but if anybody wants to ask questions, join the Serverless Insiders email list and you can ask questions.

So the first question is from Michael and they said, on your website you say that MongoDB is serverless, as long as you don't have to manage it yourself. And they said they haven't fully defined the term yet. I don't think anybody has or try to. And they're curious whether or not, where you would draw the line, would you consider a COBOL mainframe to be serverless if somebody else is managing it for you?

Sven: So how I perceive serverless is if there's a service I'm consuming and I don't need to worry about servers, patches, networking. If I'm locked out of the thing that the service runs on, but I can only interface with the service via an API, that serverless for me. And what we say on our website is it's specifically targeted for MongoDB Atlas. So that is the cloud offering that MongoDB has. And there it's a mouse click. I need a Mongo database that is, I need the Mongo cluster deployed to this region and they manage everything. They update my MongoDB, they scale it, because it has an auto-scale function. I don't need to do any load balancers, any networking stuff there. Backups are just a one checkbox. So its essence, if you look at MongoDB Atlas, and if you look at Aurora serverless, it's not very different. And this one has even serverless in the name. So as long as you don't worry about the servers and all the bits that connect to that, that's how I perceive a serverless.

Jeremy: Right. And I've been trying to define serverless for many, many years, and I have not come to it either. So don't worry about it, Michael. None of us know what it is. All right. So another question, this one's from Mark and you kind of addressed this in the beginning, but when do you think Webiny will fully support DynamoDB as a data store?

Sven: So we get that question asked a lot. So we would like to have DynamoDB, and it is one of the items really high up on our priority list. We're looking at it, we're working on it. That the challenge with DynamoDB is that although they say they're no SQL database, they are so different than a no SQL database. It is a third category. You have SQL, no SQL, and then you have DynamoDB. Purely because they're so different, we have to work around some of its quirks and challenges and we're working on it. We hoping to get DynamoDB support sometime this year because a lot of users have asked for it and we personally also would like to use it. So stay tuned.

Jeremy: Awesome. All right. And then he had one followup question as well. And that was about the business model for Webiny. And you had mentioned, some sort of business where you'd be able to deploy it for people. But the question sort of says, because serverless is so cheap to host yourself, or do you think that there's any way that there could be a managed hosting service that you provided that would be profitable?

Sven: So it would be really hard to do a profitable managed hosting service there because if you look at for example, Netlify, they're a managed AWS hosting service. Sorry, but they are. And I don't see us going in that direction. So what we are providing is, where we're going to probably call the ... It's not yet a full name yet, but we call it for now Webiny Control Panel. Essentially, you would hook in your AWS API user. And that interface would be one-click deployment of Webiny, but to your own AWS cloud. So you can kind of create 20 different Webiny projects and there's going to be reporting, monitoring, audit logs, and bunch of features in there that compliment to the features you have in AWS, but simplify it and scope it to your Webiny project.

Because the problem with AWS is, "Hey, this region has costed you 500 bucks." Yeah. But what is the cost per project? Good luck finding that out. So that's what we're going to be providing with this commercial offering and it's going to also have a free tier and there's going to be a really cheap tier and bit more expensive tiers for more enterprise business users. And I understand the background of this question. A lot of people approach us, "Hey, you're just open source, giving everything for free. You're going to go bankrupt in three months." We have a plan people. So ...

Jeremy: Well that's good to know. Well, that's awesome. So Sven thank you so much for building Webiny and for everybody that's contributing to that project. I think this is the right direction that we're going, trying to abstract that away. Get to a point where it's just more approachable and like you said, thinking at that application level. So again, thank you for that. If people want to find out more about Webiny or contact you find out more about you, how do they do that?

Sven: So you can find everything about Webiny on our website, webiny.com and on our GitHub, which is also github.com/webiny. And you can contact us on Twitter, on the Webiny platform. If you want to reach out to me personally, sven@webiny.com is my email and you can find me under @svenalhamad at Twitter. And again, I also want to kind of thank you Jeremy for having me as a guest here, it was a pleasure. And I also want to give a shout out to the whole Webiny community. That is kind of what keeps us going and why we're putting all this effort because we really love our community and our community helps us grow at the same time.

Jeremy: Awesome. All right. Well, we will get all of that contact information into the show notes. Thanks again Sven.

Sven: Thank you, Jeremy. Thanks everyone.

View Details

About Rafal WilinskiRafal Wilinski is the founder of Dynobase, a professional GUI Client for DynamoDB with mission to onboard thousands of developers to Cloud-native and Serverless world. In addition to founding Dynobase, Rafal distributes a weekly newsletter “This Week in DynamoDB” and is a Serverless Engineer at Stedi. Prior to his current roles, Rafal got his start in developing mobile games, and is a cloud native engineer, AWS certified architect, and Serverless Framework contributor and maintainer.

  • Twitter: twitter.com/rafalwilinski
  • Dynobase: dynobase.dev/
  • This Week in DynamoDB: dynobase.dev/newsletter/

Watch this episode on YouTube: https://youtu.be/41YJAflfnP4
Transcript

Jeremy: Hi everyone! I'm Jeremy Daly, and this is Serverless Chats. Today I'm chatting with Rafal Wilinski. Hey Rafal, thanks for joining me.

Rafal: Hi, thanks for having me.

Jeremy: So you are the creator of Dynobase and an independent AWS consultant. Can you tell the listeners a little bit about yourself and what Dynobase is all about?

Rafal: Yeah, sure. As you mentioned, I'm founder of Dynobase, professional graphical interface for DynamoDB, and also right now an independent AWS consultant, mostly focusing on serverless solutions. I'm deeply passionate about AWS and mostly serverless, since, I think, 2016, because I attended the first Serverless Conf back in London, and that's why I became so excited about this whole field. I'm running my own blog which is called Servicefull, because we had so many discussions about what serverless is and how bad a name that is for a paradigm of technology.

So I decided to actually steal the term coined by Patrick Debois, which is Servicefull, because full of services. I'm writing about serverless, about cloud, and I recently merged that with my own page, but you can still go there. It's going to redirect you. Before going all in into AWS, I was actually making mobile games. I've made a few of them, one of them became even quite popular, which was called Voxel Rush. But my parents were saying that making mobile games isn't a real job, so that's why I transitioned into making web and cloud, and now I'm here. Less than a year ago, I started Dynobase, which was trying to solve my problems with UX and UI with DynamoDB, but I guess we'll talk about it a little bit later.

Jeremy: Right. Yeah. And actually, that's why I want to talk to you about today is Dynobase. So you and I have been communicating for quite some time. You know I'm a huge fan of DynamoDB. I love just the scale of it. I love what you can do with it. Rick Houlihan opened up, I think, everyone's mind or minds with what you can do in regards to relational structures in there and how you can access data in different ways in the single table design stuff, which is quite fascinating. But I really want to get into Dynobase, what it does, what's the purpose of it, but maybe we just start at the beginning. Why did you create Dynobase?

Rafal: Yeah, sure. Actually, there is sort of I think quite interesting story behind it because it all started more than a year ago when I was working at X-Team. I was working on kind of quite a big project for the educational space. I was working as AWS DevOps engineer, I was setting up infrastructures, it was all set up on containers. But we received a new requirement to create a fully real time community platform. Our architects, Raynard, which is a great friend and also an engineer, approached me and said that, "Hey, I know that you're super interested in serverless. I know that you've contributed some pieces to serverless framework, and you write about it. So maybe we can try evaluating all those new tools that AWS released, like AppSync, Amplify, DynamoDB, and try to create something on top of that."

And that sounded super good because I could finally use serverless technologies at my day job. So we immediately rushed into evaluating those tools. We watched a lot of random session including Rick Houlihan about single table design and it was mind blowing and charming at the same time. But when we were evaluating those tools, we realized that if you would like to use AppSync, we probably have to use Amplify and if we go with Amplify, we can't go with single table design, because if you're using Amplify, you're creating graphical schema, which is then translated to separate DynamoDB tables, which is not working with single table design super good. So we decided to use serverless framework DynamoDB single table design to create everything. And we rushed into the implementation without having thought all about testing processes, about debugging because we decided to learn as we go because we had that opportunity.

And when we started implementation, obviously, there was a lot of bugs and there was a lot of mistakes. We had our single table design schema, designed very good, but it was a new field for us. So we obviously committed a lot of beginner mistakes. While we were doing that, we started checking a lot of data, a lot of records inside DynamoDB just in console, because we needed to check if that specific record was inserted correctly if the data that it's inside database is actually good, or we wanted to modify some records manually. And while I was doing that, I realized that I'm spending definitely too much time switching between browsers, between regions, between AWS accounts, because, for instance, you can't have open two separate AWS regions, two separate AWS consoles for two regions in one browser. It was actually a pain. The same for regions. So you had to actually kind of hack your browser. You had to have opened many browsers and it was just super messy. There was no bookmarks. There was no history. Scanning speed was quite bad.

So I decided, I'm an engineer, I don't like wasting my time fighting with software. I like automation, I like solving stuff. So I decided to hack something really quickly using React and Electron, like a tool which is going to allow me just query in a little bit easier fashion, and it's going to be easier. Yeah. And I've asked a few of my friends who are also working with DynamoDB, which are engineers, if they are sharing the same pains as I am, and it appeared that a lot of people actually have the same problem of DynamoDB, that the DynamoDB itself is super good. It's super powerful database. It's enabled things that you could never imagine before. But the way that you access it, the way you modify records directly when you debug it, it's not super good, especially if you're working with the local version of DynamoDB. You can't easily access what's inside.

And while I had that confirmation that it's not only my own problem because I initially wanted to open source the solution. I felt that, hey, I was working in open source space for the past five or even more years, I've committed so many lines of code, I've committed so many things to serverless framework and to other tools, and I haven't received any money for it. It was greedy approach, but, you know. So, I decided actually to turn it into a product and try creating a business of it because it seemed that many people would pay for solving their problem. And when I realized that it can actually work. I had this moment of big excitement and I gained a lot of power just to work tirelessly, even after my day job, just to finish this UI. I think I did the first prototype in a month.

It took me something like 100 hours. I was about to release it. I announced on Twitter to everyone like, "Hey, in a week, I'm going to show you something really good I've created ... to contract with myself and with my audience that DynamoDB is about to get much, much easier." And before releasing the alpha version, while browsing Twitter casually one day, I realized that AWS just released NoSQL workbench.

Jeremy: Right.

Rafal: And the moment I saw NoSQL workbench, I was devastated because I've spent just the last one month or two. and tireless countless hours, and so many nights working on something that I truly believed in. And then comes AWS with some of the world's best engineers and working on the same idea at the same time probably killing your idea because if AWS is doing the same thing as you, they probably made it better, right? So even without the loading the software, even without starting it, the day that NoSQL workbench appeared, I was devastated. But the next day actually I decided like, "Hey, maybe let's try, maybe let's download it, maybe let's run." And one thing I've realized is that NoSQL workbench appeared to be a super good tool for designing your data model, for designing your single table design, for doing all the things that you're doing before actually jumping into implementation and actually working in development.

While my pains were a little bit different. We already had our single table design set up. We only had problems with inspecting the data that is already in the table. And I felt that my NoSQL workbench was actually replicating many of the quirks and different weird things from the AWS console. So, after trying it out I'd say, "Hey, it might sound similar, but these are two definitely different products, they are solving different use cases." And that feeling gave me enough confidence to actually push forward and release the alpha. And I released it, I think like few days after NoSQL workbench, which might sound a little bit weird, and you know, AWS released day too, I released my own, but, hey. And their tool is free; mine is paid, so this is even more weird.

And I release it and after I released it, I've let it go. I said, "Hey, I can finally take a breath. Now I can just rest and see how the money is flowing." And you know, guess what? The product wasn't super polished. And not many people were interested in buying it when you have a free alternative, right? I've gradually started to implement the fixes, the changes, because obviously it was not a super polished product at the very beginning. But somewhere around beginning of this year, one of the customers approached me saying, "Hey, you've created quite promising piece of software, but I think you can do much more when it comes to productization of it when it comes to the growth, to the sales, to the strategy." And he proposed me a quite ridiculous thing, which is, "I can help you with that, maybe we can become co-founders." And without actually thinking too much I've jumped on a call with this guy and he seemed fine.

So I decided like, "Hey, what's the worst thing that can happen? I can only just waste some time, but there's so much to gain." And we decided to actually collaborate, we signed a contract which was less than one page saying that all the expenses and all the revenues is split 50/50. I take care of code, he is taking care of growth, and we started working. We designed the application from scratch. We started thinking about both power users and users who don't know what's even a GSI.

Jeremy: Sure.

Rafal: And we released a version 2 and I couldn't be more satisfied from what we've created so far. And I'm super happy about the state of Dynobase. It's not only helping me, it's helping a lot of my friends. There are already over hundreds of engineers using Dynobase. There are enterprises and teams saying, "Hey, this is great." And yeah, it was a super great thing to make and each of the single dollar I make on the Dynobase is much more satisfying than even $1,000 you make consulting from your day to day job. And that's how it rolling right now.

Jeremy: Right. Awesome. Well, I mean, that's a great story. I mean, and that's one of those things, too, where I mean, I'm a big fan of open source. And I've done the same thing. I've put a lot of code in open source. And it's great to see people benefit from it. And you wouldn't be a true AWS user if they didn't rebuild something that you already built, too. So that is a common tale. You implement something and then AWS comes up with something that just solves it for you. But, so, I think what's really interesting about what you're doing with this product, again, is you're taking a different tact from what I think NoSQL workbench is doing. Because it's very much so design specific and I think you have a lot of those tools as well within Dynobase. But what about the differences between the console because exploring data is just a super pain. So what am I going to see differently if I'm using Dynamo, sorry if I'm using DynamoDB console versus using Dynobase?

Rafal: Yeah, sure. So I think the first thing I've mentioned during my story is that working with multiple regions, multiple accounts, and multiple tables is much easier because we actually have tried to replicate the experience that you get in a regular internet browser. You can really easily switch the profiles, switch the region, switch the table, you can open multiple tabs, you can take a look at 20 of your tables across many regions, without switching contexts, logging out, logging in. You know, it's a pain. So that's the first thing you can much faster access your data.

The second thing is that we know that many people are not super aware of DynamoDB specifics for instance, indexes. So we decided to abstract away the concept of indexes, LSIs, GSIs, and stuff like that. The way you query data inside Dynobase is like, for instance, you would like to get a user with an email johndoe.com, let's say that's an example. And as a beginner DynamoDB, you don't know if that field is actually already indexed, because probably the table was provisioned on someone else. You don't know how it works. So the only thing that you know is that you would like to see all the records with the association email, johndoe.com. And in DynamoDB console, you have to switch between query and scan. And I even don't know what this skip query and scan is. I have to choose some kind of indexes. I don't know what's any of that. So what Dynobase is, is that it automatically figures out if we can use query instead of scan because it's much faster for the given attributes, for the things that you're looking for.

You can enter many attributes and we can find what's there fastest way to actually get that data. Once you have those fields filled and we are telling that, "Hey, we will be using query instead of scans to find this data because it's much easier." Also, in AWS console, you are capped to 100 items and you have to switch, you have to click the arrow to get the next page. And it's also slow. If you're running a scan, it sometimes can take even hours. Our solution is much faster simply because we use different algorithm used, we use different pagination settings, we are fetching, we are using search, sorry, scan segments, and it makes this whole thing much, much faster. The last thing is also editing the data because in DynamoDB console, you have to open this weird model and in Dynobase you can also edit the data in similar fashion, but you also have the same editor that you see in Visual Studio code.

Visual's great IDE. So when you click just edit this item, you see the same editor you see in VS code, and you can edit the JSON directly. But you can also click on attribute, double click it, and change the value and save it. We are also doing what most database tools is doing, that we are not immediately committing all the changes to the database. We are just kind of dry running, you're making modifications on your data. Once you've made modifications, for instance, 10 records or 10 entities, you can then decide to save that and I think it's just much safer. Yeah, and we have also a variety of other tools that our DynamoDB console does not have. We have a history of queries, we have bookmarks, we are generating the code, because sometimes you don't know how to write those expressions attribute values, expression attribute names. I remember the first day I started messing up with DynamoDB SDK, I had no idea how to write a single filter expression, right? And I think that's the problem many developers have. They know how to write SQL, but they don't know how the API works.

Jeremy: Right.

Rafal: Yeah.

Jeremy: So, yeah. So, that I mean, that's one of those things to where I just find that to be super interesting. Where it's like, there's all these little tiny features that can make a product better. And they're simple things, like the query history. That's one thing that drives me nuts, is that I'll be in the console, and I'll search for some ID, and then I'll get something back. And then you can't just open a new window with another thing or try to use a new tab to open something new. It's all just embedded. It's really tough to use sometimes. And then you have to go back and make all these different queries. And then you might have to go back, find something, change a record, or whatever.

And it usually involves opening, again, multiple windows and things like that. So I love how that just speeds up that development time or that debugging time, as you said. So, you mentioned this idea of writing queries and that you have the ability to generate the code for them and Dyno and the NoSQL workbench does as well, if you've sort of gone through a more, I think, more lengthy process. But I think just in general, it is not super easy for someone to just write a query. And that's just one of the challenges that I think developers face. So what other challenges do you see DynamoDB developers facing as they're trying to build out a solution?

Rafal: I think the there is just one big challenge. And the challenge is interconnected with many smaller challenges. And the challenge is to definitely change your way of thinking from the relational thinking because many developers use to work with Postgres with MySQL. Then you go to the Dynobase, well, to DynamoDB space, and there are no joints. You can't cross reference some tables, you probably hear from some people saying crazy things that you should put all your entities into one table. And it's crazy, right? The first time I've heard that I should put all the data inside one table, it was really crazy. So there's a educational gap. I think this is a challenge for developers. And there is the required, you need to think differently. You need to unlearn what you've learned already about relational databases. And yeah, we think that they're still needed, we still need educational resources. We still need tools and there is a massive gap to be closed. Yeah.

Jeremy: Yeah. I mean, it's not just the education either. I think that there's material out there, whether it's Rick Houlihan's, videos or whether it's Alex DeBrie's book now, there's some training courses. Certainly it's an investment to learn DynamoDB and do that. But what about the tooling? I mean, obviously Dynobase is one tool, NoSQL workbench is another tool. Are more tools needed do you think for people to really embrace DynamoDB?

Rafal: That's a good question. I think NoSQL workbench was definitely needed. Because if you are aiming to create single table design, that's definitely make it a little bit easier for you. More tooling, I think that, yeah, we still need more tooling because I kind of treat DynamoDB as a low level service. I mean, in the future I would like to see some kind of abstraction over the single table design because if you're already inspecting the table and see all the different entities in one table, it kind of feels wrong, even after Working with DynamoDB. I would like to see some kind of abstraction, which is separating all that.

We need better tooling. And I think that also your modules are also solving that problem. For instance-

Jeremy: Trying to.

Rafal: ... Dynamo toolbox. It's a brilliant tool, which is mapping the attributes and yeah, its super useful. And I think that also, I think you also wrote it on Twitter once that there is always need for more resources and for more tools because the same phrase rephrased in a different way can resonate with a person that is reading it, can finally click for that person. So the more resources we have, the more tools we have, the more freedom we have. And yeah, that's only a good thing, right?

Jeremy: Right. Yeah. I actually think that's a... I love this idea of repackaging, even if it's the same content, but just slightly a different way or a tool that works a different way. And as you mentioned, the DynamoDB toolbox is only for Node.js, right? It's just JavaScript. So that solves a specific problem for a specific group of people. But there's Python utilities, there's Java utilities, there's other things and there needs to be more because not everybody likes to deal with the same levels of abstraction. So I totally agree with that. So where is Dynobase going, though? I mean, what's the roadmap look like? Are you planning on doing some of these other abstractions or what's on the roadmap for you guys?

Rafal: Yeah, sure. So we also think we have kind of a mission of closing this gap between how great DynamoDB is and how few developers already know about it. We are aiming to solve that issue by two things, which is tooling and education. When it comes to education, we are constantly repackaging the content. I mean, like we have great sessions by Rick Houlihan, we have a great D-book about DynamoDB. And it's great. But as I mentioned before, I think some things that are repackaged are also kind of beneficial. Maybe just this diagram will work for someone, maybe this sentence said differently will work for someone. So we are constantly writing guides, we are making educational resources to make sure that more people understand, the more people use, and, yeah, so that's when it comes to resources.

When it comes to the tooling, we also think that educational gap can be partially closed by tooling. Imagine you use DynamoDB for the first time, I think there is a huge, you're going to approach a huge cliff because everything is super different. So what we are aiming to do with Dynobase is abstract away complexities and different things about DynamoDB and just let them start working with DynamoDB, and then figure out all the details later. For instance, you now are able to query the data and then you can only learn what is GSI. You can query the data and you can also generate the code that is ready to be pasted into your application code. Once you have a code that is following the best practices, that is working, you can just go back and see how it's working.

Actually, that's the way I got into programming. I started modifying some source files from games, maybe some configuration values. And that's how I learned programming. So I think that when Dynobase is generating code to query or to scan, it's also helping because it lets you use the database without actually knowing what's happening inside. It's definitely important to know how the database works, but at the very beginning, we can make the process a little bit easier for the people. One thing that we also identified is that most of the back end developers are already familiar with SQL or "S-Q-L." And you cannot use that language with DynamoDB. So what we have on our roadmap and which is the biggest challenge for us is to enable querying DynamoDB with SQL. And that way, we can just make people use DynamoDB easier. And yeah, hopefully it will drive bigger adoption.

Jeremy: Yeah, well, that'd be really interesting, too, is just if you could take some sort of T-SQL parser and give it auto complete, and then be able to start typing things in and then maybe even make the suggestions of where if you wanted to query the data this way here might be your optimal GSIs, or here might be optimal way to store the data and so forth. I think that's one of the tools that I would love to have, where basically just like, copy my ERD into some system and then have it do some thinking for me and come back and say, "Okay, so here are the entities that we want to create from this, here's the relationship between the entities and so forth."

But the other thing, and this is, again, maybe to the education side of it, where I think you've got a lot of developers who think that working with DynamoDB and, maybe I should take a step back, because I think I want to ask this question a little bit differently, the challenges of working in the cloud and working with serverless, this is something I think that is very, very new to a lot of developers. I've interviewed a lot of developers in my day, especially a lot of young kids coming out of school and no offense, calling them kids, but they're kids, and they come out of school and they know nothing about the cloud, or they know very, very little about it, like "Oh, I used Firebase one time."

But they've never developed anything on the cloud and serverless is this foreign concept to them because, again, their professors are still teaching them how to develop on servers. You know what I mean, and that level of thinking and that sort of on prem type of thinking. So what are some of those challenges that you see, maybe from people just going to the cloud and using serverless technologies?

Rafal: Oh. So there's... There are definitely many things. Because if you are using cloud and if you're using serverless, you definitely need to understand the IAM. And that's really hard thing to get all the policies, roles, users, groups, maybe some even SSO, and there is a lot of things that you need to wrap your head around. And there are many other primitives inside the cloud that are working with serverless. You cannot just think that I will have a Lambda function because it's serverless. Many people think that actually serverless is just Lambda. No, there is a series of services that are working together and there is also a whole foundation on top of that. There's this distributed way of thinking, I think, where you don't have this local, when you don't have that drive, where you don't have networks. It requires so much more.

Jeremy: Right. Yeah. No, I totally agree. I mean, that's, I think there's just that step from this idea of writing everything monolithically to being able to separate all these whole pieces, and having those things work together. So, let me change the subject a little bit, because I think having your experience and knowledge of DynamoDB will be helpful here and get some insight into a question that seems to come up all the time where people say, "Well, serverless is great for spiky workloads or workloads that barely run, or run every once in a while." And so that argument I don't necessarily agree with because I think that serverless works great because sometimes you have spiky workloads, sometimes things just roll straight and you have predictable traffic. And I think that's a smart way to build an application so you don't have to think about the underlying infrastructure.

That being said, there's a similar argument that is against DynamoDB, where they will say, "Well, DynamoDB is overkill for a small little project. So if I'm just building a little side app or small project, internal or something like that, I'll just spin up a MySQL database. And I'll just write it that way. Because it's not going to get a lot of traffic." I disagree with that because I like the fact that with DynamoDB, I don't have to think about database backups. I don't have to think about scalability, if for some reason it does, maybe I need to do a bulk load of data into it, and maybe the database isn't powerful enough for whatever those reasons are, but what are your thoughts on that scale argument? I mean, do you think that people even if they're building something small should default to DynamoDB or should they be using something else until they get to a point where maybe they need that scale and that NoSQL back end to handle massive amounts of traffic?

Rafal: I totally agree with you. I think that they should already go to DynamoDB because I feel like this is the database of the future. And if you're, for instance, starting a small business, small startup, I think there are two things that you definitely will not like to care about. And I think it's bureaucracy and maintenance and DynamoDB is distinct, which lets you set your database and forget it, you're done. You don't have to care about maintenance, patching, security, tuning the performance, checking everything works, making it highly available and stuff like that.

You get that out of the box. And it's going to be future proof because AWS engineers will take care of that and your database probably will upgrade itself many times throughout the project and you don't have to do anything about it, you just have a reliable data store. Yeah, and this way, when you don't have to care about all those things, you can focus on those things that are making you differentiate on the market. You can innovate, you can build, you can focus on application logic. And I think that's the core of innovation, that we don't have to do the grunt work, we can be just creative. And serverless and the cloud makes this creativity easier.

Jeremy: Right. Yeah. And actually, I think that, for me, the biggest sort of pro to using something like DynamoDB is, if you are not using it, it costs zero dollars. You know what I mean? And so if you set up, even if you spin up an aurora serverless database cluster and I think you can do one ACU now, but it still costs me $30 a month or something like that to keep that constantly running. And granted, you can shut it down and have it sleep and some of those things and certainly save yourself money that way.

But it's the same argument I think with spinning up an EC2 server in order to write a Node.js app or Python flask app or something like that, where it just seems like if you're getting barely any traffic, or you're experimenting, you're trying all these other things, that that cost argument is huge. I mean, I can build as many DynamoDB tables as I want to and it's most likely in that free storage, which is awesome, right? So I'm not paying anything, even to store data, I think get 20GB of storage for free or something like that, which is insane.

So, yeah, I really like that idea of just being able to do these things very quickly and very easily, and very cheaply because in the past, I would spend, I mean, I remember this in the days before, I mean, I would be spinning up multiple EC2 instances. You'd have a SQL database running or MySQL database running and as soon as you put that into production, you couldn't be running just one, right? You had to have some sort of replication there. Then you're always worrying about that, you're thinking about failover, and all that kind of stuff. All of that stuff goes away when you start using serverless applications.

Rafal: Yep, that's true.

Jeremy: All right, great. So let's move on to another topic that I think would kind of ties into this. And that's this idea of again, changing your thinking from relational databases to DynamoDB. So what are the mental shifts that developers have to do in order to go from writing T-SQL and just saying select star from whatever to dealing with NoSQL queries and really the limitations that are added to the types of queries they can run?

Rafal: Yeah, sure. So I think that there are two challenges actually, the one is that we are learned to always normalize the data. We have the second normal form, the third normal form, we aim to de-duplicate the attributes the data to store them in separate tables in MySQL or Postgres, and you have to unlearn that. DynamoDB works totally different if you want to use it efficiently in one table. And that's just simply hard. If people were using relational databases for past 10 years and someone says that's totally different, you shouldn't do that here. It requires a lot of effort, but the second thing is that it requires you to be involved in the creation of the application a little bit earlier because if you're going to use DynamoDB, you need to know the access patterns because the access patterns are actually shaping your data models.

And if you'd like to take care of the designing data model responsibly, you need to be involved in a business process. You need to understand the client because if you understand the client, you can build your access patterns accordingly. Maybe you can interact with them, maybe you can suggest some kind of change. Because once you've committed to the data model, or if there is a requirements passed to you from the top, maybe you will realize that some time after going that route you cannot change something, you cannot alter some decisions. So I think being a cloud engineer, as opposed to software engineer, it requires you to have this broader knowledge and to have take broader responsibility. And actually, cloud allows us to have more responsibility in the business process because we no longer have so much responsibility in maintaining those underlying services and tools. So yeah, I think it's good and it's fun to be involved in business and in shaping those access patterns.

Jeremy: Yeah, because I think I totally agree with you where when we think about building data for a SQL database, its usually just okay, well, what fields do we need for this particular entity and then we can always join them afterwards. So deciding on those access patterns is important. But the other thing, I think where I guess the shift needs to be made is, NoSQL might not be right for you or NoSQL might not be right for you, depending on what your application is, right? And I know Rick Houlihan talks a lot about this idea that if you're building something where your queries are changing all the time and then you move to NoSQL and you say, "Okay, well I want to be able to select star from this or want to be able to join this..."

Which you obviously can't do joins, not the traditional way anyways in SQL, but that developers will become disappointed if they put data into DynamoDB and then realize they don't have that query flexibility. So what are your thoughts on, how do we tell developers that? Because it's really hard, I think, for some of them to grasp. It's like, "Well, if it's a database, I should be able to query it." So what's that advice that we give those developers about when they should choose NoSQL?

Rafal: So I think that there is a general answer to that question because I see developers so many times rushing into implementation without properly researching the topic and evolved properly knowing the requirements and the limitations of technology. And that also applies to this specific problem. You need to know what are the limitations, but from the technology perspective. And you need to know what are the requirements from the business perspective. If you immediately rushed it in implementation, you can realize that, "Hey, I've made some bad decisions and it's not going to end well." You're probably going to hack some things and it's going to end badly. I've seen that and I've been put in projects like that in before. So yeah, lesson learned. Take your time, and spend more time on research.

Jeremy: Right. Yeah. Because I think that's the other thing. It just hits people in the face if they implement something in NoSQL and then they're like, "Well, why can't I do X or why can't I do Y?" So what about ERDs, right? Building your entity relationship diagrams and things like that. That's still something we want people to do before they jump into a NoSQL design?

Rafal: So I think that it's not going anywhere. We still need those for productive discussions, for working on application layer, for proficient communication, but just this concept of translating that to actually how it's going to be stored in DynamoDB. That part is only different. And I think actually we need some kind of better, I don't know, spreadsheets, abstractions, tools to visualize how those things are evolving from ERDs to different shapes and forms, how the data is structured and stored and then translated back to a business domain.

Jeremy: Yeah, totally agree.

Rafal: So that's another tool that we can solve.

Jeremy: Yeah, that'd be great, right? That's what I said. I mean, I would that. I would love that ERD input tool that just spits out, "Hey, here's how you want to structure your DynamoDB table, your NoSQL table." So, all right. So another thing that I think it comes up a lot, especially with single table design, is this idea of, well, how many entities do you put into a single table? So if you're building some really large application, are we putting you know hundreds of different entity types in the same table? And I always say, "No, we want to use a separate table for each microservice." So what are your thoughts on that?

Rafal: I think it all depends on the project and all the things that are specific to do to your use case, to your requirements. You can definitely interact, I think, many microservices can interact with one single table. Because thanks to a really granular IAM policies, you can, for instance, restrict the access from one Lambda function to only specific DynamoDB records inside a table using, I think, leading keys and attribute types or something like that. You can tell that this Lambda function has an access to this grand single table with all the entities, but it can only interact with the entities of type, which begins with ID for instance, I don't know payment or invoice or something like that. So it's definitely doable. Also, I think there is also a sentence in AWS documentation saying that the most, the best designed applications require only single table. So it kind of contradicts, but it also contradicts with the what's Amplify is doing. I think AWS is not having one singular statement on that. And it changes case by case.

Jeremy: Yeah. Well, I mean, it also depends on how you define application, right? So I mean, if you have a service that has a payment service, or you have an application as a payment service and a user service, things like that, each one of those services could be considered separate applications and you'd be storing the data differently that way. I mean, certainly what you don't want to do, at least, I guess, more best practices from a microservices perspective is you don't want to be storing data across bounded contexts in the same table or in the same database. You want to keep those separate so that one service can't update data in another service without using a formal contract through an API or some other method to do that.

All right, so what about some of the patterns though, that you can build off of that? So I mean, we know we've got DynamoDB streams. So if you are building separate tables for individual microservices or individual applications, obviously you need to be able to potentially share some data back and forth. But what are some of the patterns that that sort of allows you to implement?

Rafal: So we can definitely use event sourcing and common query response aggregation, because thanks for instance, to DynamoDB streams, you can react on the changes that are pushed to the DynamoDB tables. And actually, I think that DynamoDB streams are also solving some of the problems of DynamoDB. For instance, there is always this analytics requirement of all the projects that you sometimes need to aggregate some value. In DynamoDB if you want to query if you want to, for instance, sum the value of all the items inside a table, it's not going to end well because you need to run a scan through all the records and probably merge it, reduce it, run some really complicated process. Thanks to DynamoDB streams, you can aggregate the value just in time and always have that attribute that updated value. And you don't have to run the query on demand, you can always have the results on some kind of aggregation, whenever you want that. It requires a little bit work and it requires a little bit of education. And there is also a change in thinking required, but it's definitely doable.

Jeremy: Yeah, no, and I and one of the great patterns that I really like, too, is this idea of just using DynamoDB streams to take the data and put it into an Aurora serverless database as well. Because you can just use a small instance if there's too much pressure on the database, then obviously that can back off because DynamoDB streams will just build up. I mean, I wouldn't use it for translating like clickstream data into a MySQL database, but certainly for applications that are just create, read, update, delete type stuff, it's very cool way to have that extra data there for you with multiple things you can do with it. And like you said, I mean, you can push that off into EventBridge or do some sort of event sourcing with it. So very, very cool stuff. So another thing about education and you've mentioned education many, many times. And one of the things that you're doing on top of Dynobase is you have a DynamoDB newsletter. Can you tell us about that?

Rafal: Yeah, sure. So each week, we are gathering some interesting articles and videos, and probably will also sharing some live sessions from AWS. And that's also kind of part of our mission to share the good content. So we decided to start a DynamoDB newsletter something like 20 weeks ago, I guess somewhere around then. And then yeah, you can sign up and we'll deliver to your inbox the best resources we can find, so you don't have to spend all day on Twitter like I do. And yeah, feel free to join.

Jeremy: Awesome. Well, I am a subscriber. I love the newsletter because again, I like reading great content about DynamoDB and unless you are just trolling Twitter all day it is very hard to find that. So that that aggregation of that data is very, very helpful and sometimes I take some of those articles and I put them in my newsletter, so thank you for sourcing those for me as well. Anyways Rafal, thank you so much for taking the time to talk to me today and obviously for Dynobase. So if people want to find out more about Dynobase or more about you, how do they do that?

Rafal: Just go to dynobase.dev, that's our homepage of our product. If you want to approach me, I think the best way is just to find me on Twitter. It's @rafalwilinski and I'm pretty sure it's going to be included in the description of this podcast because it's different, it's difficult to spell for non-Polish people. And yeah, just go to dynobase.dev base or Twitter. And that's it.

Jeremy: All right, awesome. I will get all that in the show notes so they will be able to spell your name. Thanks again, Rafal.

Rafal: Thank you.

View Details

About Matthieu Napoli:

Matthieu Napoli is a software consultant and founder of null. Matthew has been developing web applications for more than 10 years with PHP and JavaScrapt as a full-stack developer, lead dev, and CTO, while maintaining several open source projects. He’s also the author of PHP-DI, Silly, Couscous, and Bref.

  • Twitter: twitter.com/matthieunapoli
  • Personal website: mnapoli.fr
  • null: null.tc
  • Bref: bref.sh
  • Serverless PHP Newsletter: serverless-php.news

Watch this episode on YouTube: https://youtu.be/H8tkZcjQxOA

Transcript
Jeremy: Hi everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm chatting with Matthieu Napoli. Hey Matthieu. Thanks for joining me.

Matthieu: Hi, thanks for having me.

Jeremy: You are a serverless consultant and the founder of Null. Why don't you tell the listeners a little bit about your background and what Null does?

Matthieu: Yes. I created Null two years ago. My goal was to be able to both work in open source as well as work for clients. I use that company to do trainings around Bref, around serverless and also to provide consulting services.

Jeremy: What about your background?

Matthieu: I started as a developer about 10 years ago, and I've been working mostly as a developer. I've been configuring servers, setting up servers for a while. That's why I'm also really interested in serverless. I've been looking at that very closely lately.

Jeremy: Great. All right. I want to talk to you today about serverless and PHP and Bref. Serverless is obviously the topic of this podcast. We've seen quite a bit of movement in the serverless space over the course of the last five years or so. PHP sometimes gets some slack on the internet, but can you give me a brief background as to why you chose PHP?

Matthieu: Yes. That's a very good question and that's a good way to start, because indeed a lot of people have opinions about PHP, and sometimes for good reason. I started with PHP just because it was simple. That's what I love about this language. At the time, like when PHP arrived, it was about 25 years ago. The web was about creating CGI applications, using C or whatever, and PHP arrived and simplified everything. It made the web accessible to a lot of people. I find that really amazing.

That's why I started with PHP as well. I wanted to build a simple website. Yeah. I started with PHP because of that, but I'm seeing the same thing today with serverless. It's making infrastructure, it's making hosting applications accessible again to developers and I find that amazing. Yeah. I started with PHP. I kind of got stuck with this language throughout my jobs and lately PHP has become a very interesting language. To be honest, it's really interesting.

If you've used PHP in the past, I really encourage you to give it another look. It's really worth it. While I do talk a lot and use a lot of PHP, I enjoy using JavaScript as well. TypeScript lately. Really, really interesting language. Life is full of things to learn about, I guess.

Jeremy: Right. Absolutely. I agree with you on PHP. I started with PERL and CGI way, way back when, and then I think I started using PHP 3.0 or something like that with MySQL databases. You're right. It completely changed things. From that we got WordPress for better or for worse, but I think that like 80% of the web runs on PHP.

Matthieu: Yeah. That's a huge market, which is interesting when we're going to talk about AWS Lambda later. Yeah. PHP is huge and I don't think this is something that we can ignore.

Jeremy: Right. Right. Okay. Speaking of this PHP, you realized that there was a gap in the serverless ecosystem for PHP, and so you wrote something called Bref. Can you tell the listeners what that's all about?

Matthieu: Exactly. Yes. I am a developer. I like writing code, creating applications. I don't like setting up servers and all of that stuff. This is why I created Bref. I wanted a simple way to put my PHP code online. At the time I was looking into serverless, looking into AWS Lambda and I discovered, of course, that AWS Lambda does not support PHP. I created Bref to bridge the gap, run PHP on Lambda and provide a lot of tools, documentation, examples. Yeah. Anything that you may lack to create those serverless applications. I would say that Bref is more than just a runtime. It's a whole stack.

Jeremy: Right. There's actually two parts of Bref, right? Why don't you explain those two different parts?

Matthieu: Yeah. I realized over time that there are two major use cases when you look at Bref and what you can do with PHP on Lambda. In the first case, you know about AWS Lambda. You know how it works. You know why you use it. The only thing that's missing is that you want to run PHP for some reason. Maybe you want to use PHP and you want to run it on Lambda. The first part of Bref is a runtime that works just like any other language on Lambda.

Yeah. You can write functions in PHP, handle queue messages, SQS queue messages, EventBridge messages react to S3 events, API gateway events as well. You know, the usual. The second use case is different. Instead of adapting PHP to run on Lambda, there are people that know PHP and do not really know about Lambda and what they can do with it. I take it the other way around and I adapt Lambda to PHP. The approach is that users don't have to change anything in their code.

They can take their Laravel application, Symfony application or whatever, and hopefully put it in Lambda and it just works. That's a second runtime. This runtime, I mean, we can go into the details. It's really interesting because the way PHP runs is very similar to how AWS Lambda runs. Making the old PHP way run on Lambda was fairly ... I mean, I don't want to say easy, but it was doable. That's the second approach where, well, people can just start using Lambda as a web host. That's how I host Lambda as a web host instead of functions.

Jeremy: Sure. Right. The custom runtime for just the first part of it, so just being able to run PHP on Lambda, this is something that's really interesting because I know there are others that maintain PHP runtimes out there. You are optimizing it for actual PHP developers, right? You're using PHP-FPM, right?

Matthieu: Exactly. Yes. The FPM runtime, so that the FPM runtime is used for the use case where you want to use AWS Lambda as a web hosting platform. The FPM runtimes actually runs PHP-FPM, which is like PHP web server inside Lambda. Bref has a little bridge that when there is an API gateway event, will take the event, convert it into a request that PHP-FPM understands. This is the first CGI protocol and so Bref does the bridge, provides the first CGI request to PHP. Then PHP runs just as usual.

You know, the PHP execution model is you have a request, a PHP process starts, builds the whole framework, runs the request, processes the request, and returns the response and then `. I mean, this is perfect for AWS Lambda. That's why it's quite easy to integrate FPM with Lambda.

Jeremy: Right. Then you have the ability to actually create individual handlers as functions and type classes, right?

Matthieu: Exactly. Yes. This is the second part where, so the first part to me is helpful to get people started with Lambda. They start with PHP-FPM. They understand that oh Lambda is great for running a petition. Cheap. It can scale really well. Once they started with Lambda, they understand the execution model and everything. Then they can look into using those real functions like using Lambda just the way it was designed to.

With this second runtime, they write function either using standard PHP functions or using classes. Those classes, those functions are inspired by the JavaScript runtime, as well as the Java runtime for Lambda. They can write classes to process SQS events, EventBridge events, S3 or DynamoDB events and so on.

Jeremy: Yeah. That's actually really cool because I think that some of the other libraries out there just essentially have function support, right? If you're building much more complex systems that are using classes and using that kind of functionality, then the runtime that Bref provides I think is much more flexible and more interesting.

Matthieu: Yeah.

Jeremy: All right. You've got this runtime now, and you've got this ability to port PHP applications into Lambda now, but what are some of the other benefits of Bref?

Matthieu: Well, the main one I see is obviously those are the benefits of serverless. You have an application. You can drop your server that you used, where you used to run PHP, put your application in Lambda and just be done with it. You can scale. You can pay exactly for what you use. Since you already made the step of running on Lambda of configuring the little details like where do I send logs to CloudWatch? How do I store files to Amazon S3?

Once you've done all that effort, it becomes easier to write those little functions. Like I have a link, why not write it as a simple function? I want to use cues, why not send that to SQS with the adapters? Those are all those little integrations that get you started really easily. Along that, there are different libraries and tooling, like Bref provides a simple logger specifically made for AWS Lambda.

It also provides a little dashboard specifically made so that you can view the logs and a few metrics. Some tooling to run colons on Lambda. This is really common for PHP developers to be able to run Chrome tasks or MySQL migrations on their server. They need to learn to do all of that stuff on Lambda as well. Bref provides a specific tool for that.

Jeremy: Right. Yeah. That's another thing too, right? You know, because I got these confused. You think about the Bref runtime versus the Bref library itself. The Bref library itself is like this opinionated HTML framework. It gives you a lot of those capabilities, like you said, like the logger and some of that other stuff that's built in. What about local development? How do you do that?

Matthieu: Yes. That's also something that was asked very early on and it took us some time to answer the problem, but we ended up building Docker images. Those images are the same that we used to build the runtimes. This is the matches we used to compile PHP, to set the correct extensions and the correct settings and everything. We use those images to create the runtimes, and then we also provide them so that developers can run them on their machine.

That's really helpful to develop locally. It's with either the web hosting runtime and the Bref for functions runtime. For both of those runtimes, we have two different images. Yeah. For both of them, you can run this locally.

Jeremy: Nice. All right. What about publishing the application? Is there a workflow for publishing?

Matthieu: Yeah. Bref does not provide a tool for deploying. Instead, it uses the serverless framework. Early on, I started working on a YAML-based tool that would read the YAML configuration that you define and create the resources that you would need. Then I realized it exists already. You know, there's a serverless framework. There is SAM. There's CloudFormation. There's so much stuff. Yeah. Throw everything away. Bref now uses serverless and it's working really great.

Throughout the documentation we also explain how to configure serverless for PHP use cases, PHP related use cases. Websites, APIs, queue workers, those are the three main use cases. Bref also provides a serverless plugin to easily use the Bref layers, the Bref runtimes.

Jeremy: Great. All right. Let's talk about the primary use cases for a second, because I think this is really helpful for people who are thinking about building serverless applications. Like what can you do with it? Obviously PHP is built for the web. I know some people use it for ETL tasks and things like that, if you're really familiar with it. Let's say the primary use case here is serving up a website. How do you build that with Bref?

Matthieu: Right. I would say the starting point will be to use the runtime that is made with PHP-FPM, so the one that will let you use PHP just as usual. That's where you can use your favorite framework. That's great. You can work locally, create your website and then deploy with serverless. I mean, serverless.yml. You would configure your framework to send the logs to CloudWatch. That's very easy thanks to the standard outputs.

You could also configure your framework to use Amazon S3 for storage. Then you can deploy that on Lambda with API gateway as the HTTP endpoint. If you are building websites, however, and that's what we document in Bref, you can either use API gateway by itself or use CloudFront in front of API gateway. That way CloudFront can serve assets with Amazon S3 and serve the usual PHP pages through API getaway.

It's very familiar to PHP developers in my opinion, because it's like using Apache or NGINX with PHP-FPM and Apache or NGINX serving of the assets. There's not a lot of difference here. We use CloudFront and Lambda and S3, but the setup is roughly the same.

Jeremy: Right. Okay. What about if you're building APIs and like ... I get it. Let's say you're using Laravel or you're using Symfony and you get your routes in or whatever, is the only way to do this to build just the Lambda proxy integration to accept everything, or could you build separate functions for different end points?

Matthieu: Yes. I would say if you use a framework that has a router inside, so Laravel, Symfony both have routing inside of the framework. I don't think it would make sense to have different functions. It could be the monolith Lambda pattern where you have a single function, which is huge and handles all requests. You could have made it two functions. For example, if you have a front end and a backend where the front end is public and the back office is only accessible to administrators.

That could make sense to have two functions. That way you can scale functions differently. You can protect those functions differently as well. Yeah. To me, this is the main use case, the monolithic approach. Yeah.

Jeremy: Yeah. If you did want to break them up though, you just wouldn't use the framework, right?

Matthieu: Exactly. I wouldn't because then the routing will be done twice. Once you start using the API gateway routing feature, to me that's where it makes sense to write actual functions, whether with PHP classes or functions. It doesn't really matter, but write actual Lambda functions.

Jeremy: Right.

Matthieu: Bref provides ... Yeah. I didn't mention that Bref provides a small integration here where API gateway events can be automatically converted into standard PHP requests and back and the same for responses. Yeah. Just to clarify on that. PHP has a standard for requests and responses, which is called PSR-7. Bref can automatically map an event from API gateway to those objects. That's pretty good because it lets you write PHP controllers just like in any framework, but with other framework.

Jeremy: Right. Which is pretty cool.

Matthieu: It is.

Jeremy: All right. Then what about like the workers and that sort of use case where, so maybe I have ... So I know like Laravel has a queuing system and there's some other things built into those. If I just wanted to build a worker function, do I do that as part of the framework, or is that something that I separate out into its own thing?

Matthieu: Yeah. That's a very interesting question. The answer to me is not really easy. It depends. The Laravel queue system and Symfony has one as well, which is called Symfony Messenger. Those are particularly pretty nicely built. They, for example, handle automatic centralization and decentralization of your classes into strings that can be sent to SQS. They handle pre-TRI data cues, all of that. If you start using AWS Lambda in the SQS integration, some of these features become useless. Like the retry in the later queue mechanisms.

You can use the one from SQS. Now, do you need the automatic centralization and decentralization of objects? Sometimes. Sometimes not. Now, depending on the case, whether you want to get your hands dirty or exchange messages across languages and across applications, you may not want to use a framework. Writing Lambda functions with class in those that's perfectly fine.

If you enjoy the high level service that Laravel may provide, then use Laravel queues. That's fine. That's why we've been working lately on integrating those frameworks and their specific queue system with Lambda and SQS.

Jeremy: Right. All right. Then what about like ETL tasks and batch processing or other scripts? A lot of use cases for Lambda functions, you see people spinning up infrastructure and shutting it down or just kicking off some processing, batch processing. Maybe reading from Kinesis or doing something like pulling data from S3. You know, obviously that's all possible to do with this, but is there an interface into those other services via Bref, or is it just a matter of using the PHP SDK?

Matthieu: Yes. To me, that's about using the PHP SDK. I haven't seen a lot of weird use cases. I think it's also related to the culture of PHP. PHP is a node language and I would say to match your language and ecosystem. Just mentioning Lambda sometimes it feels like a buzzword. People can be reluctant to look into those things. I understand that. It's perfectly fine to go with mature and boring technologies. Yeah. I think going with Kinesis and DynamoDB, that's a use case that is not that popular in PHP.

Yeah. I think that's the thing. It's also really frightening when I say to people that Lambda has a maximum execution time of 15 minutes. To me, it seems huge and you can split large tasks into parallelized smaller tasks, but it's still sometimes a huge step for some teams to refactor their code and change it so that it fits on Lambda.

Jeremy: Right. Yeah. That's interesting because I do think that that PHP mindset is wrapped around relational databases like using MySQL or something like that. I'm just built into that mindset, but that'd be interesting. Do you think that there'll be a culture or a cultural shift that PHP people start embracing DynamoDB more, or do you think eventually Bref will have first class support for DynamoDB?

Matthieu: Yes. I definitely want to do that. I started working on a DynamoDB, I don't want to say ORM, because it doesn't make any sense, but a DynamoDB SDK. In PHP there is obviously the AWS SDK, but for DynamoDB, for example, it's not as good as the JavaScript implementation. It's really, really hard to send objects and get back objects from DynamoDB for example. There are some things lacking specifically in PHP.

I want to cover DynamoDB. I also really, really want to cover EventBridge and its schema registry. I think it's really, really interesting. Yeah. I think there's a lot of stuff to do here that could be really interesting, but that's a lot of work to do.

Jeremy: Yes. I know. Well, I wrote the DynamoDB toolbox, which is a layer on top of the JavaScript version, which tries to make that even simpler, which is ... Yeah. Getting data in and out of DynamoDB is fairly simple, but being able to serialize it the right way and make sure that you do all the right formats and stuff can get a little bit confusing. Let's not even get started on querying and single table design and all that kind of stuff.

Matthieu: Yes. Yeah. Oh-

Jeremy: Go ahead.

Matthieu: Yeah. Just thought about a use case I heard about like six months ago. I thought it was really interesting and really like PHP. It was a team that had a really old Legacy PHP application and that's an approach I found really interesting. The application was so old and so Legacy that they couldn't just edit the code, add new features and they were really desperate about it. What they did instead was migrate the MySQL database to use Aurora and then use the Aurora trigger to run PHP code whenever there were modifications on some specific rules.

That's how they managed to breach a very old Legacy application with Lambda and with their new stack where they could forward information from the old database into the new system. I thought it was a really clever way to do it.

Jeremy: Yeah. That's interesting. Speaking about Legacy applications, so Symfony and Laravel, we talked about this earlier. There are a lot of those applications out there, right? Laracon and Taylor Otwell, I had him on the show a long time ago and he had just released Laravel Vapor, which is the serverless version of Laravel that allows you to deploy it there.

That was something though that was proprietary and you had to host it in his environment, or I guess you would deploy to your environment, but there's a deployment engine in there. Bref though, you have support for Symfony and Laravel. Can I just take an existing Symfony app for example, and drop it into serverless using Bref?

Matthieu: Yes and no. I mean, it should be that easy. It's not that easy. You would have a few very easy things to configure. For example, logging and sessions and the cache. I think that's about it. Those are three things. That's a single line of code to change most of the time. It's really easy to do. Then there's usually a few more things that actually you write to the file system. That's the main thing to the file system. You have to change that.

If it's information you want to keep, you have to use Amazon S3. Thanks to abstractions in Symfony and Laravel, it's usually fairly easy to change. Yeah. These things can take time. I would say, depending on the complexity of the application, you could spend half a day, maybe a day or two on that migration. If you have a very old ... I mean, a Legacy application that's much harder. That's where I think it would be maybe too much work.

Jeremy: Right. Let me ask you this question, because I always find that as you're building Greenfield applications, that you've got a lot more choices. Old and boring, I don't know if I would consider Laravel and Symfony too old. I mean, they are kind of old at this point, but boring. I mean, for people who are in that ecosystem they certainly love them. Is that something where if you were building a new serverless application and you were going to use PHP because that's the language you're familiar with, would you still suggest that people build using one of these frameworks? Or do you suggest that they start thinking single-purpose functions?

Matthieu: I have seen a lot of web agencies building two or three websites every month. They use Laravel. They know their tools and they know everything that they have to know in Laravel. They are really productive. For them using AWS Lambda is mostly a question of not having to deal with infrastructure. For them, it makes total sense to keep using Laravel and start using Lambda. Then Lambda will become boring. Then they will have to write a small Chrome task or a small worker, and they'll get started with actually writing proper functions.

To me, that's a very valid way to do that because if Laravel works on Lambda and if it's cheap and if it scales well, and if it just does the job, then why not? Now, I don't know if you were a startup and you want to invest into the future or you're a large company and you want to write microservices then yeah. It would probably make more sense to get started into a proper serverless architecture. I think each of those has its use cases.

Jeremy: Right. Yeah. Now, what do you do? You do a lot of consulting work. When you're writing Bref applications, do you do proper functions and just use or you don't use those extra libraries and frameworks?

Matthieu: Yeah. I would say it depends. If I am creating the application and I know that the team that will maintain that in the long run is able to pick that up and use it, then yes, I'll do as serverless as I can. In some cases that's just like I write the prototype or I do a migration and I will hand that off to a team that doesn't know a lot about Lambda. I really adjust based on the people that will work on that project later. For my own project, I use serverless everywhere.

Even I'm starting to write some experimental runtimes that go even further than what I'm doing at the moment. I'm exploring as much as I can. I think there's still a lot to do. I follow a lot what arc.codes and people at Begin are doing. Yeah. I forget the name of-

Jeremy: Brian LeRoux.

Matthieu: Exactly. Yes. I loved the approach of actually changing infrastructure and changing completely the way we create and organize our applications. That makes sense. I think it's a step that is really high. I'm sure we will get there, but I think it's also fine to take time into adapt to the actual needs for right now.

Jeremy: Right. Yeah. I think that you bring up a really good point, and that's just that this idea that the complexity of serverless ... Like a few years ago it was easy. It was simple, straightforward. Then as more use cases started popping up and more people were like, "Oh, I need to be able to do this or I need to be able to do that." It becomes more complex. This is an ongoing conversation that I have with a lot of my guests. What are your thoughts on the complexity of serverless in general?

Because if you're just uploading code or you go to the console, type in something for node, it's easy enough. Now, you're talking about deploying with a framework using a custom runtime. Maybe potentially building your own framework where you have to use your own framework or use a framework like Bref to do that. I mean, I think there's a value to just having these simple onboarding experiences. Like I already know Laravel or I already know Symfony. What are your thoughts on where this is going in complexity?

Matthieu: Yes. Actually, I wanted to say that when I think about Laravel Vapor, for example, I think that's a really good approach because with Laravel Vapor, Taylor has complete control over the framework and over what the framework will do. He designed the framework with the new versions to run specifically ... I mean, to be completely compatible with Vapor. Vapor is super easy to use. I think this is a really good approach. It looks like a very good product. Same with Begin.

I think those approaches make a lot of sense. Just like with serverless components that we've seen lately, it makes so much sense. As you said, it's simple yet it's getting more and more complex every month. Maybe we need another radical simplification yet again, but this time around maybe the codes or how we set all the things up.

Jeremy: Right. Yeah. I mean, obviously as the author of a framework as I know when I build tools, open source tools, I do it because it's something that I need. It's something that was missing. Do you think frameworks are the answer? I mean, because again, I've used frameworks all the time. It's not something I shy away from, but the question is, do we need frameworks for serverless or should we get to a point where we don't need them because the cloud provider, whoever, is handling most of that complexity for us?

Matthieu: Exactly. Yes. That's a very, very interesting discussion. The more I use serverless and services from AWS, the more I realize that these services replace parts of our frameworks. That's what I've been trained to use more and more lately with ... You have API gateway doing the routing. You can drop that off your framework. EventBridge with this even schema registry and schema validation and mapping to actual TypeScript objects or whatever, this is actually what Symfony Messenger is about.

This library is being replaced by a service, an infrastructure service. There are so many examples of that. You can do CQRS with again services. I think eventually the framework may actually move into the cloud and the code framework will actually be an infrastructure framework. I'm not sure where eventually we will arrive, but that's really, really interesting.

Jeremy: Yeah. No. I agree because I think it's one of those things where we try to ... So AWS gives us primitives, right? These simple things that we can use. Lambdas is a primitive. DynamoDB is a primitive. SQS is a primitive, but we have to glue all these things together and you have cloud formation, serverless framework as an abstraction on top of that. You've got obviously Terraform and some of these other things. I'm just wondering though.

You know, I'm trying to ... I think you have unique insight here because as you're building a framework, what you're trying to do is you're trying to build a level of abstraction. You're trying to find a way to say, "Okay. Here are all these loose ends and I'm going to tie them together for you so you don't have to worry about that interconnection." I think that you just have a unique perspective on this. Do you think it's possible?

You mentioned about maybe moving the framework into the cloud. Is there something higher level, like a serverless components that either the cloud provider creates or something like serverless components or CDK?

Matthieu: That's exactly what I was-

Jeremy: Are those the answer?

Matthieu: Yes. Exactly. That's exactly what I was going to say. I don't know if it will be the serverless components. I don't know if it will be the CDK or a Begin-like solution. Myself, I've been working on a framework, a PHP framework built on Bref that actually configures a bit like serverless components, but specifically made for PHP and that looks a lot like in the Laravel approach, except there is no code. It's just setting up infrastructure.

Yes. There is a missing abstraction here. It will happen. I don't know how, but just to answer your previous question, I don't think the framework will die. It will just change a lot and it will change in shape.

Jeremy: Yeah. I think that makes a lot of sense. All right. Then just on serverless in general, because we have a few more minutes, I'd like to pick your brain if I could. You've been building a lot of serverless applications. Obviously a lot of use cases you can solve with it. You're focused on PHP, which I don't think limits in any way what you're doing. There's certainly a lot of other use cases that serverless or Lambda isn't ready for yet, or it can't handle yet.

Where do you see the future of serverless going? Do you see this eventually replacing containers completely or do you think that we're going to live in a hybrid world for a very long time?

Matthieu: That's a good question. To be honest, I'm not sure. It would make sense to me that eventually we would be assembling bricks and not doing that container thing or setting up servers, containers, whatever, just assembling stuff. Just like we can now directly connect some AWS services together without even having to write Lambdas together. The glue is just configuration that we write in yellow. For use cases like machine learning I've never used that so I'm not really good to speak on that, to be honest. Yeah. Getting up to a higher level of abstraction is just the way things go the way we go.

Jeremy: Right. Awesome. All right. Well, listen, Matthieu, thank you so much for joining me. This was a great conversation. If people want to find out more about you and more about Bref, how do they do that?

Matthieu: I have a blog, which is mnapoli.fr. They can go there. I have a few case studies, sorry, case studies about PHP websites migrated to Lambda with Bref. I do have a few more of them to write. There's obviously Bref website if they want to get started with PHP on Lambda. I do run a web ... Sorry, a newsletter as well that is related to serverless and specifically to PHP. When there's anything new in serverless I look at it and wonder, "Is this related to PHP? Is this useful to PHP developers?"

If so, I share about it. Finally, I am at the moment working on an interactive course. It's not ready yet. Hopefully, it will be in a month or so. My goal is to show developers, not just PHP developers, but developers that do not understand why serverless is interesting, what use case can be actually solved with serverless. That's something I've been working on and I'm really eager to finally release it.

Jeremy: Awesome. All right. You also had that cost calculator too, which I thought was pretty interesting.

Matthieu: Yes. Yes.

Jeremy: Yeah. All right.

Matthieu: Perfect.

Jeremy: I will put all of that in the show notes so that everybody can see all this stuff. Again, this was awesome. Thanks again Matthieu.

Matthieu: Thank you.

View Details

About Joe Duffy

Joe Duffy is cofounder and CEO of Pulumi. Prior to founding Pulumi, Joe was a longtime leader in Microsoft’s Developer Division, Operating Systems Group, and Microsoft Research. Most recently, he was Director of Engineering and Technical Strategy for developer tools, where part of his responsibilities included managing groups building the C#, C++, Visual Basic, and F# languages. Joe created teams for several successful distributed programming platforms, initiated and executed on efforts to take .NET open source and cross-platform, and was instrumental in Microsoft’s company-wide open-source transformation. Joe founded Pulumi in 2018 with Eric Rudder, the former Chief Technical Strategy Officer at Microsoft.

  • Twitter: twitter.com/funcofjoe
  • Pulumi: www.pulumi.com/
  • Pulumi Twitter: twitter.com/PulumiCorp

Watch this episode on YouTube: https://youtu.be/MYOGfK9PHM8

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Joe Duffy. Hey Joe. Thanks for joining me.

Joe: Hey Jeremy. Thanks for having me.

Jeremy: You are the CEO and founder of Pulumi. Can you give the listeners a little bit about your background and tell us what Pulumi does?

Joe: Yeah, happy to. I founded Pulumi three years ago. Before that, I was an early engineer on the .NET framework at Microsoft. I was actually at Microsoft for a hearty 13 years working in and around developer tools the entire time, managing groups. I actually led the languages team before leaving, helped with the open source transformation at Microsoft, which was really cool to be a part of, and then founded Pulumi. Pulumi is a modern infrastructure as code platform that really brings everything we know and love about application development using great programming languages, great tooling, and actually brings it over to the infrastructure side of the house and really trying to help both infrastructure teams be super productive with great tools but also empower developers to use more of the cloud as part of their application architecture itself.

Jeremy: Awesome. You just mentioned there cloud engineers or the infrastructure team and then serverless developers, so I look at this and I tend to think... especially with smaller organizations... they're almost becoming one and the same. But as you get larger organizations and you start to separate that responsibility... and whether you have separate cloud teams and separate developers or different cells that do that kind of stuff, there is kind of this separation between the infrastructure dev ops people and the serverless developers. Can you explain that difference?

Joe: Yeah. And I agree. I think developers are doing more infrastructure now than they've ever done in the past and I think serverless is really forcing this issue a bit. Is a serverless function infrastructure or is it code? The line's a little blurry. But there's clear things that are still in the infrastructure domain: for example, setting up a virtual private cloud in Amazon, setting up a network, setting up a Kubernetes cluster. Even if you're going to run serverless functions within Kubernetes somebody's got to manage the cluster. Somebody's got to think about security. Somebody's got to think about monitoring. Some of that actually falls on the applications side. A lot of it falls on the infrastructure side.

I think of it as there are deep domain experts in the infrastructure space just like there are deep domain experts in the applications space. I think the magic of what we're seeing with serverless in particular is that the line is getting a little blurry. It's more of a policy decision, I like to say, than a technology decision about who does what. The infrastructure team is going to do the network because that's what they know how to do. They're experts in that. The development team probably doesn't want to become domain experts in how to set up networks. Similarly, the infrastructure team doesn't want to become domain experts in how serverless functions work, so it's better if the developers can be self-serve and really own their own destiny there. I think the tools and workflows really need to support the concept of these two disciplines working closer together going forward and I love that you used the phrase cloud engineering because that's really what I think of. It's the best of developers, the best of infrastructure engineers, really collaborating together.

Jeremy: Right. That's interesting, because I totally agree with that. Where, then, are these challenges, or what are these new challenges that you're seeing both sides face? Because as a developer, like you said, you need to get a little closer to the infrastructure. As a infrastructure person, you need to get a bit closer to the configurations that the developers are setting up. What are the challenges for the developers?

Joe: Infrastructure's hard. For developers specifically, I think no serverless function is an island. A serverless function is only interesting when it's paired with the infrastructure that triggers an event, whether that's a bucket, you want to do something every time a file gets added, or an API gateway where you're actually using serverless functions for infinite scale on the back end. It needs to be connected to infrastructure and historically what that's meant... Actually, one of the reasons we founded Pulumi was I got really excited about serverless and containers and I wanted to create a serverless application and it was great until I had to configure the infrastructure. I wrote 100 lines of JavaScript for a nice little serverless application. I'm like, "All right. I'm ready to go. What do I do next? Oh. For every 10 lines of JavaScript I have to write 100 lines of YAML." That was not pleasant.

That was one of the problems we wanted to solve, was really, "Let's just make it feel like we've got a real programming model, an application model, where we're just building serverless applications and the infrastructure is part of that." Not all of it, again, but a lot of it really should be closer to the applications. Most of the technology today doesn't have that worldview and so there's kind of a fundamental friction and mismatch.

Jeremy: Right. I think that's one of the things that you see more of now, especially with developers. They have to take the infrastructure into consideration now when they write in code. They didn't use to necessarily need to worry about that. Now it's like, "My code I know is going to interact with some other piece of infrastructure," which I think is a challenge there too. So what about the infrastructure side or the dev/ops people? Or I should say the ops people. What are the challenges that they face now that they've got people kind of meddling with their infrastructure?

Joe: It is a challenge to empower developers to manage more of the infrastructure, because there really is this... I think most technologies today and most teams today assume that there's this hard fall between the two sides of the house. As you pointed out, for smaller companies and companies who are born in the cloud, who are starting today, they have the advantage of not creating those silos, but for most teams there is a hard wall between them and for good reasons: because of these domain specializations and expertise. To your point, 10 years ago if I was just doing virtual machines I talk to my infrastructure team once a year when I was doing capacity planning. It's like, "Well, I need to go from three to four VMs and I need an extra database. Can I have that in a quarter or a month?" Things are so fast-paced these days the only way to keep up is to really empower the practitioners, the developers, to control their own destiny, and you need to think about security when you're doing that.

Joe: Right?

We hear all the time about, oops, a bucket was open on the internet and somebody slurped up all the credit cards. Those things happen all the time. So how do infrastructure teams let their developers control their own destiny but still make sure we're compliant, we're secure, costs are under control? All those hard problems are still hard problems.

Jeremy: And you have this overlap, right? So if I'm a serverless developer and I'm writing a Lambda function, maybe I'm not thinking about concurrency limits, but if I'm on the ops side then I'm thinking, "Hold on, we only have a thousand... or maybe we bumped it up to 5,000..." Now you've got one particular function that could go rogue and consume all of my concurrency capacity. That overlap... I know there's challenges there, but how do you manage that overlap between those two competing things when they're working on the same infrastructure probably at the same time?

Joe: It is tough. It's funny, in the early days of Pulumi we actually were playing around. We created a Lambda that created Lambdas and for every Lambda it created five more Lambdas and we couldn't kill the thing fast enough to actually stop the thing from spreading. It quickly added up to like 1,000 dollars in two hours. It's a real challenge.

There are tools out there that allow you to enforce constraints or guardrails, if you will: to say, "You need to stay within budget," or, "You need to stay within compliance: these guardrails." We offer such a tool, called policy as code. Just like there's infrastructure as code there's also policy as code. That's one tool in the tool belt that you can use to enforce these things. I think also there's just smart things you can do with setting up your accounts properly so that developers have their own sandboxes that they can play in and those are different than production. Some of these new CICD capabilities where you can really test things before rolling out to production I think is also a key element of this.

Jeremy: Yeah. And then, like you said, infrastructure as code is sort of that common language between both sides of those. That's always tough managing it, too. I've seen that where somebody checks in something on one side, somebody code reviews it, and then they want to change it. Next thing you know, something's not working right. Certainly managing that is kind of crazy. But I also see not just serverless developers building tools or building applications for serverless, or I guess applications, but you have sometimes the ops engineers that are using it as well to write automation things.

Joe: Absolutely. I think dev ops kind of started this whole trend that really laid the foundation for these two sides of the house really coming together. That led to a lot of this automation that you're mentioning where on the operations side we often do use code to automate things, whether that's Bash scripts... that's kind of a form of code... or Python scripts. So I think infrastructure teams are used to using code to solve some of these challenges. I think what we try to do is... Okay, as you say, the infrastructure as code platform. Let's use general purpose languages so that application developers now have access to infrastructure and can do infrastructure as code. It's not intimidating. You don't have to learn a new DSL. You can use familiar tools and approaches. But then because of dev ops and because infrastructure teams are used to automation, now infrastructure as code using real languages doesn't seem so foreign and now we're speaking the same language. We have a common foundation to start from.

Jeremy: Right. Let's talk about infrastructure as code a bit more. You have a ton of experience with this and whether you're doing a DSL cloud formation or something like that or you're using the CDK or Pulumi or any of these other more scripting, familiar type... Now there's the new one for Kubernetes. What's it's called? Like CDK8s or something like that that just came out. Whatever you're using, how much do developers need to know infrastructure now?

Joe: I think really the more you can learn the more powerful your abilities will be, honestly. If you look at the building block services, Amazon has 200 hosted services. Azure, Google Cloud, they have a large number as well. The way I think of it is you just think of those as building blocks that you can use to build more powerful software. If you want a data store, great, you've got a data warehouse at your fingertips. If you want a hosted AI, if you want to do speech recognition in your application, well that's just a service now. You can just take that building block and use it.

All of those things are infrastructure, right? So I think infrastructure has an intimidating connotation to it where it's like, "Infrastructure is like virtual machines and these networks and everything." Infrastructure these days really has become a lot more of these hosted services that developers can harness to increase the capability of software. That said, again to my previous point, don't feel like you have to go super deep in networking, like public-private subnets, route tables. Some of these things are really not at the level of abstraction that most developers are thinking of and that's okay, but I would say don't think of the word infrastructure as a frightening term. It really shouldn't be daunting. It really is a superpower that you can use in your application.

Jeremy:I think something that has changed certainly... and like you said, we're not talking about setting up VPCs or routing tables or some of that stuff. We might be talking about using an SQS queue or using blob storage or something like that. We're talking about using these components and that's great but we're no longer given the option of just saying, "We have a Kafka service running so we just use that as our service or as our event bus, and we have a database cluster set up and that is the database that we connect to." Now we're talking about setting up separate DynamoDB tables, an SNS topic, SQS queues, EventBridge: all these other things that we're just adding so many different components, and like you said, building blocks, but that's something that I think goes beyond coding now. Now we're architectural design. We're talking about how we architect these applications.

Where do developers fit in there? Because obviously you can't just write one Lambda function... I mean, you maybe could but you shouldn't... one Lambda function that does everything. You want to separate these into different building blocks and use all these different components. How much do developers now need to start thinking about architecture?

Joe:I think if you were an architect... I think senior developers always think about architecture. I think the kind of architecture just changes now. If you go on a whiteboard and you draw the diagram of the architecture and connect all these boxes, what are the boxes? In the past, they might've been monolithic applications with maybe little components within them. I'll date myself if I say what I was about to say, but comm components or J2EE... java beans or whatever. It's no longer those things. Now it's these services and microservices and they connect over RPC or what have you, but many of those building blocks that when you draw that system architecture now are going to be, quote, "infrastructure." That's great because infrastructure means you don't have to build it; you can use something off the shelf. Like you say, the database, the queues, the SNS topic: you don't have to go hand code your own pub sub system. You can just use one off the shelf and that's really powerful.

Jeremy: The other thing about architecture and assembling multiple things is in order for me to connect to DynamoDB or in order for me to connect to EventBridge, there are a lot of permissions: IAM permissions. It's the same in pretty much every cloud you use. There are going to be sets of permissions that need to be configured. That really opens up a lot of the security stuff. We know the cloud is really great at perimeter security. We don't have to worry about somebody getting in and messing around, but a poorly coded application can expose issues. If I'm using third party libraries there's all kinds of security issues. Saving the data: am I encrypting the data? Different compliance things. That's another thing. Where does infrastructure as code, developers, and ops people... where does that all come together to make sure that we've got not only things like reusability but also compliance and security?

Joe: I think that's, frankly, the hardest part of this whole transformation. I think IAM is something everybody has to think about. Security is something everybody has to think about. You can't ignore it. But there's levels of security. I think IAM is extremely fine-grained.

Jeremy: Maybe too fine-grained in some cases.

Joe: Right. It's kind of overkill for some... Like, once you get to the application tier do you really need to think about literally every fine-grain permission for this Lambda or is it okay if your infrastructure team gives you a sandbox and says, "Here's the permissions I'm comfortable giving my developers, and they can do fine-grain permissions within that box, but I'm going to give them something that's a reasonable starting point and assume that if they got everything wrong it's not the end of the world." I think that's the key for infrastructure teams and developers working together, is the infrastructure team needs to figure out, "What are the IAM permissions I'm willing to give to my developers?" And then developers can think about it, and they should think about it, but then it's not as much of a catastrophe if they get something wrong to the fine-grain details. It's almost like imagine if your job application every single object you had to have ACLs on your object. That would be insane. That's kind of where we are with IAM in some ways.

Jeremy: And I like this idea of infrastructure as code and CDKs for the reusability piece of it, because I feel like that's where I want to go, where it's like, "I need to build some system that processes a queue. What are all the things I have to do? What are those permissions?" And so forth. I don't want to have to write that by hand every single time. I want to pull that off the shelf and say, "This connects to the queue, this does all of the correct permissions and all that kind of stuff, and then here's my code that processes the actual data or whatever." It would be great if we could get to that point. I think we're moving there but not quite there yet.

Joe: And that's the direction that we're heading in for sure. We have lots of libraries. Actually, a lot of our customers use Pulumi to do exactly what you're saying, which is maybe a developer, they don't want to think about all of the different pieces of Kubernetes microservice. Maybe there are some security permissions that come with that. Maybe there is an RDS database. Maybe there are some services in an EKS cluster, if they're in Amazon. A developer may just want to come up and say, "Give me a new microservice." They don't want to think about all of these pieces and if you use the code to create these abstractions... real code with infrastructure as code... then the infrastructure team can build these abstractions that have built in best practices, hand it off to developers, and not only know that it's going to be secure and reliable and cost efficient but now the developers don't have to think about every little detail of how that building block was created. I think that's definitely the direction we're heading in.

Jeremy: Awesome. So we talked about developers and cloud engineers or infrastructure people working nicely together, and we know the dev ops thing is pretty strong. I think there are a lot of companies that have a good dev ops culture where they are working together but you see silos all the time. You always see it's like the developers over here and the infrastructure people over here and they're like, "We want to do this," and they're like, "Nope, because there's a security issue or there's some other reason why you can't do that," and I think it gets even crazier in the cloud because it's so easy to just grant people permission to something and let them do something but you've just got I guess maybe market forces is the right word that kind of prevents these teams from working together.

What are some of those... that's happening right now. What are some of the market forces that are keeping these things siloed?

Joe: I think, frankly, the tools and technologies are very different. Developers use a very different... Every day, their day job, they use a completely different set of tools than folks on the infrastructure side, so even if we want to collaborate it's actually kind of hard. And this is something that we're trying to change with Pulumi, where, okay, let's use Python, let's use JavaScript, let's use Go, let's at least speak the same language so we can start making it a policy decision, who does what, rather than one implied by technology.

I think that causes some of the silos. I think it really depends on the organization as well. I think the most disruptive companies that are transforming entire industries using the cloud as a competitive advantage have figured this out and they are figuring this out. I think that is forcing some of the larger, maybe more established, companies, let's say, to start reinventing how they're doing things. I think the literal market forces are pushing people in that direction. It is uncomfortable too. I think we've actually put more dev in the ops than we've put ops in the dev and I see now it's going in the other direction. You look at observability and application performance management and infrastructure as code. These are things that developers now are thinking about every day and even just five years ago they weren't. I think we're heading in the right direction but it's definitely an uncomfortable, difficult transformation for a lot of people.

Jeremy: And dev ops as a culture I think is really an important step. Like you said, there are a lot of companies who have figured this out and it's great because you get that agility, you get that speed, you remove all those bottlenecks, but has it... I guess my question is has dev ops also created more silos in a sense because now you've got really strict processes in place?

Joe: I think it's helped connect the operations team with the developers, unquestionably. The interesting thing is actually if you look on the infrastructure side of things I see silos within the infrastructure side of the house. Like, dev ops is sort of a different silo than sys admin, where dev ops is happy to write some code, happy to write some scripts, and sys admin maybe not so much. Maybe that's more point and click ticketing kind of stuff. I think now you're seeing the emergence of SRE and these more advanced infrastructure teams, which is even like a step beyond the dev ops approach where dev ops was kind of started 10 years. It's come a long way, but SRE is a relatively new practice that's actually more like software engineering than dev ops was. So each of these are slightly different factions and that does definitely cause a little bit of challenge because that just means we're speaking five different languages instead of one.

Jeremy: And I think that's part of the problem, too, where you start seeing these silos within just the organizational team. That's why a lot of companies create these cloud teams that are strictly dedicated to the cloud. But as serverless... One of those things I think maybe companies aren't ready for is giving developers more control and giving them more access. Is just this shift to serverless, like is that enough of a driver for people to be like, "Okay, serverless is going to make it faster, give us a faster time to market, more productivity from our developers, faster development cycles, or whatever"? Is that enough of a driver for these companies to change, do you think?

Joe: I think serverless on its own for some companies would be. What I see is for some companies serverless is really important to their entire strategy, their architecture. It's a naturally event-driven architecture. It's way more cost effective for them to adopt serverless. I think for other organizations it's a combination of things. I actually think containers is another forcing function for a lot of these things where building and publishing a container seems like it's something you can do without touching infrastructure until you start doing private registries and hosted load balance services and ECS or Kubernetes. Now you need to actually consume... So that line starts to get blurry as well.

Joe: To me, it's the combination of serverless and containers combined with just the rapid pace of innovation and the fine-granularity of these services because you kind of pointed out earlier it's not these monolithic things any more. It's just lots of little pieces that you need to stitch together and that means things move a lot faster and at a very fine granularity, which is even more difficult to stay on top of. I think all of those combined together [inaudible 00:24:18] the only way to keep up with the competition, frankly... the competition being the ones that are the most innovative and have already figured this out... is to really empower developers.

Jeremy: I wonder if this on ramp of containers, which I think is great... I'm not a fan of lift and shift. I think that just transferring everything into the cloud isn't going to give you much savings other than not having to manage that infrastructure any more. Well, manage the physical infrastructure at least. But the shift to containers, I like that. I think that it's a really good on ramp. But I don't see containers replacing all of these other cloud services too, right? So even if you're building your application on containers, you're still likely going to want to use SQS and RDS and DynamoDB. You're still going to want to use those things, so even that shift to me seems like it still opens up all these cans of worms with security and everything around that.

Joe: Absolutely. We talk with customers all the time that are at various points along this journey and any time somebody says to me, "I'm going to run a MySQL database in my Kubernetes cluster and manage the persistent volumes and backups and everything on my own and I'm running in AWS," my first question is, "Why aren't you using RDS? Do you really need that level of complexity?" For some people, yes, the answer is absolutely that makes sense. For most people it makes more sense to start with the hosted service. Now, like you say, you're having to manage lots of these moving pieces and stitch them together and it is infrastructure and infrastructure as code is the way to tame that complexity and chaos.

Jeremy: I totally agree. What about developers' responsibility? We talked a bit about it and maybe learning some infrastructure and learning some security and some of those things, but how much of that falls on them now and how much should we... I understand you can put in guardrails for certain things and you can do code reviews and you can have another dev ops team or an ops team that's looking at some of these things. Maybe you have a sec dev ops or dev sec... What is it called? Anyways, you have some other team, some fancy team name, that is looking over their shoulder and trying to do this, but how much of that do you expect those people to catch and how much of that responsibility now falls on the developer?

Joe: I think the unfortunate thing is secure by default would be the ideal world to live in where it's principle of least authority, which is generally regarded as the place to start from because then if you don't need a permission you don't get it. That's not the case today. The example of S3 buckets that I mentioned, the default shouldn't be that an S3 bucket is open to the internet and Amazon is definitely going that direction by adding controls and access blocks and things like that. That's the thing that a developer needs to be careful about, is just know that the defaults aren't always secure. In fact, often they aren't, so if you assume... But in the early days you would write threat models, right? Developers think about security. It's not like we never think about security. It's just kind of a different threat model. It's a different set of concerns, but it's a very transferrable set of concerns. I think it's not entirely foreign but you have to know going in and eyes wide open that there are a lot of foot guns out there and I think when you have these sister teams like the security engineering team and the infrastructure team, dev sec ops or sec dev ops... I always forget the ordering as well.

Jeremy: I don't remember either, exactly.

Joe: Lean on them as well because ideally those teams... kind of what I was saying earlier... they would set up your environment so that you can't shoot yourself in the foot. That's easier said than done but it is possible.

Jeremy: And I think the other thing you have with serverless now is that you can launch a Lambda function that's not in some private VPC, right? It's in the general VPC. Well, it's technically in VPC. It's in Amazon's VPC. But you can launch that function without having all of those other security things in place.

I agree with you. I think that you want to lean on other people as much as possible but I always now... more than I ever did before... when I'm building something I'm thinking about the security and I'm trying to think, "What happens if this happens and what are the worst case scenarios?" I don't know. I agree that it's good to have those people to lean on but I feel like developers maybe need to go a step further if they are building in the cloud and that might just be this new thing called the cloud developer or a cloud engineer as you said earlier. That might just be the new normal and where you need to be as a developer.

Joe: Yeah, and I think the complicated thing is the execution environment of the cloud is very different. Most developers are used to writing code that runs in one monolithic context, like on a server or on a desktop. In the cloud, your code is... especially serverless... spread across lots and lots of different servers or you may not even have the concept of a server if you're doing serverless, but that thing has permissions. That execution context has permissions and you need to think about are those the right permissions and what if somebody were able to get code to run in that context that I didn't expect and is that possible, and then you need to think about the network perimeter: your VPC example. You can access this thing? What are the APIs that are exposed? What are the capabilities of those APIs? Is there authentication? Is there authorization attached to it? How does that work?

I really think you got to think, to your point earlier, like architecture. You have to think architecture. You need to draw it on a whiteboard and think, "What is the security threat model for this overall architecture," and that's the way to go. I agree: you can't exclusively lean on a separate team, especially some companies don't have those teams, so you really need to take matters into your own hands.

Jeremy: Totally agreed. Let's move on to some tips, because I think that you have been working with a lot of companies. What are some of these processes that companies can put into place that help them adopt serverless faster by creating a better relationship between the developers and the cloud engineers or infrastructure people?

Joe: I think you have to figure out the tools and the workflows: the tools, the workflows, and the processes. It comes down to those three things. I'm biased. I think a tool in a workflow that works great for developers and infrastructure teams means that if you don't get it right on day one you can always change your mind down the road. It's not like, "You folks over there are going to use this set of tools and you over here you're going to use that set of..." Once you make that decision and people start building stuff using that it's incredibly hard to reverse that, so that's important to get right on day one.

From a process standpoint, I really do think finding some way where guardrails are in place is really critical because you don't want developers to always have to come to file tickets to get... You just want to empower them to run full speed ahead and know, and sleep soundly at night knowing, that nothing bad is going to happen. I also think by using infrastructure as code you can use familiar coding techniques like code reviews, like change management. You can actually just use code reviews and pipelines and a lot of these CICD platforms these days just raises the visibility for the whole organization in terms of what code is running where, who's pushing what change, because in the event that something does go wrong you're going to need to go and find out when it happened, where did it come from, who do I go talk to. That's also important.

I think it's really the tools, the workflows, and the processes and they really need to all gel and ideally not fundamentally different for the infrastructure team than the developers.

Jeremy: Right. I actually really like... One of the greatest things about infrastructure as code is I remember back in the day I was uploading a Pearl script to a web server somewhere that was running in a CGI bin and I would just upload that file. Oh, I needed a new server. I'd have to go configure that server separately, set up Apache, and do all those configurations. Things got better as we moved towards things like Ops Works or Puppet and Chef and those sort of things because it helped repeat those infrastructure deployments. But now it is so easy to spin up a new environment, especially if you're all in serverless. If you're using DynamoDB and SQS, you can spin these things up and tear them down.

I love that approach, too, where you give developers a lot of flexibility to put something out there in sort of a test or QA environment or something like that or maybe just a dev environment and then have that move through a CICD process, go through a code review if it needs to, and some of those other things. I really like that process because I think that that gives the developers that freedom to play around and actually get stuff up and running in that sandbox environment but then still put in those checks and have all that change management and that process management in there as well.

Just one more question on that, because I think this is something where, like we said, you've got developers thinking about architecture and then you might have ops people thinking about architecture as well. Where's the delineation of responsibilities?

Joe:It's interesting, especially with Kubernetes. One thing we're seeing is there's sort of like the infrastructure operators and the application operators and it's sort of like this natural divide is happening where there's like the base layer of any architecture and often times it's shared. Maybe it's company-wide and shared amongst lots and lots of applications. Sometimes it's maybe more fine-grained than that, but it's sort of the networking, the base security, the cluster, maybe some of the data services, encryption services.

There's this fundamentals layer that definitely the infrastructure team is, if you're in a larger company, going to be the one who manages that. Got to get that right 100%. It moves a lot less frequently. You change it occasionally but it's pretty stable. Once you get it up and running you might need to scale it up. You might need to go to new regions, things like that. Then on top of that it's all the application services and I think the application services, that's the stuff you want the developers to manage. That's serverless capabilities. It's data stores: Aurora, S3, Cosmos if you're on Azure, those level of things. And services, like load-balanced services, those really belong at the top.

Unfortunately, at the very front from a networking standpoint you sometimes have CDNs and some of the load balancers can get complicated and public subnets, so that cuts back over to the infrastructure team. But basically most of the stuff above the line should really go to the developers ideally.

Jeremy:I think that's great advice. I mean, especially it's like you don't want a developer going in there, setting up an RDS cluster with all the security groups and everything that's surrounded there. I mean, they certainly can, but if you've got a larger team and you've got somebody that can handle that that's definitely the way to go.

Another thing that Pulumi does is it deals with multiclouds. I hate the term multicloud because it's one of those things where it depends on what people are trying to do. Are we trying to be cloud agnostic, which probably a really bad idea, or are we just trying to find the best services in cloud A versus cloud B and use them all together? In your experience, because I always love hearing this feedback, how have companies that are using Pulumi and the customers you've talked to embracing or using multicloud?

Joe: It's a pretty broad spectrum. As you say, usually trying to abstract over what makes each of the clouds special and unique is probably not a great idea but there's some areas where that works. Kubernetes is a good example where we're finally kind of agreeing on what it means to run a container in one of the cloud environments, so I think of that as almost the POSIX or Unix API of running container-based compute... but it doesn't go much beyond that.

For multicloud we see a number of things. One, as you say, each cloud has different services. You might want to use S3 in Amazon and then machine learning in Google Cloud, for example, and that's totally fine. Furthermore, it's not always just the major clouds, right? You might be using CloudFlare. You might be using Datadog, New Relic, Mailchimp. There are these infrastructure service providers that are part of the infrastructure and you need to manage the infrastructure on those as well. There's very basic reasons. Like, we work with a customer. They were running in Azure and they get acquired by a company that runs everything in AWS. Did that company want to force them to rewrite everything just because they acquired them? No, they didn't. It wasn't the most cost effective thing, so now they're multicloud. They didn't really plan on it but they are.

The other pattern we see is companies selling a SaaS. If I'm selling a SaaS that runs in my customer's cloud, do I want to say I can only sell to Azure customers or AWS or whatever cloud I happen to pick? Probably not. You probably want to architect it so you can be flexible and sell to customers running in all of these different clouds. That's a common pattern. I don't see much folks talking about that but we see that quite a bit with our customers.

Jeremy: Do you see any vendor lock in concerns? Is that an argument that comes up?

Joe: It does sometimes. For us it's actually part of why Pulumi's interesting. It's not tied to one particular cloud. But that's just more admitting the reality that many people have to multiple clouds and they have to move some day down the road. We see some people wanting to avoid lock in, some people using it for maybe price negotiation at a very high C-level conversation. But most of the time when a CIO makes that decision or something their entire team is grumbling because it just adds so much pain and so much friction because they have to abstract over everything. I think the workflows being cloud agnostic is a good thing. Policy as code, infrastructure as code, that being consistent no matter which cloud you're going to go to is great. But once you get down to the actual building block services they're very different in each of the cloud providers and trying to abstract over them is usually a fool's errand.
Jeremy: It's the lower common denominator thing, right? We want to pick the best service for our application and if you're trying to do something that you can duplicate across multiple clouds or with the fear that some day I might need to move this thing, I think the amount of investment you make in that is more than it would be to re-engineer it to move it to a different cloud later.

So what's next for infrastructure as code or just for serverless? We talked a lot about CDKs and being to repackage things, reuse things like that, but that in and of itself is still kind of a problem in terms of learning curves. You still need to know all the individual services. You still need to know exactly how you're stitching these things together. Is this something where there's going to be more collaboration? Is there going to be a higher level of abstraction? What is that next step?

Joe: I totally agree. We're very early days. Pulumi, what we've done is we've taken those building blocks and we've exposed them in general purpose languages and given you an infrastructure as code platform where you can manage infrastructure reliably and you can bring that closer to your applications, and we've added some abstractions but definitely the average developer really just wants to get up and running very quickly and not have to worry about a lot of the low level details. The cloud APIs really were designed for infrastructure circa five years ago, 10 years ago. They're not really designed for great usability. You think of a developer. It's almost... my horrible aging analogies... like you go to build the Windows application. Are you going to program and see against the Win 32 API or are you going to use No Jazz or Python or something that you're hyper productive in?

We're still very much in the Win 32 C days. I think we'll get there. I think Pulumi... one of the things we're really excited about is it gives this foundation, so we started building these higher level abstractions and it's not like a Heroku or a Paz. I love Heroku. The challenge most customers run into is once they hit a level of complexity they say, "Now I need to abandon the platform because it's too high level." The question is, can we have that high level while still connecting to those lower level building blocks in a way that works at scale in some of the largest organizations? I think that's where we're trying to get to but it's definitely super early in that journey.

Jeremy: Yeah, and I think one of the things for me that I see a lot, especially now that everybody is working from home and you've got a lot of remote workers, is people trying to collaborate on larger blocks of infrastructure or larger applications where maybe you have 10, 15 microservices. Maybe you have 100 microservices and each one of those has 50 different services in it or Lambdas and queues and all these other things that are happening. It gets really hard to manage. Even if you're using CloudFormation or you're using the CDK or you've got a really good code repository and a good workflow for all that, you still have a lot of different things to manage. What do you see maybe being the tool of future? Besides Pulumi, maybe. What's that tool? How are people going to be able to collaborate across the world on all these different things and organize and better than just text files in a repo?

Joe: I think what we're really on the verge of is distributed computing. I think, finally, we're getting to the stage where we're moving from monolithic single computer programs. We went through concurrency with the multicore era and figured out how to do asynchronous programming, so now every language in the world supports asynchronous programming with async/await and tasks and promises and these things. Now we're about to do that with distributed computing. These fundamental concepts of having lots of little pieces that communicate with each other: our programming languages need to better support that and our programming models and it needs to be more first class. Ironically... I don't know if I should say this... I feel like we're almost converging with developers in infrastructure and if we really could figure out some of these distributed computing programming models I think they might diverge a bit because at that point developers really don't want to think about literally every small building block. They want to think about these higher level programming patterns and application models.

There are some folks that are trying to do this already and they're very exciting. I think it's going to be like a 10 year journey to get there. It took most of the 2000s, 2010s, just to figure out async, so these things take a while, but I think that's probably where we'll end up.

Jeremy:Awesome. Joe, I appreciate you being here. I do want to give you a minute just to explain what Pulumi is doing to solve this problem.

Joe: Pulumi... by choosing general purpose languages for infrastructure as code, you can build reusable abstractions. You get everything we know and love about languages, which is a great foundation... testing, four loops functions, basic abstraction... but you can really build these reusable components. I think for serverless, the team's a bunch of ex-compiler nerds. Our CTO is one of the two original guys who founded the Typescript Project, for example, so we figured out some cool ways on how to do serverless computing where Lambdas really are Lambdas in your favorite language. You don't have to do this 10 lines of code and 100 lines of YAML.

It's really exciting. As we've discussed, it's super early days, but I think we've laid a solid foundation. And it's open source. We've got a great community, support every cloud provider you can imagine: over three dozen other infrastructure providers. So very powerful. Great for developers. Also great for infrastructure teams who are trying to build that bridge between the two sides of the house.

Jeremy: Awesome. Joe, again, thank you so much for being here. I appreciate you sharing all this knowledge. I love what Pulumi is doing. I think that, like you said, early days but hopefully things will continue to progress and people will move more towards serverless and companies will figure these things out.

If people want to get a hold of you or learn more about Pulumi, how do they do that?

Joe:Pulumi.com is one stop shopping for everything. That blue getting started button will take you to download the open source and then it's easy to go from there. Great tutorials for different clouds depending on what you want to do next. Then follow us on Twitter: @PulumiCorp. We've got a great community Slack where the whole team hangs out if you want help or talk about things. Then I'm on Twitter: @funcOfJoe. Always happy to chat with people. DM me. They're open. But definitely if you run into any questions, want any help, want to talk about anything, I'm always here to help.

Jeremy: Awesome. Thanks again. I will make sure I get all that in the show notes.

Joe: Awesome. Thanks Jeremy.

This episode is sponsored by: Amazon Web Services,Serverless Security Strategies: Under the Hood Fireside chat: https://pages.awscloud.com/AWS-Fireside-Chat_2020_FC_e06-SRV.html

View Details

Watch this episode on YouTube: https://youtu.be/hFGrKjtWsQg

Transcript:

ON BEST PRACTICES...

Episode #1: Alex DeBrie
Asking about best practices and the reality of implementing them...
@4:24
Alex: I think it's pretty fascinating to see. Like you say, if you're on Twitter and you're following a lot of the big time people doing serverless architectures in this space, they have a lot of great tips around best practices, and this is what you should be doing, all that stuff. But I find, as I'm building serverless applications or as I'm talking to customers and users that are building serverless applications, there are times when there's tension between what the best practices are and what their circumstances are. This could be because maybe they're not coming in with a green field application, or maybe they have a data model that doesn't fit DynamoDB or something like that. It's difficult on how you sort of square that with recommending something that you know isn't the best practice or the most pure serverless application, but you also gotta help people ship products, right? I think balancing that tension can be tough at times.

Episodes #18 & #19: Michael Hart@30:25 (Ep. 19)
Michael: There's nothing special about Lambda in this respect. It's like, this is just sort of best practices if you were calling any API or if you're writing any API that if you're waiting for many, many, many, many seconds, then you might want to deal with that. And those are the sorts of use cases where I think, okay, fine, that's perhaps not a good practice. You actually, you asked me. We use this at Bustle. So we have a Lambda that renders our frontend HTML code. It's a preact app. It does service side rendering of the HTML, but it delivers to the, you know, to the browser via API Gateway and a CDN and things like that. But it calls our other Lambda directly, which is a GraphQL backend. It calls that to pull in the data that it needs to render the HTML page. Now in the browser, it also will call that GraphQL backend , but it will do it via API Gateway. Because it's coming from the browser, so it needs to make some an authenticated HTTP request into the function. But when you're in the Lambda world, well, that Lambda can just call that Lambda directly, and call the GraphQL Lambda, and that goes to Redis and Elasticsearch wherever it needs to pull the data and send it back. And we just make sure we have the timeouts tuned such that, you know, I mean it responds within milliseconds. It’s not even a thing we would really run into.

ON INSTRUMENTATION...

Episode #2: Nitzan Shapira
Asking about the need to automate instrumentation...
@29:35
Nitzan: Yeah, by the way, it's not just worrying. It's not just the fact that you can forget. It's also just going to take you a certain amount of time - always - that you're going to basically waste instead of writing your own business software. Even if you do remember to do it every time, it's still going to take you some time. Some ways that can work [are] in embedded in your standard libraries that you work with. If you have a library that is commonly used to communicate between services, you want to embed that tracing information or extra information there, so it will always be there. This will kind of automate a lot of the work for you. That's just a matter of what type of tool do you use. If you use X-Ray you're still going have to do some kind of manual work. And it's fine, at first. The problem is when you suddenly grow from 100 functions to 1000 functions — that's where you're going to be probably a little bit annoyed or even lost, because it's going to be just a lot of work and doesn't seem like something that really scales. Anything manual doesn't really scale. This is why you use serverless, because you don't want to scale service manually.

ON APPSYNC DATA OWNERSHIP...

Episode #3: Marcia Villalba
Asking about what service owns the data when using AppSync...
@36:20
Marcia: Well, then, it's the question on who owns the data, and that's something, at least with AppSync, I'm still trying to figure out how to really architect my application, my graph qualifications, because I've been using GraphQL with microservices, and usually I do the filtering in the microservice because the microservice knows the data, knows who can see it, and I don't want to leak that information out. But with AppSync, at least applications and have been building, they are mostly contained into Dynamo tables and Lambdas. So I think when I'm coding this that AppSync is the owner of this data and, then I do the filtering in the resolvers. So I think it's always a question of who owns the data and who is able —where is the level that you want to leak the information out? I don't know if it's clear.

ON WORKFLOW COMPLEXITY...

Episode #4: Chase Douglas
Asking about the complexity of the development workflow...
@6:35
Chase: Yeah, for all the benefits you get from serverless, with its auto scaling and its capabilities of scaling down to zero, which reduces developer cost, you do have some things that you have to manage that are a little different than before. One of the key things is, if I've got, like, a compute resource like a Lambda function in the cloud that has a set of permissions that it's granted and it has some mechanism for locating, the external service is like SQS queue or an SNS topic or an S3 bucket. So it has these two things that it needs to be able to function the permissions and locations. So the challenge that people often hit very early on in serverless development is if I'm writing software on my laptop and I want to test it without having to go through a full deployment cycle, which may take a few minutes to ah to deploy the latest code change. Even if it's, ah one character change up to the cloud service provider. How can I actually test with proper permissions and proper service discovery location mechanisms from my laptop? What mechanisms are there to do that? That's something that we are always evolving.

ON EVENT-DRIVEN ARCHITECTURE...

Episode #5: Mike Deck
When asking about event-driven architecture...
@27:35
Mike: I think that it's probably easiest to understand it when contrasted against kind of a command-driven architecture, which I think is what we're mostly sort of used to. So this idea that I've got some set of APIs that I go out and call and I kind of issue commands there, right? So I maybe have an order service. I'm calling create order or I've got downstream from that. There's some invoicing service now, and so the order service goes out and calls that and says, "Create the invoice, please." So that's kind of the standard command-oriented model that you typically see with API-driven architectures. An event-driven architecture is kind of, instead of creating specific, directed commands, you're simply publishing these events that talk about facts that have happened, you know these are signals that state has changed within the application. So the order service may publish an event that says, "hey, an order was created." And now it's up to the other downstream services to, they can observe that event and then do the piece of the process that they're responsible for at that point. So it's kind of a subtle difference, but it's really powerful once you really start kind of taking this further down the road in terms of the ability to decouple your services from one another, right? So when you've got a lot of services that need to interact with a number of other ones, you end up kind of with a lot of knowledge about all of those downstream services getting consolidated into each one of your other kind of microservices, and that leads to more coupling; it makes it more brittle. There's more friction as you're trying to change those things, so that's a huge kind of benefit that you get from moving to this event-driven kind of architecture. And then in terms of kind of the relationship to serverless, obviously with services like AWS Lambda, you know, that is a fundamentally event-driven service. It's about being able to run code in response to events. So when you move to more of this model of hey, I'm just going to kind of publish information about what happened, then it's super easy to now add on additional kind of custom business logic with Lambda functions that can subscribe to those various different events and kind of provide you with this ability to build serverless applications really easily.

Episode #30: James BeswickAsking about thoughts on embracing asynchronous patterns...
@15:15
James: I think a lot of what we're building makes distributing computing just easier for developers, and when you think about the scale of lots of developers now have to face with their applications, even things like mobile apps, these are complicated problems to solve when you get spikey workloads and just huge numbers of transactions coming through. So a lot of these tools just make it that much easier.

But the mental hurdle is going from this synchronous model to this asynchronous model. And so if you're used to building synchronous APIs, initially it can seem a bit alien trying to figure out the different patterns that are being involved. But it seems like the natural evolution given the fact that you've got all these services in the middle that have to handle all of this traffic, and the timing issues involved, you know, start to evolve from where you are in the synchronous space, but I think what's been put in place is not too difficult to understand.

Once developers start using this, they find actually for many cases, it's the right way to go, but it's interesting to watch this because I know that just even 12 months ago people were talking about the API Gateway, this 29-second, 30-second limit problem, do all this stuff throughout your infrastructure. Or you heard about the Lambda limits of five minutes, then fifteen minutes because people were trying to work this way.

I think now we're going back to thinking about how do we break up these tasks. So it's shorter-lived tasks that run between services in an asynchronous fashion. So the whole model is really evolving.

Episode #40: Eric Johnson and Alan TanAsking about storage first...
@37:00
Eric: Yeah, and I do call it storage first, or sometimes I do a presentation called Thinking Asynchronously, but the idea ... Let me go back to my earlier statement. The most dangerous part of an application that I'll ever build is my code. Right?

So, when I build an application, I want to get that data stored first. That's the thing. I tend to go DynamoDB because that's what I like, that's what I use, but there's different purposes.

I know Jeremy, you and I have had this conversation before, and you're an SQS guy, so that's where you tend to go, and we do this because we look at okay what's the pattern for the retry or the DLQ or different things like that.

For me it's because I'm going to continually write back to Dynamo. On the app, I'm specifically thinking about it. But the idea is if API Gateway can directly integrate with the storage, be it S3, be it DynamoDB, something like that, then I've stored the data and I don't have to go back to the customer if my logic fails, right?

So, in an application I've stored the data, let's say I'm using DynamoDB, I do a stream, it triggers a Lambda, I start processing that data. If somewhere in there, something breaks, and again, it's going to be my code, but let's say something breaks, then I don't have to go back to the customer and say hey guess what, I blew it.

Can you give me your data again? Can you resubmit that? And continue to trust me, because I'm sure I won't lose it again. Instead, I've got that data stored, and I can write in some retry or take advantage of the retry from an SQS or an SNS or something like that.

So, I think it's a really cool pattern for building resilience into our application. Serverless comes with a lot of resilience anyway, that's how AWS has approached this on look as much as we'd like to say nothing ever breaks, let's write as if it does, right?

So, let's degrade gracefully. I think this adds even another layer of that, where I can degrade in my code and know hey I've still got the data. I can write some retry logic. I can use existing retry logic. I think it's a safer pattern.

It does require ... The storage first is the pattern I call it, but it requires thinking asynchronously. What can I do after I've responded to the client and how do I work with them?

Episode #41: Paul SwailAsking about the things developers need to think about with asynchronous applications...
@40:50
Paul: There are a few things. Number one, I would say distributed tracing. If you have a multi-step use case where there's a lot of data processing going on in the background, you probably now have multiple log files to search through. There are... If you were doing that synchronously, you could just look in the one place, more often than not. There are strategies around using correlation IDs within each message so that the same, say in CloudWatch you can query on for that correlation ID and get an aggregate of any log entries across your different log groups, which have that correlation ID within it. You need to build that in to your application, your Lambda code, that doesn't come out of the box. Another consideration is testing, writing automated tests. It's just harder for asynchronous workflows, so if you're writing asynchronous... Say I write in Node.js and Jest test frame work. If I have a synchronous Lambda, it's generally pretty easy to write an integration test for that. You just hit the end point, invoke the Lambda function, and just verify the response. If you have a multi-step asynchronous data processing workflow, then you need to test each one of those individually, during an actual...Writing an end-to-end test is difficult. But I just like having a wait step, that just waits until background processes have, you hope, have completed and then you can do whatever verification steps you need. Just generally understandability, it's not a thing in itself but it's just for if you've got a new developer on your team... A lot of teams that I've worked with are more full stack web developers which are used monolithic synchronous workflows. Got a new guy on your team and it's just explaining to them how each piece of the pie fits together. That's just going to take time and documentation really is the only solution to that. It's just... Good documentation is important when you've got these asynchronous workflows.

ON DISTRIBUTED TRACING & MONITORING...

Episode #8: Ran Ribezaft
On the importance of distributed tracing…
@28:02
Ran: Up until recently, like recently, like two or three years, I would say that distributed tracing is not a mandatory thing that each R&D team needs to have as part of its arsenal of tools. Today, I think it's almost like a crucial or vital thing that you need to have. The main reason is that we already know that applications are becoming more and more distributed. So, for example, once a user is buying something at your store, you want to make sure that it gets the email to him with the receipt and the invoice and so on, as soon as possible because otherwise he's hanging there, waiting for confirmation or waiting for something to get to him. And in a monolithic way, it's been pretty easy because you had something specific, a single thing that will take care of everything. But now we've got, like between 3-300 services that might take care of this operation: one that will get the API request from the user, from the Web server, the other one that will parse the user request. The third might be something regarding billing that will charge through Stripe or through another service. The fourth one could be something that is mailing users, and all of them are connected to each other, with some messages that are running from one to another. It could be like a star, or it can be like 1-to-1, all the way up until it gets to the email service. And without distributed tracing, you wouldn't be able to ask yourself this question: how long does it take for a user once you buy something until the moment he gets his confirmation. Because if it takes, let's say, for example, a ridiculous number. Let's say one minute. It's not good. I don't want my user to wait one minute in my website for confirmation. I want it to be, let's say, sub-second or let's say sub-five seconds. Other than that, it doesn't meet my SLA. And only with distributed tracing can I really measure end-to-end traces and not just a single trace every time.

Episode #12: Emrah ŞamdanAsking about when to take action when there is an error...
@16:30
Emrah: You know, most of the tools that, both with CloudWatch and with the other monitoring tools, did the alerts are just for a single error. So you're just having a one Lambda invocation, then an error happens, and most of the time they're paging an alert. But we thought is this something that is actually wanted by people. Is this something that prevents people from alert fatigue? You know, we are coming from OpsGenie, and that's why we are very, very careful about not putting people into alert fatigue. So we ask people, “What is the definition of failure for you?” Like we asked tens of people,”What do you think? When do you think that this serverless architecture has failed?” And the response is that not a single error, like most of the time, it’s not an error. So when I call something an incident, when it causes something cascading failures. So I have a problem with Lambda function, and this Lambda function should should have triggered another Lambda function through SNS, and this triggers another Lambda function to, let's say, SQS. I'm just throwing out a scenario here. So this first Lambda function fails and the others couldn't even start. So in this case, we can understand that we are in very big trouble, that we lose some transaction there. That’s a failure for most of our people that we talk with and the other stuff is that for, at least, especially for upper management, the invocation duration, invocation count, any kind of abnormality about these metrics, are not very important. And they are seeing cost as a signal of failures. So let's say when they want to allocate $10 per day in to the serverless architecture and let's say, one dollar per a function. In such cases, they want to get alerted when the cost exceeds this threshold. They are not interested in if the function is running more than expected because of a third-party API. They're not interested in if the problem happens because of an input error. They just wanted to see if the cost is exceeding something, some threshold. Because all off their motivation was, when joining to Lambda, to save cost. And they don't want to read that with a problematic situation.

ON SERVERLESS COST...

Episode #6: Erik Peterson
Asking about cost...
@14:32
Erik: There is that tension there. I mean, sometimes it's a healthy tension, but there is that tension there, and it kind of, I mean it goes like this: imagine you needed to explain, let's say, a very complicated system that you constructed, and now you're trying to explain it in French to the Germans, right? You're speaking a different language. And that's the hard part, right? You know, you can ask the question, "Well, why did we spend $20,000 this month on EC2 more than we spent last month," for example. And well, it's because the product team had a new initiative. We had to do a migration. We had to do this. We had to move data from over here. We had security requirements, so we needed to encrypt the data. So we're calling the KMS API a lot. And then that resulted in a whole bunch of new storage and processing. And you're talking, talking, talking, and then you look up and there's just a glazed-over look on the finance guy's eyes and they're going, "Yeah, no, no why did we spend $20,000 more this month? And how much are we going to spend next month?" And they go, "What? I can't talk to you. Get out here." Right? And ultimately you want to tie it back to well, look, this product initiative cost this much money, and we forecast it to be X. And we have an idea, before we actually go down that path, how much it's going to cost and cost has been a part of it. Because there, I mean, for a long time in engineering has been a notion of non-functional requirements, right? What kind of performance requirements do you have? What kind of uptime requirements do you have? And the hard question that I think organizations need to ask themselves is, "Well, what kind of cost or budget requirements do you have?" And at what point are you going compromise the the budget for the user's experience or vice versa? You know, you are you gonna go, "You know what. User experience matters at all costs. Even if it's $1,000,000 in extra spend this month, our users must be absolutely happy." Okay. Make that decision consciously. Today, I don't think anybody's consciously making that decision.

ON RESILIENT ARCHITECTURES...

Episode #9: Gunnar Grosch
Asking about common things missing from resilient applications...
@33:00
Gunnar: One common thing that I see when we're performing these types of experiments is that, like we said before, we don't have graceful degradation, so that the systems they show ever messages to the end users or, parts just don't work but are still there. So we don't have UIs that are non-blocking. And that's a perfect use case for a chaos engineering to be able to find those on and, well, then fix them.

Episode #51: Adrian HornsbyAsking about acceptable levels of staleness in the data...
@56:19
Adrian: Well, even if you claim your application is very dynamic, and you claim that, no, I need to, I can't cache because for example, it's a top 10 list of real time trends on Twitter. Let's say Twitter trends, right? People expect that it's real time. So, I would say by default, if you think about real time, people wouldn't think, "Okay, I need to cache that." But, if you have millions of clients around the world requesting that data, absolutely you're going to fake it real time. It's, you might query your downstream server or service that tells you the trend, but maybe if you have thousands of clients connecting at the same time, you don't want each of those clients to query your service, you will just serve it from cache, or make sure that the requests are packed into one single request and then that's the downstream service and then serve back the content.

So it's just this idea of like any application out there, even if you think it should be a, must be real time, it's very important to think about the staleness. And staleness is how real time my data needs to be, even if it's maybe three seconds old, is it really that old or is not usable? Because it's also something you can fall back. So if your database is not accessible, it's like, maybe you can serve back the trend of Twitter, that was maybe an hour ago, and just instead of... and you can say to your customers, you can say, "Oh, we're experiencing issue, this is a trend one hour ago." And that's fine, it's a good UI. It's good use of stale data. And why would customer say, "Oh, you're cheating on us?" No, it's like... I think it's a good example of that.

ON TESTING...

Episode #10: Slobodan Stojanović
Asking about local testing…
@19:32
Slobodan: So in the early days, when I started working in serverless, I tried to just do the things that I did with non-serverless applications. So my first try was like to install DynamoDB locally and use it as a local database and test against it and things I got. But then there were so many different services such as like Cognito or SNS or many other things that they can, cannot just install locally. So there are some things that can simulate them. But these are simulations. You don't want to run your tests against simulations because they will not give you the right results. And then the second, my second try was basically to mock these things. So I tried to find some complex mocking libraries and things like that that will mock everything and return realistic results and things like that. And even that is like leading to so many errors. And some things are not mocked. You have some special things in your code or we're using the old library, so you need to send some pull requests and who knows what. So these things become more and more complex, and in the end we just decided not to do that. Instead, we want to run our tests locally. Unit test mostly. And whenever we want to test, to run integration tests, I want to test my code against real DynamoDB, which is on AWS or real SNS or real services that are on AWS and that they're working in the cloud.

Episode #47: Mike RobertsAsking about integration tests...
@16:06
Mike: Yeah. And what integration tests are about are testing your assumptions effectively. So when I talked before about functional tests, I said, "So we're going to mock or stub the response that comes back from DynamoDB and make sure that we're doing the right thing with that." That makes an assumption that we've correctly defined what comes back from DynamoDB. And so what integration tests do is validate those assumptions, they validate how you expect your code to run within the larger environment and the larger platform.

And we absolutely advocate for doing that, but remembering that running and maintaining integration tests is a costly exercise. They take a long time to run and they also take a long time to maintain because things change over time. And so, we put a lot of work into the integration test section of the book. And John did this extraordinary thing with Maven, and those of you that are Java developers understand this, but where we run Maven test, which is one command line, and what that does is it brings up an entirely new stack of all of the components in our application, runs all the integration tests against it, and then if the tests work, then it immediately tears that stack down.

We wouldn't have gone into that amount of effort to get that stuff working if we didn't think integration tests were valuable, but we also understand that because they're expensive based on in terms of our time and computer time, that we want to minimize the number of those that we write. And so we're looking normally at just a few, but capture hopefully a number of cases, but again, we're thinking, we're not testing the code when we're writing integration tests. We're testing our assumptions about the larger environment. If we want to test the code, that's when you write a unit test or a functional test.

ON SECURITY...

Episode #11: Hillel Solow
Asking about tools that can help protect developers from runtime attacks…
@32:14
Hillel: My number one cliche would be: let's focus less on mitigation and more on prevention. I know that's super cliche in security, but I think here, one of things that we see a lot is that you get a lot more mileage out of trying to make sure that the things you're deploying are deployed with least risk than you do at trying to chase after attacks. And again, that's not to say we don't need to do both. We will forever need to do both. No amount of proper configuration hygiene is going to prevent every type of injection attack or cross-site scripting attack, or whatever is on our infrastructure, right? We need to mitigate all of those things. But in cloud applications and particularly in serverless cloud applications, the value of hygiene and posture is much greater than it was in the past. You know, whether it's the things we talked about earlier, like just leaving around old stuff that could put you at risk but you don't need, or it's getting IAM configured properly, or it's things like setting timeouts to their minimum threshold, if you can. All those things are going to give you a tremendous amount of value in making it hard for attackers to do what they want to do on your system, before you even started looking for a SQL Injection, right? You still need to look for a SQL injection. We still need to run the tools that we’re going to run. But before we get there, before you start worrying about blocking and mitigating kind of runtime attacks, spend a significant amount of energy on: What do I have? Where is it? Do I need it? Is it configured in a way that gives me the least risk? Have I isolated the things I can isolate? Am I doing all that continually? That would be my number one focus.

Episodes #23 & #24: Ory SegalAsking about the legitimacy of serverless security threats…
@32:02 (Ep. 23)
Ory: Exactly. But you have to remember that as a security practitioners and specifically as researchers, we are trying to flag potential future risks. If we were to only look at what's being used and exploited today, we will always be in a dog chase with attackers. So, I think it's very good that security experts and security researchers look for the next attacks in a new technology and finding it before it's being exploited. And so you can then teach developers how to avoid these and hopefully, reduce the attack surface.

So I think it's not necessarily bad that we're pointing out things that haven't been exploited yet. And as you mentioned, regarding the evidence, usually attackers there aren't web forums where attackers share war stories of how they hacked into a system. So obviously, attackers don't publish anything about their techniques. And especially if you talk to the application owners and companies, most of them also don't like to share information. In fact, usually when they give the server a security conference talk, at the end, there's that five minutes that you save for questions. Usually, I know that nobody's going to actually ask a serious technical question because they are embarrassed. It's like something that you don't want to talk about around other people from maybe competing organizations.

So there's no resource to go and look at and see how people are exploiting and what are the vulnerabilities. We collected information from customers and prospects, I've reviewed dozens if not hundreds of serverless Apps at this point, and we collected this information to see what are the most repeated risks that people do.

ON DYNAMODB AND NOSQL...

Episode #17: Brian Leroux
Asking about why Architect chose DynamoDB…
@38:58
Brian: Yeah, I mean, it's a decision making process. And it's one that a lot of people are aren't comfortable with. It’s a managed database, which is a nice way of saying that it's a proprietary database. It's owned and run privately by Amazon. And, you know, after, our history has a, or our industry has a long history of being gun-shy of these databases because of Oracle, frankly. And I don't blame anyone for painting Amazon with that brush. “Oh, my database. That's my data. I don't want them to have that. I want to control it.” The only people that say that, by the way, are people that have never sharded a database, you've sharded a database once, you are happy to let someone else manage that for you. You are more than happy. How much does it cost? Fine. Less than a DBA. So that's going to be a good deal for me. So once you get over that initial concern, which isn't a real concern, by the way, that free tier is extremely generous. You could run a local instance of this thing yourself headlessly if you want for testing and building out locally, so you don't have this requirement of the cloud. And the free tier’s insane. I think you get something like 20GB in the free tiers. So, like you could build a lot of app with 20GB. A lot of app. You could put images in there, you don't want to, but you could. Yeah, it's a great DB. I guess the other thing that people get a little tripped up on is the syntaxes. It’s a bit strange. It’s coming out from a different world. I don't think it actually is that strange, for what it's worth. I'm pretty sure if you'd never seen SQL before and I showed it to you, you’d be like, Well, that's strange. I think just what you're used to. It's a sadly verbose query language. It takes a lot of directives in JSON form to make it do pretty trivial things. We've written a few higher level wrappers for it to make it a bit nicer to work with, but it's all about the semantics. Single digit millisecond latencies for up to a MB at a time querying, no matter how many rows I have? That's unreal. But we've never had a database that can do that. And I'm happy to pay for that capability.

Episodes #34 & #35: Rick HoulihanAsking when NOT to use NoSQL...
@4:50
Rick: So NoSQL is really suited and as we talked about, we have to denormalize the data, right? Which does that means I have to structure it and tune it to the access pattern. So if I don't really understand those access patterns, if they're not really well-defined, then maybe what we're looking at is a different type of application that's not necessarily so well-suited for NoSQL, right?

And that's really what it comes down to. There's two types of applications out there. There's no OLTP or online transaction processing application which is really built using well-defined access patterns. You're going to have a limited number of queries that are going to execute against the data, they're going to execute very frequently and we're not going to expect to see any change or we're going to see limited change in this collection of queries over time.

And that's a really good application for NoSQL because as I said, we have to kind of tune the data to the access pattern. So if I only have a small number of access patterns, then it makes sense, but if the customer comes in and tells me, "I don't know what questions are going to be asked. This is maybe my trading analytics platform and who knows what the brokers are going to be interested in today or tomorrow and I look at the query logs of the server and there's a thousand different queries and some of them execute once or twice and never to be seen again and others execute dozens of times."

These are things that are indicative of an application workload that may be, is not so good for NoSQL because what we're going to want is a data model that's kind of agnostic to all those access patterns, right? And it has that ad hoc query engine that lets us reproduce those results. So lucky for us in the NoSQL world, that's actually a small subset of the applications, right? 90% of the applications we build have a very limited number of access patterns. They execute those queries regularly and repeatedly throughout day. So that's the area that we're going to focus on when we talk about NoSQL.

Episode #36: Suphatra RufoAsking about the adoption of NoSQL databases...
@4:43
Suphatra: Yeah, yeah. I think the way that consumers behave, the retail industry is a good example. You probably didn't know that Sears, Kmart, Barneys New York, Party City, I can name a dozen more retailers that just last year, either completely closed down or had to significantly reduce their number of stores, just last year. It's because retail isn't done the same way anymore. Those spikes are now a common part of life and people are having a hard time figuring out how to handle it. Tesco, which is the largest grocery chain store, I'm not sure in America or in the world, I'll have to check that... But they, in 2014, crashed on Black Friday because they couldn't handle the spike in the demands they were getting online. So they lost an entire day of business on Black Friday because they couldn't handle that workload. And then the year later, they went on a NoSQL database and now they can handle that load.

I think what people are seeing is that normal day to day business operations are fundamentally different. For example, the fashion industry used to have only four clothing seasons. Your mother probably remembers buying a new outfit every season... So winter, spring, summer, and fall. And so women's clothiers would go and create new clothes four times a year. Now the fashion industry has 52 seasons, so every week is a different season of women's clothing, which means there's a spike every week for every launch of every new clothing line. So that's another big database problem that's now just becoming a regular part of life. A decade after NoSQL databases are invented, it's really not a new invention anymore. Now this is just the way of business.

Episode #39: Lynn LangitAsking about the tools enterprises were using for big data...
@9:44
Lynn: Well, change comes slowly and change is usually induced by some sort of pain. And so the pain in my case and my customer's case was through IoT data because IoT data increased the amount of data exponentially because the event based data. So I had some customers, some of the big, big like the biggest appliance manufacturer in the United States. Customers, I can't name, but you can guess who they are. And this was maybe eight years ago, so it was still a while ago. They wanted to IoT enable their devices.

And again, to be very clear, the majority of the enterprise applications that I would work with would be SQL plus NoSQL because they would have a need for transaction. And again, that's really important because I saw those startups go just directly to NoSQL and then they would call me and they would try to tune their transactional consistency of their Mongo and it would be clustered and it would be a mess. And then we just pull that out and put it in MySQL. Just the whole space was super interesting. So meanwhile the cloud vendors are evolving and Amazon of course comes with DynamoDB. And I have to tell you that initially I was super resistant. I was like, how do you even query that? I actually did some time tests and blogged about ... this is like seven, eight years ago.

You write SQL query, everybody knows how to do that. You write a Dynamo query, it takes 15 minutes because you have to research the query and how much is that in your dev time and dah, dah, dah, dah, dah, dah, dah, dah. So there was resistance including me. The service on the cloud that really stunned me and still does is BigQuery because BigQuery offered SQL querying, which I think it's extremely important when you're evaluating different kinds of database solutions to look at what is the ramp up time to understand how to get data in, how to take data out. And the more different the query languages are, the more errors you're going to have too, and this is your data. So I've seen a lot of bad things where developers overestimated their abilities. And because the query languages were really idiosyncratic or esoteric for the NoSQL databases, it was all kinds of problems.

But BigQuery's idea of, okay, you get around the scaling problem by using files and you then just use SQL and you just pay for the time. I mean, I literally, I got goosebumps. I was stunned when BigQuery came out. I was stunned. I really got it from the beginning. And I've written about it, I've used it. I would often add it for customers as sort of an incremental rather than NoSQL.

Episode #44: Alex DeBrieAsking about migrating data and changing access patterns…
@50:10
Alex: Yeah, absolutely. And this was actually a late addition to the book, but I just got so many questions about, I don't want to use Dynamo because what if my access pattern's change, or how do I migrate data? Things like that. So I actually went through it, I think it's not as bad as you think. And I split migrations into two categories, basically. First off, they're just additive migrations, where if you're just adding a new application attribute to existing items, you just change that in your application code, you don't need to change anything in DynamoDB. Or if you're adding a new type of entity that doesn't have any relational access patterns with an existing entity, or if you can put it into an existing item collection of an existing entity, you don't need to do anything. It's just a purely application code change there.

The second type of migration is, I need to do something to existing items, either because I'm changing an access pattern for an existing items or I'm joining two existing items that weren't joined together, or I'm adding a new entity type that I need a relation and there's no existing item collections to use there.

So now you don't only need to change your application code, but you need to do something with your existing data. And that's harder. It seems scary, but it's actually not that bad. And like, once you've gone through one of these processes, these ETL migration processes, they're pretty easy. It's basically a three step process. You're going to have some giant background job that's going to scan your table. You're going to look for the particular items you need to change. So if it's an order, and you need to add GSI2PK and GSI2SK for it, you find your orders in that scan for each order that you find, then you add these new attributes on them.

And then if there are more pages in your scan, you loop around and do that again. So it's just this giant wild loop that operates on your whole table. Depending on how big your table is, it might take a few hours, but it's pretty straightforward, that three step process: scan your table, identify the items you want to change and change them.

ON TRANSITIONING TO SERVERLESS...

Episode #20: Sheen Brisals
Asking about how LEGO developed confidence to go serverless...
@13:10
Sheen: Yes, that's a very, very good and important point because when we often look at a monolith, we often get confused. Where do we make a start? Because it looks everything big. But the thing is, we need to start looking more closely part-by-part. So then we will be able to identify some small entry point into the system that will give us the comfort to, you know, try out something new. And also, when you have, when you work in the organizations, you need to prove or showcase these things to stay stakeholders, you know, to get their buy-in. So for that, it's important that we identify a part of the system that is not complicated, small enough that we can experiment with the new ideas and show them, show them the proof that it's working and that’s feasible for us to go forward to everyone around — not just the engineering team. Bring the business stakeholders, everyone together and yeah, so that's very crucial when, especially when we start this monolith to microservices serverless journey.

Episode #22: Bret McGowenAsking about the transition of existing applications to serverless...
@21:52
Bret: I think one of the messages I want to kind of get across is that maybe moving to something like containers and Cloud Run, it feels like you're taking on way more than just getting started with functions as a service, right? Your Lambda, your Azure functions, or Cloud functions. And there is. There is a little bit more to manage. I don't think it's a huge amount of overhead, but I think what this really does is it enables a lot of your existing workloads to start to be serverless. Because we all have apps that have we started before serverless was a thing. Even if it's not perfect, we would love to get them to be serverless.

Episode #15: Mark McCann and Gillian ArmstrongAsking about advice for people thinking about adopting serverless...
@45:15
Gillian: So I tell people, especially in big enterprises, the same for both serverless and AI, which is: start now. You’re already behind. If you haven't started, you need to start now. It takes a while to learn all the things that Mark’s just said — a lot of things. It takes a while to move your mindset from highly architected things before to serverless. Serverless is very different, even the microservices. So even if you're familiar with microservices. This is still a different paradigm. So it just takes a little while to learn. It takes a little while to move all your existing practices and thinking about how you build your systems. So you need to start. You need to find places that are sort of safe-to-fail places where you can try things out and then gradually scale up. I think the big thing is, if you run the company, do you create time for people to learn. Do let them have that space and, you know, find your people who are really, really passionate about it and then let them loose.

Episode #32: Ken CollinsAsking about advice to other companies adopting serverless...
@36:36
Ken: Yeah, I think it's always to look at what your business is doing right now. And, where you need it to be sort of performing at first of right. So, always drives my success. I'm a very huge believer in DHH the sort of creator of Rails. That you do the majestic monolith first. You build an application out, and then you sort of look at where it needs to either be performing or broken apart. For Custom Ink. If, your story is anything like ours, it would basically be starting with the monolith. Looking to where sort of business units lie in that monolith and then breaking out into what we sort of call key domain services. So, we would extract the design lab from the monolith. We would extract the product catalog from the monolith, we would extract, a group order form and quoting systems and things like that from the monolith.

So, that to me is a really good way to sort of adopt the cloud. If you know you're going to be breaking up to this monolith into smaller parts. Some of them could be Lambdaliths, some of them could be say Fargate or EC2 instances, whatever. But, I think when you look at what's happening with your current application, let your success drive your architecture. And, I think Lambda is a good place for either moving apps, but it's also a really good place for, if you're in AWS. To question if that's an opportunity for you to look at doing data, and events and units of work in a different way.

Episode #33: Yan CuiAsking about the biggest roadblocks that companies adopting serverless are running into…
@9:46
Yan: Well, the biggest one I feel is by far is just education. Like I said, Lambda itself is getting more and more complicated because of all the different things you can do with it. Other roadblocks includes for example some organizations are still holding onto the way they are used to operating. With centralized operation teams, cloud teams. The feature teams don't necessarily have the autonomy they need to take full advantage of all these different tools that you get and all these power and agility you get with serverless, your team can build a new feature in a week, but it's going to take them three weeks to get anything they need done, provisions and to get assets to resources they need. Then again, you're not going to get the full benefit of serverless.

So a lot of that legacy thinking at the organization is still there and is still a prominent problem and roadblock for people to take full advantage of serverless. But in terms of actual adoptions, a lot of it is ... In terms of technical roadblocks, there's some, I think the last question you had was around some use cases that just doesn't fit so well. When you've got a really high throughput system, the cost of serverless can become pretty high. So imagine you've got something that's relatively simple, but how to scale it massively like your Dropbox, not a super complex system, but have to scale to massive extent. So for them it makes perfect sense to move off of S3 and start to build their own given hardware so that they can start to optimize for that cost.

For a lot of companies, they do have that concern as well. They may not have a very complicated system that requires a hundred different functions on this massive event driven architecture, maybe they just have five end points. But those five end points are running at 10,000 or 50,000 requests per second. So in those cases, the cost for using Lambda and API gateway would be excruciating and you'd be much better off paying a team to look after your community's cluster or your containers cluster than have them running them on Lambda.

But that's always a tricky balance. Because, oftentimes you can always get the reverse argument whereby, "Well, Lambda is expensive, so I'm going to just do this myself." But then you're hiring someone for $10,000 a month to look after your infrastructure, and your Lambda bill is going to be, I don't know, $100.

Episode #37: Peter SbarskiAsking about the challenges of training people how to build using serverless…
@2:55
Peter: Look, I hate to use the word paradigm, but it does feel, it is really a paradigm shift. Because serverless, it feels like, this is what cloud was supposed to be all along, right? You're not dealing with low level infrastructure concerns. You're not provisioning your servers and thinking about memory capacity, but you're thinking at a high level of abstraction, you're thinking in terms of code, you're thinking in terms of functions and services and event driven architectures. That's interesting. It's different and it requires people to really think in new ways.

Look, I think, honestly, the adoption of serverless will hang on education. If it can educate people, serverless as a concept as an idea will be successful. I think that's what we're all working towards. This is what you do nearly every day, right? You educate people on serverless. You blog, you talk, because this is the way we get people to understand.

ON THE SERVERLESS COMMUNITY...

Episode #25: Farrah Campbell and Danielle Heberling
Asking about the serverless community being more welcoming…
@11:02
Farrah: I definitely do. Serverless is a new approach. It's inventing a whole new way to build software applications and it's sticking. And this community, I feel like everybody has a lot of work to do, a lot of big things to accomplish and everybody's at a starting point. Everybody's willing to have open conversations without putting others down or explain to you why you're wrong about something. Everybody's at the same starting point and just trying to learn from one another.

ON FULL STACK SERVERLESS DEVELOPERS...

Episode #28: Nader Dabit
Asking about the idea of the full-stack serverless developer...
@5:55
Nader: Right, right. We think that what we're doing is a little different than anything that's kind of been out there before, I think. And we don't really have something to compare it to, but we talk about it in a couple of different ways. One of the things that we talk about is this idea of a full stack serverless development, where you're a developer, or you're a team, or you're a startup, or you're a company and you want to be able to enable a developer or a team of developers to build the front-end and the back end, versus having the traditional maybe engineering team where you have a backend developer and then you have a front-end developer.

We're looking at it like, what if a developer could just be looked at as a full stack developer like we've seen forever. But instead of the traditional full stack developer where the backend developer might be in charge of creating servers and creating a database and patching and dealing with all of the different backend resources, we could take the serverless philosophy, use that and then apply the front-end developers and merge that together and enable a single developer to build out these full stack apps, or a team.

Episode #50: Guillermo RauchAsking about the shift from compute on demand to pre-rendering...
@5:00
Guillermo: Yeah. I think a lot of people in the industry have over focused their attention on computing on demand, which is what Lambda enables, right? You're literally firing up a VM. It's amazing how easy AWS made it. It's almost like a miracle that you deploy your function so quickly, and it executes so quickly, and it's secure and in a VM sandbox, and their underlying technology is absolutely incredible with FireCracker, but the question that you have to take a step back and ask yourself is that do I really want to be computing so much? Do I want to be burning electricity and competing cycles so much?

This is where when we really sat down to analyze this problem, we realized the vast majority of pages that you visit every day on the internet can be computed once and then globally shared and distributed. So it's like the technique of memoization and functional programming where you compute once and then of course, you want to read it from that intrinsic automatic cache that you get, is different from caching because caching requires a lot of developer effort and thinking. Memoization gets closer to what I envisioned to be the foundation of serverless front end, which is basically static generation where the computation happens once probably as a result of some data pipeline, something that changes, computation happens. HTML is spit out.

That is all, and even in the case Vercel, it's powered by functions too by the way, but the funny thing is that the developer never even thinks about functions. They just think about building pages that then get pushed to the edge and then consumed by visitors. Now, that's not to say that the on-demand use case doesn't have any merit. Not everything can be computed statically. There's lots of pages where you sign in to a dashboard, and you have to query data that could absolutely not be cached. A great example is you log into your bank, and imagine that you were trying to statically generate your dashboard with your bank account balance, but you just want to check that your payment went through for utility.

You're not sure if what you're reading is up to date or not. You would go crazy, right? The movement of front end has also led us to where that dashboard is a single-page application, most likely, that is also served statically from the edge. Then there's JS code that runs on the client's side that then queries that back end. What we found is that front end is really powered by this set of statically computed pages that get downloaded very, very quickly to the device, some of which have data in line with them. This is where the leap of performance and availability just becomes really massive, because you're not going to a server every time you go to your news, your ecommerce, your whatever.

You're just downloading it from your very own city, but even in the case of like, "I may have to make a strong read, not a read that could be stale," you're basically also downloading static content that then runs JavaScript on the client, and then that goes to a server. Then the question becomes, "Who's writing that server, and how much of that server are you writing?" This is like the other big question that is, I think, coming up. We're confronting that in the serverless world is like, "Okay, I have all these amazing primitives to build everything in the world that I could imagine from scratch, but does it make sense to build everything from scratch?"

Does it make sense for you to build your own authentication function with Lambda if you could be reusing a standalone authentication service? That's why this interesting world is coming up where there's a rise of the front end, but then there is a rise of the API economy. What I mean by the API economy is that we have services like Stripe and Twilio and AWS Cognito and Auth0, and MagicLink, and all the services where you're just making some quick API calls sometimes directly from the client side, right?

That is a serverless world that seems so much more in my mind attuned with the actual ideal and the actual, original promise of serverless. I think we are too much. You're giving the example of SQS and Dynamo. We erred too much on always rebuilding from scratch a little bit, so focusing on the front end allows you to reprogram your product strategy in a way, where like, "Okay, I'm going to think about the customer first. I'm going to think about building my back end very, very low in my priority list, right?"

ON VOICE AUTOMATION...

Episode #31: Aleksandar Simovic
Asking about voice automation
@12:00

Aleksandar: Naturally, of course, you can order things from Amazon or whatever, but you're now slowly starting to see, for example, print, give me some report, or send this or send a message to someone else. And hidden and seen in the appearance of these Echo Shows in the past two years, where you can even see when something is happening in front of you.

Jeremy: Right, it's like a visual interface that's on top of your voice commands?

Aleksandar: Exactly, and what my kind of prediction, and things that I'm working on, I'll talk about it later, is that we're going to come to a point where things are going to be automated using Alexa on many manual things that we're already doing right now, I don't know, office, office manager thing, office manager tasks, like I don't know, send somebody a reminder or schedule a meeting.

Actually we already can see that using Alexa for business, but all these small pieces are starting ... Like people are starting to discover how can they easily use it, but as serverless evolved, and now people are actually building huge applications on like enterprise scale applications on serverless, and we saw that on Reinvent, at this last year.

This is how things are going to evolve with voice as well, so we're going to see an explosion of higher order kind of software, like more complex software, that's ... You might have an intelligent agent or an Alexa skill that's going to be able to do some financial or maybe do your taxes, you don't know, you know?

So, things are slowing building. These building blocks are appearing. We see AWS is building this whole serverless ecosystem around itself, where it's going to be a piece of cake actually combining these components and creating something out of the box.

ON BUILDING BETTER SOFTWARE...

Episode #42: Susanne Kaiser
Asking about building better software with domain-driven design...
@5:20
Susanne: So your domain driven design comes with a core statement that in order to build better software we have to align its software design with the business domain, with the business needs, and the business strategy. So domain driven design helps you with aligning your software design with the business domain needs and the strategy and it's very crucial for building your software solution, because otherwise, you are building something that, for example, are matching the requirements of your users. Instead you have to collaborate intensively with your domain experts to gain domain knowledge and to understand the problem first before you're solving it. We are tending to jump directly into solving a problem technically and, yeah, yeah, we can just let's deploy it on a Kubernetes cluster, but we have not understood the problem first. That's really crucial to the build better software.

Episode #49: Jared ShortAsking about building code for transparency...
@48:20
Jared: Sure. Yeah. And I mean, a lot of companies are building towards, I'd say short-term right now value versus long-term stability. And I get that. It makes a lot of sense, especially financially in certain cases. If I need something right now versus, what's this going to look like in six months, when we circle back to it.

But ultimately building as if you're going to open source at any time, I think forces you to at least think if somebody was looking over my shoulder right now, right? If I was building something and say, "If Jeremy's looking over my shoulder right now, am I really going to put like my GitHub magic key, or whatever into this line of code and just hard code and deploy. And be like, 'I'll fix that later?' I feel like Jeremy's going to be back there and be like, 'Really man? I respected you and now no, like there's nothing.'"

So I think it helps you justify the few extra minutes or in some cases, hours or days, to make the right technical decision. And I get tech that's a thing. Look, go look at open source code. There's tons of stuff out there that's like to do actually make this work appropriately or optimize this thing. That's fine. Nobody's going to judge you. We all get it. People write software and they understand software is hard. But they are going to judge you pretty hard if you make poor security decisions or poor architectural decisions where it just doesn't make sense. And I think having that fictitious open source gazer over your shoulder, it just helps you make those decisions and kind of think to yourself.

ON USING CLOUD BUILDING BLOCKS...

Episode #48: Linda Nichols
Asking about trusting cloud services...
@19:26
Linda: Yeah, absolutely. And I mean, I think in the cloud in general, because when we were talking about standing on the shoulders of giants before, we were talking about using NPM packages or Ruby gems or whatever. And like you kind of just don't even know who's written those. And there's some security kind of considerations there. And also, like, you just don't know how much they're tested, you don't know if that maintainer is just going to go find another job somewhere. When you're talking about cloud services. This is true for all the major clouds. I mean, they are tested by millions of people, they are used billions of times. So I mean, if you... No one's going to say like, oh, I just don't think that... Fill in the blank servers, like Cosmos DB. Like oh, I just don't think it's really tested that much. Or like, what if the person who works on it leaves. Well, there isn't one person, right? It's a whole team of people, and there's a company that supports it.

So I mean, I think it's somewhat it's kind of ridiculous when people don't trust cloud services. If you don't want lock in that's a whole other discussion, right? But you if you're already in a cloud, I mean, a lot of people are multi cloud also, and they kind of spread things around to try to minimize that sort of lock in feeling. But really you get locked into libraries too right? If you're writing... If I'm writing a Node.js app, and I'm using some NPM package, yeah that thing's going to stick around forever. How often do you go back and switch out your NPM package?

ON THE FUTURE OF SERVERLESS...

Episode #7: Taylor Otwell
Asking about the future of serverless...
@31:26
Taylor: Yeah, I think the next five years will be huge for serverless, I really do. I think it is the future, because what's the alternative, really? Like more complexity, more configuration files, more weird container orchestration stuff? I don't really think that's the future, you know, that people are gonna naturally gravitate towards. I think people want simpler things. And I think at the end of the day, serverless is simpler. It's going to only get more simple as the tooling gets better, as the platforms get better. And to me, it's the real endgame, you know, of the whole server thing, just deploy your code and you focus on your code and let the provider focus on the infrastructure.

Episode #38: Ben EllerbyAsking about what serverless looks like in five years…
@24:10
Ben: I think it's going to be more abstraction and more consolidation around how to do things. So in five years it's going to be more obstruction around that configuration so you're not having to manually configure retry policies out of the box, you're sort of being able to sort of, well maybe not pointing click, but in a very short amount of yaml we'll be able to have an event driven architecture and maybe that becomes formalized, an event sourcing service rather than just an event bus. Maybe other areas of event driven become more formalized, but it's always going to be an increase in abstraction. We went from virtualization to containerization to function as a service and other things as a service. Now we're sort of building more event driven. We can have obstructions at different levels, so there might be obstructions in particular services or abstraction of the whole architecture.

Things like the serverless framework have tried serverless components, AWS has tried the serverless application repository. And those things have varying degrees of success, but built into all of these services, although we're going to give it more, we seem to have had a spike of complexity recently as so much has been announced and so much has been released. I think we're getting more abstraction. If we take the amazing work you did about integrating RDS with Lambda and you built a whole sort of open source project that really helped people with that. Recently, although there's still a need, AWS has abstracted a lot of that with the RDS proxy. They've seen a need from the community and that abstracted that. If we take EventBridge, people were doing CloudWatch custom events kind of before and hacking it and then they formalize that and provided a level of abstraction. Now is that abstraction going to be driven by the cloud provider or by the community? Well I think it's going to be a bit of both. AWS and other cloud providers are going to add more services but also increase abstraction as it goes on and the community is going to build amazing open source projects that increase abstraction. So I think the move to more abstraction, which means less configuration and configuration is just code. So it means less code to achieve the same business value. So for me it's more abstraction, but right now it feels like less abstraction.

Episode #52: Tim WagnerAsking about if not addressing state will prevent serverless adoption…
@37:36
Tim: I have these two strong reactions to that statement, right? One of them is I would say in some ways the most successful thing Lambda has done is to challenge thinking, right? To get people to say, do you really need a server stood up, turned on taking 20 minutes to fire up with a bazillion libraries on it and then you have to keep that thing alive and in perfect condition for its entire life cycle in order to get something done in terms of a practical enterprise application? And challenging that assumption is one of the most exciting, important and successful things that I think Lambda and other serverless offerings have accomplished in our industry. The flip side to this is to be useful, sometimes you have to be practical. And it's equally true that you can't walk up to an enterprise and say, "All right, step one, let's throw all your stuff away and then step two, you're not going to get past step one."

It's funny, we talk about greenfields, brownfields, it's all brown in the enterprise. Even if you write a net new Lambda function, it's running against existing storage, existing data, existing APIs, whatever that is. Nothing is ever completely de novo. And so I think to be successful and be as adopted as possible in the long run, serverless offerings are going to also have to be, they're going to have to be flexible. And I think you see this with things like provision capacity. I mean, when I was at Lambda still, we had long painful debates about is this the right thing to do? And for understandable reasons, because it is less stateless. It took the ... it's obviously optional. We don't force anyone to use it. But by doing it, it makes Lambda look more like a conventional, well, server container, conventional application approach because there is this piece that is a little bit stateful now.

And I think the arc here is for the serverless offerings to not lose their way, to find this kind of middle ground that is useful enough to the enterprises that still challenges assumptions that gets people to write stuff in a way that is better than what came before and doesn't pander completely to just make it feel like a server. But is also practical and helps enterprises get their job done instead of just telling them that ... because just sermonizing to them is also not the right way to do it.

View Details

About Tim Wagner:

Tim Wagner is known for starting the serverless movement with the original business plan for AWS Lambda, and served as general manager for three of their central serverless offerings: Lambda, API Gateway, and the Serverless Application Repository. After AWS, Tim helped lead another bleeding-edge movement, driving forward blockchain innovation as the VP of Engineering at the digital currency exchange platform Coinbase. Tim is currently working on a new stealth startup, Vendia, with more information to come on June 26th.

  • LinkedIn: www.linkedin.com/in/timawagner/
  • Twitter: twitter.com/timallenwagner
  • Medium: medium.com/@timawagner
  • Vendia: www.vendia.net/

Watch this episode on YouTube: https://youtu.be/M6I0ay5R884

Transcript:

Jeremy: Hi everyone. I'm Jeremy Daly and this is Serverless Chats today. I'm chatting with Tim Wagner. Hey Tim. Thanks for being here.

Tim: My pleasure. Thanks so much for having me.

Jeremy: So you have a lot of history. There's a lot of stuff that we're going to get into today, but right now you are the CEO and the cofounder of Vendia. So I'd love it if you could tell the listeners a little bit about your background, your history, and then what Vendia is all about.

Tim: Sure, sure. So last few jobs here. I mean, I started what eventually became AWS Lambda at AWS. Joined there back in 2012, we launched that in 2014. And that taught me a ton, not just about how to run a business in the cloud, but also about how you build these massive horizontally scalable cloud services. Then I spent some time down here in San Francisco at Coinbase, a US-based cryptocurrency exchange. And I learned a lot about a different kind of scale, which is how you run these massively scaled ledgers that can hold really important information, for example like somebody's bank account. And then Vendia is in some sense kind of the combination of these two things.

I took everything that I've learned over the last seven years and my cofounders Shruthi Rao and I have brought that together to create a business to help companies break down some of the data silo and information exchange problems that they've got today. So we're still in stealth mode for a few more weeks, but I can tell you a couple of things about it. For one, when I sold AWS Lambda, customers were always excited about the product, but they also always had two concerns. First, it was an inherently proprietary technology specific to AWS. And then secondly, while it was this awesome solution for compute, it didn't kind of come preset for data solutions or a solution for state. And so with Vendia, we're trying to reimagine how companies can go serverless and then at the same time solve some of the biggest baddest challenges they've got around data silos and vendor lock in at the same time. By the way, speaking of serverless, Vendia's also proudly server and container free.

Jeremy: Awesome. So that's awesome first of all, and I'm excited for Vendia. I really am interested. Anything that you do is just gold. So I think that this is going to be pretty exciting and I can't wait for it to come out. But what I'd really like to do today since I have you, I mean, for all intents and purposes and I think you always say this lovingly, but you're really the father of serverless, right? I mean, Lambda is what kicked off this whole thing. And I know that there were other companies that this sort of like a fast type thing, but not anywhere near to the scale that that Lambda did. And I would love to hear that story. As a fan of serverless, as a fan of AWS Lambda, could we go back to the beginning and just maybe give me a little, some insights into how this all started?

Tim: So a little bit of the Lambda origin story, huh?

Jeremy: Yes. Please.

Tim: Yeah. So we roll back the clock. It's 2012, I get hired into AWS and it's my first day there. And my boss Alyssa Henry, who at that time is running all of storage, so S3, EBS, like the whole storage division for AWS sits me down at lunch and says, "Okay, Tim, so here's the deal. We heard from customers that they love S3. It's simple, it's easy to use. It's a different kind of way of thinking about the cloud. They love all of that, but it's just a storage solution, right? There's no way to ... Let's say you store an image, there is no way to make a thumbnail of it. You pull out a compressed file, there's no easy way to decompress it on the fly plus the other million things developers might want to do with the stuff that they're storing in here.

So they've told us this in customer advisory meetings and one on ones, see if he can do something with that. Okay. I'm busy, got to run. Good luck." So this is day one for me at AWS. This is literally my very first conversation coming out of the sort of the onboarding and signing up all the paperwork. So I'm like, "Okay, grow a business in the cloud. Make it easy and think about S3 as a kind of inspiration." And it's funny because a lot of people think that Lambda grew out of EC2 and it's obviously a natural extension of thinking about compute in the cloud, but it really came out of the S3 organization. And it was this kind of kissing cousin to the idea of making storage super simple. Back then S3 basically did PUT, GET and LIST. That was it.

And so the idea is what is the ... this is sort of the remit that we had. What is PUT, GET, LIST for compute? What does that ... What if you could just say run or what became invoke in the cloud and you could make a service like that? So we got started. We did, I think as Amazon is famous for doing, we worked back from customers. I did just dozens and dozens of calls with some of the folks who were some of the biggest and frankly some of the smallest AWS customers at the time. And we asked them, "How would you like this to work? What would you want it to do?" And we went through lots of, as anything finding product market fit, the false starts. At one point we thought maybe this is like a scripting service. It should be a scripting language. We could call it Amazon simple scripting service. And then we realized the acronym maybe didn't work the best for that.

So from domain specific imagery stuff to scripting, to finally landing on, no, really the challenge here is make compute simple. Then we realized we were onto something when we realized that the first million developers using AWS are not the ... They're not the next 10 million developers. We had to make the cloud as easy for someone who does applications and business logic as it is for someone with a PhD in distributed systems. And that's when we realized like there was some there, there. And so we got excited about that. We came up with this idea for event hookup and we were kind of off to the races.

Jeremy: Awesome. So I love that. And now obviously you mentioned product market fit, so there's no way you got this thing right on the first shot. Right. You must've had to go through a million different iterations. So what did you get right and what did you get wrong?

Tim: Yeah. It is funny like you think where's the crystal ball clear and where was it maybe a little bit muddy here? I think one of the things we got right and I say this without ego, I mean, because this was a lot of us working hard on this was the event piece of this. We realized that there's a lot that you can do to make asynchronous event generation and handling really easy. It's a super powerful paradigm. And if you look at the stuff in Lambda that's probably been the most, some of the first things that accompany adopts around things like cron jobs and simple events coming out of S3 and also sort of where Lambda's got a lot of its scale and initial success, a lot of it has been around those asynchronous and event handling mechanisms.

And obviously AWS has continued to double down on that with services like event hub that make that even easier to do. So I think then on the what did we get right, easier way to compute events. The idea of making it multi-lingual not tying it to a single language or necessarily a single paradigm. So making it as broad as possible. Things that we got wrong. Well, I've told this story before. I remember sitting in Andy Jassy's conference room. And for those of you who haven't ever worked at AWS, Andy's conference room was called "the chop." So different conversation about why it's called the chop, but it's called the chop. And so when you talk about going to the chop, it's this big thing. Andy Jassy, all his directs are there. It's this high pressure environment.

And I remember sitting in the chop and the guy who was running sales at the time asks me, so he was like, "I got to sell this crap you're about to make here dude. So I got to know what's it good for and what's it not good for? Tell me something a customer will never do with it." And I'm like, "Oh Adam, no one's ever going to use this for video transcoding. That kind of dense compute, we'll never get that. We'll never have that with Lambda." So of course a cloud guru, has been up on stage at the Serverlessconf talking about how fantastic Lambda is for doing video transcoding.

One of the, in fact, the fastest known algorithm for video transcoding beating Google and all other kind of in practice mechanisms is based on Lambda. Some great research out of UCSD and other places. And so this is a good example of getting it wrong, where we thought it was going to do one thing and in fact, the developers showed us that it could be so much more and really just a much, much broader set of use cases than we had ever imagined.

Jeremy: Yeah. And I know you've been away from AWS for a while now, but during those early years of Lambda, were there ... I mean, obviously you're rolling this thing out. There are people adopting it, the adoption curve has been somewhat slow. I mean, think it's sped up now, but like were there missed opportunities early on? Are there things you could have done better you think that maybe would have sped up that adoption?

Tim: Yeah. It's a great question. And one of the hard things to balance, I mean, certainly I'm encountering this again with Vendia is how do you blend the top-down and the bottom-up, right? Your fastest path to revenue is picking a few very large enterprises and trying to sell them something and your best path over the long haul to a broad successful adoption in anything IT or developer related is to get millions and millions of developers to love an experience. And so of course the best of all is when you can do both of these things, but that takes time. It takes time and energy. And I know one of the things we wrestled with in our first couple of years here was how do you balance those trade offs?

Every minute you spend on evangelism and developer education and docs and ease of use features is a minute that you're not spending helping a John Deere or a Nike or a Nordstrom or somebody else become incredibly successful at making their business soar. And so that trade off was tricky. And I think in our first couple of years we had some missteps there in terms of trying to figure out how to blend those kinds of activities and it took a little while. I mean, the other practical reality is that the more innovative something is, the more different it is. And the more different it is, the harder it is to get people to understand it, adopt it, integrate it. And I think you're still seeing some of that. Containers are a small step away from ... they're baby servers, right? So they're an incremental and organic step away from what people were already doing with serverless and more broadly with managed services.

We were asking people to forget everything they've learned about the cloud and to some degree about backend software development and start all over again. And some of the things we screwed up. We were slow to even just adopt the word. And so we had what I now call the Voldemort problem. We had a thing we couldn't name. So people would say, "What do you use Lambda for?" And now we would say to build serverless applications, right. But at the time we were trying to say, "Well, to build applications that use events to do stuff which is simple, but it's cool." Hence the Voldemort problem. And so once we allowed ourselves to start using the word, and I'm not going to defend serverless as the best term, but at least it is a term. I'll tell you one of the hardest things to do as a business owner is sell something that you're not allowed to actually name. So I learned my lesson with that. Won't repeat that particular mistake in the future.

Jeremy: Right. Yeah, no, definitely. So I guess maybe a question I have for you too, and this is something now that you're away from AWS maybe you can answer this. I get what you said. You have to sort of focus on these enterprise customers, right? The enterprise customers are important. They're the ones who pay the bills. But that broader adoption, that sort of ground swell, right. The developers figuring out a better way to do something, that's why all these frameworks, that's why all these JavaScript frameworks become so popular because you get all these developers using them. I mean, is that something like with Lambda early on, were you really pushing that towards just enterprise customers? Or was that something where you thought this could be like a ground up approach?

Tim: Yeah, it's a great question. And I think one of the things that we at AWS at the time really dragged our feet on and to ill effect was coming up with some of these frameworks. And look, great kudos to the serverless framework guys and Austin and others there for even stepping in and doing that. There are tools now, I mean, Stackery has done a great job of making serverless I think consumable by the enterprise, something that we kind of miss. Look, AWS is fantastic at focusing on things like the availability, right? The nines of the service, latency, jitter. These key kind of golden, what people would call the golden metrics of a service.

It lives and breathes that one of the most important things you do in the course of a week at AWS is you go to the ops meeting. And the ops meeting is where you show your dirty laundry, you learn from your mistakes. You reveal your metrics to your colleagues, right? You hold yourself accountable and these are the things that you focus on. And all of that's amazing. But there's no equivalent of, there's no ease of use meeting every week at AWS. Right. And so the idea of helping developers be productive, of making things simple and making them consumable. When you started with things that were just infrastructure, that wasn't really necessary. But as AWS moved up the stack into these managed services, it's had to learn that that's actually a big piece of the equation.

Something that Microsoft has known for years, right? And sort of great job in actually helping developers not just know of something or keeping something up and running, but helping them actually figure out how to use it. And so that's a systemic learning for AWS as a whole. And I'll certainly say like in terms of being vocally self critical, I didn't get that right either at first. And so we waited way too long to do things like Sam. We didn't put enough wood behind some of those arrows. And so I think we kind of left the community to sort it out. And you still see the aftereffects of that. I still talk to people who say, "I don't know how to deploy it. It doesn't really fit into my CICD pipeline. It seems simple to run, but it's not simple to kind of build and test and operate in the same way that other things are." And so the fact that somebody could find Kubernetes easier to deploy than a Lambda is-

Jeremy: It's kind of scary.

Tim: It's unfortunate and a bit of an indictment that the tooling and especially some of the kind of the broader enterprise usage patterns weren't first and foremost in our thinking when we brought this to market originally.

Jeremy: Yeah, I totally agree. Because I mean, I think that's one of the things that has been the biggest complaint that I hear is just this lack of I guess coordination or organization where you can deploy it with the serverless framework, which is great. You can deploy with Stackery now, but back before it was like Cloud Formation. It was using Terraform. I mean, even when serverless framework came around, that made it a lot easier, but there are still people who write blog posts about how they write this custom deployment script that generates a cloud formation or something like that, or uploads them manually and triggers an API. I mean, and certainly those are all valid ways to do it. It's just seems like there wasn't a way that was put into place early on that would have been really helpful to build off of that, as opposed to like you said letting the community kind of figure it out on its own.

Tim: Yeah. And some of this was a learning curve for us at AWS too, right? In terms of understanding. Because if you thought of a Lambda as something that you hooked up to an S3 bucket, then maybe it didn't need a whole kind of development paradigm or CICD pipeline mechanism or application construction framework around it. And it quickly grew to be obviously so much more than that. And I think had we known how far and how fast it was going to go back in the day, we would have given more credence to the idea that we need our own ... we need a framework here and we need client side support for this.

We also had a little bit of that AWS-itis where you're like look, if the service is great, people will do whatever they need to do. And we didn't realize, we didn't think hard enough about the fact that, hey, if that's hard or even if there just isn't a simple way of doing it, it's going to actually make the service difficult to consume because it's not the least common denominator plugin piece of infrastructure here. It's something that is very, very different in that regard.

Jeremy: Alright. Missed opportunities aside, the past is the past. What we've gotten to now is an absolutely amazing ecosystem that allows people to build applications without thinking too much about the infrastructure. You still got to think about it a little bit, but for the most part, all of these amazing tools. So where are we now? Like where, you said what? It was 2014 when it was in preview, went live in 2015, right? So it's been over five years. Where are we now with serverless?

Tim: Yeah, I call it the ... I always say like we're in the terrible teens now. It's far enough along, it's no longer an infant. It's obviously become something that millions of developers are using, that the majority of the fortune 500 have some kind of serverless technology or solution in place from some cloud vendor or another. So obviously in that sense, it's been a remarkably successful introduction of a net new technology and paradigm. On the flip side of that, look, you can see the shape of the adult that it's going to become, but it's not adult in all ways yet. So I'll take an example here, a DTCC, fantastic example. So the US financial system has a lot of safeguards in place, and one of them is that it has to be possible if something happens on the East coast, if there's let's say a flood or something in New York that you can keep the stock exchange and other kinds of key financial capabilities up and running.

And so that means you have to have capacity that you know you can get to in another region. This is the kind of thing that's really tricky to do and the way this would have been done kind of formerly with servers is you just point to them. You're like, "Okay, look, we got a thousand servers. They're sitting in us West too, we're good to go. They're RIs or DIs or something on AWS or the equivalent on another cloud. But we couldn't really do that with the Lambdas, right? There was no way to say, "Well, these are your Lambdas, right?" Because Lambdas aren't capacity. And so things like the provision capacity feature now, it's not just a response to developer needs. It's also the kind of thing that makes enterprises able to deliver these regulatory compliant capabilities.

And so it's a great example of what I call kind of serverless growing up. All of a sudden, key economic and financial mechanisms that power, not just the US but the global economy can run on Lambda. And that's a huge step forward and very, very different from where we were back in 2014 when we released this as a kind of simple scripting mechanism to thumbnail images coming into S3. So really, really, really real game changers there.

Jeremy: Yeah. And I think that you're right about the enterprises showing up, right? Like finally you see more stories. I mean, it's just like Liberty Mutual and Lego. And I mean, just so many of these stories now that are fascinating of them like rapidly moving to only serverless or as much serverless as they possibly can. So the other thing though I think that's interesting, and this is something that was sparked. It was sort of like almost like an arms race or a space race where it's like who can develop serverless better or do more serverless things? And you got a lot of the big ones in there.

So you've got Microsoft obviously doing it, and you've got IBM taking over Apache OpenWhisk and you've got Google in there. But then you have all these other like sort of fringe edge providers, like the Cloudflares and the Fastlys. So this has created a whole new sort of ecosystem. So what are your thoughts on like how is that driving maybe the complexity or maybe the, I guess, the confusion or the adoption? I don't know what the right way to say that is, but it's like the wild West.

Tim: Yeah. Or the terrible teens. Right. I mean, it's tricky. I mean, I actually think a lot of those things are positive. It's one of things I've said before like Google for example is doing a really bang up job of thinking about the customer use cases. And if you kind of position them, the different cloud providers have taken very different tacks here. AWS, we asked the question, "If we just let go of everything that we've ever done or ever known in the cloud and made something completely new, how should it work? What would it do? How could it best serve developer needs?" Right. And that's an interesting question to go answer. I think Google has asked a very different question. They've said, "What is the kind of minimal risk, maximum insurance solution we could give people that adds some value over where they are today?"

And you see things like Cloud Run, which is a relatively narrow technology, but an insanely useful one. They've taken this really important use case of building a stateless front end and they've gone out there and nailed it. And in some ways it's thematic with what they've done with their app engine, with Anthos and others and Knative. They've said, "How do you get some additional value in the world that you're already in today?" And that world might be on-prem for example, or it might be an existing monolithic application or it might be a container. And so they haven't stepped as ... They haven't gone nearly as far or stepped nearly as aggressively as AWS has, but arguably they're giving a lot of people a lot of value, even though it's perhaps not as far away from that.

And then you get folks like Cloudflare who I think are, obviously they're building what they do best here, but it's amazing. They've taken this challenge of how do you do compute on the edge, even if it's a stripped down modest kind of compute. But making it almost ubiquitous so that literally kind of in line with every HTTP call. It's like if every HTTP call in the world could be scripted, what would that look like? And Cloudflare is doing a fantastic job of making that a reality. I'm envious. We wanted to that with Lambda@Edge and I would say we never quite got there with that product for all that it does some super useful things, but it's not that kind of ubiquitous inline everywhere on every edge cell that Cloudflare has created.

And I'm really impressed by what those folks have done. In fact, AWS if you're listening, think about this for your edge devices. You don't want to run a thousand EC2s in every Verizon pod sitting up on the street and on the street corner, what I want is something that works like Cloudflare's. So I think they're serving a useful role as challengers here and for customers who have that particular need, just like with the Google Cloud Run, I think it's a really nice product.

Jeremy: Yeah. So another thing I think that we're, or where we're at a point with Lambda and with serverless in general is we've got all these frameworks, we've got a lot of tools now. AWS has built a whole bunch of tools in that help with deployment and things like that, but you're still either using APIs or in many cases configuration files. So that undifferentiated heavy lifting, a lot of that's gone. I don't have to write my own queue anymore. I don't have to manage my own database, but I still got to connect those. And the only way to do that is with configuration and often YAML files. Right. So is that still a friction that you see hindering or slowing down innovation or is it something that you think that just needs to be abstracted away at some point?

Tim: Some of this is definitely a consequence. It's the square peg in a round hole kind of challenge, right? The irony of course is that serverless was supposed to make the cloud easier to use, but when you take an existing tool and you try to reapply that, sometimes it can actually make things harder. And a lot of the CICD mechanisms out there, things like CloudFormation was designed to take a bunch of servers, configure them in a particular way and put them in an environment that would allow them to run and get something done. That's a very different problem from saying, "I want to build an application out of fully managed components, and I want to wire it up, ensure that it has least privilege, make it a femoral so I can stand it up and tear it down again."

I'll just give a little anecdote. In building, so Vendia is not just serverless in its kind of runtime, it's also serverless in its CICD deployment. So what I've done is I've taken the AWS CDK. For those who haven't used that, think cloud formation, but like turn into Python or JavaScript node form. So you can programmatically construct things. And so I essentially wrapped a compiler around it, which makes it really easy for me to stand up our code base and create as many different test cases or production deployments in parallel as I need. It's great in that I was able to accomplish that. As one guy here, I could do a prototype of something that would normally have taken a team of 10 people to do, thanks to tools like the CDK. The flip side of that, like I had to go build a compiler around the CDK to make this really easy to use.

And even with all of that technology, this took a lot of work and a lot of energy. So I think there were folks here. I'll put a plug in for Stackery. I think they've done a fantastic job of helping people find a very different and much easier way of constructing a serverless application. And then starting with CloudFormation, even if CloudFormation kind of sits in the background of that, just as it does for the CDK and others. But treats it more like the assembly language of the cloud. And I think Reed Hastings said this best back in AWS reinvent in like 2011 or 2012, where we're at the assembly language level. I think like with some of these tools, we've come up to maybe the C level. We're not quite up to Python and any other language here yet, we're getting there.

Jeremy: Well, you mentioned CICD and it's funny because I have this running in my newsletter that I write every week that it essentially always calls out like it's another week, another custom CICD process for serverless. Because it seems like every time I see a CICD process, it is written differently. And I know AWS has now put into place, obviously they have code build and code pipeline. They've added more features to that. You have the amplified console, which will do CICD for you. Plus they do like these bootstrap templates now with Lambda applications where you can set those up. But that's the thing too, like just getting through that process, implementing CICD in serverless is just, it's not easy.

Tim: It's not. And look, I think we kind of went from let a thousand flowers bloom to the problem of tyranny of choice here, right? Where just as you say, just keeping up with the set of ... in the space of options can be problematic. And I know, when I was at Coinbase for example, we went back and forth on this. We ended up writing some of our own custom stuff to make this work. At Vendia, I essentially, having tried with some of my serverless networking pieces to use off the shelf stuff, I ended up doing my own custom CICD and I feel the pain because it's challenging. The flip side of this is if we get it right, it's also amazing. And this is the piece that I also want people to like have this takeaway here, right?

It's not just that whacking this into sort of old school tools is difficult, but part of the reason you see people experimenting is because of the incredible potential. The ability to run a thousand different tests. To literally stand up your production infrastructure, not a stage, not a strip down developer workflow, not something that runs in some kind of wacky emulation mode on my local machine. But to literally in real time run a thousand different tests in production, in a real production environment, in the cloud, at scale, with all of the capabilities that they will actually have in production and then just as easily tear them all down five minutes later is unbelievable. Right? You never do that with servers. It's way too expensive. It's way too hard. You'd have a team of 200 working on this for years, right? Nobody does that.

And that is the great opportunity of managed services. It's also the thing that is furthest away from what the existing tool sets are capable of giving us. And I think one of the reasons you see people experimenting here. It's not just that they want CICD and deployment and testing to match the simplicity of the underlying services, it's also that the incredible opportunity here isn't fully exposed by a CloudFormation or some of the other tools that are out there. I'd say the CDK makes it possible. You can now do things like put a for loop around your deployments, which lets you do incredible stuff, but only if you're willing to write the code for it. And so I think that's where we are with this. The kind of that golden future Nirvana where all of this is not just possible but easy, we haven't quite gotten to that yet.

Jeremy: Right. Yeah. And I love some of the next generation tools too that are like sort of working on this stuff. And I mean, I'd like to say they're getting it right, but I still am not quite sure what right is yet either. That's something which is part of the sort of the conversation. And actually speaking of that, I'd love to move on to this thing that I don't know if we've got it right or we've got it wrong. But that's this idea of state in Lambda functions or state in serverless or I guess any type of FaaS. So state versus stateless. I love stateless because I feel like it gives me a lot of control it. I know it's like pure functions with a functional programming language. Right. Just feel like you're not bound by the state or your applications don't get confused by that state. And so I really do like the stateless aspect of a Lambda function. On the other hand, lots of applications need state. So where are we with this?

Tim: Yeah. I mean, I think in some ways this is sort of the, I call this the great philosophical debate of our time. Right. And look, we talked about the origin story. The original concept was motivated by 12 factor app design and so forth, was separation of concerns. Let S3 be an amazing storage service, let Lambda be an amazing compute service, like let Dynamo be an awesome NoSQL database. And then there's this cool thing called the internet and network cables, right? Wire it all up and put it together and it'll be awesome. And I think in some cases that is exactly how it plays out. I mean, if you've got a relatively simple use case where you want to store something in S3, you trigger Lambda function, operate on that thing, maybe put it back again, rock on, few lines of code, almost hard to imagine how you make that thing much simpler.

I mean, CICD aside because we've discussed, right? I think we've squeezed out about as much of the cost complexity and so forth of that and we've transferred as much of the operational hassle back to AWS on that as it's probably possible to do. So I think that part's all great. Where this gets tricky though is as you say, lots of applications have state and even sometimes that's macro state, sometimes it's micro state. One of the big things that's tough about using Lambda with Kinesis for all the hard work that has gone into making that integration soar, but it's still the case that there's no state there, which means there's no affinity, which means you can't do something simple like just add up the values on a particular channel within that broader data stream because you never know which Lambda function is going to get it.

And so you end up doing things like copying it into Dynamo in the back out again or something, which is pretty strange. Or taking one persistence mechanism and then copying it to another mechanism, rendering it to disc, going back again. It's a huge waste of money, of time, of opportunity to do something that could obviously be done simpler. And there's a good place where state is meaningful, it matters. And we obviously haven't quite nailed the way that that gets put together. I also think it's just confusing. The other thing about state and serverless is, look, I had actually had this conversation with a developer and this guy came up to me and said, he's like, "I looked at Lambda. It looks simple. It looks really cool, but my code has variables and I store stuff in those variables. So I don't think I can use it because they say it's stateless."

So look, we can chuckle at that a little bit, but it is a good example of folks who have a hard time understanding what does stateless mean here. Because it's obviously it's got memory, it's got disc, right. It can hook up to things like Dynamo. Making that easy and approachable is something that I don't think we got quite right. And Sam has tried to make that easier. And obviously there's a lot of education out there on serverless design patterns. But one of the things that is really tricky is it's still not easy to hook up Redis to Lambda. The standard mechanism of storing, of durable but not persistent state. I mean, the thing that everybody uses to build their like every B2C application out there, right. And it is one of the hardest things to do with a Lambda.

And so I think that's a good example and Lambda is not specific here. Azure Functions, Google Functions all the same. So here's a good example where I think the challenge to the cloud service providers is make the practical kinds of state easy to do, whether that's Redis integration, conventional file system integration. Like I love S3, but sometimes you really just want to a Linux file system hooked up. And we still don't ... EFS from AWS for example, it's like you have an infinite disc drive and then in Lambda you have an infinite computer, but then it's like there's a wall in between them. The cable hasn't quite connected on the floor there. So I think there are some of those pieces were we to kind of get them together would give developers just a phenomenally better, easier, more tractable way to handle some of the problems that they have of writing practical applications. And we can call that state.

Jeremy: Right. Yeah. I mean, and I think one of the things that I love about serverless and the fact that it has been stateless, and you're right, stateless in this sort of context does not mean that there is no state at all. I mean, obviously you can use variables and all that stuff. And it's very simple. Like you said, call DynamoDB, rehydrate some object or whatever it is that you're working on. I mean, all of that is very possible in there. And I like the fact that that kind of forces you to think a different way about how you build your applications because it also helps when you start thinking about scale. Because I think a lot of people who build stateful applications are not thinking about scale.

So there are still going to be people that do that. And as much as I would love to see us just change the mindset to say anything that you build in a stateful manner you could probably build in a stateless manner, there are still going to be people who are going to want to do that stateful stuff. So does serverless, if it doesn't get there, if it doesn't add that statefulness, is that going to hinder it from becoming sort of the default paradigm for building applications?

Tim: I have these two strong reactions to that statement, right? One of them is I would say in some ways the most successful thing Lambda has done is to challenge thinking, right? To get people to say, do you really need a server stood up, turned on taking 20 minutes to fire up with a bazillion libraries on it and then you have to keep that thing alive and in perfect condition for its entire life cycle in order to get something done in terms of a practical enterprise application? And challenging that assumption is one of the most exciting, important and successful things that I think Lambda and other serverless offerings have accomplished in our industry. The flip side to this is to be useful, sometimes you have to be practical. And it's equally true that you can't walk up to an enterprise and say, "All right, step one, let's throw all your stuff away and then step two, you're not going to get past step one."

It's funny, we talk about greenfields, brownfields, it's all brown in the enterprise. Even if you write a net new Lambda function, it's running against existing storage, existing data, existing APIs, whatever that is. Nothing is ever completely de novo. And so I think to be successful and be as adopted as possible in the long run, serverless offerings are going to also have to be, they're going to have to be flexible. And I think you see this with things like provision capacity. I mean, when I was at Lambda still, we had long painful debates about is this the right thing to do? And for understandable reasons, because it is less stateless. It took the ... it's obviously optional. We don't force anyone to use it. But by doing it, it makes Lambda look more like a conventional, well, server container, conventional application approach because there is this piece that is a little bit stateful now.

And I think the arc here is for the serverless offerings to not lose their way, to find this kind of middle ground that is useful enough to the enterprises that still challenges assumptions that gets people to write stuff in a way that is better than what came before and doesn't pander completely to just make it feel like a server. But is also practical and helps enterprises get their job done instead of just telling them that ... because just sermonizing to them is also not the right way to do it.

Jeremy: Right. Yeah. And I also wonder too, I mean you mentioned Cloud Run earlier, which I think Cloud Run is an engineering marvel. I mean, it's probably not that complex, but it really is... I really like what they did there to make, to basically take something that wasn't serverless, a container and then give it those characteristics. And I feel like you have that bleeding back and forth between those. And obviously you've got Fargate with AWS and they call it serverless containers and that, I don't know how I feel about that. Right. Because does it put you ... If we blur the lines too much, does that lose? I mean, do we redefine what serverless is? Does that even matter? I mean, what are your thoughts on that?

Tim: Well, look, and I say this with love because I want things like Lambda and serverless, sort of fully serverless apps to be super successful. But if all that we accomplished was to help the cloud providers make things like containers have fewer infrastructure artifacts, fewer things to have to set up and configure, less painful maintenance and deployment and security overhead so people could get the jobs done faster, that would still have been a success. And I think to the extent that we can also sort of challenge the dominant paradigm as it were and get developers to build applications that are easy, fast and fun, all the better. So I think it's all good.

I think it's also the case that some of these are interim states and some of these are end states. And one of the ways I've talked to people about this in the past is when you write an application as maybe you write it on a server and you think like okay, well, at some point I'm going to containerize that. And then you containerize it at some point you think, you know I should really be running this on whatever, Google Cloud Run or perhaps Fargate for my application and my cloud choice because I'm doing too much low level. Like I'm still responsible for keeping that underlying piece of infrastructure alive and running, and that's just kind of a waste of my time and energy. AWS or Google or Microsoft can do that better than I can.

So there's a sense of progression. But once you've built something out of a set of managed services, you're done. That's the sort of the end of that state machine, right? It's kind of the final, final. It's the end game. And this is something that I think is going to take us a while to get to. I mean, we will have as an industry, you always iterate organically. That's why, I mean, not everybody's even in the cloud yet today. So you can see those trend lines and you can see where that's going. And I think realistically as providers, the CSPs are going to have to do both.

They're going to have to provide people the organic incremental steps that help them do something a little better than they did yesterday. And they have to work on this thing, which is: what is the ultimate end game if you could get everything to be the way that you wanted? And those two things are going to run in parallel, which is why like it's never an either or. It's not like one's going to win, one's going to lose in the same way that VMs are still around and will be for certainly for our lifetimes.

Jeremy: So you mentioned this terrible teens idea and obviously innovation in serverless is not done. There is a lot more that we can do. We have to figure out this blurred line, we have to figure out maybe networking and some of these things. So you have a blog post that you put together a couple of weeks ago. So which I thought was great. It was like a 2020 re:Invent wishlist for AWS Lambda or for serverless. And we've talked about a couple of these things, but there's a whole bunch of them and I'll put the link in the show notes because I encourage people to go and look at this if only to figure out what is missing. Because I think a lot of people don't even know what's missing until they read your post and be they'll be like, "Wow, I didn't know you couldn't do that or that it wasn't part of it."

But I'd like to go through a couple of these because I think maybe the more interesting ones, at least interesting to me, because I think that these sort of strike a chord at least with me in terms of things that you know you need or we know we need in order for, like you said earlier, like this to become just the way we build applications. So the first one was this idea of doing one millisecond duration granularity. What's that about?

Tim: Yeah. So look, to understand this one you have to also remember the context too. When we were putting Lambda together, so say 2012, 2013, you couldn't have an EC2 instance on for less than an hour. And most companies, most of the time were still provisioning on-prem servers that they had to go buy and stand up. And it was usually a month lead time to get new hardware into those racked and stacked and stood up. And so the idea of 200 milliseconds, 100 milliseconds, I mean that just that blew people's minds. It was just astonishing. In a world where you can run an EC2 instance of or a Fargate instance or something else for as little as a minute of duration, however, and where people are using languages like go that might do something useful in just a handful of single digit milliseconds.

And frankly also as hardware continues, I mean, the whole sort of efficiency curve there has slowed down a lot. But still even from Lambda's incarnation to where it is today, the hardware has gotten much faster that it runs on, networks have gotten faster. And so the idea of running maybe an application that takes two or three milliseconds per call, and then spending a hundred milliseconds starts to look like exactly the sort of waste that Lambda was designed to get rid of. One of the most successful things about Lambda and serverless in general is that it collapses this cost structure. Companies, enterprises famously, analysts will tell you maybe 10% utilization. So they radically overspend on the amount of hardware capacity that they need and serverless helps them get rid of that.

But there are still these other forms of waste in the system and this is a good one, right? Where it just is impossible for something is very, very fast, and it opens up a new set of ... If you can do that, it also opens up a new set of applications. Because if you have something especially that fits on a front end that maybe takes two or three milliseconds to run, you're probably not going to do that on a Lambda today. With the improvements that the Lambda team has made with the speeding up the latency, running on the new firecracker architecture, it is also possible now to write these very low latency applications. And I think part and parcel of that is being built fairly for those low-latency applications.

Jeremy: Right. And I totally agree, because the last time I was in Seattle and I was talking to one of the Lambda PMs. Like, what would you like to see? I said, "Lower units of billing for it because a hundred milliseconds is just crazy." Now, when you think about it, it never used to be like you said, but now it just seems crazy. I mean, even if you did, it had to be at least 50 milliseconds at least that, and then it was per millisecond after that or something like that. I mean, even that would be better than what it currently is. Again, a hundred milliseconds is still not very much time. So it's still pretty amazing, but I'm totally with you on that. The other one, this is funny because I know you have some history with this, is EFS integration with Lambda because Fargate has it now.

Tim: Yeah. Look, and I say this like I'm not part of the team anymore. I have no special insight into this, but I can certainly speak to the need for it. One of the things that's really challenging is especially in a world where people want to try to do things like ML training or running kind of some of those ML outcomes off larger data sets, where they want to be able to start to use, think of some of these applications where you do want to use lots of Lambdas to process data in parallel. So you might want to work on a large dataset. As we start to expose Lambda to data scientists and there's a whole other conversation about how you do that. Eric Jonas has written this fantastic paper, this computing for the 99% all about how that should be a focus and a priority.

As you start to move in that direction though, you've got to make it really easy for the compute to line up with these large data sets. And one of the challenges with the S3 model is that you've got to pull them all into memory on data and then kind of shove them all, or into Lambda and shove them all back out again. And so this is where I think if we can get a ... it's going to be the ultimate compute meets the ultimate storage here, right? Like you put it together and you get, it is the mainframe of our day. And I don't mean that as a slur. I mean, that as the highest possible compliment. It is the processing engine that emulates what a supercomputer can do, but only if you can plug these pieces together. So I'm really excited for this. I think it opens up a lot of applications and frankly, it makes some of the stuff that's just really hard to do today like managing that small amount of slash temp space effectively. Those problems, if not go away, at least they get a whole lot easier.

Jeremy: Yes, definitely. Yeah, I think the use cases with that are, it just opens up a whole new world of things that you can do. Because I know I've worked with large files and with S3, you're always streaming data. And then you can't, like you said, it's hard to ... just that little bit of state would be nice in those sorts of computing situations. So another one, and this is funny because this is again personal to me. I give a lot of talks. And oftentimes when I give talks, one of my new ones is how to fail with serverless, which is all about the failure modes in the cloud. And one of those things that I find that I end up introducing this term to people. And this is not a term, I think if you're a computer scientist or maybe a more traditional background, you would know this term.

But I think there's a lot of people, especially front end developers, people getting into serverless that aren't quite as technical on that level and nothing against their level of technicality, but is this word idempotency. A lot of people, I don't think know what that means and also don't realize the impact of it when you are using events in serverless applications, because one of the things on your list was this idea of idempotency protection. So I'd love it if you could explain what your solution to that is.

Tim: Yeah. This is a ... Look, if there is an unfortunate garden path in Lambda, this is probably it. Because it gives you the illusion and it is mostly true and the mostly of course is the devastating part of this. It's mostly true that when you call the Lambda, it runs exactly once. And so take something simple, like you're going to go off and let's say using this Lambda to implement a bank account, right? So you're going to go off and you're going to compute and add interest to somebody's account, or you're going to go do a transfer with that Lambda function. So you call it once, it runs once, it does its thing once, everything looks good. The problem is the actual semantics of Lambda are at least once, which means every once in a while, not all that often, but not zero times either, it'll run more than once.

So maybe it'll run two times or even three, which means you'll end up adding more money to that bank account than, you expected. And while not everybody's using Lambda to manage a bank account, it turns out that just lots of things in the enterprise are important to do exactly once. Not maybe a couple of times, not occasionally a couple of times, but exactly once. And this is a good example where I would say by focusing a whole lot on the service metrics and dynamics, it's possible to sort of lose sight of some of the things that developers really need. And this was sort of ... this was something I feel like certainly I got wrong in not giving people a simple solution for this earlier on.

Because as someone who has run a team at Coinbase trying to adopt serverless, someone who has been trying to build a business around it, I can tell you that some of these things not being there are incredibly challenging for adoption purposes. And one of the most common of them is just you want your code to run once. It seems like a really simple ask and it turns out that building that practice is a bit of a mess, and you've got to stand up. In addition to your Lambda function, you need a full powered step function with a couple of nodes in it that you can go and run a task, and then it'll give you the exactly once. And it's not that it's impossible, it's that it's expensive, clunky and time consuming. And you require every one of the 10 million developers you want to go use this to go figure it out on their own. And that's really painful, right?

It's something that should be as simple as like go and check a box here. So this is one of my big asks for AWS to think about is yes, it's possible to solve. No, it's not easy. Please, please, please, please give us this check box that says make it run the way I've kind of always wanted it to run. And we've all heard about the cap theorem. We know that's not easy to do. It's okay if it costs a little more, but it's something that would really make Lambda just dramatically simpler to use for virtually all of us who try to get something done with it on a day to day basis.

Jeremy: Right. Yeah. And I think that's one of those things. There are some of those patterns where the developer is left to figure them out on their own. Right. So you're putting in something to try to battle that idempotent operation, and it can be quite a headache. So we talked about serverless Redis a bit, and I'm sure we could talk even more about that. But the last thing that I wanted to talk about on your list was serverless networking. And so you've done a lot of work on this and like a lot of work on this, which is pretty cool in terms of being able to have running Lambda functions communicate with one another. And this is something, again, I feel like we could talk about this more, but the ability for those running functions to communicate with one another, that is what gets us to this idea of the Lambda supercomputer, right?

Tim: Yeah, so a little bit of context on this. Some of the most exciting research and innovation that's happening in the space right now is happening in academia. And you've seen these, we touched on this earlier for the video transcoding work that's gone on, on top of Lambda. There are some researchers here, Eric Jonas, Sadjad Fouladi, Johann Schlier-Smith, Vikram Sreekanti, who are doing just incredible, insightful work into building just massively scaled systems that are often combining the state and the compute together and doing really interesting, massive data parallel applications on top of a serverless architecture. So that's the good news. The bad news, it's a struggle. It's really hard. And if you ask yourself this question, like could you go rebuild some of the big infrastructure solutions of today like MongoDB? Could you go recreate MongoDB or Aurora DB on top of Lambda? And the answer is probably not.

And so what is making this hard for researchers? What makes this hard for somebody who might want to construct an infrastructure style service on top of a serverless base? And the answer is complicated as it always is, but in some of these, it's a few missing pieces, right? It's the fact that you can't do cross calls with Lambda. So that's the serverless networking piece of this, right? If I've got two Lambda functions, I can use them to call out to other services, but services can't call into them. And we talked a little bit about the origin stories earlier. One of the reasons we did this in Lambda was to keep people from using it as a conventional web server, because that was a failure pattern, right?

We would trick them into thinking that that state was there when it wasn't. So we turned that off. But in turning it off, we made it impossible to build some of the ... use some of the standard techniques that you use to build high speed, data dense multi computer parallel applications, right? And so all these things data scientists want to come and do now get way harder. So serverless networking was an attempt to solve some of that by doing NAT punching and some of these other techniques down at the low level, so that Lambdas could actually communicate and take advantage of the high bandwidth network that sits between them. And that's just one of the several things that you need to do.

To really make this serverless supercomputer a reality, you need not just a distributed networking solution, you need low latency as in single millisecond style choreography. You need high-speed key value stores like we talked about with the serverless Redis. You need a way to hook this up to immutable inputs and outputs. So there's a whole set of things that you have to do there so that you can ultimately build something like say CRISPR or a MongoDB on top of this, with the kind of outcome that you could get if you were to grab a bunch of servers. And doing this unlocks a whole new set of people and a whole new set of applications that can start running serverlessly. So I remain incredibly excited about this. I think we've seen enough evidence in the research community to suggest it's all possible.

I think we've seen enough evidence and direction from the cloud providers to suggest it is all doable, but there's still a long road ahead to make all of that possible. So I wanted to help out with that, hence some of the open source stuff that I've done with the serverless networking piece. But really that is just one of some of these foundational elements and we really kind of need to get them all in place to make this happen.

Jeremy: Right. So is that something you're going to keep working on or is it something where you think like, you've proven this is a valuable or viable thing and now the cloud providers just need to go and run with it?

Tim: No. I mean, it's probably a yes and a yes, right? Like I continue having some great conversations with folks who are working on this in the research community, continue sort of working as just kind of time and energy permits in some of the open source parts there, and look forward to some exciting collaborations with others on thinking through some of these challenging problems, like the choreography and the key value stores. So with Vendia, I've chosen to put my energy into a commercial enterprise that's helping to solve a slightly different set of problems. I think this space is going to be one where honestly, and my great hope here is that this is also part of where open source works for serverless. Obviously open source is not going to mean that we pull Lambda out of AWS or we pull Azure functions out of Azure.

It's going to be that people can create these frameworks and these mechanisms that help them get incredible new things done in the cloud. And that's where I think you can see the university research and the research community coming together with the cloud providers, coming together with this growing ecosystem around serverless to produce something that is amazing. So I think that's probably the best role for me, and that is not to try to be the commercializer of those pieces, but to help be a human choreographer of some of that work and energy.

Jeremy: Yeah. And I mean, speaking of open source too, I mean, think about Kubernetes, right? Kubernetes has taken the world by storm because obviously containers are the, I guess the standard that most people are now considering to be cloud native, even though there's plenty of people doing stuff with serverless. So is that something where you see something like Kubernetes ... I mean, obviously it's going to be around for a while because so many people have started to adopt it. But is that something where you think that's going to continue to gain steam or are we going to see serverless and maybe some more open source serverless options kind of take over for that?

Tim: Yeah. This is one of the things that Kubernetes got right. And I would say like everything's a mix, right? It's complicated in a lot of ways, but it's also open and portable, which is a key requirement for a lot of enterprise use cases. And I think that is part of the direction that serverless needs to move in. Because one of the key buying objections, anytime I would talk to a customer, they'd always be like, "Wow, this is just, I'm like a kid in a candy shop. I love all of this. On the other hand, I'm afraid I'm going to get cavities in the form of vendor lock in here. So help me out with that. What do I do about it?"

And one of the things that has helped give Kubernetes momentum is the fact that that question has an obvious answer in a way that today Google Cloud Run and a Lambda and an Azure function don't have an equally simple answer. So stay tuned for more from Vendia, perhaps on some of those topics. But I also think this is a place where the open source community has to come together and think about what's the right way to make this work. It's not going to be trying to run to ... It's not going to be trying to emulate the services like Lambda on a bunch of individual machines. Right. And you can see some of the challenges of doing it.

I can tell you for example having been at Coinbase and watched distributed ledgers go that one of the big difficulties for them is that everybody's running this stuff on kind of stock hardware. It doesn't use the best and brightest of the cloud, and it's the least common denominator. You wonder why Ethereum is slow? Well, that's because you can run it on a laptop. Imagine running Amazon S3, literally S3, like for everybody on your laptop. And that's kind of why Ethereum is running at the pace it is. So there's a lot to do there. I think it's not going to look like ... Open source solutions for serverless will not look like the Kubernetes model, but it is still a missing piece. And I think if we could get there, we'd also create a collaboration forum that people could lock and latch onto in a way that is never going to be quite as well developed if it has to be a single cloud provider running the show.

Jeremy: Right. Yeah. I totally agree. All right. So speaking of Vendia, I know you can't tell us a ton because you're still in stealth mode-

Tim: A few more weeks.

Jeremy: But you said you built it all serverlessly and you mentioned a little bit about using the CDK and some of that stuff, but any success stories of building it serverlessly?

Tim: Well, proudly serverless, one of the nice things about that is you can do a lot incredibly quickly, right? And here's a good story of developer productivity and progress because like you are looking at the moment and we're hiring by the way. But at the moment you are looking at the developer team for Vendia. So there are like a dozen people behind me furiously typing away, right? This is me and me and my spare time on nights and weekends primarily. But think about just pick one thing here like regional build-outs. So you scroll back to when I started at AWS 2012, building a new region for let's say a service like S3, six months on a good ... if you're lucky, right?

Because hundreds of people, everything from surveyors and electricians, hundreds of vendors, supply chain in the thousands. You've got to go get this thing stood up, filled with servers, filled with racks. Fast forward to Lambda. So now circa 2016 let's say. So build a new region in six weeks with a dozen engineers who are able to use things like EC2 and take advantage of the cloud. Fast forward to Vendia. I launched not just one region, but regions all over the world, a large subset of them, all the ones in which the services I needed were available in about six minutes, because all I had to do was list the names, type the names into the CDK and wrap a for loop around it and I was done.

So you go like one guy, six minutes, launching a production service at scale worldwide with essentially three lines of code. Now that's an amazing, amazing example of the kind of productivity success that you can get out of serverless and a well-matched set of tools like the AWS CDK. And I think that's kind of the story here. It's getting rid of undifferentiated heavy lifting, but it's also this idea of capital efficient value creation, which is really what we're all about.

Jeremy: It's amazing. Well listen Tim, thank you so much for one, speaking to me and taking the time today, but also for serverless. I mean, this is my livelihood. This is the livelihood of a lot of people that I know. What you and your team did at AWS in those early days was just absolutely incredible. And it's just sparked this thing that I think has completely changed the way people build applications. I know it's changed the way I do. And I look forward to everything that happens after this, including all this new stuff you're coming out with, with Vendia and that sort of stuff. So again, thank you. If people want to find out more about you, more about Vendia, all the stuff you're working on, how do they do that?

Tim: So we've got a website stood up. It's a coming soon website at vendia.net. Tune in. We come out of stealth mode on June 26th. I'll be doing the keynote at the AWS serverless community day for Australia and New Zealand on that time. And that's also when we'll stand up more of our ... kind of take the wrappers off as it were and tell the world what we're all about here. So can't wait to tell that story and have a broader conversation about it.

Jeremy: Awesome. And you are a course on Twitter, Tim Allen Wagner. Your blog on Medium @Tim A Wagner and of course LinkedIn and all that stuff. So we will put all that into the show notes. Thanks again, Tim.

Tim: My pleasure. Thanks so much for having me, Jeremy.

This episode is sponsored by Dynobase and Datadog.

View Details

About Adrian Hornsby:

Adrian Hornsby is a Technical Evangelist working with AWS and passionate about everything cloud. Adrian has more than 15 years of experience in the IT industry, having worked as a software and system engineer, backend, web and mobile developer and part of DevOps teams where his focus has been on cloud infrastructure and site reliability, writing application software, deploying servers and managing large scale architectures. Today, Adrian tends to get super excited by AI and IoT, and especially in the convergence of both technologies.

  • Twitter: twitter.com/adhorn
  • Medium: medium.com/@adhorn
  • Dev.to: dev.to/adhorn

Watch this episode on YouTube: https://youtu.be/6o2owe2VHMo

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm speaking with Adrian Hornsby. Hey, Adrian, thanks for joining me.

Adrian: Hey, Jeremy, how are you?

Jeremy: So you are a principal developer advocate for architecture at AWS. So why don't you tell the listeners a little bit about your background and what it is you do at AWS?

Adrian: Okay, cool. So first of all, thanks for having me on your show. I'm a huge fan of your show. As for my background, it's a mix of industry and research. Actually, I started my career at the university doing some research, and then moved to Nokia research and eventually some startups, always around distributed systems and real time networks and things like this. And then let's say the particular things is much of the work that I've done was always on AWS since the very beginning. So it kind of felt very natural eventually to join AWS, which was about four years and few months ago. And I joined as a solutions architect, and then quickly moved into an evangelist role. And mostly doing architectures and resiliency and a lot of breaking things kind of chaos engineering type of things.

Jeremy: Awesome. Well, speaking of resilient architectures, that's what I wanted to speak with you about today, because you have on your Medium blog, which is awesome, by the way. I mean, I go there-

Adrian: Thank you.

Jeremy: Every time I go, and I read something there, you think you know it all, and then you read something by Adrian, and you learn something new, which is absolutely amazing. But so I want to talk to you about this, because this is something I think that ties into serverless pretty well, is this idea that I think we take for granted, especially as serverless developers, we take for granted that there is a bunch of things happening for us behind the scenes.

And so we get a lot of this, infrastructure management out of the box, we get, some failover out of the box, we get some of these things. But that really only scratches the surface. And there's so much further we can go to build truly resilient applications. And you have an excellent series on your blog called the resilient architecture collection. And I'd love to go through these because I think that this is the kind of thing where if you start thinking about global distribution, you start thinking about latency. You and I have been having a lot of latency issues trying to record this episode, because you're all the way in Helsinki and I'm over in the United States. These are things to start thinking about.

So I want to jump in first with this idea of embracing failure at scale. And I love this idea because when we build small systems, we think about reliability, right? We try to get as many nines as we possibly can. But when you get to the level of global distribution, distributed systems that are sending messages between components, that are sending messages across the Atlantic Ocean or the Pacific Ocean, this data is going all over the place, this idea of failure, or at least partial failure has become the new normal.

Adrian: Yeah. So yeah, I think it's things have changed a lot in the last few years. I mean, before you were on the monolith application, and you were trying to make sure your monolithic application was always up and running, right? I think there was even some competition into uptimes it was very popular back then to look at uptimes of servers and say, "My server's been up for 16 years, wow, awesome." But now, we've moved away slowly from monoliths to micro service architecture, and especially I think as we move even to the cloud, and we use more third party services, systems become naturally more distributed, and they go over the internet, which is everything but a reliable source of communication.

So, you have network latency, you have network failures. So there's a lot more things that can go wrong. And I think understanding and accepting that anything, at any time can fail is actually a very important thing. Because it means that you accept failure as a first class citizen for your application. And then you need to write code and design applications so that at any moment in time, there can be failures, and that's called partial failure mode, as you said. And it's very different concept than what it used to be back in the day and that means that you need to design your application with different characteristics and different behavior.

Jeremy: Right. And so if you're designing your system with these different characteristics, and you're, you're forward thinking to this idea of resiliency, and again, you have a whole bunch of stuff that you do on chaos engineering as well, which is this idea of injecting failure into the system to see what happens when something breaks. But that is quite an investment, not only an investment in learning, right? You have to learn all these different parts of the cloud, and all these other failover systems and what's available from that standpoint, but also an investment in terms of building your application out that way.

So you mentioned in the article, this idea of the investment of building in this resiliency versus what that lost revenue might be if something fails. So if your billing service goes down, or your payment service goes down, and you can't charge credit cards anymore, if that's just the end of it, right? Like you just say, "Hey, we can't charge billing or we can't charge your credit card, so our site's down." Versus building something that says, "Well, we can't charge your credit card right now, but we can take your credit card number and we can calculate the order total and those sorts of things." So what is that trade off that companies should be looking for, in terms of, as you put it, lost revenue versus the investment in building these resilient architectures?

Adrian: Yeah, it's a very good question. I think first and foremost, it's always start from the business side. It's like understanding what are the requirements in terms of availability because as many nines of availability you want, a matter of fact, the more work you're going to have to put, and the more resources you're going to have to use and that resource is money, right?

And especially I think the work around availability and reliability is not really linear. At the beginning, it's you have a lot of gain with small work, but as more nines you want is actually I would say the investment, versus the investment you have to do to gain more nines becomes a lot bigger as you have more nines, right?

So it's increasingly hard to reach more nines. So, you have to really think what is it that you want to achieve as a business? And I always tell customers to start from a customer point of view as well. Like, what kind of experience do we want? And as you said, maybe the ultimate experience is a fully working site, but what are the possibility for you to maybe degrade an experience when you have an outage? And still being able to deliver service. I always take the example of move a website into a read only mode, whether it's Netflix or Prime Video or even Amazon, when something doesn't work, what kind of features can you still provide to your customer without having them giving a blank screen, and say, "Oh, sorry, our database doesn't work. Therefore, you cannot use anything on our website." I think there's tons of things that you can do in between. And it's all things that you have to take into consideration.

And then of course, it's like, where do you invest it? A lot of people start with the infrastructures, but you have to realize that actually the resiliency is not only infrastructure, it goes from the infrastructure, of course, but it goes to the network, the application, and also people. We've talked about people resiliency for some times and it's also very important.

Jeremy: Yeah, no, and I think that's interesting, too, about this idea of redundancy in there as well. Because obviously, redundancy is still a big part of it. I just, even if we build a system that says, "Hey, if the credit card system goes down, we can still accept credit cards." Really, what you'd like to be able to say is, "Well, the credit card system goes down in this region, and we can maybe failover to this region and still provide that service." And that degradation might be a latency increase, for example, right? And so I really love that idea of duplicating these components.

Obviously, there is a lot that goes into that when you think about duplicating components. You have databases that need to be replicated, you've got all kinds of other things that become-

Adrian: It's more complex.

Jeremy: Yeah, exactly it gets to be more... And it goes back to the investment and time, right? Like, what is the investment, is that something we want to be able to do is provide four nines or six nines of uptime, or whatever it is.

You actually outlined in this post the formula for this, and I don't want to get overly technical around this. But essentially, just to sum it up, if you want to get four nines, you need to have three separate instances or three separate components running, I guess, or regions running in order to get that. And so, is that something though, that some of those multiple nines are built into existing AWS services?

Adrian: So, well, you're touching a very big thing. I think the formula says simply that if you have one component, and this component is, for example, a billable 99% of the time, right? Which is not really good, because it means that you're accepting about three days of downtime per year, which is pretty, a lot.

So let's say you have an instance running somewhere. The simple fact that you actually duplicate that instance, increases its availability to almost to four nines, right? So you go from two nines to four nines. So that gives you 52 minutes of downtime, and then you do this another time, it gives you six nines, which is 31 seconds. So I mean, this is not a new formula. It's been used in electric components, in many industries, in nuclear industries. There is sometimes like six levels of redundancy to make sure that the electricity always powers the plants and all this kind of stuff.

Now, and as you pointed out, there's also a problem with redundancy because it adds more complexity, right? So there's a trade off between how many nines you want, and how much complexity you're willing to accept. So the key there is automation, right? And, of course, if you think about AWS managed services, on AWS already have this idea that they are using three availability zones under the hood, so exactly it's duplications of redundancy, to provide this kind of service so people don't have to use it.

So it's some service provides it, some not. There's regional, there zonal service. I think the most important is to understand it a little bit. So even if you use managed service, I think being curious a little bit, how things are built under the hood gives you a very good idea of your levels of availability or possible availability because of course AWS just provides infrastructure availability, right? It doesn't provide your application.

So even if you use three AZs under the hood, but your application doesn't use it at the full extent of the capability, you won't have four nines of availability, right? So you have to go through the entire stack from, you have to use the infrastructure, you have to use your network, you have to use the application, then of course, people. Because if you deploy an application and that your application is across three AZs, but when you deploy it, you break it, well, it's like, you lose the benefits, right?

So it really is a synchronizations of the entire layer, and what you want to do. That's why it's complicated. And I think that's why it's also important to understand how things work.

Jeremy: Yeah. No, I mean, I think this idea to have repeatability of deployment, this is the infrastructure as code idea. And the other thing you go into, and you have a bunch on your blog about this as well, is this idea of immutable infrastructure, which is just, again, an entire podcast in and of itself, because it's a whole other probably deep thing that we could go down. But I think the basic idea behind that is just this thought that rather than me trying to update in place, which you always have problems, and of course, with EC2 instances it's an even bigger problem. But certainly with serverless applications, if you're just switching out a Lambda function, or something like that, but can you explain just quickly this idea of immutable infrastructure and how it relates to serverless?

Adrian: Right. I mean, yeah, for immutable infrastructures, or immutability is a problem in computer science in general, but not only cloud, but also programming languages. If you look at Python versus SQL, for example and how will you assign variables, and how the state can be shared between variables, it gives headaches to developers every day. So the idea of immutable infrastructure is very similar to that, is you have a variable, you have a state in the cloud, you have an infrastructure running, why do you want to change it? Don't change it, keep it there. If you want to modify it, is deploy something next to it, parallel to it, like a duplication of it with the new version and then slowly move traffic to that new version. That gives you two things. That gives you that hey, you have a working version here that you protect, so if anything goes wrong during your deployments. Sometimes your deployment might work but after an hour of traffic, cache warms up and all of a sudden you have issues.

Well, you can very fast rollback, you just need to move the routing back to the existing infrastructure instead of having to redeploy your old application, having to redo something. When you have an outage, I think the most important is to react without reacting, right? So that the idea that you have something there that is safe to go back to, and that is protected is actually very, very nice. And this is what we call immutable infrastructure. And there's many ways to do that, whether it's a Canary deployment, AB testing, or Blue-Green. I prefer Canary, because it's a progressive rollout. But there's many ways to achieve this kind of things.

Jeremy: Right, yeah. And those Canary deployments are built into API gateway if you're using Lambda functions, for example. But you still have other components that aren't Lambda functions, right? So let's say that you deploy a new version of, I don't know, an SQS queue or maybe a DynamoDB table or something like that. I mean, that's also where things get a little bit hairy, right? Where you start sharing things that are data related.

Adrian: Right, yeah. And I mean, and this is very, very true. I think when you make a deployment it's very important to understand what is your deployment going to affect? If it affects database, definitely, you're going to have to do something else, you cannot exclude the database from your Canary. So you might not do a Canary deployment, you might be doing, say for deployment something with progressive rollouts within your existing infrastructure, because you need to do a schema update. But I think it's not all white or black. I think there's, if most of the time you do deployments, you do not do database schema changes. It's definitely important to try to make those as safe as possible, right? And it's not only serverless, it's also any other infrastructures and on AWS, you can do this from many different ways. Whether it's Route 53 with weighted Round-robin, an ALB supports the weight for target groups, an API gateway support stages with Canary. And actually even a Lambda function supports Canary with aliases weights, right?

So I think the most important is to understand that it's not all black and white. And sometimes you might want to do a deployment that is more problematic, but the important is to limit the number of those, right? So, at least that's why I feel like this. Sometimes you can do a mutation, sometimes not, but I think most of the time, you should be doing it.

Jeremy: Right. Alright, so let's move on to the next one which was avoiding cascading failures, right? And so this is another thing where from a small level, it's really not that big of a deal. It's like, "Oh, the queue backed up, and then so it maybe is more aggressive in trying to call some other third party API. And maybe that gets overwhelmed. And we could put some circuit breakers in there, we could do some of these other things." But there is a lot that can happen at scale, right?

Like if you have maybe 100 queue messages per second, or 100 queue messages per minute versus 10,000 queue messages per second, when those start backing up, and those start retrying, and then some of those get through to the next component, and then that component backs up and starts retrying. I mean, you just have this very, very vicious cycle that can happen. And a lot of that is not built in for you.

Adrian: No, it's a very good point you mentioned. I think the deadly thing, the deadly part of the architecture in distributed systems is that you have many layers. And very often each of those layers have their own timeout and retry policies. And most of the time, let's say the default timeouts are absurd. They are... In Python, the request library, for example, is infinite, the default timeout. So that means if your third party doesn't answer, it will keep the connection open indefinitely, right?

Jeremy: Right.

Adrian: So that means if your client has a different time out, very often they do because client libraries is very often within five to 15 seconds. So that means that your client will have a retry policy and eventually will retry to get the data. And so that means all of a sudden you exhaust the number of connections from the back end side, your connection pool is running out of free connections, and that means that this server is unreachable. And what does the client do? He does a retry to the other servers, and eventually you have this cascading failures because all the clients are going to be retrying to all the servers one after another, and eventually that runs out.

So it's very important in distributed systems to really understand first the timeouts and set them, not, you know, when you do an NPM install or a Pip install, you are installing a lot of libraries from other people that maybe didn't think about your particular use case. And very often, I see teams not looking at those timeouts or they use system default and what is system default? Well, it's...

Jeremy: Don't even know.

Adrian: No one knows. So I mean, I think it's very much related to operational excellence in a way that a, how is my application behaving? What are the defaults? What are the retry policies? Are you going to retry hundred times? It makes no sense, you know?

Jeremy: Right.

Adrian: So you might want to retry once, twice, maybe three times. But that's, don't create more problem if your system is experiencing issue, I think failing fast is very important, especially in distributed systems. And if you really have to retry, don't retry aggressively, maybe retry with an exponential back off, and especially give the system some time to recover.

So, the problem a lot of the time the library is retry maybe even few times per second. It's like, "Oh, you didn't get a request, let's retry, let's retry." It's like what the kids are doing in the car, dad, are we there yet? Are we there yet? Are we there yet? It's very annoying for the drivers, for the backend. So it's the same in distributed system. What you want is either do you do a pub-sub, you say, "Okay, let me know when you get the data," or you asked it like maybe you make a retry, and then the next time you ask after 10 seconds or 20 seconds, and then the longer you wait, the longer the interval between the retries. And that's what's called exponential back off.

Jeremy: Yeah, well, the other thing with exponential back off too, is it can be tough if you have like 1000 requests to try to go through and they all fail, and then they all retry again in one second, and then they all retry in two seconds, and then four seconds, and then 16. So they keep doing the exponential thing. That's why this idea of using something like Jitter, which just randomizes when that request or the retry is going to be is a pretty cool thing as well.

Adrian: Yeah, exactly. And it's important, especially in distributed system, because you don't want to have all your distributed system to retry at the same time exponentially as well, you explained it very well.

Jeremy: So the other thing, you mentioned idempotency earlier, we talked about immutable infrastructure and that sort of stuff. Item potency is another main issue when it comes to retries. And I talk about this all the time, because essentially, if you retry the same operation, just because it looks like it didn't complete, doesn't mean it didn't complete, right?

Adrian: Right. Exactly.

Jeremy: Because there's also a response that can fail, not just the request itself. So, that's one of those things where I think if people are unfamiliar with item potency, just the idea that you can retry the same thing over and over and over again, anytime you start seeing these failures, you need to be able, from a resiliency standpoint, if we go back to resiliency standpoint, we can buffer those and we can retry, and we talked about that, but what are other ways that we can deal with these failures in a way that we can respond back to our customer to let them know what happening?

Adrian: Well, one is degradation, and you can degrade with two things stale data, right? So maybe you can't access the database, but hey, what was the last known version of or version of that data and maybe serve that, for example using cache, right? Cache is a good way to serve requests to customers, even if your database is not working. And that's why also we often use CDNs or actually you have caches every layer from your client, the CDN, the backend, even the database very often have cache.

So, having all those layers of cache can actually add also some complexity and some problems. But the idea that if you can't serve data immediately, maybe there's a version of that, that you can serve. And then maybe dynamic data can be replaced with stale data. A good example, Netflix has a very nice UI, with a lot of different microservices for each of their recommendations. If one of them doesn't work, or if a few of them doesn't work, they fall back into data storing cache, for example, a most popular topics in US today. Well, it's not dynamic. It's something that can be processed once a day, and then you serve this from the cache.

So that means that if your system is experiencing issue, instead of having all your customers query the database for their particular personalized profile, well, you serve stuff from the cache, right? So that gives you a way to free some resources from your backend. So that's one thing to do it and that's used all over the place as well. And nothing, it's the idea of circuit breakers as well, right? You have a dependency that fails, and then, okay, if that dependency doesn't return, what do you serve? At Amazon we love serving, it's quite funny, but we serve the cute dogs of Amazon. I don't know if you've seen this?

Jeremy: Yeah, I have, yeah.

Adrian: When Amazon doesn't work, service doesn't work, we return cute dogs and all that is on cache as well.

Jeremy: Yeah, so I think circuit breakers are one of those things where I don't think enough people use them, right? Because that's one of those things where when we start overwhelming a downstream resource, we have to do something to stop overwhelming it, right? So even with those exponential retries, or the exponential back off, and the retries, and the jitter, and all that stuff. If we keep trying the same thing over, and over, and over, and over and over again, eventually we're just going to build up so much load in our queues that it's going to take forever to work through it.

Adrian: Yeah.

Jeremy: So there is this thing called load shedding, a whole other crazy-

Adrian: Rejection.

Jeremy: And rejection and things like that. Can you explain that a little bit? Because I think that is definitely something that's not built in that you would have to manage yourself.

Adrian: Right. So I mean, the idea is to protect your backend as much as possible. And there's few ways to do that. From, of course, the clients can try to protect the back end by doing retries and back off, but the server, the backing itself at some point, if it really is overwhelmed by requests can do a few things, right? It can simply reject requests, and say, "Okay, no, now I'm at capacity, and your API is not the priority so I'm not going to deal with it," and that's rejection. And you can do load shedding as you say.

You know how much time it takes to process a request, right? So basically, you say, my request to not reach a timeout, I need to process it in at least seven seconds, right? If it starts to take too long, so if this latency for handling the requests start to increase, again, you can simply shut the load. So you remove, you say, "No, I'm not taking any requests now, because my latency for my request is at maximum," right? So that's one possibility, is to do as well.

And then, I mean, another very important thing is rate limiting, right? And that's, again, it sounds very simple. And I know people don't very often implement rate limiting from their own services, but they should. Because sometimes, one day their own services might do something wrong and go into an infinite loop where because you do a deployments, a configuration was not right. And then you have an infinite loop of requesting stuff from a back end that totally destroy your back end. And if you would have had the rate limit in place for your internal services, that would have avoided this. And I like this idea of rate limiting even your own services, because then you can establish contracts between different services and different parts of your system. You say, "Okay, my backend is this. This is the API, and these are the contracts. That when other teams accept to use my service they agree on that contract." And then if they need more requests, then they have to modify that contract.

So that means my backend team knows what is happening with my service. Because I've seen this happen a lot. You have a distributed architectures with different teams handling different services, and your service becomes popular and all of a sudden other teams start to use it, and they don't tell you about it. And then, it's fine, or and then you have a marketing campaign and no one tells you about it. And the marketing campaign all of a sudden is worldwide and everyone downloads or connects to the same endpoint at the same time. That's because there was no contract between the teams. No one agreed, "Okay, my service can only handle a thousand requests per second for you. And if you want more, you need to modify your limits." In fact, this is also why we have a lot of limits on AWS because we have so many distributed services, that teams are forced to negotiate to make sure that we don't kill other people's service, right? So it's, I like this idea of API contracts.

Jeremy: Yeah, no, that rate limiting thing too is, this is something that I don't think people think about. You're right, like if I have a service that is my customer service, and some team is responsible for building that. And also the fallacy that serverless is infinitely scalable too if we think about that, like nothing is infinitely scalable, right? Things can be designed to scale really well, and handle load, and scale up quickly, like that is possible to do still a lot to think about. But the rate limiting point you make is really, really good. Because if I'm a team, to go back to that customer example. I build the customer service, and our marketing team comes along and says, "Oh, well, I need to look up this, you know, every time somebody signs up with some form, I need to check to see if they're already a customer and do something with that."

If that is some massive thing, where all of a sudden, now you're getting 10,000 requests per second. Well guess what? Your Ecommerce system that's also hitting your customer service, now all of a sudden, that can't get the data that it wants, right? And it's this noisy neighbor type effect in a sense, where you're depleting services, or you're depleting resources from your own services. So I love that idea of contracts, rate limiting. I mean, even giving, if an internal team is accessing your own service, handing out API keys with special rate limits and quotas and things like that, I think that makes a ton of sense. So I love that idea.

Adrian: Yeah. And it gives the team building the service a good understanding of what's required in terms of scalability. Because if all of a sudden, if you have only 1000 requests per seconds, it defines the kind of architecture you can do. But if that becomes a lot more, so if all of a sudden you realize, hey, you've given out a lot more API keys, and each of those API keys have a thousand requests per seconds, it can easily go to hundred thousands per seconds maybe on the very large companies, and that's a different architecture. It might actually change the entire architecture because all of a sudden, you have other kind of consideration to take. So it's super important because it gives the ability for the team to understand the service they're building, its scalability patterns, and prepare for it. And it's what we call a cell at Amazon. You've heard the term sales, right?

Jeremy: Yeah.

Adrian: So, we define the size of a sale based on this kind of thing. The scalability patterns, the rate limiting, all these kind of things that are necessary to serve customers well.

Jeremy: Right. And so speaking about system availability, we do need a way to know whether or not our systems are available, and that way is typically using health checks. So you have a whole another article on health checks in this series. The thing that I really liked though, was your description of shallow versus deep health checks. Because I think this is something that not everybody, it's not necessarily intuitive to some people.

Adrian: Yeah. So I'd say the shallow health check is you check for example I'm asking you how are you? And you define the you is only you and nothing else, right? So you tell me I'm fine, but you could also decide, hey, no there's all my family members as well in "you" and tell me, "Oh, no, I'm fine but my wife is tired. My kid is at school." And this is a deep health check because you go much deeper into what is "you."

So it's the same for an instance, the same for a system is when you ask the health of an instance for example, you can say, "Oh, is my instance up and running?" Yeah, that's a shallow. Okay, cool. You have access to local network, you have access to local disk. That's shallow, right? It's like your immediate environment. But if the instance needs to talk to database to cache, can send API queries to third party dependency, well, that's kind of second level, right? So that's kind of also part of its health. Because if it can't reach the database, well it can't maybe do everything.

So the deep health check is around that, it's really understanding the dependencies and the second level, and sometimes even third level dependencies of what you're trying to contact and then report that. Because once you report that, then you can adapt your query. You can say, "Okay, my service doesn't have any database. So I won't do queries that are changing state, for example. I won't try to change my profile picture or change my name." Or, I can't offer that, but maybe I can offer API's that can read only. And so this kind of health check gives a capability for the client to degrade more wisely, right?

But of course, you have to be careful what you tell the client what is available. Because then you have hackers that can also understand how the system is built. So that's why actually, when you build the health check, very often it's built in terms of like, it's unique to different companies, how they work, and what they report to the client and things like this.

Jeremy: Yeah, I just think it's interesting, because I've seen a lot of people build a health check for like an API or an API Gateway, where they have the health check is just a Lambda function that just responds back and says the service is up and running. I think it was that like, I'm not sure what you're checking there other than that-

Adrian: Lambda works.

Jeremy: That Lambda works, right? And that the Lambda service is up and running, which is funny. But that's the kind of thing where if you were building like a serverless health check, like, you think, well, the infrastructure is up and running. But you could do things where if you are connecting to a database, like what's the... maybe you're collecting some metrics, what's the average database load? What's the average response time, though or the latency there? Are you able to connect to a third party service? How many failures have there been to a third party API in the last minute, or the last real rolling five minute window or something like that?

So I think that's important to understand, because then you can build rules around that, in order to decide whether or not a service is healthy enough for you to keep sending traffic to it.

Adrian: Exactly. I mean, and even if it's healthy, so for example, even if you can query the database, but if it answers after seven seconds, is this healthy? So you can have this deep health check that answer, I'm okay, but it takes seven seconds, and so then it forces you to define thresholds as well. And this is what we discussed about later is like, "Okay, how fast do you want your service to answer." And that that defines your, so this is the business, it's a business requirement.

You say, "Okay, my customers needs to be able to access data in four seconds." If it's not, then you shed, you do something else. And that defines a lot of the default, or that you're going to have to put in your systems and then, and it just helps you understand and build the system a little bit more predictably.

Jeremy: Alright. So let me ask you this question. Let's say we build in these really great health checks that and we set some thresholds, we say, if the data doesn't come back within four seconds, or whatever it is, then we want to route that to a different service or to a different region or something like that. What happens if all of your services are coming back with bad health checks?

Adrian: Yeah, this is a good point. And let's say you can have bad health checks like this when sometimes you make configuration mistakes or you do a deployment and something doesn't work. And sometimes, if everything fails at the same time, you have to assume that it's not broken, right? So it's called failing open. So it means okay there is, it might be a health check problem, so let's continue sending traffic to the environment and hopefully things will work. And this is what we have in places on AWS, if you look at all the systems that implement health checks or Route 53, the ELBs, Lambdas, API Gateways, and all this kind of things, if all the health checks fail at the same time, we fail open, so we assume it's more like a health check problem versus an infrastructure problem.

Jeremy: Right, yeah. And then the other thing too, that's kind of cool. And this is just something where I don't think people understand how powerful Route 53 is. Because if you think about your normal load checks, or your health checks, I know I always would think about application load balancers or elastic load balancers. But that is region specific, right? So if, again, I can't health check across 10 different regions or five different regions with an ELB. I need to do that at a higher level, and Route 53 has a ton of capabilities to do this.

Adrian: Right. So now, Route 53 allows you to do health checks on many different levels. And what's the nicest feature of Route 53 is that when it checks the health check, it uses eight regions by default from around the world, right?

So sometimes on the internet, you have regional outages, and sometimes the route on the internet doesn't work, but it doesn't mean other routes don't work from outside. So, Route 53 allows you to go around these kind of regional outages or intermittent regional outages over the internet because it's checks over actually, eight regions, and then three availabilities for each of those regions. This is where the 18% comes from, that's written in the blog. Which is a bit weird, but it gives us an idea that if 18% of the system is at least answering is... if 18% or fewer of the health checks report that is healthy we'll consider unhealthy because it's not enough. So it means something is wrong.

Jeremy: Yeah, so alright. So then you figure out that a particular thing is healthy, you get this consensus, which is crazy, because you're right, this 18% is a weird number. So basically, a lot of them can fail but as long as it's running. But so once it decides that something is healthy and it starts routing traffic there, there's also a bunch of other capabilities too where it's not just route it, based off of it being available. That's one part of it, but then also you can route things based off of geographical distance, you can route things based on latency. Yeah, so how does some of that stuff work?

Adrian: So that's the case where you have two systems, or two environment and then you want to switch between one environment and the other. And I would say, maybe, actually multi region maybe in that case, is the idea that when you are a customer, you want to have data fast. And if you want to have data fast, then you want to have low latency. To have low latency, you have to have your data as close as possible to the end user, right? So in the last, let's say, five, six years, we've seen an explosion of multi region architectures because now we have global customers, right? App stores have exploded, basically we have customers around the world and each of those customers, game is a very good, good example of that they want to have as small ping as possible, right? So the latency should be as small as possible.

So we have to have systems that are deployed very close to the customers, so in multiple region. And then you need to figure out from that user, how do you route the user to a particular backend or particular environment. And then you have different policies, right? So, these are the policy you were mentioning, in Route 53 of whether it's geographic, you have latency, you can have weighted Round-robin. And then, of course, all that supports what's called a failover. So, if any of those fail you can failover to another region. So, Route 53 gives you very complex sets of possibility to make very complex sets of routing, and very flexible as well. But it can be complex. But yeah, so this is the idea behind that.

Jeremy: The thing that's important, though about this multi region or these multi region architecture. And of course, if you're using latency based routing, or you're using geographic based routing, or even just Round-robin routing, you're looking at sort of this active-active type environment, right? So it's, this is not if this one fails, then shift all the traffic to this, this is I have a region in Europe and I have a region in the US, so I want to minimize my latency for customers based on that. So when you're designing those types of multi region active-active systems, especially from a serverless standpoint, you want to be using regional API's instead of edge optimized ones, correct?

Adrian: Correct. Yeah. So I mean, this is especially if you use API gateway, right? So API gateway when it was released, came with an integration with CloudFront, so basically you got a domain name, which was, well, not regional, it was global. So you couldn't basically use Route 53, which is also a DNS provider to actually route traffic to that particular API gateway. So I think it was about a year and a half ago, API gateway released a regional endpoints. So that means now you can have an API gateway without CloudFront integration. And that means now you can use Route 53 to actually route traffic to directly the API gateway in your region. Actually, you can have several API gateway in one region, as long as they have regional endpoints. So you can have multi, I would say multi API gateway routing in one region via Route 53.

So, and this goes into more complex discussion because I see a lot of people design an application per region, right? So you use one API gateway for your serverless application per region. But, if you think about it, the blast radius is high because you have one API gateway for one region. What if that API gateway has an issue? Well, you could have several API gateway, you could have two, three, four, and so that means, I think you need to think about how do you shard basically an application, and the idea of sharding is okay... On Amazon we call that a sale. We say, "Okay, we have a sale," which has an API gateway, maybe Lambda, DynamoDB. And that sale will take, let's say, 100 customers, right?

As long as they're less than 100 customers, we only have that particular sale. But if we have more customers growing, instead of growing the sale and which we know how it behaves, because we've been testing it, we understand the pattern, we can deploy it. It's repeatable, it's very well understood, well, we replicate that sale. So that means at some point, we might have hundreds or thousands of sale in one region. And you can do this as a customer, also, today. You can say, I want to have 10 API gateway, Lambda, and DynamoDB pair per region, why not? Now, it doesn't mean you need to do it, you need to really understand the thing. But that's the idea of a regional endpoint for API gateway is that you are not limited to one for your region, right?

Jeremy: Right. Yeah, no, and I think that's one of those things too where just I've always fallen back on the edge optimized ones, because it was just there. But now with the new HTTP APIs, you don't have those. So I like that idea though, of being able to say, I have more control now of which endpoint my user gets to, I can control that latency a little bit better. So, interesting thing to think about, certainly, another thing to put on your list of things to learn and things to do.

Alright, so we talked a little bit about caching earlier. But you have a whole other article on this about caching for resiliency. So there are obviously a million different things to think about when it comes to caching, how many layers of caching do we need that kind of stuff. But there are some really good reasons for putting, let's say, a CDN in front of your application.

Adrian: Right. So yeah, that's.... I think the CDN is probably the first layer of cache that people tend to use simply because it has massive security implication, right? So CDN, to explain a little bit what a CDN is, is a collection of servers that are globally distributed, they are globally distributed closer to the customer. It's way more, basically a Point of Presence or I think if people I've seen this for CDN, it's called PoP. It's a Point of Presence. It's not matching the AWS regions, there are way more PoPs around the world than there are AWS regions. So that means they are way closer to the customers.

So each of these PoPs is basically an entry point into an application, right? So when you use a CDN, the CDN uses some traffic policies that basically makes the request of the customers come to the closest PoP available to the customer making that request. Okay. So it also allows you to have hundreds of entry point into your application dispersed around the world. So that means, instead of having one entry point, for example, the API gateway, you put a CDN on top of it, you have hundreds of entry point, hiding your API gateway. So that means that doing a DDoS attack on the API gateway becomes a lot harder because all of a sudden, your attacker needs to attack hundreds of Points of Presence on the CDN to be able to do DDoS attacks, right? So people use CDNs, well, to improve latency of static content, but also to make it resilient to DDoS attacks. Because well, it's much harder to attack 160 point versus one, right?

Jeremy: Right and CloudFront and AWS WAF, like have some of these things built into them too to protect against like UDP reflection attacks and SYN flood and some of those other things too. So, a lot of that is good to protect just from I guess a resiliency and uptime sample, and we talked about overwhelming systems, right? So if you have a DDoS, or something that's happening that is overwhelming that system, being able to shed some of that at the CDN layer, because AWS is smart enough to pick that up is just great. It's great and a good layer to have there. But beyond just the I guess the hacker attack or the protection that you get there, depending on what type of data you're serving up.

So let's go back to the Netflix example, that top 10 movies in the US or whatever it is, top 10 shows in the US, that particular thing is pre-generated, and it is served from cache. And that gives us this idea of I guess, page caching even, like even if it's a short amount of time. So what are some of those things that you can do where you can reduce pressure on that downstream system or on your system? Like some of those different techniques for caching.

Adrian: Right. So, to continue on the CDN, right? Like what you have explained, what you said was CDN is actually a cache, right? It's a cache layer to serve content. Primarily people use CDN to serve video, or pictures, or static files, things like this. They often forget that actually CDN, CloudFront for example can also cache dynamic content, and dynamic content, even sometimes cache for a couple of seconds can save your back end. If you... let's take the marketing campaign, for example. Let's say that the marketing campaign is actually a list that is dynamic of something that changed. And if people press refresh every second, because they want to see who is winning the competition or something. If you don't cache, even those dynamic content, the content of the list, even few seconds, that means everyone is going to query the backend, right?

So I think caching dynamic content even for one seconds or two seconds is very, very important, right? They think about the possibilities of doing that. So that's kind of also something that CDN's can do and often people forget about it. So, does it answer your question?

Jeremy: I think it does. And I mean, I guess where I'm trying to go with this, too, is that there's just a million different things that you can do to cache, and there's multiple layers of cache. There's multiple strategies or caching patterns that you outline in here. And I think that, people need to go read the article, I think in order to really understand this stuff. I don't know how much justice we're doing it trying to explain some of it. But I think one of the things we should touch on, just in terms of caching in general, and this is a quote that you have in your article is that for every application out there, that there is an acceptable level of staleness in the data. So just what do you mean by that?

Adrian: So, it's exactly the idea of caching dynamic content. Well, even if you claim your application is very dynamic, and you claim that, no, I need to, I can't cache because for example, it's a top 10 list of real time trends on Twitter. Let's say Twitter trends, right?

Jeremy: Right.

Adrian: People expect that it's real time. So, I would say by default, if you think about real time, people wouldn't think, "Okay, I need to cache that." But, if you have millions of clients around the world requesting that data, absolutely you're going to fake it real time. It's, you might query your downstream server or service that tells you the trend, but maybe if you have thousands of clients connecting at the same time, you don't want each of those clients to query your service, you will just serve it from cache, or make sure that the requests are packed into one single request and then that's the downstream service and then serve back the content.

So it's just this idea of like any application out there, even if you think it should be a, must be real time, it's very important to think about the staleness. And staleness is how real time my data needs to be, even if it's maybe three seconds old, is it really that old or is not usable? Because it's also something you can fall back. So if your database is not accessible, it's like, maybe you can serve back the trend of Twitter, that was maybe an hour ago, and just instead of... and you can say to your customers, you can say, "Oh, we're experiencing issue, this is a trend one hour ago." And that's fine, it's a good UI. It's good use of stale data. And why would customer say, "Oh, you're cheating on us?" No, it's like... I think it's a good example of that.

Jeremy: But that's actually, that brings me to the caching patterns, because that is one of those things where, like you said, there's acceptable levels of staleness in data. So again, if a data is five minutes old, and we don't know how many tweets are about, I don't know, Tiger King, or some popular thing is happening on Twitter or whatever, if we don't know how many, what the most accurate count of tweets is for that, that's probably not going to kill us if that's five minutes old or a minute old or whatever.

But so there's a couple different patterns here for caching. So obviously, we have cache aside, inline caches, things like that. But the ones that I think are more interesting and I'd like to talk about is this idea of soft and hard time Time To Live explain that.

Adrian: So a soft Time To Live is your requirement in terms of staleness, right? So you say, my Twitter trend lists, I want to refresh it every, let's say, every 30 seconds. So you give it a TTL of 30 seconds, a soft TTL of 30 seconds. So if my service requests the cache and the TTL, the soft TTL is expired, and everything is fine you go and query the service, right? But if my service doesn't answer at that moment, so you are, you've passed the soft TTL. Now, your downstream service doesn't give you the data. What do you do? Do you return a 404, or you actually fall back, and you say, alright, my soft TTL is expired, but I'm still within the hard TTL which is it's one hour, right?

And then you say, okay, your service returns the hard TTL and you say, "Oh, sorry, we just have one hour old data, because we're experiencing issue." So again, it's a possible degradation. And actually quite often cache could be used like this. I think it's all about how you create your cache and things like this and how you define your eviction and policies and things like this.

Jeremy: Right. And I think that that is something that a lot of people don't think about. I think the most common way to get rid of your cache is just to set a TTL on a Redis key or something like that. And then that expires, and you're like, "Oh, wait, now I can't fetch new data, what do I do?" So that's really interesting that those two ideas, there are something that people should be thinking about.

Alright. And then the other thing was this idea of requests coalescing, because this is another problem you have is if it takes one second to repopulate the cache or to run that request, if you have 1000 people requesting that before the first one completes, you need to be able to handle that in a certain way.

Adrian: Right, yeah. It's exactly like, you go, you have one requests, and then a thousand requests asking for the same data, but you don't have it in the cache, what do you do?

Jeremy: Right.

Adrian: So one way to do it is you say, "Okay, I'll take one of these requests, they're all the same, and I'll get the result and then use the same result for everyone." So that you park basically all the 999 requests on the side and say, "Wait a second, I went to ask, and I'll give you the data." This is very related to idempotency, and things like this. So that understanding what requests are you doing, what data can be returned from your API, and if it's dynamic or static content. But actually, some framework supports this kind of things out of the box. So it's important at least to figure out if your framework supports that.

Jeremy: Yeah. And I think there are some patterns that you can build fairly simply in order to do that even if you're using a Lambda function, for example. So, but anyways, very, very cool stuff.

Adrian: And if you use DynamoDB actually, with DAX, it supports this kind of request coalescing hard to pronounce for me.

Jeremy: Hard word to say. I think that, the last thing you point out is just be aware of where your caching happens, right? Because there's just so many layers of caching sometimes that it can cause a lot of problems.

Alright, so the last thing I want to talk to you about and then I'll let you go, is this article you wrote about building multi region active-active architecture on AWS that was serverless, right? So if you were doing that, just give me a quick overview how would you build an active-active multi region serverless application on AWS?

Adrian: You're asking me for a pill give me your pill to create an active, active... So first of all, if you define active-active I think before giving you the solution, if you think about active-active a lot of people think about active-active and they think about data replication, right?

Jeremy: Right.

Adrian: So active-active doesn't necessarily involve data replication, right? So you can have active-active but federated as well. So that means you have local database. And if you look at Amazon retail site, you have Amazon UK, Amazon US, Amazon Germany, they're all it's an active-active system. All the regions are active, but they are federated.

So as a business you need to understand, okay, first, what is it that you want? Do you want a multi region with Federation, so local database or do you really want to replicate all your data under the hood, right? So, if you really want to replicate the data under the hood, because this is what I used during in this blog post I was using Dynamo global table which allows you to replicate the data across multiple region. So, that means that if one region is experiencing issue, my data is replicated synchronously to other regions. So then I can failover to another region to get that data.

And again, this is when you say multi region, it doesn't mean multi continent, because in the US you have multiple regions. So when you do multi region, you need to be very, very careful about your compliance in which region of the globe you're operating. If you do this in Europe, there's GDPR so you probably don't want to have Dynamo in Europe and US being replicating data from customers. Because all of a sudden you have different kind of governance on the data and regulation and laws and stuff like this.

So, you can use multi region within one continent, because you want to maybe serve customers faster, because you want to decrease latency. Now, at the end of the day, when you do, we did test on the retail side hundred milliseconds latency, reduce the sales by 1% on the retail side, Amazon retail side so 100 milliseconds latency is very, very little, right?

So when you are between let's say, Germany, Ireland, France, or other regions in EU or even east side or west side of the US, you can easily have a 100 miliseconds latency improvement by choosing the right region closest to the customer. So once you've done that, then you can define multiple regions by which you want to serve your data within the respectable law and governance. And then, of course, you can use Route 53 to direct traffic between different region, or there's also the new global accelerator that doesn't use DNS, but it uses IP Anycast to move traffic. So you avoid all the DNS caching, which is a problem. It's a whole other podcast if you want to talk about DNS caching and the problems of DNS. But yeah, you have multiple solutions to do this.

Now, word of warning, multi region is, it adds complexity as well, right? So it's very, very important that if you decide to go multi region that is a very, very strong business case, or that do you know really well what you're doing. And I always say to customers start with maybe one region, automate it as much as possible so that you can just move your automation to another region, and then avoid data transfer between regions, right? Try to federate the data, and then, you know, I think Federation works great because if I have a service like Amazon.com, I'm using mostly the German Amazon because it's closer to me, and then the delivery is obviously better. I rarely go to US. So, there's no reason for Amazon to transfer my data or my shopping cart between Germany and the US.

So if I go to the US store, I need to recreate a cart and re-authenticate again. So it's still an active-active, it would serve me better if I moved to the US but how many times a year am I in the US to shop very little at the end of the day. It's important to understand that maybe it's acceptable if I'm in the US that actually I'm routed to the German Amazon, and I do my shopping from there, maybe twice a year with increased latency, and that's okay. I don't need to have very complex systems to support the three days in the year while in the US and replicate all my data and make sure everything is there. So it's very important to really have a strong understanding of if you're building multi region, are you doing it for the right thing?

Jeremy: Right. Yeah, no and I think that you made a really good point there where it's like that the Federated aspect of it, especially depending on where it is. Like if you're using three regions in the US, for some reason, then federated might not work, because you could get routed to different regions, you might want to use global tables in that case. But if you were just doing one region in the US, one region in Europe, one region in South America or something like that, then maybe just replicating, say, the login data, right? Like just the authentication data using a global table for that, but then federating the other data...

Adrian: The customer data.

Jeremy: Yeah, right. So some of that other stuff. So I think that's a really interesting approach. People have to go read your articles seriously. And I'm going to put all this into the show notes. So honestly, thank you for sharing not only here and dealing with the technical issues that we had, which might just have to be another blog post at some point, we'll discuss. But seriously, thank you for writing all those articles and sharing all that knowledge. And just giving people the insight into some of this stuff, which I think is not publicly available, it's not readily available, you kind of got to dig through that stuff to find out what is important and what's not important. And you do a great job of summarizing it and going deep on those things.

Adrian: Thank you very much.

Jeremy: So thank you very much for that. So if people want to get in touch with you and find out more about what you're working on, and your blog post, how do they do that?

Adrian: So I'm pretty much everywhere on the internet, Adhorn, A-D-H-O-R-N, whether it's Twitter, Medium, or even Dev.to. So yeah, I'm pretty much there. And you can... anybody can ping me on Twitter. I have open DMs. So if you want to talk about anything, I'm happy to answer. I'm much better at writing than I am at answering live questions. So I hope I did justice to what you expected. But again, thank you very much for having me on your show. I'm a huge fan of what you do, Jeremy. So thank you very much for everything.

Jeremy: Thank you. I am a fan of yours as well. So thanks again and we'll get all that information in the show notes.

Adrian: Thank you very much.

THIS EPISODE IS SPONSORED BY: Amazon Web Services(Innovator Island Workshop)

View Details

About Guillermo Rauch:
Guillermo Rauch is the CEO of Vercel, but before starting the company in 2015, he was CTO and co-founder of LearnBoost and Cloudup, acquired by Automattic in 2013. Guillermo is also the creator of several popular Node.js open source libraries like socket.io, mongoose and slackin. Prior to Node.js, he was a core developer of the MooTools frontend toolkit.

  • Twitter: twitter.com/rauchg
  • Vercel: vercel.com
  • Next.js: nextjs.org

Watch this episode on YouTube: https://youtu.be/iRNxV9vRg6o

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and this is Serverless Chats. Today, I'm speaking with Guillermo Rauch. Hey Guillermo, thanks for joining me.

Guillermo: Hey, thanks for having me.

Jeremy: You are the CEO of Vercel, which was formerly ZEIT, so I'd love it if you could tell the listeners a little bit about yourself, your background, and what Vercel is all about.

Guillermo: I'm the CEO and co-creator of Next.js, which is the React framework for front-end development and JAMstack development. Vercel is the platform for deploying projects like Next.js and many other frameworks. Vercel focuses on making the lives of front-end developers really, really easy, allowing them to push their pages to our edge network, and have a very delightful serverless development experience.

Jeremy: That's what I want to talk to you about today. The last time, I think, we saw each other in person was back in... Was it back in Milan, I think, right? Almost two years ago at this point, maybe it was last year. I don't even remember. Quarantine has lasted so long at this point that I can't keep track of time. The last time I saw you, I was speaking about this idea where I felt like serverless was getting harder and harder and harder.

That was or it seems to be the wrong approach, right? We want serverless to become easier. This is something where, I think, this idea of maybe I think you call it front-end serverless or serverless front-end is where you're trying to go with Vercel. I'd love to just get your thoughts on that, just that complexity that we're now pushing towards the back end, and where you're trying to go with the front end.

Guillermo: I think you nailed it. I think the serverless world is big and complicated. I think when we first met, we really connected on this idea of like, "What is even the right definition of it?" We were both presenting at Milan trying to give a definition for it. It's a pretty silly game to play to try to even fight that fight. When I think about serverless, I think about wanting to give people a very good recipe for leveraging that kind of technology.

I think anything that relates to serverless or infrastructure really needs to disappear. It has to be all about letting people focus on their products, focus on their pages, focusing on the things that they're publishing to the internet. That's why front end really is the place where, I think, all the serverless action is happening, and the techniques and technologies that we're using in some ways are the original serverless because much of what we're doing today is this idea of taking pages, generating them statically and putting them at the edge, which means...

To me, the most fundamental serverless technology out there is CDN. They've been around for a long time even they predate a lot of the serverless movement. Yet, they had that critical idea that there is no management to do, that it accelerates you, obviously, because it's putting your content next to your customers. The very technology that this accelerates is the front end. I think what we're about to see is that a lot of what we've been advocating for in the serverless world is really starting to become much of a reality with front-end developers.

Jeremy: I think that actually makes a ton of sense, because whenever I was thinking of serverless, I would always think about the actual computations that were happening behind the scenes, so whether that's something where you're running a Lambda function, and it's pushing it into SQS, and you're connecting to DynamoDB, and you're doing all these different things with the data. A lot of that is still necessary, right? There's a lot of complexity that has to happen behind the scenes in order to make a full-fledged serverless application run.

But I think the funny thing is that a vast majority of the applications you see out there are just a collection of static pages. That's, I mean, with a little bit of API happening in the background, but that shift, that thinking of compute versus static pages, isn't that really where we want serverless to go to is just this super easy precomputed system?

Guillermo:Yeah. I think a lot of people in the industry have over focused their attention on computing on demand, which is what Lambda enables, right? You're literally firing up a VM. It's amazing how easy AWS made it. It's almost like a miracle that you deploy your function so quickly, and it executes so quickly, and it's secure and in a VM sandbox, and their underlying technology is absolutely incredible with FireCracker, but the question that you have to take a step back and ask yourself is that do I really want to be computing so much? Do I want to be burning electricity and competing cycles so much?

This is where when we really sat down to analyze this problem, we realized the vast majority of pages that you visit every day on the internet can be computed once and then globally shared and distributed. So it's like the technique of memoization and functional programming where you compute once and then of course, you want to read it from that intrinsic automatic cache that you get, is different from caching because caching requires a lot of developer effort and thinking. Memoization gets closer to what I envisioned to be the foundation of serverless front end, which is basically static generation where the computation happens once probably as a result of some data pipeline, something that changes, computation happens. HTML is spit out.

That is all, and even in the case Vercel, it's powered by functions too by the way, but the funny thing is that the developer never even thinks about functions. They just think about building pages that then get pushed to the edge and then consumed by visitors. Now, that's not to say that the on-demand use case doesn't have any merit. Not everything can be computed statically. There's lots of pages where you sign in to a dashboard, and you have to query data that could absolutely not be cached. A great example is you log into your bank, and imagine that you were trying to statically generate your dashboard with your bank account balance, but you just want to check that your payment went through for utility.

You're not sure if what you're reading is up to date or not. You would go crazy, right? The movement of front end has also led us to where that dashboard is a single-page application, most likely, that is also served statically from the edge. Then there's JS code that runs on the client's side that then queries that back end. What we found is that front end is really powered by this set of statically computed pages that get downloaded very, very quickly to the device, some of which have data in line with them. This is where the leap of performance and availability just becomes really massive, because you're not going to a server every time you go to your news, your ecommerce, your whatever.

You're just downloading it from your very own city, but even in the case of like, "I may have to make a strong read, not a read that could be stale," you're basically also downloading static content that then runs JavaScript on the client, and then that goes to a server. Then the question becomes, "Who's writing that server, and how much of that server are you writing?" This is like the other big question that is, I think, coming up. We're confronting that in the serverless world is like, "Okay, I have all these amazing primitives to build everything in the world that I could imagine from scratch, but does it make sense to build everything from scratch?"

Does it make sense for you to build your own authentication function with Lambda if you could be reusing a standalone authentication service? That's why this interesting world is coming up where there's a rise of the front end, but then there is a rise of the API economy. What I mean by the API economy is that we have services like Stripe and Twilio and AWS Cognito and Auth0, and MagicLink, and all the services where you're just making some quick API calls sometimes directly from the client side, right?

That is a serverless world that seems so much more in my mind attuned with the actual ideal and the actual, original promise of serverless. I think we are too much. You're giving the example of SQS and Dynamo. We erred too much on always rebuilding from scratch a little bit, so focusing on the front end allows you to reprogram your product strategy in a way, where like, "Okay, I'm going to think about the customer first. I'm going to think about building my back end very, very low in my priority list, right?"

Jeremy: Yeah. No, and I think just this idea of what could be static content versus what needs to be generated dynamically, and I mean, I think of an ecommerce site, for example. Every product page, every category page, every set of recommended products for a particular product or related products, or things like that, all of that stuff could be precomputed and pushed out to the edge. Then the developer never has to think about processing the scale of that, because if you think... I know you used to work at WordPress, right?

That was one of the things you did before. As you know, WordPress loves to query that MySQL database on every single page load.

Guillermo: That is a great example. That is a great example. I think this is the difference, you just nailed it, between ahead of time computation or generation static pages versus just in time. With the just in time model, which is what WordPress is doing every time you go to index.php or blog.php, you're creating all this load. You're sometimes issuing dozens of queries. Something you and I were talking about before the show is that because we're at peak cloud in the amazing power and capacity that we have our fingertips, anything, it seems, could scale today. If you use the new serverless MySQL service, I'm sure Jeff Bezos will sell you enough MySQL on-demand capacity.

With his incredible database engineers, they'll make MySQL scale so much that you might actually make that work, but the question is like, "Do you actually want to? Do you actually want to, first of all, pay all those database bills? " Secondly, it seems like we're increasing the entropy of the universe, and we're producing all this heat and carbon emissions for no reason. The point is the writes that happen are here and there to those product pages, to those blog posts, to those marketing pages.

Somebody at a marketing team might go and say like, "Today, I'm going to edit the headline of this page, or today, we're going to work in a blog post," so you write and the writes are not that frequent. Then that is what creates this asymmetry. If you can take the opportunity to convert that write to the database into HTML when it happens, and then share that super easy HTML stream that gets downloaded from an edge server, you can't compete with that with any other serverless architecture. You can't compete with that from a speed of light perspective, but also, cost wise, you can't compete.

We've talked a lot about over the years about how amazing it is that Lambda gives you 1000 concurrency and whatever, but at the same time, just imagine 1000 VMs in a rack firing up to respond to your blog post. It doesn't seem very appealing. Then from a developer experiences standpoint, this is really what we're enabling with Next.js at a very large scale is that we also don't want people to necessarily have to think or remember to apply caching. This is why we took that idea of the CDN, but now we're really taking it to the next level because CDNs always require this calibration of components where the front-end layer has to coordinate with several layers of caching.

Over the years, I've talked to so many people that have front ends that combine a Redis cache, and then beyond the Redis cache, there is the CDN cache. Then there is all this brittle purging and invalidation of strategies all over the place. Then when you peel all these layers of complexity, you remind yourself, "Oh, I was just working on this simple page that had this simple content." If you think about the ecommerce example that you just talked about, the underlying JSON data structure for that page that renders the ecommerce item, and recommended products and so on, it's very simple.

The idea that you could convert it into HTML, and serve every market in the world with that precomputed HTML is extremely compelling.

Jeremy: Yeah. I mean, and if you think about, like you said, this multiple layers of caching, the last big company that I worked at, every time there was a problem, the engineers were always like, "It's a caching issue. It's just a caching issue," because it was like CDN. Then in front of all of the application servers, there was a Varnish cache, and then there was Memcached in the back and all kinds of these other things that were just layers and layers and layers, and you never knew where it was.

Guillermo: Yes.

Jeremy: It's funny that you mentioned this idea of these infrequent writes to basically massive reads. I think about going back to the WordPress example. There are people who might argue well, but you need to keep your comments up to date. You need to see the freshest comments. Well, I think about any installation of WordPress that's getting comments. Even if you are getting comments at a very rapid click, which is probably unlikely for a WordPress installation, you could just take the write, once that write happens, and then generate the static content or the comments list or whatever it was, and push that back out to the edge.

We're a point now where I feel like the edge has become the only cache we might actually need if we do these things right.

Guillermo: Yes. Yes. Yes. Thanks for using... Also, reminded me of the Varnish example because the cleanest architecture you could think of is one where there is no layers of caching that need to be coordinated, right? That's why I don't think about this new wave of edge as necessarily a cache. For those of you that have done versions of this by hand with S3 and CloudFront, when you put content that gets generated into a bucket, and you know that this bucket is super highly available, and it's super easy to think about how you could invalidate the edge once this specific writes happen to that bucket, you don't really think as a complicated caching scheme.

You think more about it's just a simpler model for reasoning about your architecture. Let's analyze that comment example for a second. Let's say that you're the New York Times, and you have a very high comment throughput. First of all, there's high value comments that they in-line with their page, because they contribute substantially to the narrative. So they're the highlighted comments. There's not going to be lots of highlighted comments. There's going to be maybe five or six. They're almost like an extension of the article at that point.

They're like just like you would want to statically inline the paragraphs, you want to statically inline the highlighted comments. Those, again, are not subject to this strongly consistent read system. If the comment that gets promoted to the highlighted ones takes a second to reflect in the global cache in the global edge, that's totally okay. It's the right trade off to make. Then when you paginate, when you read the long tail of trolls or whoever, you can [inaudible 00:17:16] in the client's side, right?

Jeremy: Right.

Guillermo: You can actually go against your database that has high read throughput without a cache there either. You could go to Dynamo or whatever, and tell, "Give me the very latest comments. At this point, I want a strongly consistent read. Give me your very latest comments." At that point, by the way, also, New York Times performs moderation on their comments, so they're throttling the writes anyways to make a better quality product for their pages. The system just fits like bread and butter, I think.

At the end of the day, we're all in this business of publishing quality content that we want our visitors to consume ideally in that first TCP packet. What I say is that I'm a really big fan of deleting code. I don't want any code to execute anywhere. If I'm going to newyorktimes.com, and I'm from Argentina, in Buenos Aires, I don't want all this Turing complete circuitry to being between me and landing page of an article. I just want to go direct to that HTML stream. Browsers are so good at rendering a stream of HTML already. Think about it, we've deleted all the code.

There's no JS that needs to run on the client side to give me that first paint of the article. Everything has already been precomputed, so there is no function execution in between me and the content. It's this crazy combination of availability, performance, greener for the world, and just overall better.

Jeremy: Right. I want to jump back to your comment on databases because you mentioned DynamoDB there, and I'm a huge proponent of DynamoDB. I love this idea of just having these super fast, single second or single millisecond latency to retrieve back data, but you're still oftentimes querying a database, right? It's still technically a database. You still have to wait for that computation to happen to bring that back. I like things like DAX, being able to put a cache in front of that, but I also find that if you do it right, you can cache whatever that GET request is or that query to DynamoDB.

You can even cache that at the edge, right, if it's responding back from an API call.

Guillermo: Totally.

Jeremy: The point that I wanted to make was people using Aurora Serverless or RDS or whatever it is, and using these relational databases, there is absolutely a need for relational databases, right? You can't run analytics on DynamoDB.

Guillermo: Totally.

Jeremy: ... these other BI tools and things like that. This is something I talked about actually a while ago, where if your front-end customer, if the person accessing those comments, for example, they don't need to sort them in 10 different ways. They don't need to join them in a bunch of different ways. They don't need the power of that. What I found is I've been able to eliminate almost all of the clusters of databases that I've had, and use something as simple as Aurora Serverless with two or four ACUs, so really, really small by simply replicating data out of DynamoDB into that.

Now, I can serve massive load with DynamoDB, but then when I need to write those queries, or even run reports, and I'm not having hundreds of thousands of people hitting my reporting site or my analytic site, that my ability now to have a smaller footprint there is...

Guillermo: Totally. Totally.

Jeremy: Again, I think this is what we're trying to do with serverless, right? We're just trying to reduce the amount of footprint and the amount of extra computing power that you need.

Guillermo: Yep. Yep. Yeah, I think you nailed it, because you're going at the core of the database problem and data access problem, which is understanding how the data is being accessed. You mentioned something there, which is like, "What is my throughput of queries for analytics and complex joins?" Like, "Let's find the top 10 most active commenters on my website and things like that." It's very rare. It fits very much systems that can respond more slowly, that it can take their time to scale up and scale down. Then you have the other layer that you touched on, which is, again, it doesn't matter how fast my database is, how real time, how scale, how serverless if my customer just wants a bunch of HTML of a certain set of comments in that case with the example that I gave of maybe the first five or the most voted ones and so on.

I think what's important for the developer to always think about is that think about that access pattern of your data from a read perspective, from a write perspective. I will say also not just the ratio of volume, but also the consistency. That's, I think, what's really important as well is that when I go to a breaking news page on COVID-19, and I work at New York Times, and I'm going to push an edit to it, I can afford for one second to pass before I can make a strongly consistent read of the typo that I just fixed.

I will have wanted that within that second to all the reads to go uninterrupted globally in the world, because this is so much more important for that smooth line of low latency access and highly available access to my story of COVID-19 that for everybody in the world to be able to read my writes in a linearizable fashion. I can fix my typo, and I can say, "It's 99% likely that within a second, everyone will be able to read my write." Now with Dynamo, they give you other characteristics. They tell you, "Well, you make it right, and as soon as you query that same API that you're servicing Dynamo from, you can immediately read your write."

But then everybody in the world has to go to that specific Dynamo cluster, right?

Jeremy: Right.

Guillermo: Again, what apps or websites need that? You have to really think hard about that because it's not going to be ecommerce, or at least for most of that front end of ecommerce. You might have some specific things where you really, really want to read your writes in that very low latency fashion. That might be the case. For example, when you go to the logged-in section of your website and say like, "Give me my latest five recent orders," you don't want that to be a weird stream of static generation that every time an order happens, you're custom making a static page for the logged-in administrator.

Then you start adding complexity to your permission system, and everything becomes chaos. That's why I said when you think about pages that need more granular data access, more stronger consistency, that have more complicated permission systems, more complicated queries, then that's better served still by a static page, but that runs JavaScript on the client that can query those APIs. Just to give you that idea of why the front-end economy relates to the API economy, now, if we continue this example of the ecommerce website, maybe that ecommerce API will be a headless ecommerce API.

Shopify and BigCommerce and WooCommerce and many others are now giving you very rich GraphQL APIs and REST APIs for querying this type of data as well. You even have to wonder, "If I'm making this really slick new ecommerce experience," maybe you're starting a new microsite. Maybe you're going after VR eCommerce. Maybe you're thinking of reinventing your front-end layer. The question becomes, "Am I going to be writing a serverless API with 10 queues, one million Lambdas, four Dynamo clusters, DAX, Aurora Replication, if I could have bought an API from the shelf?"

Guillermo: That's, I think, a question that a lot of people will be facing in the coming years.

Jeremy: You made the point about the top 10 comments. This is where this is something, I think, people don't get about serverless. I'm not trying to be like, "I understand it better than anybody else." It's just for me, some of these things or at least for me, the way I feel about serverless is a lot of it has to do with asynchronous operations. It's not responding immediately to a request, and that is one way in which we can get the latency down. The edge is one piece of that, but with your top 10 commenter thing, that's the kind of thing where, again, "Do I maintain a database cluster that has 50 instances running so that I can calculate on the fly who the top 10 commenters are, or is that something I could delay and maybe run every minute if it was really added to be that much and just run that off of that small cluster and then push that to a cache somewhere or to the edge somewhere so that that is pre-calculated?"

Guillermo: Totally.

Jeremy: Anyways, I love the idea of precalculation. I just think that this idea of being able to access stuff as quickly as possible is just insane. The point that I want to make too, because this is something I noticed, I was on the Vercel site the other day. I went down. I was looking at all of your edge locations, and you've got a really great page on the site that shows you all the different edge locations. It shows you where you are assuming based on IP address or whatever. Then it gives you the ping and the latency to each one of these edge locations. I think the closest one to me...

I'm up in Massachusetts in the US. The one closest to me, I think, was in Montreal, and it was like 25 milliseconds was the latency or something like that. Alright, here's the problem with calculations or computations. They have to run somewhere. You're not going to necessarily run your computations at the edge, so that's another huge disadvantage is that if you're running your DynamoDB cluster and all your Lambda functions in us-east-1, and you're trying to access it from Brazil or from wherever-

Guillermo: Absolutely.

Jeremy: ... then there's going to be a huge delay in latency out there.

Guillermo: Trip. The only time where you can justify that trip is where you need to very strictly read the writes that happened at that origin, right? I can't tell you I'm going to cache your latest five stock orders or your bank account balance. I'm going to cache it in Brazil so that Brazil customers have a better time, right? No, I'm just going to give you a static page that then gives you a skeleton placeholder of your balance, and then goes and fetches it from wherever the brain is of that bookkeeping database. That goes at the heart of like, again, the vast majority of pages on the internet should already be within 25 milliseconds of you.

You can accomplish this with layers of caching, but things get really tricky because sometimes, that cache gets a miss, and you have to go to the origin at that point. You were talking about how there's been a rise of the usage of functions for background processing. Functions are notoriously... They have a cold problem unless you're provisioning or like things get even more complicated, so you have to think about what happens when you go to the Montreal edge, and you miss. Again, we host customers with cardinality of pages in the orders of millions. Then we also see millions of deploys per week as well.

That's why I went to that idea of think about static as something that you put in a bucket, and Vercel is the process that automates that process. Don't think about like, "Well, there's always a server there that needs to be hit," because you nailed it. When you're doing this kind of background computation as a result of events, as a result of your data changing, we now generate a static HTML, and now we can put it in a highly available bucket that we're also able to distribute around the world. From a perspective of performance and availability, there's this idea that now, every time you go to the page, no computation ever happens.

We're just manipulating very basic static objects. The fact that you can run incredibly large websites with just this primitives is very reassuring from the DevOps perspective. Every time I talk to people, and I tell them that whatever they invested in that was a server could have been static, there's always this incredible desire to go toward that kind of place. Even if AWS has done this incredible job where Lambdas have 99.99% SLA, and every system that AWS monitors is automatic and they have an incredible track record for reliability, you still want to pick the architecture that has basically the fewest number of moving parts.

You can think of code as a moving machine.

Jeremy: When I think your point about latency, I mean, you mentioned you're not going to cache somebody's banking records all across the world in case they happen to log in. I happen to be in Europe. I want my banking records immediately. There is a certain amount of latency that obviously is acceptable depending on what it is that you're doing. I don't know if you've given this example before. I think we've talked about it before in the show. For every 100 milliseconds of added latency, Amazon.com loses 1% of their sales or something like that.

Latency is a much more important number that I think a lot of people give credit to.

Guillermo: Yes. Let's take that example, that exact figure. There's this famous internal memo from Google about numbers that every developer should know, and one of the key numbers is the coast to coast Netherlands to California. It's somewhere around the ballpark of 150 milliseconds just to do the complete round trip. We've improved our routing networks so much that we've optimized California to the Netherlands so much that it's just close to the physical limit. When you think about incremental static generation and putting pages next to customers, you already have there...

You mentioned you from Montreal, you said 25 milliseconds. We've already have a leg up of 125 milliseconds. That is unsurmountable for the traditional serverfull or function plus CDN case that event sometimes has to go to origin. It's absolutely unsurmountable. Then we don't stop there, because we make massive investments also in the Next.js layer, for example, to make sure that that content also when received by the web browser renders as soon as possible. We have several integrations with Lighthouse. Now, we ship the integration with Chrome web vitals. That allows the developers to measure the time to the first content full paint.

This is where I want to stress that I want to delete all the code from the world. I don't mean no code in the webflow sense. I mean, no code in that... If you have a stream of HTML coming in from the Netherlands, and it's some product that you want to buy, then when the browser starts interpreting it, if it has to boot up into the V8 VM to start executing JS code, then you're going to waste another 100 milliseconds for sure. 100 milliseconds, V8 loading JS and starting to interpret it in 100 milliseconds, that sounds like Nirvana. You know what I'm talking about?

If everyone goes to their terminals right now and they run "time npm --version", you can see the V8 warmup time in real time. You'll see it. I'm going to run it while we're talking now just to get my own measurements here.

Jeremy: Sure.

Guillermo: "time npm --version". What's happening here? We're executing node, which is booting up V8, which is executing a bunch of JS. That code hit that I just performed in my machine, which is also a V8 because it's executing the stream that we're doing, 812 milliseconds. That just sounds insane, right? 812 milliseconds for V8 to boot up, MPM's code to get parsed and compiled and returned back to extend the route. I run it again. The universe seems hotter, is 174 milliseconds. Still awful, right?

This is why I don't want to run code. I don't want to run it at the edge. I don't want to run it in a worker. I don't want to run it in a function. I want the function to have been executed at some point in the life cycle, but not in between my customer and the page. Then I want to have as little JS as possible also when that page runs in the web browser. This is why we started with why Vercel is focusing so much in the front end. I suspect that a lot of your audience, my audience, also sometimes over indexes in measuring back end, measuring Dynamo latency, measuring ELB latency, and then they forget that there's this universe of complexity that we're shipping to the web browser, that is adding that 100 milliseconds that you just talked about times 10.

It's not even a couple 100 milliseconds like what we just talked about with that MPM version exercise is saying, "The best case scenario of blocking your entire page on JS booting up, we're talking about downloading the JS." Hopefully it's a hit from a cache. Hopefully it's a hit from a cache on the local computer, which by the way, a fantastic essay just came out that we all know especially all of us that have worked extensively with AWS, the disks are pretty slow. IOPS are expensive, but also the distinction between a hard drive and a SSD, and what Google Cloud calls the local SSD like the one that's wire right into your instance.

We're talking about lots of milliseconds there, even in your web browser, retrieving a cached CSS asset and JS asset from the local computer. This is why we're so obsessed also about edge precomputation. We don't even trust, and we have the data to back it, that even if you have a stable JS resource and CSS resources that's been cached from the device, we don't even think that we can afford to revive that asset, bring it alive, and execute it very quickly. This is why I'm now endeavoring towards getting rid of computation, because if I can give my customer a stream of HTML, and inline CSS for the critical parts of that page, then what I end up is with something that can actually rival the performance of amazon.com when it comes to their own products, because what do I get?

I get from the edge. I get the precise image for the size of the device that I'm serving for the product that I want to buy. I get the styling for the buy button, which is what I want my customer to press. I get that first paint in 100 milliseconds. We've altogether removed JS from the equation. JS is not being executed at the edge, is not being even executed by the page on the local device, so we can make that dream happen of your product is in front of your user's eyes in 100 milliseconds.

That's doable even for 2G connections or legacy Android devices. It's totally possible. It's just that we really need to shift our thinking and our obsession from back-end architectures and AWS charts connecting one million pieces into now thinking about what we're serving to our users. That is a big shift that's happening.

Jeremy: I love that idea because I think there's been a movement lately of sites that run completely JS-free. You do not need JavaScript to make your dropdown menu work.

Guillermo: Absolutely.

Jeremy: You don't need it to place an order. You don't need it to do a lot of these things. It's funny, but they invented this thing called HTML and CSS that allow you to do a lot of really interesting interactivity without using JavaScript. Now, obviously, Vue.js and React and all these other things add really cool features, and if you have it enabled and once you've got it loaded on your machine and interacting via the APIs, that's great.

Guillermo: That's why with Next.js, we're giving you that, but we have to find that balance. You're right, for the first paint of most of the pages that you visit, there's very little need for JS, but even when it's needed, it has to be consumed in small amounts. This is why one of the big hits that we had was that we eliminated the idea that you have to configure bundler and webpack when you use Next.js, and when you produce these pages, because that's when things start to go really bad.

We want the bundler to be so ingrained into the system. Let's say that that buy button, when you press it, you do want to use some JS because you think that, for example, just like Stripe does with their credit card modal, you think that you're a PM at ecommerce company, and you think, "If I transition them to another page, and if I didn't memorize their credit card details, and if I didn't auto complete their credit card and whatever, we're going to lose sales, right?" That's a valid argument, but that's the point is that JS needs to load just in time for that specific interaction that's going to happen.

We have to be smart so that we bundle JS in minimal amounts only for the interaction that's likely to happen, which in this case is buy. Maybe as you start scrolling, it's loading product recommendations or loading the carousel of related products. That's why static generation also always gets combined with loading strategic amounts of code on the client's side with JavaScript so that you can bring interactive experiences. Another example is, I believe, I've seen that Amazon loads some 3D navigations of some of their products. You don't want that bundle or that feature, which is this long tail feature that maybe some customers use for that to be blocking what we call the time to interactive of the buy button.

That's, by the way, what's happening to every website of every visitor of most every website in the world today is that the bundler is... We talked about that MPM example that my computer took 800 milliseconds. What's happening there, for most of the seasoned JS optimizers in the audience, you probably know this, is that V8 is crunching to vast amounts of code that have nothing to do with rendering the version. MPM is not being bundled in such a way that all we did was execute, return version 9.2.3.

Instead, we're probably doing lots and lots of unnecessary stuff before. This is what happens with the web at scale today is that in order for that buy button to become interactive, we're still downloading the code for the 3D navigation carousel. Again, this is what, I think, makes it really, really compelling for companies to stop thinking so much about full stack development, and instead focus most of their attention to this kind of problem, start measuring what actually happens when their products are being delivered to users.

Jeremy: Alright, so we talked a lot about front end. I think there are so many optimizations that could be made there. There's so many cool things that we can do like the pre-rendering, the pushing stuff to the edge, getting those latencies down, not trying to load stuff from local cache even. I mean, there are a million things to think about there. In a perfect world, we could do everything we wanted to do that way, but the reality is like you said, sometimes we have to load, pre-load credit card data or something like that for a user.

Obviously, sometimes we're going to have to make an API call. We're going to have to access the database or DynamoDB, but as you said, how much of that calculation do we want to be doing ourselves? How much of it can we do directly from the front end? Vercel's got the serverless function capability, and I think you had mentioned this as a use case where maybe you're doing authentication and you need to send in some private key into Auth0 or something like that to trigger the workflow.

Guillermo: Correct.

Jeremy: That's something where, I think, you'd advocate build a really simple function that just triggers that thing there, but don't do all that calculation yourself.

Guillermo: Correct. Yes. Basically, I think the... You mentioned... Okay, how do I execute my functions? Do I execute them as a result of pipelines like SQS, SNS, whatever, or do I execute them just in time? Both of them are absolutely amazing, right? The executing just in time case throughout the years has had ups and downs, let's call them, because we know that P99, you have to be smart about the size of your function. You might want to provision your function for very, very, very amazing P99. In some ways, I feel like the background async computation usage of functions has been the home run, and the just in time has been more of a incremental adoption.

It's definitely going to happen, but it's been more of an incremental path for a lot of people. The one that we've found that is awesome in the just in time space is giving the developer team, especially the front-end developer team, a way of gluing services together. A great example would be, what you mentioned, is, "I have to create a function that talks to API systems that already exist, because I want to aggregate a bunch of API calls together, because I want to talk to Stripe, for example, in a private manner with an authentication token from the user."

Let's say they want to commit a charge. Stripe gives you products that you can invoke directly from the front end, but you start hitting some limits at some point. You want to customize their UI. You want to do something more fancy. Maybe you have some recurring charge. This is where, I think, the world of serverless infrastructure is getting really neat, because we don't have to reinvent Stripe from scratch. We don't have to reinvent Auth0 from scratch or Cognito or whatever. We can now use functions as a way of mediating between the front end and the services that already exist.

That's not to say that you can't use this function to talk to Dynamo directly, for example, right? That's still a use case that'll exist, but I think what's more compelling for a lot of people is to not necessarily have to reinvent the wheel and be smart about their investment into these functions, because the difference broadly between functions and also in pre-generated content is that functions now have an on demand cost. You have to be careful about their availability as well.

Now, your uptime for those functions depends very much on the uptime of the services that you depend on. We talked a lot about, Okay, how do you even ascertain that those services that you're depending on are functioning correctly, and why it's so much more appealing for you to be interfacing with a high level API like Auth series for the users table rather than using Dynamo as your users table, right? Then you start worrying a lot about rate limiting. You start worrying a lot about, we talked about, Dynamo auto scales until the end of time.

But do you actually want to let people have this function that is invoked on demand by themselves endlessly, and that now you scale with them for both Dynamo and the function? Things can get hairy, so function I see as this important tool in the tool set that has to be used when it's necessary to use, necessary from a data consistency perspective, as we talked about. Another need by the way that's very strong, no pun intended, is when we talk about precomputation, we're talking about vast categories of pages that are public in nature.

We gave that New York Times example, that amazon.com example. Of course, you would want to push those pages to the edge. They're all public, and they're all shared by lots of users. Even if they have some, what I call, page variants, which is something that we're incorporating into Next.js, which is like, "This page will be in a certain language for the Netherlands, and this page will have a built-in promotion for Texas." Those are what I call variants of pages. But what they all share in common is that they all address vast numbers of users, and that from a security perspective, they don't contain anything that is sensitive.

Now, let's think about product recommendations that are user-personalized. Now, let's think about your order history. Now, let's think about your credit card, inputting your credit card and whatnot. Those are all things that no longer fit that neat world of precomputation from many perspectives, one of them being that it becomes prohibited from the explosion of combinations of precomputations that it could make, but also security.

Jeremy: Sure.

Guillermo: Again, I don't want to cache personally identifiable information at the edge and from a data strong consistency perspective. Again, I don't want to go to stale cache for performance and availability reasons. I want to go directly to the data source. That's where functions come in. Now, some of those functions will be written by your team. Some of those functions will be assisted by other teams, because those functions can call to all those other functions, and some can even go directly from the front end.

Those three are all amazing, legitimate use cases that get enabled by executing code on the client's side.

Jeremy: Right, and I think that when you start talking about Cognito, and you start talking about Lambda functions and DynamoDB, there are a lot of primitives that exists in the cloud right now that you can stitch together. As you said, functions are great for gluing these things together. There's other ways to glue these things together. Even though there are these amazing primitives out there, though, it doesn't mean that building a serverless back end is easy.

Guillermo: Correct. Correct. I think what's amazing about serverless is that it's exposed the essential complexity of the problem. It stopped developers from sweeping hacks under the rug. The best example that, I think, from this is you can no longer do async computation as a result of invoking a function that easily anymore. In the world of no JS, I would see a lot of customers just put lots of state in a process. When they respond, they continue doing things behind the scenes in that same process.

Functions have altogether made this impossible, but for a great reason, right? They were exposing, "Hey, that side effect that you were computing, you should have not been doing in that same process. You should have used a primitive like a queue to put your side effect, your event there, and then use other functions that respond to that event. Then it's so smart that they also put the developer into this state of success of saying, "Well, if it's a side effect that now can no longer be retried by the client," because the client is executing the function.

The function is responding with 200, and it queued the side effect, so there's no reason for the client to retry. Now, the side effect is loose in the universe of computation. That means that we need a system that can retry it because we want that side effect to run to fruition. Now, it forces you to put that into a queue, and queue can retry and then eventually also fail and go into a dead letter queue. So now just like going through all this in my head, I'm going crazy about the amount of complexity.

But here's the thing, and this is why I love serverless. That was to begin with the essential complexity that had to be managed to begin with. What we were doing before was chaos, was side effects that maybe sometimes run correctly and sometimes not, was unscalable systems and so on and so forth, but it is a complicated world.

Jeremy: Yeah, I mean, and I think you mentioned this idea of scalability. That's one of those original promises of serverless like just everything can scale infinitely. You know what I mean?

Guillermo: Yeah.

Jeremy: I think you mentioned the DynamoDB can just keep scaling up, but there are limits. Eventually, that does stop. There's soft limits in place for Lambda functions, and of course, there's wallet limits, I think, for anybody out there who is eventually going to say, "This is more than I want it to be." You mentioned on-demand serving of data versus the static piece of things. You had shared an example with me before about a site that was mostly front-end serverless, had a little bit of back-end serverless, and even though it scaled up, I think it was 10s of millions of hits or something like that, that it was able to scale gracefully.

I'd love for you to tell that story, because I think that is the perfect example of how we should be thinking about building serverless applications, because even if you think your application will scale infinitely, there are a lot of reasons why it will not.

Guillermo: Yeah. I love to give this example because on one hand, we deal with customers that have very traditional websites that fit under this incremental static generation umbrella. Our most recent we onboarded a couple weeks ago is barstoolsports.com. They fit under this category of the New York Times website that we're talking about. Their business is assisted by subscriptions and ads. They want those conversions to happen very quickly. Frankly, they want their customers to get their news as soon as possible. This was a larger team.

They were proficient. They had already chosen Next.js, but I'd like to give this other example that happened that same week of a meme that went absolutely viral throughout the entire internet, and where the promise of serverless really came to fruition in a very... I mean, again, this meme is weird. It's called billclintonswag.com. What's incredible about this meme is that you would go to Twitter. We noticed a spike on our edge network because we get alerted when there is abnormalities, and this was a very large abnormality that normally we would confuse it even with an attack, because it was like, "Holy Moly."

We went from this little thing went like this, and which went completely vertical. This meme is very basic. It's a static page that presents a photo of Bill Clinton holding three albums. Then the visitor can auto complete. This is where clientside JS comes in, auto complete and find their three favorite albums. Billclintonswag.com, it should obviously still be around, but the meme has faded a little bit like the serverless spike patterns. That's when computations are happening on demand.

Going back to the glue pattern, the one individual developer that... By the way, again, this is a meme that is quite controversial to begin with, so that's why he got all this traffic. Think about this. You can create with one person a thing that dwarfs the traffic of most websites on the internet for a very short amount of time in this case. We're talking about 10s of millions of hits per day. It is scaled infinitely, and it was created by one person. But here's the thing, he could afford this because he designed it with this static first mindset.

There was no function that was getting executed when people were first going to the website. Then he didn't write a database of records. He was using the Last.fm API. Now, guess what, the Last.fm API for whatever reason was not consumable from the front end directly. What he did is he created a serverless function in Python, also hosted on our platform, with 128 megabytes of memory, so very lean, also caching at the edge. You mentioned this earlier, by the way, but what he discovered was that these meme makers all like the same albums.

We're talking about a very large number, so we're talking about trending topic on Twitter large numbers. They were all auto complete into the same things. You know what I mean, like Beatles? Well, actually not that one, but like, I don't know, who's the famous new... Kendrick Lamar. You know what I mean? K-E-N. The function was lean. The infrastructure was actually Last.fm. He also added his own caching at the edge of... The query parameter was K-E-N for Kendrick Lamar, so he was responding from his serverless function with cache control.

We're caching that at the edge as well. Meaning that you are getting the suggestions for albums for Kendrick Lamar in milliseconds. It was really amazing to see how this combination of what a lot of people call JAMstack. The first paint is a static. It affordably went to hundreds of millions of hits that week. Then when he needed to do a little computation, he was careful to pick the best provider for his data source that exposed an API to begin with. Then he also cached the computation of that.

It was honestly amazing because with these two pages, he navigated the entire world. If you think about doing this with the old patterns that we had, he would have collapsed several times over, or have been incredibly expensive like we talked about, because if he had made the landing page a function, it would have been a pretty hefty bill. If he had made it a server, it would have not scaled very quickly, not very easily. Frankly, the ease at which he also put this together was just a great validation as well.

He happened to be a Python developer, so he chose Python for his functions. A lot of our customers use Next.js, where index.js is your react page that gets built statically, and then you have your functions on the side. It was like a micro example of how... I think we're going to see this a lot at scale in the future, where very small teams that use primitives in a very smart way for front end can now go from experiment to world takeover. Again, this is a silly meme, but this is the same thing that can happen to your product.

This is the same thing can happen to your news story. This is the same thing that can happen to your ecommerce site. In fact, this is also a micro story of that because he was selling some product as well. Going back to the 100 millisecond thing, he created a delightful experience that was very fast to load. Then he was selling something as well. This is the whole story of the internet. Publish fast. Make it accessible to everybody in the world. Have a great success story, where your back and your front, everything is actually working, and it's not collapsing.

Then find that fitness function that allows you to evolve your business. Like I mentioned, for a lot of the publishers of stories that have chosen the Vercel edge, for a lot of them, it's optimizing so that they can render an ad or a paywall or whatever it is that they need to render to make money. What you need to think about is like, "Okay, what..." You mentioned that 100-millisecond rule. There's so much greatness that went into the development of amazon.com, but for me, that was one of the biggest ones is the realization of the idea of latency in correlation with business success.

Jeremy: Yes.

Guillermo: When we talk about serverless primitives, you can't just over index on wanting serverless if it's not really enabling that business success for you. When you think about 100 millisecond, 100 millisecond is just what it takes for a function to cold boot in a really, really good case, right? It's probably more than that. I think this is what the big lesson is, is that you have to really shift around your mindset to that customer finding you whether it's... in this case it's Twitter or LinkedIn or ads that you're buying, and then how quickly and effectively can you go to that page?

Jeremy: Yeah. It's funny, I can verify the first time a few years ago that one of my articles made it to Hacker News, it immediately killed my WordPress installation and knocked that down. It's funny because I feel there are so many people now talking serverless first, serverless first, serverless first. I love that idea. Yes, serverless first, but static first is a really, really interesting twist on that, because I think that... I mean, if you just think about that for a second, it's like, "What can I pre compute? What can I put on the edge? How can I reduce the number of computation that I need to do?"

Then anything I have to go beyond that, how do I build the infrastructure as cheaply as possible without reinventing the wheel?

Guillermo: Yeah, and also what's more serverless, right?

Jeremy: Right.

Guillermo: That's a big thing.

Jeremy: Very good point. Very good point. There are a bunch of other things that I had wanted to talk to you about, but we've already spent quite a bit of time. So the last thing I want to touch on simply because we're in this new COVID-19 world, and we've got all these people working from home, and we've had all these big companies tell us in the last few years like, "Oh, working from home doesn't work." I know that Vercel is a huge fan of this. I know you came from several companies that were also a fan of this, so I'd love to get your take on distributed teams, and how effective those can be.

Guillermo: I think we're very lucky to be a distributed-friendly company, remote-friendly company before COVID happened. We're so well positioned to provide the best tooling possible to people that now are in this new same position. We dog food our product very extensively. I love what you just said about being on top of Hacker News because yesterday, Deno, the new JavaScript runtime was at the very top of Hacker News in a very top way. I think it's one of the most upvoted things on Hacker News in a long, long time.

Their website is not a Deno server. Their website is a Next.js website precomputed at the edge, and hosted in Vercel. How did they choose this? How did they start using Next.js, learn it, they created their website and import it into Vercel. Well, we never talked to them. It just happened. It's because we're out there designing this distributed systems of collaboration, especially with the open source community, people are using Vercel every day without even noticing it, because they go to GitHub. They push to a repo, and guess what?

That repo is connected to Vercel, and it automatically builds and deploys your website at the edge. Then you get back a URL. We're enabling this workflow for teams all over the world to collaborate in some cases without even the need to teach each other anything. Somebody comes in a team that is using GitHub, installs the Vercel app. Now all of a sudden, their website is getting built for every push, and you get back your deploy URL. Now, you can share that deploy URL on Slack, on Zoom.

You can give that deploy URL to your end-to-end testing service, so now we created this new world of it's a distributed workflow where the primitive is the URL to the front end that you're building. It's inherently incredibly shareable. It's obviously fast for everybody in the team, right? This is something that we always obsessed about a lot was like, "Hey, we're building our website with Vercel. It better be fast for everybody in the team, right?"

Our own team is in Japan, and it's in China behind the firewall. That is an amazing fitness function, by the way, I have to say, because if you try to use the internet from China, you better be very, very lean, static, precomputed and cached right outside of China, because a lot of people escaped that firewall through Singapore or Hong Kong or Tokyo. We were so well positioned for this world. Now, I'm not celebrating COVID, but I have to say the week that COVID was... The stock market was imploding.

We saw the biggest peak in creation... We measured this particular metric which is builds, builds of this pages. We have built concurrency. That is quite a sophisticated system, and builds obviously take a lot of CPU power, so we're constantly monitoring it and so on, but that week of COVID mayhem was our largest week in build concurrency. That same week, the CEO of Slack shared that he also saw his biggest week in a number of concurrent connections to Slack, number of WebSocket connections or whatever concurrently to their servers.

So really, what's happening is not a recession like a lot of people think, is really an acceleration of teams now being exposed to the right primitives to publish pages to the internet faster, to build pages faster, to collaborate faster, to collaborate from their homes. My daughter just interrupted a meeting, but that's okay. Soon she'll learn what I'm talking about. There's all these advantages to this new world. I'm excited about it.

I'm excited about giving our tools to everybody that needs them that can find themselves in this world that it's harder to collaborate in just because you're not used to it. It's not hard to be productive in a distributed manner and a remote matter if you install the right tools.

Jeremy: Totally agree. Awesome. Well, listen, Guillermo, thank you so much for spending the time with me and talking about front-end serverless, because I think this is something that is not on a lot of people's radar. They're not thinking static first, and it should be static first. Serverless second maybe is the new way we should think of it. Anyways, if people want to get a hold to you or find out more about what you're doing and what Vercel is doing, how do they do that?

Guillermo: For sure. About me, on Twitter, @RauchG is my handle. I talk a lot about these topics, so it may be entertaining. Second, a lot of people asked, "Okay, I love these abstract ideas that you talked about, static first, whatever. How do I put it into practice?" Next.js is our open source framework. If you go to nextjs.org/learn, we'll walk you through creating your first page of this manner. Then a lot of you are more advanced in this trajectory. They already have... They're already using Next.js, already using Gatsby, Vue.

They're already making single-page applications. They're making... They're using static site generators, so they can deploy them to Vercel. I think what really sets Vercel apart is this global distribution of pages that is faster for the customer. Builds are much faster than you trying to do this yourself. I talk a lot with customers that they've created versions of this with complicated pipelines from GitHub to Circle to S3 to CloudFront.

Sometimes they purse. Sometimes they forget a purse, sometimes they cache their CDN. Sometimes they use their CDN as a dump pipe. What I tell them is try importing your repo into Vercel. It might simplify your life quite significantly. It might speed up your visitors quite significantly. That's the three things I recommend.

Jeremy: Awesome. Alright. Well, I will put all that into the show notes. Thanks again, Guillermo.

Guillermo: Thank you so much.

THIS EPISODE IS SPONSORED BY: Amazon Web Services(Serverless-First Function May 21 & 28, 2020) and Stackery

View Details

About Jared Short:
Jared has been building and operating serverless technologies in production at scale since 2015, and is laser focused on helping companies deliver business value with a serverless mindset. Jared is currently Senior Cloud Engineer, Developer Accelerator, at Trek10, Inc. but was formerly Head of Developer Experience and Relations at Serverless, Inc. and an early contributor to the Serverless Framework. In his current role, Jared's day-to-day is serverless all the time, as he helps people build and operate cloud native architectures.

  • Twitter: twitter.com/ShortJared
  • Email: jaredlshort@gmail.com
  • Website: jaredshort.com/
  • 3 Guiding Principles for Building New SaaS Products on AWS: trek10.com/blog/guiding-priciples-for-building-saas-on-aws
  • 3 Big Things I Wish Someone had Told Me When I Started Using AWS: dev.to/trek10inc/3-big-things-i-wish-someone-had-told-me-when-i-started-using-aws-2d0n

Watch this episode on YouTube: https://youtu.be/rA4eVtpFnVs

Transcript:

Jeremy: Hi everyone, I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Jared Short. Hey Jared, thanks for joining me.

Jared: Hey, pleasure to be here. Thanks.

Jeremy: So you are a Senior Cloud Engineer and Developer Accelerator at Trek10, Inc. So why don't you tell the listeners a little bit about your background and what you do at Trek10, Inc?

Jared: Sure. So my background, I think starts similar to a lot of people, where I dabbled in the basement on the all the Apple II, I learned how to program actually from a book from the library on that Apple II. And then throughout college... Well, high school and college kept keeping up with technology and building things and exploring and learning. And eventually that led me to kind of the cloud back in 2014 or so. I was big into Docker in the early days, in the cloud, and eventually found serverless while I was at Trek10.

So Trek10 is of course an AWS consulting partner. And as part of that, I get to help companies design and build serverless and cloud-native systems, with different kind of verticals all over the world. SaaS companies, enterprise companies, all of that kind of stuff. So that's where I'm at today. And I'm mostly focused on helping people learn and understand the cloud through our developer acceleration program. So taking all of those things that I've learned while helping people build things, and now helping people just learn what all they need to learn to build successfully in the cloud.

Jeremy: Awesome. Alright, well, so I've been following you for a very long time. I mean, you and I have known each other now for a while. Met up at a few conferences and so forth, and you always do great stuff. So I love the Trek10 blog, love the stuff that you've been working on. You've done a lot of stuff I know with Forrest Brazeal and some other things that have been very popular. There's a whole bunch of great stuff out there by you. So definitely search for Jared Short, serverless and go check out your stuff.

But I saw an article from you a couple of weeks ago. That was the three big things I wish I knew before I started working with AWS, or something like that. And that just struck a chord with me, because as I was reading through these things, I was like, "Oh man, this was the article I wish I had when I started working with the cloud way back in 2009." And since then, it's like exploded a thousand times over. So this is a great article and I'm going to put the link in the show notes, because I do want people to go read it. But I think it'd be awesome to just go through and talk about this article and kind of hit on some of these points.

The article is very in depth that goes deep into some of these things, but this is something that really warrants a conversation. So the first point that you made, the first learning or the thing that you wanted to that you wish you had known, was this idea that AWS is just this massive ecosystem and it's basically pretty much impossible to understand all of it.

Jared: Right. Yeah. It's a massive ecosystem that shows no signs of slowing down. It's pretty similar to the ever-expanding edge of the universe, it just keeps growing and consuming.

Jeremy: It's like, S3 was the big bang and then it just kept growing from that point. Right, right. So you point out a couple of things though about this that I thought was sort of really interesting. Where it's like, there are all of these different services and you had said, you could explain what most of these services do, at a high level. Like what is Amazon Sumerian or AWS Sumerian, who even knows the names of some of these things. You can explain that at a high level, but then understanding the nuances and the limits. And that's like a graduate level course in and of itself.

Jared: Yep. Yeah. Right. And in fact, the fact that I can't even tell you how they name Amazon versus AWS in front of something tells you a little bit of something. Right? I think I would guess it's Amazon Sumerian I have no idea. And the fact that I can tell you a little bit about this, I can tell you at a high level, what it does, is I think you have to know that in many situations, if you're an architect or someone building stuff on AWS. Because you need to know at least which tools I need to go read the docs on to understand if I need to use it, or it could be useful in my particular scenario.

What I can't do, with for instance SageMaker, I can tell you it's their machine learning product and things like that. I couldn't tell you what models are preexisting in SageMaker. I can't tell you what limits might apply to SageMaker endpoints that I've deployed. Things like that. If I were to need to build a machine learning product or have some feature for that, I know I could go look at that, and then I would have to learn those specifics.

And I think that applies to the vast majority of services that exist in AWS. You can certainly know what they do. You might not know how or why you should use them. But knowing the what for the core services, it's at least I think a starting point, right?

Jeremy: So one of the things you mentioned too, is that again, reading the docs, right? This is something that you've publicized on Twitter. And I think it's a brilliant idea, and if only we all had the time to do this. Where you take a different service and you read through all the documentation once a week, which is... I probably should be doing this too. But this idea of being able to read the docs and get a really good understanding of a single service. I mean, obviously there are hundreds of services, and even beyond that, I mean, there's sort of hundreds of sub-services, right?

And like things to understand and then the interconnectivity between them. So what's the suggestion there? Like, do you try to learn it all or do you just pick a few things that's going to work best for you?

Jared: Yeah, I think I would start at least in terms of consuming documentation. I always suggest people start with the stuff that's relevant to them right now that they're looking at, right, so Lambda or S3. S3 I think S3 is the most applicable service probably that exists in cloud today. But I would consume the documentation. Now it's important to realize that there's multiple kinds of documentation out there that exist AWS at varying levels for each service, right?

You kind of have your narrative documentation, which is the one that I think most people read where it kind of goes through, it explains the features, the services, how to use them. You have the technical documentation, which you only read if you're really intending to implement something or low-level API docs, things like that. I would consider the Boto docs, the AWS SDK for Node.js, things like that. And then you have the blog posts, the examples, the explainers, the how tos.

You have to, I think read and synthesize all three of those sources, to be able to construct in your head, a cohesive or nearing complete model of what that service can do for you. And that's what I try to do when I'm going out and reading the docs for a particular service, and that's extremely time-consuming process, right? Even just finding all of that documentation can take a couple hours to really build all of that out.

And I guess my suggestion to folks is like, you don't have to do that for every service, do it for the ones that you use regularly, do it for the one that might make the difference in your particular product. So for instance, if you're doing machine learning, I would go do that for SageMaker, to understand if it meets my use case. And if it doesn't necessarily meet my entire use case, what problems am I going to run into, or limits am I going to run into? And understanding those makes it easier to build to your use case.

So I guess ultimately look, nobody has the time to read all those docs. I get that. But I think the investment upfront is absolutely worth it for those core services. And you'll just have an easier time later on if you're willing to do that small upfront investment.

Jeremy: Right, now and one of the things about investing time in anything outside of actually programming or doing something that is maybe making money for the company you're working for, obviously is this learning time is a huge investment, right? So digging into some of these document... You know, the docs and going through like you said, three sets of docs, probably for every service that's out there. Plus other like non-specific AWS affiliated stuff where other people are writing examples and things like that, open source libraries. What effect does this extra learning have on not only developer productivity, but maybe on like team productivity in general?

Jared: So I think it's compounding. I think the more folks that you have spread across various services, you quickly establish subject matter experts, right? SMEs. Even inside of Trek10, for instance, we have something we called the SME matrix, where we all kind of go through and individually rank our familiarity and skill level with particular services or technology. And ours is on a scale of: have no idea to have multiple production systems I've built that are using that service, right?

Like three in the middle is I've read all of the documentation, I understand it, I've experimented with it. Finding people that are fives in stuff that's outside of a Trek10's core competency, right? You can expect if you were to put a five on Macy or something crazy like that, we would be going to that person right now and saying, "Okay, tell me everything you know." Right? And having those people on your team and kind of distributing knowledge of more niche services across your team is super valuable.

Being able to go to somebody and say, "Tell me about API Gateway WebSockets, because I didn't want to build a real time product. Having that person that's invested that time, that's actually built something is invaluable. The 30 hours, the 40 hours that they spent reading documentation or prototyping, something I think pays back exponential dividends over the course of history.

Jeremy: Right.

Jared: It takes time to make back that investment, which I understand.

Jeremy: Yeah. And I think this idea of SMEs is a really, really interesting concept when you look at an ecosystem like AWS. So you have mentioned in the past this idea of the T-Shaped Engineer, right?

Jared: Yep.

Jeremy: And I'd love for you to explain that because I think that, that might take the pressure off of some people thinking they need to learn everything as deeply as possible.

Jared: Yeah. So, so the concept of a T-Shaped Engineer is really go wide on many topics, right? If you're coming into the AWS ecosystem, you're not going to know a ton right off the bat, but it's valuable to go wide in a lot of topics and gain a wide breadth of topical knowledge. Understanding what Lambda is and does, and maybe how to deploy Lambda functions, SAM, serverless framework, things like that. Understanding this is how they work. This is what they do. Understanding EC2, understanding S3, Dynamo, RDS, all this stuff.

Understand a fairly minimal level, just what they do, how to use them, but then that's the top of your T. And then you have the vertical where that's your specialty. That's the thing you're really good at the thing that, you know the limits of DynamoDB On-Demand capacity will work for zero to X amount of spike, right? Zero to 12,000 will just work. If you go above that in the span of 60 seconds, you might have some throttling or something like that. Right?

Most people don't know that but having those people that do know that for their prospective services, and you as an engineer makes you invaluable in that particular vertical. And then having that topical knowledge, really just helps you understand what questions you need to be asking of people as you start investigating other areas.

Jeremy: Right. And that idea of going really deep on specific things. I think this is where, sort of the next point you brought up in the blog post was, is that understanding a service's use case versus sort of just the service itself or the baseline, like you said, those right questions to ask versus exactly how they work. That's another thing that you sort of pointed out is that for every service out there, there are probably a ton of different use cases. And if you just learn one service, then everything starts looking like a nail.

Jared: Right. Yeah. I think part of that is the consequence of AWS being very good at building service primitives. Right? You have SNS, SQS, Kinesis, DynamoDB streams, EventBridge. I can make them all do very similar things.

Jeremy: Right.

Jared: The question then becomes what are the constraints of my business systems or my business logic, and that's how I start picking the service that actually fits the best. Right? Does my event need to go to multiple people? Well, okay. Maybe that's SNS, SQS probably won't work anymore, because that's pretty much a one-to-one. Does my event or a thing that I'm publishing meet ordering guarantees. Well, if I need ordering guarantees, I'm pretty much stuck on Kinesis or something like that.

But those are the questions that going and asking those subject matter experts or things like that... Like, I wouldn't expect somebody to know that if I had strict ordering guarantees, that Kinesis is pretty much my only option. Unless I need SQS FIFO, which is now an option, right? That wasn't an option until very recently because it didn't have a native integration with Lambda.

Jeremy: Right.

Jared: That of course evolved, but...

Jeremy: But speaking of evolving, right? Things that change over time, I mean, that's one of the things that's nice about using any of these services that even if there are limits in place, even if there are other types of limitations that prevent you from using it for certain use cases, eventually over time, it seems like these products just keep getting better and new features are added and next thing, you know, you can use it for some other use case.

Jared: Yeah. And I mean, that's I think part of the beauty of building on the cloud. There's very few other instances or circumstances I can think of, where building some piece of technology and letting it sit for a year or two, means it gets better and cheaper without touching it. And we've legitimately had, at Trek10, systems that we've built, where we kind of just let it sit. We had built it, we might do some occasional dependency level package management, you know, addressing CVs or things like that.

But largely, we just left the architecture untouched and we would go and look at metrics. And over time the metrics just got better. It got more performant, it got cheaper over time and that's kind of weird. But that's one of the benefits of building, I would say, not just serverless, but cloud-native in general you know, Macy just got cheaper by some crazy number, things like that, right? If you're already using these technologies, they tend to get cheaper in most circumstances.

Jeremy: So the other thing, I think that's interesting, especially we put this in the context of things I wish I knew. These service limits exist and there's been a very, I don't want to say a... I guess, a dogma in a sense around this idea, "Oh, well, serverless is infinitely scalable. And if I use this service and I use that service, then I can just scale up whatever." But there are soft concurrency limits for Lambda, there's Kinesis shard limits, there's throughput limits for, you know DynamoDB and these other things.

So there are a lot of limits in place. And I guess my frustration around this is you really do need to go deep on some of these topics in order to understand what those limitations are. And then it really hurts when you hit those limitations in production. And you're like, "I didn't even know this existed." I mean, that might be on you, but is that something that cloud maybe needs to do a better job of? I mean, not just Amazon or AWS in particular, but other clouds, do they need to do a better job at managing these service limits for us?

Jared: I think it's fair to say that those limits are not put there maliciously. They're put there to stop people from doing silly things and protect, I think the broader ecosystem in the cloud. Right.? There can still be a noisy neighbor. Like if somebody would accidentally infinitely call a Lambda function, that recursively called itself, I feel like letting someone eat up the entire capacity of Lambda in a particular region, because we didn't have those established limits, would be a scary thought that somebody could do that either intentionally or on accident. But you have to pick some arbitrary number.

Jeremy: Right.

Jared: And I think the cloud... Many of the providers could do a better job at providing guide rails while still managing where those guide rails are. And I think in a lot of cases, they do do a pretty good job at that with certain services. And I'm of course speaking from the Amazon perspective, the AWS perspective, I can't speak too credibly to most of the other clouds. Once again, I kind of know what the other clouds do, but I know Amazon really well. But I think it's absolutely fair to say that the cloud providers could be doing more there. And I think that's just, we're all learning this stuff this as we go.

Jeremy: Right. We're like learning those limits as we go, sort of.

Jared: I would argue that the cloud providers are learning where that number is, right? The default limits for Lambda have been lifted, I think at least once, maybe twice. And default limits change as we're discovering. "Okay. It's most valid use cases probably can get up to this amount." And they do good jobs at, at least providing the support that you need, if your cases are scaling beyond those defaults. They don't do a terribly proactive job in my experience of some of that, but they do a very good job in most cases.

Jeremy: Right. Well, one of the things I like what you said there is that they didn't put those in place maliciously. And one of those reasons that they're there I think is to protect against costs, right?

Jared: Yeah.

Jeremy: And like out of control costs. This denial of wallet thing that has sort of become a standard saying, I guess, in the serverless community, I think is very, very real. Right? And that's actually the third thing you said you wish you knew, is that cost is just really, really hard to understand. And not only just understand what certain things cost, but that dynamic flexibility, like what effect that has on the procurement departments and the business decision makers that have to understand these costs.

Jared: Yeah. Yeah. I think there's a couple points to that and it's really hard to do. Like people like to do apples to apples comparison. Well, I can buy this particular server or some similar server that could run Lambda functions like this, or whatever, put it in my data center. And my capital expenditure is going to be like X, right? Versus Y for the cloud, and X is less than Y therefore clogged up. Right? I have yet to see an actual good calculation, because it's so darn hard to quantify, not just the server expense.

Sure. That's easy. Like apples, apples, you can get close sure, but the power, the electricity, the real estate, you can kind of get put numbers on that. Sure. You can get close, but then you say, "Okay, so what's the opportunity cost of someone having worked on building out that server or putting stuff on that server, patching that server, versus what features could they have built if they weren't working on that? Right? What's the opportunity business cost there? I have no idea. Right?

So I wish we could put numbers on that. And then I would imagine, I don't have any really anecdotal evidence, let alone empirical evidence, that shows that serverless is of course always cheaper.

Jeremy: Right. And I think you've got this thing too, where what I really like about the fine grain billing aspect of serverless, or just as the cloud in general. I mean if you're just using EC2 instances. You know, with pretty good certainty, if you measure the number of users that 20 of these servers can support, or how many invocations we have on a given month or given week, and the number of users that supports. Being able to take those numbers and break them down, and then being able to see, "Okay, every time a user is in our system and they do X, it costs this much money."

And knowing that even though that may be variable, I think is really handy to know. Because that's some of the stuff you would never know if you were either on-prem or even in some cases just using virtual machines.

Jared: Yeah. No, I think it's definitely a point I've heard actually over the past couple of weeks, in really interesting conversations with people that are trying to get essentially to this idea of zero unallocated costs within their bill. They know for every penny that passes through their billing framework, try to understand which tenant or which feature that that is allocated to, which is very cool. And it's I think, you know... That's a hard problem to solve.

Just, "Okay. AWS billing or any cloud billing," that's a hard problem to solve. Yeah, I will say one of the most interesting things I have seen, I think my record right now for crazy Lambda spend, which I wish we would have had some of those soft concurrency limits we had raised them for other testing in an account. But I think the highest I've seen was like $12,000, like an hour. Because of like an infinite recursion thing. And we're like, "Whoops." Right?

Now to cloud providers credit, most of them are pretty good about being like-

Jeremy: Right, yeah, we know you made a mistake, yeah.

Jared: No sane person would do that. We're sorry. Here's some money back. Right?

Jeremy: It would be nice for some of those other things, those out of control cost things, where even if the limits are raised, that there'd be something that would detect that and would be like, "Well, hang on. This probably isn't right." So then the other thing about this though, is we're talking about costs, right? Which is funny because if you think about most developers, they're writing code and they're uploading into a server or checking it into their code repository, and then somebody else takes care of it.

How much it costs to run that code is generally not a huge... It's not the main focus for a developer, but you shift to serverless. And then all of a sudden, it's like the choice between using Kinesis and using SQS, or SNS, where there's a big cost difference depending on what that scale is. Right? If you said, "Hey." Your boss comes to you and they say, "We need to make sure we do detection in our S3 data, to get rid of all of our credit cards or any PII in there."

And you're like, "Oh great. I'll just flip on Macy" $200,000 later, right? You're like, "Well, wait a minute, maybe we could have done this a better way." So how much should developers be thinking about costs now that they have so much control over the services that they use?

Jared: I think developers should be cost aware. It's interesting, there's kind of been this FinDevOps or whatever you want to call it. And I'm not sure that's necessarily the answer. I think developers should absolutely be cost aware, and at the very least be able to make cost effective decisions. Now that gets hard at scale, right? At a very low scale, Kinesis it's exponentially more expensive, right? It's like a per shard thing where if I'm sending one SQS message per month, or using one Kinesis shard per month, that's not comparable.

Now, if I'm sending thousands of messages per second. Kinesis is probably a much more interesting option, because there's additional features and functionality that also come out of that. So I don't want to say that developer's core job should be cost optimization or anything like that. I think in any sufficiently large organization, you approach the point where having those subject matter experts or architecture experts, or even dare I say, cost optimization experts or cloud economists.

Jeremy: Yeah, there you go.

Jared: Few of those, but I'd argue some of those. I think they are positioned to be able to help developers understand the cost impact of decisions they're making in their architecture. I wouldn't pin that all on the developers, I think that's unfair, but I think sufficiently large organizations should have those people in place.

Jeremy: So what about organizations though that maybe aren't that large? I mean, even small... I mean, I've worked in very small startups before where it's like... I mean, after I finished building the CICD pipeline and writing some new facial recognition thing, I go, and I empty the trash. I've been at that level of diversity of job tasks. But for maybe midsize companies that do have a little bit of a separation, especially a separation between the developers and the technical people, and then some of those business decision makers, is there a good way or an effective way that you know of that developers can communicate costs sort of up that chain to those business leaders?

Jared: I think you can at least take some of those usage costs and say at this order of magnitude, this is what this will cost, right? And you can kind of put some of those rough numbers together. Especially in this granular world of serverless and many of these managed services, you can at least get an order of magnitude and you can kind of get close with the numbers. You can say, if one user is using this, this is what it costs us, if 10,000 users or whatever that is.

But I would say what... And this is kind of like a little bit like the scream test, but it's also my general suggestion to use DynamoDB On-Demand capacity or reserve capacity is just keep using a thing until it hurts. And this is not great advice in the enterprise ecosystem necessarily, because no one's going to come to you and be like, "Why are we spending $10 million on S3?" And you're like, "Well, because we just put everything in normal storage and we're not glaciering or life cycling anything."

Jeremy: Right.

Jared: But you have a pretty good idea of what hurts your wallets. And if you go look and say, "Why does my DynamoDB cost 10 times what everything else is?" Well, we could probably go make a more cost-effective decision at this point. And the reason that I like the, just pay for this thing until it hurts, is it's a moving target, and it also depends on your company, right? For a startup spending, a $1,000 on DynamoDB, might hurt. And we can go figure out how we can solve that problem, for an enterprise spending a $100,000 on DynamoDB, might be perfectly fine. Right?

And the nice part about most cloud-native architectures is it does give developers the knobs and turns and buttons too, for the most part, pivot their architectures or adjust their architectures to cost optimize when they're ready to. You don't have to prematurely optimize in most cases.

Jeremy: Right. Yeah. Good point. Alright. So there were a few other things you mentioned this article too. So those were sort of the three big ones. But you had a couple of other ones you mentioned in here. Just things like good AWS account hygiene, what's that about?

Jared: Yeah. So one of the most frequent things that I'd say we as a consulting partner end up doing, we come into new engagements is, let's do a quick audit of what our AWS accounts landscape looks like. Are we all in one AWS account or do we have multiple accounts? Do we have each environment its own account? Does each product have its own suite of accounts? Things like that. And I would say there's different levels of maturity that organizations can pass through. There's different tools out there.

Organization formation is a pretty cool new one. There's of course Control Tower from AWS themselves who have recognized that this is a struggle. But I would say investing in your core AWS account landscape is probably one of the most core things that you can do to set yourself up for success later. And that's where I really... It's just important. It's so hard to reverse bad decisions.

Jeremy: Right, yeah. Especially once you have things in production, start moving things around, and re-pointing things.

Jared: I'd say it's your best utility for that security blast radius that you have.

Jeremy: Right, right. Good point. Alright, so you have another one in here: follow new products and services. I think that's pretty straight forward. Proof of concept early and often, that's another I think interesting thing. Just this idea of staying up to date with new products and trying something new and seeing if that can help. Also this idea of preparing for failure and regional outages, that's an entire podcast in and of itself. But another one that I thought was really good was this idea of having a plan to turn people into cloud-natives.

Jared: Yeah. I think I'm biased towards that because that's very much my day job. Is helping people accomplish that. But it's very difficult to succeed when you are adopting AWS. If you have as Forrest likes to put it, one of those cutely named like centralized cloud teams, like the Cumulus Nimbus team or something like that. If you're trying to disseminate knowledge... Or actually I should not say to disseminate knowledge, I should say disseminate architectural decisions and authority from a central point. It's very, very hard because you haven't won the respect of the people that are not on that team, to actually adopt and follow those practices.

Jeremy: Right.

Jared: They're going to say I'm familiar with EC2, and maybe I am familiar with Ansible or other of those solutions like that. From my data center, they work just fine in EC2, you can take your SAM template and just go shove it in an S3 bucket, I don't care. I'm going to use the things that I know work. And I think you really have to invest and train in folks that are helping you on your cloud journey. So you can't just do that with a centralized team, without having that hard-won knowledge, like pushed across your entire team, it's just very difficult to succeed.

Jeremy: Right, yeah. I think you need to get buy-in from people and it's easy for people to get comfortable. And you're right, if my Chef scripts and ops works and all this other things working just fine, why am I going to change over to something else? So I totally agree with that one. But that actually I think is another mindset that is very damaging for people that are trying to do these cloud journeys. And those are the people who think this runs this way in my data center, and basically Amazon or Google or Microsoft or whatever, one of these cloud, these public cloud providers, they're just a big data center, and so I'm just going to take my stuff, I'm going to lift it, I'm going to shift it into the cloud and everything's going to be perfect. You think that is a terrible idea?

Jared: Yeah. So I would say in my short tenure in doing some of these lift and shift things, about half a decade or so at this point in doing some of those practices or kind of being called in after that was attempted. It's never a short-term strategy. It always ends up being a long painful slog of, "Okay, can we make this very specially designed server work in EC2?" "Oh no, it turns out they don't have the correct processors to run this optimized code that we have." That's a very edge case example, but my goodness it never works.

You can do it short-term and spend a ridiculous amount of money doing so, and still not have what I would argue is, as good of a solution as you would take a little bit more time, a little bit more money up front and have a better solution long-term. And you can have it fast, you can have it cheap and you can have a good brand too.

Jeremy: Right. Right. Yeah. I mean, and that's one of those things too, with lift and shift where, I mean, I don't think you have to get everybody to embrace microservices. Right? You can build a lot of distributed monoliths if you need to do that. I mean, just switching over to something that already like RDS versus trying to run your own database cluster or any of those things, just starting to use more cloud-native services, I think is a huge step in the right direction, even if you're still running your application on EC2 instances.

But the one last thing I wanted to mention on that article, and you brought it up a little bit earlier, and I thought this was really good advice is especially for someone like me, I do a podcast, I write a newsletter. I try to keep my finger on the pulse of serverless. That includes Cloudflare, and AWS, and Microsoft Azure, and GCP, and Fastly, and Kubernetes and Kubeless, or Kubeless or however you pronounce it. There's just so many... Who knows? There's just so many, OpenFaaS, right?

There's just so many cloud providers that are doing this now. The Adobe I/O Runtime, I mean, there's just so many of these that are doing this now. And trying to keep up with these just from an information dissemination standpoint is really tough. I know nothing, nothing about some of the services in these other clouds, other than just a tiny bit of it. So if you asked me, "Hey, you're reporting on all this stuff on GCP and on Azure, can you show me how to set this up?" "No, I can't. I don't even know the first thing about it."

And I think that is good advice for people that are trying to build something. You can't learn AWS completely, so don't try to learn four different public cloud providers either.

Jared: No, I mean in the very early days of Trek10, both personally, and then also as a company, we were asked to do work on Azure, I believe. And we attempted it and it worked, right? Like what we built worked, and then we pretty quickly decided AWS is already huge. The market share is plenty, and I think that's true for most of the other providers as well. You could probably make a living consulting on GCP or consulting on Azure but doing a good job in delivering services or even your own product, on all three of the cloud providers, the main ones I should say, there's tons more, Azure, CloudFlare, all of those.

It's extremely difficult. And I think you just have to optimize for the minimal time and brain capacity you have, pick one and commit. And if you're wrong, okay, just pick another one and commit. I'm sorry, you wasted the time, but... Even today, if you were to tell me Jared pick one, it can't be AWS. That's fine. I would just go pick one and that's where I'd go. I don't think I would try to distribute around too many of them right now. You can't do what I would consider a good job by spreading yourself so thin.

Jeremy: Yeah. And I think you've got all of these major cloud providers and it doesn't matter if it's Tencent or Alibaba or any of the big U.S. ones. There's just a growing ecosystem around every single one of these. So yeah, so AWS great. If you like it. I mean, it's got a lot of stuff. I love AWS because there's so much stuff there, but Azure is pretty cool. Right? And they've been doing a lot of stuff around that. And GCP has Cloud Run, which is a very cool way to do serverless containers.

So I totally agree with that. But I think if somebody is looking at this overwhelming number of clouds, that's just really good advice. Do not try to learn them all, pick one, go deep on a few services and learn that.

Jared: You'll never build anything, you'll spend all of your time learning. Right. So...

Jeremy: Which is important, but at some point you're going to have to write some code. Alright. So I want to move on to another article that you wrote that was about the guiding principles for building a SaaS. And I thought this just tied in nicely to this other article that you wrote recently. And one thing reading this article too, is I don't think this is just about building SaaS. This is about building any application in the cloud. I know Trek10 just recently got your SaaS competency with AWS, which is awesome by the way.

And by the way, I love these new Lambda ready programs and some of these things where it's just basically like certifying providers and consultants and partners that know what they're doing and kind of have that sign off from AWS. I think that's super important that those exist. So back to this article though this was... You know I guess there was about three principles. What are the three guiding principles? So I want to go through these, because I think this is super important for any company building an application.

Obviously I think it's very much a bias here towards building an application cloud-native and more so serverless cloud-native which is... I mean, again, this is the advice that I would give as well. So let's talk about these because this first one was really interesting, build as if you may sell at any time. What did you mean by that?

Jared: So I've actually experienced a couple folks have come to us and said "Hey, we're spinning off of this product, where we're selling the company, we're selling this branch. It is so tightly bound to all of our other infrastructure and practice, that spinning this off is like its own entire effort, that is nontrivial when it comes to actually needing to sell that product or branch or whatever." So I think the guiding principle there ultimately is consider your dependencies back to the company and things like that.

So if I'm building a new product line or a new branch of the company, I'm giving them their own AWS accounts, their own segment of an organization where, I can hand off the keys fairly easily, right? Even in terms of financials, I would run if I could through their own bank account or something like that. So auditing the books, right? Like it's good financial hygiene to be able to audit those books kind of independently for that product.

And then come back to your mainline business as well, and you have a better understanding of what my cost allocations are. Being able to do the hand off the keys, say, "Here you go. Here's the keys to your brand new car or product or whatever it is." That's ultimately where I'm trying to get is auditability and understanding, and segmentation away from anything else in your company. Try not to entangle too many things.

Jeremy: Right.

Jared: And that's really where the guidance there comes from. And the reason that I think that's important, even if you never plan on selling that thing, is there are security benefits. And you get your own internal audit benefits, and you get so many compounding benefits that I think it's worth that what I would consider small upfront investment if done correctly at the beginning.

Jeremy: Right, yeah. And then the other thing that I think this ties into, and you mentioned it in the article, is also the idea of sort of setting up the developer. I don't know if we'd call it the developer experience or just sort of the developer I guess interface into this piece of the cloud. And you quoted Ben Kehoe in the article, when he said "Move your development environment towards the cloud, do not try to move the cloud down to your dev environment."

And this is me interpreting this and you can correct me if I'm wrong, but I see this as basically saying, "Don't build some sort of overly complex local development system that is going to be really hard to migrate if somebody else takes over. Utilize as many tools as you can to again, move that development experience more towards the cloud."

Jared: Yeah. Right. I mean, if you can sufficiently mark AWS on one machine, you should be starting your own company. But I would say you can mark well enough locally to do some very fast unit test or things like that. But you cannot really sufficiently mark or simulate the cloud locally enough in such a way that I would consider it, even if you would run end-to-end test or something like that. Locally, I would never consider that sufficient as compared to doing it in the cloud.

There're so many things that are interesting about your application running in the cloud, whether it's network or even IAM permissions or things like that. There's so much complexity up there that I would much prefer my developers are working kind of with those resources natively for their end-to-end or integration tests and things like that, all of that, that should just all be happening up there. Now of course, it can be painful if you're using SAM or things like that. And you're like: type a line of code, try to push it, type line of code, build, push, I get that, that's painful.

And that's where I think fast, local unit tests are acceptable. That's fine. I'm never going to ask to take that away from a developer. But giving developers their own AWS accounts, or their own kind of small team environment, ephemeral AWS accounts, there's some cool tooling out there, that's kind of enabling that, that stuff's really cool. And that's where I would invest company and some engineering time into providing that better engineering experience for the rest of my teams.

Jeremy: Right. So then the other thing that I guess, ties to this idea of being able to sell it at any time is, what you title it as "build as if you may open source at any time." And I can tell you, I write a lot of open source projects and it's a bit scary at first when you write an open source project. And you're writing documentation and you're letting people look at your code, because I've worked for a lot of organizations where you would not want to pull back that curtain and see what was behind there.

Lots of duct tape, lots of Popsicle sticks, hamsters in wheels, keeping things running. And I think that is true of a lot of companies. And I think the way you get there is because technical debt builds up over time, right? You're moving fast. Like, "Oh, we have a proof of concept." Next thing you know it's productized, and we were missing some things there. But this is I think a really, really good point is, you should build your company in a way that says, "Look, we could be transparent tomorrow if we needed to be."

Jared: Sure. Yeah. And I mean, a lot of companies are building towards, I'd say short-term right now value versus long-term stability. And I get that. It makes a lot of sense, especially financially in certain cases. If I need something right now versus, what's this going to look like in six months, when we circle back to it.

But ultimately building as if you're going to open source at any time, I think forces you to at least think if somebody was looking over my shoulder right now, right? If I was building something and say, "If Jeremy's looking over my shoulder right now, am I really going to put like my GitHub magic key, or whatever into this line of code and just hard code and deploy. And be like, 'I'll fix that later?' I feel like Jeremy's going to be back there and be like, 'Really man? I respected you and now no, like there's nothing.'"

So I think it helps you justify the few extra minutes or in some cases, hours or days, to make the right technical decision. And I get tech that's a thing. Look, go look at open source code. There's tons of stuff out there that's like to do actually make this work appropriately or optimize this thing. That's fine. Nobody's going to judge you. We all get it. People write software and they understand software is hard. But they are going to judge you pretty hard if you make poor security decisions or poor architectural decisions where it just doesn't make sense. And I think having that fictitious open source gazer over your shoulder, it just helps you make those decisions and kind of think to yourself.

Jeremy: Right. And it's beyond just code though. I mean, I like to litter my open source stuff with to dos, because I know if I don't put it in there, I won't go back to it. And I also feel like you put some to dos in an open source project, somebody that has... Excuse me, that has a little bit extra time might come through and be like, "Oh, hey. I can do a PR for it. Great." But internally in your own company, I mean I still think that's a good practice.

I mean, if you say, "Look, this thing doesn't check the string the right way, or it needs more parsing or more validation." Great, then just put it to do in there and say that. But I think another thing that almost every company I work with and every company I sort of dealt with that isn't a open source company, that's publishing closed source software. They are terrible when it comes to documentation.

Jared: Yep. And I would argue, I mean, even most companies that aren't just all internal or anything, many open source projects have terrible documentation.

Jeremy: It's very true.

Jared: You've just never heard of them, because they have terrible documentation, because nobody knows what they do. Can you tell the theme here? Documentation. It's very important. So I think that when you're building most successful open source projects, if you go look have... They're very good at these technical documentation, you can go and understand how the product works and the technical documentation. And then also they usually have decent narrative documentation.

You can understand what the product does, how you can leverage the product, how it solves your use cases, things like that. And I think having some of those internally as well, can help with onboarding a new employee faster. They can help with, when a client comes along and says, "Hey, can you explain more to me about feature X? How does this feature work?" Right. If you can hand over some decent documentation around that particular feature, even if it's not the technical documentation, but it's the narrative docs and say, "Here's how this thing works. Here's how it's designed internally a little bit."

Give them a little peek behind the sheet. That's fine. And customers will value that, your employees will value that, people that are trying to build that contextual awareness of your product and how it works, they're going to value that documentation. And I think building as if you make open source at any time is there's the embarrassment aspect of the code decisions that might not be great. But you wouldn't want to open source with no explanation. And I think that documentation is part of your explanation.

Jeremy: Right, yeah. And one of the things I like to... And I don't think I'm the only one who does this, but if I make a decision, if I say, "Okay, I'm going to use SQS versus Kinesis or something like that." You make that decision and then six months later you go back and you're like, "Why did I choose SQS again?" Or, "what was the one..." I think you do that a lot. And so oftentimes I try to put in just justifications of why I made a certain decision. I do this a lot too.

I mean, I know recursive functions generally are pretty bad if they go very deep and they can cause all these kinds of overflows and whatever. But stack overflow, where the site got its name from, but the thing that I will do, sometimes if it doesn't go that deep and I know there's a limitation to it, I'll write a recursive function because it's faster, it's more compact. It's easier to probably reason about, especially if you had to write out some long interpretive or imperative version of it.

So justifying that though, and putting a note saying, "I did this, this way because," and I didn't do it this way because I think those notes can be really helpful as well.

Jared: Yeah. And I mean, I think even kind of going back to this case of the cloud improves around you, or things improve around you. Even going back here in and say, "This tactical decision was made before this other thing existed," is completely valid to do, right. Someone might be like, "Why in the world would you have used Kinesis here, when you can totally have used SQS FIFO." And you'd be like, "I made this decision two years ago. I can't be held viable for improvements made to the cloud while I wasn't doing this."

Jeremy: Yeah, dates in your code. I guess the other thing, I always date comments in my code, because it's helpful to have those when you do that. This other concept or this idea where people think... Or I've heard a lot of companies where they're like, "Well, the code is the documentation," right? And we're not talking about well-commented code that can use a doc generator, which we can talk about in a second. But the thing that I tend to see, especially when people write code is they like to get cute. They find their own shortcuts, right?

They find their own different ways to do it. I was into functional programming for JavaScript for quite some time. And then I realized, "Okay, the speed benefit of the interpreter probably doesn't matter. It just looks more compact and it's much more confusing when I go back and look at it later." So it's even looking at my old stuff from like two years ago. I have no idea what that even does." So speaking of that, do you think that the code itself is good enough documentation if it's well-written or what are your thoughts on that and what are your thoughts on adding a doc generator?

Jared: Yeah. So I think that the best written most elegant code in the world cannot compare even remotely to well-done natural language documentation when it comes to communicating the intention and context of what was built, right? How this thing works? Why we built this thing? Sure, the code can explain what it does and at a technical level, how it does it. But it doesn't explain the business reasons for some of that code, and the business impact, or even necessarily the other systems that might depend or how they depend on that.

And of course you can get some of that in, I think doc generators, can get part of the way. You can explain definitely the technical library API or the technical documentation can be codegen, or docgen for a lot of those. Sure. I think you know Py docstrings and JavaScript docstrings, all that stuff fantastic, we need it, it's important. It does not replace the narrative documentation that helps people actually read for context and understanding. And I think even what a lot of people underestimate is the value to the doc writer, of having to sit down and write out those docs.

Jeremy: It's not easy.

Jared: It's not easy. I'd say it's not easy, but also you as individual sitting down and writing that documentation, might discover things about what you've just built, or you're going to build that you're like, "Oh, wait, I missed something," right? Like, it could even be like trivial things. "Oh, you know what? We have this input for color, but it's actually enumerated and you have to give us certain kinds of colors. But that's not explained anywhere. Are we just going to let people pass in a hex code for colors or are we going to expect strings for colors?"

Things like that, where it's like, you have to explain that. And also just in general, here's limits of the service or things like that. That's not always explained terribly well in the technical documentation, whereas narrative docs, as we all know from AWS, you have to go to the limits page or something like that to really find where that is.

Jeremy: Yeah, totally agree. So, alright. Last thing on this subject, because this is another thing that is painfully obvious with most companies you work with, is the fact that they don't write any tests or if they do, they write very few tests.

Jared: Yeah. So, I mean, that's definitely true. And I think this is open source gazer over the shoulder, but oh my goodness, you have to have some tests in there and this is just pure embarrassment aspect. And I think tests are another level of documentation that most people don't consider to be documentation. But I can tell you that at least me personally, probably the first way that I go and understand if I should use an open source project, is to go look at the test folder, the test directory. Because it's going to tell me a couple of things.

It's going to tell me A how well tested are they, how solid is this system. But beyond that test code and the test themselves, explain to me the API of the system and how the library works in most cases. I can go look at that and say, "Okay, this is what the system is capable of. This is what they're testing for. This is how I interact with this library or this system. This is how they think I should be interacting with this library or system." Things like that. Right? And that's stuff that you can't get from any other form of documentation, but also it's just system stability is it's so critical.

Jeremy: Right. Especially with I mean, just regression testing and any changes they made to the code. And I think I just said this on the last episode was, I see code where you change something and you're like, "I have no idea if this will break everything or what? Because there's no way for me to thoroughly test it." So I think that's hugely important. And I mean, even just some level of testing even at a high level, even if you're not going deep on unit testing, even testing at a function level and more complex things happening under that.

But it is just such an important thing that... But I get it, it takes time, right? It's an investment. It's another thing you have to do that takes away from feature building.

Jared: And I do think that it's kind of interesting once you move more to this cloud-native managed services world, I would say that end-to-end testing is more valuable than ever, right? I can't unit test, or I shouldn't really be unit testing does S3 ranks?

Jeremy: Right.

Jared: I think I can safely assume it probably ranks. And I'm not going to like unit test the core functionality of that. But what I can test is if I call my API, that's supposed to write an object to S3 and then can I call the other API endpoint that is supposed to mutate or map that object and give me some kind of response out of it. And if I can do that, if I can end-to-end test, like a few different things, through a few different endpoints, my goodness, I can get to like 80% code coverage with probably 10 tests in a sufficiently large system.

And it takes you a day, two days, let's say a week, because you have to learn a whole new end-to-end testing framework to get to a pretty high level of confidence that my system is at least working in the happy path. And that's invaluable. So that's where I would start at least these days.

In fact, when I go to work on new client systems, if I don't know how they work, and I don't have the docs, and I don't have even doc generator docs or anything like that, the first thing I do is sit down and go, "Cool, I'm going to run some end-to-end test, just so I understand how your API works or your system works. Then the side effect is I understand this thing and we haven't done tests, so I can start changing stuff and at least know if I broke something big.

Jeremy: Yeah. So I totally, totally agree with that. Alright. So last thing, and this was the one I was really looking forward to getting to. And I think this ties back in with the note you said earlier about just sort of building cloud-native people. Is this idea of just building with a cloud-native mindset, right? I mean, because this is the thing where your opinions on lift and shift, I totally agree with that. I think it's a bad idea, might be a great maybe onboard thing, but you know that stuff's going to get left that way.

So if you start thinking about building things in the cloud and using those native services and you actually had a quote in there, something like, "If Cloud Formation doesn't exist for it, then is it even available," or something like that.

Jared: Right, right. Yeah. If it's not supported in Cloud Formation, does it actually exist? Questionable.

Jeremy: Right, right. But yeah. But you tie i into a bunch of other things and you've always had this really great, I don't know, sort of like your serverless credo or something like that. Where it's if the platform has it, use it, if the market has it, buy it, if you can reconsider requirements, do it. If you have to build it, then own it. And I love that because I think that is such a... It takes what we're trying to do with serverless and just wraps it up into four quick sentences, which is great. But your thoughts on that overall, what are your thoughts on this building with a cloud-native mindset?

Jared: I think it takes practice, right? You're giving up a lot of fundamental control that I think people are used to having, right? I can't walk into my data center, open a rack and turn off or turn on a server or pull wires or things. That's a huge fundamental shift for a lot of folks. And as we're migrating to people now, these days, that have never even walked into a rack of servers, we're having people that are coming out of college that AWS and going into ec2 and clicking launch instance is their concept of a server.

I think what we're starting to build towards in terms of this cloud native mindset is, we fundamentally can trust these larger providers to provide mostly good experiences, let me be careful there. Mostly good experiences around these cloud primitive services. And we have S3, which has kind of been referred to as one of the seventh or eighth wonder of the world. It's like this modern Marvel, right? That thing holds so much data and performs so well, and it's so scalable.

When it goes down the internet is just basically done. That's incredible that they have this service and we're trusting it. As cloud-natives, we're trusting these providers. I don't care if it's Azure or GCP or anybody, to provide these primitives that we can build on top of it. I think cloud-natives look at those primitives and you have an implied level of trust, and you're willing to build businesses and business value on top of them.

And I think it's control and being able to trust somebody else with giving up that control, so you can accelerate what you're doing and looking to build in terms of business value, is more of a cloud-native mindset than anything else.

Jeremy: And I'll bring Forrest into this again, because you brought him into this. He's becoming very popular as a side topic on this podcast. But I think you included one of his cartoons about the regret index. And that is brilliant, because it perfectly captures exactly what it is. If I write something myself for me to be like, "No, I'm just going to get rid of it," is so hard. If I just buy something for me to switch from X to Z, for some product that I bought much, much easier, and it's no skin off my teeth to do that. But if I built it myself, I really, really want to hang on to it. And that is, I think just a huge problem.

Jared: Yeah. I mean, we see organizations and when I say we, I don't just mean Trek10, I mean, universally. I think all of us have seen organizations that have built something that some platform does better, some product does better, some open source product does better. And pretty much everybody says, "Why in the world do you keep using that thing?" And it's like the not-built here syndrome, right? Interestingly, I think Netflix suffers from the not-built here thing, but also they have this side effect of the stuff they build is also really good.

And all other people are using it, so outlier. But to take it to an individual level, right? If you're willing to build something or invest in a hobby, or craft brew, or something like that, you're not just going to throw that out if it's not great.

Jeremy: Right.

Jared: Right? You're going to be like, "Well, I've made this thing, I'm going to drink it now. This is terrible." But you take a sip. You're like, "Wow, that's bad. Hey, Jeremy, do you want to try this thing that I made?" Whereas if you go buy a six pack or whatever from the store, and you take a drink and you're like, "That's terrible." You're like, "Eh."

Jeremy: Yeah right, exactly, exactly.

Jared: I'm not going to drink that. I'm going to go get another drink.

Jeremy: Just dump it out and yeah.

Jared: And I think that just generalizes to building stuff as a company. And Forrest, I think really brought that to light when he was like, "If I buy something off the shelf and it doesn't work out, my regret is, "Well, I spent money on something that I could have spent money on something else." If I invest engineering resources and build my own thing that turns out to be not great and the wrong thing, I'm much more invested in saying, "Well, I've already made that decision. I don't want to look like an idiot to my company. I don't want to waste company resources. We're just doubling up."" Right?

Jeremy: We're going to make it even worse. We're going to keep on working on it until this thing is even. Yeah. Well, so I think there's a couple of key takeaways you had in here. I mean, one of the thing was just like this idea of cloud-native mindset, or building with that, is you're giving your team that autonomy so they can build things on their own, right? That they have more flexibility, right? You're not trapped into some system. You can try new things. You can experiment.

You can use products that get better around... They get better like you said, they just get better with time, because somebody is upgrading those things and you don't have to do anything. And then the greatest piece of that whole bit of using somebody else's stuff, is you don't have to write documentation for it. Right? Because they've already written the documentation. So anyways, Jared, thank you so much for being here.

Jared: Thank you.

Jeremy: Those articles are awesome. I will put those in the show notes because I do think you need to go check out not only those, but also everything else that you're working on. So if people do want to find out the other things you're working on, how do they get ahold of you?

Jared: Yeah, I'd say the best way is of course Twitter, the universal complaint box. So @shortjared on Twitter, and then you can also go to my website jaredshort.com, which just uses a notion page because the market had it, I just used it. And then you can always email me as well, I guess we can put that in the show notes, I guess. But yeah. Thanks Jeremy. It's been an absolute pleasure.

Jeremy: Awesome. Alright. Well, I will get all that information into the show notes. Thanks again.

Jared: Alright, thanks.

THIS EPISODE IS SPONSORED BY: Amazon Web Services(Serverless-First Function May 21 & 28, 2020)

View Details

About Linda Nichols:Linda Nichols is a Cloud Solution Architect at Microsoft. In addition to creating software solutions, she has a passion for community involvement and education. She is a co-founder of Norfolk.js, NodeBots Norfolk, and RevolutionConf. She also enjoys teaching local classes and workshops.

  • Twitter: twitter.com/lynnaloo
  • LinkedIn: linkedin.com/in/lynnaloo/
  • GitHub: github.com/lynnaloo

Watch this episode on YouTube: https://youtu.be/e5oFSIMuvcM

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Linda Nichols. Hey, Linda, thanks for being here.

Linda: Hello.

Jeremy: So you are a cloud native technical specialist, and a member of the global black belt team at Microsoft. So why don't you tell listeners a little bit about your background and what you do at Microsoft.

Linda: Sure, sure. So first of all, I have like the most awesome title at Microsoft. And most people don't understand, but it sounds great. Especially you join a call with a customer and they're like, the global black belt is here. But, essentially, we're problem solvers. If someone has a problem, or they want to know how to build something, or they want to have a conversation about maybe like, something that's out of the norm, like, it's not your typical service that a lot of the cloud architects within Microsoft know, or it's something more in the open-source side, something that's heavier into serverless or Kubernetes, or just something that's maybe out in the community. There's a lot of open-source tools and things that we talk about, we could be just the open-source blackbelt team. And my background is development. I was thinking about this today. Like, I'm not afraid to say I'm in my early 40s. And so now, half my life has been developing.

I've had a professional job as a developer for half my life. So it's really like kind of ingrained in me being a developer, even if I'm not coding every single day now, because I'm on the phone a lot, just like chatting with people, but I still, like really enjoy kind of hacking at things and thinking about methodologies. And, that's part of what we do too on our team I mean, maybe someone calls us up and says, I just can't get this working and we help them through it. But also, maybe we just kind of talk about, like why are you doing this this way? And, what you think about this? And how about these tools? And that sort of thing.

Jeremy: Awesome. Alright. So I've seen you give a number of presentations actually in a lot of the presentations that you give are around DevOps and serverless, right? And kind of how those things connect. And speaking about being in your early 40s, one of the things I love about your presentation, I'm in my early 40s, as well, I love your, like 80s and 90s References because I get all of them and it is absolutely amazing. So, but your talks usually are around DevOps and how it kind of intersects with serverless. And a lot of times about the serverless developer themselves. And I remember back at Serverlessconf, it was like serverless developers are developers or something like that and it was great talk. So I kind of want to talk to you today, though, about the culture, right? Like this culture around the serverless developer. Because, if you look at people using things like Amplify, there's this whole new thing like a full-stack serverless developer.

And then you've got some people who are kind of focused more on the, I guess, on the infrastructure side of serverless, which is maybe a bit of an oxymoron, but maybe understanding at least how some of these configurations work. So maybe you just give us a quick overview like, what is the overall culture look like for serverless developers?

Linda: Sure, sure. Well, first of all, you threw like an AWS term at me. And I was like, Amplify, which one is that? But, yeah, I mean, I think what I keep trying to kind of drill in, is it like, yeah, serverless developers are developers. And I keep saying too, serverless was made for us, right? I mean, serverless wasn't really it didn't come out and become popular because ops people were like, "No we don't really want to do our jobs." Like we hate infrastructure. No, no, they love it, they've been skeptical this whole time. They're like," Oh, so the developers are going to push to production now. Okay, have fun with that." so I mean, it's essentially for us, so we shouldn't be the ones that are distrustful, we should be the ones that are saying, okay, here's our process, which is what we love to do. These are the things that we've done to be really successful at development all these years. And now we're carrying it over into this ecosystem where we have a little bit more control, but also less control kind of. I mean, we're not having to hand as much over to the ops people. But we don't have to worry about things.

Like, I think there was a period of time there for a while, where I started to have to care about Docker containers more than I wanted to for a while in development. And I mean, I'm at the point now, where I kind of, I understand the process a lot more, because I've just been in cloud for so long at this point. But there was a point where I just, I really, I had a lot of strong opinions about my development environment and like and libraries and tools, and then suddenly I'm like, Okay, well, I'm just going to push to, whatever past service and then the ops people are like, okay, but like, we're going to need you to like, put a Docker file in there. And I'm like, "Hmmmm." And then there's suddenly there's all these troubleshooting steps. So when serverless kind of took hold, I was like, oh, okay, everyone, this is now the way forward. Because I don't have to care as much, I'm just using some command line tools. And just as simple as I push to GitHub, I push to the cloud.

And I really got on board with a lot of tools like serverless framework especially, too that even abstracted some other things away. Now, I think we're at the point where like, you can use different IDEs, and push different cloud platforms and be really successful too.

Jeremy: Right. And that's actually one of the things I want to ask you about, too, is that again, I've been developing for 20 some-odd years, I think I started in 1996, or something like that. So, this idea of having your tools, right? Your IDEs, I mean, I remember way, way back when using Eclipse and things like that, and some of those other ones. Obviously, there are a lot more now. But what are those tools that the developers were using in the past and are still, or can they still use those now?

Linda: Yeah, well, it's funny, I was talking to my husband yesterday, so my husband works for GCP so we're like a multi cloud household anyway. But we have a very similar background in that we were both Java developers, and we moved to kind of...

Jeremy: I'm sorry to hear that. I sorry to hear that.

Linda: Imagine that. I still love Java. I don't know .NET at all, which is funny. I always like, I'll tell customers, I'm like, "Hey, Microsoft hires traitors like me too." like, I don't know anything about .NET. But I can look at it and say, like, okay, this is enough like Java that I can figure things out. But, so we still talk about kind of development culture a lot. And we're both in cloud now. And so we were just talking about Eclipse yesterday, because I just loved Eclipse. And when I moved to Node, I was like, Okay, so now to kind of assimilate into this hipster culture, I need to like use just a text editor with no highlighting. Okay. Alright. So I struggle with IDEs for a long time. And I was saying yesterday, I just got to the point where I feel like VS Code is at a point that I love it as much as Eclipse. And I'm not like, a VS Code salesperson, but like I was already loving it quite a bit before I even joined.

Jeremy: I was actually, I used Atom for quite some time that was sort of my go to and everyone was like, oh VS Code VS Code. And I'm like and then I started using it, then I got my new laptop, I only installed VS Code and other than needing to create an alias in my CLIs so that when I type Atom that it opens code right? That was, other than that, the transition has been relatively easy.

Linda: Yeah. So and I loved Eclipse, I think VS Code is there. But there's all this other like additions too. Like I said, there's so many great extensions and like trustworthy extensions, you can really push it to anything. And, I think, because my interest is developer culture and deployments and just thinking about how people interact with cloud, like I think about multi cloud all the time. So I work for Microsoft, and I love that Azure but I think a lot about okay, well how would AWS approach this problem? How would GCP approach this problem? My husband and I talk a lot about, okay what's like your priority? Or like, how would you tackle this if you were using Cloud Run or Knative or something like, how do you approach this problem? And really, as developers, that part's the same. I mean, once it gets to the cloud, things are architected slightly differently. But as developers, we still have the same sorts of opinions. And we're still essentially using our same process that we did when we were both developers, but we just pushing to some other place.

Jeremy: Well, and that's something interesting too where when I get into the mode of sort of enterprise development, or you're developing code, you're checking it into a CIC, or you're checking it into some sort of code repository. It's going through code review, going through CICD, you got Jenkins doing builds and tests and all these kind of things, and doing the deployments for you. That was something that became very, I just think became ingrained in a lot of developers that were working for larger enterprises. And then serverless comes along, and all of a sudden you've got the serverless framework, it's like serverless deploy and next thing you know, it's in production. Is that something... Obviously, there's good practices around code repositories and GitFlow and things like that, and CICD and testing, but we can talk about that later. But is that something you see is part of the culture where just people are I mean, I guess they're skipping some of those things?

Linda: Yeah, yeah. And I mean, that's kind of what I talked about a lot. And I was making fun of we as developers in my talk at Serverless Nashville and also at Serverlessconf. But, it's just funny to me, because especially coming from a Java background, I mean, Java developers are just notorious for process. And I mean, your static code analysis and going through I mean, it would take forever for a build to run. I mean, literally, it's just like the XKCD comic, where the guys just sword fighting while the build's going on. I mean, it really was like that. I remember sitting there and being excited when smartphones came out because I could get on my smartphone while my build is running. And it was because it was just running through so many tests, unit tests and integration tests. And like this whole process of even being able to check code in and then moving through all the different environments was a pretty elaborate pipeline. I mean, I didn't have access to testing or production, I only had access to development.

And so when I moved to Node.js, it was a little bit more relaxed, just because like Node was super brand new when it started and we just didn't really know what to do exactly. There wasn't even really like testing frameworks out. And so things were a little bit wild, wild west, but people still knew that they were doing things wrong. Like they were like, Okay, this feels really dirty that I'm not unit testing. This feels... Like how do I move things through the stages? But I started noticing though, when I was in consulting before I started working for Microsoft, that our customers were not feeling that guilt at all. That something, there was some like switch that flipped. And like they're in the cloud, but more importantly, like they're in serverless, and suddenly it's okay to just push to production all the time.

And so we talked to customers, and they would say, "Oh, I have 40 or 50 Lambdas." And, okay, great, great. Okay, how are you keeping track of them? Or how are you testing them? Or how do you develop them? "Oh, well I log into AWS and I go to the portal and I type some code in, I hit submit." And I'm like, "Okay, so you don't lint anything, you don't test anything. That's production. You don't have a separate subscription or account for these different?" "Nope, mm-mmm no." And then they want help debugging something. And I'm like, "Okay, well, first of all, like we need to start from scratch." Because I mean, at that point, I think I was kind of saying during my talk it's not just even pushing to production, but you're essentially like using the cloud as your revision control, like, yeah, your portal i, AWS or Azure, GCP. Like that little box that has a certain amount of protection in there, that's not your revision control. Like GitHub needs to be involved or GitLab or like, whatever your choice is, but.

Jeremy: I remember when I first started developing websites, and I was uploading Perl scripts to cgi-bins, but what was great about that was, you make a change, you upload it, and it's immediately available because it's running on one server, right? We didn't see the volume that we had now. And you do things like that, and it was fast, like, oh, something's not working, right. Okay, fine, change it, upload it, it's there. And then we went to this culture, like this Java culture, like you talked about where we're essentially like, okay, you need something changed, it may take four seconds to make the code change, but it's a good three hours before that thing's going to make its way into production. Then serverless comes along, and then all of a sudden, I feel like I'm uploading things to a cgi-bin again and it's like great, because I'm like, "Oh, I can just put this into production really quickly."

And that's obviously a better practice, I think, than just using the console. But certainly being able to just write quick scripts, especially if they're like DevOps Scripts. Things that aren't really client facing production, just being able to push those up is a lot of fun. But now we go back to this thing where it's like, we have all this control, the power is in the developers' hands, but the importance of going through that CICD process, that's something we have to get back to, right?

Linda: Yeah, yeah, absolutely. And I mean, I will say if you're using serverless framework, or you're using VS Code to push, you're at least slightly better than the portal developer, because you're at least in your IDE, you're in your comfort zone. And you probably have some sort of linter or static code analysis running because, in VS code, or atom or whatever, there's all these extensions. So you're like, you have a little bit of protection engaged there. So that's at least somewhat good. But then yeah, of course, if you're using serverless framework, just pushing production, then you're going to possibly run into some issues, especially because you're offline there too, right? So you're not considering all the other services

There are so many tangents that I could go off on here, but I think that even if you have like a baby CICD process, like, even if you're pushing to revision control, and you're keeping track of those changes that you're pushing, and then maybe you have something simple from there, like it's a GitHub action, maybe you go with like Travis CI or pipelines and Azure DevOps and then you have like a really simple pipeline and that pipeline, just runs some additional checks for you. Maybe it does some additional linting or runs through your test suite or something. Like that's better. Like you don't have to start out, you don't want like to get analysis paralysis, right? Like, okay, I'm writing a Twitter bot and so what is my path to production look like? No, not really but like, but you need to like at least have that muscle memory engaged where you're developing.

And I think that's kind of the thing is with these developers, like when they, you cross languages or you cross tools, you still kind of keep that muscle memory of this is what it feels like to develop and this is the right way to do things. But there's just something about going to serverless that just became like processless. It's like, it removed that muscle memory because it suddenly felt like it wasn't development anymore, but it's like more development than ever. It's like most development you can do in the code or even the cloud is serverless, essentially.

Jeremy: But that's the other thing though that I think that's kind of interesting, too. And maybe we can get into just the testing bit a little bit more in a minute. But this idea of being able to write code on your local machine, and then test it locally has always been a thing. We'd always be running a JVM or something locally on our machine where we could test that code, and we knew exactly how it looked. And then we moved to Docker and we'd run things in Docker containers and if it ran on the Docker container on your machine then it was even more you never get that it works on my machine TM statement all the time. But the thing about serverless that is sort of interesting now is that you oftentimes need the cloud and you need all of those interconnections in the cloud to actually see if something is running correctly. And so that's why I really like this rapid development thing of sort of like SLS deploy to Dev. Like just keep putting stuff up there and then be able to test that, but because that's so easy, it seems like it's also like well, wait, I could just change dev to prod and then all of a sudden it would be live. So, getting rid of that process.

But I mean, that's the thing, too, is like, how are people supposed to just iterate? You know what I mean? Especially when you need the cloud for a lot of the nuances and the interconnectivity that you're going to have to deal with.

Linda: Yeah, I mean, you still need your playground, right? And that's fine. And you can still do whatever the heck you want in your playground. But the problem is, is that, and we've all seen this, as developers to POCs always become production apps. Always. Every time a customer says, "Oh, we're just making a POC." Oh, yeah, That POC just-

Jeremy: Oh in production. It means, the P stands for production.

Linda: That's right. It always turns into some enterprise application. So I mean, that's fine if you're truly playing around, you're truly testing things, you're getting things going in Dev. And Dev is Dev for a reason. But once things start going to test is when you start thinking, okay, I'm building something real here. This, like, there's going to be an end user, there's going to be some consequences here. And then that's kind of when you really need to start to buckle down on that process there. But, I have seen that so many times, oh, this is just a POC, we don't need to worry about, DevOps not for our POC.

Jeremy: Well, I always find that as soon as you put something up there, that anybody who's not a developer can play with, it's automatically in production. They just assume it's productized and it's ready to go and they start pointing customers there so I've had that experience as well. So the other thing about serverless, I think that is really interesting is less as a technology and more as a mind shift or mind change, or whatever it is, or what's the other word that we usually use for it? Like, I guess it's sort of a culture shift in the sense. But this idea of saying, look, there's all these services out there for you, right? So if you're in Azure and you want to use a database, use Cosmos DB or whatever, right? Like, you've got all these other things that exist for you, why would I create these things ourself, or recreate these things ourselves? But that, I guess that's less serverless developer culture and more overall developer culture. Everybody wants to build their own things. But is that something you think might shift with serverless?

Linda: Yeah, absolutely. And I mean, I think in the cloud in general, because when we were talking about standing on the shoulders of giants before, we were talking about using NPM packages or Ruby gems or whatever. And like you kind of just don't even know who's written those. And there's some security kind of considerations there. And also, like, you just don't know how much they're tested, you don't know if that maintainer is just going to go find another job somewhere. When you're talking about cloud services. This is true for all the major clouds. I mean, they are tested by millions of people, they are used billions of times. So I mean, if you... No one's going to say like, oh, I just don't think that... Fill in the blank servers, like Cosmos DB. Like oh, I just don't think it's really tested that much. Or like, what if the person who works on it leaves. Well, there isn't one person, right? It's a whole team of people, and there's a company that supports it.

So I mean, I think it's somewhat it's kind of ridiculous when people don't trust cloud services. If you don't want lock in that's a whole other discussion, right? But you if you're already in a cloud, I mean, a lot of people are multi cloud also, and they kind of spread things around to try to minimize that sort of lock in feeling. But really you get locked into libraries too right? If you're writing... If I'm writing a Node.js app, and I'm using some NPM package, yeah that thing's going to stick around forever. How often do you go back and switch out your NPM package?

Jeremy: Right, and if the NPM package, speaking of the developer going getting another job or something, if the NPM package is used by thousands and thousands of or has thousands of dependencies, and they make a change that breaks a whole bunch of stuff like that just happened a couple weeks ago, I mean, that's what I think was the isPromise package or something like that. I mean, that's kind of scary. I mean, then lock in, like you said, that's a probably a longer discussion. But I mean, for me, I feel like there's this sort of 80% rule, right? where if another service gets you sort of 80% of the way there, to me that just makes a lot of sense. Like, I like building my own stuff too like, yeah, I mean this system is good, but maybe I could make it a little bit better. And you might get to that point, but I just, what are your thoughts around just this idea of this, I don't know, I call it the 80% rule, and maybe it's called something else in actuality. But just this idea, like, just good enough. Is that something we should be embracing?

Linda: Yeah, yeah. I mean, yeah. So I mean, I have so many thoughts here, I have to give a plug to Forrest Brazeal for his talk at ServerlessDays Virtual this week, because his talk was basically a talk that I also was submitting to conferences about. But my talk was like, don't write code, because I just started thinking about the fact that if I was developing something for the cloud, or just in general, if I start typing a lot, I pause and I go, okay, somebody's already written this. I'm not that clever. Not really, there's a lot of smart people in the world. There are a lot of people that code all the time. This is already done somewhere. It's either in a library or it's a service. And I talk to so many customers and people who they're like, "Oh, here's my great idea of a thing." And almost always I'm like, "No okay, so that is this." And I mean, even like messaging systems like Service Bus on Azure. I mean, there are so many developers that have tried to write messaging systems. And there are so many out there, there's so many people that tried to write Kafka. And they still are.

And sometimes I talk to people that are trying to create something, and they will say, "Okay, well, I'm going to put this in a function, this in a function." I'll say, "No, you don't need functions here. This is already a service, or you can already use something like Logic Apps, like you don't have to write any code." And, you just kind of connect some things together, or there's already built in services and that's still serverless, right? Like serverless is not just fast. Like, I don't have to write 100 Lambdas to be a serverless developer.

Jeremy: Could just write figuration code. I mean, in many cases, it's just, it's YAML or JSON or something like that, right?

Linda: Yeah. And for me, if that's the case, I've won. I feel, I'm like wow, I just saved all this time, I don't have to test anything, because I didn't write anything.

Jeremy: You don't have to maintain it. I mean, that was the point that Forrest made, I think that was, is pretty smart to think about that as the amount of technical debt that you take on. He basically said something like the first version of anything you build is the worst version of that thing that's ever been built, because it hasn't been tested, it's missing all those features. I mean, and that's brilliant, right? And then not only that, but then even if you keep working on it and adding features, and so forth, someone's got to maintain it. And you know that there's always that one developer who leaves and you're like, "Hey, how do we manage this project? Or how do we control this server?" Whatever it is. And it's that one person who left who did it and all that institutional knowledge is gone.

Linda: Yeah, exactly. And I mean, what if that person left and everyone was like, "Oh, look at this architecture that he built, we only have to change it slightly and swap out a few services or Microsoft or AWS came out with something else that deprecated this thing he used and we just have to swap something in and out. And we don't have to worry about any bugs hidden down in code. I mean, what about that, that sounds beautiful, right? Instead of just hating Larry, who like left the company and like, why did he write all this code? And, I mean. I just, it's funny. I mean, this was, it was funny. It was like, the week that we got married, we went to Canada, and we're on vacation, and I was cooking up this idea for this "Don't code" talk the whole time we were up there.

I mean, it's just how my brain works, because there was some conversation I had before we left for Canada, and then the whole time I was there, and I'm like, and I'm sort of talking it through and sort of writing notes in the hotel because I'm just sort of thinking about, okay you're not that clever; your codes not that important. Like code doesn't define you, you're not a coder, you are an engineer. And like when you're given a bunch of building blocks that are solid, then use those building blocks to build something more efficiently and something that's stronger than if you built your own blocks that are potentially full of technical debt essentially. And so I think it just... It's a matter of changing our idea of like, what a developer means. And a lot of developers are still like running by this idea of like, Okay, well, my salary is based on how many lines of code I write or something stupid like that.

Jeremy: And the CEO who asked me how many lines of code we wrote, I said, "Why?" Why would you want to know that? The more lines of code, the more of a risk we are, right? I think that's-

Linda: It's like a Dilbert comics, right? That's like the pointy eared boss, like, well. So I think it's like, it's a matter of... So like my background, my original original background was art. I was an art major in college. And I spent many years struggling to figure out how the heck I was going to like pay off my student loans. And then I was repairing computers on the side, I was like, oh, okay, this is how I make money. But I'll be an artist. My identity as an artist, but like I'm going to make money doing this, like computer thing that like is stupid. Well, and so I remember when visual artists started transitioning to digital. And it was really difficult because they really felt like they were losing their identity or this is like I'm cheating, and I'm not getting my hands dirty and I'm not like using paint, and there's not a canvas and someone can't own a physical thing and so this isn't real. And there was like this division between the digital artists and the visual artists.

And so I feel like it's a little bit similar. But it's the same thing that you're still using your same part of your brain. You're still an artist, you're still an engineer. It's just, you're just evolving essentially. And you're embracing technology.

Jeremy: Yeah. And so there's certainly a shift, though. And I think this is something that you mentioned, where you write less code. But it certainly as a serverless developer, you might write less code, but you write more config. And it's more about, and I guess it's less about knowing how to code a switch statement and more about knowing how to connect, I'll use a AWS terms, because I'm more familiar with those, but how to connect a Lambda function to an SQS Dlq or something like that. So I mean, and knowing how that works, there are scaling characteristics, right? There's still stuff that you need to know like, how many concurrent functions do I want running? Or what are the redrive policies? I mean, there's a lot of stuff that you need to know that is not "if this equals that then do this."

Linda: Mm-hmm. Well, I mean, and that's when we bring the ops people back in like, "Hey, hey, guys, you still have a job. We didn't put you out of business with serverless because there's still servers and there's still config and there's still some infrastructure." So I think a lot of developers are still only going to go but so far. And I think really it's still our job to write like the infrastructure as code or to make sure that infrastructure as code is existing and that's part of our process too. Maybe we're not just writing Terraform or whatever all day long. But that's part of it too, is when you have all these configs and you have all these pieces together then that's when this repeatability becomes more important. It becomes as important as unit testing or integration testing, is make sure that you have a repeatable deployments and that includes your, all the ways these systems connect together.

Jeremy: Yeah, I mean in certainly infrastructure as code is just the new normal right? Like I can't even react... I remember again, this is how old I am but again, if you had to move it to a new system or you had to transfer something to where it was like everything from the database configurations to the patchy config mean, like all that stuff, you had to duplicate it on all these other systems. I mean, actually, not that long ago I, that so I guess it does make me that old. But even just using like OpsWorks, or any of these Chef scripts and Puppet and those sort of things, in order to recreate these infrastructures that just to me was this crazy idea where now with serverless, it's just like, serverless remove, and it pulls down the whole, or tears down the whole stack, and then serverless deploy, and it recreates a new one and it's just there for you.

So I do think learning infrastructure as code and at least the developers, they need to embrace that first piece of it, right? Like he or she has to say, "I am willing to at least give a rough architecture, and then write any code that has to support that." But then I do think Ops people come in there, that is something where understanding the scaling of that and monitoring it. Like that's a whole other thing that we didn't talk about. But I do think that developers need to really embrace that IAC stuff.

Linda: Yeah, I think so too. I think so too. And I mean, these are lots of conversations I've had as well because there's always that struggle between the development team and the ops team of like, something's broken. And whose fault is it? Is it the fault of the code? Is it the fault of the infrastructure? And if you're always deploying everything the same way, if your infrastructure looks the same on Dev, as does test, and then prod, it doesn't matter, like you should be deploying the same way you're deploying the same code, and that repeatability just sort of, like keeps you honest. And I guess that kind of goes back to like writing that code in the portal or whatever. You're just completely bypassing everything and you're not even considering, like, what configuration changes that someone might have made in other environments which you just give your Dev subscription or whatever. So yeah, there's just, there's so much outside of that.

I mean, it feels good when you get something working, when you just like you type something in there and like, I mean my talk at Serverless Nashville, I had a video of the doofus developer that's like typing in the portal. And then at the end he's like, "Yeah!" Because it like it works, that feels great. But for longevity reasons, like you need to embrace your process. And process feels good too right? When you have a process that just works, and like a bug comes up, and you know exactly where it is, you know exactly what the problem is and or like, tests fail, and you're like, oh, okay, that's why these tests were here, because it's going to catch this particular problem before it slips through to production, that feels good too.

Jeremy: Well again, I mean, again, I do want to get into testing just quickly, but that is one of those things too where, knowing that if you change something, I have seen spaghetti code, companies I've worked with or clients I've worked on code for, where there's just so much like, I don't know, if I change this, what happens, right? Like, why is it checking for a string and an object you know what I mean? And trying to figure out why it does that and then having no test coverage at all to know what breaks in production? So I think, yeah, I totally agree with that. But one of the thing though, about the Dev ops or the ops side of things. So obviously, a lot of configuration and like you said, the less code you have to write the better.

So what about going the other way? What about ops people that are very familiar with infrastructure and have been using Terraform, have been using CloudFormation, things like that, over the last couple of years, have gotten very familiar with this, the configuration piece of it. Now then moving into a developer role, only having to write a few lines of code in order to make an entire an entire system.

Linda: Yeah, and that's funny, because I've had a lot of DevOps conversations with infrastructure teams. And a lot of them are like, "I don't need this. We're not a development team. I don't need to have this discussion." And a lot of the discussions have been around our tool, which is Azure DevOps and really and like, I mean, I guess like, that could be seen as a developer only tool. Because you can run builds and you can run your static code analysis and unit tests and all that. But I mean, you can also deploy with your IAC of choice. And so there are a lot of... You can and there's so many like plugins and things, even if you're not using Azure DevOps, if you're using like some other completely different lesser tool. There's a lot of like we have a ton of plugins, so like, if people want to use like, Ansible, or whatever, we support all that. So a lot of discussions I've had, I've had to have conversations about tools that are way outside of the realm of like, what Microsoft produces, or even that I've used all that heavily, because I'm talking to entirely infrastructure teams, and I'm trying to teach them about the culture of DevOps, which isn't necessarily around code development.

You can have no developers whatsoever, like let's say you have some off the shelf tool, but you just need to make sure that that tool is getting deployed in the same way in a way that like supports, I guess, whatever configurations, this like code needs, this software product let's say. And that's sort of a bit... Like in healthcare, that's a huge use case. There's a lot of industries like legal or finance, where they have software products that they use and even have in house developers, but they do have an Ops team that has to make sure that this gets deployed. And as they make upgrades, as cloud platforms change, as maybe they're migrating to a different cloud platform, maybe they're going multi cloud, they still have to make sure that this software product keeps getting deployed and keeps working, and keeps integrating with all these other tools. So they have to start now thinking about testing, they have to think about repeatable deployments, writing terraform.

It's a little bit of a struggle because I've heard of it tons of times. Well, "I'm not a developer." Like, well that's, okay, like, I mean, anybody can be a developer and, it's not like you don't just get a stamp at birth "developer," "not developer." I mean, it's like, goes back to that identity that I was talking about before. But, and I think like, obviously, they have this great benefit of knowing infrastructure really well, and understanding like networking and security way better than I do even in the cloud. But it's more like teaching them like, you can't log into the portal and just even change a little configuration setting. No, because you don't know that you're making that same change in every environment. And then what if you make that change, and then you leave. And then everyone's forgotten that like Sally, like set some flag once upon a time, and then everything breaks the next time you deploy.

Jeremy: And it takes forever to hunt those things down. So I mean, another thing and this is maybe a little bit off topic, but, this idea of like low code or no code solutions and things like that. I do you see those, I mean, I think some people equate serverless to that idea of like what happens when we don't have to write any more code? And now it's just configuration, and why can't we just put that in UIs and things like that? I mean, do you see these tools becoming popular?

Linda: Yeah, absolutely. And I mean, I think serverless kind of started this, at least for me with things like API Gateway. Because for the longest time I was writing these Node apps, and I was using like, Happy or Xpress. And there was so much code that I was just repeating constantly, just to set up like an HTTP server, and to set up routes for like an API. And I mean, a lot of it, I was just copying and pasting code and copying pasting code into handlers and just like to add to like an API and it just like, it took a long time. And now either those types of things are just becoming separate services. Like, we have API Management in Azure. A lot of the things that I was repeating and a lot of the problems I had, I can kind of just do in the portal now and it's just kind of its own service. Or I remember the first time I wrote a web app that was like a Lambda with API Gateway. And I was like, wow, I don't have to worry about any of this HTTP BS. Like, I just wrote my code.

And also too, I mean, I would just take the libraries that I had written before, and I could just sort of call those still from a Lambda. And I really like didn't have to rewrite that much to make a serverless application. And that was really nice. And I still kind of write things like that, like, I'll write a library first. And then the actual amount of code that's in my function is really small. Because I'm just calling either other services in the cloud or I'm calling some library or something. but now I almost never feel like okay, I got to copy and paste this again, and copy and paste that again. Like maybe I have a scaffolding in GitHub and I just have a like a project template that all reproduce to start out with things. But, I never feel like, oh man, like, it really sucks to be a developer to have to, like, make sure you don't like copy paste this thing wrong.

So that I mean, so that was how it kind of started for me with serverless is like, okay, I'm writing way less code, and the things that are replacing the code that I was writing these services, they were so well tested and trusted, and I don't have to worry about it, like I never would be concerned that like, API Gateway or API Management, like just aren't working correctly, it's probably some configuration that I've made. But I think like in your next point about like, sort of low code, no code, then you start to move into tools, like we have Logic Apps. If you look at kind of on the Office 365 side, we've got Power Platform and Flow. And so that's kind of interesting because not only are you is it low code, but we're now taking the shadow IT people and the citizen developers kind of, end like, pulling them out and saying, "Hey, you can now create applications in the cloud."

So it's kind of like democratization of coding now. And so that like, this is even a bigger topic. But, okay, now coders are now engineers. And then citizen developers are developers. And then it opens it up to a lot of people being able to create things really quickly, get them into the cloud, faster. And like on... I don't know what the sort of equivalent would be kind of in the other clouds, but for us, like, I mean, with Logic Apps for example, I've built a ton of Twitter bots. I have so many Twitter bots out there. People have no idea. Like if there's an account out there is tweeting like adoptable cats it's probably mine. And who knows what cloud platform it's on because I've deployed them all out to different cloud platforms. They just run forever, they don't cost me any money.

But for the longest time, though, I was having to like copy paste my Twitter API code across all these apps. But now I just, I'll just use Logic Apps. And I just sort of like, drag it over my Twitter connector, and then just connect that and then maybe I'll write like a function to do something really simple and then... So in it, I have tons and tons and tons of bots now coming out of Logic Apps. And, you can still, I still set up, I have like an ARM template for my Logic Apps and so I still have like, somewhat of a, like a baby CICD process where I deploy those. And I don't really have much code, so I'm not really writing any unit tests there.

Jeremy: And I think that's interesting, too, because I mean, if you look at services, maybe like a kind, I don't know, like Airtable, for example, where you can just drag it and you create a spreadsheet, essentially, or a database and then create forms for it, and then reports. I mean, you mentioned API Gateway earlier. And that's one of those things where, API Gateway and then AppSync, which is the GraphQL, sort of I guess version of that. You right now you can set up your authentication, right? And you use Cognito, or OcZero or something like that. And then you can set up all your routes, so it knows where it needs to go. And you can do service integration. So you can say just pull data, right from the database and or put it right into an SQS queue or something else. And I think that that is really interesting, because a lot of those things that we do now where we write a Lambda function, or we write an Azure function or something to process some data just to move it somewhere else. All of that stuff is starting to go away.

And so that's why I just I bring it up, because I think that those are really interesting tools. And if you can solve those problems, where it's like, you don't need a developer now to write a simple API to maybe, put data in and things like that, but you need them to do something more interesting afterwards. I mean, I'm not saying developers are going away. I think there's more need for us now than there ever was. But it was I find that interesting. And then I think, with your Logic Apps example you know like AWS has like step functions, which is like their state machine type things. And they also have their serverless application repository where you can create like little reusable snippets that you can do all kinds of processes for you. And so I think that's really interesting, because you're right, you don't want to be using the same Twitter API over and over and over again, copy and paste that in, because then it changes in one. And then all of a sudden, you're adopted cats things breaks because the Twitter API...

Linda: Yeah. And like, all those cats don't find homes because I had a terrible deployment process.

Jeremy: And that would be a tragedy.

Linda: Like if there's any reason to have a good developer process, it's for those kitties okay?

Jeremy: It's for the cats. Make sense. Alright so, one other thing. So I do want to talk about testing and I know there's holy wars on this. Especially now with testing in the cloud and things like that. So we don't have to get into the specifics of exactly how we might do it, but just testing overall. I mean, obviously a good idea, and I just don't think enough people are doing it.

Linda: No, I mean, like, a lot of people aren't doing it at all. Maybe most people aren't doing it at all. I mean, I think I mean, I guess most things that are public on GitHub are not production applications. I mean, the production, especially enterprise applications, like those are going to be private. But you almost never see unit tests on GitHub. Like when I go to people's projects, and the usual things that I'm using I don't...

Jeremy: You do on mine. My projects, I write tests, I write too many tests.

Linda: For my demos, like for my conference demos I'll have a test that runs but usually it's just like printing something out. I'm like super lazy about conference talk demos which goes against everything I'm going for here, right? Because like if that demo breaks before a talk, that is devastating. I mean, that's my equivalent of a production app breaking. But yeah, I mean, I think that kind of goes back to, this like muscle memory. I mean, you don't have to go full on test driven development or I think Forrest was saying you need to, like, have those tests in there first. And I mean, that's really good practice. But I think it goes into that analysis paralysis, where people are like, I don't know how to write a test for this. And so you just put something in there, and then it will start to come to you, as you're like writing code, or you're connecting services or you're building something, it will start to occur to you, okay, if, you don't need to test the services themselves. But you start to say, okay, this is where things might break down and then you kind of insert tests there. You don't necessarily have to have it planned all out it in the beginning from my opinion, I mean, I think that's great.

But, I've fallen into a lot of traps with development teams where you try to teach them best practices, and it just gets to a point where they can't move. Because they're like they can't develop because they don't know how to write the test first or... Like I worked on a team that tried to implement required pair programming. And all the introverts were like, "No, we're calling in sick." So I don't... You have to understand human nature there to some extent. But yeah, I mean, as much as you would test any application, it's no different with serverless. And if you've been testing zero, then okay, you should at least test 5% or something. But, I think it is really important. And I think people get into a real rabbit hole too, about unit testing versus integration testing, and the cloud and testing frameworks and all that. And that's like, the thing that I don't want to argue with people about but just do something.

Jeremy: Do something. Yeah, I totally agree. And the funny thing is, is that, for a while I was big into TDD. And I was like, I write my tests up front, and I was very good. Did it for a big project and it worked really, really well. But the problem is, is that you're just... You're right, it traps you, you're like, now I'm writing tests, like I just wrote of these tests and I haven't written any code yet, right? But I just spent like five days writing tests. And often what I do and this is just something that if you have tests are you haven't written tests right now, this is just some advice I'll give. I love using like a code coverage tool to just go back and show you which statements and which branches haven't been run, and then just go in and just write some tests that make sure that those different parts of the code are tested.

As I know, for me, I'll stare at something and I'll have like these nested ternary operators, and I'm like, wait a minute, how do I even trigger that branch, like I don't even know how to get there. And so being able to write that and reason about and then sometimes it's good to go back and look at your test to figure out how code works sometimes, which is always a good thing. Alright. So one more thing about testing in general. Like do you think that the testing culture has changed? I mean, is this due to the cloud? Is it due to serverless? Containers?

Linda: Oh, it's definitely changed. I mean, but it's changed in the same way of people selling their soul essentially as I was saying in my talk. I mean, if people weren't testing before, then they're not testing now. But there are a lot of people that were TDD enforcers that just suddenly don't write anything, because they just like, don't know how kind of. But I mean, really, I think if you start from the beginning, if you start with unit tests, you can still use whatever your testing framework of choice is like JUnit or Mocha or whatever. If you start there at least writing unit tests and writing them to test your code, then you can at least that's something, that's like a good building block, and you can work from there to start and say, okay, what does integration tests look like? How do I test these interactions?

And this is sort of back to what we're talking about before too. Like, the basis of serverless, in my opinion, Is that it's event driven? And so for me serverless isn't Lambda or functions or Logic Apps or whatever. It's about like this event driven paradigm. And it just happens to be fully managed. And it happens to be kind of to support developers writing less code. There's all these other features. But it comes back to the to event driven architectures. And I think that create... That's kind of a different, like kind of a mind switch, I think for a lot of people, too. And so that makes it a little bit harder to wrap their brain around how they might test those things. But that also allows for you to go more service to service to service with events and not have to write a lot of code, but at some point maybe you're ingesting or using some code to kind of figure out like what's inside of this event? What you want to do with it and pass it on. And those are the cases where you definitely want to be testing that.

You don't want to lose some parts of your message because like you didn't test your message in gesture.java or whatever.

Jeremy: Well, the black box of I think that's one of the reasons why people love to write functions in serverless applications, is because there's just, it seems like a black box, like if it goes from service A and service B to service C, and nobody ever sees it I mean, it's kind of like, what happened there? And it's funny...

Linda: Like, why do I commit to GitHub if I'm just doing ABC? I mean, I've fallen into that before, too. And It's just like a Terraform.

Jeremy: It's just a config file, exactly. That's kind of crazy. I mean, and I think the advice always there has been like use functions to transform not to transport, right? So if you're just moving down from service to service that's not the best way to do it. So speaking of the best way to do things, just maybe wrap this up. I mean, off the top of your head, what are the sort of the most important, I guess, processes or best practices for a serverless developer?

Linda: Well, I think it's the same as a developer developer, I think that all the things that we've been kind of drilling into developers forever. And I mean, you're the same age as I am, so like, we've seen everything. We've seen every shift possible and even shifts within like Scrum and how Scrum affects development. And there's all these weird trends in development. I was thinking earlier, when you're talking about testing, that there was a trend where developers would have the customers write the unit tests for them. And that so that way, the developers would know what to write. And I mean, that was weird. I mean, I understand why they were doing that, but that was weird. And so but I think like, do what works for you, right? You're a developer, like you're a trustworthy individual who knows how to write applications and you're transitioning to the cloud.

Everything that you've been doing, it is probably correct. And now we just need to tweak your process. We want you to be more efficient. We want you to use services that are trusted. And but ultimately, I think if you're using your IDE that you love, you're doing your static code analysis in one way or another, you're writing some amount of unit tests and integration tests, and you have a repeatable CICD process, revision control, of course, is always in there. I have to stress that too because that's like, there are a lot of people that write serverless applications that never commit code. And that just blows my mind. Because if you're writing a traditional web app that maybe you deploy to a pass or something, you'd always commit it somewhere. So I think just yeah, I mean, go with what you know. I mean, that like, Forrest talks about the familiarity bias, like there's some good to that too, because like you've been doing things in a repeatable way that have been very successful, continue those, even though you're in this scary cloud environment.

Jeremy: Awesome. Alright, well, listen, Linda, thank you so much for being here today, and sharing all of this awesome knowledge. If people want to get a hold of you or find out what you're working on, how do they do that?

Linda: So I'm pretty much "lynnaloo" everywhere. L-Y-N-N-A-L-O-O. I don't I'm... I find that I'm like bad about checking Twitter DMs, but like, and they're not open because there's just like, not nice people in the world. But if you just mentioned me on Twitter, I'll find it. You can email me. GitHub. I'm out, out in the world. I think probably just a Twitter mention is probably the best way to poke me. And I'll find it.

Jeremy: Alright. Well, I will get all that into the show notes. Thanks again.

Linda: Alright. Thanks.

THIS EPISODE IS SPONSORED BY: Datadog

View Details

About Mike Roberts:
Mike Roberts is a partner, and co-founder, of Symphonia - a consultancy specializing in Cloud Architecture and the impact it has on companies and teams. During his career, Mike’s been an engineer, a CTO, and other fun places in-between. He’s a long-time proponent of Agile and DevOps values and is passionate about the role that cloud technologies have played in enabling such values for many high-functioning software teams. He sees Serverless as the next evolution of cloud systems and as such is excited about its ability to help teams, and their customers, be awesome.

  • Twitter: twitter.com/mikebroberts
  • LinkedIn: linkedin.com/in/mikebroberts
  • Website: mikebroberts.com
  • Symphonia: www.symphonia.io
  • Symphonia blog: blog.symphonia.io
  • Programming AWS Lambda: Build and Deploy Serverless Applications with Java book: shop.oreilly.com

Watch this episode on YouTube: https://youtu.be/16en-TTGNhk

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. This week, I'm chatting with Mike Roberts. Hey, Mike, thanks for joining me.

Mike: Thank you very much for inviting me, Jeremy.

Jeremy: So, you are a Cloud Architect and DevOps Consultant that specializes in serverless and AWS, and you're also a partner at Symphonia. So, why don't you tell the listeners a little bit about your background and what Symphonia does.

Mike: Yeah, that'd be great. So, I've been in industry now for 21 years and in that time, I've been an engineer or a senior engineer, manager or CTO, sometimes consulting, sometimes working for product companies, so a whole mixture and sort of up and down the manager versus technical ladder.

About four years ago, I was a VP of Engineering at an ad tech company here in New York and we started using a lot of sort of much higher level AWS technologies and especially at the end of that year, we were using a lot of Lambda, so I really thought that serverless was really interesting and so I wrote an article four years ago now about serverless. That proved to be really popular and I was like, "Oh, wait, other people like this, too. Maybe I should start a company about this kind of stuff." So myself and my business partner, John Chapin, we decided to start Symphonia as a consulting company to help people with the kind of technologies and lessons that we'd sort of seen over the last few years. And that's what we've been doing now for three and a half years.

Jeremy: Awesome. Alright. Well, so recently, you and your business partner, John, wrote a book called Programming AWS Lambda, and great title, right, there it is. He's got it. Okay. Now, the thing that struck me though about it was about Java. And so I'm just curious, it's 2020 and so, why would you write a book about serverless programming in Java?

Mike: Mostly because my writing is terrible and I didn't want anyone to actually read the book. No, that's not the reason. It is weird and a lot of the things that you read about Lambda, the examples are in Python or JavaScript or Go and then there's this Java thing. And who actually uses Java with Lambda? Well, it turns out a lot of people use Java with Lambda and the other thing was, it's how we got started with Lambda. So when John and I started using Lambda, which was about three and a half years ago, the Java support has just come out and we were working for a Java shop, so we had a lot of engineers who were very Java savvy. We had all of our Java tool chains all sorted out and so we decided to use Java and Lambda and see how it worked and it worked brilliantly.

And one of the reasons it worked brilliantly was that the system that we were building was pretty high throughput, like we were processing millions of messages a day with Lambda and so we never hit, and even back then, any of the concerns with cold starts or anything like that, and so yeah, it really just fitted in like a glove for us and. And so, when we decided to write the book, we knew that we weren't unique and we knew that there were a lot of other people out there who have built up this knowledge in Java and the ecosystem that surrounds Java and we wanted them to have a book for Lambda, just like JavaScript developers and Python developers and all that kind of thing.

Jeremy: Awesome. Well, so the funny thing is, is that I saw this book come out and I immediately was like, "Oh, no, it's a book about Java." And I haven't programmed in Java, and I don't remember how long, but I said, "I know Mike and I know John." I've been following your work for a couple of years now and I know you produce good stuff. So I said, "I'm going to look at it. I just want to give it a look." And what I found was that it's not really a book about Java. It's really a book about building serverless applications with the examples in Java and there are a few very Java specific things in there, which I think is actually great and we'll get into some of those reasons why.

But yeah, but I mean, the book covers everything. All those core concepts like the execution environment and invocation types, logging, timeouts, memory, CPU, environment variables, all those things that you would want to know and it gets into detailed explanations about deployments, infrastructure as code, security, event sources, so it really is a much more complete reference. So, if you pick up this book or if you don't pick up this book, because you're like, "Oh, it's about Java."

I actually would really suggest that you pick it up and just read some of the core concepts because I really like your take, you and John's take, on just some of these different concepts because I think there's a lot of, I don't know if dogma is the right word, but there's people who approach things a certain way and they just sort of think that's the way to do it, but when you start applying those things to real world situations and real applications, you start to butt up against some of the limitations. Anyways, you had some really interesting thoughts on that.

Mike: Well, thank you.

Jeremy: So I want to get into the book because there's some really interesting things and I know I'm talking more than I probably should be here, but this is something that I found really interesting was your approach to testing.

Mike: Yeah, and it's interesting, because John and I, we actually only met about five years ago and we both been working 20 years, but the way that we approach Development and Engineering in general is extremely similar. I mean, I come from 20 years of extreme programming and test-driven development and that background and John does to some extent, but not quite as in... I was working with people that were speaking and writing about this stuff back 15 years ago, but yeah, we very much felt the same way.

And while I don't, I'm not a test-driven developer all the time, I don't always write my tests first. What I learned from when I did write that way 15 years ago is to rely on unit tests and functional tests that are of a specific type and John feels exactly the same way. In fact, John was the primary author on that chapter in the book and it's interesting because I completely agree with everything he wrote in there. And so the way that we think is that when it comes down, especially with Lambda functions, Lambda functions are just code. A Lambda function is a piece of code that accepts some JSON, might interact with some downstream systems and might return some other JSON.

That's it. There's nothing else to it and what that means is that you can run unit tests and in process functional tests and the word - the naming gets a bit weird - very, very simply. And so what we rely on a lot with our testing is writing tests that run within the I.D. or within the process, the same process with the tests and the code under test. We use other forms of testing, which we'll get into maybe in a little bit, but our primary goal is to say, "Let's prove that what we've done works using unit tests that all run in one process."

Jeremy: So yeah, because you read a lot about things like the testing pyramid, right?

Mike: Yeah.

Jeremy: And this idea of integration tests become really expensive where unit tests are really, really cheap because you can run them quickly and you can get that feedback, but then there's this sense that in order to do the testing, you need to have all of this stuff running in the Cloud and you need to do these end-to-end tests and I think I'm with you here, where I like to write a lot of unit tests that make sure my business logic works because I mean, really, the business logic is the interesting part of your application, isn't it?

Mike: Yeah. I have a feeling that I know where some of this comes from, which is that a lot of the complexities, especially when you're learning serverless development, are not about the code. It's about this brand new platform. It's about all the services that you're integrating. It's about how you deploy all that kind of stuff and that's the stuff that's new. And so people I think, naturally go, "Well, that's the stuff that we want to test," but that's the stuff that is kind of hard to write automated tests for, that's why you need all of these long running integration tests and people focus on that.

And that's understandable when you get started with this stuff, but as engineers, we have to really think about wearing two hats when we're when we're writing software that's going to last a long time. There's our experimentation hat, which is, how does this stuff even work at all in the context of the platform that I'm using and then there's the, "I'm now writing software that is going to last a number of years." And those are two different modes of thinking, figuring out how this is going to work and then writing the production code.

And when I think about testing, what I'm talking about, is writing tests with the production code in mind and really separating out those two things. I think what the trap that quite a lot of people get into is they do this experimentation stuff and then they write some code, and then they don't sort of switch gears. They don't come back to the "Okay, what is the actual code I need to write? Where is the domain logic in that and what are the tests around that?"

Jeremy: Yeah. And so then when would you suggest though that people write things like integration tests, and actually we should take a step back because you mentioned a term in the book called "Functional Test," which is something where I don't think it's a standard term maybe and maybe it is, I mean, but it's something that I typically don't hear and it's this idea of basically, it goes beyond the idea of the unit tests to test the business logic, but goes more to testing, not really the integration, but maybe the integration point. I mean, can you explain that a little bit better?

Mike: Yeah. So the term it definitely is one that has been around a long time. However, I would say that it's a term that not everyone agrees on the definition of it. The way that we've defined the term in the book is that like unit tests, when you're running a functional test, everything runs within one process, so you're not calling out to external processes from either your code or your test code or your code under test. And so from a structural point of view, when you're actually running the test, the unit test and the functional test feels very similar. However, what it is that you're testing is quite different.

With a unit test, you are testing an individual method like a language method or language function within the code and it's completely in isolation. And what you're doing when you do functional testing, which is a little bit different, is you are testing how a bigger part of the application is responding to the system around it, but the important thing is that you're not actually running the system around it, you are stubbing out the external environment. So in that case, we would write an internal stub for DynamoDB, if we had a Lambda function that was calling DynamoDB. That's not anything clever. That's not a DynamoDB stub library, that's just literally us like saying, "Okay, what is DynamoDB? What are we assuming that DynamoDB returns and is our code processing that response correctly?"

Jeremy: Right. And that's another point you make in the book, too, where you say things like local stack, or any of these sort of local mocking libraries are essentially a bad idea and you go into it a little bit more in the book, but I mean, I feel like using those sort of systems for building local tests, great for experimentation, right? Like if you just want to check and do something quickly, but once you rely on those then you have more complex testing setups and things like that. But I mean, if you are just calling an API, an API is returning JSON to you, right? So just simulate that JSON and test against that and then you know that JSON that you're testing against is always the same and it isn't going to change because of some update to local stack or some other local mocking library.

Mike: So I think local stack has its place and local stack is excellent when you've got your experimentation hat on and you're wanting to do lots and lots of really quick iteration and you don't want to be constantly deploying your system to the Cloud. And I get that, like deploying to the cloud is 10 times longer than deploying to local stack, and so, using local stack as an experimentation system is brilliant, but that's not testing. That's experimentation and that's where I'd like people to take that experimentation hat off and put your testing hat on. And when you put your testing hat on, you're writing a system that is probably going to last years.

Local stack is a simulation and an occasionally good simulation of it. There are things that local stack doesn't simulate properly about the Cloud and it's also a lot slower than running functional tests that are all in process and so you sort of have, it's great for experimentation, but it's almost like the worst of both worlds when it comes to regression testing.

Jeremy: Right. Yeah.

Mike: I know some people might find this word offensive, but it's a little bit of a, and this isn't about an age thing, this is about a maturity thing. There is a maturity of doing serverless development. Once you've got used to it a little bit, you need to be like, "Okay, now we need to think about testing versus experimentation is a different thing." Because I've seen people get into all kinds of messes where 90% of their testing relies on local stack and A) it's slow, once you've got like 100 tests and B) they're relying on shifting sand and for regression tests, that's really too much of a risk in my opinion.

Jeremy: Yeah, I totally agree. And then the value of integration tests. And again, I think to clarify, I mean, the integration tests are important with serverless applications. I mean, there are a lot of different connectivity pieces or a lot of services communicating with one another, right? So you have API gateway calls a Lambda function that writes data to DynamoDB, that triggers a stream that loads another Lambda function that sends a message to EventBridge that triggers four more functions or something like that. So, there are certainly complex workflows, but simulating those locally with these functional tests is basically saying, "If this Lambda function gets this, does it do what it's supposed to do?" That is relatively easy with those functional tests, but what about those more complex like actually seeing that go all the way through?

Mike: Yeah. And what integration tests are about are testing your assumptions effectively. So when I talked before about functional tests, I said, "So we're going to mock or stub the response that comes back from DynamoDB and make sure that we're doing the right thing with that." That makes an assumption that we've correctly defined what comes back from DynamoDB. And so what integration tests do is validate those assumptions, they validate how you expect your code to run within the larger environment and the larger platform.

And we absolutely advocate for doing that, but remembering that running and maintaining integration tests is a costly exercise. They take a long time to run and they also take a long time to maintain because things change over time. And so, we put a lot of work into the integration test section of the book. And John did this extraordinary thing with Maven, and those of you that are Java developers understand this, but where we run Maven test, which is one command line, and what that does is it brings up an entirely new stack of all of the components in our application, runs all the integration tests against it, and then if the tests work, then it immediately tears that stack down.

We wouldn't have gone into that amount of effort to get that stuff working if we didn't think integration tests were valuable, but we also understand that because they're expensive based on in terms of our time and computer time, that we want to minimize the number of those that we write. And so we're looking normally at just a few, but capture hopefully a number of cases, but again, we're thinking, we're not testing the code when we're writing integration tests. We're testing our assumptions about the larger environment. If we want to test the code, that's when you write a unit test or a functional test.

Jeremy: Right, plus that feedback loop is just so much faster.

Mike: Yeah.

Jeremy: So you mentioned TDD a little bit earlier and I am a big fan of this, but I am a horrible practitioner of it because it's one of those things where it's like, "Yeah, great. If I don't know what I want the code to look like yet, but I know what I want it to do, then it sort of makes sense to do this." But what are your thoughts on TDD and especially as it applies to serverless because you mentioned this idea of experimentation versus production mode and when you're doing experimentation with serverless, I feel like trying to do TDD is really tough.

Mike: Yeah, it is. And I think that there's a difference there between experimentation around, how do we expect the larger environment to respond versus how do we want to write out the main logic? Now, again, we hit these things slightly weird when we're writing Lambda code because when we're writing a large container-based app, it's really easy to see the domain logic, like it's all that. There's a little bit around the edge that's not domain logic, but most of it is domain logic. But when we're writing a Lambda function, it feels like there's all this other stuff around the edge and there's only a little bit of domain logic and sometimes that's true, but oftentimes, it's not. Oftentimes, that stuff around the edge is something that as we write more and more Lambda functions, that's going to get refracted into libraries or whatever.

Jeremy: Right.

Mike: And so the experimentation part is like, "Okay, so what is the JSON that Dynamo is going to respond to me when I make a request to it?" or whatever. That's the experimentation part. And then the TDD part is, "Okay, given that I'm going to get this request from the user and I'm getting this response from DynamoDB, what do I want my code to look like given that that's what's going to happen?" And so sometimes, I do TDD occasionally. Sometimes it's like, I have no idea what I want my code to look like, but I know I have those inputs, so let's start with a test. And remembering that TDD is test-driven design as much as anything else. It's about: how do I design my code for testing. And unit testing is great and we need unit testing, but TDD is a mode of thinking that I use occasionally where I want to be like, "Okay, I don't know how to write this code for testing."

And the good news is that you don't have to make your code testable, you don't have to do TDD. If you haven't done a TDD, you can always refactor it for testing later. And one of the things that John did in the testing chapter in the book is he took one of my earlier examples that was not written with testing in mind, whatsoever, and what he did is he actually updates the code first to allow it to be more easily testable. And then we have the best of both worlds. So effectively, I wrote in chapter five, I've written the experimentation mind part, and then John sort of have done the switching hats in chapter six.

Jeremy: Yeah, and I think that's super important, too. I mean, just thinking about when you're writing code, is it going to be easy to test and I mean, that's where things like hexagonal architecture comes in or your portion adapter, things like that, where you really are separating out so you don't have to call that Lambda function handler in order to invoke business logic that you can test that outside of that and test those different things separately. So, yeah, I think that's super interesting stuff. Alright, so I want to move on to another thing you mentioned in the book. And this is something that comes up all the time, and this is cold starts.

Mike: Yeah and it's funny. So just as an example, we did a signing of an early version of this book, a conference back in February. Do you remember conferences, Jeremy?

Jeremy: Yes. I remember. Have you ever watched something on TV and you see a crowd of people and you're like, "I don't think they're supposed to do that. What?" It's just now that's the mindset.

Mike: I saw a trailer for a movie that was just coming out, and I'm like, "When did they film this?"

Jeremy: Exactly.

Mike: Wow. Anyway. Yeah. So, I was at a conference in February and we were doing a signing of an early version of our book and about 50 people came up because O'Reilly were giving away free books and people love free books.

Jeremy: Sure.

Mike: And if I had a nickel for everyone that said, "But what about cold starts?" And it was like 60% of people said, "What about cold starts in Java?" And cold starts is this... I mean, you know this Jeremy, it's like, "Is the Boogeyman in the Lambda, well, anyway." And so yes, everyone thinks that basically because of cold starts Java is a non-starter for Lambda and obviously that's not true. Otherwise, we wouldn't have written the book because we would never have had production Lambdas. And I think a really good case in point is I was working with a client last year and the year before and they were just writing Java Lambdas, they were writing Scala Lambdas and for those of you don't know, Scala is another language that runs within the JVM and Scala is even worse for cold starts because not only do you have to start with JVM, you also have to start the Scala runtime within the JVM.

Jeremy: Right.

Mike: It was not my idea, but they were very concerned about cold starts, but they were a team of Scala developers, they knew Scala really well. So very, very concerned, and I'm not sure it was going to work. Anyway, they put it in production and a month or so later, I saw some announcement about AWS cold start, this was pre the VPC stuff, so it was something else. And I went up to the tech lead on the team and said, "Oh, hey, by the way, there's this improvement to cold starts coming out." And he looked at me, he was like, "What?" I'm like, "Well, you all are worried about cold starts?" He was like, "Oh, no, no. We put it in production. It was fine."

And this happens nine times out of 10 when we talk to people. Cold starts are these big scary things because when you're in development with Lambda, every time you run your new Lambda function, you see a cold start.

Jeremy: Right.

Mike: And there's this like feeling that that's like every single time your Lambda function is going to get called in production, you're going to get a cold start. Well, if your Lambda functions are running frequently enough for big, serious applications, that's true, then you're actually going to be getting cold start like one in 100,000 times or whatever it is. And so when you amortize the cold start over how many times your Lambda function is actually running, normally in many, many, many situations, it's not a problem even if the cold start was 10 seconds every time, it's not a problem. So that's one thing.

And then the other thing is that cold starts are not as bad as people think they are, especially, we have a number of mechanisms in the book that we recommend people use. We don't completely dismiss cold starts and we spend a lot of time saying how you mitigate them. But if you think a bit about how you're going to package your code and write your code and architect your applications, because you do have to do that. You can't just throw the whole of your typical way of thinking of it because you will end up with 15 second cold starts and that's not great.

Jeremy: Right.

Mike: With a little bit of thinking and a little bit of smarts then you can then you can fix that and this has nothing to do with Java, but especially now Amazon had fixed the VPC issue with cold starts.

Jeremy: Right.

Mike: Cold starts are just not nearly as much of a problem as they used to be.

Jeremy: Yeah. And one of the things, too that I always notice is when it comes to front-end colds starts, I mean, those are obviously more obvious, too. If you're connecting to a Lambda function through API gateway or ALB or something like that, you get that cold start and it's fairly noticeable. Like you said, on an application in production that gets a fair amount of traffic that you usually don't get those cold starts, but even when you do, and I know I think I've said this like a 1000 times like how many times have you typed something into Google and it just didn't respond for some reason, right? There's a network hiccup or something happened. I mean, if it's 10 seconds, it's kind of insane, but I think users will be like, "Why is this not responding?" And then they click refresh, or whatever and then what do you know, it comes up just fine.

Jeremy: But one of the things that I always noticed is, I think the vast majority of my Lambdas now run asynchronously in the background, right? So they're not even hitting user face or you're not hitting them directly and really, when it comes to that, cold starts, they don't matter at all.

Mike: Right. Exactly and again, this is a little bit of the problem that comes from a lot of the places that people start doing serverless development is not how they're thinking when they're writing production applications. The tutorials that you see are all APIs, because those are easy to test.

Jeremy: Right.

Mike: We understand we can just hit it with a web browser, so it's easy, so it's very similar to this testing issue. It turns out that, you know this, Jeremy, but, where Lambda really shines is in large-scale, back-end asynchronous systems and that's how John and I got started with it. It was kind of lucky in some ways and that's how we got started, because it made it gave us this mindset, like our first real Lambda round was processing events from Kinesis and was processing millions of events a day of Kinesis, and a bunch of events off S3, like we weren't doing API stuff there. And then if like one in 100,000 of your invocations takes 10 seconds instead of half a second, do you care? No, you don't care.

Jeremy: Yeah.

Mike: One thing I would say for people listening to this, if all you've ever used Lambda for is synchronous APIs then you're missing like 95% of what Lambda is about.

Jeremy: Yeah, and the other thing, too, and I'm not a huge fan of Java. From my first class in college of remembering public static void main, I've just had nightmares about it ever since.

Mike: Yeah.

Jeremy: But I will say this, the cold starts in Java are certainly higher than something like Node or Python and so forth. They barely ever come into play, but the other thing is that once a Java function is initialized and it's warm, it is fast.

Mike: Exactly. And that's one of the other reasons that we liked using it for high-throughput systems because obviously, Go is going to be Java because Go is compiled down to real code.

Jeremy: Sure.

Mike: Yeah, you compare the JVM with JavaScript or Python, it's going to be faster over time and the thing that people forget about that is that that means it could be cheaper, like if your Lambda function takes 800 milliseconds in Node and 700 milliseconds in Java and you'll run-

Jeremy: Times 10 million or whatever it is.

Mike: You're saving 12% on your compute costs.

Jeremy: Yeah.

Mike: Purely by using a different language. Now, would I tell people to use Java based upon that, solely that? No, but if you have Java experienced engineers on the team, then that's a really nice benefit if you're writing high-throughput Lambda systems.

Jeremy: Absolutely. So then another thing too, about cold starts, is because so many people have, I think, complained about them or I guess, thought they were a problem and there's other reasons for this, too, but AWS came out with provisioned concurrency and you write about this in the book. You have some interesting thoughts about this.

Mike: Yeah. It's interesting. I'll start off by saying, I'm glad that AWS did this because I've met people in this world who are like, "No. Cold starts are a problem. Must always have absolute guaranteed latency." I'm like, "Okay." And very occasionally those people actually need that, and that's fine, but normally they don't. And so I'm glad that AWS have come around because now if needs be, I can just point those people at this and say, "Fine. You have your escape hatch. It's called Provisioned Concurrency."

But oh, my word does Provisioned Concurrency come with some caveats. And the first one was my very first experience with it. It's really slow to deploy. I can't remember the numbers now, but I did some testing and this was in December, so it was only just after it came out, so I'm sure this will get better.

Jeremy: Sure.

Mike: But it took like an extra minute and a half, two minutes to deploy a single provisioned concurrency Lambda function and it took an extra four minutes to deploy something where the Provisioned Concurrency was set to 50. So that was really annoying, because I'm used to my little Lambda apps taking, well, less than that in total for it to deploy.

Jeremy: Exactly.

Mike: So, that was really concerning. The next thing is that the costs around Provisioned Concurrency are troublesome for two reasons. One, and lots of people have already talked about this, which is that the nice thing about Lambda is it's pay per use. You only pay for what your Lambda function is actually doing whereas with Provisioned Concurrency that is broken, like you are always paying a flat fee for your Provisioned Concurrency fee to the point where when AWS launched Provisioned Concurrency, they also showed how you could manually auto scale Provisioned Concurrency.

I'm like, "Wait, we're going backwards here. This is the wrong direction." So that's another part of the problem is like you have to start thinking in terms of like old economics and one of the benefits of Lambda is we don't think about those economics anymore. The other problem with the cost of PC is it's expensive.

Jeremy: Yeah, it is.

Mike: It can get really expensive.

Jeremy: Right.

Mike: So that's problem number two is the cost. And then the third problem, this is frustrating because this is my OCD kicking in a little bit where when I set up a SAM template or whatever you want, the difference between the development configuration of my Lambda versus my production configuration for my Lambda is often precisely the same, maybe different environment variables. With Provisioned Concurrency, you don't want to be using the same Provisioned Concurrency settings in development as you do in production and so you're mixing up this whole thing down into. It's just, yeah.

Jeremy: Yeah, I agree.

Mike: I love the fact that they managed to make it work without any code changes and that was very clever.

Jeremy: Yeah.

Mike: I love the fact that there is now an escape hatch for people that really, really can't have cold start, that's great, but it's something that you should almost, almost never use. And I think the major benefit for someone like me and I think this would apply to many others as well, is just the sort of ramp up that Lambda functions can do. They only scale so much so fast, like it takes five minutes or something like that in order for you to go up to the next 500 of them or whatever it is. And so, that's something that we think about Lambda being infinitely scalable, but in actuality, there's some limits to how fast that can scale.

So having something like Provisioned Concurrency is great to say, "Hey, I need to warm 2000 functions for some big flash sale that I'm having at noon," or something like that, that that it would come in very handy in cases like that, but I just was playing around with it and I'm like, "What if I just kept one function warm or one container warm?" And I forget, I was either $14 or $17 or something. Basically the cost was or maybe $10, whatever it was, but it only got hit when I actually hit it, it would cost me like s$0.6 to run that Lambda function for an entire month. If I use Provisioned Concurrency, it would cost me $10, right? Yeah.

Jeremy: Which is not a lot of money, except if you multiply that by 1000 functions then all of a sudden things start to get more expensive, plus, if you multiply it by saying keeping 50 warm as opposed to just one, so it does get pricey.

Mike: Yeah, and to be fair to AWS, I don't think they particularly wrote it for, built it for cost conscious companies. It's like I think they built it for big enterprises, frankly.

Jeremy: Right. Yeah.

Mike: But that's my guess.

Jeremy: Yeah.

Mike: And so, that becomes less of a deal there, but it does become a deal when you're doing that for provision currency of 50 and you're deploying it five times for multiple environments, then it can ramp up, so yeah.

Jeremy: Yeah. Interesting. Alright. So the other thing you mentioned in the book is sort of when to or when to not use custom runtime. So what are your thoughts on those?

Mike: Yeah. This is interesting. I was actually just using one of these this morning. So yeah, for those of you that don't know that Java comes with however many, it's like 10 standard runtimes now normally for different languages and different versions of languages. In fact, if you take them to multiple versions, you're up to like 30, 40, whatever it is. About a year and a half ago, I think it was, the Reinvent 2018, Amazon came out with a capability where you could basically write any runtime that ran on Linux. And so a number of people came along with specialized runtimes for different environments.

And so one possibility that you have with cold starts is to write your own or to deploy your own custom runtime that is configured in a different way and maybe solve some of these cold starts issues for you. And there's a couple of ways of thinking about this. One is if you're a large organization, and you have a standard VM setup that you want to use, virtual machines, Java stuff, that's different to the Amazon way of doing things then you can use your organization's Java runtimes as opposed to the Amazon runtime. The other option, another sort of way of using these things, which I haven't dug into, but I'd like to, is that there are alternative ways of running Java code other than the stock JVM.

Jeremy: Yeah.

Mike: So, one that's talked about quite a lot is a thing called Graal, G-R-A-A-L and what Graal does is at build time, it will take your Java code and actually produce something that doesn't run in a regular VM, it just runs as a regular executable and so the idea there is that your startup times using Graal are significantly faster. I think I've got that right. I think that's what Graal does. It's one of those things I want to dig into it and try it out.

Jeremy: Right.

Mike: There's also other alternative VMs that just start super quickly, so I think Graal is one that compiles down to real code, but in those situations, so yeah, that could solve it. So then your thing might be, "Okay, well, if cold starts are ever a problem, we'll just use all of these." Well, the thing is, then you have to maintain the use of these runtimes. If something about the platform changes, the Lambda platform changes, you have to update your runtime.

Whereas if you use Amazon's Stock JVM, when they want to update the underlying Linux environments, whether there's like, we get another specter or meltdown or something or doing something else to get a bunch of performance improvements, we get those improvements automatically if we use a standard runtime. Whereas if we use a custom runtime, we probably don't get that and probably want to have to go through a whole bunch of testing against those new environments before we roll out on new runtimes, so nothing comes for free.

Jeremy: Yeah.

Mike: So, it's one of those things. For some people, it's going to be worth looking at the tradeoff.

Jeremy: Yeah, I mean, I think that's one of those things, too, where it's like with serverless you're trying to minimize all of that undifferentiated heavy lifting, so why would you want to go and maintain your own runtime? If you're a big organization, like you said, I think this makes total sense where you can bake in security and other things that you might want to do, but certainly for the average developer or the average company, I think it's something big to bite off, to chew, that's not the right way to say it. It's too much to bite off, I don't know, maybe that's the right way.

But anyway, so I want to get into some really geeky Java stuff here. And like I said, I'm not a Java person, but I did work for a company that everything was written in Java. So I did have to look at it quite a bit. So one of the things that's very, very popular with Java, especially when it comes to building APIs is the spring framework. And AWS has spent, I think, or has invested a significant amount of time and energy into something called the Serverless Java Container project. And this is something they maintain that makes it easy for you to write Java Spring Boot projects or whatever they're called on AWS Lambda. You are very, very clear in the book that you think this is a bad idea.

Mike: I am. So a little bit of context for those of you that have never and will never write Java applications. Back in the dawn of time, otherwise known as about 2001, those of us that wrote software for a living would normally write Java and we normally run Java in these large things called application servers that would take minutes to start up and we'd all run these on our laptops and run them in production and they were horrible and very, very slow. But they allowed doing certain things that at that time would otherwise be a lot of work for us as developers.

Over the course of a few years, people were like, "These are getting way too big and slow and heavyweight. Can we come up with something simpler?" And along came this thing called Spring and Spring tried to still be this idea of running a large application and doing a bunch of stuff for you, but started up in 10% of the time, perhaps even less, and it really became over the sort of next 10 years, the de facto way of writing Java applications. But it's still based on this idea where you are starting an application bringing in a whole bunch of things and dependencies at startup and then your application is going to last a long time and you're going to make requests over the course of days or weeks.

So, people just got used to writing Java apps in that way. However, those assumptions that it was based on, don't make sense in a world of Lambda. We're not building a huge application. We want to write an individual Lambda function. We're not depending on like a whole bunch of different environmental dependencies. We're normally just depending on two or three and those assumptions don't apply, but even if we take those assumptions out, there is still a cost to running something like Spring and those costs come normally at cold start time, and that there is a lot of stuff that we have to bring in, a lot of libraries and whatever that have to be loaded and instantiated and all that kind of stuff.

And also Spring does a bunch of stuff at startup using reflection to dynamically load code, which makes sense when you're writing one of these larger applications that's going do a whole lot of different stuff and is going to last a long time, but it just doesn't make sense. So this framework that you speak about, I'm sorry, Stefano if he hears this, because I know that he's put a lot of work into it. It's one of those great occasions where AWS, they use this phrase like meeting people where they are.

Jeremy: Right.

Mike: And so it's one of these situations where people, they know that there's a lot of Java developers out there that have got all these Spring apps, and they would be like, "Okay, well, we can meet you in your writing your Spring apps and still allow you to write Lambda functions." And I admire AWS for doing that, I think that's great. But there is a real problem there, that Java developers who are using Spring and use this thing will get to a point where they go, "Oh, this is how you build Lambda apps." It's not how Lambda is designed, right?

Jeremy: Right.

Mike: You are missing a big, big trick by sort of locking yourself into the Spring way of thinking when you're building Lambda applications and you will be far more effective as a Java Lambda developer if you got rid of all that Spring stuff and just thought about the underlying function that you're trying to write.

Jeremy: Right. And that's one of those things were like you said, I think AWS has actually done a really good job of giving people on-ramps to serverless, where I mean, "Look, just throw your Flask app, throw your Node.js app or your Express app." Throw it into a single Lambda function, you get this mono Lambda or there's Lambda lift, it works with Spring Boot and all that other stuff. It's not the most efficient way to do it, but it certainly gets you started.

But rather than just complaining that this is not the way to do it, you outlined a really interesting solution, and again, I will never do this because I'm not going to write anything in Java, but for people who are listening, explain this whole idea of multi-module, multifunction Lambdas.

Mike: Yeah. And a lot of this applies to any language as Jeremy just said, shockingly, it's a little easier in Java because Java's tooling is better to handle this, but this is not specific to Java, so what these Lambda lifts or mono Lambdas do is they say, "Hey, we're going to express all of the logic of one application, which might have 10 different types of requests come in, we're going to we're going to put that in one Lambda." And that's to us, is sort of missing the point of a lot of a lot of what Lambda does.

And so what I would rather do and what John and I would like to do is have if we have 10 different types of requests then consider having 10 different Lambda functions. Each way, each Lambda function only has the security that it needs, so we talk about the principle of least privilege a lot in the book when we talk about security. So each Lambda can only access the things it needs, partly that's about bad actors, but mostly, that's about reducing the blast radius so that we don't shoot our own filth. When people think about IAM, sure think about security, but also think about it as a safety blanket.

It's like IAM is about safety. It's also about security. So if we separate out all our functions into 10 different functions, we can have much, much smaller IAM scopes. Cold start is reduced because we're only loading up the code and the libraries that each Lambda function needs. We're not loading up that for everything and that does make a difference. Like the difference between a 45 megabyte distributable and a five megabyte distributable, like you'll notice it and even if it's not important to production, it's nice at development time to have that speed up.

Jeremy: Okay.

Mike: So, you can separate it out, great, but then everyone says, "Yeah, but if I had 10 different functions, then it's 10 different repos and 10 different deployment scripts, and how do I share code and blah, blah, blah, blah. So, the way that we think about this is, well, first of all, just because it's 10 different functions doesn't mean it's 10 different applications. You can have one application and one repo that has 10 functions in it, and serverless or SAM or whatever, will support you having multiple functions in your template. That's fine, so you don't have to have multiple applications and multiple repos.

But then there's the point about, "Okay, but what if I have some shared code among these things? Maybe five of these functions are going to go out to a database. What if I don't want to have to rewrite that database code in all my five functions?" Well, wouldn't it be nice if we could have like our 10 functions and then also some shared code in the same repo. And one of the things we do in the book is show how you do that is how you build like a mini library, that's all in the same application and then your Lambda functions can rely on external dependency libraries, but also can depend on these internal little bits of shared code.

And so we have a whole system that works and when you use it, it's just a matter of running Maven package. It just does the thing. But it uses this this Java tool called Maven under the covers, and Maven, trust me has its drawbacks and it's been around 15 years, and it's XML and oh, my goodness.

Jeremy: Ah, XML.

Mike: But its semantics around modeling dependencies are far advanced from any other main language on Lambda, as far as I'm concerned. If I want to write some quick code, I write it in Node or Python. If I actually want to model some dependencies, I'm going to run away from both of those screaming and use something like Maven.

Jeremy: Yeah. I mean, reading it in the book it was it was actually really, really interesting and my mind was like, how could this be applied to other languages that make it as easy as this does, maybe without XML as the configuration for it. But yeah, but definitely check that out. If you are building AWS Lambda functions using Java and you're writing things in this monolithic fashion, this solution, I think, it's brilliant. It really, really works well or it looks like it works well. I mean, obviously, you've experimented with it, but breaking these Lambdas up is definitely what we want to do.

Mike: I think another sort of metaphor around this stuff is that when you're bringing your Spring based applications into Lambda, it's a little bit like lift and shifting.

Jeremy: Right.

Mike: But you're not writing code in a Lambda native way and just like when we when we lift and shift from a data center onto the Cloud, we then need to go through a second activity, which is building Cloud native apps because Cloud native apps aren't typically lift and shift. And so when you create a Lambda lift, you're effectively lifting and shifting and then what you need to do is learn some new skills and using new techniques to build some Lambda native code.

Jeremy: Yeah, definitely. I totally agree. Alright. So, another thing you have in the book, I think is a great section is sort of your gotchas section and that's one of those things to where it's like, I think, serverless developers get all excited, they're like, "Oh, I can do all this great stuff." And then all of a sudden, you hit some limitation or something and it kind of kicks you a little bit. I mean, that's sort of true with all Cloud native development, but you call out a couple things. One of the things that's actually wasn't in the gotchas section, but I kind of classifying it as it, is this idea of using versions and aliases.

Mike: Yeah. So I mean, this comes from towards the back of the book. So there's a couple of chapters towards the end, which are really not Java specific at all.

Jeremy: Right. That's another important point. This applies to all serverless developers.

Mike: Yeah. So basically, chapters eight and nine are purely about architecture. And so yeah, if you never want to read any Java, then just skip ahead and read those chapters of the book. Yeah, so versions and aliases are, for those of you that have been using Lambda a while you'll know you've been there since very early on, not quite the beginning, but very early on and they are useful when you're using Lambdas, the AB testing, Canary testing, Canary release.

Jeremy: Right. Canary deployments and things like that, yeah, sure.

Mike: They're really useful for that, but when AWS created them first, I think what they were trying to do was it was a sort of another one of these sort of meeting people where they are things where when people deploy applications, they're deploying the test version of their app and the production version of their app, that sort of treating them as one thing. So, even back then with API gateway, you'd have multiple stages. You would have your production stage and your testing stage and development stage because people weren't sort of used to this idea of ephemeral or isolated stacks, as we now think about it.

But what I tend to use now is instead of just deploying a different version of a function, I'll deploy an entirely different set of resources. So if I want my production Lambda versus my test Lambda, that could have probably been in two different AWS accounts, let alone two different stacks. And so I think where people sort of got stuck a little bit with versions and aliases was like, "Oh, so I have to have my test and my production are the same Lambda functions, and it all got very confusing." And I'm like, "Yes, it's all very confusing. Don't do that."

Jeremy: Right.

Mike: If what you want to do is gradual release of production Lambda, sure, use the same Lambda function, but if you want your acceptance testing Lambda versus your production Lambda, just two different Lambda functions is to deploy them separately. And especially what that requires is having your infrastructure as code down, like you really need to have automated deployment, but if you're not doing that, you really shouldn't be doing serverless development anyway.

Jeremy: And the other thing about aliases or versions is that Provisioned Concurrency has to attach directly to a version.

Mike: Yes. Yeah, I forgot about that. That's another Provisioned Concurrency thing. And again, if you're using the canary release stuff that comes with SAM and Lambda, that requires an alias as well, but that makes sense and I'm down with that usage of it.

Jeremy: Yeah, and I agree. The traffic shifting stuff is actually really, really cool if you've never played around with it, and there's a bunch of great plugins that just handle it automatically for you, too, so it's definitely something to checkout. So another one of your gotchas in the gotchas section was at least once delivery, which is something that I think people who are new to distributed systems in general can get bitten by.

Mike: Yeah. And Amazon have started taking more stick for this and I'm kind of with the people giving Amazon stick for this. I think there should be a switch for this now. So yeah, so the idea is that Lambda, your Lambda functions respond to events and Amazon guarantee that should an event that is configured to be attached to your Lambda function occurs, when that event occurs, your Lambda function will be called. They guarantee it will be called. What they don't guarantee is how many times it will be called.

Almost every time there'll be a one-to-one correspondence between the event occurring and your Lambda function being called, but sometimes your Lambda function will be called twice. Okay, what's the big deal? Well, what if your Lambda function was charging your credit card.

Jeremy: Right.

Mike: Right? You wouldn't want to be charging people twice? Well, not if you wanted to keep the business anyway. So, that's obviously a little extreme, but there's a lot of places where you don't want to do something twice. That's what your initial thinking may be when you're designing code. And so when you're writing Lambda, you have to be cognizant about the fact. If your Lambda function is making any change to an external downstream environment, in that if it's not just returning something, if it's actually like updating DynamoDB or writing a file out to S3 or calling an external service. If it's doing any of those things, you have to be aware that your Lambda function may be called multiple times.

Jeremy: Right.

Mike: And there are ways that you can manage this, which we go into in the book.

Jeremy: Yeah, I mean, that's the thing is just this idea of building item potent operations.

Mike: Yeah.

Jeremy: And understanding that, and it's funny, I mean, the billing your credit card twice thing, I think Stripe has this figured out pretty well. They have an item potent ID that you can send in with every with every API call and you can use something like a message ID or something like that that would be unique to each individual event that it won't double charge that card or whatever. But I do think that's interesting, and you outlined some strategies, like DynamoDB, lock tables and some of those other things, so that's certainly interesting.

Now, the last thing I want to get to, though, because this is something that I'm passionate about. I do an entire talk on this when it comes to serverless and that's the impact of Lambda scaling on downstream systems.

Mike: Yeah, this is this is a big deal where you're building what we call hybrid serverless, non-serverless systems. So really, really easy example to describe is if you have a Lambda function and it's in front of a SQL database. One of the really awesome things about Lambda is that it will, by default, scale 1000 instances wide, so, 1000 concurrency. One of the terrible things about Lambda when it's connecting to a SQL database, is that it will automatically scale 1000 instances wide.

Jeremy: Exactly.

Mike: Right? And, and if you're not careful, that could take down your non-serverless infrastructure components.

Jeremy: Or your downstream APIs or what your code is.

Mike: Or whatever, yeah.

Jeremy: Or whatever else. Yeah.

Mike: Or logging systems that you haven't thought about and all that kind of stuff. And so yeah, it can have a real impact. It's one of those things again where that's what it is. You just have to be aware about that. Lambda is a different way of architecting systems. And this is the kind of thing I've talked about for years and even talked about some of this. But when you look at Lambda code, it looks like the same code that we've been writing for years.

Jeremy: Right.

Mike: When you look at Lambda architecture, it's drastically different to what we've been building for years. The difference between writing stuff in a container and writing stuff that was running on bare metal, from an architectural point of view, there's been this gradual evolution through Cloud native through VMs to containers, not that much different.

Jeremy: Right.

Mike: When you're architecting for Lambda, you have to think very, very, very differently. And I think people don't necessarily realize that because the code is easy and it's like, yes, but we've shifted the mental effort from the code to architecture. And now everyone needs to be an architect and I think that's a good thing, right?

Jeremy: Right. I do.

Mike: I think that making architecture not this thing that exists in this ivory tower and bringing it to all engineers is a wonderful thing because engineers that are building these systems are much better able to make optimization decisions than someone that's completely removed, right?

Jeremy: Right.

Mike: So that's a good thing, but the flip side is that people need to learn architecture now. We need to learn about the architectural trade-offs that come when you build distributed systems.

Jeremy: Yeah, I totally agree. I think that's a good exercise, though, for developers who are getting into serverless to start looking at that. And then the thing I loved about the book was that you outlined I think four or five different strategies of being able to mitigate those downstream issues, which are, are something you definitely have to think about because I don't know many people who are building entirely serverless applications where every piece of the infrastructure scales, just like Lambda does. So certainly something to think about. Alright, so anything else? Anything else about the book we should know?

Mike: It's amazing. It's awesome. Everybody should buy it. It was interesting. We spent a long time on it for various reasons, but one of the nice things that came through at the end because we delayed it was that Serverlessconf New York October last year, I think it was.

Jeremy: Yeah.

Mike: I bumped into Tim Wagner, who I know pretty well, by this point. Tim developed and ran the Lambda team for those of you don't know for a number of years. And so I was chatting to Tim and I was telling him about the book, I was like, "Hey, would you be willing to write the foreword to our book?" And he said, "Of course." And he meant it. So Tim wrote the foreword to our book, which in and of itself was great and I'm very appreciative that we have that introduction there with Tim, who if anyone knows anything about running Lambda in the enterprise, it's Tim, because he that's what he designed Lambda for, right?

Jeremy: Right.

Mike: But what also happened was, this is the story. So I'm sitting on the plane to re:Invent, and I'm doing the final read through of our tech draft. We've had all the written input from our tech reviewers. John and I applied all the changes. I'm reading the final thing that we're going to give over to O'Reilly, so they can start doing proofreading. I turned on my plane at the end of the flight, there's an email from Tim, which has his foreword to look at. What it also has that I wasn't expecting was 15 pages of Technology Review, well, technical review for the book. So the book would have come out six weeks to eight weeks earlier, but it didn't.

Jeremy: Sure.

Mike: Apart from the fact that Tim gave us this extraordinarily useful tech review feedback.

Jeremy: Right.

Mike: And we incorporated a lot of those changes and things like for example, Provisioned Concurrency, which was announced around that time. That came into the book because Tim's like, "Yeah, you really should include Provisioned Concurrency in here," and a number of other stuff as well. So yeah, Tim was very much involved in it and we're very appreciative to him for that.

Jeremy: Yeah. Well, if you're going to have somebody write your foreword or review your book, I would say Tim Wagner, about Lambda anyways, is a very, very good source. And I actually read it. I read the foreword. The foreword itself is just worth reading because it actually is really interesting and gives some good insight. So seriously, I mean, the book itself, like I said, if you're a Java developer and you want to write Lambda functions, go pick it up. If you're not a Java developer and you just want to learn more about serverless, and get some insight from two very, very smart people in the industry, pick up the book, take a look at it. It's on oreilly.com. If you've got a subscription, you can read it there.

Excellent book, very well written lots of awesome stuff in there. So thank you and thank John for me as well, for writing it as I think anytime there's good, accurate serverless content out there, it really is a gift to people who are trying to adopt this crazy new thing. So if people want to get a hold of you or find out more about what you're working on, how do they do that?

Mike: Yeah, we have a couple of ways, so our website and the thing that's updated most, my website is our blog, so if you go to blog.symphonia.io, there's a bunch of stuff on there and just symphonia.io shows what we do, has a link to the book. If you want to take a quick look at the book and not sure what you want to do, O'Reilly, who it's published with, have this nice thing where you can get a one-week, I think, free subscription to their online platform. So you can go on there and look at our book for a week and then if you like it, you can either subscribe fully or buy a copy of the book.

I'm on Twitter at @mikebroberts, which I guess will be in the links for this podcast, which will be a combination of tech stuff and New York theater and my cat and not as much theater at the moment because, obvious reasons, but yep. And then we're also on Twitter at @symphoniacloud.

Jeremy: Awesome. Alright. Well, I will get all of that into the show notes. Thanks again, Mike.

Mike: Thank you, Jeremy. Thanks, everybody.

THIS EPISODE IS SPONSORED BY: DatadogandAmazon Web Services(Serverless-First Function May 21 & 28, 2020)

View Details

About Gareth McCumskey:

Gareth McCumskey is a web developer with over 15 years of experience working in different environments and with many different technologies including internal tools development, consumer focused web applications, high volume RESTful API's and integration platforms to communicate with many 10's of differing API's from SOAP web services to email-as-an-api pseudo-web services. Gareth is currently a Solutions Architect at Serverless Inc, where he helps serverless customers planning on building solutions using the Serverless framework as well as Developer advocacy for new developers discovering serverless application development.

  • Twitter: @garethmcc
  • LinkedIn: linkedin.com/in/garethmcc
  • Portfolio: gareth.mccumskey.com
  • Blog Posts: serverless.com/author/garethmccumskey/

Watch this episode on YouTube: https://youtu.be/5NXi-6SmZsU

Transcript:

Jeremy: One of the things that I know I've seen quite a bit of is people using just the power of Lambda compute, to do things right? And what's really cool about Lambda is Lambda has a single concurrency model, meaning that every time a Lambda function spins up, it will only handle a request from one user. If that request ends, it reuses warm containers and things like that. But, if you have a thousand concurrent users, it spins up a thousand concurrent containers. But you can use that not just to process requests from let's say frontend WebSocket or something like that. You can use that to actually run just parallel processing or parallel compute.

Gareth: Yeah. This is one of what do they call it, the Lambda supercomputer.

Jeremy: Right.

Gareth: You can get an enormous amount of parallel... Try to say that three times quickly. Parallelization with Lambda. I mean, like I said, by default you get a 1000 Lambda functions that you can spin up simultaneously. And if you ask nicely... Well, you don't even have to ask nicely, just ask them and AWS will increase that to 10,000 simultaneous. And it's really impressive how much compute you can do, to the point where, at one point I was working with a company looking to try to do some load testing of an application.

They had an instance where, on Black Friday, the tech kept falling over. They wanted to try to get some load testing in beforehand to make sure that it can handle at least a certain amount of volume. Because you can never entirely predict what your traffic patterns will look like. But at least let's try something. And they spend a lot of time looking at commercial solutions out there because there are a few of them out there that try to help with that.

And they normally try to do about 500 to maybe a 1000 simultaneous users or simulated users, which is impressive but not quite good enough when you're an organization that's going to be having 10,000 to 20,000 simultaneous users on your site at a time. That gets a bit rough. So the move was then to try and build some load testing application ourselves. And this was initially tricky to do because we were trying to do this using the traditional VMs, virtual machines, and containers in some way, try to get EC2 instances up and running to try and run multiple simultaneous users at a time, in a single VM, using essentially a combination of these end to end testing tools where you can simulate a user flow from loading the homepage to going to a product page, adding to cart, going to checkout. Doing all of this on a staging environment so that you could simulate the whole user for all the way to purchase, the sort of main line to purchase as it were, make sure that you could get a few thousand users all the way to there without issue.

And what ended up happening was these virtual machines just couldn't cope with the load of all these simultaneous users running on a single machine even with inordinate amounts of CPU and RAM on them. So the idea came to us to try and do this with Lambda instead. So what ends up happening is, because you have a thousand simultaneous Lambda functions, AWS also architects this in a way that the noisy neighbor effect of all of these Lambda functions is almost nothing. You can't say nothing.

There has been some research I've read that shows there is a bit of a noisy neighbor effect between Lambda functions. But one interesting thing that we found was this is reduced when you increase the size of your Lambda functions to the maximum memory size, which is pretty cool. Because then uses an entire machine essentially or virtual machine as it were. So now you're limiting the effect of that noisy neighbor effect happening. Which means you can then also run 10 to 20 simultaneous users on that single member function with that enormous amount of size.

And if you have a thousand of those, well now you've got a thousand Lambda functions with 10 to 20 users per Lambda function, running an end to end test, pointed at a single staging environment. That's a pretty powerful bit of load testing you can perform there. And Lambda being as flexible as it is, we needed to import a binary to execute the end to end testing framework that we were using.

So you can use Lambda asynchronously to help you spin up the required binaries, import all of these items in and then synchronize the start of the tests through SNS, for example, which can just fan out the go command to all of these Lambda functions waiting to execute. And that was it. We have 15 to 20,000 users, load testing and application. And that's going to tell you whether you're ready for Black Friday or not.

Jeremy: Right. Yeah, no, I think it's an awesome use case. And I mean the parallel load testing, I mean, just the amount that you can get. I mean even you try to run something on your local machine and you're trying to do just to simulate some things. You only have so many threads that you can open up, so many users you can simulate. And to do this reliably, I mean, some of these testing sites, you can go and get some of these... Use some different sites to do it.

But they get pretty expensive if you want to do regular tests. If you run a thousand concurrent Lambda functions, even at the maximum memory, and it takes maybe five minutes to run your full load test, you're talking about a couple of dollars every time that runs. Right? I mean, so the expense there is amazingly low. I think that's a super useful use case. There are some more specialty things though that you can do with Lambda.

And Lambda is very, very good at, or I should say Lambda is built to receive triggers, right? Serverless is an event driven system. So there are all kinds of triggers for Lambda functions. And I'm sure we probably could talk about a thousand different ways that you could trigger a Lambda function and do something. Everything from connecting to an SQS Queue to like you said, the DynamoDB streams and things like that.

But one of the interesting triggers for Lambda functions, is actually getting an email that is received from SES, which is the Simple Email Service that AWS has. and I find this actually to be really interesting. You did something interesting with that.

Gareth: Yeah. We worked with an organization who essentially they handled requests for medical insurance. So other companies would send this organization an email with information about users who needed medical insurance. And these emails are usually pretty similarly formatted. They were structured almost exactly the same, just with user information that was slightly different every time. And it was getting very tedious for them to have to constantly go into this inbox, troll through all these emails and then manually insert them into a CRM system, so that the sales team could later get back to these folks and help them sort out their medical insurance.

So one of the things that they initially did before we came along, was they had a virtual machine essentially running that was a regular old email inbox and a script that ran every five minutes on cron that would then log into this inbox pull all these emails out, and try to parse them and then insert them into their CRM binaries API. Anybody who's done that kind of thing would realize there's quite a few flaws in that potential process, because not only do you potentially have thousands of emails you've got to pull in in five minutes before the next cron runs, but how do you keep track of which email you've already read.

And the issue there is as well, this inbox was used by humans as well. So you couldn't use the red flag on the inbox as well, because a human might've clicked on an email and then you completely miss this lead in the first place. So it was kind of a problem to solve. So ultimately, the solution ended up being, registering a new sub domain on their email domain. And then just informing the partners that the email address to send these leads to had changed. And this email address was actually created inside of SES, Simple Email Service, which has a way for you to create a way to receive emails.

You can create inboxes in SES to receive mail. And then you have a process you can set up and how to manage these emails. So, various methods you can do with these emails, the one that we ended up choosing was taking the email and storing it as an item in an S3 bucket. And this is where these triggers then happen. Anybody who's looked at serverless has seen the Hello World equivalent of serverless, where you can use S3 buckets to create thumbnails of images. But you can trigger anything in S3.

So if you drop an email into an S3 bucket, that can trigger a Lambda function. So what's useful here is that we have a system that's receiving an email, puts those into an S3 bucket and that specific object put, spins up a Lambda function with all the detail of what this item is. The Lambda function that's triggered can read that email straight out of the S3 bucket and then process it just like it was doing before. It can pass through this email, get the user's contact information and then put that into that CRM that they need in order to get in touch with folks.

And again, there's no worry here about have we read this email before. This isn't a human readable inbox. This is only used through SES. So there's none of that concern. And again, this is all entirely serverless. SES is going to receive your email at what pretty much whatever quantity you need, and they were receiving a few thousand emails a minute. So it became quite a big deal. And S3 as well has enormous scale that you can just use. You can just insert all of these items. Lambda can just scale out and process all of these items individually, pretty handily.

What actually ended up being the problem was that their downstream CRM couldn't handle the load at one point, so they had to have an upgrade. But that's a different story.

Jeremy: Well that's a common problem I think with serverless is that it handles so much scale that the downstream systems have a problem. So that use case though, this idea of receiving emails, dumping them in S3, reading them in with Lambda, there are just so many possible use cases around that. So the medical thing I think is interesting, parsing it, trying to get it into a CRM.

But if you wanted to build your own ticketing system, like a support ticketing system. Now again, I wouldn't suggest you do that unless you're like building a SaaS company that's going to do it. But if you're building a SaaS company that has a ticketing system component, this use case is perfect for it. I mean, it's great. And then I actually saw quite a while ago, somebody built an S3 email system, like an entire email server using just S3 Lambda and SES.

So essentially when the message came in, it gets processed by Lambda function, the Lambda function would read it. It was just sort of a catch-all address, Lambda function would read who the "to" was from and put it in the right box for it. It's amazing. So I think that's a really, really cool use case. I think you could handle attachments and all kinds of things like that that you could do with that, run algorithms on them, send them into SageMaker and do machine learning. I mean there's all kinds of things that you could do that would be really, really cool around that.

Gareth: There's also the idea... I mean, the ideas that have crossed my mind and you think about all these triggers that you could potentially do, and all the services available in AWS to do them with. And I even picture an idea where you could talk about a ticketing system and combine this with a CRM type system where you're receiving a constant communication with customers. And you can send this through, I forgot the name of the AWS service that does sentiment analysis on text.

Jeremy: Yes. Yup.

Gareth: So you can just use that to do sentiment analysis. And you can have managers in a customer services team get notified when there's a certain proportion of customers that have a sudden mess of a negative sentiment. And you can start investigating before, even though there's a problem you've picked up from customers that there is a problem to go and solve. You could even do this with voice, because of AWS's sort of call center service. You can pass that through a sentiment analysis machine.

And again, all of this stuff is built in serverlessly. You can trigger all of these things just as they happen automatically event based, really.

Jeremy: I hope this has got people thinking, because we clearly are not going to have enough time to cover all of these use cases. But there's a few more though that I'd like to talk about, because I think these are broad that you could use for... And you can think of your specific use case for them. And one of those is cron jobs, right? Like cron jobs are...They're the Swiss army knife for developers. Like we use them for everything.

We use them when we're like, "Hey, these log files keep filling up on this server. Let's run a cronjob and clean them up every couple of days or whatever it is." We use them to trigger ETL tasks. We use them to trigger all kinds of different things. And that is a really, really good use case for serverless, especially if you want to run something sort of peripheral to your main application.

Gareth: Yeah. And cron jobs funny enough is probably the second most common use case I think we've seen with serverless applications, because I think every developer has been in that situation where you've got your main stack of stuff, sitting there doing your web stuff that you need. And you suddenly realize you don't really have any way to run schedule tasks. You spin up a little T2 small, EC2 instance somewhere to run some basic cron jobs, that might just call it an API end point at some point.

But you need that capability to schedule things on a permanent, hourly, daily whatever basis to run those things. And that's where services like Lambda for example, become incredibly useful, because you can just schedule a cron job or a schedule onto a Lambda function, and then have it access all of the AWS services that you'd normally access. And a lot of the times the cron job's fed everything from sending regular email, because often you'd have a management that wants a status update sent for certain metrics. So you build a cron job for that. A lot of the times before you realize the wonder of SQS as a queuing system, you might build your own little queue system in a database table and you use a cronjob every few minutes to run over there. That kind of thing happens. So again, Lambda functions become really useful for managing all of these sort of scheduled cron jobs that you need to execute.

Jeremy: And combining cron jobs with other things. Like let's say every hour or something like that, you want to trigger something that anybody who's on your website, something gets pushed to them or whatever. If you've got WebSockets set up right, like can just run a cron job and you do that every hour. The ETL tasks I think are an excellent use case for the cron job things. And then you actually did something with some XML feeds, right?

Gareth: Google Shopping feed is one of these things that Google provides for you to advertise your products. Again, this is part of the e-commerce platform that I was working with. Google Shopping has the ability to read an XML feed of your products, but this feed needs to be built. And one easy way to do that is because the details of your products don't change all that often. I mean a shirt is a shirt, and a pair of shoes is a pair of shoes. You can build this feed ahead of time.

So cron jobs is a great way of pre-rendering this XML feed so that when the Google Shopping spider comes along to read the XML feed, it's always available for you. And in this particular case the organization was using Magenta as their e-commerce backend. So instead of building the features on top of the existing stack, we were able to build a serverless sort of side project to handle this, so that we didn't have to make these changes to the existing stack and potentially cause issues there.

And Google could just come at any time and constantly read this shopping XML feed or because of a XML data built ahead of time with a cronjob.

Jeremy: Yeah, and I love that too, where you use Lambda to do the compute, and it doesn't touch the rest of your stack. I mean like if you're generating reports or something like that every night, do you really want to be using CPU power that runs alongside your application? That's handling requests from your users to generate what could be very CPU intensive. And that actually leads me to this next use case, which is this idea of sort of this offline or async processing. And you had mentioned, asynchronous in the past.

Like that's the idea of, API Gateway sends a request to SQS, SQS which is a Simple Queue Service grabs the message, replies back to the API, "Hey, I got it." Which replies back to the user and says, "Okay, we've captured it." But then you've got something else down the road, that might require more processing power than you would want to do synchronously, right? Like, you don't want to generate a big PDF or convert a bunch of thumbnail images.

You don't want to do that in real time, while the user is waiting on the other end of an API. You want that to happen in the background. Right? So as a background task. So this idea of offline or async processing, like what are some of the other sort of things you can do with that?

Gareth: Well, you've mentioned a few of the use cases already, and one that I've ended up working on was a project where users were able to upload images into the application. And one of the things that had to be done to these images was essentially a uniqueness hash calculation on them. And this essentially scans through the pixels of the image, and then calculates sort of a string-based hash that you can very quickly determine if you have another image in library that has a certain similarity level.

So you can also tweak how similar you want all these images to be. But this is a pretty intensive process and can take 10 to 20 seconds in some cases depending on the size of the image. And you don't want this kind of thing happening synchronously on upload. So use it as an upload the image and sit around waiting for 10 to 20 seconds, until this hash is calculated. So what's here is that you have for example, an S3 bucket and we kept talking about S3, but it's the workhorse of AWS. It does so many things so well.

But again, you can trigger asynchronous offline style processing by dropping this image, for example, into this S3 bucket. And then either through a cron job as you mentioned before, or just triggering off of that put objects action that gets generated by S3 bucket to a Lambda function. You can then trigger these calculations. And this can run the gamut. It's not just this hash calculation I'm talking about, but you mentioned PDF generation. So you can have dynamically generated PDFs made available to the public when things like a DynamoDB table is edited.

Now in the background, it receives that event trigger to a DynamoDB stream that the data has changed and it starts rebuilding PDFs. Maybe there's multiple of them. And you can combine this with the power of something like SNS or EventBridge that you can trigger multiple Lambda functions, each rebuilding a specific PDF because of one data change that you made. Very powerful ways of doing these things.

One of the other useful ones that I've used in the past is we were talking about the whole Jamstack style process before. But a lot of cases, if you look at a lot of web frontends, there's many pages that often never change in their content, or very rarely change in their content. And in those rare circumstances where somebody has come to edit the contents in a CMS style. You can, instead of having a WordPress style CMS that you click a save button and the content is instantly changed.

You can instead save that in some kind of headless CMS system, for example, that can then trigger off a Lambda function to pull in this new content and rebuild the static HTML, JavaScript and CSS that page consists of and push that into an S3 buckets. And this is a CMS, a type of system that we built in the past with asynchronous processing, because you don't really need that page to be updated, the instance somebody hits a save button.

But you do need it to be updated within a reasonable amount of few... Maybe a few seconds is more than enough. And then you have the entire power of a Jamstack that can manage this enormous amount of load, loading static content, but still have that asynchronous process and to make pages dynamic as well. Pretty useful.

Jeremy: Yeah. Love that. Love that. So the other one that I really like, you mentioned that image hash. In order to figure out the differences between images. Actually did one of those again several years ago. But I know really has nothing to do with serverless, but it is a really, really cool little algorithm that you can write that essentially you reduce the quality of the image to 10 x 10 or something like that, take an average of the pixels and then you can use each pixel to determine... Make it black and white and so forth or gray scale. That's a very, very cool... It's a very cool algorithm that you can write. So definitely check that out if you're interested in running image hashes.

Gareth: The other interesting thing is if you combine these kinds of scenarios with something like CloudFront and Lambda@Edge this isn't necessarily asynchronous processing, but this is a way to... AWS actually has an example architecture where they combined Lambda and CloudFront to do thermal generation on demand. Which is a very interesting pattern to also take a look at where you have a base image. It might be your monstrous 4K image that is 20 megabytes in size. You don't want to serve that in a frontend, but you want this to be your source for all of the other images.

And you could use you can use a CloudFront and Lambda@Edge to receive a request for this image. And with Lambda@Edge you can intercept that request for the image. And often this is done with a unique URL. So you can have a URL that says something like thumbnail/ and a specific size, written somewhere in that path. Extract that information out of the path. Realize that the image that this URL references is this enormous one sitting in your S3 bucket.

And pull that out of the S3 bucket, resize that to the correct size you want, so that it's much smaller and return that to the user, and immediately CloudFront is going to take this much smaller image and cache that. So the moment that the next request... Because the first request might take a second or two for that whole process to happen. But the instance you do that the first time it's now cached in CloudFront and the next request that comes in for that size dimension, it's already done but you haven't consumed any extra space in S3.

It's all just an item sitting in CloudFront. So you could even clear your CloudFront cache and reset all those images. But again, you haven't incurred the cost of additional items sitting in your S3 bucket that you need to worry about managing last cycle. This is all just managed in CloudFront for you.

Jeremy: Yeah, that's great. We didn't really talk at all about edge use cases. But obviously there's the ability to do things like, I shouldn't say obviously, because this might not be obvious to people, but you can run a Lambda function at the edge. So you can do AB testing, you can do redirects, you can do blocking, you can do blacklisting, you can do... There's a million use cases around just the edge itself.

But anyways, so we're running out of time here and I do want to get to one last one, which is the, I guess the proverbial thorn in serverless's side, if that's the right way to say it. And that is machine learning because anytime you say, "Well, serverless can do pretty much everything." Everybody's like, "No, no, I can't do machine learning." Which is true to some extent. So there are some use cases where machine learning does work. And then there are some that they don't, but I don't know, maybe you have some more insight into this.

Gareth: Well, there's a few angles you can actually take on that because it all depends. I think one of the biggest Achilles heels of Lambda when it comes to machine learning is really Lambda's disk space that's available to you. Because a lot of the times with Lambda you need to import additional libraries in order to run machine learning models. You also need to import your models, which can often be enormous amounts of megabytes inside, yeah.

And that means you've got some limited space to work with there. And if you can actually fit those libraries in those models onto a Lambda function that turn out 250 megs of limited space well then you could probably run that in parallel, like I mentioned the Lambda super-computing, you can run those in parallel and potentially get a lot of work out of Lambda functions. But serverless as an architectural concept isn't necessarily just about Lambda functions either.

I mean the whole point of serverless is to look at the managed services that are available to you so you don't have to rebuild everything from scratch, remove that undifferentiated heavy lifting. So again, there's a couple of angles on this because if you want to build an image recognition model yourself, well, maybe reconsider that. Because there are image recognition models out there that you can use.

If you're doing text to speech, well, there are text to speech engines already available in AWS, and they might be good enough to do what you need to do. Of course, if you're trying to build your own product that is a text to speech product, well, okay, I get it, then you might want to build it yourself. But if you find that the model you're doing isn't quite provided by these services. There's one additional service that you can use in a serverless context.

And that's Fargate, which is pretty cool service to look at. It's different to Lambda in that it isn't quite as responsive. So if you're looking for something that's low latency and can really get things done really quickly, Fargate might not be the tool for you, but if you're doing ML models, that's probably not your concern. And Fargate's for anybody who isn't aware... Fargate is a service that lets you run Docker containers without worrying about any of the underlying orchestration and management of them.

You essentially say, I need a Docker container. This is the image, this is the parameters of what I need to execute and Fargate will spin up that infrastructure for you in the background. I don't know how AWS does it. But again, that's the beauty of it. I don't need to. They manage all of that for you and you allocate the disk spaces as well. When you're building these images, you set up the disk space you need, you import the libraries you need, the models you need, and it'll just execute in AWS's backend. So that's a great way to run your own models. And another angle is SageMaker. So there's many ways to take this where AWS provides a service that lets you run them on models. SageMaker is a way...

Jeremy: They have the models already built for you in most cases too.

Gareth: Yeah. So you can just import your models and run them. And there's an entire set of infrastructure back there to let you run your machine learning models anywhere that you like.

Jeremy: Yeah. I totally agree too. I mean, there's just so many options for doing that. And like you said, there's a ton of these media services that they have. They have Lex, they have you know, Rekognition with a K that allows you to do image recognition and some of those things. The sentiment analysis, I think it's Comprehend, right? And we've talked about that a little bit earlier. That's machine learning.

That stuff is just there for you and it's just an API call away right from your Lambda function or whatever you're doing. And so unless you have some really unique machine learning model that you need to build. There are still options to do that in a fairly serverless way or close to serverless way, just maybe not on Lambda functions. But anyway, so listen, this has been a great conversation.

I think hopefully people have learned a lot from it, but before I let you go, I do want to ask you. I mean, you do work with some customers at Serverless Inc. You sort of help them figure out how they want to move to serverless and what they're building. So what are you seeing as sort of those first steps that companies are taking as they're starting to migrate or think about serverless?

Gareth: Yeah, it's interesting. One of the downsides I think of serverless is that, it is such a new way to build things that it initially seems a little daunting. So a lot of organizations, especially the older ones have come from the idea of having your own servers on premise, and now this new cloud thing has happened. So we need to move to the cloud and they can essentially take what they have on premise just lift it and shift it into the cloud. And things are good. Things are familiar, there are some slight tweaks here and there, but pretty much it's what we know.

But just running in their person's data center and the lift and shift seems to work. Serverless mostly that isn't a lift and shift type operation. But there is some limited ways to do some lift and shifting. So an example of this is, if you're already running an Express backend for example, you can pretty much take your existing Express backend and fit it into Lambda. And we have a lot of customers who do do that. And if anybody who has used all of the available services in serverless and built their Lambda functions from scratch, this might seem like an odd way to do things, but it's actually a really nice way to quickly get into serverless.

And see some of those benefits that you get with the automatic load balancing and the disaster recovery and the less maintenance and so on, lets you quickly get into that. And we see this across the board. There's even a project now, it's a project out there called Bref for example. If you're building PHP applications where you can just run your Laravel or Symfony application on Lambda functions for example.

So we see that a lot as the initial use case, where folks want to take what they have existing and lift and shift it into serverless. And then ideally what we find is that they understand that there's limitations to this, because you're just taking what you already know and putting into something brand spanking new, and then you start realizing there's a lot of benefits to serverless that you're not really getting by doing that.

So things like making full use of all of these services that are available to you in AWS because you can't necessarily get your Express backend to get triggered by an S3 bucket, for example. And Lambda function does that and it doesn't really speak to Express really cleanly. That's when we start helping organizations with those POCs where they're building a sort of cloud-native, serverless-first style application or just one small element of their application as a serverless-first item. And this runs the gamut.

We have a conference coming up that's going to have thousands of attendees. We want to build a mobile application that's only going to exist for that weekend. Let's build it serverless and see how that works out. Or we have a review system on our site that customers, sort of reviews to this third party, but let's rebuild our integration into our frontend using serverless for example. There's so many different use cases on these POCs that folks are trying to do.

And ultimately you find that we then end up in a situation where they realize that serverless is incredibly powerful. Their Express or their Django, or their Flask app is running really well. And their POC for the serverless-first application is running incredibly powerful and reliably. Now they start looking at re-architecting their entire stack a lot of the time using serverless as the primary way to do this.

And again, it's very difficult to put a use case there because again, this runs the complete gamut from everything we've spoken about tonight, whether they're integrating WebSockets with parallel compute and Jamstack style applications and so on. And it's a really exciting field to be in because all of this growth in serverless that we're seeing and all these different use cases with organizations out there.

Jeremy: Yeah, no, totally agree. I mean, that's just awesome. I mean, and that's what I love about sort of how serverless works, is that there are a lot of really easy on-ramps, right? I mean, the DevOps piece of it running some of those cron jobs and doing something that's peripheral to your application, or building out a separate microservice that does the reviews or does some sort of integration, or does your PDF generation or that kind of stuff. But it's not touching the main system.

But then starting to build more complete tools, taking advantage of sort of that Lambda lift or that fat Lambda that does a lot of processing for now. But then start breaking it up, use the strangler pattern, start sending things to different services. Yeah that's just awesome. Thank you so much for doing this episode because this is one of those questions where people are like, "Well, what can you do with serverless?" And really it's what can't you do with serverless? And right now we're getting to a point where it's just really there are very few limitations here. Yeah, I just think it's amazing.

Gareth: Yeah. I've had folks ask me, so what can't you do with serverless? And that actually, it's one of the most difficult questions for me to answer. In the past, it used to be that you can't run really long compute and then AWS increased the timeouts to 15 minutes. So then that removes a lot of use cases that you couldn't do anymore. And then they introduced Fargate, so if you'd really had something that you needed to run in the background for a very long period of time, now you can do that with serverless.

So with just the whole industry and the whole architecture is advancing so rapidly and so many new services are coming out. AWS keeps listening... The other vendors too. We've been talking a lot about AWS. The field is growing enormously with other vendors too. Like Azure for example. And Google even making a lot of inroads in their serverless infrastructure that they're building out

Jeremy: Tencent is doing a lot.

Gareth: Tencent is actually busy deploying a lot of cloud services right now. And in fact Serverless Framework, we have support for Tencent because they approached us and said, "Listen, we want to make sure we can do serverless, stuff because this is the way of the future." And that's what they're focused on now.

Jeremy: Yeah. No, and I mean, that's I guess my last point would be just for people that are moving into serverless, trust the services in the cloud. Right? The cloud can do things better than you. So just moving your Express app over into a single Lambda function. Everything like retries and failure modes and some of those other things. There's so much stuff built into the cloud.

So it's not you having to do all of it yourself. There's just a lot of support there. So anyways, Gareth thank you so much for being here. If listeners want to get a hold of you and read some of the great blog posts you have. How do they do that?

Gareth: Well, most of my blog posts are written on the serverless.com websites. And that's easy to find serverless.com/blog. We pretty much update regularly about the serverless framework, new features we're bringing out. And then all the work we're doing at Serverless. For me personally, if anybody wants to get in touch personally, they can get ahold of me on Twitter, it's @garethmcc on Twitter. Nice and easy to find. And yeah, that's really the best way to get in touch with me and see what I've been writing.

Jeremy: Awesome. All right. We will get all of that into the show notes. Thanks again.

Gareth: Awesome. Thanks so much, Jeremy.
THIS EPISODE IS SPONSORED BY: Datadog

View Details

About Gareth McCumskey:

Gareth McCumskey is a web developer with over 15 years of experience working in different environments and with many different technologies including internal tools development, consumer focused web applications, high volume RESTful API's and integration platforms to communicate with many 10's of differing API's from SOAP web services to email-as-an-api pseudo-web services. Gareth is currently a Solutions Architect at Serverless Inc, where he helps serverless customers planning on building solutions using the Serverless framework as well as Developer advocacy for new developers discovering serverless application development.

  • Twitter: @garethmcc
  • LinkedIn: linkedin.com/in/garethmcc
  • Portfolio: gareth.mccumskey.com
  • Blog Posts: serverless.com/author/garethmccumskey/

Watch this episode on YouTube: https://youtu.be/Q3tbdlHH0Mg

Transcript:

Jeremy: Hi everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with Gareth McCumskey. Hey Gareth, thanks for joining me.

Gareth: Thanks so much for having me Jeremy.

Jeremy: So you are a solutions architect at Serverless Inc. So why don't you tell the listeners a little bit about your background and what you do as a Solutions Architect?

Gareth: Sure. So, going back a bit, I mean I've been a web developer for a few years now coming up to 15 years. It doesn't feel quite as long as that, and I actually started back in the days of building a PHP web frameworks and so on. And my first start with serverless was back in 2016 where I actually, was taking over the lead of a team at the time.

And part of my job there was to try and help modernize this aging WordPress monolith that had been the company's entire online presence at the time. And the company, they sold tours online. And online was the only way that they sold their product. So it was quite important to have this product working well. And then I was going through the usual steps, just taking a look at how we could potentially modernize things, looking at the Laravels and Symfonys of the time.

And I was chatting to one of the guys at Parallax Consulting who had helped this company set everything up on AWS. Get all the VMs up and running in the load balances and so on. And one of them suggested that I take a look at the serverless thing that one of their team had spotted. So I thought, well, let me give it a try. Let me give it a... We see what this thing is.

And that really ended up being my road down into serverless, because the moment I picked serverless up and started looking at potentially building a RESTful API out of serverless to help modernize the architecture for the company, that was me. I was down the road and started building a POC. And the POC we had was just to take one small portion of the existing stack and replace it with something completely based off of serverless.

Something that received reasonably high traffic that wasn't super critical for the running of the organization. So if it failed, it wasn't a train smash. But if it succeeded, it would give us a great indicator that this was something we could definitely move forward with in the future. And ultimately the POC was a raging success.

Everybody in the organization was incredibly impressed with how well this serverless tech that we built. And to be perfectly honest, it wasn't even the best architect serverless tech in the world, but it still performed incredibly well, which was quite impressive at the time. So yeah, we were really happy with that. That essentially solidified serverless for me and the way forward for me in the future.

Jeremy: Awesome. And so then you started working at Serverless Inc as a solutions architect. So what are you doing there now?

Gareth: So now I'm involved with the growth team and being a startup the roles are quite mixed, so I'm called a solutions architect. But I end up doing a lot of different things. One of my main roles is involved in support of our paid product. Serverless Framework Pro dashboard, I help users who are using our product and helping them deploy it and set things up. We have a number of users who need support and help and assistance in setting up their serverless architectures and designing those sort of architectures around their use cases.

That's a really interesting job where you get to see quite a variety of ways that organizations are using the Serverless Framework. And it also means I'm working on content all the time, so I'm writing blog posts, producing videos, talking to the community, doing talks, all the usual sort of developer relations side of things as well. Keeps me quite busy.

Jeremy: Awesome. All right. Well, so since you are so deep into this stuff now, and again, you're working with a bunch of different clients with Serverless Inc, you're writing these blog posts, you've been doing this for quite some time now. I mean, I think you started what? Around 2016 or so working on serverless. Is that about right?

Gareth: Yeah, I started in 2016 building serverless applications for the first time, and last year I joined Serverless Inc themselves. Yeah.

Jeremy: Right. So in terms of experience with serverless, you probably have the most amount of experience you can possibly get, right? Because this is such a new thing. So you've been doing it for a while, you've been seeing all these different things. And one of the things I think is really interesting for people to be able to see and especially people who are new to serverless is this idea of what are the use cases that you can solve with it. Right?

And it's funny, if you're familiar with James Beswick, he has this sort of joke that he used to do it and in one of his presentations where he thought serverless was just for converting images to thumbnails. Like that was sort of like a very popular use case way back, when this first started to become a thing. And obviously you see a lot of things like web APIs and some of this other stuff.

But I'd love to talk to you about that today. Because I think there are a broad range of serverless use cases. And I'm probably in the camp of you can basically do anything you want with serverless. There may be a few exceptions here and there. But maybe you can just give us a... What do you see as like the most popular serverless use case?

Gareth: All right. Now by far the most popular use case is using serverless to build APIs. Whether that'd be a RESTful API or even a GraphQL API. And that's hands down the most common use case at the moment. And I think that was primarily pushed by the fact that API Gateway is actually such a great technology to use for building APIs, specifically RESTful APIs because it just takes away so much of that headache of trying to manage where web servers, load balancing them, a whole bunch of features that it includes to help you build your APIs, including things like JSON Schema requests Syntax, API keys that you can use to throttle users on your APIs and a bunch of others.

I mean, it's an amazing technology and then you combine that with the power of something like Lambda in the backend that you can use to receive these requests, process them and glue all the other managed services that you may need like DynamoDBs and so on. And you have a very, very solid wrestle API backend that you can very easily use. And then when you combine that with something like the Jamstack, which is a relative... It's odd how this relatively new phenomenon that's coming out is essentially a regurgitation of an old phenomenon that we used to do in the old days.

Where the static files are stored on a server and it's just serving HTML, CSS and JavaScript. And we have an API backend that helps us manage all the dynamic data that gets populated in aesthetic files now. It's become known as the Jamstack essentially. And that becomes a very popular use case for building web applications on serverless.

Jeremy: Awesome. So let's dive into the HTTP or the API use case a little bit more because just so if people are listening to this, I'd love when we can kind of teach things on this podcast and we can get people to sort of just understand or make it click. Right? And so you talk about an API Gateway, we talk about Lambda as the backend. Maybe just explain exactly what you mean by, what is API Gateway and then how does Lambda really tie into that?

Gareth: So API Gateway essentially is, AWS is a solution to give you endpoints. So you need some way to expose an HTTP endpoint to any client. And in this case, when I'm talking about Jamstack, I'm talking about a web client and a browser, but it doesn't necessarily just have to be a web browser either. It can be a mobile application. And that's another very common use case we see, where web APIs are reused in a mobile application to provide data to a web app client. And API Gateway is the front facing feature that allows you to receive data from your users.

So if you think of a front end that's using React or Vue or any of these jobs for frameworks, it's going to send a request to the same point, be it a GET, POST, PUT, DELETE whatever it might be to help manage that data. And in that way, it's building, it's hydrating a UI off of this API Gateway backend that you built. And I think gateway is essentially a replacement for what you would normally traditionally knows as roots in your web framework.

If you've used any of these MVC style web frameworks, you'd have a roots configuration that you'd apply, that would then point to a controller, potentially with some actions in them that then handles those requests. And API Gateway essentially removes all of that work for you. You just need to configure a path pointed at a specific Lambda function and your code then receives an event object from API Gateway that contains all the details you need for this request.

Including everything from, your headers that have been received, a few added in by API Gateway to help you make some analysis on your request, potentially, to the body content that's being sent. If this is a post request, for example. So it's essentially it replicates the effect of having an HTTP request come through any old web server like an Apache or an Nginx. But without any of that concern about configuring this very complicated piece of technology on an EC2 instance that you could mis-configure which I've never done, I promise.

Jeremy: Or not secure properly.

Gareth: Yeah, exactly. So yeah, just combining those two really makes it incredibly powerful. Especially because when you look at Lambda. Lambda is an event driven way to run code. So Lambda by itself is kind of useful, but if you drop something like API Gateway in front of it that can trigger the Lambda function and pass data, you now have a match made in heaven essentially, something that could receive your data and then process it in real time.

Jeremy: Right? And so, I mean, that's the thing that's powerful about Lambda, right? And so Lambda is this one part of serverless Lambda functions-as-a-service if people are familiar with that term. And that essentially allows you to run code, whatever it is. And like you said, in response to events. But API Gateway is one of those really cool tools. Right now, there's a new version of it or there's the REST APIs, which is the existing version.

There's the HTTP APIs, which we talked about on the show a couple of weeks ago. But what's cool about the REST version, and this is something that will be coming to the other version as well, is this idea of service integrations, right? So Lambda functions are great and they can do all kinds of processing. But maybe explain, "Hey, maybe I don't want to use a Lambda function, maybe I don't need to use a Lambda function." What can you do with service integrations?

Gareth: Well, one of the really interesting things is if you have a client application that's making API requests, sometimes you want that response from the actual request to come back really quickly to the client side, because you've set things up in your infrastructure in a way that you know the request is going to get processed eventually. And this is the idea of asynchronous processing. So you make the request, the AWS or API Gateway essentially comes back and sends 200 success, don't worry about it. We've got this now. And in the back, what you've done is you set up this integration with another service, like an SNS or an SQS or something like that, that receives that data from API Gateway.

So instead of directly diving into a Lambda function, which becomes asynchronous request, this means your client then has to sit there and wait for the Lambda function to complete execution and return a success or a failure code. API gateway can immediately send us data into a service integration of some kind. And that data will eventually get processed, an optimistic style of coding your client, your frontend, and so on. It's a very, very, very useful and efficient way to handle your API requests, so you don't have to keep diving into Lambda.

And then you also, you don't end up having the additional cost associated with the Lambda function. Because then the functions do build unit, for every 100 milliseconds of execution time. And in the past, if you had just been taking the event data from an API Gateway request and dropping it into SNS or SQS service anyway, well let's just saves you having to do that in the first place.

Jeremy: Right? And a lot of those good use cases around something that you would use an HTTP endpoint for, that would be perfect for SQS or maybe Kinesis or something like that, would be like the webhook. Right? You know, you've got a lot of data coming in. You don't need to respond to the webhook with anything that says, I've done all the processing. You just need to say I've captured the data.

So that's where I think you get a lot of really useful benefit out of using something like that asynchronous pattern. Like you said, storing that data, throwing it into SQS, responding to the client immediately and then worrying about processing that data later on down the line.

Gareth: Well, this is one of those interesting situations where... It's one of those things that didn't click for me for a while, that you have this powerful asynchronous processing capability available to you in AWS. Because traditionally if you're building a web application, you receive a request, you process things synchronously, you put things in the database, you handle those things and you return a response. Maybe you'll drop something in a separate queue.

But generally things are done synchronously. Whereas with AWS, you can run things completely asynchronously if you wish. You can drop things into SNS queues, EventBridge, all sorts of different services. And that means that in the background, things are processing while your latency sensitive applications on the frontend, for example, are complete and it gives a great user experience. So it's a very interesting way to build an architecture.

Jeremy: Right? And another thing that you see that's becoming very popular and I think this probably started with the idea of building mobile applications and trying to minimize the amount of data that you're passing back and forth, is this idea of GraphQL. Right? And the benefit of GraphQL is the client can make those requests and they can just request the data that they need.

So that way they're not over fetching data. But also they can combine multiple bits of data together so that they're not under-fetching data either and need to make multiple calls. So API Gateway is great. You can actually build a GraphQL server with Lambda and API Gateway if you wanted to, "server" in quotes. But there's actually a service for that called AppSync. So can you just explain sort of what that does?

Gareth: Well, AppSync is a great tool to take away the headache of managing your own GraphQL server, which is no small feat. It can become quite a hairy situation to do that yourself. But AppSync essentially lets you tie all sorts of different resolvers to your GraphQL query. So you could link one portion of your GraphQL query to a Lambda function, which will then go off and do some interrogation. They'd be inspecting these three buckets build a data model and return that to your AppSync bottle.

You might talk straight to a DynamoDB table and pull some data in that way. You may even just grab items right out of an S3 bucket as part of your GraphQL query. And again, this gives you the power of an asynchronous feature where you're not necessarily incurring that latency involved in running code all the time. AppSync is always available. And just like a lot of the other services, it's pay-per request so you don't have a VM or a container sitting idle at 2:00 AM in the morning when there's no customers.

AppSync is really in waiting to receive a request and only bills when there's actual requests and only executes in your architecture with there's actual requests coming in which is pretty useful as well.

Jeremy: Right? Yeah. And the thing that's nice about AppSync is that it is just massively scalable. The throughput on it is insane. I mean, the same thing with API Gateway where you might be setting up load balancers in the past, right? Even if you're using elastic load balancers or application load balancers, you still have to worry about what it's hitting underneath. And if you connect these things right with serverless, that AppSync or I should say that GraphQL use case or that HTTP API use case is just massively scalable and it's very great.

So the other sort of, I guess common use case that we see is this idea of real time communication and AppSync has a way of doing sinking. It does offline sinking and some of that stuff, it's kind of built in. There's a whole new data store thing that they built, which is really cool. But I think a lot of people are more familiar with WebSockets. Right? So that's another thing we can do with serverless now.

Gareth: Yeah. And I've actually worked with an organization to help build out a eCommerce product using serverless WebSockets and the biggest advantage that WebSockets gives you is maintaining those connections to clients. And that really becomes one of the trickier things to do if you have to manage a WebSocket set of infrastructure yourself. But with WebSockets in API Gateway essentially part of API Gateway V2, AWS calls it. You have WebSockets available to you.

And the product that we ended up building was essentially a product counter on a frontend that as a user is sitting watching a screen, you just see a counter counting down as items are sold. And this is this was part of a larger Magenta backend again. So you have a data store storing the quantity of products in your warehouse and as items are sold, it's updating through a sequence of Kinesis dropping into DynamoDB table, which can then trigger your Lambda function again.

So this is all part of that asynchronous side of the things that I was talking about. Again you have a WebSocket connection that's set up by a client when they connect to a page. So they do a product page. The website and connection is created to your WebSocket backend that you create through API Gateway for example. And in the background you have Kinesis receiving data, but items that are sold that are triggering, that have been installed in DynamoDB.

That DynamoDB is using streams to trigger a Lambda function which can then go and look at the current existing active connections on their product page and send that data to the frontend. And the frontend then receives the data and can update the DOM with the actual value of products available. And WebSockets are great for that kind of use case because you're not dependent on a 100% perfect liability.

WebSockets do have a reputation of not being perfectly reliable. But if you need to give people and a rough estimate of the amount of products available in a product page, it's a perfect use case. It just gives that kind of nice solid feedback to the user that they're somewhere useful that they want to potentially maybe they want to buy because they can see the product running out or whatever it might be.

Jeremy: Right. Yeah. And obviously things like real time chats and anything where you want to be able to push data back and forth, multiplayer online games. I mean there's all kinds of crazy things you could do with WebSockets, but I think something that probably confuses a lot of people when they look at the WebSocket piece of things, from a serverless perspective anyways, is the fact that Lambda functions, if people are familiar with how they work, they are stateless, right?

So a WebSocket creates a connection, a long polling connection in a sense, and keeps that connection open and remembers who it is that's connected to it. But you're not running a Lambda function in the background that just sitting there and waiting. So can you explain how, because it is through API Gateway, but how API Gateway handles those long-lived connections while still using a femoral compute with Lambda?

Gareth: So WebSockets on API Gateway. So essentially API Gateway manages that long-lived connection somehow with the client. To be perfectly honest, I don't know the integral details of how AWS have configured this and personally I don't really care. It works. It does the job and that's the point of serverless as well. I don't want to have to worry about that undifferentiated heavy lifting of running a WebSocket service.

They do that really well. But how this works from an implementation point of view if you're building it yourself is that you essentially configure your WebSocket connection almost exactly the same way as you do a regular RESTful API with the serverless framework for example. And when that user connects to the WebSocket connection through API Gateway. You can set a specific Lambda handler to be triggered on a connection event with a WebSocket.

And this is useful because you can set your Lambda to receive these connection attempts. Which gives you a unique client ID and this is negotiated between the browser and API Gateway itself. So you don't have to worry about the details of it. You get essentially a UUID of that user's connection. And at that point you now have a unique reference if you do need to communicate back with them.

But as you said Lambdas are stateless, so we can't just use a session token or anything like that. So one easy use for that is DynamoDB, which if anyone's not familiar with it, it's a fantastic key value data store, that is incredibly useful in the serverless context and because of its low latency and high throughput, DynamoDB is fantastic for storing these UUIDs. Essentially because the only thing you're storing is a UUID and potentially the location that the person connected from, so you have some way to refer back to where they've connected.

And then on the other side of things, when you have data that you want to send. So again in my example, like I said, you have a user on a product page. As they load their product page in the backend, your JavaScript saying, all right, you're on the product page, create a WebSocket connection for this user to this API Gateway endpoint that you've already pre-configured. When you built your infrastructure, your architecture.

And at that point either it sets up the connection, triggers your Lambda function that gets triggered on connect with the UUID and the location of the page the person's viewing. And you can just log that into DynamoDB as is. When you have a product information updates come along, you can see that a specific product has a stock level change, you know the location in your frontend where this product is viewed, what the page URL is, what the path is.

So you can just query DynamoDB for all users that are on that specific page. And at that point you now have all the UUIDs you need to send an updated quantity. And at that point it's up to the client. Again, the client receives this data across the WebSocket connection on that end, and that's your JavaScript on the frontend that will update the DOM with the correct value. And it sounds really, really simple and is actually as simple as that.

There's none of the concerns about the load again, because all of the infrastructure we've used behind the scenes is completely load balanced because of how AWS manages this for us. We're not dependent on any server-full infrastructure that might need load balancing that might run out of capacity because DynamoDB is an absolute a monster when it comes to providing you capacity.

API Gateway itself, as we've said, has all the capacity you might need. And most of the work then is done by the client, which updates the DOM in real time when the data comes across the WebSocket connection.

Jeremy: Right? Yeah. And I actually, one of the startups I was at, well, several years ago now, we built a real time interface. You could comment on things, comments would appear in real time. We started using long polling, right? This is constant polling, which is just a terrible, terrible idea. You'd always setting up and tearing down connections on. So we ended up installing the software, I think it was called Ape, A-P-E.

I had to modify some of the backend because it wasn't a load balanced application. So we had to make it so that it would work from a load balance standpoint. And so we had multiple EC2 servers running. We had an elastic load balancer in front of it. And we had to bounce them back and forth between those. We had to use sticky sessions. It was a nightmare. And I don't know if anybody's ever tried to build a chat application at scale, but it is a lot. It is a lot.

And so this WebSocket use case from API Gateway just... I mean I really like the way they set it up. I think it's really, really interesting. I do wish there was like a broadcast channel that you could use to maybe like broadcast to everybody that was in group A or group B, without having to run some of those things. But it is very possible and like you said, it's actually not too difficult to set up.

Gareth: One of the interesting things is when I first started building back in 2016, we needed a WebSocket-style set up as well at that time. And back then, API Gateway didn't have the WebSocket capability built in. But there was, you'd jerry-rig it with a bit of, using the IoT service, in AWS and MQTT protocols and WebSocket connections and get it working. And that was one of those situations where we went down that rabbit hole, we bought all the stuff out, it looked really great. It performed really well. And as soon as we were done, WebSockets came out with API Gateway. So that's one of those lessons you learn in serverless that the moment you want to build something yourself AWS solves the problem for you.

Jeremy: Well, what you need to do is you just need to pretend that you're building it and tell everybody you're building it and then AWS will come up with it a few months later. And then you don't actually have to build it. But, no that's...

Gareth: Tweet it a couple of times. Block it a couple of times, let AWS know.

Jeremy: Exactly. All right. So those are really great frontend use cases, I think. I mean the API Gateway obviously or the API use case you can use for internal APIs and some of that stuff as well. I think that makes a lot of sense. But there are a lot of other use cases that go beyond just maybe interacting with a website, and then this ties into this, but this could be used for other things as well, and this is this idea of clickstream data.

Gareth: Yeah. So clickstream data is an interesting one because, a lot of the time, and we find ourselves, I was working with an organization who's finding that they wanted to get more information about what users of their product were doing. And this was more a case of they weren't personalizing. So again, this was an eCommerce platform and they wanted to provide some kind of personalization. Some personalized recommendations on the platform.

And it was funny, tricky to do this because they weren't super high volume in sales. So it's difficult to pinpoint what the personalization would be, because they didn't have thousands of products that somebody would have bought and then understand what people like. So they wanted some way to determine if someone's viewing this, more often than not, they click through on this, they click on this, they select this, and this ended up with a project view sort of capturing clickstream style data.

Clicking on a DOM element and sending that data to a backend to let process. And so there's a small element of frontend to this where you need to capture this data from your frontend, whether that's beyond your app or on a web frontend. And this is quite simply done with just some JavaScript that can record those click events and then eventually push those into a backend.

And again, depending on your volume, and in this case the volume is reasonably high. You need a tool, you need something like a Kinesis for example, which is probably one of the least serverless serverless products that you very often find because Kinesis still has this element of you need to allocate shards or you need to allocate some form of capacity to it. But it still handles an enormous quantity of data which is pretty impressive.

And it's really good at handling this time sensitive data that gets piped in constantly through a stream, exactly as the name suggests. And that's what we were finding. We needed some way to capture a lot of data very quickly and constantly all the time. So Kinesis is a great way to manage capturing this click stream data. But you also need somewhere to store this, once you're done.

So Kinesis has a great feature called Firehose. It's actually called Kinesis Firehose, where you can capture all this clickstream data and just point at an S3 bucket and say, "Put all the data there." Instead of trying to find ways to process the data, once it's in Kinesis. And this prevents you from having to spend a lot of time and effort on Lambda or any other compute platform processing vast quantities of clickstream data, but you still want to capture this and store it somewhere.

And then what this helps with is there are other services for example that you can use. I think services like Glue and Athena, two completely unrelated names that work together. AWS naming scheme hard at work. But Glue and Athena work really well together because Glue allows you to do things like introspect the format and the structure of your data, in a way that it can pass that to Athena as if it's a SQL table or SQL database. And lets you run SQL queries essentially on top of data just sitting in an S3 bucket which is incredibly useful.

And that was eventually what ended up happening, was there was a personalization engine, in the background that was receiving all of this clickstream data that was being funneled into the AWS backend, dumped into an S3 bucket and then regular Glue and Athena jobs that were running, I think it was on an hourly basis. Even at that, I was triggered by a serverless Lambda.

So you have these jobs rendering on a regular basis and with Athena running queries, you can now take useful information out of raw data and push that into any other BI platform you might have at that time, including a backend for the application itself to then start building personalization into the platform, which ended up being a pretty useful project at the time so.

Jeremy: Yeah, no, and I think that the really nice thing about this Kinesis, to S3, to Athena, and you don't even necessarily need Glue. I mean, depending on how you're writing your data. Especially with Kinesis Data Firehose you can automatically convert it. You can actually run conversions while it's processing before it puts it into the S3 bucket. But what I really love about that is essentially is a 100% serverless and you're just really paying for when you run the queries, obviously you're paying for the shards and stuff like that for Kinesis.

But you're not paying to just store this ton of data in an expensive storage place. I mean, S3 is relatively inexpensive for that amount of data. If you were to use something like Elasticsearch, which is a popular analytics tool and so forth. We tried that at one point and we were just doing some of the projections with the number of clicks we were collecting per day. And it would've just kept on adding more and more data, more and more data.

We had to keep making the drives bigger and bigger and bigger in order to handle all this data. And we were thinking about, well, maybe aging some things out, doing some roll-ups, aggregations, and that's all possible. And you could still do that with S3 as well. But I really, really love that use case because, that is such an important thing now is capturing that data to understand what your users are doing on your site.

And whether it's for personalization or whether it's for other types of optimization, or you're collecting clickstream data for AB testing or any of that sort of stuff. It is just a really, really good use case. And I think with the tools in place now serverless just handles it so well.

Gareth: And the interesting thing is we actually... To start off with, we looked at two models and we went with the Kinesis model, but we even investigated using just API Gateway with a Lambda function, dropping items into S3 as one potential method to handle this. Just because that was what were familiar with at the time. And that works and the scale that you can get with that is pretty impressive.

Kinesis just ends up being a more performant and a cheaper way to run these kinds of operations. But if API Gateway and Lambda is something you need as well, it can still handle these kinds of clickstream events. Just there are often specific tools made for a specific job, which is often a good one to go with.

Jeremy: Right. All right. Again, clickstream sort of does fit into that frontend piece. But so as we move past this, one of the things that I know I've seen quite a bit of is people using just the power of Lambda compute, to do things right? And what's really cool about Lambda is Lambda has a single concurrency model, meaning that every time a Lambda function spins up, it will only handle a request from one user. If that request ends, it reuses warm containers and things like that.

But, if you have a thousand concurrent users, it spins up a thousand concurrent containers. But you can use that not just to process requests from let's say frontend WebSocket or something like that. You can use that to actually run just parallel processing or parallel compute.

Gareth: Yeah. This is one of what do they call it, the Lambda supercomputer.

Jeremy: Right?

Gareth: You can get an enormous amount of parallel... Try to say that three times quickly. Parallelization with Lambda...

ON THE NEXT EPISODE, I CONTINUE MY CHAT WITH GARETH MCCUMSKEY...

THIS EPISODE IS SPONSORED BY: AWS (Optimizing Lambda Performance for Your Serverless Applications)

View Details

About Alex DeBrie:

Alex is a trainer and consultant focused on helping people using cutting-edge, cloud-native technologies. He specializes in serverless technologies on AWS, including DynamoDB, Lambda, API Gateway, and more. He’s an AWS Data Hero and recently published author of The DynamoDB Book and the creator of DynamoDBGuide.com. He previously worked at Serverless, Inc., where he held a variety of roles during his tenure, helped build out a developer community, and architected and built their first commercial product.

  • Twitter: @alexbdebrie
  • Blog: https://www.alexdebrie.com/
  • DynamoDB Book: www.dynamodbbook.com (Discount Code: SERVERLESSCHATS)
  • DynamoDB Guide: www.dynamodbguide.com

Watch this episode on YouTube: https://youtu.be/GZTLFWlEnaw

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly and this is Serverless Chats. Today I'm chatting with Alex DeBrie. Hey, Alex, thanks for joining me.

Alex: Hey, Jeremy. Thanks for having me.

Jeremy: So you are actually returning to Serverless Chats. You were my first guest, and you are also my first returning guest. So I don't know if that's an honor, but thank you very much for being here.

Alex: Yeah, it was an honor to be here the first time and honored to be back as well.

Jeremy: So a lot has changed since you were with me almost a year ago. You went out, you used to be working at Serverless, Inc., where they created the Serverless Framework. You went out you started doing some consulting, you were named an AWS Data Hero. So why don't you tell the listeners a little bit about yourself and what you've been doing over these last few months?

Alex: Yep, sure. So as you mentioned, I used to work for Serverless Inc, creators of the Serverless Framework. That's how you and I got hooked up initially. Worked for them for about two and a half years. And then last fall I was named a AWS Data hHero specifically focusing on DynamoDB which was a big honor for me. And then in January I left Serverless Inc to go on my own to do a few different things, some consulting, some teaching and also finished up this book I've been working on.

Jeremy: Yeah and so speaking about this book, I'm super excited about this because I remember we were out I think in Seattle at one point several months ago and I looked over your shoulder, I saw you typing and I asked, "What are you doing?" You're like, "Oh, I'm writing a book on DynamoDB, of course."

And you obviously you created the DynamoDB guide at dynamodbguide.com which is a really great resource for anybody looking to get familiar with DynamoDB. It's much more approachable I think, than the documentation on AWS. It's really well written and there were a couple of modeling strategies in there and things like that but this new book, which I've had a chance to read which is awesome, by the way, so congratulations really, really well done. But this new book, just you know it is not DynamoDB guide repackaged.

This is a whole new thing with tons of strategies, tons of information. So why don't you tell us a little bit about this book?

Alex: Yeah, sure thing. So as you mentioned, I created dynamodbguide.com. That was about two and a half years ago now. And it was basically, I'd watch Rick Houlihans re:Invent talk over Christmas one year, and just had my mind blown and rewatched it so many times, and scribbling it out in Notepad on how it all worked, and then wanted to share what I learned. So I made this site DynamoDBguide.com.

That did pretty well. And I've stayed in touch with the DynamoDB team since then. But I really wanted to go further than that. Because I think like you're saying there are some missing stuff out there. So I've been working for the last almost about a year on this book.

I started I think, last June or July. And really, we just want to go deep on DynamoDB and not just the basics, all that stuff really introduce this idea of strategies, introduced some data modeling examples to show that you can really handle some complex access patterns. It's not just about key value storage, you can do complex relational data in DynamoDB.

Jeremy: Yeah, definitely. And so just in case somebody doesn't know what DynamoDB is, let's just give them a quick overview of what exactly that is.

Alex: Yep, sure. So DynamoDB is a NoSQL database offered by Amazon AWS. It's a fully managed database. I'd say, it got started where amazon.com their scaling needs just you know, they were out scaling their relational databases. So they built this underlying storage mechanism that... they built this underlying database to replace their relational databases. That was used internally at Amazon. They released some of the principles behind it in this DynamoDB or this Dynamo paper.

That eventually became a service in AWS called DynamoDB. Fully managed service, works really well for highly scalable applications. In fact, all the tier one services at Amazon and AWS are required to use it. So if you think about the shopping cart or the inventory system or IAM or EC2, all that stuff that's all using DynamoDB under the hood. But also it's gotten really popular in the serverless ecosystem just because the connection model, the permissions model, the provisioning model, the billing model, it all works really well with everything we like about serverless compute.

So a lot of people have been using it there. And that's how I got introduced to it mostly and just wanted to go deeper on it and really use it correctly.

Jeremy: Yeah, right. And so one thing that's super important to remember is DynamoDB is NoSQL, right? Or NoSQL, it is not like your traditional RDBMS system. There's no joins, right? You're not doing any of that sort of stuff. And there's reasons for it, obviously, it's a speed thing, and I did a whole episode or actually I did two episodes with Rick Houlihan himself and he went through a bunch of those things. So if you want to really learn or get a good audio overview, I guess, of DynamoDB I suggest go back, listen to those episodes because I want to use your time today to actually go through a couple of things in the book that I found to be just really helpful, like things that I don't think they pop out to you when you read the documentation.

And I've talked to so many people, because I love DynamoDB. I have my DynamoDB toolbox. I'm working on a new version of it right now that I'm thinking like, it's just going to make my life easier. Hopefully, it makes other people's lives easier. But I just use it so much. And the problem always is, is I think a lot of people think that it's just a key value store, right?

And it is to a certain extent, but there are ways to model data that are just, I mean, they're fascinating. It's absolutely amazing what you can do with some of these things. So, I'd love to point these things out. Because I think like I said, these are things that will not jump off the page at you when it comes to documentation. So the first thing that I think you did a really good job explaining was the importance of item collections. And this is something for me where I always think about them as folders with files in the folders and try to think about it that way. But you probably do a better job explaining it.

Alex: Thanks. I hope so. So, yeah, I introduced the concept of item collections and their importance pretty early on. I think it's in chapter two. And it was actually one of the solutions architects at AWS named Pete Naylor that that turned me on to this and really made me key into its importance.

But the idea behind item collections is you're writing all these items into DynamoDB, records are called items in DynamoDB. And all the items that have the same partition key are going to be grouped together in the same partition into what's called an item collection. And you can do different operations on those item collections, including reading a bunch of those items in a single request.

So as you're handling these different access patterns, what you're doing is you're basically just creating these different item collections that handle your access patterns. And that can be a join like access pattern. If you want to have a parent entity and some related entities in a one to many or many to many relationship, you can model those into an item collection and fetch all those in one request.

You can also handle different filtering mechanisms within an item collection, you can handle specific sorting requirements within an item collection. But you really need to think about, hey, what I'm doing is I'm building these item collections to handle my access patterns specifically.

Jeremy: Yeah, and I think that is something where people like, and I don't want to speak for other people, but these are the questions that I'm getting. And I'm sure you've gotten these similar questions where it's like, well, how do I join data? How do I how do I represent things like, I don't know, one to many relationships and things like that? And you explain all that in the book. I actually want to talk to you about that a little bit later.

But I think this comes down to this idea of like you said, it's putting these things into collections, it's knowing which groups of items that you actually need to have available on the same partition, right?

So again, I'm probably not doing a great job explaining it, which is why you need to buy the book and read what you wrote. But I think this brings us to another thing, just this idea of partitions and consistency, right? Because there are, the way that this works is a lot different than your relational database lookup would work. So understanding partitions and consistency, I think is actually really important to properly modeling your data as well.

Alex: I think that's right. I mean, these are more underlying architectural concepts, but they really do help your data modeling and really your mental model of how DynamoDB works and how you should organize your data. So let's start first with partitions. If you're working with a relational database, generally all your data is going to be all together on one node unless you're doing some pretty complex sharding strategy if you have really huge amounts of data.

But with DynamoDB, what they're going to do is they're going to try and split your data across a lot of different nodes. And each of those nodes are called partitions. And what they do is they use that partition key that we talked about to create these item collections. They use that partition key to find out which node that data should go to. So when you're writing an item to DynamoDB, they're going to look at that partition key, they're going to hash that partition key and then assign it to a particular node based on the value of that hash.

And then when you're reading data out of DynamoDB, what they're going to do is they're going to look at the partition key that you want to read from, they're going to look that up in their little hash table and figure out where you need to go to find that data. And what that does is it transforms. If you have a 10 terabyte table, instead of having to scan that whole table, it chops it up into these little 10 gig chunks, makes an 01 lookup in a hash table to figure out which node it belongs to. And now you're working on a much smaller amount of data, maybe 10 gigs or less there. And that makes for a lot more efficient operations down the road within that particular partition.

Jeremy: And then, what about consistency, though? Because I mean, that's the thing, whereas you start writing data across multiple nodes. If you, you know, there's the whole CAP theorem, right? I mean, if you are to access a node that hasn't been replicated to yet, then the data wouldn't be there.

Alex: Yep. Great point. So within each of those storage nodes, or those partitions that I was mentioning, there's actually going to be three copies of your data. So there'll be a primary and two secondaries of that data.

When that data gets written, it's going to write to the primary and one of the secondaries before it actually acknowledges that, that write was successful. And that's just for fault tolerance there to make sure that it's committed to a majority of those nodes.

Then after it's returned to you, they're also going to asynchronously replicate it to that third node, which is that second secondary. Now, when you're doing a read from DynamoDB, by default, it can read from any of those three nodes. So then you get into that situation where, if you've just written an item, and you do a read request, where it hits that third node, that second secondary, where it hasn't asynchronously replicated that yet, you have a chance of getting some stale data.

So this is, by default, DynamoDB gives you eventually consistent read guarantees, which means, that data is going to get there eventually, but you might not see the latest version of that data on a particular read.

Now, if you want to, you can opt into what's called a strongly consistent read, which says, "Hey, give me the most recent version of my data with all the rights included." If you do that, it's going to go straight to that primary node that has all those rights committed and read the data from there. You can guarantee you'll have all the latest rights.

So if you're doing a banking application or something like that, you might want to do that. But in a lot of cases where that replication lag is pretty small, a couple hundred milliseconds or less. Usually, you can handle that eventual consistency with DynamoDB.

Jeremy: Yeah, I think it's even less than the 100 milliseconds. I think I know, for GSI it's something like sub 10 milliseconds or something crazy like that.

Alex: I mean, they're pretty good most of the time, yeah.

Jeremy: Yeah. So besides, like you said, the banking use case, which is interesting. I haven't personally found a lot of use cases for strongly consistent reads. Are there a lot of them out there? Because for me, I almost feel like I'm never just writing data and then immediately reading it back.

Alex: Yeah, I would say in most of my applications, I just use that default, eventually consistent read, which if you do that, you're going to pay half the read capacity units of actually making that read so it's cheaper if you opt into eventually consistent reads. And all your rights are going to be actually strongly consistent, because in that case, it's going to that primary node. So you don't have to worry about writing the wrong data in some sort of way.

So in most cases, for me, eventually consistent has worked, it depends on your application needs, but I think for a lot of people it works, especially given how small the lag is on that replication.

Jeremy: Yeah, makes sense. So the other thing that you talk a lot about in the book is this idea of overloading keys and indexes. And obviously, if you're doing a single table design, and you've got multiple entity types in there, you only have one PK and SK at least for the primary index. So you have to put different, maybe different types of keys or whatever in those or different types of identifiers in those. You might be reusing things like your GSIs, your GSI PKs and SKs and things like that. So what are some of the strategies for maintaining your sanity with that maybe?

Alex: Yep, sure. I mean, one thing I do is I make what I call an entity chart whenever I'm making an application. And what I do is I create my ERD that has my entities and relationships, all that and then I list out all the entities in a chart.

And I say, and then I have a column for PK and a column for SK, which are the two elements of my primary key. And as I'm working through modeling that I just state out this is my customer entity, and it has this PK pattern, it has this SK pattern, I write that down.

This is my order entity, it has this PK, this SK and I write that down, all the way down the list. And if I'm adding secondary indexes, I add new columns, I have GSI1PK, GSI1SK. And I'm adding out what those patterns are. So then when you get to the end, you have this chart that says okay, this is the pattern I need for each of these entity types. This is how I'm going to handle all that stuff. And then you actually go into implementation and your items, you're just making those patterns and decorating your items with these GSI1PK, GSI1SK values as you need them.

Jeremy: Yeah, so that is another really interesting thing that you point out. And this is something I talked to Rick Houlihan about was this idea of not reusing just any old attribute, as secondary indexes or as the PK and the SK for your secondary indexes. And one of the reasons for this was, you can just bang your head against the wall when you're trying to model data. And you say, "Okay, well, let's see, if I make my date the SK on GSI1, then that'll fit these patterns, but then I've got a bunch of extra data in there that maybe doesn't need to be there. Or maybe I needed to be different with a different access pattern."

So I really like... I don't think I articulated this well when I was talking to Rick about this, but you do a great job. You basically say just separate out your attributes from your indexing attributes, I guess.

Alex: Yeah, so I split them into two buckets, like you're saying application attributes. And those are things that are meaningful in your application, in your code. So if you have a customer, the customer name, the date of birth, the address, all those things are what I call customer attributes, or application attributes that are useful in your application. I contrast that with indexing attributes, which are attributes that only exist on your items to index them properly in DynamoDB so you're putting them in the right item collection to handle your access patterns.

So this is going to be the PK and SK for every item, but also GSI1PK, GSI1SK. And for me, I say separate those completely. Don't try and cross the streams, as they say. If you have your PK and SK and even if your PK for that customer has the customer username or whatever in it, and I want to try and parse out the username from that PK and SK, I would just duplicate it, have a username attribute that has that application attribute in there, and don't worry about it.

And then also like you're saying, if the PK and SK pattern for your user is the same as your GSI1PK and SK pattern for that user, don't try and duplicate those across because like you're saying, you're really just going to tie yourself into knots trying to make sure that it works for all your different kinds of items that you're doing there.

Rather, I would just say duplicate those attributes over, even if it is the exact same value somewhere else, just add those over so you're not pulling your hair out trying to make these indexes fit.

Jeremy: Yeah, no, I totally agree with this idea because I mean, I've built tables in the past that would use like this fancy nesting structure or hierarchical structure for like the SK. You'd use that hierarchy sort of pattern. And then what you try to do is parse those out when you want to read those with your application. And even if you've built a layer in between there, that does it for you, it's still... There's just a bunch of things that could probably go wrong, right?

So if you have state and county and city and things like that, and you're laying them into a hierarchy, because that's one of your access patterns, and you need to use that for sort. That's great, definitely do that. But I would like you said, store each one of those things separately so that at any point, you can one, easily parse them out of the table without having to worry about trying to figure out what your index pattern is, or what the sort key pattern is. But even more so then if you need to recreate that sort key at any time, you've got all the data right there that you can do that with.

Alex: Yep, absolutely. And one other thing in this vein that I used to do is, I would have these different access patterns, say on orders where I want to fetch orders by date for a particular customer. And what I would do is I'd create a new index based on those application attributes for my order. So I'd make the partition key for that index be the customer ID on the order and I'd make the sort key on that index be the order date, rather than using these more generic attributes like GSI1PK, GSI1SK.

And I would say just get into those very generic indexing attribute qualities there, put them all in there, and then it becomes very methodical and more scientific on how you're creating these item collections rather than having all these ad hoc indexes all over the place with weird attributes and hard to understand what's going on.

Jeremy: Yeah, definitely. All right, so the next thing that I thought was really interesting, although it does make me question your sources, because I apparently you've learned this or you at least were inspired by something that I said, which is to add a Type field to every item in your table.

Alex: Yep, that's right. So I was actually complaining on Twitter about how difficult it is to export your DynamoDB items into an analytic system, because now you've got this single table that you need to renormalize into different things. And Jeremy's just like, "Hey, why don't you put a type attribute on every item that says, this is a customer, this is an order." And I'm like, "Okay, that makes it a lot easier." So yeah, the advice here is just on every single attribute that you're writing... every single item that you're writing into that table just include this type attribute that indicates what kind of item it is.

So whether it's a customer, whether it's an order or something like that. It's going to help you in a few different scenarios. Number one, if you're in the AWS console, just sort of debugging your table, it can be easier rather than trying to parse your PK and SK pattern and figuring out what that was, how it translates to an item type. If you can just see that type there, and you say, "Okay, this is a user, this is a order," whatever that is.

Number two, if you're doing a migration, which is common, especially a migration, where you need to add new attributes to existing items. What you're often going to need to do is scan your table and find particular items that you need to update. And if you have this type attribute on there, you can use a filter expression to just get those types of items. And now you're again, not parsing out some weird PK and SK patterns to handle that. And then finally, I think that the place where it fits best is the one you said where you're exporting it for analytics, because DynamoDB is not great for OLAP type queries. These big aggregations of saying, "Hey, what items sold the most last week," or, "What's our week over week growth."

So something like that you'll export to an external system whether that's something like Amazon Redshift, whether that's S3 and use a thematic query on that. But then having that type attribute there is going to be really helpful as you filter down to the particular items you want or maybe do additional ETL to split them out into different tables to where you can then do joins in your relational analytic system.

Jeremy: Yeah, and it's funny, because the way that I actually came up with adding the type thing, and I'm sure other people have thought of it as well. So it's not just me, but when I was building DynamoDB toolbox, I was thinking ahead to single table designs with multiple entity types, and then trying to figure out how to do the parsing of those types when they come back out.

Because again, you have a lot of overlapping attributes, right? You sometimes can't tell just from the PK and the SK exactly what type of item it is. So having that little bit of data is great, even if you build your own data access layer to say, "Hey, I need to take this item. I know it's a certain type. Maybe I want to maybe unmarshal it into some other format that's easier for me to understand." That is just a really, really great strategy to be able to do something like that.

Alex: Absolutely. Yep, totally agree.

Jeremy: Yeah. So speaking of strategies, this is another thing that I thought was really well done in the book, because you can learn everything you need to know about DynamoDB. You can know about consistent reads, you can know about partition keys, you can even understand how, yes, I can put all of these things into the same partition and all that kind of stuff. All the query capabilities, the condition expressions. I mean, there is a lot to learn with DynamoDB, just on the mechanics side of it. But when it comes to modeling, that is where I think there is just not enough information out there that gives you really, really good strategies to be able to do that.

And then there's a really helpful document on the AWS site that is like DynamoDB best practices. Gives a couple of overall points of some of that hierarchical stuff, some of the edge nodes type thing, and those sort of strategies, but they really don't go into deep detail.

You actually like, I mean, you have a whole big long section just on strategies here. So maybe start by what's the importance of having good strategies when approaching modeling?

Alex: Yep, sure. This is something that I evolved to over time because my original draft of the book was seriously going to be like six chapters of basics, and then like 20 chapters of examples, because I had all these little things that I wanted to show. And I'm just like that's like way too long. And what you see happen is they're all these little patterns that you see a few different times.

And the big thing for me is that this is very different than a relational database. With a relational database, there's generally one way to model your data and you make your ERD and then you actually put your data into your tables. You have your different table for each entity, you structure the relationships between them. And then you write your queries based on this one way that you've written your table.

You might add some indexes to help some things out. But generally, there's one way to do things in a relational database. With DynamoDB, it's different. You're going to model your data very much based on your application needs. So then there are different strategies and patterns for how you actually do it. And I go through that in the book, a few different chapters, there's one to many relationships, there's many, many relationships. There's sorting, there's filtering, there's migrations, and then just some additional grab bag strategies as well.

Jeremy: Yeah, so you actually outline a number of different strategies. So rather than just saying, "Oh, yeah, strategies are important." You actually go in and write about them. So I think there were five different ways that you outline to handle one to one... Sorry, five different ways to handle one to many relationship. So I mean, I would love it if you could just give us a quick overview, and then again it's there's a lot of information, so you have to dig into the book if you want to find out more about them, but just maybe an overview of each.

Alex: Yep, sure thing. And I think this is like the clearest way to explain strategies to people. Because there's one way to do strategies in a relational database. And it's that foreign key relationship. But like, when you tell people, there are five different ways, and here are the situations you might want to use them. I think that really opens up people's eyes.

So that first strategy I use is called, I call it denormalization plus using a complex attribute. And so DynamoDB has this notion of complex attributes where you can store a list of data or a map, or a dictionary of data on an item directly. And so that's denormalizing your data. That gets you away from that normalization principle in relational databases because now you're storing multiple elements in a single attribute. But the example I give here is imagine you have a customer in an E-commerce store and they are going to have multiple mailing addresses they want to save. They want to have their home address, their business address, their parents address because they send them a gift sometimes.

What you might do, instead of splitting that out into different items, if that's small enough, you can actually just put that as a complex attribute directly on that item. And then whenever you fetch that user, you'll get all those saved mailing addresses with them that you can show in the context of their user profile, or their order checkout page, or whatever that is.

Jeremy: Right. But that strategy works great for very, very small bits of data. Because one thing is, is that again, you pay every time you read data from your table. So if you have 100 addresses stored in there, and all you're trying to do is just get the customer name, then you're loading a lot of extra data and paying for that read that you don't need. But I like that strategy for very small, small groups of data.

Alex: Yep, yep, absolutely. So that one works. Yeah, if you don't have any access patterns on that related data itself, and also the number of related items are going to be small. So it's like you're saying, if you can eliminate them and say, "Hey, you can only save 10 addresses," that works well here.

Second pattern that is in the book is, I call it denormalization by duplicating data. And again, this is denormalization because in relational databases, you don't want to be duplicating data. If you do you split that into a different table and sort of refer to it via these foreign key relationships.

But the example I use here is maybe you have movies and actors or books and authors or anything like that, where on each of those book items, maybe you replicate some information about the author, such as the author's name, the author's birthdate, things that aren't going to change, especially. Stephen King's not going to get a new birth date. So rather than referring, or having to join up with his author profile every time, you can just store it directly on that book and show that data if you want to.

So again, you're getting away from normalization that would have in a relational database, but it's okay in this particular situation.

Jeremy: Yeah, and denormalization when it is something that is highly immutable, right? Like you said, the birthday is not going to change.

Alex: Yeah.

Jeremy: Let's say for some reason that somebody changes their name or something that you think isn't going to change. There's still plenty of strategies to go back and clean that up if you need to.

Alex: Yeah, absolutely. If it's one off updates, you know this duplication can work and maybe you just have to handle those one off updates. But if you actually have consistent access patterns where you're going to be needing to update information about that parent item, then you want to use some different strategies.

The two most common strategies I see for one to many relationships are the primary key and the query and the secondary index and the query. And both of these rely on that concept of item collections that we were talking about before. Where you're assembling different collections of items into a single item collection and using that same partition key. And then you can use this query operation that lets you fetch multiple items in a single request. And you're basically pre joining your data.

So imagine you have an access pattern where you want to fetch a customer and the customers most recent orders, which that would be a join in your relational database, but joins aren't as efficient once you scale. So DynamoDB doesn't have joins. So what you do is you're pre-joining them for your access pattern into this particular item collection to handle that access pattern that you have.

Jeremy: Yeah, and then what about the fifth strategy?

Alex: Yeah, the fifth strategy is kind of an interesting one. This one's called a composite sort key. So usually you're using what's called a composite primary key where you have a partition key and a sort key making up that composite primary key.

You can actually use this composite sort key pattern where you're encoding multiple values into that sort key to represent a hierarchy. And this works really well if you have multiple levels of hierarchy, and you want to query across those levels. And maybe sometimes at a high level, sometimes at a low level.

The example I always give here is think about like store locations, right? Starbucks locations where there's a hierarchy. Starbucks has some locations in different countries, the US versus France versus Mexico. But then, and maybe that's your partition key, but then your short key they have locations in different states, in different cities, in different zip codes. So maybe you'd encode each of those into that sort key. And now you can search at any level in that hierarchy.

You can search for a particular state, you can find all the Starbucks within New York State. You can find all the Starbucks within New York City or you can find it within a particular zip code to get all those Starbucks as you need it. And then it allows you to be querying at these different sort of levels of granularity as you need it.

Jeremy: Yeah. And that's a super powerful use case. And I actually find that works really well sometimes too, for secondary indexes, because you can, it's easy to write that data in, where you might have a more, you have a different type of identifier, maybe for the primary index, but then that secondary index, you just copy that in.

And of course, we didn't even talk about projections. That's a whole other thing to talk about. But that's something to think about too when you're building your secondary indexes.

Alright, so those are the one-to-many relationships, but I think we see quite a few many to many relationships in different types of data and different models and you give four strategies to handle those.

Alex: Yep, sure thing. So yeah, like you said, there are four different many to many relationship strategies that I outlined. The first one is one that I call shallow duplication. So the example I have here is imagine you have students in classes where a student can be in multiple classes, but also a class has multiple students that are enrolled. You can think about your access patterns, and for each side of those access patterns think like what do I need to show about these related items?

So if you're showing that class, maybe you're showing information about the class, the name, the teacher, the code, whatever that is, and you also just have a list of students in your class, but you don't need all the information about the students, maybe you just need the student name, and then it has a link to that student, and you can go click on them and go over there if you want to find more information about the student.

If that's the case, we'll do something similar to what we did in that one to many relationship section with that denormalization, that complex attribute and we'll just store a list or a map on that parent entity in that many, many relationship on that class item and you just have a student attribute that could be a list, that could be a map, whatever it is, and it just has the list of students in there.

And you don't need to know their GPA or their graduation date or any of that other stuff that could change. So you don't have to worry about updating all those different items. That's shallow duplication. That works sometimes, again, think about your access patterns and what you're going to need there.

The second one, and the more common one that I see. This is what I see mostly is one that's called adjacency list. And basically what you're doing is you're representing each side of that relationship in a different secondary index or a different primary key. So maybe, on that main table, on your primary key, you have your each side of the relationship. So let's, I think the example I use here is movies and actors where movies have many actors in them, and actors can be in many different movies.

So I have three different entity types. I have the movie, I have the actor, and then I have a role item that actually represents a actor in a particular movie. And once you do in that primary key, you set it up so that you can handle one side of that relation.

So maybe you put the movie and the role items all together. So it allows you in that single item collection to use that query and say, give me the movie, and give me all the roles in that movie, which is going to give you the actors and actresses that played roles in that movie. And then you have a secondary index where you're flipping that and you're putting all the actor or the actress and all the roles that they've been in, in the same item collection as well.

So now when someone clicks on Toy Story, they can hit that primary key and get Toy Story and all the roles in that one, or when someone clicks on Tom Hanks, and they can go to that secondary index, get Tom Hanks and all the roles he's played as well.

Jeremy: And it's one of those things that I think is really hard to visualize in your head, and even hard to visualize if you're trying to do it in Excel or like Google Sheets or something like that. And I know you love this tool, and I have been using it quite a bit recently is the NoSQL workbench for DynamoDB, which allows you to create those indexes, put some sample data in there, you can even use facets to define the different types or the different entity types, and then be able to flip that data so you can see how those relationships work.

So if you were totally confused by what, Alex just said, it's not uncommon, because it is hard to think about switching that. But yeah, so that's awesome. So what else?

Alex: Yep. So the next one is, this is probably the hardest thing that I have wrapping my head around in the entire DynamoDB ecosystem. This is called the materialized graph. And this is if you have a lot of different maybe node types or different types of entities and a lot of different relationships between those nodes and you want to be able to ritually query across those.

What you do here is you might in your primary key, maybe for a particular node, you put them all, you put a bunch of different items into an item collection, but each item in that item collection represents maybe a relationship that that node has to something else.

So for example with me I might have an item that represents, I have a relation to a particular date, which is my anniversary date to my wife. So that's a particular item.

Jeremy: And you don't want to forget that.

Alex: Exactly, I'll get in trouble if I lose that one. So I also have another node that represents where I live. And I have another node that represents my particular job. And so I've broken me as a person and all these different relationships that I have. And then I have a secondary index, where now maybe I'm grouping together all these different relationships into other things. So I can see the nodes and relationships to them.

So if I look at that date, that anniversary date of mine, I can see all the people that are related to it, and maybe that's someone else's birthday. But also, my wife has an entry in there too, it's her anniversary date as well. And we can see both of those relationships there. Or I can see all the people that are also developers. I can see all the people that live in Omaha.

So it allows you to query against that and say, what is this node, I can get a full look at that node if I want to, but also who has relationships with that node. And if I want to go find information about those nodes, then I go back to that primary key and query and get that full information about that node.

So, it's tough. I'd say if you're deep into graph theory that can be good, if you have a very highly relational model that you want to represent all these relationships that can help. But I don't even have like a great example for it in the book because it kind of twists my mind so much.

Jeremy: And there's Amazon, was it Neptune?

Alex: Yeah, probably use a graph if you're really going to be going down this road, so.

Jeremy: Right. Was that all the strategies?

Alex: There's one more, there's one called normalization and multiple requests. And this one sounds anathema to DynamoDB. Because DynamoDB is about denormalization. It's about trying to satisfy these access patterns in a single request. But sometimes it's pretty tricky.

And I think especially with many to many relationships, where if you need to fetch all the related items, and those related items can change, it can be hard to handle that in a single request consistently.

So with this, when you sort of normalize your data you might have a pointer back to two other items you need to go fetch, but then you need to make a follow up request to actually fetch those items.

And the example I use here is Twitter and like the people that I'm following, right? I can list the people I'm following on Twitter, I can follow multiple people that can be followed by multiple people. But when I'm viewing all the people that I'm following, I need to see their latest profile picture and their latest profile description, and their name and all that stuff, which can change a lot.

And that can be hard to duplicate that out if you're duplicating that data a ton. So instead, what I'll just say is, hey, this user happens, or I happen to follow this user, I get that list of people that I'm following, and then I make a follow up request to Dynamo to get all the information about them. Their profile picture, their name, their profile description, all those things.

Jeremy: Yeah. And that's one of those strategies too where it's like, you'd rather not have to use that one, Because you don't want to make all those different requests. And I think this is another thing that is super important with especially from a get perspective, is the idea of caching, right? I mean, just if you are fetching data like somebody's profile from Twitter or something like that. If you've got that stored, the more you can cache that kind of data so that you don't have to pay for that round trip to the database to get it, I think is really, really helpful. And then just having good timeouts or making sure that the data doesn't get stale, and things like that are great ways to do it.

So, alright, so those are awesome and definitely, I mean, check them out. Because these are things that are probably really hard to wrap your head around without seeing those examples. You give examples of these things. And it just makes it so much easier to digest. So that's really great stuff.

Alright. So the other thing that you talk a lot about in the book, and this is something where, again, super powerful when it comes to secondary indexes, is this idea of sparse indexes.

Alex: Yep, I think sparse indexes are one of my favorite patterns and strategies to use. So just to add some background. If you create that secondary index DynamoDB is going to replicate that data from your main table into that secondary index with this reshaped primary key that enables different access patterns, it's grouping into different item collections. As you're doing that Dynamo is only going to replicate it into that secondary index, if that item has all the elements of your secondary index as primary key.

So if you've got this GSI1PK, GSI1SK defined on your secondary index, but then you write an item that doesn't have those attributes, it's not going to replicate it into that index.

So what happens is, you actually end up with sparse indexes by default, because some of your entities won't have additional access patterns. So you won't be copying these entities into those secondary indexes. But sometimes you want to use a sparse index strategically. And what that sparse index is, is you're filtering out items specifically to help you with an access pattern. So you're applying a filter as part of that index. And there are two different strategies here. One is when you're filtering within a particular entity type based on some attributes about that entity. So the example I have in the book is imagine you have an organization with a bunch of users in it, and a small number of those users are administrators for the organization. And say you have an access pattern where you want to get all the administrators.

Now, you could do that query on your primary key and get all users and filter through that, but maybe you have 3000 users, and that's going to take a long time. So instead, what you do is on the users that are admins, you add these attributes, GSI1PK, GSI1SK. Now only those administrators will be copied into that secondary index. And you can use that access pattern to say, "Okay, give me all the admins for this organization." You can do that quickly without getting all the users that aren't admins.

Jeremy: Right, and that's and just I mean, that has so many use cases, right? I mean, if you're looking for error conditions, or orders that are in a certain state, or tickets that have risen to some severity level or things like that. That is, the copying of that data is very inexpensive, and you don't have to copy all of it right, again, back to the projections thing. But you copy a tiny bit of data over that just gives you quick access to that list. And then yeah, I mean, I actually love that strategy.

Alex: Yep, I think it's so cool. The great thing about this, since you're just filtering within a single entity type, you can use this with that overloading keys and indexing strategy. So you can have different entities that you're also copying into that secondary index and handling different access patterns on. So it works really well, it keeps it efficient, and limits the number of secondary indexes that you're using.

Now, the second pattern with sparse indexes is where you're projecting only a single entity type into that index. And the example I use here is imagine you have an E-commerce store. And you have a bunch of item types in your table. You have customers, you have orders, you have inventory items, all that different stuff.

Now your marketing team says, "Hey, occasionally, we want to find all the customers and send them an email because we're doing some cool new sale." And you're like, "I don't want to scan my 10 terabyte table just to find those customers because I have way more orders, I have way more items than I have actual customers, use that sparse index here." So what you're doing on this one is on each of those customer items, you add some attribute. It could just be customer index, or whatever you want to do there, just on those customer items, you make a secondary index that uses that attribute. And now in that secondary index, the only items that are getting put in there are those customers.

So you don't have orders, you don't have inventory items, any of that stuff. You can run a scan directly on that secondary index, you'll get just the customers and you can send them all an email or whatever you want to do, and that works really well. The difference with that previous strategy is that you can't use this with that overloaded keys and indexing strategy because you can't be jumbling up other items in there and handling other access patterns. The point of this index is to only hold those types of items so that you can find all of them as we need to.

Jeremy: Right, and maybe even a better way to handle that use case would be to use DynamoDB streams every time a customer record changes, push that into Marketo, or whatever your marketing software is, and do it that way.

I mean, that's I think another thing too, where we need to think about what are the use cases when you're accessing DynamoDB, right? Because if your use case is I want to get a list of all my customers, how often do you get that list of all your customers, and that's something that you need the power of DynamoDB to do. Especially, you can only bring so much data back, I think it's one megabyte per get request... or sorry, per query. And so you're going to be limited to how much data you can pull back, and then you got to keep looping through it. And that's kind of a pain.

Whereas just dumping that data for internal dashboards that you're going to use, put it in mySQL or put it in Postgres or something like that, and just replicate off of that stream. I love using those strategies, and try not to overthink it. I try to think of, if I have thousands and thousands of users banging up against my table, what do they need to do quickly? And that's what I try to optimize for. But certainly, I mean, again, there's other strategies or those strategies you mentioned are useful for obviously forward facing things or user facing things as well.

Alex: Yep, absolutely. Good point.

Jeremy: Alright, so another thing that I think people get hung up on with DynamoDB is that I mean, you are limited with your querying capabilities, right? I mean, you can't really... calling it a query is probably the wrong thing.

I mean, essentially, what you're doing is you're requesting a partition, and then some little bit of control over that sort key, I think the most, the best one you have is the begins with the, right? That's the search capability. Other than that, though, you're looking for values that are between a certain thing, or you're looking at value that's greater than or less than, or something like that.

So you are a bit limited in your querying capabilities when it comes to that. But one of the things I think, again, where you see this use pattern, or you see this use case quite a bit is needing to sort by date, right? I mean, that's just an obvious thing. But the problem is, is that if you just put a date in as your sort key, it's always a tough way to access it when you're like, "All right, I need to know that date that I put the item in there. Otherwise, I don't have a way to access it.

So it's not as easy saying, "Oh, I need customer one, two, three for example. But there are some different ways that you can do that. And you mentioned in the book, this idea of these case or user IDs or these KSUIDs. And so can you explain that a little bit?

Alex: Yep, sure. So KSUID, this was created by the folks at Segment and it's pretty cool. And like you're saying what you need is some sort of random identifier for an item. So this could be an order ID or a comment ID. Ideally, it's something that can be easily put in a URL path as well if you're doing some restful stuff there, but you want some sorting across that data set. Because you might want to say give me a user's most recent five orders or the most recent 10 comments for this Reddit post, anything like that.

So if you're using a generic UUID, you're not going to get that sorting capabilities. So this KSUID really interesting. It basically combines the uniqueness of a UUID with some sorting mechanisms where the prefix on this UUID is an encoded timestamp, basically. So then it's all going to be laid out in order, within your particular item collection, you can get the most recent five comments or the first five comments or anything like that. You can get chronological ordering for it. But also you get uniqueness, you get URL friendliness. It works really well that way. So I found that like, pretty late in the book and just fell in love with it and started using it everywhere.

Jeremy: And I think that's something that Twitter did with their tweet IDs, as they called it like snowflakes or something like that.

Alex: Yeah.

Jeremy: Yeah. Which I always thought was really fascinating. I remember, I was working on a startup. And I remember spending about a day and a half trying to figure out how to generate the right IDs with ordering and going down that whole snowflake thing thinking that of course, we're going to have as many records as Twitter does. I mean eventually when this thing takes off. Then it lasted for three years, but anyways.

So the other thing that probably, I guess, maybe this is just me. But I know when I'm thinking about building a table, I'm always trying to think of, "Well, what if I need to change this thing?" Like, what if I have a new access pattern that I need to add, or I need to sort the data in a different way. And with a couple of hundred thousand records, or whatever, moving things around or copying the data over isn't that difficult. But if I was to get a terabyte of data or two terabytes of data on a table, and then all of a sudden somebody comes to me and says, "Hey, we need to add this new access pattern."

That is one of those things in the past where people have said, and even Rick had said this in one of his talks, where it's not a flexible data model. But I think that tune has changed quite a bit. I mean, Rick, and I actually talked about this a little bit where there is some flexibility now. But so migrations in general, whether it's migrating from an existing table or adding new access patterns, or even just migrating data from your existing workloads, you spend a lot of time in the book talking about this.

Alex: Yeah, absolutely. And this was actually a late addition to the book, but I just got so many questions about, I don't want to use Dynamo because what if my access pattern's change, or how do I migrate data? Things like that. So I actually went through it, I think it's not as bad as you think. And I split migrations into two categories, basically. First off, they're just additive migrations, where if you're just adding a new application attribute to existing items, you just change that in your application code, you don't need to change anything in DynamoDB. Or if you're adding a new type of entity that doesn't have any relational access patterns with an existing entity, or if you can put it into an existing item collection of an existing entity, you don't need to do anything. It's just a purely application code change there.

The second type of migration is, I need to do something to existing items, either because I'm changing an access pattern for an existing items or I'm joining two existing items that weren't joined together, or I'm adding a new entity type that I need a relation and there's no existing item collections to use there.

So now you don't only need to change your application code, but you need to do something with your existing data. And that's harder. It seems scary, but it's actually not that bad. And like, once you've gone through one of these processes, these ETL migration processes, they're pretty easy. It's basically a three step process. You're going to have some giant background job that's going to scan your table. You're going to look for the particular items you need to change. So if it's an order, and you need to add GSI2PK and GSI2SK for it, you find your orders in that scan for each order that you find, then you add these new attributes on them.

And then if there are more pages in your scan, you loop around and do that again. So it's just this giant wild loop that operates on your whole table. Depending on how big your table is, it might take a few hours, but it's pretty straightforward, that three step process: scan your table, identify the items you want to change and change them.

Jeremy: Yeah. I mean, and that's one of the things that I've come to realize too, and certainly, where you get a lot of flexibility in adding new indexes or even just making changes to the underlying data in the table, it goes back to the strategy of using separate attributes for the indexing and separate attributes for your application. Because when those two things don't mix, then it's so much easier to just go in and change the shape of certain bits of data.

Alex: Yep, absolutely. I agree completely.

Jeremy: Alright. So then another thing too, you mentioned, it could take a couple of hours. One thing that's great about something like Lambda, for example, is the fact that you can run a very high, you know, a high level of concurrency when you're doing these jobs, and DynamoDB actually supports parallel scans, which you talk about as a way to help with these migrations.

Alex: Yep, parallel scans are awesome. Basically, it's a way for you to chop up your scan into a bunch of different segments as you need to and operate on them in parallel, without you needing to do any state management around that and figuring out which ones you've already done or haven't done.

So when you're doing a scan operation. If you add in these two additional parameters, one is total segment which says, hey, how many different segments do I want to chop this up into? Do I want to have 10 different workers? Do I want to have 100 workers? Do I want 1000? But you're going to, for that total segments parameter, you're going to use the same number across all your different workers.

So let's say you have 10 segments, and then there's also a segment parameter, which identifies what segment that particular worker is. So if you had 10 different workers, each one of them would get a different value for that segment, which would be zero through nine, each one for one of those 10 segments. And then they can just operate in parallel, and they're just going at it. You can set this up as a step function with Lambda and just use that map to fan it out across a bunch and just let it roll and it works really well.

Jeremy: Yeah, no, that is something that using parallel scans just is as soon as you figure out how to do that, all of a sudden, like all these daunting ETL tasks are much, much more approachable, I guess, is a good way to say it.

Alex: True.

Jeremy: Alright, so then another thing that you mentioned in the book, and I thought this was good to mention, because I tend to see this in my own design sometimes where I have maybe like a user ID or a username. Maybe it's the email address, right?

And then I have another item that also has that email address, but it's for a different entity type, like maybe one's for the authentication and one's for their user record or something like that.

And when you want to make sure that you've got uniqueness, because sometimes if one of those records doesn't exist, you want to make sure you have uniqueness across multiple items, and you have some strategies for that.

Alex: Yeah, absolutely. So if you're doing any uniqueness in DynamoDB, you're going to need to build that into your primary key pattern. And then when you're writing that item, you'll write what's called a condition expression that just says, "Hey, make sure there's not an item that already exists with this primary key pattern." So you'll do that, and if you have a user, often that username or something is built into that primary key patterns where you can assert that it doesn't exist.

But, if you have multiple uniqueness requirements, like you're saying here and you have, you want to make sure your user doesn't have the same username, and also that a user hasn't signed up with this email address, you don't bake both of them into the primary key, because that's not going to work. That's actually just going to make sure there's not a unique combination of that item in your table. What you need to do is make two separate items. One, which is tracking the username, one, which is tracking that email address, wrap that in a transaction. And transactions are really cool, you can operate on multiple items in a single request. And if any of those operations fail, the entire operation is going to get rolled back and fail for you. So that's how you can handle both of those uniqueness conditions if you need to.

Jeremy: Awesome. Alright. I have one more question for you.

Alex: Yep.

Jeremy: Why on DynamoDB can't I just do select count(*) from something where x, y, z equals whatever. I know the answer to that. But that is a question I think that comes up where people want to be able to count items, I want to count the number of downloads. So I'll just scan the table, or I'll run a query for all my downloads to match a particular query condition. That's a terrible idea, right? I think we know that. So how do we do that? How do we get reference counts?

Alex: Yeah, it's a good question. I think going back to why can't I do that is DynamoDB is going to strictly enforce that you cannot write a bad query. So unless you really try hard, they're going to make it so you can't do stuff. And something like that aggregation is totally unbounded. What if you had three million related items in this particular thing, or 10 million or whatever? That count would have to read through 10 million items, they could each be 400 kilobytes each, and it would just be a mess on how long that would take.

DynamoDB is going to cut it off and make sure that you can't do crazy stuff like that. So instead, what you need to do is you're going to need to maintain reference counts yourself. And the common example I use here is in the book, I have GitHub repos and the number of people that have started, right? Where like, I think 150,000 people have started the React repo, or if you have a tweet, and people can like it or retweet it, and that could have hundreds of thousands or millions of likes.

And what you do is you just store a reference count on that parent itself. So again, you'll be using that transaction, like you just said, and doing two operations at once. Number one, you're inserting that item that tracks that operation that happens. So if someone starts a repo, you're inserting an item there to make sure they don't start multiple times, or if someone likes or retweets a tweet, you insert that item to make sure they don't do that multiple times.

But in that same transaction, you're also incrementing that count on that parent item. So on that repo, on that tweet, whatever it is, you're just incrementing that counter by one. And then when you fetch that repo, that tweet, you can show the number of people that have started, or liked it, or retweeted it or whatever. And that's how you handle reference counts rather than sort of doing this count star when you're grouping by something huge.

Jeremy: Yeah. And I also, I'm always a little leery of transactions and DynamoDB. I know they work really well but they're not your traditional transactions. I know they're a little more expensive than you would normally get. So even in situations like that, depending on how important the account is. I mean the account has to be 100% accurate then yeah, transactions or maybe DynamoDB streams, but sometimes just writing those in a batch - you know what I mean - will probably, unless there's some error or something that happens, but for the most part that should be a fairly safe operation.

Alex: Yeah. The one thing I say there like on the bat, it depends, like if you're worried about the hot path and slowing down that hot path, then like doing it in streams later on and batching it, that can be a good way to increment those counts. But if you're trying to use batch to save money on your rights, it's not really going to save you unless you have a small number of items to where in a particular stream batch you're going to be able to group some together, right?

But if you get 25 stream records and it's 25 different items you need to increment counts, you're still going to pay the same write cost there as if you did in a transaction.

Jeremy: Yeah, I was just thinking like a batch item write. You know what I mean? Doing that to do multiple writes or something like that. But anyways, so listen anything else we should know about this book?

Alex: I mean, I don't think so. I'd say you know, I think it's been well received. Rick Houlihan who got me started on this journey, hero of mine. He wrote the foreword, he endorsed it. I think a lot of folks on the DynamoDB team and within AWS they're really enjoying it.

I think a lot of people have enjoyed it. So, get it. There's a money back guarantee if you're unhappy, but I think you'll be happy with it. I just think there's nothing else out there like it. There's maybe 25% of it, you can get out there cobbling together, and it's easier that's in one place, but 50, 75% of it, I think, is really just brand new stuff that I'm pretty happy with.

Jeremy: Yeah, no, I totally agree. And you have all those extra examples. You have videos that go along with it depending on which level you buy, and things like that. So yeah, so definitely. So listen, Alex, thank you again for coming on. And I mean, really, thank you for writing the book because it's my new reference, my new go-to reference for DynamoDB.

But I'm always like, rather than searching on Google I just search the PDF and find it there. It's a little bit easier and I know that the information will be accurate and up to date. So hugely important piece of work that you've done and I think that the community is very much so appreciative of it. So if you appreciate Alex's work go to dynamodbbook.com and reward him by paying him for that work, because I think that is hugely important.

So other than that, if people want to get a hold of you what's the best way to do that?

Alex: Yep, you can hit me up on Twitter, I'm @alexbdebrie there. I also blog at alexdebrie.com, I have dynamodbguide.com, I have dynamodbbook.com. So you know if you google me or look for me I'm out there. I'm always around in Twitter DMs or email anything you want to hit me up. I'm happy to chat with folks.

Jeremy: Awesome. Alright. I will get all that into the show notes. Thanks again, Alex.

Alex: Thanks, Jeremy.

View Details

About Stephen Pinkerton:

Stephen Pinkerton is a Product Manager at Datadog. Stephen has also held roles in product strategy and software engineering, working with teams at Google's Nest Labs, Facebook, Cloudflare, Square, and Monzo to ship products, build distributed microservices, debug real-time embedded devices, develop features for modern frontend apps, and create data pipelines.

  • Twitter: @spnktn
  • Site: pinkerton.io
  • LinkedIn: https://www.linkedin.com/in/spnktn/

About Darcy Rayner:

Darcy Rayner is a Software Engineer at Datadog, and previously worked as Lead Software Engineer at at Two Bulls, a boutique software development firm, running the front-end chapter. Darcy’s projects boast clients like Disney, PBS, LIFX, Verizon and the Linux Foundation.

  • Site: www.darcyrayner.com
  • LinkedIn: https://www.linkedin.com/in/darcyrayner/

Datadog’s Research Report: The State of Serverless: https://www.datadoghq.com/state-of-serverless/

WATCH THIS EPISODE ON YOUTUBE: https://youtu.be/VC1CjUsKqBI

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Stephen Pinkerton and Darcy Rayner. Hi, Stephen and Darcy, thanks for joining me.

Stephen: Hey, how's it going? Thanks for having us on.

Darcy: Hi.

Jeremy: So Stephen, you are a product manager for Serverless at Datadog, so why don't you tell the listeners a little bit about what Datadog does and a little bit about your background.

Stephen: So Datadog is a company that lets you monitor all of your servers, all of your application performance, if your website's up or down, all your application logs in one place and it joins all these disparate data sources together, so that whenever you need to explain something or debug something, you can join all of this different data and really figure out root cause quickly. So I've been working here about a year, focusing on our serverless integration, so that's helping our customers running products like AWS Lambda, to be successful in deploying new services built on top of serverless and then debug issues with their applications.

Jeremy: Awesome. Darcy, you are a senior software engineer for the Serverless Team at Datadog, so why don't you tell us about your background and what your role is at Datadog.

Darcy: Sure. So I've been at Datadog about a year in the Serverless Team, before I joined Datadog, I was working at an agency. So we were massive serverless adopters in everything we did. Everything was about getting stuff off the ground running very quickly, and low cost to our customers. But while I was there, I realized that there was still a bit of a gap in terms of the monitoring story. So I joined Datadog about a year ago to work on some of the integrations that we're building here with services like Lambda or Azure Functions or GCP Functions.

Jeremy: Awesome. So a couple of weeks ago, Datadog came out with this very, very cool report called the State of Serverless. You basically looked at a bunch of your clients, went through and figured out how they were using serverless, broke it all down. This is really great, can you, maybe Stephen, can you give me some background on what was the reason for running this or for putting this report together?

Stephen: So we're in a unique position where customers of all different sizes, with all these different use cases are sending all their data to us, and we frequently get questions when we're on the phone with them of, "How do I run serverless successfully?" So this is from customers who are moving workloads into serverless, or they're 100% serverless and they're asking us, "Which metrics do I pay attention to, or how do I get data out of serverless, or how do I run this in a cost efficient way?" So we frequently get these questions, and the report was a way for us to look at data across all of our customers, across all these different data sources that we have and say, "Here's exactly how people are running on serverless." Which technologies are they using, what are they monitoring with it? So it was a really interesting opportunity for our customers to see how other people are using serverless really in a data driven way.

Jeremy: It's interesting, because you do mention in the beginning of the report that you're saying "Serverless," but you are just focusing on FaaS. So actually, I'd love to get your thoughts on this, Darcy. Just considering what serverless is as a whole, what do you consider that to be, because it's more than FaaS.

Darcy: In terms of the things we're looking at specifically in this report, it's very focused on Lambda. I think in general, we have people who come to us asking us for solutions for things like ECS, Fargate, things like Knative or Google Cloud Run, which they're not necessarily following the pair invocation model in terms of pricing and cost structure. And they're not necessarily building containerized services directly around functions, but they're adjacent. They have some of the pieces. The ability to very quickly on-demand spin up resources and the ability to have event driven architectures, like this is something I think we see across different solutions.

Jeremy: Nice. That makes sense. I want to get into the study itself, but I do think it's important, because I know someone who's tried to run a survey in the past, that the methodology is important. We want to know who the people are that are answering these, which way they might skew based on their population and so forth. So the report actually did a great job outlining this, but just because I'd like to go through these findings, it would be great if we could just talk about that methodology for a second. Let's start with the population, so this was just Datadog customers?

Stephen: Yeah. The claims that we make in the data that we looked at is across all of these Datadog customers. We don't have data on someone who's not a Datadog customer, so for all of this, we looked at our customers' metrics, their trace data.

Jeremy: It just seems, and your customers are obviously more cloud savvy. So we're not looking at all enterprises here, just the ones that are probably much more cloud savvy, using Datadog.

Stephen: Yeah, that's correct.

Jeremy: Great. Then we also talk about Lambda adoption in here, and that's one of those tough things too where, what does it mean to adopt Lambda? Can you explain what that means?

Darcy: So Lambda adoption, we consider it to be any account in AWS, which is running more than five Lambda functions a month. That was the cutoff point where it's like maybe they have one or two people experiment around with it, after five, we considered it being regularly run. We considered that to be a company that's adopting Lambda.

Jeremy: And that ties into the AWS usage, right? In order for a company to be using AWS, they would have to be running some workloads in that cloud?

Darcy: Yeah. Our broader definition of AWS usage included both anyone who's currently using Lambda, but also we looked at any organization that had more than five EC2 instances running in a given month.

Jeremy: So this is still covering fairly small customers too. Five EC2 instances is relatively low, so this gives a nice broad perspective of these small customers plus large customers. Then that brings us to scale of the environments, so how did you estimate the scale of the environments?

Stephen: When we talk about customers being small, medium or large, we look at the scale of the other infrastructure that they're running. So we might be looking at companies that just have five Lambda functions or five EC2 hosts, but the way that we talk about them in the report is based on how much other infrastructure do they have. So they might have a small Lambda footprint or they might have a small EC2 footprint, but they might have a very large footprint using containers, ECS, Fargate, et cetera.

Jeremy: Awesome. Let's jump into this, so the first finding in this report was that half of AWS users have adopted Lambda. So that means that of all the AWS customers you have, 50% or more, I think it's 53% or something like that, that they are using more than five Lambda functions.

Stephen: Yeah. I think this one was very surprising in that organizations are aware of Lambda, and more importantly, they're using it. So from a business perspective, we found this really surprising, because these leaders in companies are bringing in Lambda because it lets their teams move a lot faster, shipping products a lot faster. It also means from a development perspective, there's not a team you need to go to talk about who's going to monitor your code, or you don't need a request new servers, new hosts from your procurement group within your company. Lambda is just a really easy way to get started, so I think that's why we see it used across so many organizations.

Darcy: I think there is a lot less red tape in getting a Lambda function approved than say speeding up an EC2 instance. So even large traditional organizations, like banks or like financial institutions, there is some adoption that's happening there, just because it's a lot easier to get products off the ground running.

Jeremy: One of the points in the data is it shows that it's up from about 21% in 2018, so that's more than doubled in two years. Is this something that, the trendline looks like it's continuing to grow, so is this something where we think the vast, vast majority of customers are going to be using Lambda say by 2022?

Stephen: I think for a number of workloads and use cases, absolutely we're going to see this used everywhere. I think not just Lambda as well, just any type of these serverless products. The report is very focused on Lambda, we do get into what products are people using with Lambda, but people are starting to realize the value of a serverless database and serverless message queues, and the whole ecosystem is definitely here to stay, but the way that you're running your code might be changing.

Jeremy: Absolutely. Darcy, what types of use cases are you seeing with these Lambda functions?

Darcy: I think there's a pretty big variety. A lot of the companies dipping their toes I Lambda, they start very small. So it could be like IT is running some batch jobs every now and again on a Lambda, it's the perfect use case. To startups that are maybe migrating existing Django apps into a single function, or migrating a monolith, to startups that are entirely driven from serverless and everything's a function. So there's a massive variety and spectrum, and I think some of the details we have in the report show that. Large enterprises with large legacy systems are adopting it in slightly different ways to newer companies.

Jeremy: So let's talk about that, so that's the next finding here is that Lambda is more prevalent in large environments. So you kind of got at why you think that is, but is that just because of cloud sophistication you think?

Darcy: I definitely think it's a major factor. When you have a large organization with several teams, there is this broader movement towards microservices and having team ownership boundaries of services and giving engineers more autonomy. I think with that, you just have an increased likelihood of one, two, three, four teams adopting Lambda, and then that being the gateway in a large organization. Whereas, if you're a smaller company, maybe you're not operating at the same scale yet, you can end up with more unified technology, which means that there's less chance that Lambda will be adopted, even though more of these startups and new companies are adopting Lambda.

Jeremy: I wonder, Stephen, for you, do you think that education or the learning curve to get people started with serverless, this could be another thing? You've got a small organization, they can't just go off and do all these skunkworks projects like Darcy said. Do you think education in serverless or the learning curve is holding some of these smaller companies back?

Stephen: That's a good question. I think there's a lot of misconceptions around serverless that might hurt its adoption in some organizations, but it definitely requires a different way of thinking that not ever organization might be ready for. But I think what we saw, the trend several years ago was that it was engineers driving serverless, people are building their own projects off of it and seeing how cool it is and how fast it lets them move, and they're bringing it into their organizations. So I think education is a big part of it, both realizing that it can solve my use case, and I might be able to solve that problem faster than I could otherwise. I don't need to go spin up a bunch of docker containers and manage scaling and everything for different services that I want to run.

Jeremy: So speaking of containers, this is one of the findings that I absolutely love, the fact that container users have flocked to Lambda. Now clearly, they're not abandoning containers altogether, not that there's anything wrong with containers, we love containers, but the idea of being able to use Lambda functions to do some of the workloads and just make it so much easier, not even to worry about orchestration or any of that stuff. But it sounds like, based on these findings, 80% of users that are using containers, AWS users, have at least five Lambda functions running.

Stephen: Yeah. This was another really surprising finding, and like you talked about with cloud sophistication and what Darcy talks about with the popularization of microservices, that it's less important where you're actually running your code. So if you're already running in a microservice architecture, then it's very easy to adopt serverless and see maybe this is a service where we want more elastic scaling or our workloads are a little bit less predictable and were spiky, and that's a great place to run serverless functions. So we see people running these together sometimes to get the benefits of both, or just because different teams are using different technologies for the problems that they're solving.

Jeremy: I wonder, Darcy, I don't know if you know the data on this, but do you see a reduction in use of containers while people are migrating to Lambda functions? Or is it just a new subsegment of their architecture that's growing?

Darcy: I don't think we've seen any hard data to indicate that, I think it's probably more likely as organizations grow container adoption, we did have our containers report come out I think a couple months ago as well, so there's probably better data there, but that's also been growing. There's so many opportunities for organizations to migrate from maybe a traditional internally hosted infrastructure or to even older cloud infrastructures to containers and serverless, that they're both still growing.

Jeremy: I think the lift and shift is much easier when you're porting to containers, a lot of your application code doesn't have to change as much. Whereas with Lambda, you are completely re-engineering and refactoring how, not just the code works, but your whole entire architecture as well. I think that gets a little bit complex. So the next finding here is that Amazon SQS and DynamoDB pair really well with Lambda, so I think that is kind of obvious. You would think people who build serverless applications, you want to use tools that are serverless themselves or at least play really, really well with serverless tools. But that was interesting, because it seems like the pay-per-use stuff is very popular with Lambdas, and not so much with SQL databases, although you still see some people trying to hit MySQL with them as well.

Darcy: Yeah. I think this is something that drew a little bit more to the legacy of Lambda. We've seen newer services like some of the Managed RDS stuff coming from Amazon, which is really promising, but there were historical reasons or historical technical reasons why using DynamoDB might have been the easier less resistance solution, versus using RDS. I think that's changing with a lot of the services that AWS has been introducing lately. Definitely adding the Managed Aurora stuff is definitely a big step up from that. So I think we might see that change in the future, there's certainly a lot of advantages to SQL or relational databases that are more eventually consistent databases like DynamoDB. Looking at SQS versus something like Kinesis streams, again, it's just a lot easier to set up and adopt. Engineers tend to prefer simpler integrations and the integrations that AWS has with SQS are a lot quicker to get up and running, versus something like Kinesis, which is more of a workforce solution.

Jeremy: I didn't see Managed Kafka showing up as one of the queues that people were using with Lambda.

Darcy: Definitely there are a lot of orgs using Managed Kafka, but there is another level of overhead. Where if you are looking for a really simple managed solution, it's not necessarily the best way to go.

Jeremy: You mentioned about the services that AWS is offering with Managed Aurora or Aurora Serverless and things like that, and I do think that, I wish there weren't as many of those solutions. I wish people were forced more to go with the more serverless type things. Because Serverless Aurora or Aurora Serverless isn't quite serverless in the sense. There is some scaling that's still required there, and you still have a connection management and all those other things. I know you've got the Data API and some of those things that can really help with it, but I think overall trying to push people towards things like DynamoDB for operational stuff ... Now granted, I get it, you're going to do your transactional stuff or you've got to do your analytics reporting and things like that. But I'm curious in terms of whether or not, because it looks like the number of data storers that are mixing with that SQL, generically SQL I guess is fairly low, with DynamoDB being much higher. Are you seeing companies moving away from SQL and to DynamoDB, or is it just one of those things where it's like whatever people are comfortable with, that's what they're going with?

Darcy: I definitely think there is a comfort level. People tend to ... Different databases for different solutions. Dynamo is not necessarily great for things like ad hoc queries or things with complex table joints, so there is a level of technical trade off that engineers make. I don't think you're ever going to see a case of DynamoDB entirely replacing relational databases. I think relational databases have a ton of uses. If anything, I think we'll just see the story of using relational databases and having managed relational databases become easier and easier, and more of the scaling overhead being taken away from engineers.

Jeremy: The other thing I'm curious about too is, especially seeing that SQS and Kinesis and SNS are so popular as Lambda triggering those that are then triggering the data sources. It seems like a lot of your customers are starting to embrace that idea of asynchronous thinking. Is that something you feel you're seeing as well?

Darcy: Yeah, definitely. I think there is this event driven microservices revolution that's happening. It does take a long time for large organizations to really buy into that idea. If you've been running a monolith successfully with 50 to 100 engineers for the last five years, then it's harder to have that organizational buy in. It takes time. But I think that is a growing trend, and I think the different architecture patterns that people are employing around distributed queue-ing and event driven and relying on very redundant and ephemeral instances for everything. Whether that's in containerization or in Lambda, I think that's growing as one of the most popular web architectures.

Jeremy: Definitely. So the next one, which I thought this was, I think it makes sense for Lambda functions to be written in Node and Python, because the cold start is a little bit lower and so forth. But it's interesting, because with the number of large clients that seem to be adopting it faster, which I would expect to be developing apps in Java or something like that, as a more popular language, Python and Node.js in a landslide are the most popular.

Darcy: Yeah. I think it makes a lot of sense to me. Things with compiled languages, like Java is, even though it's running in a VM, it is compiled, so things like Go and Java, there is more overhead in setting that up and deploying that. You can't really use, and I'm sure a lot of Lambda usage, or a surprising amount of Lambda usage isn't people with big infrastructures, code deployments, it's somebody copy and pasting a Lambda function directly into the AWS console. I think a lot of the adoption probably just happens from that, it's a lot easier to copy and paste code, adjust it in the console. Then for people who are using, a lot of the orgs using actual infrastructures code and deploying in more like a rigorous way, it's still a lot easier to control and deploy Lambdas for those languages.

Jeremy: I think too, those languages are, because they're not compiled, because you can launch them so quickly, and I think the cold start time was another thing that was hugely popular. But I don't know, Stephen, maybe you can shed some light on this too. Just thinking about which engineers might be using this. Is it that your main, your hardcore app engineers who are coding the main infrastructure or whatever, the main app are doing it in Java or something like that? Whereas some of these other side use cases, maybe ETL tasks, maybe some data transformations, maybe some DevOps things are just quick and dirty Lambda functions that are written in these languages by DevOps engineers?

Stephen: Yeah, absolutely. So we really see the use cases all over the place here, but for actual run times that we see used a lot, for a background job it's very common that we'll see it written in Python and Node. But I think what's been really surprising is that even at these really large organizations who are adopting Lambda that we talked about before, they're still using Node and Python a lot, which I think would surprise some folks that they're really at the cutting edge of both how they're running their code and the way that they're actually writing it. We do see some more Java use on the enterprise side, of people moving these Java microservices into Lambda functions. I think that fits their use case pretty well. But in general, we do just see the most in Python and Node everywhere.

Darcy: I think the other maybe interesting thing is you would think Ruby would fit into that paradigm of dynamic language, but we're not seeing a lot of Ruby adoption in Lambda, and we're not really sure what the reasons why. I think that probably would be the only outlier, whereas Ruby was traditionally one of the more popular languages that was being monitored by Datadog as a whole.

Stephen: It could even be skewed by people have these Rails monoliths and if they lift and shift that into Lambda, they just have less Lambda functions.

Jeremy: That's also true. But that actually could be, that leads in I think to the next finding, which could also be based on the use cases. Is the fact that the median Lambda function runs for 800 milliseconds. Half of them run for less than 800 milliseconds, which could be front end, maybe synchronous calls from an API Gateway or something like that. But those longer running ones, that sounds like there are other tasks that those are performing.

Darcy: Yeah. Some people have probably decided to use Lambda in a way that is actually like running more computationally heavy workloads. Things like web scrapers or like jobs that are running continuously over a period of time. We anticipate some of that is coming down to, I don't want to say misusage, but trying to convert a workload, like a square workload into a circular hole.

Jeremy: You can say misusage, because I think that's right.

Darcy: It might be the case that some workloads aren't necessarily entirely appropriate for services like Lambda, and containerization and something like ECS or Fargate would be more appropriate. But there are some cases where people are just making it work I think. For the lower use cases, I think the majority of events that we see are probably like HTTP events, so low latency is really, really important for an API, which is probably why we see such a heavy skew towards fast implications.

Jeremy: The other part of that too was that basically it says one-fifth of Lambda functions run for 100 milliseconds or less, which is interesting, because of course that is the billing threshold, the unit of billing for AWS. I know there have been some calls, including from myself, to get that granularity down maybe 50 milliseconds as opposed to 100 milliseconds. But that's interesting, because again, you can't run much in 100 milliseconds, unless it's something like responding to an API Gateway for example.

Darcy: Yeah. I think with a lot of these asynchronous patterns that exist in Lambda, you can push something to be written off to a database, a more eventually consistent way in that amount of time.

Jeremy: That's true.

Darcy: So there are patterns that exist that really reduce the amount of latency that can go onto a web request and give you a very fast average response time. A lot of Lambdas are doing just processing work, they're transforming output from a queue, and then that's going off to the next queue. So there are some mixed use cases I think.

Jeremy: Yeah, definitely. The next one is half of Lambda functions have the minimum memory allocation, and speaking of misuse, this is probably one of those.

Darcy: I think the memory allocation, it probably comes more down to miseducation, although if your Lambda's executing with 100 milliseconds, maybe it makes sense to leave it. I think people don't spend a lot of time thinking about how to optimize these services, we've removed so much of the thinking about overhead and infrastructure that developers are just putting things in Lambda and not even spending the time to tweak it and reduce latency, and potentially cut costs.

Jeremy: Stephen, do you think that is just something where we need to develop better best practices around?

Stephen: Absolutely. This is a question that we hear a lot from folks about, "How do I optimize my Lambda workloads?" I think just because Lambda and serverless has let us write code, upload it to a cloud provider and let someone else run it, there's still knobs that you can turn to adjust performance. I think there's a lot of miseducation or complete lack of education around what do these knobs mean. So we see this as well when customers run into issues with concurrency limits, memory's another big one, and one of the latest ones has been provision concurrency, that's both how do these work and how do I use it affectively for my application and to be most efficient with cost.

Jeremy: So another best practice maybe is setting good timeouts, because you certainly don't want ... With Lambda functions, you're paying while they're processing. So if you have something that hangs for some reason, if you expect it to end or the processing should end within 10 seconds, and you set it for 10 minutes, and it just keeps on running and running, then obviously you're paying for something you don't need to. So the report points out that two-thirds of defined timeouts are under a minute, which is probably good, there's probably more granularity in there. But I think the scary thing was a bunch of them were set for 15 minutes, the maximum.

Darcy: I can see why some developers would choose to do that. Their reasoning is maybe this is an intense job and we can't afford for it to fail. But the truth is that any piece of infrastructure you build, you should have the ability to recover and retry for the sake of scalability. I think it's maybe being used as a crutch by some developers or engineers to try and guarantee something's going to finish, without spending the time to think about how to engineer their workloads to run quickly.

Jeremy: I tend to lean more towards the crutch side of things, because I think that it's just like, "I can let it run for 15 minutes, I'll let it run for 15 minutes." But have you or has your team, and Stephen maybe because you're closer to the customers, you could see this, but have people had concerns over this idea of the denial of wallet, this idea of flooding Lambda or API Gateway so that your Lambda functions are just running, is that something that you hear about a lot?

Stephen: We hear some security concerns occasionally similar to this, and I think they're split between timeouts and concurrency limits, where they're both things that can starve resources in your account. So if you have a high timeout, and you get flooded with requests, you're going to have a Lambda function executing for a long time. It's just going to cost you money unnecessarily. On the concurrency limit side, we see this as an issue where if you have unbounded concurrency and your function gets too many indications, then it can actually starve the concurrency in that region of your AWS account, and actually cause your other functions to not run. So I think that's a common gotcha that we see, so we typically recommend that customers set timeout limits and when appropriate, concurrency limits just to prevent these kinds of attacks.

Jeremy: Actually, the concurrency limit was the last point here, or the last finding that only 4% of functions have a defined concurrency limit. I think for a lot of internal communications, maybe if you're using it for DevOps or different use cases like that then having a concurrency on it probably isn't that big of a deal. But certainly some of these ones that are forward facing or processing off of queues, or anything where you could just flood these things, that seems a little bit scary to me. Although, it does say that 88.6% of organizations have at least one function with the concurrency limit defined.

Darcy: Not 100% sure why people aren't setting them. I think there is this promise of serverless, of it just scales, and you have a sudden demand or peak in usage, it just scales. This might be coming down to engineering teams not asking themselves questions, or not really planning the maximum capacity to their systems. Lambdas that are just running Chrome jobs in the background, they don't necessarily need that kind of thought put into them. I think organizations that do use concurrency limits generally are thinking about scaling, they're thinking about scaling in a much more granular way. Again, that account wide concurrency limit across every function is something that's very easy to get bitten by early on when you're dipping your toes in Lambda, and getting a bit of usage.

Jeremy: Then maybe this is an education thing too, because I've been saying this for quite some time where your average developer usually didn't know anything about the infrastructure. They're writing code and you have an ops team or you have some, you eventually get to the DevOps culture where working together to make sure that the code you wrote had the right infrastructure behind it. But when it comes to things like Lambda functions, there's just a lot of questions that you probably never had to ask yourself as a developer before. As you said, as teams become more agile and they're able to just publish directly to a Lambda function, as opposed to have somebody set up a Kubernetes Cluster for them with all kinds of defined rules. I guess your average developer's probably not thinking about this.

Darcy: Yeah, absolutely. We've built a really, really powerful abstraction, but it is an abstraction and there are leaks. Concurrency is absolutely one of those things, how the process of Lambda works leaks a bit, how the language and the runtime works within that context leaks a bit. If you just take the mentality of, "It will just run my code as much as I need it to," then you're probably not thinking about it quite enough.

Jeremy: Because the other thing too is, as you mentioned, where this promise of serverless is we'll just scale, we'll just keep scaling and scaling, and nothing is infinitely scalable. Everything has limits at some point. Obviously, the per region concurrency limit is an artificial limit that AWS puts in, you can increase that. So if you have 10,000 concurrent requests, you can have them increase that for you. But I think your average person probably doesn't think about that, at some point is too much scale not good? Do we want to limit scale because of downstream systems, because of billing concerns, because of all of these other things.

Darcy: Yeah. When we talk about event driven architectures, it's the same as any infrastructure, there's bottlenecks in the pipes and the connections between different services. It's very easy to create downstream pressure. I think the more people buy into the serverless promise, which is every piece of your infrastructure is meant to be something that can scale automatically for you, like Dynamo, like SQS, the more easy it is to miss the finer print on how your system could fail. It's why monitoring is still extremely important in these environments, you can't really skip that, because you could find yourself in a situation when your services start failing, even though you've got all the settings to say, "Scale to a million. Scale to a billion," or whatever. One small thing might not be able to keep a promise, and suddenly you have a failure somewhere in your system.

Jeremy: Definitely. Distributed systems in and of themselves are very difficult, and now when you start talking about all of these little components, all of these little building blocks with serverless, and you've got Lambda functions and queues and databases and streams and all that stuff that's communicating with one another, being able to understand where those failures are is a hugely important thing.

Darcy: Yeah, absolutely.

Jeremy: Awesome. So those were the findings in this report, and if you haven't seen this report yet, DatadogHQ.com/state-of-serverless. I'll put it in the show notes as well. But while I've got you two here, you work with a lot of enterprises, it's great to get some insight into what other people are doing. You're an enterprise yourself, or you work for an enterprise yourself, so I'd love to start and maybe just get a little bit of context from you. How did Datadog get started with serverless? Actually, how serverless are you actually? That would probably be a good question.

Darcy: That is a good question. We're about a 10 year old company I think, it's either 2009 or 2010, that we started. So we were pretty early as part of this whole DevOps revolution, I think that was part of our mission statement was to be a part of that movement. But our infrastructure was built on the EC2 instances, it was glued together with all the things you would do in 2010 cloud architecture. We've scaled, so we brought in a bunch of different teams and products and different services, acquisitions from different companies that have been brought into the fold. So like happens, we've had internally a lot of adoption to move to a more microservice based approach, and having teams own individual services which they maintain and deploy and monitor. So that's been a big trend internally. We have a lot of Kubernetes adoption and we have serverless adoption as well for a lot of our teams. On the Serverless Team, surprisingly our compute loads don't run on serverless, because we use the same infrastructure as a lot of Datadog has historically. But when we went out and started talking to people in the company and finding places that serverless was being adopted, we found it was actually adopted pretty widely in the company in very different and diverse use cases.

Darcy: We have a lot of internal tools and products built on top of Lambda, a lot of tooling around. CICD is using Lambda, Slackbots, in some production workloads, Lambda is a piece of that as well. It was really surprising just to find the different sort of use cases that existed within our company. We have a commitment funnily enough to be a vendor agnostic infrastructure, so a lot of what we do and build is run on Kubernetes. But a lot of the glue pieces that we have to run our business are running on services like Lambda.

Jeremy: Awesome. Stephen, is that something too, again, talking to your customers, and obviously you've seen this serverless adoption with the report here, but it seems like no matter what you're using out there, serverless is probably a part of it somewhere.

Stephen: Yeah, absolutely. That's something that we took away from the report as well, just the customers that we talk to every day, that serverless is all around us and that things that we use, applications, websites that we interact with every day are using serverless somehow. It's not even just in internal tools, it's your banking application, the way you buy movie tickets, the way you sell something on the internet. In some way, all these different companies are using serverless, are using Lambda workloads or they're using some sort of serverless technology, either with serverless containers, databases, on other cloud providers. It's really everywhere, and I think the adoption is so much higher than any of us expected.

Jeremy: That's awesome. Another thing that always seems to be a popular topic of conversation, especially with companies like yours that have to work with customers that are probably faced with this internal battle, is this idea of multi-cloud. Again, there's a million different definitions of it, multi-cloud versus being cloud agnostic for example are probably two different things. But just I'd love to get your take, Darcy, on this multi-cloud thing, and what you're seeing from your customers and what you need to support I guess, from a serverless standpoint.

Darcy: We definitely see a lot of very large enterprises that are multi-cloud, and sometimes it just comes down to which team within that company is building something. You have large multinational corporations that tend to go, or be more likely to go multi-cloud, because they might have been acquiring a company here, or bringing a different provider in here. I think I haven't seen personally many examples of companies choosing multi-cloud as their primary architecture. I think you do have people mixing and matching software service solutions, maybe outside of their primary cloud platform. It might be pulling an Auth0 or another managed service on top of that, but I haven't seen too many examples of companies primarily splitting their infrastructure service by service on different clouds.

I do think you do see bridging technologies, so there are some of these providers that do multi-cloud, but in an abstracted way. So you have code that you want invoked, and maybe they'll find a way to bundle up the Lambda or Azure Functions and GCP Functions, and distributed across different cloud providers that way. We do see maybe mixed adoption between something like Cloudflare workers for web traffic specific flows versus running the majority of your infrastructure in AWS. That's pretty common. Then there's really exciting things with Knative and Google Cloud Run for instance, where trying to think about how to build serverless applications in a platform agnostically, which I think is going to be very cool in the future.

Jeremy: Again, platform agnostic would maybe be a great dream if all of these platforms weren't doing things differently, because to be serverless on AWS probably means something, or being serverless on AWS means something completely different than being serverless on Azure for example. There's a lot of overlap functions and service, things like that, but just the different services available, it seems like there's so much different that at least for quite some time, until there's some standardization, that it just doesn't make sense that cloud agnostic is going to be a thing.

Darcy: Yeah. I think with something like Cloud Run, it wouldn't tick all the boxes for how you define serverless with ... Like Knative probably wouldn't tick all the boxes for how you define serverless compared to Lambda. So if you're running a Knative workload, that's not ticking the box of paying the invocation model. You're managing a cluster still, you are building functions, and you are doing event driven workloads. Versus something like Google Cloud Run, where you do have more of the invocation model. So I think there is a definitely mix and match, companies tend to buy really heavily into one platform's set of conventions and services, and I think unless you have a high priority on having multi-cloud availability, that's generally the way companies that I've seen would choose to go.

Jeremy: Stephen, I'm curious too again, being close to the customers, in terms of how people are approaching to build serverless applications, it's more than just choosing technologies obviously, it's avoiding choosing the lowest common denominator, trying to choose services that are scalable, that are easy to plug in. But are you seeing the serverless mindset or the serverless first shift?

Stephen: Absolutely. Obviously, that's biased by the customers who I talk to, but we see these customers who come in and are adopting serverless for all new products that they build, or they're moving existing infrastructure over to it, or they've even decided when they started as a company that they were going to be 100% serverless. So we're seeing this mindset really increase more, and we're seeing it mean more than just functions as a service, it's really how are we storing data. So for example, we see a lot of these customers using technologies like AppSync, so they're really at the cutting edge of using GraphQL, data stores, these event driven architectures. We typically see all of this together.

Jeremy: Awesome. I want to ask you one more question, because I'd love to get your thoughts on this and what you're seeing from adoption from your customers. I know it wasn't in the report, but something, people always ask, "What's next? What's after serverless, what comes next?" And I tend to believe, and I think other people agree with this, that it's edge computing. That's the next place where you're going to see a shift. So this idea of not having to manage infrastructure is great, it's going to be even better when your execution environment is no further than two miles away from the customer that's trying to access it, because it's just in every pot throughout the world. What are your thoughts? You did mention Cloudflare workers, Darcy, but what about Lambda@Edge and things like that? Are you seeing that adoption?

Darcy: Yeah, I think Lambda@Edge is becoming pretty popular. It's primarily just used for use cases around, most of the time serving web traffic and things like adjusting HTTP headers, or sometimes if you're serving a static site, you might stick a CloudFront plus Lambda@Edge in front of that. And Lambda@Edge is a great way to do things like AB testing or create, customize a static site in a way for your users. I think there will be a move past this concept of regions, because something like Lambda@Edge doesn't really fit into the idea of regions in AWS very neatly. So I think the idea of infrastructure and actual heavy workloads that are maybe still talking to your databases and still distributed geographically is going to be more common. At the moment, it's still tricky to do a multi-region infrastructure, you're still going to be limited. You're still going to be making geographical trade offs, like how you reach your customers and how you store data in a way that there's no proximity to reduce that latency. So I think there is an evolution to be had there, but we're only at the tip of that I think.

Jeremy: I think the data piece of it is probably going to be the hardest. With Cloudflare workers, they've got the global KV store, which is great, but still, what data do you need to replicate to certain regions. I think making the serverless jump is hard enough for a lot of developers, nevermind ... Not only are you running concurrent functions, you're now running concurrent functions at 180 points of presence across the world or something like that, and you have to manage the data separately, and your app has to be smart enough to do it, or the system has to be smart enough to do it. I think that's crazy.

Darcy: Developers are struggling to join data across multiple tables, how do they do that across multiple regions?

Jeremy: That's right. I can see that getting very complex. Listen, thank you both for being here and going through this report. That was awesome. A lot of great information I think, and like I said, if people want to go and check out the report itself or check out Datadog, they can do that at DatadogHQ.com. So why don't we just, if people do want to contact either of you, I know, Stephen, you're on Twitter?

Stephen: Yeah, I'm @SPNKTN on Twitter.

Jeremy: And then the Datadog blog is just DatadogHQ.com/blog. I think Darcy, you're in the shadows on Twitter, you don't do much of that?

Darcy: No, I'm not a massive Twitter user.

Jeremy: You're too busy building stuff to support serverless, so that's great. So again, thank you both, I will get all this information into the show notes, and it was great talking to you.

Darcy: Yeah, great talking to you.

Stephen: Yeah, thank you for having us.

THIS EPISODE IS SPONSORED BY: Stackery & AWS (Amazon EventBridge Learning Path)

View Details

About Susanne Kaiser:

Susanne Kaiser is an independent Tech Consultant from Hamburg, Germany, and was previously working as a startup CTO transforming their SaaS solution from monolith to microservices. She has a background in computer sciences and experience in software development & architecture for more than 15 years and regularly presents at international tech conferences.

  • Twitter: @suksr
  • Site: www.susannekaiser.net
  • LinkedIn: https://www.linkedin.com/in/susannekaiser1/

WATCH THIS EPISODE ON YOUTUBE: https://www.youtube.com/watch?v=eGYlTfBJBJQ

Transcript:
Jeremy: Hi everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Susanne Kaiser. Hi Susanne, thanks for joining me.

Susanne: Hi. Thanks for having me.

Jeremy: So, you are an independent tech consultant, so why don't you tell the listeners a little bit about your background and what you've been up to lately.

Susanne: Mm-hmm. So, as an Independent Consultant, I am helping organizations within the broad spectrum of software architecture and design including development to software delivery, or in other words or in shorter terms, helping organizations in building and shipping their digital products. I was also previously working as a startup CTO and I have a background in software development and software architecture of more than 17 years and I also regularly present at international tech conferences as a speaker.

Jeremy: Well, speaking of international tech conferences, I saw one of your talks at ServerlessDays Belfast before we shut down all conferences so no one can get together in person anymore. And, you did a talk about domain driven design and how that applies to serverless and serverless microservices. So, I'd love to talk to you about that today, because I think that is one of those things where software developers ... I don't want to say all software developers ... but a lot of software developers are just really bad at software design, and not just design of architecture and things like that which are in our wheelhouse, but more from the business side of things and understanding what the business needs are, understanding what the technical needs are, and then where that comes together in the middle, and what people should actually be building to solve those customer needs. So, I think that is domain driven design in a nutshell, but maybe you could tell the listeners a little bit about what domain driven design actually is.

Susanne: Yeah, so domain driven design is a software philosophy or methodology created by Eric Evans and it's about to capture the business domain as closely as possible into your software and it comes with a lot of strategic and technical design patterns and practices that I am happy to share with you in a moment. But, it's also, I would like to mention also, from the very beginning, it's not applicable everywhere. So, you should focus on your core domain ... I will explain it in a minute, hopefully, too ... where it makes sense that focusing on complex business logic and have to do with solving problems that have complex business logic behind.

Jeremy: Yeah, and so you had mentioned in your talk, the cost of poor software quality, and the number you gave here was, I want to say it's $2,840,000,000,000 a year in poor software quality, so what are some of these indicators of poor software quality?

Susanne: So, there are no simple measures for bad or good software quality, but there are several metrics that can be used as indicators. For example, an increasing curve of defect trend over the time is an indicator of poor quality software or low test coverage, assuming that as there is good test quality or cyclomatic complexity or large dev of inheritance and high degree of class coupling could also be indicators. Also the amount of effort it takes to understand a piece of code or badly engineered software resulting, for example, from immature or undisciplined practices and using less qualified software engineers or also and one thing is also really important to mention is the lack of domain knowledge and also poor communication and coordination issues and teams, specifically if one of the teams are growing.

Jeremy: Yeah, so that lack of domain knowledge, I would think that sort of gets to the crux of it, right? Because like I said in the beginning, we think we know how to solve a problem and we know how to solve it technically but what we're really trying to do is solve a problem that's very specific to a group of customers, whatever that group of customers might be and those different models could be your inventory system, right, and your inventory teams think of ... They think of inventory in a certain way and then you need software engineers to be able to build a system that makes sense for them -- that uses the same language, that uses the same sort of communication patterns or styles or things like that. What goes into building good software, then? Like what are the main components of building good software?

Susanne: So your domain driven design comes with a core statement that in order to build better software we have to align its software design with the business domain, with the business needs, and the business strategy. So domain driven design helps you with aligning your software design with the business domain needs and the strategy and it's very crucial for building your software solution, because otherwise, you are building something that, for example, are matching the requirements of your users. Instead you have to collaborate intensively with your domain experts to gain domain knowledge and to understand the problem first before you're solving it. We are tending to jump directly into solving a problem technically and, yeah, yeah, we can just let's deploy it on a cubinated cluster, but we have not understood the problem first. That's really crucial to the build better software.

Jeremy: Well, you mentioned the word strategy, which is one of those things where, I don't think a lot of people know what that exactly means. Like what is the strategy and you had talked a lot about Wardley Maps and sort of understanding this landscape and being able to use that to sort of plan your strategy and we can get into some of that more but can you explain Wardley Maps. We've talked about it before but just sort of in the context of domain driven design, how does that help you?

Susanne: Yeah, I really like to combine domain driven design with Wardley Maps, because at first, like when you start the journey to a domain driven design, it's really overwhelming. It was for me, very overwhelming, because there's a lot of new terms and it requires some time to grasp and understand it, and since it's not applicable everywhere, Wardley maps helps you to visualize the journey to domain driven design and so Wardley Maps has been created by Samuel Wardley, a researcher from the UK and a Wardley Map is a representation of the landscape the business is operating in and it's really simple. So it consists of an Y axis for the value chain and then X axis for the evolution stages. A Wardley Map visualizes the evolution of a value chain, so the first question is so what is a value chain? So behind every user need, there is a value chain and start off with the questions like who are your users, who are going to approach you to get their problem solved, and what kind of user needs do these users have. Like what kind of problems they would like to get solved by you and what are the components and activities that are necessary to fulfill these user needs directly or indirectly by facilitating other activities. You have on the Y axis, at the top, we have those components and activities that are visible to your users, so where your users are touching your system and then at the bottom of the Y axis where it becomes more and more invisible to the users. And why you have to identify every component and activity that your value chain is composed of, you take these components and activities and plot it along an evolution axis, the X axis, going from left to right. And at the left of the X axis, we have genesis with brand new things that have never existed before, then custom built, then product and rental such as off the shelf products or open source software and then on the very right like commodity and utilities. The movement of a component or activity along the X axis is determined by its stage of evolution. So it comes also with, Wardley Maps comes also with patterns and principles, too, and for example, one of the pattern is that the map is never steady, they are very dynamic, so everything evolves from left to right, so with the force of supply and demand competition and as the components evolve from left to right, their characteristics change. For example, from the uncharted domain from the left, would be undefined market, it's uncertain, unpredictable, constantly changing and becoming more and more industrialized, going to the right where we enter the domain of the market of known, mature, widespread, commonly understood market itself. And also, one other pattern is like that efficiency enables innovation, so that means that the industrialization of one component enables the, for example, new features of existing products to appear or the evolution of other components or that new components can emerge enabling new user needs. So the evolution of one component and its efficient provision enables the innovation of others. But also that it comes with principles. For example, that you should use appropriate methods per evolution state, so what component should be built in house? These are the components that were residing in the genesis, in the custom built evolution stage, or where to use, buy off the shelf products and open source software so that other components and activities that are residing in the product and rental evolution stage and where to outsource to utility suppliers and these are the components and activities that are located in the commodity and utility evolution stage. And it's really important in order to know, like to use the appropriate methods because it's also like we don't want to ... We'd like to avoid too custom built commodities because we would like to custom build those components and activities that belong to our core domain and that's why I can make then the link to domain driven design later on. But it's also important to know your users and focus on your user needs.

Jeremy: Right.

Susanne: Because the user needs, they are the subject area of what we build software for, right? They are the reason, they are the why of our business domain but before we develop a solution that solves the user needs, we need to understand the problem domain first and that's why I would like bring in domain driven design where the ...

Jeremy: But before we get to that, though, because I don't want to get too far ahead ...

Susanne: Yeah.

Jeremy: Because I think like what you just said, again, like I said, we've talked about Wardley Maps quite a bit before and this idea of moving from this custom built stuff to all the way up to where you just outsource it and it's electricity, right? It's a complete utility. This is really where I think Serverless comes in with how we're planning on building some of these products, because it always used to be, again, you write some custom software that did something. Like right now, it would be crazy for someone to write their own database software, unless for their own company, unless they were building a new product but if you're just, if you need a database, you're not going to say, "I'm just going to write my own database software." Like that's done, like there's a lot of solutions out there and you're going to pick something off the shelf and you're going to run that and then it even has gotten to this point now with a lot of the tools that AWS has is to say, it would be crazy for you to install your own MySQL Cluster, for example, on EC2 instances or on VM somewhere. That would just be crazy because you have RDS and that whole process is not only commoditized in terms of the software that's running but then almost, sort of a utility at this point, by running the actual RDS cluster in the background and managing all the updates and all that kind of stuff for you. That's sort of that great ... and you had mentioned this, you said, this idea of building versus buying versus outsourcing, this is the thing where there are so many pieces of the infrastructure now that have become commoditized in a sense, whether that's SQS or it's even compute with Lambda or Fargate or things like that, that now really I think we're at a time where as technologists, we can focus on building just the things our customer needs and let all that heavy lifting go up to the cloud provider.

Susanne: Yeah, exactly, so it shifts the focus back on the things that are really important to our business and to our customers, right? So instead of handling all these infrastructure and operational complexities, without serverless, for example and that for example, if we are a small team and handle all these infrastructure operational complexities all by ourself, then we don't have time to solve our users' problems and to focus on our core domain and provide, yeah, and focus on our competitive advantage, instead our focus is blurred away and handling all these infrastructure and operational complexities. With serverless, we have the possibility to shift our focus back solving our user's problems first and getting, as you mentioned the undifferentiating heavy lifting and outsource this to cloud providers using serverless technologies for the time being.

Jeremy: Right. Right, so I interrupted you, you were getting into more about domain driven design.

Susanne: Exactly, so one of the principles of a Wardley Map is that we know our users, and focus on our user needs and because that's the why of our business domain, right, but before we develop a solution, because I can talk of myself. I tend to jump in, right into the technical solution without understanding the problem first and-

Jeremy: Right, I think a lot of us do that, yes.

Susanne: It's like, "Oh, yeah, yeah," like first to wire something together and then, "Oh, no" it's not really matching the requirements that our users are looking for.

Jeremy: Right.

Susanne: So we need to understand the problem domain first and that's why I would like to bring in domain driven design, where the collaboration between the domain experience and development teams is an essential part to obtain domain knowledge, and which is described in terms of a shared language, the ubiquitous language. And the domain knowledge is really free of any technical terms itself. It's just the shared language between domain experts and development team. The domain knowledge is very crucial because as mentioned earlier the lack of such can have a huge impact caused by poor quality software.

Jeremy: Right.

Susanne: Yeah, so that's the reason why...

Jeremy: Yeah, no, I was just going to say that's one of those things you keep talking about these sort of two sides, right? You've got the domain driven model or the domain expertise and then you have the technical expertise, and so in terms of taking an approach to building a product, right, there's that ... and this is something you mentioned in your talk where there's sort of those two sides to it, right? There's strategic design side and then there's the tactical design side.

Susanne: Yes.

Jeremy: So maybe we could get into those a little bit.

Susanne: Exactly, and that's also why I would love to combine it with Wardley Maps, domain driven design, combine it with Wardley Maps, because ... so domain driven design comes with patterns and practices and those can be categorized into strategic design and tactical design patterns and practices. And when we entered the field of strategic and tactical design and would like to combine it with Wardley Maps and use the Y axis to visualize that position, and the value chain going from top, the strategic design to further down the tactical design patterns. As I mentioned before, we tend to jump directly into the tactical at the bottom first, before we have covered strategic design and I would like to avoid it by putting Wardley Maps in place, starting from top to down, starting with the strategic design first. So we start at the strategic design within the problem space and that's where we analyze the problem domain and then like discover sub-domains. And problem domain or business domain is a business overall activities, for example, the services a business is providing to its customers and the problem domain is composed of sub-domains. The sub-domains represent a set of interrelated use cases or business processes where ... and we are distilling this problem domain and partitioning them into smaller sub-domains and by partitioning it, the problem domain into smaller sub-domains, we are on the one hand reducing complexities but on the other side we also have to consider that not all sub-domains are equal and some sub-domains are more important and more valuable to the business domain than others. So we have three different kinds of sub-domains, so we have the core, the supporting, and the generic sub-domain. The core sub-domain, that's the essential part of our problem domain, providing the competitive advantage, that are those parts of the system that makes it a success and should be hard for competitors to copy or imitate, so they are supposed to be quite complex, they are supposed to have a complex business logic and they tend to change often. That's the core domain, that's where we have to strategically invest most and innovate on. And that's the sub-domain we need to build in house and this, the core domain that's supposed to go into the genesis and custom build evolution stage of the Wardley Map.

Jeremy: Mm-hmm.

Susanne: And then the other sub-domain type that we have is the supporting sub-domain, so that helps to support the core sub-domain, but does not provide any competitive advantage. So they are quite simple, but do not change very often and if possible, you should look out for buying off the shelf products or open source software. If that's not possible for what reason ever, you should not, you're supposed not to invest heavily in these systems. They should be quite simple, but look ... So these are the supporting sub-domain that should go into the product and rental evolution stage of a Wardley Map. Then the generic sub-domain, these are sub-domains that many large business have. For example, authentication, payment handling, or something like that, so they are on call usually, and provide no competitive advantage but the businesses cannot work without them. So you cannot work without authentication in most cases, but they are generally complex but they are already solved by someone else. The generic sub-domains supposed to go into the commodity and supplier evolution state of a Wardley Map, so we should focus on buying ... or to open to outsource or to utility supplier or if that's not possible, at least have it then part of the ... Can be combined like going into the product and rental or commodity and utility evolution stage. So either buy off the shelf products or use the open source software or outsource to commodity suppliers. So that's how I would like to handle the strategic design combined with Wardley Maps and further down we enter then the tactical design patterns that I will probably talk later about it like that supports making low level design decisions to architect and implement a solution but first we should go, when we go from top to down, we should start with the strategic design first and then go further downward into the technical design later. I mean, it's not necessary to use the technical design, but at least we have to understand the problem domain first, using the strategic design. After we have analyzed the problem domain and distilled the problem domain into sub-domains and also discovered our core domain where we strategically, we have to strategically invest most, we can go further down and switch to the solution space of strategic design and that's where we come to the domain model and later on to the bounded context. So the domain model is at the center of domain driven design and within each sub-domain, a domain order can be created representing the domain logic and the business rule that are relevant to that area of the system, and the domain model comes in different shapes. For example, in the beginning it's formed as analysis model during the collaboration between the domain experts and the development teams. For example, it could be UML diagram or product sketches and later on it can result in a code model when we come to the technical design. And the domain model is described in terms of ubiquitous language and is free of any technical complexities but a model, a domain model cannot exist without a boundary and that's where we come to the bounded context so in that--

Jeremy: Sorry, before we get into that though, because this is one of the things that I think some people get hung up on. It's this idea of ubiquitous language. So I think I mentioned this earlier but I think we should be clear on what exactly we mean by that, because I think for a developer, if you say, "This is a customer," or "This is an order," like that seems pretty straightforward but to a domain expert, something that is, say, a particular status of something, like an order that is pending, that may mean something completely different to a domain expert than it does to a developer. Like a developer says, "Oh, the order's pending." That means, "I just flag it with a pending status." But there could be a whole bunch of other stuff that's involved around what may seem like a very, very simple term. Can you expand more on that ubiquitous language thing, because I think that's really important?

Susanne: Yeah, so it's also like to build in the common understanding, the shared understanding of the domain itself and then through the collaboration, in terms of collaboration between the domain experts and the development teams. During that collaboration, so this misunderstanding has to be to cleared out like for example, what do we understand for like what is a pending status? Like in terms of it's not only a flag in a database, it also comes with business rules.

Jeremy: Right.

Susanne: For example, if an order is pending, then it means, for example, we cannot delete it or something like that. So there are some invariance that we have to check when ... and that are applicable to specific ... yeah, to its domain model and so that means that it comes with a lot of new ... With business rules or invariance that's what they call it in domain driven design, and we have to make it sure that we keep the domain model consistent to protect its integrity within this bounded context and so that means that not only a flag, but also it comes with new, yeah, with a lot of rules and invariance that we have to check.

Jeremy: Alright, I interrupted you. You were talking about bounded context.

Susanne: Yeah, so as mentioned that a domain model cannot exist without a boundary, then that's where we come to bounded context and a bounded context provides different types of boundaries for a domain model so it forms, it form a consistency boundary around the domain model and protects its integrity and it could also form a linguistic and semantic boundary so that the language's terms are only consistent inside of its bounded context. So for example, pending in one bounded context could have a different meaning than another bounded context, for example. And it also serves as an ownership boundary, so for example, bounded context could be implemented and evolved and maintained by one team only and a single team can, on the other hand, can also own multiple bounded context but it's really, really relevant that multiple teams are working on the same bounded context, because this enables a ton of teams working at their bounded contexts independently at their own pace and with minimal impact across other teams.

Jeremy: Yeah.

Susanne: And this also serves as a physical boundary and can be implemented as a separate solution and can be deployed independently as separate artifacts and also enables separate data stores which are not accessible by other bounded contexts and, for example, also the source code could, of each bounded context, can be maintained in separate git repositories with their own CICD pipeline. And also each bounded context can have separate architecture patterns applied and that's where we enter the tactical design area.

Jeremy: Right, and before we get to that though, I just wanted ... Because I think we hear this term bounded context quite a bit when we're talking about microservices and this is exactly what we're talking about here. But that is certainly one of those things where when people move to microservices, it's always how do I ... What is a bounded context? What does that mean? Should I create a billing service? Should I create an inventory service? Should I create an alerting service? Where are the separations there? This is exactly what you want to follow, right? You want to think about all these different ideas, you want to talk about the language barrier. Like if you have two domains that are trying to work with one another, you don't put those into the same bounded context, because the language could be different. The infrastructure could be different, all kinds of things like that. The physical boundary, the ownership boundary, these are all great things, so if you're thinking about building microservices start paying attention to everything that Susanne is saying right now, because this is super important. Sorry, continue.

Susanne: Yeah, no problem at all. So yeah, and each bounded context can have separate architecture or patterns applied. That's what I was mentioning that we enter now the technical design patterns and practices and, for example, one bounded context can go with the layered architecture or where you split your source code into layers such as presentation, business logic and persistence layer, or as a hexagonal architecture, a specific form of the layered architecture ... I'll come to this in a minute ... or for example, I'll CQRS command queries responsible segregation where you split your source code into command models and read models and for updating the state of your domain model or reading the state of your domain model. I would like to highlight the hexagonal architecture a little bit. I know that you have already covered this in a previous podcast already, but it's specifically also combined with serverless when you transform the value chain going from open source to, for example, to serverless. Combined with domain driven design, hexagonal architecture is really helpful because it aims for separation of concerns. It's also called ports and adapters and what it does it creates a loosely coupled software components that can be easily connected to the software environment and by using ports and adaptors. And your softer components are categorized into an outer and inner part where your ports and adapters are used to connect from the outside to the inside and from the inner part to the outside and its core, the inner part, there is the business logic.

And hexagonal architecture makes your components exchangeable and so matter what kind of infrastructure you are using at the outside, the business logic at the inside can remain the same and makes your business logic adaptable for future evolution. So for example, when you started with your core domain, using open source software like MongoDB and using Key Cloak for authentication, something like that, so and you are now shifting to AWS using DynamoDB and using serverless Lambda functions to provide and back end API and in domain driven design your domain model remains the same so it's representing your business logic including your rules, but the infrastructure can be exchanged without touching the inner part of your system.

Jeremy: Yeah, so I mean, ports and adaptors and the hexagonal architecture like we did talk about that before but this is a super important way to build software, especially with serverless. I started doing this actually, after I talked to Slobodan Stojanovic. He was a big fan of it and I said, "Ah, I'll give it a try." Because I would often build in a lot of logic right in Lambda functions or you have some of those calls that aren't quite separated out and I've always been a big fan of building like data access layers, right? So you always build some simple interface to your database so that you don't have to always be writing sequel queries or queries against DynamoDB or something like that. But this great, because it does help you evolve, right? Like as you said, like if you are using MongoDB and then you decide, "Hey, I want to go to DynamoDB," forget about moving data which is a different problem, but in terms of your software, and the architecture that you build or I should say the models that you build, those can stay the same and then you just have to again, sort of swap out that adaptor and have it access X Database, rather than Y Database. So yeah, so then how do you implement some of these business pattens themselves?

Susanne: Yeah, so ...

Jeremy: Or I should say, how do you implement some of these business logic patterns?

Susanne: Each bounded context can be implemented by different business logic implementation patterns. So for example, for our core domain, it could be implemented as a domain model related to our domain driven design, with its building blocks of aggregates, entities, value objects and so on. And so the domain model is an object model of the domain, which incorporates both behavior and data ... but it should not be applied to everywhere. So it copes very well with cases of complex business logic and complex business rules and it's very well suited, as I mentioned for implementing the core sub-domain, which goes into Wardley Maps into the genesis and custom built evolution stage. And so the domain model should be free of technological and infrastructure complexities and I come a bit later to the building blocks of the domain model but I would also like to highlight other business logic implementation patterns that are possible. So there are a lot of ... I just want to highlight two others like the active record, for example. This active records represent a role in a database table of you and they encapsulate the database access and also adds domain logic to the data and it's ... so it carries the data and the behavior and the access to the database itself. So the active record supports cases where business logic is quite simple, but operates a more complex data structure and it's very usable for the supporting sub-domain, and domain design so ...

Jeremy: Yeah, that's what I was just going to say is that a lot of ... That's one of the things I think you mentioned earlier was you don't use domain models or core models for everything, right?

Susanne: No.

Jeremy: Like that's the whole point of the build versus buy versus outsource type thing is you want to make sure that when there's something simple that needs to be done like simple stored database record and there's not a lot of complexity to that, that using something like active record is a better choice than trying to ... Because you don't want to reinvent the wheel, essentially.

Susanne: Exactly, and so if you try to apply domain driven design to cases with less complex business logic, then you're over engineering it.

Jeremy: Right.

Susanne: And so domain driven design is really very well suited for complex business logic cases and these are supposed to go, these are supposed to be your core domain, right, which is in the genesis and custom built evolution stage of Wardley Map. So there it fits very well, but domain driven design, I would not recommend to you to use it for generic sub-domain. Sometimes they can use transaction scripts, for example, not a business logic implementation pattern where you organize your business logic by procedure and where each procedure handles a single request from the presentation, so it works very well with small applications that does not implement any complex business logic and they can just, yeah, for example importers or authentication using Cognito or something like that.

Jeremy: Yeah.

Susanne: So you don't have to apply domain driven design into that generic sub-domain case.

Jeremy: You mentioned the building blocks of domain model, so what are those?

Susanne: Yes, and they could be kind of like overwhelming, because there are a lot of building blocks but I guess it's ... and if you start, like it could be the case that you have the impression, "Oh, this, my first model does not look very nice," and that's okay, because when you start at the beginning, first versions, it's kind ... an iterative approach, right?

Jeremy: Sure.

Susanne: While you're collaborating with your domain experts, while you're gaining more knowledge, your domain model is evolving as well and also like, you can't do it right from the very beginning but ... Well, that's at least my experience. I don't want to ...

Jeremy: No, I think with anything, it's like designing a DynamoDB table. You're going to get it wrong the first time, and you just got to keep working through it until it ... I mean it's with all software. An iterative approach to anything it typically what we do because we think we're better than we are and we just make a decision and then we realize, "Oh, that was a bad decision," a couple of weeks later.

Susanne: Exactly, and you're always smarter after we have gained all the experiences, right.

Jeremy: Exactly.

Susanne: Specifically with DynamoDB details.

Jeremy: Right, exactly.

Susanne: "Why is it so expensive?" And so, yeah, domain driven design comes with building blocks and I would like to highlight a few of them. So for example, we have ... I guess the most simple building block is the Value Object. A Value Object is an immutable object, which can be identified by its value and when you change this value, you will replace the entire object instance. For example, a common example that the people are bringing us, like the address of the a customer, that could be a value object, but I also like to go in more fine grained but I talk to that one later one. The next building block is an Entity and an Entity is identified by a unique ID and it can change its state over time. So like, yeah, that's something that comes, it's incorporated into an aggregate, so an aggregate itself is composed of, represents a hierarchy of objects and it's composed of one or more entities and one or more value objects and where one entity is called the aggregate root and an aggregate root is designated as the aggregate's public interface, that you can use like when you, for example, when you use a domain model that is an aggregate, when you load it from the database you are going through its aggregate root by its unique entity ID, load it into a ... and also ply then the ... Like renaming it or rescheduling or something like that that's applicable to the domain the logic. That's usually then the aggregate roots matters that you are using. Then you also have the repository that can save and retrieve entities, aggregates from the underlying storage mechanism and you also have an application servers. These are only, their purpose is to only orchestrate use cases and manage transactions. So they do not contain any business logic but they are the gateway to the domain model, to the aggregates that are building your domain model, and another thing also is domain event. This is a message that describes a significant event that has happened in a new business domain, for example. So there are a lot of building blocks and they are kind of like, oh, really, not really tangible at the very beginning, but the more you try to use them, for example, also like, during a workshop I was building a conference solution and I tried to compose my domain model of aggregates and the aggregates were composed of value objects and each of them is taking care of keeping the invariance, the business rule was its integrity. So that you can make sure even though you might have a lot of classes and compressed together with each other and it says, "Yeah, why do I need so many classes?" Because each of them is ensuring that your aggregate composed of these components, of these value objects and entities, you can make sure that they are consistent and they are valid every time in your system, because they take care of it. They take ... Whenever you compose it, at that time, they are valid. They have matched the validation restrictions, for example.

Jeremy: Well, I mean, even just object-oriented design, I mean that software pattern of creating a new object that has a very specific structure, and then specific methods that can interact with it. It's very similar I think in a sense to what, sort of what you're talking about with these individual bounded context, only a lot more complex, but that ability for you to create these things and make sure that when they're created, and the things that they can do. But these are all very specifically prescribed, like it's very restrictive in terms of what these things can do so they can follow all those business rules.

Susanne: Yes, and so for example, when you have, when your domain model that is composed of ... it could be one aggregate, it could be multiple aggregates, for example, and they are having methods, for example, that are really matching consistently with the terms of your ubiquitous language.

Jeremy: Right.

Susanne: You don't have a domain model that consists just of setter and getters and changing state, instead they have methods that are reflecting the shared language. For example, if you have ... I don't know, for example you have a customer aggregate and you are ... The company name is changing. You will have a method that's called rename and containing them the name as a parameter or for example when the address has changed, you're not saying, "Okay, set address." Instead you say, "Relocate," or something like that.

Jeremy: Mm-hmm.

Susanne: So that it's really reflected to the domain language, to the ubiquitous language so when you talk to your domain expert, you share the same language so that you don't have to translate it, your technical terms to the domain experts' terms. Instead you are using the same language and this also allows you to get rid of a lot of misunderstandings.

Jeremy: Well, and you had mentioned, too, like that relocate. Like sort of let's say I have a functionary method that allows me to relocate a customer. It's not as simple as just saying, "Okay, change the address," right? Like there may be rules in place where you have to keep X number of past address or you have to make sure you have to keep all addresses or you can only relocate the customer if the customer is active or they've been active within a certain amount of time.

Susanne: Yeah, yeah.

Jeremy: So all of those business rules, all those things that tie together with that domain model, all of that ties together with these methods that you talk about, so I think that is ... I think that's super important for people to think about it that way. It's not just update a record in the database.

Susanne: Exactly. It's also like trying to ... It's also matching the invariance, the business logic, the business rules.

Jeremy: Yeah, that's important. Alright, so if we want to use domain driven design, how can we use it to help us build say serverless like back ends, for example?

Susanne: So when I entered the field of serverless, I was with its fine grained short list function triggered by a variety of events, and I was kind of overwhelmed in terms of if you design a system composed of countless functions that might be the risk of losing sight of each function's purpose or it also be is the risk that we accidentally introduce a big bowl of mud. You two coupling functions together by accident or accessing resources that we should not access from various functions itself and one approach could be to introduce domain driven design. Because domain driven design allows you to organize the serverless functions and Lambda functions, for example, if you build a back end API. So it allows you to organize your functions together so that you, as I mentioned, like bounded context provides different forms of boundaries in terms of like also physical boundaries like that you, for example, you can organize all your functions that belongs to one bounded context, your module in terms of like your system, and it allows you to organize your functions, for example, that they all go in one good repository, for example.

Jeremy: Yeah.

Susanne: And that also allows you to organize your Lambda functions if you go on AWS, it allows you to organize them in terms of like what access rights do I have for the DynamoDB table, for example that I'm using, so only those functions that goes into one bounded context and located at one good repository, they are allowed to access the DynamoDB table that is saving the state of our domain model, but not other serverless functions that are relevant or that are responsible for another bounded context. They are not allowed to access my DynamoDB table directly. Instead, they have to figure out what kind of integration patterns we have between these bounded contexts like event driven, like for example submitting an event or into an SNS topic and our other boundless contexts and external bounded context is subscribing to this topic and then is reacting on these ends appropriately. But it's not accessing the data storage components of an external bounded context. So that allows us to get an overview, which serverless functions are solving what problem, what part of our problem, what bounded context, and they help us to organize which function skills into one good repository and also help us to organize what access rights do we have, that we are not allowed to access resources that are where other bounded contexts are in charge of.

Jeremy: Right, and that actually creates sort of that physical boundary, right? I mean in the sense where you can have ... Well, physical boundaries and ownership boundaries that we talk about earlier, because you can have a team that goes and just works on one bounded context and a group of functions and I think we've been seeing that quite a bit, that's sort of been the approach when people are taking a microservices approach to serverless that what they're doing is they're taking a single stack and they're putting all their related functions for that service in the single stack, have a separate DynamoDB table or other whatever their data store of choice is that they're sort of doing that. I think that makes a ton of sense and certainly the domain driven design approach to all of this and really thinking about that ubiquitous language. I know I'm hung up on this, but that to me is hugely important because I think a lot of people miss that and I think as you said, you know, sort of what's important for the customer is the most important thing and I know a lot of technologists who say, "All right, we just want to focus on what's good for the customer, but what we really want to do is figure out how to install Kubernetes," right?

Susanne: Exactly.

Jeremy: Like we want to figure out how we can do this cool technological thing, whether or not ... I mean it eventually will have some benefit to our customer and that's I think what slows down a lot of software and I think what you said right from the beginning is that's why we have so much poor software. Anyways, if people want to learn more about domain driven design, I know there are some books out there, but you're working on a workshop for this, right?

Susanne: Yes, yeah, so I have created a workshop but it's still under development and I have created a workshop that is focusing on an example, because I guess you can learn best when you have a practical example.

Jeremy: Sure.

Susanne: It's also a simple example, but it's for example a conference solution, where we identify the bounded contexts that are necessary to, for example, manage a conference event, including starting a CFP, a call for papers, evaluating the submissions, and create a schedule and choose and submit a session proposal from the speaker and something like that. And this, I was creating a workshop out of it in terms of like, okay, how to use ... One thing, like how to use Wardley Maps, how to identify what shall be core domain and combine with domain-driven design and what goes into the other evolution stages and also evolving the value chain going first from, okay, "What happens when we implement this solution using also the tactical design patterns after we have started with the spreadsheet sheet design." But how can we evolve the value chain in terms of like starting first with an open source solution, going on premise, and how does it evolve going then at the end to serverless, using ... providing serverless and back end API with Microfunctions, for example. And, yeah, it's a kind of like multi-stage workshop where we try to, yeah, going also from the strategy, like how can we outsource those components and what does it mean for our ... so for architecture and our software design and what kind of impact does it have and then we at the end, we figured out, okay so the business logic remains the same but ... and hexagonal architecture but the infrastructure components they are then replaced. For example, in one case, we are using a rest API to kind of communicate with another bounded context. In, for example, in the serverless example, I'm using then ... an SNS topic, as in forms communication via events and still the core, the inner part remains the same and I try to make it tangible for the audience, for the participants, how it looks like. Yeah, evolving the value chain.

Jeremy: Awesome, well that's going to be a very, very good resource, so thank you for putting that effort in and thank you, Susanne, for coming on the show and talking about this, because I think this is something that isn't talked about a lot on a lot of development teams and it should be more because we would be building better software and of course, I'm quite fond of serverless, right? So the idea of the things that you can do to reduce that build or that sort of implement yourself type projects is something that serverless just blows everything else out of the water, at least in my opinion, so anyways. Thank you very much for coming on the show.

Susanne: Thank you.

Jeremy: If people want to get in touch with you or find out more about your workshop and some of the other things you're working on, how do they do that?

Susanne: They can contact me via Twitter, for example, so my Twitter handle is @suksr. So I am also my direct messages are open and also I am having my own website it's www.susannekaiser.net and also you can reach me over LinkedIn as well.

Jeremy: Awesome. I will get all of that into the show notes. Thanks again.

Susanne: Thank you.

View Details

About Paul Swail

Paul is an independent cloud architect who helps development teams make the transition to serverless. He has almost 20 years’ experience delivering software solutions to clients across a wide variety of industries. In addition to client consulting, Paul has been running his small SaaS business on AWS for the past 6 years and is slowly migrating it to a fully serverless stack. He writes in-depth articles on serverless in his email newsletter and on his blog at winterwindsoftware.com.

  • Twitter: @paulswail
  • LinkedIn: https://www.linkedin.com/in/paulswail/
  • Blog/Newsletter: http://serverlessfirst.com/

WATCH THIS EPISODE ON YOUTUBE: https://www.youtube.com/watch?v=gf__z3K8LBI

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week I'm chatting with Paul Swail. Hey, Paul. Thanks for joining me.

Paul: Hey, Jeremy. It's great to be here. Thank you.

Jeremy: You are a cloud architect at Winter Wind Software. So, why don't you tell the listeners a little bit about yourself and what you do?

Paul: Yeah, sure. I'm an independent cloud architect. I work primarily with AWS and I specialize in helping development teams ship their first serverless application into production. I've been focused specifically on serverless for two years or so now, although I have been doing software development professionally for about 19 years in total.

Jeremy: Wow. Still not as old as me, but that's okay. One of the things that you've been focusing on lately... Or, I think you're making this transition into serverless first. So, why don't you tell us a little bit about that?

Paul: Yeah, yeah. My website is currently in the process of being moved to ServerlessFirst.com, so I think that... Basically, serverless first is a methodology, which I've been taking with my clients that I've been working with over the past couple of years. They're used to more traditional ways of building serverless apps, and they see the value of using serverless, but they can't use it for everything. By default, your architectural decisions start with serverless services unless you can justify using something else.

Jeremy: Awesome. Alright, I wanted to talk to you today because... I don't know what happened, but somehow you've become one of the most prolific writers in serverless over the last couple of months, releasing a few articles or a few blog posts every week. It's been awesome because every time we get more content, and you answer more questions, and you get deep onto one particular subject, I think it's super helpful. One of the things that you focused on quite a bit, I think, has been this idea of these communication patterns in serverless applications. You wrote two articles recently. One was called Seven Ways to Do Async Message Processing in AWS, and another one was Interservice Communication Channels for Serverless Microservices in AWS. Both great articles. Definitely go to WinterWindSoftware.com, check those out. Very, very interesting stuff. Why don't you tell us a little bit? Maybe you can get us started, in terms of what are the main communication patterns in serverless, and why is it so different, maybe, than a traditional application?

Paul: Yeah. I think a lot of my clients have come from the monolithic architectural background and asynchronous stuff. They may be aware of it but they haven't used it a lot. AWS has a lot of services which does... Around asynchronous messaging patterns and it can be hard to understand what to choose. So, a lot of it is just documenting questions that clients have had for me. A lot of the writing is around that area. We can go through the details of lots of individual services that I discussed in the article.

Jeremy: Yeah. Let's start, though, maybe just thinking about the difference between asynchronous and synchronous because I think most people are very familiar with that monolithic approach of... I should maybe take a step back. They're used to that request-response type mechanism, right? I make a request to a website or to an API and that data comes back to me. There's that one part of that immediate response. That's not going to change whenever you have a customer-facing or a web-facing side of things, but it's where the backend... The backend is what gets different. That is one of those things where, I think, when people are familiar with monolithic applications, they think, hey, I've got 15 different methods, or functions, or whatever, that are all in one big application and I can say, "hey, I need to process the order. I need to pull the inventory. I need to send the message." And that's all in one app, or one, I guess, big chunk of code, really. But when you start moving to this asynchronous thinking, we're starting to separate out these components separately. So what do people have to think about when they start building that type of application?

Paul: There are quite a lot of things to think about. I guess based on the workloads, firstly. If it's like a task or a job-based type workflow, you may want to do, or maybe that you need to notify a lot of other systems, so based around... That's a decision in itself, around what service you use, based on the nature of the workload. There are a lot of operational considerations just around throughput, and concurrency requirements, and latency requirements, and scalability of any downstream systems that you may need to talk to, and message durability, and your error handling, and retry mechanisms. These are all things... Oh, cost, of course, as well. These are all things that you need to consider around how you would structure any asynchronous messaging patterns that your workload requires.

Jeremy: Yeah. I think that makes a lot of sense. I think what you get people coming from the monolithic space or the traditional system space, and they move into distributed systems. Now, there's this whole different idea of passing messages around, right? We're no longer using a single system, so now we're trying to communicate with multiple systems, and as you said, things like durability of messages, that becomes a huge concern. Or, something like the error handling. What happens when you send a message off into the ether and you don't know what happens to it? Does it ever get to its destination? How do you know a message was even sent if that information isn't recorded correctly. I think maybe that's another thing that is interesting to me though, is that even people who maybe come from distributed systems and think about, oh, I've got to set up a Kafka cluster, or I've got to do something like that. Traditionally, there has been a ton of stuff that you would need to do just to create the messaging components between these distributed systems, but serverless changes that quite a bit.

Paul: Yeah, it does. I can echo those sentiments you said, I used to, back in my Microsoft.net developer days, setting up BizTalk servers, something actually I knew how to do, and we did it in a few projects, but just the provisioning and management overhead of doing it just... even though it was a nice distributed messaging pattern, it was just so much effort to manage. Whereas with AWS serverless and async services, it's simple. It's just quite a few lines of YML and SLS deploy, or whatever it is you're using to deploy it, and away you go. The operational overhead is just significantly less.

Jeremy: Yeah. I think that makes a ton of sense. When you're setting up these distributed applications using serverless, and you're using one of the pub/sub services, for example, like SNS or EventBridge, it makes it really, really easy. Maybe let's talk about pub/sub for a second.

Paul: Yeah. For folks who don't know, pub/sub is a system where you're a publisher, not a subscriber, and so it's a Web decoupling and separate services, if you even think of your application as having separate services, and you can simply... something happens in service A and you can just publish an event. Service B, may need to do some processing on that. It can consume that event, but the actual message communication between those is managed by a pub/sub system. Within AWS, we have two main pub/sub systems. We have SNS and EventBridge. I've been using SNS for a few years now, and I've just started using EventBridge. Very similar feature set, although EventBridge seems to be the preferred solution amongst most serverless experts these days. I think mainly around it offers more event targets. A particularly nice feature is around the schema registry, and so a big thing with pub/sub is your different services, and there is a small coupling around the actual schema. They obviously couple, but they do need to know what shape the message is going to be. The schema registry that EventBridge provides gives you a way, if you're using a tech language such as TypeScript or Java and you can download type definitions based on the events that are getting sent out. So, it means your consuming application knows exactly what to expect and you can, at compile time, you how to work with your messages.

Jeremy: Yeah, yeah. And I think as you said, it's... I think things are going to move to EventBridge. I think we've seen a lot of noise around this and I've been using EventBridge, in production now, for several months and there are some really, really cool patterns that you can do. Just with the huge throughput. I know there's some concerns or some talk about some latency issues and things like that, but it also depends on what your work flow is. SNS still has its place, right? So, SNS, still a great pub/sub service that you may need to use. So, let's talk about the benefits of using asynchronous versus synchronous, right? I think that most people, again, going back to that monolithic example, it's, alright, an order is placed, so I have a function or a method in my system called process order. And then process order has to do the inventory. It has to do the billing. It has to do the messaging. It has to do the invoice creation or some of these other things. That all happens by using internal method calls, or whatever, and it's easy to call those. When you start switching over to distributed systems, you might have a billing service. You might have an invoice service. You might have an alerting service and that is all a simple way to think about it, if you're doing a container based system and you might want to do asynchronous communication that way, but when you're using serverless for breaking it down into even smaller functions and things like that. What are the benefits of being able to split those up into jobs that don't have to immediately give you a response?

Paul: There's several benefits. One benefit is that it's simply... It's easier to reason about it if you've a single task. As a developer, you can see just by reading it, what it's actually doing, rather than having a huge, as you said, a single function, which does the orchestration of all the things within code. From a logging point of view... From a monitoring point of view, say within a AWS client watch logs, you know exactly... You can go into the functions logs just to see exactly what happened or what went wrong. A major benefit is around retries. So, say you've a multistage work flow and say steps A, B, C, and D. So, say step A succeeds, step B fails, and you don't want to... If that was a single Lambda function carrying out all those four steps, then you would have to retry it all or you would have to build it all into your logic, how that gets handled. If you split it up into four separate asynchronously invoked Lambdas, the Lambda service itself will do the retries for you. In my example there, step B, say it fails, Lambda: A will have completed successfully, that's done, but B will be automatically retried if you so wish it to be. So, that's another benefit around the errors and retries. The most obvious one, which I forgot to say, is around latency. If you have a Interface and API call, if this is all behind, then the user probably doesn't need to wait for all four tasks to complete. Whereas, if you just have an API Gateway call, you can just write it to a queue or write it to a pub/sub system and just return to the user. All the rest of that processing happens in the background and the user gets a quick response.

Jeremy: Yeah. I know, I totally agree. I think that's one of the... That idea of immediate response is huge, cause everybody wants stuff back quickly. There's no reason why that, that invoice or that credit card has to necessarily be processed immediately. I mean, you submit an order on Amazon, and it says we've received your order and then you usually get something later on that says oh, we couldn't bill your credit card, or something like that. The other thing though, I think that is really great about splitting things up into separate functions, is this idea of, not only the security aspect of it where each one can be finely grained tuned or finely tuned security, but also this idea of being able to scale each one of those things independently.

Paul: Yeah, that's right. So, if you have... Say one of your tasks, it talks to its answering system, say in the RDS database or a third party API, which may not scale as well as serverless services. You may want to throttle it, using the async. Say for the likes of an SQS queue. You could use that to throttle your throughput to that system. That's scaling down and such but you can also... If you don't need to do that, you can use the SNS and just have a fanout pattern to distribute it as widely as possible. Say if you have a task which is split right into DynamoDB, or something, which can match the scaling that Lambda would give you.

Jeremy: Right. Yeah, I think that... I mean I do a talk on this about downstream systems not being able to handle the amount of pressure that might come from certain workload. I think that's actually a really important thing to consider... Is the great thing you have with just putting concurrency on a particular function, is to say, "Look, I only want 50 concurrent connections". Or, "I only want 100 concurrent connections". That can seriously help when you're trying to throttle against a downstream API or, like you said, a RDS cluster, or something like that. That's one of the things, too, that I think, is where people get a little bit confused. They say "Well, normally, if I make a request to an API and I've got to call the billing service", you know? Somewhere in between there, if that thing is being throttled and I have to respond back to the customer and say "Hey, I'm throttling this". That's a bad experience but with the async piece here, and putting something in between, like SQS for example, to be able to store and create message durability, that is something where people get a little bit lost on how that works.

Paul: Yeah. Yeah. It can be quite confusing and I guess, if you want to do a full asynchronous back to the user, in that example, you could even introduce WebSocket. You can have the initial API call from the client just writing to the queue. Have all your async processing happening server side and then separately then, do a WebSocket push back to the cloud, once it's ready.

Jeremy: Yeah. That makes total sense. All right. So, let's talk about billing for a second. How are you seeing with your customers, their reaction to this pay-per-use billing? Cause I think in most cases, we're not afraid of these $0.05 Lambda bills but when you start adding on API gateway, and things like DynamoDB, and SQS... I saw an article the other day where it was this application where Lambda cost $0.01 and SQS cost $1.83. So, again, it gets... Other services, besides just Lambda get kind of expensive. So, how're your clients seeing this and sort of planning for the cost?

Paul: Yeah. It's more... With serverless, it's more but the variability than the actual absolute price. So, it's generally low and it's generally negligible in pre-production scenarios. It's the variability of... Especially with large asynchronous work loads that can fan out to lots of invocations. A lot of the times, the client comes... One of my clients is - they're a dev agency themselves. So, they build a lot of apps with totally different workflow patterns. At the start of each project we would often just... It's in an excel spreadsheet, just plugging figures in to see, based on expected usage or what our current architectural design there is and how much it could be. It could vary quite significantly. API gateways is one of the most variable, off the serverless servers, but sometimes clients have existing infrastructure as well, which gets hit. API gateway... I know you did an episode recently with the API gateway team, with the new HTTP API. I haven't tried those out, so hopefully that will... I think they're 33% of the cost, so hopefully that will help out on that front.

Jeremy: A little bit less, actually. It's $1.00 per million, as apposed to $3.50. Yeah. That and... I mean, actually, you bring up API gateway... I mean, I think one of the things we can't lose here, and we mentioned this a little bit earlier, that there are still synchronous use cases, right? We can't just throw away synchronous altogether. Sometimes synchronous might not just be the front end web API. We might need synchronous invocations, and in most cases, even in asynchronous process, needs to make a synchronous call to something like the Stripe API, or maybe to another microservice. What are some of the pros and cons of using... I mean we certainly don't want to try to chain synchronous invocations, right? We don't want to call our API through API gateway, then make a call to some other service that makes a call to another service. Decoupling those, I think, makes a lot of sense but you can't always get around that. So, what are some of the pros and cons of some of those synchronous patterns?

Paul: I guess it's the client gets an immediate response. If the client needs an immediate response you can't... If it's just simply fetching from a database that it has to be synchronous. If it's a GET request, HTTP API... I guess a benefit... You can say it's easier to reason about. So, from a developer debugging point of view, I still would find synchronous calling patterns. The logs are generally easy to find. You don't need to look through log files for separate Lambda invocations often. So, from that point of view, synchronous is still easier to monitor and to debug, I would say. You could possibly argue that it's maybe easier to author in the first place. Well for developers, who like writing Javascript code or Python code, or whatever, lots of YML to configure. Those guys would be probably happier writing simple synchronous code but we try to move folks away from that but... Generally, I would recommend, if you can do it asynchronously, do it asynchronously. Write something too if you're going to HTTP user API, just write something to a queue or to SNS or to EventBridge, or to DynamoDB. Just return and have any other processing done in the background.

Jeremy: Yeah. I think especially when it's POST request, right? Any type of write request, you can return something back, even if it's a... If you have to return an ID back to somebody, then generate a UUID and return that, and also submit that with the job so that, that gets associated with that job and you can look it up later, or something like that. But I mean, if it's a GET request, right? I mean, there's a lot of caching you can do. I think people don't take advantage of the CloudFront caching through API gateways. How you can actually change the amount of time something gets cached for, so you can minimize impact on the API. A whole bunch of interesting tricks that you can do there. Another thing, though, that comes up, is once you start using multiple Lambda functions, is this idea of function composition, right? We talked about the asynchronous patters. Where, sort of, one function generates an event. Maybe that goes to EventBridge and then another one is listening to it. But what are your thoughts on Lambdas invoking Lambdas? There's two ways to do it, synchronously and asynchronously, but maybe your thoughts on each?

Paul: Okay. Let's take synchronously first. Generally avoid against this, but the reason why I would generally say don't do synchronous Lambdas is because say, you have Lambda A invokes Lambda B and it's waiting on the response. As soon as that invocation happens, you've now... the clock is running on the invocations of two functions so you're paying twice for both functions being running. Also, there is... It can be quite difficult. You've also two places to look for the logs, as well. People who are fond of reusing functions, like functions at a code level, rather than the Lambda function level when they first code the Lambda, they may think oh I've got this piece of functionality which I want to reuse. I can just create a Lambda function for that and call that from all the other Lambda functions where I need it. Generally, that's not what you want to do. Just use a code module or a code library and reuse it in that way. There's one exception where I have done it, invoked it synchronously. That is when I'm using VPC. It's pretty low. I had a use case recently where I had the scheduled job that runs nightly, it's for a SaaS app that sends out nightly e-mails, but to do that it needs to query a RDS database, which is inside of VPC and send it out to all the users with a certain... Who matches a certain workflow. There's, sort of, two things going on there. There's a cloud watch reel to trigger Lambda then there's a database query and then there's... I think I was publishing to SNS or using TSCS to send the se-mail at the end. If I just had that as the single Lambda, running inside a VPC, I can't then call the e-mail service and so, in that case, I put a Lambda inside the VPC, just to do the query, the database query and the calling function, invoked that, got the result back and then it was able to send the e-mail because it had internet access. So, that's the only sort of exception where I think it's valid to do that, but some people may argue that's not even valid and-

Jeremy: So, I actually think that's a really good use case. I mean, the only other thing I would suggest is maybe use the RDS data API if you were in Aurora Serverless database. Then that way you could use... and actually that's one of the things I do. I have a reporting service, I think it's a cool set up, where basically all these operational front ends are DynamoDB. They have DynamoDB streams attached to them that replicate the data to an Aurora Serverless cluster and then there's a reporting service that runs queries against that through the data API. Surprisingly, it's extremely fast, the data API, that's not the fastest sometimes when you're doing synchronous stuff but it works pretty well, and then of course, it avoids that VPC issue. You mentioned this idea of separating functions, as you said, maybe that you might consider code modules or code libraries. Where again, you might a function that is charge credit card and you might have another function that is create invoice. I'm using these same examples, but there's a really, really good argument. I know LEGO does this, I know that Bustle does this, where they create these fat Lambdas, right? Where Lambdas do more than one thing because you need to have that synchronous component. Now, maybe the one I just mentioned is probably not the right idea, but there are times, I think, where you don't want to be calling Lambdas between Lambdas in most cases. If you do find yourself needing two separate pieces of logic that need to happen asynchronously, this idea of fat Lambdas is, I think, really interesting. Where you just put these two things together and say "I'm going to run these two snippets of code in the same Lambda function because they get the benefit from that".

Paul: Yeah. Yeah. I would do it like that. I don't even know if I would... Cause sometimes, I guess fat Lambdas can be misunderstood to be, I won't have even call that a fat Lambda such as... I guess it's just...yeah.

Jeremy: Well I break Lambdas into three categories. I have the single purpose function, right? Which is the one that we try to favor. Then you have the fat Lambda that takes multiple bits, or multiple actions and puts them into a single function to optimize it. Then you have the Lambda-lith, which is your entire application, runs in a single Lambda function. So, I think fat Lambdas are an optimization and not necessarily a bad pattern, because sometimes you need to do that.

Paul: Oh yeah, absolutely. Absolutely, yeah. There's no point in being dogmatic about just... I would generally say single purpose Lambda functions, but yeah, if it's something which will always have two natural steps, and you can't really... It doesn't make sense to split them apart into their own Lambdas then go for it.

Jeremy: When it has to be synchronous. The other time I think that a synchronous API call to another Lambda function isn't the worst, is if you're doing interservice communication and you need a synchronous request. If you have a customer service and you have an order service, sometimes that order service needs to look up that customer. When that happens, you definitely don't want to route that back out through the API gateway, or something like that. That just gets overly complex. You do build coupling in, but you build coupling in on all kinds of things, because think about if you're calling the Stripe API or you're calling the Twilio API, your service is bound to, or is coupled to, that other API. I think that's okay in certain circumstances, as long as you have the fallback methods in place. If you're going to process an order and the order API or the order service has to call the customer API, that is fine as long as you, again, maintain some sort of contract between the two. But also, knowing that the order service could fail that particular call and then retry, maybe once that other service comes back up. Of course, you'd want to do circuit breakers and all that kind of stuff in there. But what are your thoughts on that?

Paul: Yeah. I think that's a totally valid pattern. If you have the two separate services, I guess, if you don't want to introduce the coupling, you could somehow have an event-based model where each service keeps a copy of the other's data, as such, so it doesn't have to do those synchronous calls. But sometimes, that introduces problems itself, just keeping the data in sync as well.

Jeremy: You've got to get the data the first time, right. So, it's okay to maintain a copy of the data but you still have to get that data the first time.

Paul: Yeah. Well, I guess you could have an event-based pattern where the service with the data, in the first place, could publish an event and-

Jeremy: Yes. You just...

Paul: You could transfer it asynchronously like that. I'm not a massive fan of that approach either. Sometimes it is just easier to make that synchronous call, as long as you don't have too many. If it's one or two between a service, that's okay. Once you have more synchronous interservice channels then that's where you're sort of losing advantages of having microservices at that stage.

Jeremy: Right. Yeah, I agree cause too much coupling can cause some problems. What about the asynchronous calls of other Lambda functions? Cause one of the complaints that you hear about this, is that even if you are chaining Lambda functions and everything is happening asynchronously, you are introducing a lot of coupling and then you really reduce the amount of reuse that you get from a single Lambda function.

Paul: I can understand the reasons for that and I guess... I don't know if you're alluding to, there's Lambda destinations recently, which have, a few months ago AWS announced it. You can configure any asynchronously invoked Lambda functions. You can configure another Lambda function to always take the result of the first invoked function and pass that into the event of the next one. That coupling, I can see it, but from a reuse point of view, I'm not sure, can you maybe directly... I think you can still synchronously call the function.

Jeremy: You could and actually...

Paul: You won't get the side effect of the destination being invoked.

Jeremy: Right. Yeah, and I should clarify this to because essentially the point that I'm getting at is that if you write it, if you hard code it into your code, right? So, every Lambda function that gets invoked regardless of whether synchronous or asynchronous, is always going to have a hard coded next step, right? I mean you could put a whole bunch of logic in there if you wanted to, or whatever, but there's always that problem, of that function getting invoked synchronously, asynchronously and something happening where then that next call doesn't happen. Which is where Lambda destinations comes in. I love Lambda destinations. I think it's a great pattern to have another Lambda function or in most cases, EventBridge, or something like that, respond or be able to handle the output of an asynchronously invoked function. I think I'm talking more about sort of hard coding that work flow into Lambda functions to say when this Lambda function is done then it calls this Lambda function, and then when this Lambda function, then it calls this Lambda function. I think you introduce a whole bunch of challenges with that approach.

Paul: Yeah. I guess you do and I guess there is step functions, which for certain workflows, that makes sense to solve that problem. If you have a workflow, which is pretty well defined, like a business workflow, it may make sense to... You could put 10 individual Lambda functions together via Step Functions state machine. From that, you can still invoke the individual Lambda functions separately, if you need to, or you can use the Step Functions to compose to pass the output of one to the input of another or to fan out the more complex mechanisms if you need to.

Jeremy: Right. Yeah and Step Functions, again, I say this all the time but for complex workflows, Step Functions are great and for function reuse... and that's one of those things where if you think about your traditional monolithic application, you say, I got to process the order, charge the credit card, create the invoice, do all these things, and all those steps, you need all of those to complete. You need that guarantee. If you can do that asynchronously, meaning that you don't have to immediately respond to the client and say, all this stuff is done, which again, the more complex systems get, the harder it is to respond in a short amount of time, then using something like Step Functions is great because then you can basically set up your retries separately. If something fails or a whole pattern of them fails, or a whole bunch of them fails, you can implement a SAGA pattern and you can go back and unwind all of them. There are a lot of really cool things that you can do with Step functions. They're all invoked synchronous too, so you know exactly what's happening as that state machines is processing it. So, when we get beyond just regular, or I guess, basic messaging between systems in distributed systems, and more so in serverless systems, we've been talking about passing messages back and forth and invoking one Lambda function at a time, or things like that. You kind of mentioned polling or you mentioned SQS and some fan out and some of those things. That's another thing that is a very powerful way to communicate or to build distributed systems, is to use queuing or streaming or some of these things. We obviously have SQS ques. We have Kinesis streams and we have DynamoDB streams, things like that. What are the... I guess, for the benefit of the listeners, what are the differences between those and maybe when and why would you use different ones?

Paul: The benefits of Kinesis is that you get a long back log of events. As an example, I have a SaaS product, which has website click tracking so the fans coming through from different websites, at a high throughput. So, we need to capture them quickly but the processing doesn't need to happen that quickly. The processing can happen gradually. We have a long back log. It gives you long back log potentially, of events and it's similar to the way a queue does but unlike a queue, where you just have a single processor, and pulling items off the queue. With the stream you can have multiple subscribers. In a way, it's a combination between a queue and a pub/sub, in that respect. You can have multiple subscribers processing messages off the stream, although, other consumers can consume that same message. Based on that DynamoDB streams is similar in that, if you already have an application, you're using DynamoDB, as your application database, but you need to react based on certain data that gets put into your system, DynamoDB streams might be a good fit. It gives you an asynchronous event model based on an item is added or updated or deleted on your DynamoDB table. You can then in a separate job, get notified about that even and do whatever processing you need. A drawback of using DynamoDB streams is that the event schema that you get is quite specific to DynamoDB. So, it's in your DynamoDB item, your consuming service needs to know that effectively your database schema. Whereas it's not unlike a nice friendly domain event schema.

Jeremy: Right. Yeah. I think the other thing too is that as you mentioned, the ability to process the same message twice, is sometimes something you want to do with Kinesis or with DynamoDB. So, you might want to have multiple consumers reading that same backlog, as opposed to something like SQS, where you can have multiple Lambda functions or parallel Lambda functions, reading those things off the queue. Once you read the message and remove it from the queue, it's gone. It just takes the message out of the queue, whereas those streaming services, like you said, they store the events for a certain amount of time. You can go back and you can read through those things. I mean, again, it depends. When would you suggest somebody use an SQS queue over Kinesis?

Paul: I guess in most cases, SQS would default. If you have a job which needs one-to-one processing, you're only ever going to have one downstream processor, then SQS makes sense. It's proper serverless price and pay-per-use, which unlike Kinesis... Kinesis charges by the hour based on how long it has a shard based pricing model and so you're paying by the hour, rather than per use. In general, I've used, in my applications, for my own and for clients, I have used SQS a lot more than I have Kinesis. I guess if you have that large volume of events coming in, that you need multiple processors, to need multiple subscribers too. Then Kinesis certainly makes sense in that respect.

Jeremy: Makes sense. Let me ask you one more question about developers building asynchronous applications. A lot of things to think about, right? Lot of pitfalls whenever you're building a new type of system, and of course lot of pitfalls in distributed systems. What are the main things that somebody that's now building things asynchronously really has to think about? What are the big ones you have to be worried about?

Paul: There are a few things. Number one, I would say distributed tracing. If you have a multi-step use case where there's a lot of data processing going on in the background, you probably now have multiple log files to search through. There are... If you were doing that synchronously, you could just look in the one place, more often than not. There are strategies around using correlation IDs within each message so that the same, say in CloudWatch you can query on for that correlation ID and get an aggregate of any log entries across your different log groups, which have that correlation ID within it. You need to build that in to your application, your Lambda code, that doesn't come out of the box. Another consideration is testing, writing automated tests. It's just harder for asynchronous workflows, so if you're writing asynchronous... Say I write in Node.js and Jest test frame work. If I have a synchronous Lambda, it's generally pretty easy to write an integration test for that. You just hit the end pointer, invoke the Lambda function, and just verify the response. If you have a multi-step asynchronous data processing workflow, then you need to test each one of those individually, during an actual...Writing an end-to-end test is difficult. But I just like having a wait step, that just waits until background processes have, you hope, have completed and then you can do whatever verification steps you need. Just generally understandability, it's not a thing in itself but it's just for if you've got a new developer on your team... A lot of teams that I've worked with are more full stack web developers which are used monolithic synchronous workflows. Got a new guy on your team and it's just explaining to them how each piece of the pie fits together. That's just going to take time and documentation really is the only solution to that. It's just... Good documentation is important when you've got these asynchronous workflows.

Jeremy: Well I mean, the good news is, that the distributed tracing there's a lot of options out there now to do that but writing good tests and writing good documentation unfortunately that's a challenge, I think, for most organizations. Anyways. Well, thank you so much. Let's leave it there. So Paul, thanks again for being here. Really appreciate you sharing all your serverless knowledge. The amount of stuff that you've been writing is awesome, really enjoy it. New tips every, couple times a week, it's great. How do listeners find out more about you, if they want to subscribe to your newsletter or some of those other things?

Paul: That's great Jeremy and thanks for having me. You can get me on social media, on Twitter, and LinkedIn, @paulswail and you can get my website. It's serverlessfirst.com. My newsletter is there too.

Jeremy: Awesome. Alright, well, I will get all that into the show notes. Thanks again.

Paul: Super. Thank you Jeremy.

THIS EPISODE IS SPONSORED BY: Stackery

View Details

About Eric Johnson

Eric Johnson is a Senior Developer Advocate for serverless at AWS. His passion is to help developers understand and employ best practices for planning and developing event driven, highly scalable applications using serverless technologies. Eric has been a software developer and architect for almost 25 years with a focus on serverless since 2014.

  • Twitter: @edjgeek
  • AWS Compute Blog: https://aws.amazon.com/blogs/compute/
  • AWS HTTP APIs Blog Post: https://aws.amazon.com/blogs/compute/building-better-apis-http-apis-now-generally-available/

About Alan Tan

Alan is a Senior Product Manager at Amazon Web Services working on the API Gateway product team. He was previously a Senior Program Manager at Microsoft and software developer before that. He holds a B.S. in Computer Science and Microbiology & Immunology from The University of British Columbia..

  • Twitter: @t_alan
  • API Gateway Product Page: https://aws.amazon.com/api-gateway/

Transcript

Jeremy: Hi, everyone, I'm Jeremy Daly, and you're listening to Serverless Chats. This week I'm chatting with Eric Johnson and Alan Tan. Hi Eric and Alan, thanks for joining me.

Eric: Hey, thanks for having us.

Alan: Hey, Jeremy, thanks for having us.

Jeremy: So, Eric, let's start with you. You are a senior developer advocate for serverless at AWS, so why don't you tell listeners a little bit about your background, and what it is you do at AWS?

Eric: Yeah, absolutely, my background is I've been a developer since 1995. Yes, that's right. I am an old guy. I've been doing serverless since serverless came out, the day they announced serverless I looked at it and said, "Wow, why would you do anything different?"

I've been following that trail for a long time. I now work for AWS. I've been there for almost two years. I am a senior developer advocate for serverless at AWS. I love all things automated, all things serverless, so I'm having a great time.

As far as our role, what we do, is we're kind of a two-way conversation with developers, and users of serverless, we like to teach on how to serverless and get the message out and the best way to use serverless, how to make it useful for all workloads, but we also listen.

We try to bring back what the developers are saying to the product team, to someone like an Alan Tan, so they know, hey here's what people are saying, here's what they're wanting, here are the paper cuts, here's what they're loving, that kind of thing. I love my role and appreciate you having me here today.

Jeremy: Awesome, all right, so Alan, you are a senior product manager, so why don't you tell the listeners about your background and then what you do as a senior product manager?

Alan: Yes, I started in computer science, a developer just like Eric for a long time, and then I went into product management building products for big data analytics, data analysis, and most recently, been doing this for two years now in the API Gateway team, product manager for API Gateway.

So, what I do, I'm a senior product manager, so I talk to customers directly. I talk to people like Eric. I talk to people like Jeremy, yourself as well, to get the feedback of what customers are really looking for, and then translating that back into the product. So, our customers think we're building the right thing, and they really love coming back to us and using the same product over and over again.

Jeremy: Awesome, all right, well, so you work on the API Gateway team, and just the other day, AWS released HTTP APIs, which is really, really cool. So, can you explain to listeners what that is?

Alan: Yeah, of course, so just to start with some context of where that came from, so when I talked to customers, they usually come back to us for improvements and feature requests, but there are a few things that are really core to products, so more generally, when people think of products and the things they use, whether that's an appliance, a car, a computer, there are certain things they look for.

Which is how can we get this thing faster? How can we get it cheaper and better? API Gateway is no exception to that. Our customers come back asking us the same things, so last year we started looking into how we can continuously deliver on these improvements for them. The results was HTTP APIs.

So, it offers the core functionality of REST API at a 71% lower price, that's at a dollar per million in IAD, a dollar per million requests, sub ten millisecond latency overhead at the P99 level, that's a 60% improvement, and way easier to use features. So, you can think of HTTP APIs as the next generation platform for API Gateway's API types a kind of V2.

Jeremy: Awesome, now Eric, you have done a ton of stuff with API Gateway, with the REST APIs. You had your happy little APIs show on Twitch, kind of getting into all the details. I know you love service integrations and all that kind of stuff that you've been working on.

Eric: Yes, sir.

Jeremy: But you're pretty excited about HTTP APIs as well.

Eric: I am, yeah. For me, as a developer, HTTP API brings a lot. I'm going to start with the better part, the cheaper, faster, those are great, and I love those, and those are huge to our clients, but as a developer, the better part for me is the UI, is the interface.

There's been a lot of work done on how do I, as a developer, interact with API Gateway and how do I develop against API Gateway? It just starts with something like the UI. When you get a look at the UI, and if you've seen it, we did announce it last year at re:Invent.

I've gotten in. You've seen it. Hopefully, you've seen wow, this is a lot simpler. One of the examples that I'll use is CORS. And if you've ever heard me talk about API Gateway, I like to have everybody raise their hand and say, "Who loves CORS?" There's good a surprise, nobody raises their hand except for one person, and he's a liar, so you know that's...

Jeremy: That's probably me....

Eric: Yeah, probably you.

Jeremy: No, I hate CORS as well.

Eric: Nobody loves CORS, and CORS is not easy, but it is a necessary thing, so what a lot of folks end up doing is they just put stars in their CORS. Hey let anybody get to it. Let anything happen. Let anything get to it, because configuring CORS is too complex.

Well with HTTP API, we've taken that, we've simplified that. I say we, but really Alan's team has done a lot of great work on this, but I take credit for it when Alan's not around. So, what we've done is we've made the CORS integration a lot easier to set up, so you can add, here's the domains that should be allowed to get to it.

Here's the methods, and it's all in a simple UI. We've also taken and extended that where we're able to simplify the return coming out of your Lambda or your backend, and we use heuristics, which is a big word that I've just learned recently, but we use some logic to fill in the CORS data and return to the customer.

So, a lot of the heavy lifting of configuring CORS and other items have been simplified, so as a developer, I love it. That's just one example though.

Jeremy: Yeah, and speaking of that, sort of reducing that complexity, another thing that is now available, and I should have said HTTP, while it is hard to say, it's also a great SEO keyword by the way. I think that you're going to find a lot of people finding HTTP APIs.

But it just actually went GA, so you're right, it was in preview for quite some time, but going back to this idea of reducing complexity, one of the things that you have to do with API Gateway is add in authorizers and some of these other things, but now you have like JWT authorizers built right into the service.

Alan: Yeah, that's right, so going on the same vein of things that both you and Eric have been talking about, about reducing complexity, JWT authorizer is one of those things we've heard our customers ask us a lot about from our REST API product.

So, how customers used to do it in our REST API product is that they would use a Lambda authorizer, and it would just write custom code just to do OIDC auth or Oauth2, and every customer has to do this, so what we've done with HTTP APIs JWT authorizer, is to build all this as a native feature inside the product.

So, customers don't have to write code. It's basically code-less, you just configure where your token is being issued from, what audience and what scopes will be part of it, and we will handle all of that for you.

Jeremy: Yeah, and one of the cool things that I've been seeing a lot of articles written about is this ... is using something like Cognito, obviously, as a way to do it, but also being able to use Auth0, I saw somebody write an article about using Google Firestore (correction Firebase), I think, in order to use the authorizer, anything that can process those JWT tokens you can plug that in natively into HTTP APIs.

Alan: Yeah, that's right, so we support the token regardless of where it's being issued from, ping off the Cognito, so it's really being able to ... for the customer to be able to bring any token they have, and being able to use that with API Gateway.

Jeremy: Awesome, all right, so one of the things that I found, sort of experimenting with, with the new HTTP APIs, which is really hard to say by the way, but it's hard to say it fast anyways, HTTP APIs, maybe we just shorten it, but one of the things that I found is obviously the feature set isn't quite ...

You can't do everything in HTTP APIs that you can do in API Gateway, so this is obviously something where you said you've gone back and you've rewritten this thing. Is that something that we should be looking for in the future is to be able to do more things with the service?

Eric: Yeah, with the HTTP API, we're looking to have feature parity with API or the REST API. Now, is that there right now? No, not yet, but it is ... It's top priority. It's looking. We know that as we ... People ask us, "What should I use?"

I think we'll get into this a little later, but, "What should I use?" My rule of thumb is generally start with HTTP API. If it doesn't have the service or the feature that you need, then go to API Gateway, REST API.

However, keep checking, because we are looking to do feature parity to continue to build out the features that are on REST API to build them into HTTP API or something like that, but again, at the better, cheaper, faster way of doing it.

Jeremy: Yeah, I guess I wonder a little bit, because I get that, and that's awesome, the idea of being able to continue to add features. This is what's great about serverless is just things keep getting added, and you don't need to do anything, which is pretty great.

But who is this targeted for? Because I think you see with API Gateway you have a lot of API management controls, right? So, you have everything from quota management to key management and all these other things that are built in there.

It's great if you want to build an API that allows customers to have access to a certain amount of data or a certain quota every day. You can't do that with HTTP APIs yet, so who is this really targeted for?

Alan: So, HTTP APIs is targeted for developers. We wanted to make it the easiest platform for developers to build APIs, and just to add a bit onto what Eric was saying earlier, because we built HTTP APIs from scratch, we are able to start from a clean slate.

So, a lot of these new features, we're going to be adding as part of feature parity. We're actually going back and looking at them and seeing how we can make those a lot easier to use, so it was really targeted for developers who are trying to build APIs.

Jeremy, you touched on API management as well, so this also a set of capabilities that's available in REST API today. For going forward, another request we hear a lot from our customers is they're challenging us to look into the API management space and bring some innovation and disruption there.

As you know, it's a very old space. A lot of our customers, especially when they're moving to AWS, are asking us to make this space easier, so this is a set of new capabilities we're looking into right now, probably later this year you'll see something.

Jeremy: Awesome, yeah, because that's one of those things too where as a developer, who mostly is building just APIs, I just want synchronous invocation of a Lambda to function is essentially what I want.

Adding on the CORS, and having to put extra stuff in the CloudFormation, and set up those response templates, is such a pain, and then you've got all that extra overhead of things that you just don't need, so I really, really love how this is super slimmed down. And once you get a couple more of those core features in, I think that's going to be pretty exciting stuff.

Eric: Yeah, it's interesting, as we talk to developers, and how many are saying, "Look, we just want to ... We need to proxy a Lambda or we need to do this. Now a lot of times, if you've seen anything I talked about, I talk about okay but are you looking at the full feature set?

Are you re-inventing the wheel? But sometimes they're not, but sometimes they're just look we just need a simple authentication and a proxy or an API in front of a Lambda or in front of a service. And that's HTTP API at the moment. I mean it's a great way to just do that very quickly. And then as we start, you're right, as we start to build out those features, then it's going to become for everybody.

Jeremy: Awesome, all right, so let's jump into some of the features that it has right now, because I think that'd be useful to know what exactly we can do with it, so basic features we have routes, right?

Eric: Yeah, with routes, yeah, so the way we've done ... and we've changed that a little bit, a route is just simply an HTTP API. It's simply a path and a method, right? So, look, I'm going to hit the route, and it's going to be a get a post, an any or something like that. So, routes are very simple. And then you can tie to that route. You can tie different integrations to different places going on, but yes, you have the basic concept of a route.

Jeremy: What are the integrations that exist right now?

Eric: So, currently, we have Lambda. And I'm ticking them off on my fingers, so you can't see it, but we do have Lambda. We also have an HTTP proxy, so you can front another HTTP endpoint. And then we're adding private integration. I want to expand on that, because I'm really excited about that, but those are the three integrations that we can do.

Jeremy: All right, so before we get into the integrations, because I know that's a whole ... that might take us an hour to get through.

Eric: Yeah.

Jeremy: We also have stages, similar to API Gateway, right?

Eric: Right.

Alan: Yup.

Jeremy: All right, so you can still do your dev and your prod and things like that.

Alan: That's right, and we also added a new usability feature for stages actually. We added the concept of auto-deployment, so as you change an API, your changes will automatically be deployed to the stage. This is also a big area of feedback we got from our customers, where as Eric was talking about earlier, right, some of our customers are really just looking for a basic endpoint.

They're just really looking for something to proxy traffic, and for them, having to learn more and more API Gateway concepts might not be the most customer friendly, right? So, we added auto-deployment, so that's another concept you don't have to learn about until you really need it.

Eric: Yeah, I can't tell you how many times I've been standing on stage demonstrating API Gateway, did all the changes, and then spent two minutes going, why doesn't this work?

Jeremy: Because you forgot to deploy?

Eric: Exactly, and I have actually said to the audience, "Hey, if I start getting frustrated this doesn't work, somebody yell, 'Deploy it,'" So, I'm really excited about the auto-deploy. I'm not going to lie.

Jeremy: Yeah, and actually, I think that it's subtle, but it is a really cool feature, because if you think about your average developer, who just again wants to put an HTTP endpoint in front of a Lambda function, for example.

The problem is that again you're adding on all this complexity with API Gateway, and it's just another step that you have to go through, and so having that all kind of wrap into one I think is pretty cool.

Now I personally don't like when every Lambda function has an HTTP Gateway in front of it, similar to maybe Google Cloud functions or something like that, where those all have their own endpoint, but I do love this, because it's a super simple way to kind of add that on top. So, what about logging?

Eric: Yup, we have access logging. Alan, go ahead.

Alan: Okay, yeah, we have access logging, just like REST API. We actually added a new context variable to our access log, which is integration error message, so you're able to see why you're backend is not working, for example, if you're using with Lambda it'll tell you things like hey your permissions weren't set up correctly. So, you're able to easily debug using our access logs now as well.

Jeremy: Which again is another huge, huge improvement, right? When you keep getting that internal server error coming back, and you have no idea why, and then you realize oh wait a minute my Lambda function processed correctly, but somehow my API Gateway didn't respond correctly or vice versa.

So, that's definitely a cool feature. All right, so let's get into private integrations, because this is super cool stuff, so I don't know, Eric, you want to start?

Eric: I do. I do. I'm excited. Yeah, so private integrations, a lot of times we get folks that are telling us look we've got ... hey. I know the three of us are serverless people, but apparently there are non-serverless architectures out there.

Jeremy: I don't believe it. I don't believe it.

Eric: I don't either, but that's what they're telling me, so they said, "Talk about it." No, so when we're looking at API Gateway, we deal with VPCs. When you're dealing with especially with non-serverless architectures, you deal with VPCs, and that's a good thing.

VPCs are a layer of security that, as in a past life, as a solutions architect, we talk about security, and everything goes into VPC. That's what we always hear. That's why sometimes with serverless they kind of say well it doesn't have to be, because it's already in a VPC, but that's another discussion.

But so dealing with VPCs, when you're thinking about EC2s or you're thinking about container servers, like ECS or EKS, you need a way to be able to talk, to basically front those with API Gateway. Now you can do it through some other ways, like ELB or ALB, but you want to be able to take advantage of the API Gateway features that we've got, like the JWT or the throttling that we've added, different things like that.

So, we've built private integrations, and I think we really did a great job on this, and Alan's team did an amazing job on this, because it's not just NLB like REST does, which is a great way, look, we can point to an NLB, and then we can get your stuff in a VPC, but we also added the ability to use ALB.

Which you can do path mapping. It's a level 4 versus a level 7. I don't know what that means, but there it is. It's on how the routing is done. Alan will expand on that, but the other one that I'm really excited about is AWS Cloud Map.

And if you're not familiar with Cloud Map, it gives you the ability to ... basically it's a service discovery product, so I can actually create a name space, that's going to maybe ... Maybe I'll just tie ... these are you can create a name space on however I want to be, or we can say look I've got an application that does this, I'm going to create a name space that matches that.

In the name space, I can make different services, so for instance, I may have a tier one or a tier two, or my different tiers, however, I want to break my services apart. And then finally, in the name, or in the service, I can register service instances.

So, let's take, for example, if I'm doing ECS, so I being up a cluster, it's got four different machines in my cluster, so I can register those instances with Cloud Map, so API Gateway can use Cloud Map to get to those instances, to connect to the clusters.

And so I've got a secure infrastructure, so I'm saying look API Gateway, for all intents and purposes, is in that VPC with those instances. I'm probably making Alan cringe right now, because it's not actually in there, but we're using technology to connect those.

So, with security groups, when I'm locking down my security groups, I can say hey, everything in this VPC security group, and that API Gateway, can be in there, because of Cloud Map and because of the way it connects through.

So, and Alan feel free to correct any of that if you want, because you're obviously the big brain here, but that's ... I'm excited about the way we're doing the private integrations and the ability to build those. With Cloud Map, you can iterate fast. You can build them fast. You can change them out fast. It gives you serverless flexibility in non-serverless world.

Alan: That's right Eric. I think you got it perfectly. I think it's totally fine what you said, no corrections needed.

Eric: Please do not edit that out Jeremy. That needs to stay in.

Alan: Okay, I just wanted to add on that while the Cloud Map capability is super cool, the ability to connect to a private ALB, actually that's one of our top feature request we've heard from customers from our REST API product, so it's going to be super great for customers, even if they continue using private ALBs.

Jeremy: So, what exactly is the use case for that?

Alan: So, a lot of our customers, when they're using ECS, because all the tutorials are focused around using ALBs, and ALBs, generally, itself, is easier to use than NLBs, because, like Eric said, it is a layer 7 product, so it has more features, more routing options. It's easier to set up. So, a lot of ECS customers just have a stronger affinity to use private ALBs instead of NLBs.

Jeremy: Awesome, all right, so what about migration, because now let's say I've got, I don't know, 200 API Gateways or something like that, that I probably have configured for a number of customers and other people, how do I go from API Gateway to HTTP APIs?

Alan: So, to go from API Gateway's REST APIs to HTTP APIs is actually pretty simple, because they both support the OpenAPI spec, right? You can go to REST API, export your API using OpenAPI3, and then import it using HTTP API, right?

But as we mentioned earlier, or at least you alluded to it, there are some features that are missing, so when you export from REST API and import from HTTP API, when you do that, if you look at the object, that gets returned, we actually tell you exactly which features are missing, so you can actually use that as an indicator for whether or not if there are any updates you need to make, what you should be aware of, so you can take a look at the info object that's returned.

Another thing we've done to help customers to move REST to HTTP is that we've made customer domains cross compatible, so with one custom domain name you can put a REST API, base path mapping and an HTTP API base path mapping.

So you can take a look at your current API, if it's REST API, bring it over to HTTP API, test that it works. And if it works, you can just swap the base path mapping under that custom domain, so your clients don't have to change any URLs. You don't have to update any URLs in your documentation or anything like that. Everything will just work.

Jeremy: Awesome, and I know that this is more about specifically HTTP APIs, but I do want to touch on custom domains for a second, because I really like this pattern, and maybe Eric, you can talk about it a little bit more, but I find this ability to create new services, so microservices, my billing service, or my alerting service, or whatever it is.

Create a separate group of functions and resources and things like that, front that with an API Gateway, and then use the custom domain mapping so now that it's myapi.com/billing, myapi.com/alerting, or whatever that is. And that gives me a lot of control, because now I can add endpoints, remove endpoints, add new services. I'm not worrying about a single, or a bunch of teams trying to publish to a single API Gateway and map those all over the place.

So, that's a really, really cool feature, so Eric, I mean this is just great, because now I can say look maybe I have one service that needs a service integration or it needs one of these features of API Gateway REST APIs, but if my other 10 services just need simple proxying, then I can use the HTTP APIs for that.

Eric: Yeah, no, I love it too. When you think about that ... We get a lot of questions on how do I architect serverless architectures or how should it look? Are all 300 Lambdas in one application, so they can be behind it? We say no.

We really encourage folks to find a logical breaking point for serverless applications or applications in general. It really is not just a serverless thing. It's anything. It's so being able to take and say okay I'm going to have ...

Instead of one large application, I'm going to have 10 micro-applications, and each one of those are going to be micro-service architectures, so then I can put one base path mapping right in front of it, and then I can just route that around as I need, rather than having to okay I'm going to try to do some funky routing.

I'm going to put an API Gateway in and sub API Gateways, and you can do that, you can certainly do that, especially if you're trying to maybe unify or consolidate your authentication, but base path mapping is a simple way to say you go there, you go there, you go there, you have a lot of control and it's another thing. And I don't know if you want to touch on this now or later, but it's also really helpful when it comes to the strangler pattern.

Jeremy: I was just going to say that.

Eric: Okay.

Jeremy: I was just going to say that, because I think it's a new way. It's a new way to think about the strangler pattern. Obviously ALB was a great way to do it, because you could route things right to your Lambda functions and start splitting off, like you said, that layer 7, start splitting them off at the path level, which is pretty cool stuff.

But using an API Gateway to do that, it was just expensive. There was a lot of overhead and things like that, but now, with the HTTP proxying in HTTP APIs, which is pretty inexpensive, and the custom domain capabilities, now you've got some more options there.

Eric: You do. Going forward, HTTP API will be referred to as API unless otherwise designated. Does that help?

Jeremy: Is that the official AWS stance?

Eric: Nope, that's the Eric Johnson stance who can barely speak English as it is. All right, no, so yeah, and that's an interesting idea, if you're not familiar with the strangler pattern, just so you know, the idea is let's say I've got an old legacy application that's a bunch of different services all in one.

I can actually put ... so, we'll talk about API Gateway for a minute. I can put API Gateway in front of it, and then as I start breaking out those little, little services, I can say okay now I'm going to reroute to serverless or some other type of architecture. Which should be serverless.

And I'm going to break those out, until eventually I've strangled that application until it's gone, so everything's broken out. But if you wanted to you could say I'm going to take my application, and I'm going to ... instead of breaking it into single service, I'm going to break it down into maybe 10 micro applications.

I'm going to say to my teams, I'm going to take 10 teams, I'm going to say build these 10 applications, and when you're ready, we'll just put a custom path, a custom domain path in front of it, and we'll route to your API Gateway.

And it could be a REST API if you need it. It could be a HTTP API if you need it. I have the ability to send to either and change as needed, so yeah, it's kind of a level up of the strangler pattern.

Jeremy: Yeah, that's awesome. All right, so let's talk about the practical use cases of this now. So, if I'm a developer, and like you said, Eric, first thing you do is use HTTP APIs. That'll be your defacto standard unless you need something else. But what are some of the reasons why I would use API Gateway REST APIs or maybe something like ALB over the new HTTP APIs?

Alan: So, between REST and HTTP API, I think if you're, as Eric was saying earlier, if you're bringing in new traffic, you should take a look at HTTP API and see whether that's a good fit for you, because like I said, it's going to be the next generation platform for us to build our API improvements.

If there are any feature gaps or feature parity type things, as I mentioned earlier, we're also going to bring those over, over time, so you shouldn't let any feature gaps worry you about whether or not we're going to support in the feature. We definitely are.

But going to ALB and API Gateway now, they're ... while there are some overlapping features, they're really quite different, so I think really it's similar to earlier the rule of thumb is just to pick which ever one meets your needs, and where you think your needs are going to be in the future.

So, for example, API Gateway offers throttling, authorization, publishing APIs, monetizing APIs, serverless web sockets, so you don't need to worry about these kinds of things in your backend code.

So, going back to the strangler pattern, for example, if you're using a Lambda function, and if you're using your legacy on-prem backends, for example, and there's some authorization authentication you want to add for all these APIs, API Gateway is a really good place to put that.

Because then you don't really need to worry about having an authorization library in your Lambda code and an authorization library in your backend code, and trying to keep that consistent over time. Because consistency is really hard, especially when you're talking about doing code manually, right?

Jeremy: Yeah.

Alan: So, that's the kind of the benefit that you get with API Gateway, you get a consistent layer, where you can put some of these logic in. And also being serverless, if you're not serving any traffic, you're also not being charged, so it's really good for highly variable workloads, where you need to scale up and down, so you don't need to maintain a minimum footprint.

ALB on the other hand, works really well if you have already ... if you already have all that logic, that's already handled, so you're already handling throttling, caching, authorization on your backend code, and you don't really care about having to maintain that over time, so it's really good if you already have those things.

The pricing model works really great if you have a lot of small requests, and you already serve like a very consistent level of traffic, because how ALB works is there is a base price for keeping the load balancer around, so it's really good if you already have a minimal footprint for your traffic.

Eric: I would add in also, and I completely agree with Alan, but a lot of times when we have the discussions, and I'm in the field, it's the, "I don't use any of the features on API Gateway, so why do I need to use it?" That's fair.

If you really are not, okay, that's something to look at, and maybe ALB works out, but what I encourage people is you ask them okay are you handling caching? And we get, "Yeah, yeah, we're writing to elastic cache from Lambda." I was like, "Okay, well can ... I can help you not do that, and just offload."

I really encourage people to figure out what are they doing in the Lambdas or the backend integrations that can be rolled off into API Gateway, because then you start ... I know people hear this a lot, but you start looking at the total cost, the TCO, the total cost of operating it, right?

So, and you look at how much code are we writing to do stuff that's already available? I'm a big fan of less code. If I build an application, the most dangerous part of an application is my code. There's just no two ways about it.

And so my encouragement is if you don't have to code it, and you have a service, and sometimes you look at that price factor, but you have to look at the price factor in light of how much developer time, how much operational time, how much management time is going into putting it into your code, versus rolling it out.

Especially with the drop in cost with the HTTP APIs, when you're paying at a third of the cost, as we reach feature parity, I just think it's the way to go. Like I said, rule of thumb, that's the first place I go.

Jeremy: Yeah, no, I think you hit the nail on the head there. The biggest thing for me too is why rebuild something if it already exists? It has to exist in a way that works for you, right? Obviously, caching in API Gateway in the REST API version is very good for the right use case, right? You can't always cache everything.

Eric: Right.

Jeremy: But same thing with request transformations, that's a really cool feature that I don't think a lot of people use, but you probably should if you want to limit... You can have the same Lambda function that responds to requests, and actually spit them back in different ways using different endpoints, using request transformations and things like that.

And then one thing too about authorizer, so custom authorizers, love that feature when it came out, but now ... with API Gateway, but honestly if you think you can write a better standard than JWT or oAuth2 or something like that, good for you, go give it a try, but honestly-

Eric: Good on you, yeah.

Jeremy: I think it's probably smarter, when it comes to security, to maybe use something that is a pretty good standard, like you said, there are a couple of things, like the usage plans and API key management that you can't do in HTTP API right now. There's also no x-ray integration yet, right?

Alan: That's right, yeah. That's also coming really very soon.

Jeremy: Awesome, all right. And then just some other things too. I have some notes here, just because I like to be thorough, but things like client certs, WAF, and resource policies, so those aren't supported, right?

Alan: Yeah, those aren't supported right now, but it's part of the feature parity bucket of work, so that's the buck of work we're going to be looking into bringing over and seeing how we can do those better.

Jeremy: Awesome, all right, so one more question about ALB though just in terms of ... in the scale up, so if you need to scale very quickly, is it ALB, is it HTTP APIs, is it API Gateway, what handle a fast scale-up?

Alan: That would be API Gateway, so API Gateway, for both REST and HTTP APIs, and even web socket, actually, your account is actually provisioned for 10,000 RPS, and you can scale up from zero to 10,000 as quickly as you want.

And if you need even higher than 10,000, we can easily increase that for you. We've gone with customers significantly higher than that, like way higher for sporting events and big announcement type deals, right?

Jeremy: All right, yeah, so that scaling is pretty intense. Again, I like ... I always like to say, if you have to scale that much then you've got some really good things going on with your application, so that's good stuff.

One of the other things that HTTP APIs doesn't do yet. I don't want to harp on this fact, because I know we've talked about feature parity, and I know the teams are working on it. I think that that ... the fact that it's built in a way that you started from scratch, like it's a new product all together ...

And as you mentioned, it's going to give a ton of flexibility going forward, but with the existing service integrations in API Gateway, very common pattern, and this is something, Eric, you talk about quite a bit, is this idea of taking data right from the request and putting it right into some service ...

Whether that's SQS or it's Kinesis or it's DynamoDB, being able to handle those service integrations, and I think you call it storage first, which I think is really-

Eric: I do.

Jeremy: ... interesting, an interesting way to think about it, so I know that doesn't exist yet. I know it's going to come at some point, but let's talk about that pattern, because I think that's really interesting feature of using something like API Gateway or the REST APIs and eventually HTTP APIs that lets you take that request, forget about the Lambda function, just make sure you capture it, store that data and then do something with it.

Eric: Yeah, and I do call it storage first, or sometimes I do a presentation called Thinking Asynchronously, but the idea ... Let me go back to my earlier statement. The most dangerous part of an application that I'll ever build is my code. Right?

So, when I build an application, I want to get that data stored first. That's the thing. I tend to go DynamoDB because that's what I like, that's what I use, but there's different purposes.

I know Jeremy, you and I have had this conversation before, and you're an SQS guy, so that's where you tend to go, and we do this because we look at okay what's the pattern for the retry or the DLQ or different things like that.

For me it's because I'm going to continually write back to Dynamo. On the app, I'm specifically thinking about it. But the idea is if API Gateway can directly integrate with the storage, be it S3, be it DynamoDB, something like that, then I've stored the data and I don't have to go back to the customer if my logic fails, right?

So, in an application I've stored the data, let's say I'm using DynamoDB, I do a stream, it triggers a Lambda, I start processing that data. If somewhere in there, something breaks, and again, it's going to be my code, but let's say something breaks, then I don't have to go back to the customer and say hey guess what, I blew it.

Can you give me your data again? Can you resubmit that? And continue to trust me, because I'm sure I won't lose it again. Instead, I've got that data stored, and I can write in some retry or take advantage of the retry from an SQS or an SNS or something like that.

So, I think it's a really cool pattern for building resilience into our application. Serverless comes with a lot of resilience anyway, that's how AWS has approached this on look as much as we'd like to say nothing ever breaks, let's write as if it does, right?

So, let's degrade gracefully. I think this adds even another layer of that, where I can degrade in my code and know hey I've still got the data. I can write some retry logic. I can use existing retry logic. I think it's a safer pattern.

It does require ... The storage first is the pattern I call it, but it requires asynchronously. What can I do after I've responded to the client and how do I work with them?

Jeremy: Absolutely, awesome, so Alan what do you think about that?

Alan: I think it's super powerful, and we see a lot of our customers doing that actually, especially after Eric published his blog that talks about this thinking, and showing a fully functional demo about creating ... what was it, like a ...?

Eric: A URL shortener.

Alan: A URL shortener, yeah, that's right.

Eric: I'm glad it impacted you a lot.

Alan: Yes.

Eric: Just kidding.

Alan: And as we look at HTTP APIs, we know actually setting up service integrations in REST APIs is actually kind of a pain, right? In HTTP APIs we're going to make it super simple, just as easy as it is to set up a Lambda integration today, so some background there.

In REST API if you set up a Lambda integration, super easy, paste in the ARN and you're done. If you try to set up anything else, like DynamoDB, SQS, there's a lot more setup steps, especially in writing a little transformation template.

Eric: VTL.

Alan: We're going to make all those super easy in HTTP API, so just as easy if not easier than Lambda.

Jeremy: That's awesome, all right, well listen-

Eric: That's me.

Jeremy: Yeah, no, the service integrations are not easy, which is why I typically use a plug-in to help me do those, because it's just a lot easier. But anyways, so listen, both of you, thank you so much for coming on and talking about this stuff.

Alan, honestly, great work on HTTP APIs. I think this is going to be, like Eric said, this should probably be the first thing you look at, once we get those service integrations in there, which are going to be easier, I love that idea, because that's the other thing... Learn from previous mistakes, maybe they're not mistakes, but certainly make it less friction to get that stuff in there, so that's going to be absolutely awesome, and I look forward to maybe checking back in with you two, and see when these new things come out.

So, if people do want to find out more about HTTP APIs, they can just go to the AWS site and click on API Gateway and there's information there. If they want to get in touch with you, Eric, how do they do that?

Eric: The best way is on Twitter, @EDJGeek is the best way to hit me. My DMs are open, and I'm always getting questions on that, so that's the best way to get me.

Jeremy: All right, and Alan if people want to bug you with feature requests, how do they do that?

Alan: They can also reach me on Twitter, @T_Alan. Eric's been bugging me a lot about getting a Twitter account, so I did that.

Eric: That's right.

Alan: And now people can reach out to me.

Jeremy: Awesome, and then you publish a lot on the Compute Blog, right, the ...

Eric: Yeah.

Jeremy: It's aws.amazon.com/blog/compute. There's like what, 17,000 blog posts a day across all of AWS or something like that? So if you can-

Eric: All by me.

Jeremy: Anyways, all right, well listen, I will get all of that stuff into the show notes, thanks again for being here.

Eric: Yeah, hey, thanks for having us.

Alan: Thank you Jeremy, thank you again.

THIS EPISODE IS SPONSORED BY: Epsagon

View Details

About Lynn Langit

Lynn Langit is a Cloud Architect who codes. She's a Cloud and Big Data Architect, AWS Community Hero, Google Cloud Developer Expert, and Microsoft Azure Insider. She has a wealth of cloud training courses on Lynda.com. Lynn is currently working on Cloud-based bioinformatics projects.

  • Twitter: @LynnLangit
  • Site: LynnLangit.com
  • Courses: https://www.linkedin.com/learning/instructors/lynn-langit
  • GCP for Bioinformatics: https://github.com/lynnlangit/gcp-for-bioinformatics

Mentioned Articles:

  • Genome Engineering Applications: Early Adopters of the Cloud by Jeff Barr
  • Scaling Custom Machine Learning on AWS
  • Scaling Custom Machine Learning on AWS — Part 2 EMR
  • Scaling Custom Machine Learning on AWS — Part 3 Kubernetes
  • Shopping with DNA
  • Learn | Build | Teach

Transcript
Jeremy: Hi everyone I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Lynn Langit. Hi Lynn. Thanks for joining me.

Lynn: Hi. Thanks for inviting me.

Jeremy: So you refer to yourself as a coding cloud architect. You're also an author and an instructor. So why don't you tell the listeners a little bit about yourself and what you've been up to lately?

Lynn: Sure. I run my own consulting company. I've done so for eight years now and I work on various projects on the cloud. Most recently I've been doing most of my work on GCP because that's what my customers are interested in. But I've done production work on AWS and Azure. And I've actually done some POCs now on Alibaba Cloud. So one of the characteristics of me and my team is that we work on whichever clouds best serve our customers, which makes work really fun. In terms of the work that we do it really depends on what the customer needs because I have this ability to work in multi-cloud. Sometimes it's me working with C levels or senior technical people helping them to make technology choices, so based on their particular vertical. But at other times I'll hire a team of subcontractors for a particular project and we might build a POC. We might actually build all the way to MVP for a customer.

Lynn: And then occasionally I take projects where I build all the way out. The longest one I've had over the past few years is I did a project for 14 months where we went from design all the way out to product. And I worked every single day I was embedded with the developer team. So I do everything from design to coding to testing. It's a fun life.

Jeremy: It sounds like it. Well, so listen, I have been following you for a very long time and I'm a huge fan of the work that you've done. I've watched some of your videos on LinkedIn Learning and just been following along with some of this other stuff that you've done. And really like you said, a lot of what you have done has been around big data and recently you've been getting into, or you have gotten into, big data and serverless. And that's really what I'd love to talk to you about today because I just find big data to be absolutely fascinating and just the volume of data that we are collecting nowadays is absolutely insane. It's overwhelming.

And I don't know if traditional systems or if especially smaller teams working on some of these specialty products have the capability or the resources to keep up with the amount of data that's coming in based off of sort of some of these traditional methods to do that. So we can get into all of that. And I have a feeling this discussion will go all over the place, which is awesome. But maybe we could start just by sort of level setting the audience and just explaining what big data is or I think maybe what you mean by big data.

Lynn: I can have a really simple explanation. I'll say the explanation and I'll tell you why. So the explanation is data of a size that doesn't function effectively in your SQL Server or your Oracle Server or your data warehouse, so your traditional systems. And the reason I say this is because that is my professional background. I've been doing this for about 20 years now and for the first five or so maybe seven, I was working in those systems. I've actually written three books on SQL Server data warehousing. I worked for Microsoft as a developer evangelist back in 2007 to 2011. And the consulting practice that I built initially was around optimization of relational database systems.

So I was literally working on systems and figuring out, oh, this could be optimized. Let's optimize it. Oh, whoops, we have too much data now, what do we do? So when I left Microsoft in 2011 to launch my consultancy, I left because I was so fascinated by what was coming beyond these systems. One of the impetus was the launching of Hadoop as an open source project. And literally when I left Microsoft, I went to New Jersey and I took a class with Hadoop Developers, which was really throwing me in the deep end because I had come out of the Windows ecosystem. Of course the class was on Linux in Java, all coding. And I learned a lot that week.

Jeremy: I can imagine, yeah. So that's maybe my question there. So big data is this volume of data, this immense amount of data that's coming in that I think as you put it, that sort of these traditional systems like a SQL Server or even an Oracle can't handle or at least can't handle at a scale that would make the processing easy. So you mentioned Hadoop and there's other things like Redshift is now a popular choice for sort of data warehousing. And then you've got Snowflake and Tableau and some of these other things I think that are ... products out there that are trying to find a way to analyze this data. But what is the problem with these traditional systems when it comes to this massive amount of data?

Lynn: Well, it goes to the CAP theorem, which is consistency, availability and partitioning. This is sort of classic database ... what are the capabilities of a database? And it's really kind of common sense. A database can have two of the three but not all three. So you can have basically the ability for transactions which is relational databases or you can have the ability to add partitions is really kind of to simplify it easily. Because if you think about it, when you're adding partitions, you're adding redundancy. It's a trade off. And so are you adding partitions for scalability? And so when adding partitions makes a relational database too slow, then what do you do? So what you then do is you partition the data in the database to SQL and NoSQL.

And again, I did a whole bunch of work back in 2011, 2012, 2013. I worked with MongoDB, I worked with Redis. And one of the sort of key, I don't know, books I guess, would be Seven Databases in Seven Weeks. It's still very valid book even though it's many years old. It tells how you do that progression and really turn the light on for me, because prior to that point it was, oh, just scale out your SQL Server, scale out your Oracle Server, which still would work but these NoSQL databases were providing much more economical alternatives. And of course I'm always trying to provide the best value to my customer. So if it wasn't a great value to buy more licenses for SQL Server or for Oracle, rather you want to get a Mongo Cluster up or a Redis Cluster up, you could partition your data if that was possible because there's cost to partitioning your data and writing your application.

So I just found those trade offs really, really fascinating. And of course during that time, cloud was launched, led by AWS. Microsoft had an offering, but they didn't really understand the market until a little bit later. So Amazon had an offering and they first started, it was really interesting. They started by just lift and shift with RDS at a PaaS level taking SQL Server and actually making it run effectively in the cloud. That was how I got started, because my customers wanted to lift and shift and maybe go to an enterprise edition and run it on cloud scale servers.

And ironically because they're kind of co located some former Microsoft employees who were kind of frustrated at that time with what Microsoft at that time was doing with SQL Server in the cloud went over to Amazon. Notoriously right after I left at SQL PasS Summit presented on Amazon SQL Server RDS. And PaaS Summit was kind of a Microsoft centered event and the Amazon people came over. And because of that, I kind of to this day have a pretty good relationship with the Amazon data services teams.

Jeremy: So then this idea of moving away from those systems, so we have, you mentioned NoSQL, or NoSQL. So we have a DynamoDB in the AWS ecosystem, you have Mongo, you have Cassandra, some of these other things. And all of them, maybe less so I guess with DynamoDB, but had some scaling problems sort of built into them. But this is something where I think you started looking at serverless tools to try to handle that big data. Now that thing is like data lakes and S3 or something like maybe BigQuery or Cosmos DB, what tools are you moving to, now to handle that scale of data?

Lynn: Well, change comes slowly and change is usually induced by some sort of pain. And so the pain in my case and my customer's case was through IoT data because IoT data increased the amount of data exponentially because the event based data. So I had some customers, some of the big, big like the biggest appliance manufacturer in the United States. Customers, I can't name, but you can guess who they are. And this was maybe eight years ago, so it was still a while ago. They wanted to IoT enable their devices.

And again, to be very clear, the majority of the enterprise applications that I would work with would be SQL plus NoSQL because they would have a need for transaction. And again, that's really important because I saw those startups go just directly to NoSQL and then they would call me and they would try to tune their transactional consistency of their Mongo and it would be clustered and it would be a mess. And then we just pull that out and put it in MySQL. Just the whole space was super interesting. So meanwhile the cloud vendors are evolving and Amazon of course comes with DynamoDB. And I have to tell you that initially I was super resistant. I was like, how do you even query that? I actually did some time tests and blogged about ... this is like seven, eight years ago.

You write SQL query, everybody knows how to do that. You write a Dynamo query, it takes 15 minutes because you have to research the query and how much is that in your dev time and dah, dah, dah, dah, dah, dah, dah, dah. So there was resistance including me. The service on the cloud that really stunned me and still does is BigQuery because BigQuery offered SQL querying, which I think it's extremely important when you're evaluating different kinds of database solutions to look at what is the ramp up time to understand how to get data in, how to take data out. And the more different the query languages are, the more errors you're going to have too, and this is your data. So I've seen a lot of bad things where developers overestimated their abilities. And because the query languages were at really idiosyncratic or esoteric for the NoSQL databases, it was all kinds of problems.

But BigQuery's idea of, okay, you get around the scaling problem by using files and you then just use SQL and you just pay for the time. I mean, I literally, I got goosebumps. I was stunned when BigQuery came out. I was stunned. I really got it from the beginning. And I've written about it, I've used it. I would often add it for customers as sort of a incremental rather than NoSQL. I called it kind of NoSQL going all the way back to SQL. Right?

Jeremy: Yeah.

Lynn: But to this day, there's still a lot of people that you show them BigQuery and they just don't believe it because you go to the query window as you know and you say ... Well, now it's more common since Amazon has Athena and I don't know, Azure has something like that. But even two, three years ago I was at a serverless conference and I was doing a talk on serverless SQL, and at the end, I do live queries on BigQuery on terabytes and I explain how it just costs pennies and as you obviously know, there's no service set up, there's no scaling, there's no clustering, there's no whatever. And people just wouldn't believe it.

Jeremy: Right, yeah.

Lynn: I was like, really. So I think it's really interesting that BigQuery, the biggest problem for a long time was that once you gave it to customers, what would happen is they would spend too much money, which means they'd love the product. And it was so interesting interacting with Google and saying, "Okay, I know you use this internally and so you guys aren't ... you're really used to cost controls and stuff, but customers need to understand what they're going to be spending. So you need to put cost estimates and cost caps and all that." And frankly that's just been coming in the last couple of years for BigQuery. And I think that that really hindered adoption, really a lot.

And it's interesting to see this pressure between cloud vendors because Google on some of their other services, like their VMs, they'll put their pricing and now that's pushing Amazon right in the console. So when you're setting it up you can click and size it and you get a more CPU use or whatever, you can see how much it will be. And I noticed, because I just did a refresh to one of my Amazon data courses for LinkedIn Learning, that Amazon is now putting this in the console. And I think that's really great. I love that.

Jeremy: Yeah. I mean, one of the things that I think is a good criticism or is a common criticism of services that are pay as you go is this idea that you don't really know how much you're going to spend until you start using it. And then you get all these benefits of not paying for maintaining some level of scale so that you have that availability. But then if you do sort of start using it heavily, then those prices go up. And as someone who is deep into the AWS ecosystem, I've used Athena quite a bit and when it first came out and I first started using it, I was very nervous thinking to myself like, wow, if you're searching through a terabyte of data and it's going to cost you $5, how many times do you run these queries and so forth.

But one of the things that I found with that, and I think it's very similar with BigQuery, is that these are all very, very optimized the way that they search. So you can limit them by date and you can limit them different ways. So you're not necessarily searching through terabytes and terabytes of data every time, especially if your queries are optimized in the right way.

Lynn: Yes. But you have to know how to do that.

Jeremy: True.

Lynn: So what I have seen now that I'm working with customers with enormous amounts of data, this actually happened and it wasn't related to me showing them BigQuery, but this actually happened at one of these customers. They got very excited and they got a BigQuery bill of $83,000 in one day.

Jeremy: Wow.

Lynn: Now this customer, because they're one of the biggest users of GCP in the world, they have the relationship where you can kind of have ... this kind of thing can be forgiven occasionally. But the greater point, and this is very germane to serverless services in general, is that there is a sliding scale. So let's take Amazon. Lambda basically costs you nothing until you're Spotify level, but what does cost you? Well, Dynamo costs you.

Jeremy: Yes.

Lynn: Well, Athena costs you. And so it's an entirely new pricing model that I find my customers are just utterly confused by. And so part of my advisory is when customers are moving into the serverless platforms is they have to repurpose some of their team members to work with cost estimates and cost management. And if they don't do that from the beginning, then they're not going to get the level of value from the serverless services that could be had. That is, it doesn't have to be a full time role, but it has to be a role. I can't tell you the number of times that I've had these dev teams, oh, well we don't say the bill. They call us if we go over a certain amount. I said, "That's just irresponsible."

Jeremy: Yes.

Lynn: In the old box software on-prem days you had licensing specialists, so what are these people doing now? In fact, again, I don't mean to be constantly plugging my courses here, but I try to make courses around needs in the industry. So I made a course called AWS Cost Control. And I often recommend it to students that reach out to me who maybe come from a finance background or come from some other background that are moving into cloud computing because I say this is a skill area that I find not covered. I think it really actually hurts the adoption of serverless actually because you sort of get the, oh, I got the shock Dynamo bill or I got the shock BigQuery bill. And again, this is just part of using the serverless services properly.

Jeremy: Right. Yeah. And actually you mentioned the cost, knowing how to use BigQuery or Athena correctly to optimize for those costs. The same thing is true with DynamoDB. I mean, I think a lot of people use DynamoDB in a way that makes it very, very expensive as opposed to storing data in different ways and having the indexes optimized the right way and doing queries instead of scans and other ways that they can optimize that. All right. So let's talk a little bit about sort of where some of this big data is going because I think we've got some serverless solutions and I guess we could get into some of that a little bit more. But I'm really interested in this sort of rapid increase in data, you mentioned IoT, but you're seeing a really, really big growth in data in a different sort of industry now. Can you explain that a bit?

Lynn: Sure. About three years ago now I had a couple of things in my life that changed my professional life. The first was my then 17 year old daughter went to Stanford Summer Program and she was at that time interested in bioinformatics. So she got enrolled in regular Stanford level classes and she was doing some bioinformatics. And I was thinking, oh, big data I'm really interested. I've heard this is a lot of data. And she came home and she said, "Mom, these people don't use the cloud." And I said, "What?" And I said, "Aren't you using Docker?" And she, "No, no, we have to make tiny data sets on our laptop or we have to SSH into the mainframe." And I was like, "Are you joking?" This is Stanford, so I was shocked. So that was the first thing.

The second thing was a very close friend of mine got breast cancer and because of some other work I had done in the bioinformatics community in San Diego, I knew there were lots of immunotherapies that were becoming available and she was not able at that time to get her tumor sequenced or even participate. And she had a very, I would almost call it dehumanizing course of treatment. And she recovered, she completely recovered, but it was just unnecessarily horrific. And at the same time, the third thing, Google came to me and said, "As a Google a partner, we have this interesting data challenge with the bioinformatics company or group. They are really interested in using GCP and starting to move their on-premise methods for research into the cloud. Can you work with them to use Docker and some of the services."

I said, "Yes, I can." And so I did and I had so much to learn and I found that the group it's called the Galaxy Project. It's a consortium, mostly based in Europe that does a genomic analysis workflows for research that will discover immunotherapies. They had a conference in Australia and I had some personal interest in Australia at the time so I called Google and I said, "Can I go?" So I went from, this is interesting to presenting it at an international conference and being the only non PhD there, which was, oh my gosh, it was intimidating, but I did get it to work because I do have the dev ops sort of chops and team. So a live dockered up little galaxy cluster on GCP, which was kind of a cool feat because at that time there were no GCP data centers in Australia.

So props to GCP, I was running it out of Singapore and I did it live. I tried to go into the world of the bioinformatics people. I participated in a five day training on the tool. And I told them why I was there. I just was honest. One of the things that I do on all the crazy stuff I did in my career when I go kind of out there at the end of the ladder or whatever, I always say default to honest. So I just tell people, I tell a story of my friend Terry, and I tell them why I'm here. And I tell them, I'm a new student in bioinformatics and I'm going to watch the Illumina sequencing videos at night and I'm going to try to make a contribution to what we're doing, but I'm a new student.

Jeremy: Well, I think it's amazing when especially the fact that you had this opportunity to use your powers for good. And so I want to get into exactly the scale of bioinformatics because this is something that I didn't even realize it was this big. Until you and I started talking earlier, I was sort of like, hey, I know there's a lot of big data like genomics, that kind of stuff, that's big data. But I did some research on this and I found this article about this system was trying to do a RNA sequencing and there was something like 640 million reads that they had to do. And it used to take I think 29 hours for this to work. And this is where serverless comes in because they took Lambda functions and they were able to optimize this down to where it used to take 29 hours, now it takes 18 minutes and it costs them $2.82 cents to do that.

Lynn: Yes. Speaking of serverless and genomic scale data, which is really a super set of big data, continuing my story of Australia, this is how I moved into serverless genomic solutions or serverless solutions for genomic analysis. At that conference where I presented, there were some researchers from the equivalent of the National Science Foundation of Australia. It's called CSIRO, the Commonwealth for Scientific and Industrial Research Organisation. So interesting, they had this burstable search problem, genomic scale to find where the edit points in a genomic sample would be. And the edit points would be for CRISPR-Cas9 editing so that you can then try to develop immunotherapies. So you want to have a fast feedback loop so you can go, okay, I could cut here, I could cut here. And the reason you have to have that is because when we think of DNA, when we non biologist think of it, we think that it's like 3 billion letters all in line, like a big road, like a straight road.

But it is not, DNA is coiled and curled and clumped. So just logically, if it's all smooshed together, it's hard to find the precise cut point. Now there's more scientific properties than just smooshed together. But that's basically, there are a set of properties that help to determine optimal cut points. So these guys in Australia, they went to an AWS Summit in Sydney and they saw the sort of classic serverless pattern for burstable websites. So they saw S3 with Dynamo and they saw Lambda with API gateway and they said, you know what, we're running out of room on our shared on-prem cluster, let's just try this, let's just try to build something. And they did. And it was published by Jeff Barr as one of the first all serverless genomic applications. And the applications is called GT-Scan2 and it's on Jeff Barr's blog and we can put the link in here, and it's just a classic architecture. And they built it really fast. And it was up and running.

So I met them in the conferences in Melbourne and they were there in Sydney, and I was going to Sydney for something else. And I said, "Can I just come and talk to you about this?" And they're very smart researchers. They researched me and they said, "Yes, you can come and help us with a challenge we have on this. You're going to get here for free, right?" This was a couple years ago now, like three years ago. I was kind of in the beginning of my journey to genomics and serverless. So we sat down and they said, "This runs, but sometimes it bottlenecks." And I said, "Oh, what are you doing about your logs?" Because of course serverless applications, is all about the logs, right?

Jeremy: Yeah.

Lynn: And they said, "Well, we're not really familiar," because they had never really done serverless before. And at that time Amazon had just released GT-Scan, like literally the month before. And so I looked like the complete hero because I said, "Let's take a look using GT-Scan. Let's instrument this." And bam, one day, one Lambda, was causing 80% of the problem. They reroute, fixed the bottleneck and we publish that too.

Jeremy: So that's one of the things though I think that's really interesting about, again, moving from this on-prem, you said they're running out of space in their on-prem solution. And I think that's the problem with all on-prem solutions is eventually either you're going to keep buying more hardware, buy more hardware, or you're going to run out of space. Or you're going to run out of compute power, which is another thing, which is why you see some of these examples of Hadoop jobs running for 500 hours to try to process some of this stuff. So that's another limitation I think that that serverless overcomes when it comes to data at this scale is for small teams, especially research teams that don't have huge budgets, that can't run massive clusters by themselves. And even if they could would still have to wait hours and hours and hours and hours to get feedback on the work that they're doing. So how does serverless and the cost reduction help sort of the smaller research teams?

Lynn: Well, again, continuing on working with this team because they were a really small team at that time. They've subsequently got more people and stuff. They said, "Okay, great. Good job, coding cloud architect. We have another challenge for you." It was a big challenge given my skill level at the time. And it was one of those 500 hour problems. They had written a customized machine learning library for analysis of genomic variants or differences between in the disease sample versus the reference. And it's called VariantSpark and written in Scala to run on their internal Hadoop Spark Cluster. And again, they had to wait for time on the cluster and then it literally did take up to 500 hours. This is a 500 hour problem. And they said, can't you do something on the cloud?

Okay, here's the tricky part. The computation was machine learning and stateful. And they were really new to cloud and they had a really small team. At that time they had no dev ops people. So exactly the thing you're talking about. So one of the things that was really interesting because it was very much a process and it ended up ... because I didn't do it full time. And we solved it. We got it down to 10 minutes with serverless. But the process is the key. I actually wrote this up on medium.com there's a series of three articles. So the first step in the process was they hadn't even used Docker, like at all. And so what we did first is we didn't even look at VariantSpark because it was too complex and machine learning and probabilistic. We took a bioinformatics tool that is deterministic.

It really doesn't matter. It's called BLAST, which is binary local alignment sequence or something like that. It basically just is the first level of analysis and it's a single executable. And we just put that in a Docker container so they could have one input, a process and one output and see how that scales, and then do that on-prem and then do that on the cloud because you have to have this ladder to learn. This is, again, a really common situation I find with customers moving to serverless that have experience in sort of traditional enterprise; BMs are on-prem clusters or whatever. It's really hard to go directly, really, really hard even for new implementations. So then the next level we did is we did EMR just almost as a lift and shift so that they could learn best practices like CloudFormation rather than clicking on the console. You know what I mean?

And then what was interesting, this was a spark-based library. Conveniently the spark team, the open source group, made Kubernetes a potential controller rather than YARN. So I said, "Oh, here we go. We can now make a data lake." So the progression was we went from on-prem 500 hours to EMR to learn cloud skills basically. And that is something that still a lot of the researchers use because it's simpler, right? Serverless isn't always the right answer. If you just are going to just kick off a CloudFormation template and run a small size test job on EMR, that's more familiar. You just click and you're done and it takes maybe an hour. And the huge scale which runs on the data lake ... So the difference in the serverless aspect of the solution we built was not in the compute layer, it was in the data layer.

We got rid of HTFS, we got rid of ... we put all of the data in S3, and then the compute layer was a Docker. We dockerized spark for EKS. And I think we were actually one of the first groups in the world to do it. And we published on it. It's all on GitHub. And Amazon actually supported part of that research because of what we were doing. So thanks to Amazon for that. And so what the CSI road teams will do is they'll use the Kubernetes for the really huge scale and when people are more comfortable. So this idea of having like a menu for this particular group was vertical where they sort of start with something that's more familiar and then they move to the serverless level. It worked in this particular case.

Jeremy: Yeah. And now, so that's interesting because you mentioned and obviously you're big into education, you mentioned sort of learning how to do some of this stuff in that progression. So obviously there's a lot of tools and there's a lot of things you need to do with serverless that aren't quite as straightforward maybe as the, I guess, the traditional sort of like you said, the EMR approach for example. So maybe in your opinion, what's the correct level of abstraction then for some of these teams to work with serverless? Because I think right now it's very much so pick and choose all the little things you want, configure them, turn all the knobs, tweak all the dials and things like that. But is there another level of abstraction above that, that might make it easier for some of these smaller teams to start with serverless?

Lynn: Yeah, it's interesting. So I've subsequently started working with another bioinformatics client and quite a large one. It's the Broad Institute at MIT and Harvard. And they are in some ways kind of like a bioinformatics incubator. They have 5,000 researchers that have various labs and they have massive on-premise compute resources because they're well funded. But because of the volume of data, are starting to exceed those resources. So they have had a multi year cloud enablement project, which includes working directly with services, whether it's AWS, GCP or whatever. But they have been collaborating with Alphabet, which is spinoff off from Google. It's actually the Verily Group.

And they have created a higher level abstraction, almost a SaaS level, which is called a tara.bio. So it's a website basically. And what it implements is not only this higher level GUI based abstraction, but within the Broad and in collaborating with other researchers worldwide. All the stuff is open source. They have, for example, created a configuration language called WDL which is Workflow Definition Language, which is designed as a higher level of distraction over like a cloud formation or a Terraform because it allows configuration of the execution environment, so the VMs or whatever. But it also enables configuration of the bioinformatics tools. So another trend that's happening in the bioinformatics industry, and I saw this in other industries like ad tech and finance previous, is as the volumes of data go up, there start to be ... instead of just using scripts or code to manipulate data, there starts to be a tool repositories.

So in the case of the Broad, they have this tool that ... they've made several tools, but there's some big tool is called the GATK, the genomic analysis toolkit. And it's a jar file that has over a hundred common tools. So for example one of the things that you might do in the analysis is you might look for duplicates in the reads and so they have a duplicates feature that you can then configure. So then working at this level, you're more configuring. So if we take it back to sort of working generally with serverless versus not serverless, the amount of configuration code is growing and growing because with serverless you have more parts and pieces. And when you go a higher level up, the configuration code is also extremely important. So it becomes, I think, almost as balanced maybe to the point where you're at the level of terra, you're not really even writing any application code, it's all configuration code.

Jeremy: Yeah.

Lynn: So again, this is like personal what is code? What is it being technical? And I really strongly believe that dev ops and configuration code is code and needs to be checked in and needs to be source controlled and needs to be reviewed and all that kind of stuff. And I think, again, this is a big problem in serverless in general, that there's still this lingering sort of bias that if it's not Java or C++ or something like that YAML is not code. Well, it's our podcast so we get to say, I think YAML is code.

Jeremy: I totally agree with you. And actually I think YAML in some cases and some of these configurations are more difficult than writing code in some cases. So I definitely agree with you on that. So I think all that's fascinating and I find this approach to big data, moving to these things where like you said the cost can be dramatically reduced and whether the computer is on something like Kubernetes or even if like this other example with the RNA sequencing is using Lambda or some other functions as a service, just this idea of being able to store this massive amount of data and being able to process this massive amount of data and use it for something that isn't ad tech.

I mean, it's great when we can serve up a nice personalized ad for somebody, but it's better if we can map the coronavirus and find some sort of a vaccine for it or something like that. So maybe this is a question for you, because you had mentioned, three years ago you really had nothing to do with bioinformatics or genomics and you kind of got into this. Now you're working with the Broad Institute there and that's really fascinating. So is this going to be a big growth sector, you think? I mean, people are saying, "Hey, I want to do good, I want to work in serverless, I want to work in big data." This is sort of the next place for people to go, right?

Lynn: I think so. I mean, there's a couple of things. First of all, even if you're not motivated by the ethical concerns, it's the most data that I've ever worked with. I mean, financial it's sort of similar. But just to make an example, and this is public information because the Broad publishes a lot of customer stories. They are currently putting in on average 17 terabytes per day into the Google Cloud. We all like hard problems, we're builders and so this is a really fun, hard problem. And the patterns used to come for big data out of FinTech and ad tech and I worked in those areas and I'm trying to apply some of those patterns. But I do think that the most interesting place to work for a big data professional is human health right now because of sequencing.

One of the really cool things at the Broad, I had done a presentation there, which again, I'm still intimidated to do on the work that I did in Australia because they were interested in that reference example on AWS.

Jeremy: Sure.

Lynn: And one of the researchers said, "Have you seen some of our sequencing facilities?" And I said, "I really haven't." And so she was kind enough to arrange a three hour tour of one of their principals sequencing facilities, which used to be a beer plant, which cracked me up. Because in the big freezer they have millions of human DNA samples now, which was the beer storage. But the most impactful thing to me throughout this tour was, because she'd worked there for five, six years. She said, "We used to take really long periods of time to just sequence parts of the genome, the chromosomes or the exomes." But now they have rooms full of Illumina sequencers, they have Illumina people onsite and they are just like constantly just processing, processing.

And this has been within the last three to five years. And now we're even getting testing, like handheld sequencers. And there's really just interesting aspects. When I was in London last year, there was a new store, first one in the world called DNANudge. And what they do in the store is you spit in a tube and within one hour you get the results back and it's a sequence. And I wrote again, a medium article about this, but it's a very small subset. Again, you have 3 billion letters. So what they do in that store, it's almost like a Regex on the genome. They go to certain spots and they say, okay, this is ... one example is caffeine metabolism, which I happen to have high caffeine metabolism.

So they go to one spot. It's not going to be all caffeine metabolism, but just a known spot. Do you have high, medium or low? So it's very much a subset. And then what they've done, it's really interesting concept, is they give you a wristband and it's like a Fitbit, so it tracks your activity and you store the results in there. It's a little tiny chip. And then they partnered with a major grocery store. And based on the eight characteristics that when you go shopping, you can actually say red, green, or yellow for those characteristics and your activity.

Jeremy: Oh wow. That's amazing.

Lynn: It's super interesting. So this idea of different kinds of sequencing, not sequencing the whole genome, just sequencing parts of the genome. And very timely. The Broad has actually published on this. They're actually working on an improved coronavirus test. Not the exact same principle, but kind of that idea so that the results can come back faster and be more accurate as the sequences are verified. Because again, one of the things that's been really fascinating to me about the coronavirus is, again, I haven't obviously working with genomics for a long time, but the fact that the Chinese sequenced their virus immediately, they released the information and they are open source repositories. One of them is Galaxy, people that I worked with that are actually published a workflow on GitHub and they've published a paper. I mean this is the dream. And of course it's all running in the cloud. And as they get new data coming in, they update the paper. So to enable citizen science, people working together faster because of the cloud.

Jeremy: Yeah. Again, it's fascinating to me and I think the impact just from a health standpoint and ... I don't know. I mean, this is going to bring us to the point where we are able to find better cures for cancer or to test drugs faster or to see how the changes in the environment ... I mean, that's something I'm interested in as well as this, what are the environmental impacts or the longterm environmental impacts of some of these things and being able to test changes in DNA and that kind of stuff based on environmental factors and all that stuff, I think it blows my mind. So certainly congrats for working on this stuff. And I think this is where we can maybe change the conversation to you being recognized as an AWS data hero because this is true hero stuff in my opinion that you're doing with data.

So again, thank you for the work that you do. But you're not only a data hero, you're also a Google Developer Expert. And I think you were maybe the first one that was announced as a Google Developer Expert?

Lynn: Yeah.

Jeremy: Which is pretty fascinating. So I'm going to make this typical thing I think that most men with daughters say. But I'm going to put it out there anyways and just say this, I have two young daughters. I have an 11 year old and a 13 year old daughter. One is really, really great at math. The other one is very, very interested in biology. And I really, really hope that they stick with those STEM type subjects and that they continue to go and pursue those things. And that hopefully this country and this world will allow them to continue down those paths with as few barriers as possible. But I think that that has been a major thing that has challenged women in the STEM education or STEM jobs and things like that. And you're one of those people who I think is truly inspiring to my daughters for example, because you just broke through all that and now you're doing stuff that anybody can look up to.

Lynn: Well, thank you. Well, it's all about teams. And one of the ways that I'm able to accomplish providing value to my customers is by working with like minded people from all over the world. And one of the successes that I've had is this idea of learn, build, teach. And again, I wrote about it on Medium because I always have to share.

Jeremy: Sure.

Lynn: And I encourage it in the people I work with. For example, I just had this young woman, I'll tell you this story because it just makes me so happy. She was one of my students from LinkedIn Learning and they write to me and they always ask me to mentor them. And I actually don't personally find mentoring valuable for me or the person. Now, I know a lot for a lot of people they do. That's great. But what I do is I consider having people work as interns. And the way that it works is they work with me remotely for a couple of one hour sessions, like three or four. And then if we both decide we want to continue, I'll actually hire them as subcontractors. And for this particular person, she came out of the finance industry, she's a grad student. And she had a dream to be a cloud architect and this kind of stuff. And so we worked together for about six months. And she's just left my internship because she got a full time gig with Amazon.

Jeremy: Wow.

Lynn: And I told her, I said, "Okay, you tell me that the way that I work has been helpful to you. Now it's your turn, you learn, build, now it's your turn to teach." So I try to get that idea going with people. Because I think that it's not about an age or about a degree or whatever, it's about what you have done. And I think that that applies even to younger people. When I've worked with some people even in high school that have done hackathons, if they built something learn, build, teach. Because what you're doing then is you're giving back, but you're also establishing your competence. Because another severe problem we have in our industry is bias in tech interviewing. I mean, I'll just say, because I think it's important for people here. I have failed many tech interviews, many. And even with my experience, it's very disheartening because I don't have a CS degree, I'm self taught.

And the interviews at the big companies test for one very small set of skills. They test for did you go to Stanford and take algorithms, and I didn't. And yes you can spend hundreds of hours and memorize that just for the interview and that might be okay for some people. But I just feel that the industry needs to grow up because we have all these very competent people that are technical and I define technical as somebody who has a curiosity and an ... I also want to say when you're driving a car you would like to open the engine and look inside, you're technical.

Jeremy: Very good point.

Lynn: Yeah, you're technical and I don't interview and I don't conduct interviews. Again, I wrote another Medium article about this. I was talking to a founder in Berlin recently and he was like, "How do we get people, everybody's failing our interview process." I said, "That's because your interview process is crap."

Jeremy: Exactly.

Lynn: I said, "You just go to hackathons and you work with people." It's all about working together. And then you understand what does the person do when they get stuck? Do they Google the problem? Do they ask for help? Do they try to compute it in their head for 27 minutes and not talk to you? Who is going to make the most contributions to your team?

Jeremy: Yeah. No, and I think it's a confidence thing as well. I tend to find that people who know a lot of things tend to think they don't know a lot of things. And people who don't know a lot of things, they know everything, right?

Lynn: Right.

Jeremy: Or at least they think they do.

Lynn: Right.

Jeremy: And that's actually one of the problems that I've found even when I ... because I've interviewed for several companies or I've interviewed people for those companies. And I mean, even just talking to people to be guests on this podcast. I mean, I think I find this where you have a lot of women who ... and it's not that they don't have the confidence themselves just, I think they think that maybe they're not enough of a voice to lend to the community.

And again, I can find men who have no idea what they're talking about that are willing to talk about anything. And I don't mean to criticize. I mean, all the guests on the show have been great, but there are a lot of people that I would love to hear from. And I think there are a lot of people that the community and the tech community needs to hear from. And just the culture and whatever else it is, is just of course hundreds of years or thousands of years of history suppress some of those voices. And that I think is a tragedy that that needs to be corrected. So when you have more people like you speaking out then it can inspire other people and like you said, this idea of saying, look, it doesn't matter what your background is, it doesn't matter what college you went to, it doesn't matter if you went to college at all, do you have the attitudes? Do you have that drive? Are you willing to put in those hours and learn that stuff?

And then you get to that point where if you share those things, yes, sometimes you get criticism. But I mean, I always found whenever I share something, I try really, really hard to make sure it's right. I do more research. I dig in deeper. And then when you start getting criticism or congratulations or whatever it is, you get feedback and feedback is always good. It makes you a better person. So I'll say this again, thank you so much for everything that you do and I really appreciate you being on the show and sharing all this stuff on big data. So if anybody wants to get in contact with you, how would they do that?

Lynn: Twitter is probably the best way. I'm really changing kind of my professional profile. For many years I was an international speaker and I'm stopping travel for personal reasons and just because it's a good time to stop travel. So I won't be doing any talks other than near where I live in Minneapolis, Minnesota now, but I'm going to be writing a lot. So yeah. And then I have, of course, I have 30 courses on LinkedIn Learning, so if you want to listen to me there, I have lots of courses and GitHub. Again, learn, build, teach. Whenever possible, when I learn something, when I'm building something for a customer, as long as I can make it generic version or whatever, I will put it on GitHub.

In particular, I have a course on GitHub called GCP for Bioinformatics where I took my LinkedIn Learning course, GCP Essentials, which is an introductory course and I just converted all the examples to bioinformatic data and bioinformatic examples. I would love to get more collaborators on that. I would love more input on that. It's GitHub so you can do pull requests or whatever. And the course is about 50 markdown pages. It's like a quick graph basically; what is the service? Why would you use it? How do you use it? Focuses on the console and then quick short screencasts like five minutes screencasts. So the idea is if you're a researcher, you can learn how to use the cloud with some guardrails of understanding cost and everything just by going to this repository. And I want it to be useful so I've gotten some feedback, but I would love to make that course even better.

Jeremy: Awesome. And then lynnlangit.com as well, right?

Lynn: Mm-hmm

Jeremy: Okay. Awesome. And that's L-Y-N-N-L-A-N-G-I-T.com.

Lynn: That's right.

Jeremy: Awesome. Well, thanks again. I will make sure that we get all this stuff into the show notes.

Lynn: Thank you so much.

THIS EPISODE IS SPONSORED BY: Stackery

View Details

About Ben Ellerby

Ben is VP of Engineering for Theodo and a dedicated member of the Serverless community. He is the editor of Serverless Transformation: a blog, newsletter, and podcast which share tools, techniques, and use cases for all things Serverless. He co-organizes the Serverless User Group in London, is part of the ServerlessDays London organizing team, and regularly speaks about Serverless around the world.

At Theodo, Ben works with both new startups and global organizations to deliver digital products, training, and digital transformation with Serverless across London, Paris, and New York.

  • Twitter: @EllerbyBen
  • Blog: Serverless Transformation blog
  • Newsletter: Serverless Transformation Newsletter
  • Podcast: Serverless Transformation Podcast
  • Theodo: theodo.co.uk

Transcript
Jeremy: Hi everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Ben Ellerby. Hi, Ben. Thanks for joining me.

Ben: Hi, Jeremy.

Jeremy: So you are the VP of engineering at Theodo, and you were just recently named an AWS Serverless Hero, so congratulations on that. So why don't you tell listeners a little bit about yourself and what Theodo does?

Ben: Ah yes. As you mentioned, I'm the VP of Engineering for Theodo. We help other companies launch digital products, be that startups, launching their initial MVPs, to large companies attempting a digital transformation. And more and more I'm helping our clients to use serverless. Be that through building their initial MVPs, but also training and upskilling their developers. So we're based in London, New York and Paris. And basically my role is to help coach our developers, and help us find the new technology areas we want to work on. And serverless has been highlighted as the main area we're trying to move towards. And many of our clients are starting to adopt serverless first architectures.

Jeremy: And what's your background?

Ben: My background, I've been at Theodo in London since we kicked off a team here about four years ago. Before that, a bit of time at IBM. And before that studying computer science.

Jeremy: Awesome. All right. So you mentioned digital transformation, and we've heard this term a lot, especially over the last couple of years. And I think some people think that means sort of moving from on-prem to the cloud, or sort of modernizing things. But you've been using this term, serverless transformation more recently. And essentially, this is this idea of going, I guess your second move to the cloud. Right? So could you explain what you mean by serverless transformation?

Ben: Yeah, sure. So what you touched on was digital transformation was that initial move to the cloud, which smaller and larger companies have managed some, with varying degrees of success. I actually helped a company called Junction launch their initial product about two years ago, which is an AI service that helps large companies plan their migration to the cloud. And that was very much a lift and shift approach. But more recently, if we take the example of Junction, they've had more and more targets going to things like SaaS, and FaaS and serverless first approaches. When I talk about serverless transformation, I'm talking about startups who are launching their initial MVPs and doing that in a serverless first approach, but also larger companies who are trying to consolidate their developer resource by building serverless first architectures, rather than managing infrastructure. And more than just managing infrastructure, common application things like authentication, moving to that as a service and really leveraging everything as a service to focus their development teams on the core business value that they're adding, the distinct business logic that makes their company who they are.

Jeremy: Right. So I mean, it's more about that lift and shift approach. And I think we've talked about this on the show a number of times, that trying to just sort of move everything as is from your on-prem into cloud is a bit of a fool's errand, right? I mean you're essentially copying this local environment, but you're not getting the benefits of the cloud environment.

Ben: Sure. And it has some benefits like virtualization was an initial move. Containerization was another move. And now we're seeing sort of function as a service, and other things as a service. As soon as a further level of abstraction. The higher that level of abstraction goes, the more business value I think he gets.

Jeremy: Awesome. Right. You wrote a post called In Defense of the Term Serverless, and nobody seems to want to have this conversation with me anymore. Because I have been very outspoken. I had a post a while back called Stop Calling Everything Serverless, where I tried to essentially define what I thought the term serverless was. And for me, I look at it as not a technology, not a managed service, not FaaS, not some sort of a spectrum or a ladder or these other things that I think are really, really interesting ways to try to classify what it is, because it's such a hard term to sort of grasp. But I look at it more as sort of this process of using these services that don't require you to really have an active involvement in the management of the infrastructure.

And to extend that even further in some cases where possible where you don't have to even worry about the scaling or provisioning a cluster of something like, I don't know, a cluster of Elasticsearch or something like that. So I think this is a perfect opportunity because in this community now we have a lot of forward thinkers and I think we want to move past this idea of what is serverless? The problem is that our community is for the most part an echo chamber. Right? And we keep having this conversation every once in a while when somebody from the edge sort of asks us this question. So, I'd like to get your definition of serverless and why you think that term actually is really important.

Ben: And I think it's a long answer. So I've recently been working on the sort of preview chapter of the book, Serverless Transformation at any scale. And the first chapter of that deals with what is serverless? It talks about that move from virtualization to containerization and then function as a service. But then it talks about how it's not just function as a service, it's using cloud native technologies as much as possible, which makes it a hard thing. It's not a binary classification between serverless, not serverless. It's a polymorphic space that keeps changing and keep adapting. I think we can place different services on a spectrum of serverlessness. So Elastic Beanstalk is obviously not serverless, but it's more serverless than manually provisioning EC2 to instances that goes all the way up to using something like Cognito, which is very little code. It's really the cloud provider providing that logic for you.

So I think you can put things on that spectrum, but I don't think that's where the value is. I agree with you there. I think the value is the sort of, I don't like to use the word paradigm shift or mindset shift, but it is, it's a mindset shift to think serverless first to basically as an engineering leader or a developer to focus on trying to defend your team from doing work, trying to offload that to a cloud provider or a third party. It's extremely competitive in the minutes of launching initial products or if you're in an industry to keep your costs down. So I think we need to focus on trying to leverage things as a service which makes development more like combining lots of different things, which then makes you have to be an expert in lots of different things. And as you touched on earlier, this space keeps changing and we're in a bit of an echo chamber internally in the serverless sort of group where we keep coming up with different ways of doing things.

AWS, for instance, introduces something like Lambda destinations and we now have three different ways of handling the same problem. I have lots of people asking me, should I use SNS? Should I use SQS? Should I use Kenisis? Should I use EventBridge? I mean the answer is yes, but it just depends on scale and many different things. So I think at the moment the space is very much in flux and we need to sort of consolidate some best practices, which is going to come with time. But for me, for now, serverless is a good North star for us to walk towards. So as a term for me, I think it captures the imagination of the place you want to get to. But you're right, I don't think it's particularly useful to try and place everything on a spectrum of serverlessness.

Jeremy: But how do you explain that to somebody who literally has no idea of what we're talking about? You and I were talking earlier and you had mentioned, hey, you're going through airport security and someone's asking you what you do and you said cloud computing and they wanted more details, so if you throw the term serverless out there, maybe they wouldn't even let you into the country because, I mean, honestly, it's just this thing where they're like, what does that mean? And there's all those jokes, there's still servers and serverless and things like that. But there truly is, I think a defining characteristic of what that term means. But it's so abstract. Right? And again, all of these other definitions. So how do you explain that to someone who maybe is new to the cloud or is trying to understand this role of leveraging third party services, the managed services and so forth. And like you said, writing as little code as possible, but probably a lot more configuration.

Ben: Yeah, it's a challenging area. So Theodo as I mentioned, we build digital products, but we do that very hands on with teams. Our goal is not to stay with a client for a long time. It's to empower their engineers to upskill in the areas, help them build their initial architecture, then let them be empowered to build it further. So I get junior developers, experienced developers, people with little sort of knowledge potentially of the cloud or of serverless, and it starts kind of as this book I've talked about, is structured. The first chapter is what is serverless? It talks very abstract, then jumps into some detail, but the ending paragraph says something to the effect of keep coming back to this chapter as you progress through the book because you need to experience some practical examples. We see things like the URL shortener that doesn't actually use Lambda.

We see things like imagery sizing or PDF generation. They're great ways to show people the value of serverless. So generally in training projects around serverless, I sit with a team and talk to them very abstractly about the different definitions of serverless and what the space means and they look at me a bit blankly, but I say, trust me, we'll come back to this in a couple of weeks. We then take part of their system and migrate that to a serverless first approach. Let's say a PDF generator is a classic example that I like to do because you can use Lambda, you can use S3, you can start to use some of the events, and you can start to see a lot of value. You can have a system that before required a complex queuing system because PDF generation is quite complex and it takes quite a bit of compute power, but now we can parallelize that across many Lambdas, so they start to see, okay cool this scale, cool, we're leveraging third parties here, we're doing different things.

And then you come back to the definition and you see them start to understand a bit more and then you withdraw it and go more complicated. You take a microservice with API gateway and you let them start to build out and I think it's really a learning by doing. But the complex thing with serverless is a space of grown is, there's many different ways of doing the same things. So people are almost scared to start because they're not sure that what the best practice is. Best practices in serverless are, and it's quite a confusing space. I don't know if you've found the same thing around best practices recently.

Jeremy: Yeah, and actually Paul Swail wrote a great article lately about, basically don't trust best practices. Especially with a space that's moving so quickly, it's just really, really hard I think for somebody that's new to this to jump in at the level we're at and I think this may be moves us to maybe be on best practices and to this idea of complexity. So we know that there are ton of configuration options when you're building serverless applications. And the hello world ones are super easy, right? You write a little bit of code, you throw it into a Lambda function. Maybe put on the other end of an API gateway, maybe get fancy connected to DynamoDB. Maybe you do some SQS, like you said, and do some queue processing. For the most part, that is fairly simple.

But then let's say that you have a downstream system that is processing off of your queue and you need to throttle that Lambda function. Right? Now all of a sudden you've got all of these settings in the SQS redrive policy where the visibility timeout has to be six times the function timeout and you have to do, set your retries to be higher so that for the initial burst of polling, there's things like that. There's the Lambda destinations, as you mentioned, right? So what happens when your functions fail? Where does that go? How do you monitor that? Where's the retries? What are the retry policies? There are a lot of retry policies, I've just done a presentation on this. There's a lot of retry policies and a lot of failure modes. So that type of complexity, coupled with the term, it becomes really hard for people to start. And maybe that's a question that I have here is, is this complexity increasing to a point where it's going to make adoption harder?

Ben: I think so. And I think people inside the serverless community taking a very sort of purest serverless first approach, which I myself have done. And I find myself a little while back saying there's no way I'm possibly letting an RDS inside this architecture. There's no way, it's not a serverless service. It's going to ruin everything. But sometimes you do have to work with existing technologies. And I think if we're throwing onto somebody who's just moving into serverless, the complexity of DynamoDB and all the other services with the retry policies and everything, we're not going to end in a good place. Which I think is what your recent talk about, adopting non-serverless components in an architecture was very valuable.

And you're right, there's still a lot of complexity in different patterns that you yourself has come up with around how we manage retries. But you don't potentially have to have the perfect retry policy from day one. I think making a start, making some errors and then learning from those. I think people adopting serverless have to learn with the community and that's a challenging thing to say, but at the same time, so many other problems have disappeared that your team can spend time looking into retry policies or all the different services they have to upskill.

Jeremy: Right. Yeah. And so I mean definitely one of the things too, just speaking of people sort of starting with serverless, I think you hear sometimes people think, well, serverless is great for spiky workloads, right? Because it can handle a large spike if you've got a Black Friday sale or something like that as an example. But if you're not somebody who needs that type of scale, then it might just be easier spinning up a lamp stack or something like that on a virtual machine or using Elastic Beanstalk as you had mentioned. But I look at serverless and I look at the ecosystem of tools around it, and it's so much more than scale. I mean obviously scale is important. I mean that's one of the reasons why DynamoDB is so great because yes, it can handle 60 gajillion transactions per second or whatever it is that AWS or that amazon.com has on Black Friday.

But what I think is interesting about approaching serverless as more than just a mechanism for scale is all of that built in reliability and resiliency that is available to you sort of right out of the box. So you see a lot of people, you're encouraging people to move to serverless. But do you see pushback on that where people are thinking that serverless is just something for these spiky workloads?

Ben: No, definitely. And I think a lot of the marketing material about serverless talks about its infinite scale, which sounds amazing, but if you have the problem of infinite scale, you also have good things in your business model. So that's one of the reasons. The book I mentioned, it has the subtitle at any scale because it is for those large companies of huge workloads attempting a digital migration, moving to a vendor driven or they have to handle all these complex retries and amazing scale and all the problems that come with that. But it's also for the small startups who can't spend their time building generic application functionality like use a signup with multifactor authentication and password reset. They need to focus on the actual business value. So serverless is an abstraction, helps them focus on that business value, but it also gives you amazing things like the developer experience.

If you have a stack that's completely serverless through project at the minute with a client migrating a legacy PHP application to a completely serverless stack. Now they were very happy to move to serverless. They wouldn't drop the PHP, but that's for their CTO to live with. I'm joking. But actually PHP has been absolutely fine on the project, because they have a completely serverless stack, they have a CICD pipeline that for every pull request spins up a complete stack runs 300 automated integration tests. And it's not a problem. It runs in six minutes. And so it's an amazing pipeline that they've built. But this also means that they have no out-of-date dependencies in their project. Because every night a cron job goes and updates the dependencies. And if a test is pass, well it propagates set up. And if a test fail, then it goes onto Slack and to ticket for the next day. The development team are empowered by the flexibility serverless gives them. But yes, that comes at the cost of a bit more complexity when it comes to configuration.

Jeremy: Right. And the other thing with complexity too is obviously, I've had this discussion with a few people before as well is, what do we call serverless applications? They're micro microservices? They're nano services? They're these sort of complex beast. But what's great about it regardless of what you call them is, you have a lot of control over configuring each individual function to do something for you, right? So this is the argument that I always make is to say, if you're putting everything in a container, this is the problem with monoliths, right? If you need to scale your PDF processing, and that's running on the same machine or the same container as your login process or your order process or something like that, then when you have a huge batch of PDFs that needed to be created, you have to scale up that entire thing.

And that means that you're potentially over scaling some of these other things. You're wasting some compute. Whereas with serverless, you can say, look, I just have a function that just does PDF processing, and if that has to do 2000 PDFs per minute, then it can, but the rest of my system can live at whatever scale. But that does create quite a bit of complexity. And it also creates sort of, this messaging or this microservice communication problem. So how are you sort of telling your clients to address sort of this new volume of messages that need to be passed around in order to coordinate serverless applications?

Ben: Yeah, so you touched on a few things there. I've seen a lot of, a lot of companies where their code base becomes unsustainable because of the amount of complexity in that monoliths. And the business as a new requirement that analytics wants is events, but then it's two months on the backlog because the monolith is so tightly coupled. Now at Theodo we've started recently formalizing code quality internally across projects and it's sort of a five point model. And one of those points is sustainability. So keeping the code based sustainable, meaning that it's always adaptable to change and people will often make the argument, small companies don't need microservices or microservices and more complexity than they're worth. But when we talk about serverless, the ability for a team, a developer to be able to spin up into production, a Lambda function that does something in production and the real production stack is amazing.

At Theodo we've recently started to, and I say recently, it's quite recently from AWS, start to try and adopt EventBridge, so formalizing the events flowing through our system. This means that we have this sort of centralized event bus events flow through there and that allows your Lambda functions or other services to subscribe to those events and third parties to push events soon in this cool developer experience stuff like a fully typed SDK that comes out of the box, but it's also a sustainability thing. It allows you to say, okay, well these things are all just listening to events. If we define our system by events from the business point of view and we have things listening and doing things with those events. We can add new events later and it's not going to be a problem. It's not going to be a huge friction because it's simply a case of publishing in new events and listening to it and our existing services can choose to listen to it, but that's optional. Our new service can listen to it or can listen to existing events and do different things.

So EventBridge what we're really seeing as the future for serverless architectures. Now like everyone we're starting to try and understand the best practices. A Sheen from LEGO has done amazing work there and recently at the serverless meetup in London, I help run with Ant Stanley, Bob Gregory from Kazoo gave an amazing talk about how his team does event storming to sort of figure out the whole system and then they use EventBridge on top of that sort of formalize how those events flow. And he made a very interesting point that everyone's talking about latency and EventBridge can actually have quite a bit of latency, a few minutes or even longer. But what he talked about was, I mean his business model is giving you a car within 24 hours. So you have a car, you sell your car to them and they ship you the new car. I think, I don't fully understand their business model.

But I mean that's a 24 hour process. Three minutes of latency isn't going to kill you and that three minutes of latency. Okay, I mean that's a worse, I'm not sure the worst case is for EventBridge but something in that ballpark. But EventBridge has given them so much. They have autonomous teams working together to build a product faster than I've seen other companies ever built and building it sustainably. And that sustainability I think comes from the fact they're event driven. And event driven, you don't need to be event driven to use serverless, and you don't need to use serverless to do event driven, but everything makes a lot more sense when you do both at the same time.

Jeremy: Right. Yeah, I totally agree. I mean I am a huge fan of EventBridge and it's funny you mentioned that about sort of best practices and what those best practices are. The EventBridge team themselves are still trying to understand what those patterns are, what are the customers using it for and what sort of makes sense for them to do that. And so for me, I agree. I think that this idea of event driven with serverless, right? The combination of those two things is huge. And before something like EventBridge came out, I was using SNS and I was using SNS basically as a standalone sort of event service where I would send all my events to this event service. You'd have to take that and then distribute that out and create all the subscriptions around that. There's so much more sort of capabilities and options and features that are in EventBridge that I totally agree. I think that is going to be the way.

I know that my latest projects, I've used it to sort of standardize how we do messages or how we do messaging between the different applications or the different microservices I guess if you want to call them that. And so essentially that brings us to this next question. So if EventBridge is this sort of, central component of all future serverless applications that you build and maybe not small ones, maybe there's a few small services in whatever you want to call those, but certainly for the systems that are more complex, like what's the next step? What's the future going to be? I mean this is probably a dumb question to ask, but what is serverless going to look like in five years in your opinion?

Ben: I think it's going to be more abstraction and more consolidation around how to do things. So in five years it's going to be more obstruction around that configuration so you're not having to manually configure retry policies out of the box, you're sort of being able to sort of, well maybe not pointing click, but in a very short amount of yaml we'll be able to have an event driven architecture and maybe that becomes formalized, an event sourcing service rather than just an event bus. Maybe other areas of event driven become more formalized, but it's always going to be an increase in abstraction. We went from virtualization to containerization to function as a service and other things as a service. Now we're sort of building more event driven. We can have obstructions at different levels, so there might be obstructions in particular services or abstraction of the whole architecture.

Things like the serverless framework have tried serverless components, AWS has tried the serverless application repository. And those things have varying degrees of success, but built into all of these services, although we're going to give it more, we seem to have had a spike of complexity recently as so much has been announced and so much has been released. I think we're getting more abstraction. If we take the amazing work you did about integrating RDS with Lambda and you built a whole sort of open source project that really helped people with that. Recently, although there's still a need, AWS has abstracted a lot of that with the RDS proxy. They've seen a need from the community and that abstracted that. If we take EventBridge, people were doing CloudWatch custom events kind of before and hacking it and then they formalize that and provided a level of abstraction. Now is that abstraction going to be driven by the cloud provider or by the community? Well I think it's going to be a bit of both. AWS and other cloud providers are going to add more services but also increase abstraction as it goes on and the community is going to build amazing open source projects that increase abstraction. So I think the move to more abstraction, which means less configuration and configuration is just code. So it means less code to achieve the same business value. So for me it's more obstruction, but right now it feels like less abstraction.

Jeremy: Yeah, no I totally agree. I mean that's one of the things where I started saying this towards the end of last year where what serverless needs now is some sort of abstraction as a service. Something that runs on top of a cloud formation. And even in a sense on top of something like the serverless framework or SAM and you're right, serverless components is one of those things. The CDK, the AWS CDK has some capabilities around that as well where you could sort of formalize, maybe call them best practices and say, look, the standard way that we want to create this particular type of solution or this outcome that we're trying to create, that we want to be able to write this in one line or two lines of configuration and say this is the base use case, this is the outcome, this is what we want it to do.

And then if someone says, oh well I needed to retry more times or I need to do this. Then you can go in and you can start tweaking those defaults. But I totally agree. I think that having a way to encapsulate some larger process that stitches together all of these individual sort of off the shelf products that AWS or other cloud providers give to you is going to create and empower people to not only create something that is highly usable, right? Which doesn't take a PhD or months of study to figure out how all these different components work. But being able to take that and then standardize that across multiple departments in their organization. And that allows you to follow standards, to implement security policies, to implement whatever best practices your organization sees fit and vet those, and then allow for the developers to just easily pick that up and go ahead and do that.

And I know Liberty Mutual has done some things around vetting different AWS products so that their developers go ahead and they could use them that they already met the compliance they needed and things like that. And now I know they're working on CDK patterns that they're going to be able to distribute so that there'll be able to implement these different functions that are these highly reusable patterns and again fully vetted, battle tested, just things that work right out of the box. Just making it super easy for them to do. So I totally agree with you on that.

Ben: So just then we both kept using the disclaimer on any cloud as sort of a get out of jail free card. Not to seem to tie to AWS, but I think we both agree that AWS are really the trailblazer in this space and that's why we're often using AWS terminology to talk about it. As we talk about abstraction and we talk about standardization of these components that are easily used by our teams and abstraction is going to help us be able to actually build applications, not spend huge amounts of time learning configuration. Do you sort of see that abstraction being across cloud providers or still specific to each cloud provider?

Jeremy: Yeah, actually that's an interesting question because I think that what we're seeing is some of these other clouds sort of now trying to play catch up. Right? And I mean there's no, I don't think there's anybody who would argue with the fact that AWS is years ahead when it comes to, at least pioneering the idea of serverless and things like managed services. They are releasing all of these individual services and they're releasing all of these components and adding features that make it easier to build these applications. So thinking about other clouds, what worries me is that, and I like diversity. I like the idea of there being another cloud that can do something different and maybe you can do something a different way in cloud A versus cloud B. But I think that actually deepens the divide and makes it harder for people to understand not only what serverless is, but to develop the skills they need to actually build serverless applications.

Like right now. And not to mention Paul Swail again, but Paul Swail just wrote an article the other day. He's been very prolific lately about venturing away from AWS and using some other services. Like maybe Cosmos DB works better for certain things or Google Bigtable or something like that. And could you use that in combination with your current AWS infrastructure? And I mean, honestly, in applications now we use IBM's NLU API for natural language processing. I use Twilio, I use a SendGrid, we use Stripe. I mean we use all these services that are not AWS services. So really using these other components is not unheard of. And that's sort of how I like to envision multi-cloud in a sense where you're using other clouds but you're not trying to create full parity or some of these other things.

But I think sort of to go back to your question, the biggest problem I see with trying to standardize things across clouds is that one, it's not going to happen. I mean, it's too diverse and I mean the functions as a service in general is such a commoditized thing at this point that it doesn't matter. I mean if you run your functions on Azure, you run your functions on Google or you run them on AWS or you run them on Spot Instance functions or whatever. I don't think that really matters too much in terms of things like vendor lock-in. But what I do think matters is if you choose SQS, which is the simple queue service in AWS, the best way for you to process events off of that is to use a Lambda function to do that.

And if you're using Azure and you're writing sort of complex workflows, then the best thing for you to do is to use logic apps, right? You're not going to use step functions to communicate with Azure functions. So I think that what you're going to see is that people are going to have to pick a provider and learn that provider deeply understand everything from IAM to like, again, all these failure modes and all of these little tricks that they can use in all these different levels of abstraction that they can do. I just don't think that you're going to see that translate to multiple cloud providers. And I don't think Kubernetes is the answer either, because again, that just handles really the compute part of it, which is highly commoditized, but I don't know. I'd be interested in your thoughts on it as well.

Ben: Well, yeah, I think when people think multi-cloud, they think the containerization dream that Docker managed it's standardized containers and therefore we can use Kubernetes to run the same container on GCP as on AWS, as on Azure, as on Alibaba cloud, really any cloud provider. And that containerization standardization allowed that portability, which for some is important. So I'm generally of the sort of view that yes, you might find a way to standardize functions as a service, but there's no way you're going to standardize DynamoDB or NoSQL databases or how are you going to standardize the events flying off them and how they interact. I mean it's just, it's not possible because of the different offerings and you're right, the diversity is good. It's good that people are doing things in different ways. Cloud providers are doing things in different ways, but when you move to a regulated industry, it becomes a bit more difficult.

So in the European Union, there's a regulation that's been passed that banks need to be fault tolerant to at least one cloud provider. Now this is because banking and the economy are very tightly coupled and when banks go down, it's not a good thing for the economy. And as every, let's say, three major cloud providers that these sort of scale of banks could work with in Europe, it's very likely that multiples of these banks are on the same cloud provider. So if one of these cloud providers were to go down, which many of us think wouldn't happen, but obviously it's something to be considered, then multiple banks could go down and the impact could be huge. So being fault-tolerant at least one cloud providers, what does that mean? Well, it doesn't mean that you have sort of a hybrid between different card providers where some stuff is done in AWS and some stuff is in GCP.

It means that you're able to deploy the same infrastructure in both. And I think at the moments, and maybe forever if you want to do that, I don't think you can fully embrace serverless. Yes, you can use Knative for the compute parts, but I think you're going to be stuck with containerization and classical databases until either standardization happens, which I don't think it does or is a more advanced way of doing this sort of multicloud.

Jeremy: Right? Yeah, because I agree with that. I mean, you can't be, this is this lowest common denominator argument, right? I, if I'm going to try to choose something that I can run in multiple clouds, how can I truly be cloud native if I'm using something that is going to be not built for the cloud? Maybe that's the wrong way to say it, but I mean that's, I think that's where you have that, that huge amount of capabilities. If you're on GCP and you're using something like Firestore, right? Like that is an incredibly powerful service that does all kinds of really great things. And if you said, well, I don't want to just be locked into Firestore, so I'm going to go ahead and just do an installed MySQL database or something like that, you're not going to get the features that you get from something like Firestore.

You're not going to get the features in some other, you're not going to get the scale features, let's say, of DynamoDB and some of the capabilities that that has. If you just say, well, we're going to install on some VMs, we're going to install a MongoDB or we're going to install a Cassandra Ring or something like that. So I totally agree with you. I think that in order for you to really embrace the cloud that you have to basically pick a provider and maybe if you've got the resources go great, pick multiple providers, figure out how to build your app cloud native on AWS, figure out how to build a cloud native in Alibaba or Tencent or whatever you're doing. But you have to, if you want to embrace that true cloud native, I think a mindset if we want to call it that, you have to choose the tools that are built to be cloud native.

Ben: And then put those teams in different buildings and don't let them talk to each other.

Jeremy: Exactly, exactly.

Ben: Sure. But yeah, that's a huge cost. But I guess if there's a hard constraint coming from a regulation point of view, maybe that's the sort of cost that has to be embraced or people don't adopt the cloud native architectures. And comparing, which would be a higher cost is probably too complex to do.

Jeremy: Right. That's probably true. But I mean the other thing to think of too is I think if you're developing for multiple clouds and I have not seen any teams directly doing this, but I would certainly think that you're going to need people who are experts on AWS and people who are experts on Azure for example. If you truly wanted to do those separate things, because even if you're installing Kubernetes and you're running Knative and you're doing some of these other things, there are very specific services on each one of these cloud providers that handled different things that kind of deal with them differently. There's different security requirements, there's different protocols, things like that, that you have to learn. So I do think that that if you truly want to take that multi-cloud approach, and I hate the word multicloud I think is a terrible idea in many cases, but I get what you're saying on the regulation stuff that, yeah, I think you, you just, you just have to have a really big team in order to do that.

Ben: I mean how often? We've not been touched on the data duplication.

Jeremy: Right. Yeah. That's another-

Ben: Which is a huge.

Jeremy: Yeah. Yeah, that's a really good point. So yeah. So talk about that. I mean the data piece of it, right? I mean this goes to this vendor lock in thing, right? And people always get concerned of, well, if I choose AWS Lambda function, then I'm locked into AWS. And I would argue, as you just mentioned, that it's really not about choosing the compute, even choosing some of these other managed services. There are duplications for some of these that you could potentially port over. But as soon as you start putting a bunch of data somewhere, you're kind of locked in from a data aspect.

Ben: Especially when you get to the scale that vendor lock in is a concern and that's an immense scale or a regulation requirement which comes at large scale. The volumes of data we're talking about migrating potentially between different storage mechanisms. It's, yeah, I'm not sure it's always going to be possible and that's where it becomes very complicated and whether you do that, incrementally or not. I think for many, my view is you need to embrace vendor lock-in and embracing vendor lock in is called cloud native. They're just different names for the same thing depending on your mindset towards cloud providers. But you need to embrace vendor lock-in and if the day comes where Amazon's going to triple its costs or whatever the fear is, because for many the availability that Amazon provides, it's better than the availability they'll be able to maintain in time obviously.

The data that comes, it's probably going to be cheaper to try and move then rather than try and always have this ability to move, maintaining an ability to move cloud provider for years and it might never happen is not the right move. I would, when that day comes, figure out how to move.

Jeremy: Well I totally, totally agree with that. Awesome. Okay, well listen, Ben, thank you so much for joining me and for sharing your serverless knowledge with the community. And congratulations again for the AWS Serverless Hero distinction. So if our listeners want to find out more about you, how did they do that?

Ben: Yeah, so I'm running a few different initiatives under this sort of brand of serverless transformation. So there's a serverless translation medium. Feel free, check that out. This same podcast recording is going to go out on the serverless transformation podcast which you'll on Spotify. Another good places to find podcasts. And there's also a few open source projects under the Theodo github. And if anyone wants to reach out with any questions, I'm always very available on Twitter.

Jeremy: And that's @EllerbyBen.

Ben: Yeah, that's @EllerbyBen.

Jeremy: Awesome. All right. I will get all that into the show notes. Thanks again.

Ben: Amazing. Thanks, Jeremy.

THIS EPISODE IS SPONSORED BY: Epsagon

View Details

About Dr. Peter Sbarski

Peter Sbarski is VP of Education & Research at A Cloud Guru and the organizer of Serverlessconf, the world’s first conference dedicated entirely to serverless architectures and technologies. His work at A Cloud Guru allows him to work with, talk and write about serverless architectures, cloud computing, and AWS. He has written a book called Serverless Architectures on AWS. Peter is always happy to talk about cloud computing and AWS, and can be found at conferences and meetups throughout the year. He helps to organize Serverless Meetups in Melbourne and Sydney in Australia, and is always keen to share his experience working on interesting and innovative cloud projects.

Peter’s passions include serverless technologies, event-driven programming, back end architecture, microservices, and orchestration of systems. Peter holds a PhD in Computer Science from Monash University, Australia.

  • Twitter: @sbarski
  • LinkedIn: https://www.linkedin.com/in/petersbarski/
  • A Cloud Guru: acloud.guru

Transcript
Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Dr. Peter Sbarski. Hi, Peter. Thanks for joining me.

Peter: Hi, Jeremy. Thank you for having me.

Jeremy: So you are the VP of education and research at A Cloud Guru. So why don't you tell the listeners a bit about yourself and what A Cloud Guru does?

Peter: Yeah. Thank you, Jeremy. So my background is in computer science. I got a PhD about 12 years ago, from Monash University in Australia. I worked as a consultant focusing on cloud projects, primarily. Then four years ago, I joined A Cloud Guru and have been with A Cloud Guru ever since. So at A Cloud Guru, we create awesome fun online education. So we help people get skilled up on AWS or Azure or GCP, or just learn cloud related technologies in a very fun and engaging and practical way as well.

We help people get certified, but also learn how do you use containers? How do you use Kubernetes? How do you go serverless? How do you do things with best practice in mind? So we focus on making sure that we produce that high quality, curated education that anyone can access.

Jeremy: That's awesome. All right, so you and I have bumped into each other, and you're all the way in Australia, and I'm over here on the East Coast of the United States. We've bumped into one another in Seattle, and at re:Invent a couple of times. Every time we get together, we are always talking about serverless education. I think last time we were together, we went maybe even way beyond serverless education. That's what I want to talk to you about today is just the state of serverless education.

And we can go a little bit deeper. But one of the things that I think is unique about serverless as opposed to maybe learning even containers, or even some of these other cloud concepts is, serverless just seems to be such a reworking or re-engineering your own mind to think about these things differently. What are you seeing in terms of maybe the challenges between training people, just on programming languages and some of these other cloud computing, concepts versus training people on serverless?

Peter: Yeah, it's a great question, Jeremy. Look, I hate to use the word paradigm, but it does feel, it is really a paradigm shift. Because serverless, it feels like, this is what cloud was supposed to be all along, right? You're not dealing with low level infrastructure concerns. You're not provisioning your servers and thinking about memory capacity, but you're thinking at a high level of abstraction, you're thinking in terms of code, you're thinking in terms of functions and services and event driven architectures. That's interesting. It's different and it requires people to really think in new ways.

Look, I think, honestly, the adoption of serverless will hang on education. If it can educate people, serverless as a concept as an idea will be successful. I think that's what we're all working towards. This is what you do nearly every day, right? You educate people on serverless. You blog, you talk, because this is the way we get people to understand.

Jeremy: Yeah, so that's actually a really good point about education, because I think there is an education gap. But before we talk about the education gap, I think from a more maybe structural standpoint, one of the things that is really interesting about cloud computing in general, and I think you're right, it's hard to draw that distinction between what is serverless and what is just eventually cloud? What we understand that to be. I'm thinking that I watch people struggle, trying to figure out, "Okay, AWS just launched some new feature that has now made some of the workaround that I was using in the past has made that obsolete."

I think that you have this speed of innovation in the cloud. It's not just AWS, it's Google, it's Azure, it's Alibaba, it's Tencent. All of these cloud providers are just going through now, and releasing all of these really cool new features. So how does the average human that doesn't read 800 articles a week like I try to do, how do they stay up to date with this stuff?

Peter: It is actually very difficult. It's very hard because the pace of innovation, especially in cloud computing like you said, is incredible, right? It's so funny. We actually do a weekly show, a round up at A Cloud Guru covering everything that has happened in AWS or Azure for that week, right? And we always have material to talk about because there's always something new. So yeah, you have to have a trusted source, you have to watch a show like AWS This Week or read a round up blog or something, because it is so hard to keep up with everything.

Look for us, it's a full time job, right? You just have to stay up to date, then hopefully, we can share what we've learned and what's important with everybody else. But yeah, it's a challenge. I don't blame you if you miss a few things. It's just too quick.

Jeremy: Yeah, no, it goes beyond that too, right? If you think about saying, "Okay, well, the great now they've released Lambda destinations, or now there's the HTTP API, or there's these other little things they do." The Lambda destinations really changed probably what the best practices for dead letter queues with Lambda functions, right? Because you get more context when you use the failure path of a failure destination, I guess. So that's the other thing. Forget about just knowing what's available, knowing the right ways to use it, or the best practices or the leading practices. That's a whole nother thing you have to keep up with.

Peter: That's it. It's so funny. I remember, I was writing my book, I wrote a chapter on the API Gateway, and API Gateway came out. I remember I finished that chapter. I was so happy. Then literally two days later, I know Proxying came out. So now you could proxy request straight to Lambda. So you no longer have the right velocity templates. And I'm okay, well, let's scrap that chapter. Let's do it all over again. All right, had to start from scratch and come up with that new best practice. It happens all the time. It's hard.

Jeremy: Right.

Peter: Yeah, this is why there are great people such as yourself, who write blogs and talk about best practice, and we produce content on that as well. Yeah, I think it's the only way right, you have to have a good source of advice You got to try keeping up to date. What do you do? Let me ask you Jeremy, what do you do keep up to date? Apart from reading 800 articles a week?

Jeremy: Well, that's basically what I do. Listen, this is the other thing too is, I don't stay up to date on these things. I pick and choose some of the things that I want to follow. There's a lot of innovation that's happening with IAM and some of the cloud map stuff, Cloud Discovery. Honestly I have not been paying attention much to it, other than watching Ben Kehoe tweets, right?

Peter: Yeah.

Jeremy: That's basically what I've been doing to try to follow along with that, because there are so many other things that you have to go deep on. Again, I think that's one of those things with the cloud in general where, as a normal developer, and if I'm a developer 20 years ago, I'm learning Java in my freshman programming class, my first CS class that I take, and then I go on to maybe learning about data structures, and I learn some of these basic things. I'm learning about allocating memory, which no modern programming language that you're going to use is making you do those things anymore.

Jeremy: But, you're going to learn some of those basics. I think that's absolutely great. But you're also coming out of school, and I think a lot of these people I know, just from people I've interviewed, they don't even know anything about the cloud at all. Right?

Peter: Yeah.

Jeremy: I know some schools have moved to changing that education to maybe learn something like Python as that opening language or that beginner language that they don't get as stuck on as Java. But, going back to your point is that, you go to be a developer in 2020, all of a sudden, not only do you need to know programming language, you need to know about data structures, you need to know about some of these basic computer science things, but now you need to know how the cloud works. You need to know about distributed computing and you need to know about caching and eventual consistency, and all these other things.

Maybe this is a good time to go back to that point on where those gaps are, what are kids ... and I say, kids, and you and now we're getting older, right? So we can refer to college students as kids. What are they learning, that's preparing them for this cloud economy that we're going into?

Peter: I think that is actually a really good point. There are a lot of people coming out of colleges and universities who are not exactly prepared for the industry and the expectations of the industry. They have to ramp up on cloud really quickly. It's a challenge, and I just want to say that, look this stuff is hard. Honestly, it is, there's a lot to know and I don't want anyone to feel discouraged, just because they don't know something right now, it doesn't mean that it cannot be learned, right? We all started from nothing, we were all beginners once. It's all doable. It's all possible.

You just have to go and spend a little bit of time watching courses, reading books, going online, finding blogs to really upscale if you're missing some of those skills, but with universities and colleges in particular, there is that question, what are they teaching? Are they relevant? Are people getting the right skills for the industry? And if not, then where do people go to, to get those skills,

Jeremy: Right. And then you've got things like code camps too. I've seen a lot of people that I've interviewed some people that have gone to code camps and I think there's value there too. I really like the people who have gotten a CS degree and then went to a code camp. I do like that mix of people. But what else ... especially for young people coming out of college, or people looking to change careers, are code camps really the answer you think to those things?

Peter: It's a tough one, right? You can definitely see some benefit with code camps. There definitely is because you do get that quick, deep dive into programming, or computing or cloud computing. But then there needs to be this ongoing program for people to continuously skill up, because it really isn't enough to spend a month for three months, deep diving into something and then you won't become an expert, basically. You need to have that practical hands on experience, and you need to continuously learn, you need to continuously stay up to date.

Really, honestly, it applies to all of us, right? Because I spent eight years at a university, but I have to stay up to date as well. Right? Everything moves so quickly, that continuous education is really key to continuous career growth.

Jeremy: Right

Peter: What's the other interesting thing is that you started saying that, we work university, we learned Java, and we learned memory allocation, but what happened was, you were born, right? You went to school, you went to college or university and then you had a job right? Then you had a job from 25 to 60, or some one thing, and then 65 onwards, you go back to retirement back to play, but things are now different. So an average american will now have 15 different jobs in his or her lifetime.

Jeremy: Wow.

Peter: That's three years per job. Imagine that, and you have to continuously stay-

Jeremy: That's a lot.

Peter: ... learning, it's a lot. It's heaps, right? So you have to continue to stay up to date, you have to learn, you have to get that next job or progression in your career. In our industry especially, if in three years you stop learning, you go out of date, right? You won't get that next promotion. Three years as all it takes if you stop learning to really start going backwards. So being able to continuously learn and figure things out and practice, is really key to being successful.

Jeremy: Yeah, no, I think you actually make a really good point about continuing education. Obviously ... because I can tell you right now, five years ago when Lambda first came out, and I started playing around with Lambda, as soon as you started mixing it with SQS and with API Gateway, and then with SNS and all these other tools that started coming up, I felt like a beginner again, right? I've been doing this for a very, very long time, and I felt pretty good about being able to build web applications.

I started working with AWS in 2009. So I'm like a cloud grandpa at this point, I feel like but essentially, I got to the point where I was able to build pretty good applications in the cloud. Then serverless came along, and it changed everything and made me feel like a beginner again. But I guess maybe one of the questions I have is, because I think if somebody's listening to this, and they are building stuff in the cloud, they're going to say, "Oh, yeah, we need to know the cloud, we need to learn all these things in the cloud and so forth."

Are there people who don't? There are a vast majority of companies who are not using the cloud for anything significant at this point, right? Obviously, the cloud market is pretty big, but the on prem market probably dwarfs it right?

Peter: That's right. Yeah.

Jeremy: For the people who want to hold out, I don't think it's a wise choice but you still need people to work on mainframes, right?

Peter: Look, that's true, but I think cloud is inevitable. It's the future of computing. So yeah, like I said, there are a lot of companies still in that beginner phase, they're still looking at cloud, they're still trying to figure out, "Is it for us? Should we migrate? How does it work?" But really, I think it's inevitable and even mainframes, they will cease to exist one day and cloud will be the thing that we all go to. It doesn't matter whether you're a bank or some other industry that has existed for decades, it will happen.

I don't think there's a choice, right? And what's interesting there serverless, I think, is the future of cloud computing as well.

Jeremy: Right, I agree.

Peter: This is why I think it's important to really understand that progression and know that if you are moving on to cloud provider, what are the serverless options? How do they work? When do I use them? Because if you're moving, does it make sense to just lift and shift? Or should you re architect for the cloud environment? And really take advantage of the power that it offers. So those are a lot of questions and they are hard questions. That's where education again comes in. Because you have to really clearly understand what a cloud provider is, what are the features that you get, what are the trade offs, cost, security, compliance, governance, those are all questions of education.

Jeremy: Yeah, and I totally agree with you. I think you should definitely learn cloud. I just wanted to make sure that I said, people who don't want to learn, give them a give them a little bit of out. But I think absolutely it is just the way things are going to be, and not only the lessons learned in the cloud or I guess the public cloud, but I think a lot of that is going to eventually translate into private clouds as well, right? I think that just the way that Alibaba is doing it, or the way that AWS is doing it, their best practices are going to become the best practices, whether it's on prem or in a public cloud anyways.

Peter: Yeah, completely agree there, right? You'll have your private clouds, public clouds hybrids. But once you learn how a public cloud works, why can't you transfer your knowledge and your skills to the way a private cloud from the same provider works? For example. You have a look at Azure Stack for example, well outpost right, you can have a little AWS center in your own data center. Right? That's cool. So it's awesome that you can transfer your skills there, backwards and forwards.

Jeremy: All right, so let's talk about serverless in general for a bit, because I think that people approach serverless. We talked about this a couple of minutes ago, where it is a completely different paradigm shift, if you want to call it that or mind shift or whatever we want to call it. It is certainly different, right? People need to start thinking asynchronously, they need to start thinking more along the lines of distributed systems. I had this debate the other day with somebody about, is it a microservice, Is it a monolith? All these things.

I'm not even sure that really matters anymore. I'm not sure you can define a serverless application as any one of these things. It's a little bit SOA, it's a little bit microservices, some people like nano services and things like that. So it is certainly a different way of thinking. So, how does somebody who maybe they have a little cloud experience maybe they I've done some stuff with EC2, maybe they've even ventured into containers, for example. What's the first step though, if you are new to serverless?

Peter: I think the first step, as with really any technology, is to get your hands dirty, right? This is what I would tell someone who was new to serverless. Create a Lambda function or an Azure function, or go to GCP, create something, deploy it, run it, see how it works, right? You have to get that initial burst of adrenaline and enjoyment. You have to get something working, activating in the cloud, right? Once you've done that, start building on that a little bit. Add a timer, make it work, make your function run, I don't know on a scheduled event. Integrate it with another service like SQS or SNS to send you an email.

So start building that rudimentary architecture just little by little, just to get that feel, that experience for the power. Suddenly you realize that, "Hey, I've done this. And I didn't need to provision a virtual machine, or I didn't need to create a Dockerfile." It just like code, right? It's code, and I've glued a few things together. It is now working, and it's scaling and it's giving me results. Then I think really, what you should do, is try to build something a little bit more meaningful.

So go find a tutorial, go find the course, something that will actually get you to build a system, right? So you can maybe build an online resume, or you can build ... I don't know, like a really small CMS or something like that, right? But do it practically. Something you can start composing functions together, you can start gluing services together, you'll start using infrastructure as code, because you have to, right? You're not going to be deploying functions manually by hand for too long. That's just going to get crazy and out of control, and suddenly you have to learn, IAM if you are in AWS.

And you'll be like, "Okay, how do I do authentication and authorization? Custom authorizers?" And so things will start slotting into place. But you really need to continuously do things with your own hands.

Jeremy: Right.

Peter: What's interesting for example, is we publish a bunch of courses on serverless. But all of those courses are practical. So you follow the instructor and you do things as that instructor does, right?

Jeremy: Right.

Peter: So we try to stay away from pure theory, because while theory is good, it doesn't really give you that hands on experience that you can then apply when you have to build something in production or for your company or your startup.

Jeremy: Yeah, I agree. It doesn't click until you actually see that work. So you mentioned the courses that A Cloud Guru does and obviously Yan Cui has done a bunch of courses, and Ant Stanley has just launched a new training platform I think and there's some other stuff going on. So there's obviously a market for training. I question is that because the cloud providers aren't doing a good enough job?

Peter: Look, I think it goes to the issue we discussed early, right? Cloud is so vast, it's so big, there's so many things happening. Just look at Kubernetes. I was actually trying to learn containers. I was like a beginner. I was like, "Wow, how do I start? What is all this container stuff and orchestration? And how does monitoring work?" It's just so large, and then you multiplied by other cloud providers, there needs to be that layer of education. So yeah, it's a very interesting space, it moves quickly. So people obviously look for interesting, engaging education from other that they can consume.

Jeremy: Yeah.

Peter: That's what we try to focus on. Right? Really interesting and engaging content that you can learn, that you can practically then try out yourself and do, and that will actually help you in your day to day job. It has to be applicable because we're all so super busy. Right? When you learn what you learn has to then be useful for you later on.

Jeremy: Yeah, I agree. I think that I've seen especially over the last couple of months, I know Azure has been putting quite a bit of energy into this. I know that AWS has grown their Developer Advocate team and they're trying to produce more content, doing more series, doing videos and all that stuff. I and I love it. I love the fact that they're producing all that information, but I also feel like, when you get some of this material directly from the source, right, so when it's AWS telling you, "Hey, here's how you do this, it's use AWS SAM, do this, do ... " It lays out their vision, which I think is good.

I think the sample serverless app that they released a couple of months ago was really interesting. Of course, it was obviously using all their services and doing it exactly the way that they would do it. I think that's very, very helpful. But I also think it's good to get that third party perspective, it's what I try to do with the podcast here to is to talk to people from different clouds, and I know we do a lot of AWS stuff, but at the same time, getting different perspectives and seeing how other people are accomplishing things, I think is just helpful especially depending on how certain people learn.

Peter: Yeah, I think it's fantastic. No single cloud provider can address all of the needs in education, it is just impossible, right? It's too large, it's too broad. For example, our training architects, they are all professionals, right? They all come from the industry and they both had hands on experience and that's what they're teaching, right? They are sharing the experience that they have had. That's interesting, right? I love learning like that. I want to learn from somebody who has actually done it and who can tell me, "Hey, just be careful. There's danger lies, there's a dragon here. That dragon will swallow you if you make the steps. You have to go left instead of going right."

That's interesting. That's what professional education should be about, it should be helpful. But yeah, look, I do love what AWS does, they have awesome people, they create awesome content. The more the merrier. It's great for everyone, I think. You too as well, the content that you create is awesome. It's so fun. It's so engaging. No wonder people come to you and follow you and watch what you do.

Jeremy: Well, I appreciate that. But I would recommend some other training courses. I write blog posts that hopefully make a point here there. But anyways, so one of the other things though I think that we see quite a bit is, you get organizations now starting to adopt serverless. It's great to have some of these one off courses here and there that an individual can take. But really, how do you scale that? And how do you learn at scale?

Peter: That is a key question that people come to us and ask all the time, because it is hard, right? First of all, how do you start? Where do you start? Where do you go once you've done the course? There needs to be a program, a systematic approach to really learning at an organization. At least from our perspective, we always recommend starting off with certifications. I think if you are an organization trying to get into cloud, trying to adopt and use AWS or Azure or GCP effectively, you have to create that common language, a baseline for everybody in the team, that you can then build on as you go ahead.

You and we know that cloud, it's a disruptive technology. Some people refer to it as that discontinuous innovation. In other words, it is that new technology that can solve an existing enterprise need in a new way. However, it does pose challenges. How do you adopted in an enterprise? How do you bring people on board? How do you educate them? How do you show best practice and do it all at scale? And so to be successful, you really need to hit a critical mass in terms of adoption and understanding of what cloud is in an enterprise.

If you take a look at the adoption lifecycle, it typically consists of five segments. So you have your innovators, then you have your early adopters, the early majority, the late majority and the laggards. I think this was originally proposed by Everett Rogers a few decades back, or we have Simon Wardley he talks about Pioneer, Settler, Town Planners in terms of how to think about organizational structure, to support long term innovation. You have your pioneers who come up with crazy ideas Settlers take those ideas and flesh them out, make them real.

Then Town Planners figure out how to scale them. As you go for that innovation lifecycle, the most difficult step is usually the transition between early adopters and the early majority. So to be successful, you need to create a bandwagon effect, in which enough momentum builds up for the technology to become a standard. Now to do that, you really need to have 10% of the population embrace the technology or embrace the cloud, right now example and become committed agents. So having 10% of committed agents in the population, is enough to create that bandwagon effect and create a hockey stick adoption of the technology.

The thing is, if you don't achieve that, you will see early adopters become disenfranchised. That's where you see that attrition of talent. Yeah, you go from your peak of inflated expectations, "Oh my god, cloud is amazing. We are going to do great things." To that trough of disillusionment.

Jeremy: Disillusionment. Yeah.

Peter: That's it. Yeah. By the way, this isn't my idea. This has been talked about a lot by others, like Simon Wardley, Drew Ferment, who is SVP of partnerships at A Cloud Guru speaks about this beautifully. So please have him on and he will tell you about this in great detail. He's fantastic. But to drive the adoption and achieve that bandwagon effect, you have to create a cloud culture right? So you have to create a culture of continuous learning. I think to do that, you need to start by building that common language, which can be created by getting people certified, getting people understanding the basics, getting the baseline, and that will help to create that organizational fluency right?

It will give people understanding of what they know and what they don't know and help them to begin speaking on the same terms, really. From there, you can build that momentum. With people learning more about cloud, you can structure a program, you can help them go from the beginner stages to that expert guru face level, that they want to get to. So those are very long winded explanation.

Jeremy: No, it's perfect. And it's funny, you again mentioning certification. For a very long time, I was on the fence with certification, because it was one of those things where, if you go out and you do the work, you're learning this stuff. The certification. Yeah, a stamp of approval. This was early on, right? This is earlier on with AWS when they first launched their certifications programs. I don't think there were a lot of job opportunities necessarily for people saying, "Oh, well, you need to have this AWS certification. You need to have this other certification."

But I think what I've come to realize over time is one, like you said, this common language for certification, I think is extremely important within an organization. But it also is a really good way for you to figure out what you don't know, right? They say there's like three learning, is the stuff you know you know, the stuff you know you don't know. Then the stuff you don't know that you don't know.

Peter: That's right.

Jeremy: The thing with cloud is it's just so vast like you said, there are a lot of things that you don't know that you don't know. Right? Taking one of those tests, even if it's just a practice exam, to show where those gaps are in your cloud learning, I think that's a really powerful argument for at least taking a certification course, whether you go through with it or not, I think it at least fill those knowledge gaps for you.

Peter: Agreed, and I think there are different advantages for individuals versus organizations.

Jeremy: Sure.

Peter: So for an individual, I'm completely with you right? It is awesome for figuring out where your gaps are. What you really know and what you don't know. I'll tell you a quick story. So I at one time decided to get certified and do three certifications in one week. So I decided to get the free associate certifications for AWS, the solutions architect, developer and SysOps, right? So I'm a developer, I come from a developer background, and I was like, "Should be a piece of cake. I'll just watch our courses. I'll be ready. I'll do practice exam I'll go in." So I went in and I did the three certifications.

With a two day gap between each one. I scored the best on the developer exam, because it was DynamoDB and Lambda and all the cool stuff that we love and talk about all the time. Then I scored poorly ... Well, I still passed but I scored the lowest on the SysOps exam. I realized that, "Hey, I actually don't really understand or know very well, some of the Sys admin sides of AWS." And I'm like I was making mental notes in the exam was like, "Okay, I need to look up this. I'm not sure about this, so I need to look it up as well. Then I need to go and really refresh this particular line."

It was eye opening, I had enough skills to pass, but I was like, "Wow, okay, I need to get back in and really continue to learn." I think another benefit of doing a certification is that, it forces you to do things that you normally wouldn't do. So even with a developer exam, you have to go and you have to try Elastic Beanstalk. Yeah, maybe you won't use it, because we're all serverless.

Jeremy: Yeah.

Peter: But it's good to know what it is and how it works and why it's there and when is a good case for it? And then I had to do ECS and deploy a container and use ECR the Container Registry, I was like, "Wow, nobody would never touch it. But my god now I know how to deploy a container and I know how Fargate will scale it and I know some of the properties." So it gives me a much more holistic understanding and perspective about cloud and the entirety of the experience. Then if I go back quickly to organizations like we said, and I know I'll get into a lot of trouble at work for saying this.

So maybe you should cut this out. But I don't think that certifications are a goal into themselves, right? The goal is really to create that organizational fluency. It's to give people an understanding of what they know and don't know, and to begin speaking on the same terms. And at the same time, I think they are a great yardstick for measuring the overall cloud fluency of an organization, right? You can actually see how much sure it is, by looking at how people are certified, at what level is the associate level is at the professional level. So there's a bit of an indicator there as well.

Jeremy: Yeah. No, I think that you made some really great points in there. Especially if you've ever had to use like ECS or deploy containers, things like that, that should be the argument for why you would use serverless, because you wouldn't want to have to do those things, right?

Peter: Exactly right. I was like, "Wow, that was fun. But my God, I want to go back to my function. I just want to focus on the code. I don't want to care about provisioning anymore." You see two instances. Well, luckily, there is Fargate.

Jeremy: Right. Exactly. All right, so let's move past just maybe serverless in general and just talk about I guess, education in general. So this is a conversation that you and I have had where, I think the value of going to college ... and this is maybe going to get way off the rails for some people, but the value of going to college and learning, getting a computer science degree, those things. I think there's a tremendous amount of value in a very practical degree. But I think there's also a tremendous amount of value in technical degrees.

By that I mean electricians or dental hygiene, all those things, that technical education hands on experience, as opposed to taking a bunch of general education classes and learning about the history of Western thought to 1600 or something like that. Some of these more general classes, and not that there's not value in those but, I think that especially for people who are maybe past college, and they want to change careers, they want to learn something new, they want to become a professional at something else, maybe they want to become a plumber or some other type of technician, or maybe they want to get into computers and they want to do some programming or something like that.

I don't think colleges are fitting the bill for a lot of those really technical hands on things that we need. This probably going to end up being the future of a lot of what we do. So where do you see the future of education going in general?

Peter: Jeremy, I think that is a great question. I think the future of education will be very different. I think this applies to all types of schooling, whether its primary, secondary K through 12, tertiary university college level, vocational education, or the professional, ongoing adult education. I think there are going to be a number of areas that will really change in the way education is delivered. How students are approached and what it's done. And let's talk about it in a second. But I think traditional educational organizations will really need to adopt and move quickly.

Because if they don't, they will fall behind. And they will find it harder, much harder to compete for students in the coming decades.

Jeremy: Yeah.

Peter: Just top of my head, I think if you want to succeed in education in the future, you have to do this right? Number one, you have to prioritize the student, and focus on the student experience. I think we can say safely that a lot of educational organizations don't really do that well, and we can do much better, all of us right/ As an industry. We need to go and start creating curated and personalized education for each individual. Let's say Jeremy Daly comes into my school and wants to learn something.

Why can't I make an assessment of Jeremy's skills and create a personalized curated path just for Jeremy? Why is everybody learning the same thing?

Jeremy: Right.

Peter: Not accounting for the person's strengths and weaknesses. Maybe Jeremy is awesome at math, but he needs more help with physics. So why can't we change our curriculum and deliver it in a way that will really help Jeremy, versus Peter who needs a different kind of help? I think next it's the deconstruction of the value chain. I think honestly, some of the top tier universities, they have a bit of a monopoly on education or that's how at least people perceive it. I think we need to break it down, I think we can really democratize education, and it doesn't matter whether you are in US or Australia or India or China, you should have access to great high quality education, regardless.

And it should be affordable, you should be able to do it and have a meaningful, successful life. So this is what we're really trying to achieve here. Subscription based learning, that basically talks about having this ongoing access to high quality resource that you can always tap into, right? Now we spoke about three years for a new job all the time. So you need to have access to a place where you can go and consistently find best practice and help. Just in time options, being able to consume education from your mobile phone, when you are busy.

Maybe you are on the train or bus and you want to look something up, you should have a resource to do that. I think another major thing is that educational organizations, they need to work with businesses, right? They need to really curate education and make sure that it's delivered for what the industry needs. It has to be relevant to the job market, and to what companies and organizations expect. Then of course, lastly you said focus on quality and currency. I spent eight years at a university right? So I spent many, many years and I did love it, right? I love doing computer science. I honestly did not want to get out.

But I know that a lot of what I learned was out of date. As much as I loved it, it really it set me up for research. I could have stayed and I could have done a lot of academic research, but it really didn't equip me for life in the industry and then I had to really quickly skill up. So being able to focus on the quality and currency and match the expectations of the industry. We have to all contribute to that. So yeah, that's a lot to get done. There's a lot to get done for all of us, I think.

Jeremy: So you talked about a lot of different things. I was taking notes as you were saying these because I'm like, "Yep, yep, yep, I agree with all these points." But one thing to know of out of date information, even the rules of beer pong in college has seemed to change, which is really weird.

Peter: Oh, no. Really?

Jeremy: I know. I went to a wedding, it was an after party for wedding but I was thrown by these changing rules-

Peter: [crosstalk].

Jeremy: But seriously though-

Peter: That's crazy.

Jeremy: Well, you and I will play old school next year when we [crosstalk].

Peter: Okay.

Jeremy: But you made some really good points and one of the points you made was about personalizing the educational experience. Now I can't tell you how much I have been just beating the drum on this because my wife is a third grade teacher, and they piloted a program at her school a couple of years ago back, where every student had a little tablet, and they would take their math tests on the tablet. Then what it would do is it would give my wife on her computer an immediate tally, of which questions people were having or which questions the kids were having trouble with.

So she could see, okay, they're having problems whatever it was, multiplying fractions or something like that, or some concept that they were having trouble with. Then she would know as a result of that, that she could teach, review the things that were the weaknesses for the majority of the kids. But that only goes so far. You're still got kids who understand that, maybe some other kids have weaknesses on something else. The ability for these systems now ... and this probably ties back to serverless I would say, is the ability of the scale, I guess, the capabilities of the cloud for you to build these machine learning technologies and the ability for you to hone in on what it is that somebody doesn't understand.

And then even what I was saying about that different perspective, maybe student A is a visual learner, maybe student B is more of a reading or hands on learner or whatever. If there is a way that you see which types of education, which types of content they respond better to, where their weaknesses are, and can feed that information to them, so that you strengthen them on those individual pieces, that's huge, right? That is, to me the future of education. That's not something that you're going to do, sitting in a lecture hall with 300 kids for a general education class.

The other thing you mentioned too about this idea of I guess quality versus quality and currency, if you go to a university now, especially in the United States, it's different in other countries, but you are walking out with a degree and likely $50,000 in student loan debt, right?

Peter: Yeah.

Jeremy: And for a couple hundred dollars, you can get a subscription to A Cloud Guru for the year or LinkedIn Learning or any of these other online training platforms that are out there. Probably A Cloud Guru. That's probably the one you want to go to. But honestly, I learned a lot of stuff in college. I know I did. Of course this was 20 some old years ago, but, when I got out of college, it wasn't until I started reading blogs and watching videos and getting hands on and doing stuff, that I actually use most of what I know today.

So this idea of being able to constantly feed education to people, in a way that doesn't put them into debt for the rest of their life ... not that college isn't an experience that maybe you have to have, but I think for a lot of people that, access to this constant education is just absolutely game changing. You're going to learn more in a six hour course on A Cloud Guru or in one of these other training platforms, then you might learn an entire semester in a college course, that you're paying thousands of dollars for.

Peter: Yeah.

Jeremy: So, I do think that's really, really interesting.

Peter: Yeah, I completely agree with everything you said. I think that's going to be what makes education effective in the future. It's that curated personalized education. You spoke about A Cloud Guru, we have full time training architects, instructors, who what they do every day, right, is they create content. Whenever anything changes, they update it. Right?

Jeremy: Right.

Peter: So when you go to the platform, you know that what you're getting is the latest version. You're getting that latest best practice. So suddenly, what you are learning in a lecture hall, right, doesn't really match what you could be learning online because, that content is much more up to date. So that's an interesting aspect as well. Yeah, the currency and the quality. Yeah. Because we can continuously iterate on it. I think that universities too have an important function that cannot be done with just an online delivery of education.

That element is really that ... It's going to sound harsh, but it's babysitting, right? Because just after you finish school, right? There's still a little bit of time for a lot of people to mature, right? They need to go through that maturation phase. Going to university, going to college allows people to do that, right? It allows them to build social connections. It allows them to learn how to work in a team, maybe better than they did at school. So it gives them that opportunity to mature before they go into the industry.

As much as I love online education, and I think it is the future, there is that element that still needs to be solved, that social element. But I think we'll figure things out, maybe it'll be some blended learning, where you do get that up to date curated delivery of education online. Then there's an additional element where you go and you socialize with your peers. So yeah, we'll see how all that pans out.

Jeremy: Yeah, totally agree. And you mentioned to this idea of tying universities or training too the job market, that is one of those things that's if you're a large corporation, and you're looking to hire developers, or you're looking to hire whatever it is, you have specific needs, you have things that you want these people to know. I think that there are a lot of overlaps between most corporations in terms of ... and again, not just corporations. These are small businesses, these are startups, these are other companies that are going to have the same type of needs.

That information has to get back down to the people who are creating the education materials. I don't think a lot of colleges want to hear that. I know there's some partnerships. I think AWS is doing some partnership with universities and trying to create curriculum for them and things like that. I think that's a great start. But I don't think we're anywhere near where we need to be, like you said, there's a ton of work that's left to be done.

And just from if you're already in a job, if your company is not giving you resources to go and learn, to get continuing education, whether it's a subscription or something like that, or they're doing regular trainings, or hackathons or something, that they're giving you some time to keep up to date, you're doing your company a disservice, right? You have to keep your employees learning because that's the other great thing too is, just motivating them to stay in their job be like, "Oh, I got to learn about ... I don't know, EventBridge or something like that and that was cool. I learned all these new things."

And maybe there's something we can do that integrates that, that will make our company better and our product better and deliver better value to our customers.

Peter: Yeah, and frankly, it's a business risk for the company as well. Right? You want to have your people, be as knowledgeable as possible, to have expertise to know what's available and how to use things and what is actually best practice. Because you have competition, right? And that competition is trying to do what you do better and faster and cheaper, and they are working hard at it. So if you're not giving your people a chance to learn and figure out what is the next thing, you're really doing your business a big disservice.

You know what's interesting, I think some companies think that, "Hey, if I give our stuff education, right? They will learn, they'll get certified, they'll get awesome, and then they will leave." And yeah, sure, it could be a risk. But I think there's a bigger, bigger, much bigger risk, to have your stuff without that education. Not give them the opportunity to learn, not give them the opportunity to lift your entire company, to lift your entire organization. We've dealt with those new skills. It's massive.

Jeremy: I totally agree. Totally agree. All right, so let's maybe close this off with some advice for learners. So we talked about how to get started with serverless early on, but maybe just general advice for people looking for the type of stuff that A Cloud Guru produces?

Peter: Yeah. Look, here's what I would do, right? If you are wanting to get certified, that's your goal. So here's how you could study. If you are a complete beginner coming in. Go to A Cloud Guru, do a certification course, watch the entire course. Follow along, do everything that the instructor does, all the practical labs, read the white papers, go through all the suggested blog posts and resources, just do as much as you can right? Then do a practice exam, and see how you have faired. See where your gaps are.

Because we have practice exams that can help you identify, "Hey, maybe my PC knowledge isn't really good. So I need to double down on that." It's really good if you don't have that much background. If you are very experienced, right? So let's say Jeremy, you want to go get certified, get that professional level certification? I would do the other way. I would actually take a practice exam first, right? I would use that to identify my gaps, right? "Hey, I'm great at storage, and I understand S3 and security, but I need a little bit of extra help with ... I don't know subnets and routing tables.

Peter: And then I would focus on that area because frankly, unless you have heaps of free time, if we need to be very efficient with what we do and what we study. So yeah, I would use that practice exam, to really help me figure out the areas, the gaps that I have and then go and double down on them. Then I could do the practice exam again and figure out, "Okay, is there anything else that I need to know? Or am I ready for the exam?" So these are the two kinds of different strategies and, they work relatively well for different kinds of people.

Jeremy: Mm-hmm.

Peter: If you just want to learn serverless, then like I said before, go do something practically, deploy function, hook something together, use SNS, SQS, send an email, do something right? Then get a course and build a bigger system. Continue building, continue doing things practically and hands on. And then share.

Jeremy: Oh, yeah. I agree.

Peter: I find that the best companies that create that cloud culture, those companies create awesome internal communities, help people share what they've learned, right? To get people contributing blog posts and ideas and suggestions and going to the community, tweet. Create a gist of what you've learned and share that around, and that's awesome. It gives you a lot of satisfaction, and actually promotes that knowledge in your brain as well. So just go and do it.

Jeremy: Yeah, and actually I totally agree with you on sharing. That's one of those things where, I know for me when I first started doing some blog posts that, I basically would write something down I'd be like, "Wait a minute? Is that right?" And then I would do a bunch of research to make sure that what I was saying was right, and then from there you learn more right? You learn more by writing that stuff and sharing it and putting it out there, and if you write a good post about serverless, send it to me and I'm more than happy to amplify that the best I can in the Off-by-none newsletter. And share that with people but I think one of the-

Peter: I love it. I think Jeremy, like what you said is spot on. If you can clearly articulate and explain the concept to somebody else, then you have really understood that idea yourself.

Jeremy: Right.

Peter: That's how you test right? In a few sentences, explain something to somebody who doesn't know what that is, then yeah, you understood, you learned you what it is now. That's how you can really check. So sharing, creating your own knowledge and sharing that is key to really validating that you have learned that material.

Jeremy: Yeah, and I think one of the other things you were getting to the point of with depending on what you're coming into, or if you're getting certified ... I guess what your goal is, right? So I think when you're trying to educate yourself on something, you should really pick that goal, right? Then target content for that goal.

Peter: That's it. You have to have a goal. Yeah, It drives you, right? A goal ... If you have something in front of you and you need to achieve it, that's important. Yeah, certification is great. By the way, if you don't have a cert it's a cool experience. But, come up with a project, do something fun.

Jeremy: That's true. Side projects are always great.

Peter: Side project. Yeah, exactly. Go in GitHub, you can contribute to open source, or maybe you just want to build a game. Or maybe you want to build a little platform yourself. Because once you start building, you'll be learning. It will push you to do more, more and more. Then yeah, you can share with us and we'd be happy to learn from you. That's great. That's why people have GitHub repos. That's why they create all these projects. That's what you do it. That's what I do. It's very effective.

Jeremy: Totally agree. So one last thing. So A Cloud Guru just recently acquired Linux Academy. So what's that all about?

Peter: Yeah, so look, we are joining forces, A Cloud Guru and Linux Academy. Well, it is great. There's going to be a lot of great content for you whether you want to learn AWS or Azure or GCP or Linux or you want to go into Kubernetes and containers, basically we're trying to bring best of both worlds together, into one. So yeah, watch the space, there's a lot of cool stuff that will come out. Not giving you dates. But I know our teams are working very hard. So watch the space.

Jeremy: Yeah, well, there's great content on both platforms. So it'll be really interesting to see those all merge together.

Peter: That's it. Yeah.

Jeremy: Anyways, listen, Peter, thank you so much for taking the time to rant about education with me, and go off the rails about colleges and universities. But seriously, you're a serverless hero, you have done a ton of great work for the community. You have a book and there's some other things so if people want to get in touch with you or find out more about some of the other things you do, how can they do that?

Peter: Look, my Twitter, LinkedIn, anybody can connect, please connect. Let's talk. If you have any questions about cloud education, generally serverless, please get in touch. I'd love to talk to you. Yeah, you can find me at various conferences and events throughout the year as well. Hopefully Jeremy we'll get to hang out very soon at a summit or an event. So yeah, please connect and yeah, happy to talk to you, to anyone, at any time.

Jeremy: Awesome. All right, well, I will get all of that contact stuff in the show notes. Thanks again, Peter.

Peter: Thank you, Jeremy.

View Details

About Suphatra Rufo
Suphatra started her career at NPR and PBS stations around the country, and quickly found her way into technology. She worked on social good initiatives like Microsoft’s Imagine Cup, a competition for young inventors; We Day, working with Selena Gomez to advocate for more young women to learn how to code; and TEALS, a program that places industry engineers in high school classrooms to teach computer science.

She has deep product experience and led the effort to create a nonprofit SKU for Office 365 and Azure and bring cloud computing as an upsell to the social sector to 93+ markets and realize a new revenue stream for Microsoft. She was part of the original team that built Microsoft Teams and saw the product from Preview to GA, all the way to v2. She worked at the forefront of cloud computing at Amazon Web Services, managing their $6B database category's developer advocacy and customer storytelling efforts. Today, she heads up solutions marketing at Couchbase, a late-stage VC-backed cloud database startup in Silicon Valley valued at nearly half a billion dollars that develops open-source, NoSQL, multi-model, document-oriented and key value databases.

  • Twitter: @skprufo
  • Couchbase: couchbase.com

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with Suphatra Rufo. Hi, Suphatra. Thanks for joining me.

Suphatra: Hey, Jeremy. Thanks for having me.

Jeremy: You recently became the head of solutions marketing at Couchbase, so why don't you tell the listeners a little bit about yourself and what Couchbase does?

Suphatra: Hi, I'm Suphatra. I'm head of solutions marketing at Couchbase, which is a small start-up in Silicon Valley that develops open source NoSQL multi-model, document oriented and key value databases. We've raised $155 million in funding and we're valued at nearly half a billion dollars. As head of solutions marketing at Couchbase, I create the company's market strategy, sales plays across different industries and solutions, and I handle all of our compete scenarios. In my typical day to day, I'm usually looking at complex technical and business challenges and trying to diagnose how we can create solutions around that, and working with our engineering team to influence product road maps so that our solutions can be integrated in features and helping our business teams determine our next go-to-market investment areas.

Jeremy: Awesome. All right. You have a ton of experience and you have a very impressive resume on marketing cloud databases... Is maybe a good way to say it. I'd love to get some insight from you into how companies, especially enterprises, are looking at migrating data to the cloud and moving away from maybe more traditional on-prem type installations. I guess maybe the best place to start is, I think most people know what relational databases are, that's a pretty common thing. And I think people have a sense of what NoSQL is, people might be familiar with DynamoDB, MongoDB, Cassandra, those sort of things. But maybe you could just give us a little bit of background on what modern NoSQL looks like.

Suphatra: Yeah, yeah, like NoSQL 2.0, but I'll just start from the very beginning too, because I think a lot of people are confused NoSQL still, which is funny because it's been around for almost a decade at this point, but NoSQL is essentially a different kind of database that doesn't rows and columns. One good example is if you think about an application like Snapchat, on New Year's Eve, millions of people want to use Snapchat at the exact same time. So 11:59 PM, millions of people get on their phone to use Snapchat because they want to capture the exact same picture at that exact moment. So Snapchat, as an application, has to be built in a way to accommodate for a very sudden and huge surge in performance for a very brief moment of time, and then scale right back down... Because once people take that picture of them kissing their loved one when the bell rings, or the ball drops, I should say, then they stop using Snapchat, so then that goes straight down.

What NoSQL databases are great for is they can handle those types of really heavy spikes because they can scale up and down really easily because they aren't constrained by rows and columns like a typical relational database. That's what the NoSQL databases really offer. Since NoSQL databases were invented a decade ago, they've really branched out to lots of different types of NoSQL. Now you have document databases, you have adjacent documents, key value databases... Couchbase is cool because they do both of those things. When I was at AWS, I helped with stories about DynamoDB, which is specifically just key value database, which is also really strong database as well.

Jeremy: The thing that's interesting about NoSQL, and we're hearing more and more about it, there's a lot of different companies that are offering solutions for it. And more importantly, I think there are companies that are starting to adopt... And specifically for the workloads like you talked about, that New Year's Eve... Billions of records or billions of transactions in a very, very short amount of time, but is this something that you're seeing companies, maybe not just your start-ups and your Snapchats, but you're seeing other companies start to adopt?

Suphatra: Yeah, yeah. I think the way that consumers behave, the retail industry is a good example. You probably didn't know that Sears, Kmart, Barneys New York, Party City, I can name a dozen more retailers that just last year, either completely closed down or had to significantly reduce their number of stores, just last year. It's because retail isn't done the same way anymore. Those spikes are now a common part of life and people are having a hard time figuring out how to handle it. Tesco, which is the largest grocery chain store, I'm not sure in America or in the world, I'll have to check that... But they, in 2014, crashed on Black Friday because they couldn't handle the spike in the demands they were getting online. So they lost an entire day of business on Black Friday because they couldn't handle that workload. And then the year later, they went on a NoSQL database and now they can handle that load.

I think what people are seeing is that normal day to day business operations are fundamentally different. For example, the fashion industry used to have only four clothing seasons. Your mother probably remembers buying a new outfit every season... So winter, spring, summer, and fall. And so women's clothiers would go and create new clothes four times a year. Now the fashion industry has 52 seasons, so every week is a different season of women's clothing, which means there's a spike every week for every launch of every new clothing line. So that's another big database problem that's now just becoming a regular part of life. A decade after NoSQL databases are invented, it's really not a new invention anymore. Now this is just the way of business.

Jeremy: Wow. I can't imagine buying something new every week. I buy a new hoodie maybe once a year or twice a year or something like that. That's the extent of my fashion choices. I think that's really interesting. I think that's where everything is moving, is that just you have to become global now, right? You have to be able to handle these workloads that are just gigantic, and obviously there's some major players in this space. We have AWS and we know of things like DynamoDB and now some of the managed services they have. Google still has their big table and a few other things like that. Obviously, you have a bunch of these start-ups and I guess start-ups that are much further along like Couchbase, but what is the concern there? I mean data itself is a huge lock-in problem, right? As soon as you put data, and you got terabytes of data somewhere, you're kind of there. I mean, do you see vendor lock-in with some of these NoSQL players? Do you see that as a major concern?

Suphatra: Yeah. I think this is where things get really interesting. When I was at AWS, I worked almost exclusively on my creations off of Oracle and Azure at AWS. And a database migration is, by and large, the most difficult thing that you can do in cloud computing. It's really hard. You've got to do a lot of data modeling. You've got to do your schema conversions. I mean, it's really just a ton of work and what I have found is that when people are charged with, "All right. We got to migrate our database." We tend to do it in multiple phases and that will take multiple years, so oftentimes they'll first just re-host. Let's say they're on Oracle. They want to get off Oracle, but they don't want to be penalized. So they take their Oracle license and bring it to a different cloud provider. They keep all their data with Oracle still. They're just moving it. That takes six months to a year, then afterwards, they say, "Okay. Well, I think we're now going to replatform." And that's a whole nother workload and that's even more work, and even harder down the line is refactoring, which is where they might actually go from a relational database to a NoSQL database.

It's much more rare that you see people to a database migration where they go from a traditional relational database on one provider to a NoSQL database on a another provider because it's a really difficult piece of work. When people make the decision which database they want to move to, it's often a very permanent choice. One interesting thing that I've seen is people choosing not to go with public clouds which I think is a bit interesting and we see that with the hybrid solutions set Azure and AWS are coming out with with Outpost. I think that's a really interesting trend that I'm definitely keeping an eye out for in upcoming years.

I think that vendor lock-in is generally a problem that's going to be most relevant to really big enterprises. I think the smaller guys, I still find that they're pretty agile, smart and quick. They figure out how to do things pretty quickly. I think for the lock-in, my recommendation would be lots of NoSQL databases, Mongo, Couchbase, DynamoDB... They support a certain amount of data for free on their databases. So for most people, you can actually get that great performance completely for free as long as you have a small amount of data. And as you start to grow, that's when you might be more pressed to make a smart purchasing choice. Then you might say, "Okay. Do I want to do a data migration? Am I really comfortable with this database or now that that we have more needs, do I need to pay for it? What do I want to pay for?" But for the most part, I would say for the life of an organization, that free level is going to be totally fine.

And where the vendor lock-in happens is something... What I found was at AWS, it was like going to Target, where you go and you say, "Okay. I just need to buy a pillow," but then you see all the other cool stuff and Target and then by the time you've left the store, you've spent $300 dollars on towels and clothes and gadgets for you family. I think that's some of the stuff that happens with AWS. Take, for example, DocumentDB, which is a document database... It doesn't have Eventing. It doesn't have full tech search, so if you need those things, you need to purchase additional resources to get that, which means you think you're buying DocumentDB, but now you also buy X, Y, and Z. DocumentDB scales by storing its data on S3, so now you are also buying S3. What happens when you go with a provider that does a purpose-built database, is you have to be comfortable with buying a lot of different services. And Azure's a bit different where it's a little bit more all-in-one, right?

Microsoft's very good at the whole all-in-one thing. I worked there as well for a long time and they love just bundling stuff together. The downside of that is you never know what you're paying for it because you just get a lot. You get everything.

Jeremy: Yeah.

Suphatra: So that's different from Target, whereas Target you and you don't know what you end up buying. I liken Microsoft to the really nice seafood buffet restaurant in town where you're spending good money because the seafood's going to be good. It's not going to make you sick and you could have as much of it as you want... But after you've paid, you realize you didn't eat as much as you think you should have and there's that sense that you're not getting that price that you deserve, which is why Amazon has a good edge on Microsoft, because Amazon says, "Oh, we're going to cost because we let you choose. Pay as you go." Microsoft says, "Well, we're good on ease and convenience because you get everything you want all in one." So that's some ways to think about those two big providers.

Jeremy: Yeah, that's actually interesting way to think of it because I've always thought of AWS to be very additive, a lot of pick and choose the low level components that you need, and I like that about it... But when it comes to airlines where you pay that baseline ticket and then if you want to bring a bag or get a better seat, pay more money. That always drives me nuts there, but I do like that control. I think that's really interesting. You mentioned this idea of re-hosting, replatforming, and refactoring, right? The migration piece of somebody moving into a public cloud or moving into one of these other tools obviously is a long process, takes a while to do. And so when they get to that refactoring point and they're starting to get rid of Oracle, for example, or get rid of Microsoft, which AWS migrated everything off of Oracle... Do you see the role of Oracle and Microsoft SQL Server, things like that, do you see those starting to go away? Is that a dying market?

Suphatra: No, I don't think it's a dying market. I don't think people are migrating as fast as the hype makes you believe and you can kind of tell... At re:Invent last month, Andy sort of alluded to that. He was annoyed that people weren't migrating off of on-prem fast enough because it just takes time and people want to be smart and careful. And you would be surprised at how much both of those companies go through switchbacks, right?

Jeremy: Yeah.

Suphatra: And that's expensive, if you're a company that's like, "Oh, I want to do this," and then, "Whoops, my bad. We're going to actually go back." Yeah, I mean you can imagine that happens with Redshift and Snowflake.

Jeremy: Okay. Yep.

Suphatra: Yeah, right? It's an expensive process, so I would say that I think it's a dying market maybe for Oracle, and I don't know honestly if that's a technology reason. I think the reason for that is the business practices are ones that consumers are not willing to do anymore because the terms of the market have changed. When Oracle was the only player in town, they could treat you any way they wanted, right? It's like the Mafia, someone running your neighborhood, they can call whatever shots they want, but it's not like that anymore. Now there's multiple players.

And I spoke to so many customers at AWS when I was pursuing these stories about people migrating off of Oracle that would say, "We got audited and found out that we were using features that actually cost us money. We had no idea, so then we got the bill that is six or seven figures. And then they tell us, 'Hey, if you signed this contract, we'll alleviate this bill,'" and then you're locked in for further. Those types of business practices that worked in the past, when you're competitors are not doing that anymore or actually not doing that at all or ever have, and people say, "Well, I don't have to go through this," and go over here, it's just not going to work. So I do think tech in Oracle is not dying, it's just their business philosophy.

Jeremy: Yeah, and I think there's so many more options now, right? I mean, I remember way back when when the commercial databases that I licensed, early 2000s or late 1990s, we were licensing SQL Server, and you're paying for that license to use that and to install it on one server. And then if you need to grow, you have to buy a new license. You have to scale back or whatever... You're always paying more money in order to have all that extra licensing. And this model that... I'm sure not AWS who introduced it, but this idea of just paying per hour for a database, for example, and for the licensing that goes along with that, and you can turn as many on as you want or turn them off whenever you need to, is really a much different approach, obviously, like you said, to that mob boss mentality of Oracle maybe.

Suphatra: Yeah.

Jeremy: And not saying anything bad about Oracle, I mean I'm sure nobody ever says anything bad about Oracle, so we'll keep that level of decorum here as well. All right.

Suphatra: I did want to follow up on that too, because a lot of people say it's a dying market. They say, "Oh, well, Oracle and SQL Serve, people are leaving that." People love SQL. They love SQL.

Jeremy: Yes, they do.

Suphatra: And I was at the PASS conference last month and PASS has 30,000 members. Those are 30,000 SQL administrators. I mean that's a lot of people and that's just the people that are SQL administrators that have joined an organization about it. There's probably a bunch that they didn't capture.

Jeremy: Oh, yeah.

Suphatra: And they really love it, and I don't think obviously SQL is going anywhere. That was one thing I really liked about Couchbase, is that query and know SQL database can be really confusing and frustrating because sometimes people treat it like a data dump so it's hard to find information in there, and Couchbase has their own query language called N1QL, and it is essentially SQL for JSON. For example, if you were to use Mongo or DynamoDB and you wanted to query something, it would take you literally hundreds of lines of code, about 234, I believe... But on N1QL, it takes seven lines. I was like, "Wow, that's actually something different and new."

One thing I learned when I worked at Microsoft and at AWS, and when I was working on these stories about Oracle and these big companies, is I feel like I was often in this situation where I was seeing our teams build something that someone else had already built. We were just trying to build something again and compete against them. For example, the Amazon-managed Cassandra that came out last month to compete with DataStax and open source Cassandra and those sort of things, that I think I just start to feel like, "You know what? I think it's time for me and my career to go somewhere where the technology's net new. Someone invented it. It's brand new."

Jeremy: Yeah, yeah.

Suphatra: I do think it's a valuable business play and it's been done forever, to go and copycat something and compete with it in the market, but I also thought it'd be exciting to try some new technology as well. I think what I've seen with SQL and with the move over to NoSQL is it's not that people are abandoning SQL, the rate of SQL to NoSQL use is about 90/10, so obviously SQL had a huge advantage. But what's happening now is people are using SQL databases with NoSQL databases, so those numbers-

Jeremy: That hybrid approach, right?

Suphatra: Exactly. The hybrid approach, because they're saying, "Okay. Well, I don't really understand the NoSQL thing, but I want to venture into it, but I'm certainly not going start migrating all my stuff over." Like I mentioned, refactoring is very difficult, so they're just using both and I think that's really great because... One good example is Pokemon that I worked with at Amazon. They were using Aurora for their authentication and they're using DynamoDB for their botnet strategy, all-in-one, for their log-ins... Because DynamoDB, you can set a time to live. They were able to kill botnets within five minutes. They reduced their botnet issues, their bot log-ins, by 93%.

Jeremy: Wow.

Suphatra: And they have a really great video about it on Twitter, I think I tweeted it. I could also send it to you, but it's really funny. It explains how they defeated their botnets. And then they're also using Amazon Aurora relational database for the rest of their authentication process, so it's that kind of pairing of two databases of two different schema types, that make magic.

Jeremy: Yeah, and I actually think that is something where you just have to move to. And that's why I feel like some sort of really large commercial database like an Oracle or like a Microsoft SQL Server, is not needed at the scale it used to be needed at, right? You get Amazon, who is doing... I don't know, gajillion transactions per second or whatever they're doing, trying to run that on Oracle and it's basically starting to choke. They have to just keep making the clusters bigger and keep adding more servers and just keep scaling up and scaling up in order to handle that because of the complexity of the queries in a relational database. Whereas when you move to NoSQL, then the complexity of the queries stays the same regardless of how much information there is and as long as you understand your access patterns.

That's why I think that you're still going to need relational databases because you're still going to need to do analytics and you're still going to need to some of these other things, but as we are starting to talk about scale, where the growth is probably going to be in the market, I think that's going to be with NoSQL where those are the things that are going to be able to support this massive data from an operational standpoint. And then maybe you take that hybrid approach and use SQL Server or, I should say, a relational database as your reporting or some sort of data warehouse type thing to run analytics on, but certainly wouldn't need to run at the transaction scale as something like a NoSQL database would.

Suphatra: Yeah. No, I think you're exactly right, exactly right. Spot on. That's why I'm here. You're a smart guy, Jeremy.

Jeremy: That's why you're here because you're smart as well. I just wanted you to validate me, is basically what I wanted you to do.

Suphatra: Done. We can sign off now.

Jeremy: The thing you mentioned a little earlier too was about purpose-built databases and I think this is an interesting strategy that Amazon is doing. I mean, I get it. If you need a time series database, then use a time series database. If you need a Blockchain for some reason, then use a Blockchain, Quantum Ledger or something to that effect. What are your thoughts on these purpose-built databases?

Suphatra: Yeah. It's funny because I can't tell you how many times people come up to me and say, "What does that mean?"

Jeremy: Yeah, that's a good question.

Suphatra: It a weird word, it's not a real word, right? Yeah, it's good. I really liked it. The old way of database was you just set it up and you threw your data in there. A purpose-built is trying to get you to... You're using a database because you have an application and that's the database that's going to be used for it. Another way to think about it is when multiple databases are being put together for one application or purpose, like what I mentioned with Pokemon. I think that strategy is good.

I would recommend for folks... The only thing is there is a little bit of learning curve if you're going to be stitching multiple databases together. You got to figure out how to make them work together, and then also, your cost is going to go up oftentimes, so you got to figure that out. But I think the most likely situation is that only large enterprises are going to be using multiple databases. I think smaller companies will probably find one that is just suitable for their needs... Because I doubt that a lot of really large businesses are dealing with a 93% bot log-in reduction they need to deploy, right?

I would think if you are a small accounting business, one database is probably fine for your needs. If you are a small retail company selling quilts, you're probably not dealing with big spikes. For most cases, again, I think... What I said, SQL's not dying. I think it's going to still be really relevant. Examples I see most common are in travel and hospitality, retail, obviously... Retail's the easiest one, and finance. And those three, I would say, is where if you are in one of those three categories, then you need to be doing purpose-built databases... Definitely for sure NoSQL databases and you got to learn how to pair relational and non-relational together. One good example here, and travel and hospitality is a good one because... I'll give you an example of Carnival. Carnival is a customer of ours at Couchbase here. They run 20, 30 nodes with us and they really take advantage of our Eventing, and I know serverless, you guys love Eventing.

Jeremy: Yes, we do.

Suphatra: Eventing lets them do... I've actually never been on a cruise. Let me preface with I've never been on a cruise so I will explain something to you, an experience I've never had, but it sounds very exciting. The engineer was talking to me about this the other day because I was like, "Hey, man. I'm going to do this Serverless Podcast. I'm pretty sure they're going to want to talk about Eventing." He's our Eventing guy and he says, "Oh, I got a great Eventing example for you. Have you ever been on a cruise?" I was like, "I've never been on a cruise." He's like, "Oh, my god. You have no idea what you're missing. I take a cruise every winter."

So Carnival does these cruises, and apparently in the cruise, there's these different rooms for different activities, and he said that when you go from room to room, they put this wristband on you. It triggers our Couchbase database to say, "They're in this room now," so all this stuff pops up to you... Discounts on drinks if you happen to wander to the bar, if there's a ballroom, you can request a song. This is a combo of both field IOT and Eventing, which is great... And it's all in real time. What Carnival also does is... For example, if there's a show on the boat and a lot of people have crowded at the bar, not only does it trigger to the customer, "Hey, you're at the bar. Why don't I give you these special deals?" It triggers to Carnival so they can send more staff from one part of the boat to another. It's really helpful for their business.

Jeremy: Interesting.

Suphatra: I'm sorry, I went on a tangent, but I thought it was interesting with the cruise Eventing. Also, another exciting thing about a cruise, my husband refuses to do a cruise so I can only live vicariously through these technical case studies.

Jeremy: All right. Well, I have never been on a cruise either because 8,000 people stuck on a boat in the middle of the ocean does not sound like a good time to me, but anyways... But hey, listen, to each their own, and to each, their own purpose-built database if they need one. The other thing you mentioned, and this is something I'm really interested in because I think when you think of public cloud providers and you hear Andy Jassy's keynote and you hear some of these other keynotes even at Google Next and something Microsoft Ignite or any of these big tech conferences, and they're talking about where some of this stuff is going, there's a lot of focus on enterprises and I get it.

Enterprises have a lot of money. They spend a lot of money and it makes sense. That's where I think the vast majority of the cloud money is made, is with enterprise and even with government cloud and things like that now. I think one of the things I love about serverless in general, and again, something like DynamoDB or NoSQL database, is the ability for it to start really small and then get really big if it grows, right? So there's a lot of small companies, start-ups or small businesses, that can utilize this type of technology. What role does pricing and infrastructure play for these smaller companies?

Suphatra: Yeah, that's a really good question. There's a couple of things. If you're choosing a NoSQL database and you're concerned about pricing, you have to look at how the databases scale it... Because if it's going to scale up, then you're going to end up paying for expensive hardware than can handle scaling up. If it's scaling out, that'll be a lot better, so essentially if it can shard. If it can shard data, that's going to be more cost efficient. If you can get the NoSQL database, honestly, as software that you can install in your own server, that's probably the cheapest way.

I think my concern for the small guys is if they go to a public cloud provider, that they end up consuming features that they don't realize they're paying for as features... So that example I gave with DocumentDB where let's say you say, "Oh, I want to try some Eventing," DocumentDB, and then you do it, and then you find out later that you actually were consuming additional services. Those are some surprise bills that can happen. I think the benefit, though, of going with fully managed cloud is you don't have to worry about it, right? If you want to just pay that premium and you don't have to worry about it and that's worth the extra cost, then yeah, I think Azure and AWS is a good option. And one way to keep your costs down, is to make sure to keep your data really clean, which means you're just going have really good data hygiene. Does that answer your question?

Jeremy: Yeah, I think so. I mean I think what I'm trying to get at is there are... Obviously we've made a big shift from this idea of on-prem to even just, I guess, general hosting providers. I remember back... Late 90s, early 2000s, where it was always the hosting provider. You pay couple bucks per month and eventually that gets more expensive, but you would just upload some code into a server somewhere, maybe they installed MySQL or something like that for you, and it was sort of this very simple approach to that. As things become more complex though and companies start bringing in their own servers or try to do something on-prem, it gets more complicated. And even if you're a start-up, I mean I can't possibly imagine a start-up nowadays saying, "All right. Let's go buy $200,000 worth of Dell servers and rent a co-location for facilities somewhere and do this on-prem installation."

It seems crazy to me that they would do that and it also seems a little crazy to me for someone to say, "Well, we need a database so let's go ahead and install directly on a VM." They're going to use something like a managed RDS or whatever it's going to be, and I guess what I'm trying to get to is, these services that are built for these massive enterprises, I mean eventually these small companies might want to become enterprises, but is there a way for them to start small with these things? I mean, do you think it's a bad choice to go with DynamoDB or CouchBase if you're a small company? It seems like a good on ramp to me, I would think.

Suphatra: Yeah. I know it sounds crazy to just buy and install your own database and put it in with whatever cloud you're in, but honestly, for a small company, that's the best way to stay portable with your data because essentially you're buying something that you can put on AWS, on Azure, on GCP, whichever cloud you choose, and you can move that around if your costs start to overwhelm you. You can easily change your provider. So that was something that I personally had never considered because when I was at AWS, I only talked to large companies ever. We were only interested in storage from large customers so I didn't get to interface that much with those small companies. And when I was interviewing with I think every single database company in the world, back still a few months ago when I was interviewing, I got a lot more exposure to smaller companies. And what I found was they're really scared of vendor lock-in, what you talked about earlier, and they wanted that portability while they were growing because they weren't quite sure how their data was going to grow, so actually that is a really smart choice for them.

Couchbase, for example, does hybrid multi-cloud, bring your own cloud, put it on a server, then do literally anything you want, and that's something that's common if you were to go to any other small database as well. And I think that when you go to a fully managed cloud provider, you're saying, "This is it. I've made my decision, done deal." Once you're on Dynamo, you're on Dynamo. Like I said, database migration is very complex, very difficult. It'd be a significant investment to change it. If you're under 100 people, I would say that find one that is portable with you so it can grow as you change. You can always migrate to a bigger player later if you find that, "Hey, we're so big. We're going to go do that." I will say the nice thing about DynamoDB though, is that I believe 80% of DynamoDB's customers don't pay for it at all because they're on the free tier, which is a permanent level of free because the storage allowance is really big.

Jeremy: Yeah.

Suphatra: That is a nice thing about Dynamo, so if you say, "Hey, actually I want to go straight to Dynamo," that's an easy way to do it and feel pretty safe there for quite a while while you're growing as a company. And then Mongo has their community edition which is totally free, and then Atlas is paid, so you can take that route as well. That's what I'd recommend... I'm sorry. It's not declarative because I don't want to be like, "No, don't do this. Don't do that," but I guess I'm trying to think if I should be more brisk about it... What I think I find is hard is I feel like I meet people who are... I had an engineer over for dinner the other day because we moved recently, and he has been an engineer for a decade and he said, "Suphatra, should I go to NoSQL because I don't want to learn anything over again." And he's like, "Which is the best one to go to because we're on AWS so I guess I should just use that. That's probably just easiest."

And I think that's often the decision making that happens for people. It's like, "What are we already doing?" I think it's actually pretty rare that an engineer gets to be in a position where they're at the very beginning and they're like, "Okay. We don't have a single database yet at all, so I'm going to choose which one we are going to start on." Most of the time, what I find is people are talking about a database migration because they arrived after something's already been built and now they're choosing whether or not to stay or go somewhere else.

In that case, it's like, "Okay. Now you're factoring in the cost of the migration, the cost of the re-host, replatform, refactor." And then the burden of that choice feels heavy because you're hoping that it's going to be permanent. That's where I say, "Hey, be portable because you're probably going to be making a database migration choice later down the road if you're a small company." If you're a large company, pay the overhead, don't worry about it. I feel that way with my kids. Some things, you just go a little bit bigger on like nice doctors because you're like, "Screw it. Pay an extra 100 bucks. I avoid more [inaudible] in the future." It's totally fine.

Jeremy: I actually think that's really interesting advice, and I'll be honest, I don't know if I 100% agree with it just because of my... And it's always good. Whenever I never disagree with a guest, I always feel like, "Well, we're just talking. We're basically, again, reiterating what each other said." I do get what you're saying and I think depending on what it is that you're building, that portability may be important, especially if you're building something internal or things like that.

But I do think that if you're building something that is consumer facing, that choosing the lowest common denominator right out the gate just because you might need that portability later, might not always be... Especially if it adds complexity, I would say, or it adds additional cost that you might not need initially. For me, that seems like something that you would have to weigh to say, "Do we really think we're going to need to move this or is it something that we can get away with for a while, and then if we happen to hit some stride where we're having some success and we have to think about migration," maybe that's a problem that comes up. I don't know. I think it's probably one of those, "It depends," type answers.

Suphatra: Yeah, and it seems a bit unfair too, because here I am giving the advice from the side of being the provider of the database and not actually being a customer ever. I've never been in a situation where I've had to buy a database, so that seems totally unfair. I feel like it'd be a really interesting thing. I'd love if anyone that's listening to this has any opinions to tweet to me and Jeremy what you think because I would love to know too... Because I definitely see it both ways and I really care about this too because I feel the pain of every person that I talk to who's in the middle of a database migration.

Jeremy: And that's very true.

Suphatra: Because I'm so excited about databases, I want them to be excited too. And I think for people who are still trying to explore the NoSQL world, so maybe that is a net new choice for them, the nice thing is they're so many free trials out there now that they can do that with almost no penalty or no burden. Yeah.

Jeremy: Well, I think the other thing you're going to run into is no matter what database you choose, which technology you choose, whether it's managed or hosted, whatever, you got to discover all kinds of things that you never knew. You're going to discover... That are going to be either welcomed or they're going to be like, "Oh, my goodness. Why did we make this choice?" I think with any technology, there's going to be trade-offs and you're going to have to find the right path for your business.

But I think for me, if I was recommending to any small company that wanted to, say, build a product in the cloud, I would just say, "Choose the things that let you prove it out as fast as possible. Get 80% of the way there, and if you prove out that model, then you go ahead and you can start building custom things that may be more tailored to your solution," but I know that's super general advice and it's one of those things where it's like, "Yeah, in a perfect world that works, but you always run into little things where you end up with other problems." But certainly don't want to be spending months of engineers' time trying to set up servers or anything like that that would potentially slow you down.

Suphatra: Yeah. This is a key business quandary for us here at Couchbase. I mean that's why we created Couchbase Cloud.

Jeremy: Oh, that's right. This just came out, Couchbase Cloud.

Suphatra: Yeah. Yep. So just came out, Couchbase Cloud. It's a fully managed cloud database and it's everything that Couchbase 6.5 provides. You've got the Eventing, full tech search, the multi-dimensional scaling. We've got Couchbase mobile, asset transactions, role base access controls... Oh, my god. I'm going to keep going on and on, built-in ..., service side Eventing. We've got all this stuff, but I think this is part of that business question of just like, "Well, okay. We know people definitely want Couchbase server, which is where they can just go buy it and they can put it on any cloud that they want." They want multi-cloud. They want private and then they want hybrid. We also created the the fully managed. Now you have every possible way to deploy Couchbase and I think what I'm going to be really interested to see is how these numbers shift in participation in those areas. I think that will give an idea as to what are people moving to. I'm sure you see it too, but I see in the news all the time, people saying, "Okay. Well, now it's going to be a move over to hybrid instead of going full public cloud." But I'd love to know what you think about that, if you think there's going to be a shift over to hybrid.

Jeremy: Yeah, I mean, listen, every application that I've built probably in the last two years has incorporated some sort of hybrid database technology. A lot of the operational things live on NoSQL or on DynamoDB as that operational side of things that you don't have to worry about throughput, and that becomes that source of truth, but things are either replicated into something like Kinesis Data Firehose, they get dumped into S3, they can be queried by Athena... Or some of it goes into an Aurora database or something like that where you have the ability to then run aggregations and some of these other things that don't need to run at the same scale as the operational side of things. I think this is something that people have been doing for quite some time now. Maybe I think people that are developing in the cloud are ahead of the curve in a sense because they're understanding that the scale becomes a real problem when you start to get a bit of usage.

And just from a cost standpoint too, I mean the on-demand side of things is incredibly flexible from a pricing standpoint in the early days. I mean I think if you really hit scale with some of these things, you're going to start to realize some of it's expensive, but you're also not paying a team of engineers to manage servers for you either. So I do think that that is the trend, but I'm really interested to see whether or not you have massive enterprises that have spent billions of dollars building these complex custom Oracle systems and have paid these consultants to come in for a year and a half and sometimes walk away with maybe not a lot to show for it, if you're going to see NoSQL technologies start really getting into the enterprise at an internal level and using NoSQL to do more complex things that maybe aren't customer facing, but is starting to replace some of these large... I don't if we would call them legacy, but some of these older ways of storing data.

Suphatra: Yeah, that's interesting. Like I said, Jeremy, you're a very smart guy. No, I think you're totally right. I think you're totally right, totally spot on on that. Yeah, I think that's something I'd love to, in a year, come back and see if our predictions panned out. But yeah, there's some sort of shift happening, it's surprising how many people are doing switchbacks. To me, it's mind-boggling too, because it just seems like an expensive endeavor to take on, but you have to be [inaudible] pretty high level frustration to say, "Oh, hey, I'm going to go back to where I was before." And there must be a pressing business reason, which I think there's not enough investigative journalism and research into that. Instead you see other stuff like, "Oh, who's going to get the Pentagon contract?" or "How much I paid on Azure, AWS, was insane," those kinds of things. But what I'm really curious about is why do people make that change back? Because the business reason there has got to be affecting more people than we realize, so if I find out, I will definitely let you know, Jeremy.

Jeremy: Yeah, that's actually something that I'm curious about. I think you see these very big contracts for AWS that get signed and obviously you hear complaints sometimes that the costs get ridiculous. You know what I mean? That the cost of running maybe DynamoDB at scale or some of these other things... Now again, some of that has to do with optimizations and I think people don't necessarily use some of these things the way they were intended to and maybe that's a pricing issue, but what are you thoughts on... Does this get to a point where maybe cost is the reason why people start migrating back?

Suphatra: I hate to say it, but yeah... I mean I don't think it's the technology, right? Like I said with Oracle, it's not the technology, I don't think, at Oracle.

Jeremy: Right.

Suphatra: I don't hear people complaining about the tech, it's more the business practices and that's why I liken something like AWS to Target, which everyone loves Target. There's nothing wrong with Target. It's just that when you're in that environment, it encourages you to spend a little bit more than you should. And then I think with Azure, it's that opposite problem where you feel like you're paying more than you're spending. I don't know if there's going to be... There's no perfect answer, but I will say that I think the deals are really important here. With enterprises, they get their own dedicated sales teams, they're negotiating very specific deals... And then these guys are very aggressive. They're going head to head often so they're trying to undercut each other in those deals.

Jeremy: Right.

Suphatra: I'm sure that they can make compelling offers to do a switch or a switchback. Service is also really important, so if you've just signed a huge migration to one provider and then you're not getting the quality of service that you need to really maintain or improve on the infrastructure that you've built with them, then you might feel like, "You know what? We don't want to make a long term investment with you if you're not making a long term investment in us." That's another reason I think I've seen people move, is they say, "I don't think your service teams care about us." Another is competition. I think that is why some of the database companies have enjoyed a lot of growth is because a lot of... In the retail space, because a lot of retailers just will not work on AWS infrastructure, right?

Jeremy: Yeah, right.

Suphatra: Because they say, "Well, Amazon is just edging us out."

Jeremy: Well, they say they don't say that, but I think we know that they do.

Suphatra: Yeah, right? It makes sense. Amazon is completely edging out their business. They're not going to go and they know that AWS is paying for and supporting the rest of that business, so they're not going to go on that.

Jeremy: Right. Let me move on to one more thing. I think a lot of people are interested in what it's like to work for Microsoft and Google and AWS and some of these other big companies. Obviously, you have experience working with Microsoft. You were there for 10 years, I think, and you were with AWS for a bit of time. I don't know, maybe you could share just what was it like? What was it like working with those two different organizations?

Suphatra: Yeah, I feel so lucky and blessed to get that opportunity and just really great experiences at both. I have so many friends from both places and I would happily work at both of those places again... Really great experiences. And I think there's a lot of misconceptions about both of those places, especially Amazon, and it's unfortunate because I think that Amazon is a really special place to work. It's difficult in the sense where the people are incredibly smart. Anyone that you're working with at Amazon is just top-notch, really smart. Hiring process is very difficult, it's a very difficult filter to get through so you're really getting some of the best, smartest people in your industry in the world. There wasn't a day at Amazon where I didn't feel intellectually challenged and that's just really fun. And same with Microsoft, Microsoft has some of the smartest people in the world as well and Microsoft is more global in a sense, and so you're really working with the smartest people literally all over the world. But Amazon, they're smart in a different way. There's a saying at Microsoft that the loudest person in the room wins because Microsoft has a bit of a bullish culture where presentations are made with PowerPoints and people have big personalities. Bill Gates built his little empire there in a way where it was just traditional 80s, 90s business culture back then.

But Amazon's very different, it's a written culture, so to succeed at Amazon, you need to be a really good writer. The best writer in the room wins. Amazon has a very strict document culture, so you write these six-page documents which are one-inch margins, single space, no pictures, no graphs, with an appendix at the end, and it has to be just a business case of why you want to do the thing you want to do... And that's for anything. That's for a website change, that's for a new program, that's to add something. I mean you really spend a lot of your time writing these business proposals, business memos, and what that does, it causes you to really think very thoroughly how what you're proposing to do is going to benefit the customer. I would say what the retail side, Amazon.com, culture brought over to AWS is really beneficial, which is this customer first mentality... With a little bit of differences, AWS is obviously a little more competitor focused. You see a lot of focus on Oracle and Microsoft, more so than what you see the retail side does.

The downsides of working at Amazon is because it is a written culture, it's actually a very introverted culture, very quiet. It's funny because Microsoft's very gregarious, full of people, everyone's visiting and flying in and you go to fun parties. I remember my third week into Amazon and I was on LinkedIn and my old coworkers at Microsoft were literally posting selfies with Will Smith.

And I'm sitting in Amazon where it's just quiet. All you can hear are dogs because everyone brings their dog in and the dogs are making more noise than the people. We're all just working and typing up our memos, and here my buddies at Microsoft are just posing with Will Smith. So it's really a truly difference in culture and it's a very introverted culture. If anyone stayed for more than four years are likely a very serious written, introverted person... And that can be a little nerve-wracking for some folks. People do not like that. I've heard people say, "Oh, well, I've heard that Amazon makes you cry at your desk."

Jeremy: What?

Suphatra: You know that New York Times article that said people were crying at their desk all the time at Amazon and living in their cars?

Jeremy: Oh, yeah, yeah, yeah.

Suphatra: I never met someone living their car at Amazon or even sleeping in the office, I will say. And I thought that there was very good balance, people were able to get home to their families. The people are really passionate about the work, so I think it's a really, really great place. It's not for everyone. It is very difficult, so I would say that, but work can be difficult. And I will say it's funny because Amazon's known for being an entrepreneurial culture, but I would say that's the one thing I didn't find it to be. I did not find it to be very entrepreneurial actually. I found the very top management sort of allowed that kind of entrepreneurialism, but for the most part, if you're below those ranks, the work is very tight. Your swim lane is very narrow, so it is run very much like the retail side, which is almost like a supply chain where you have your role. You do your role. This is what you do. You have to excel at that and do it at max speed. It's very, very efficient.

Microsoft's different. You could come up with some sort of crazy idea. They'll give you a million dollars, you can go and do it. And they're like, "Hey, we should bring Will Smith in for the staff meeting." They'll do it. So it's really, really different. It's just different at Microsoft. Microsoft's just been around longer and it's all over the world and they've got big parts of their business that just print money, Windows, Office Suite-

Jeremy: Yeah, of course.

Suphatra: Prints money without really having to do too much for that, and most people don't know, but Microsoft's number one customer... I would love for you to guess, what do you think Microsoft's number one customer customer is?

Jeremy: Is it the U.S. government?

Suphatra: It is governments.

Jeremy: Governments in general.

Suphatra: Governments... It's not a consumer. It's not consumers and it's not businesses. It's governments and when you get a government contract, you are in that country for a decade or more.

Jeremy: Yeah.

Suphatra: I remember at Microsoft working on something that was for North Dakota. We are in every single kindergarten through graduate school... Is running Windows and Office for the next 12 years.

Jeremy: Wow.

Suphatra: I mean when you get contracts like that, like I said, Microsoft likes... You pay the premium, you get everything you want, but you pay for the full buffet. But the nice thing when you work there, it's a very different, lush culture and I think people at Amazon would say, "Oh, well, Microsoft's a country club." It's kind of true, a little true... But then people at Microsoft would say, "Well, everyone at Amazon's crying at their desk." Not totally true, but it's a very difficult work culture, so there's a lot of trade-off.

Jeremy: Well, I've always thought it was very generous of companies like Apple and Microsoft to donate computers and technology to schools, but then when you think about it, the more cynical side of you say, "Well, they're just educating consumers," right? They're just grooming consumers for when they graduate and go to work which devices they're going to choose. I know a lot of people who work at Amazon and I've heard similar things. I mean I think that people enjoy the work that they do like you said and they're very passionate about it, but I think the bottom line is if you like dogs, work at Amazon. If you like Will Smith, work at Microsoft. All right. We've been talking for a while so why don't we wrap this up? But listen, Suphatra, thank you so much for being here. If people want to find out more about you, how do they do that?

Suphatra: You can tweet me on Twitter. My Twitter is SKPRUFO. I'm really active over there. I would love to hear from you. I would love to hear what you thought about what I shared today. I love to hear your opinions and your thoughts because I'm always trying to get more information from customers and users and developers to help better inform my ideas and my opinions, so I'd really love to hear from you. Come and tweet, "Hi," to me.

Jeremy: And if you want to check out more about Couchbase and the new Couchbase Cloud, you just Couchbase.com, right?

Suphatra: Yes, Couchbase Cloud, we just launched it recently, so please go check it out. It's fully managed cloud database for your server folks. It's got all the service side Eventing you're going to ever want. And I really appreciate, Jeremy, the chance to talk with you. You're such a cool guy and I love this podcast.

Jeremy: Oh, thank you.

Suphatra: Everybody loves this podcast, so thank you so much.

Jeremy: Well, I appreciate that and the next time we bump into one another, we'll hit up the seafood buffet and we will continue this conversation.

Suphatra: Thanks. Thanks, Jeremy.

View Details

This is PART 2 of my conversation with Rick Houlihan. View PART 1.

About Rick Houlihan:

Rick has 30+ years of software and IT expertise and holds nine patents in Cloud Virtualization, Complex Event Processing, Root Cause Analysis, Microprocessor Architecture, and NoSQL Database technology. He currently runs the NoSQL Blackbelt team at AWS and for the last 5 years have been responsible for consulting with and on boarding the largest and most strategic customers our business supports. His role spans technology sectors and as part of his engagements he routinely provide guidance on industry best practices, distributed systems implementation, cloud migration, and more. He led the architecture and design effort at Amazon for migrating thousands of relational workloads from Oracle to NoSQL and built the center of excellence team responsible for defining the best practices and design patterns used today by thousands of Amazon internal service teams and AWS customers. He currently work on the DynamoDB service team as a Principal Technologist focused on building the market for NoSQL services through design consultations, content creation, evangelism, and training.

  • Twitter: @houlihan_rick
  • LinkedIN: https://www.linkedin.com/in/rickhoulihan/
  • Best Practices for DynamoDB: https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/best-practices.html
  • 2017 re:Invent Talk: https://www.youtube.com/watch?v=jzeKPKpucS0
  • 2018 re:Invent Talk: https://www.youtube.com/watch?v=HaEPXoXVf2k
  • 2019 re:Invent Talk: https://www.youtube.com/watch?v=6yqfmXiZTlM

Transcript:

Jeremy: So one of the things that you have never mentioned or at least I don't think I've ever seen you mention it, at least not in any of your talks for your modeling is local secondary indexes.

And I used to think, "Hey, this is great. They've got really strong guarantees and then it's sort of this great use case if you want to do a couple of different sorts." But LSIs are not quite ...

Rick: Not the panacea you might think they are.

Jeremy: Yes, correct.

Rick: So LSIs, I'm not exactly sure. I mean, I think you're exactly correct. The biggest value of LSI is the strong consistency, right? But the limiting factor of the LSI is it doesn't really let you kind of regroup the data, right?

Jeremy: Right.

Rick: You have to, you have to use the same partition keys to the table. So the only thing you can really do is resort the data, right? So right there, that's a limited set of use cases, right? There's not a lot of access patterns. I mean there are, but there's not necessarily a ton of access patterns or applications that only required me to resort the data. Most applications are going to require to group the data on multiple dimensions so that limits the effectiveness of the LSI. The other thing about the LSI that kind of stinks is they have to be created at the time the table is created, they can never be deleted.

So if you mess it up, then you've got to recreate the table to get rid of them and I find them to be extremely limited use. I mean, most developers can tell you that strong consistency is an absolute requirement, but when you get down to it and started looking at the nature of their application, yeah, what they really need is read after write consistency, right? It's worth kind of talking about the difference, right? Strong consistency implies that no update to the database is going to be acknowledged to the client unless all copies or all indexed or copies of that data are also updated, right?

Jeremy: Yeah.

Rick: That's strong consistency. That means if I'm in a highly concurrent environment, that no two clients could read different data, okay? Unless the read is not, or the write is not yet fully committed. As long as the right hasn't committed, you're not going to get two copies of the data. Well, most use cases are really more about like if I make the write and I read back, did I get the right data?

So what we're really talking about is read after write consistency. If you think about the round trip between the client and the system, if I have a let's say in DynamoDB, GSI replication is 10 milliseconds or less, it's highly unlikely that you're ever going to be able to return to the client, that the client is ever going to be able to returned to the server and ask for the same data back in 10 milliseconds.

Jeremy: And honestly, if you do, welcome to distributed systems.

Rick: That's exactly right. I mean, that's the other thing I was going to say and most distributed systems, what you'll find is there's a propagation delay on configuration data. So oftentimes, even if you get to the point where the developers are going to tell you that there's going to be concurrent access on this data, when you back up a step, you're going to find configuration data is going to live in multiple entities. So hey, all bets are off, right?

So let's take a look at that need for strong consistency and not make arbitrary requirements because as developers when we make arbitrary requirements, it's like hooking a fire hose up to our wallets. Let's make sure that we're actually making requirements that are meaningful to our business. 90% of the application workloads I work with, I would say even maybe even higher don't require strong consistency. So let's just use those GSI. They're much more flexible, right? They can be completed anytime, they carry their own capacity allocations. They don't pillage capacity from the table. Overall they're just a lot more flexible.

Jeremy: Yeah, and you've got more control. I mean, that's one of those things too. If you are doing the single table design and you're using all those different entity types and so forth, what are the chances that all those LSIs and the sorts all align with one another too. It seems like a lot of wasted capacity.

Rick: Inevitably, you're going to end up using GSIs, right?

Jeremy: Right. Exactly.

Rick: You may be able to use an LSI for one use case, but you can't use them for all of them.

Jeremy: Yeah, and I mean, and I think just the important thing about LSIs too is regardless of the inflexibility of them, there's also a doubles the costs, right?

Rick: Well, all indexes double the cost, right? I mean [crosstalk 00:49:35]

Jeremy: Of course, yeah.

Rick: Because actually, one of the things people kind of ... It's kind of an incorrect assumption about LSIs is that customers believe that, "Oh, they use the same capacity as the table. Oh, they must be free." No, they're not free. You still pay for the storage, you still pay for the capacity. I'm just going to have to allocate twice as much capacity to the table now.

Jeremy: Moving on from LSIs and GSIs, the other thing that always comes up is this idea of hot keys or hot partitions where you basically have one key that gets access quite a bit. You sort of pointed this out in your slides where you see sort of as big red marks and sort of heat, this heat map where you get one partition that is red or is being accessed quite a bit.

So we can talk a little bit about the performance of those things, but I'm actually curious what happens if a partition exceeds that 10 gigabyte partition limit?

Rick: Oh sure. Okay. So yeah, so as you pointed out, there's a partition size limit in DynamoDB, there's capacity and throughput limits in DynamoDB and the reason we chose to do this is because we wanted a system that was responsive and scales in minutes, right? If the larger systems like you look at a MongoDB or DocumentDB that has used very large storage nodes, they have large capacity storage nodes, it takes them a long time to be able to add new capacity and what we wanted was a system where a user can come in and say go from 10 WCUs to a million WCUs and do that in realtime, right? Not months, literally months.

So what happens when a partition exceeds its 10 gigabytes is the system behind the scene is going to say, "Okay, I need to move this data into multiple storage partitions." So the way that NoSQL databases scale is they're going to add partitions. When they add partitions, they need to copy data to those new partitions in order to be able to bring them online so to speak. If I use extremely large storage nodes, then it takes me a long time to copy that data, okay?

Jeremy: Right.

Rick: So in large NoSQL clusters, I mean, the largest MongoDB cluster is about 64 shards right now. They're adding shard 65, they started in November, they expect to be done sometime in the next couple of weeks and that's no joke. That is no joke.

Jeremy: That's really scary.

Rick: Yeah, it is scary. It's really scary for your business. I mean, what happens if they see a surge in traffic in the meantime, right?

Jeremy: Yup.

Rick: They're DOA. So and they're actually talking to us to migrate because they know that when they go to add shard 66, it's going to take them nine months, right? So it's not something that's going to work for their business. So anyway, so we want to be able to scale in minutes and we can do that because we use small storage nodes. When a storage nodes hits 10 gigabytes, is going to split into two nodes.

Now, I can copy five gigabytes of data in literally seconds, right? And I can do that in parallel and number of times. So that's how DynamoDB table scale gracefully is they have these large number of smaller stories nodes when you want to add capacity, we just split those storage nodes very quickly in parallel and we can bring that capacity online in minutes and that's the advantage there. So that's kind of what happens.

Jeremy: Yeah, and so that only works though if you are not plagued by a local secondary index though.

Rick: Yes. So the local secondary index, again, another one of those limiting factors, the local secondary index. Since we only allow you to resort the data, not regroup the data, that's what gives us the ability to support strong consistency. But the only way we can do that is to ensure that all the data between the local secondary index and the table actually live on the same partition. So if you have a partition in it's single logical partition in DynamoDB on a local secondary index that exceeds 10 gigabytes, it's going to throttle the table and stop the rights because if you think about, if I resort the data inside of a logical partition and it's larger than 10 gigabytes, then that's automatically going to mean that some of the data lives in two ... The data lives in two places and maintaining consistency on two physical hosts is hard.

So we punted on that idea and said, "We'll give you a consistent indexing, but don't ask for more than 10 gigabytes in a single logical partition." That doesn't mean that an LSI can't be larger than 10 gigabytes, it just means that a single logical partition value can not contain more than 10 gigabytes of data.

Jeremy: Right. And the other thing is in order for it to split data across multiple nodes, it has to have a sort key.

Rick: That's correct. Yeah, yeah, yeah. If you're going to split a single logical partition, I mean, if you don't have a sort key than that, you're limited to 400 kilobytes in a single partition because that's the item size limit in DynamoDB.

Jeremy: Got you. Okay. All right. So then in terms of throughput performance, if you actually are on multiple nodes, wouldn't you have better throughput?

Rick: If you are on multiple nodes, yes, you have better throughput and that is another advantage of DynamoDB with lots of small storage nodes, right? We can increase the throughput of the system more easily. Now, we do need to maintain some proportion of throughput to capacity with storage allocation, right? So if I have a storage device, a storage node that has X terabytes of data, and I'm carving that up into 10 gigabyte chunks, I kind of also need to carve out the IOPS as well because otherwise, there's no way for me to guarantee that that capacity will be there for you when you come ask for it, right?

So that's kind of what you're doing. When you reserve capacity in DynamoDB, it's guaranteed you're going to get it and it's up to us to make sure there's enough capacity on the system that to satisfy your request, but whatever you allocate, there's ... Nobody is going to be able to take that from you and nobody's going to brown out your workload because they're too busy and you're sharing a storage node.

Jeremy: Yeah, so now if you were to create partitions with hundreds of gigabytes of data, it's going to spread and split itself across multiple nodes. There's sort of a throughput benefit, I think they are performance game because you're sort of doubling or increasing that throughput, but is that something we should avoid? Should we try not to create partitions more than 10 gigs?

Rick: Well, I mean, it's going to be hard to do that, right? I mean, most applications it's going to be ... You've got to have the data to aggregate. I mean, if I'm partitioning data and I'm saying I want orders by a customer, the 10 gigabytes of orders could be a lot of orders.

Jeremy: Right.

Rick: That's really what it comes down to. I wouldn't say avoid it. The one thing to be aware of when you're working with the data moving in and out of these individual logical partitions is you do want to be aware of velocity. How fast am I moving the data in and out?

Now, having 10 gigabytes of data in a single logical partition is really no big deal, but if I have to read it really quickly, that's going to be a problem because you're only going to get 3,000 RCUs, that's a megabyte a second. So you can only really at one megabyte a second. If I have a gigabyte of data inside of that logical partition, that's going to take me a hundred or a thousand seconds to read that gigabyte of data. So if I had 10 gigabytes, yikes, I'm going to be reading for a while. Right?

Jeremy: Right.

Rick: This is where we start to talk about right sharding and read sharding.

Jeremy: Yeah, sharding. So yeah, so let's talk about charting for a second because that is something that I think some people see that as, "Oh, if I want to be able to read a bunch of data back quickly or whatever, I have to have to split it up, I have to use some sort of hashing algorithm maybe, I have to figure out how much I want to sort of spread out that key space." But actually, there are quite a few benefits to doing that, right? Because you can read it in parallel and things like that, right?

Rick: Right, absolutely. I mean you want to increase the throughput of any NoSQL database, you'd talk about parallel access, right? So in DynamoDB, what we're going to try and do is if your access pattern exceeds, 1,000 WCUs or 3,000 RCUs for a single logical key and now bear in mind that I had ... It sounds like that's not a lot, but I have architected, I don't even know how many thousands of applications at this point on DynamoDB and right, sharding comes into play like, I don't know, less than 1% of the time.

Jeremy: Okay.

Rick: Most workloads are just fine with those capacity limits, right? And if there was a problem with that, then we would be working to adjust them. We just don't see that as being a problem. We see that as being more of a concern that developers might have when they start learning about the system, but when we actually start going into the implementation cycle, what we find is nobody writes shards.

Now, that's not to say nobody, some people absolutely need to and when ... But that is just a nature of the beast when you're dealing with NoSQL, right? Because we're dealing with a partition data store. If I want to increase throughput, I need to increase the number of storage nodes that are participating. This is true for every NoSQL database. It's just that the individual throughput of the storage nodes and the legacy NoSQL technology is higher because they're using entire physical servers is the storage node whereas DynamoDB takes it as physical server and chops it up into a thousand storage nodes.

So, and again, the reason we do that is we want to scale gracefully and we found that the write throughput and read throughput settings that we've adopted tend to accommodate the vast majority of workloads. So again, if your throughput requirements are higher on a per logical key basis, let's talk because it's not that hard to do, right? It's just a mechanical chore. Once you've kind of implemented that mechanism underneath the data layer API, most of your developers don't even know what's happening.

Jeremy: And I think if you're an application developer and you run into a problem like that, it's a good problem to have because obviously, you're doing pretty ... Your application is being used.

Rick: [crosstalk]

Jeremy: Exactly. All right, so let's move on to denormalization, right? This is another thing I think that the trips a lot of people up. We talk about third normal form and stuff like that, that when you're optimizing a regular SQL database, we want to split everything up into separate tables and so forth. But in in DynamoDB and NoSQL databases, we often have to denormalize the data, we have to put logical data together. Sometimes we have to copy things to multiple records, sometimes we have to copy things into the same attribute and things like that. So what are of the advantages though to de normalization?

Rick: Time complexity on your queries. I mean, that's really what it comes down to and we'd started talking about NoSQL, we're talking about cost efficiency, right? We're talking about the low latency consistent performance at scale and the way we get that is through denormalizing the data, right? Because now, everything instead of select star from inner join, inner join and inner join, it's select star from where X equals, right? Now a single table filtered select from a relation database is blindingly fast, right?

Jeremy: Mm-hmm.

Rick: And that's what we're really doing. NoSQL reduces every query to a single table filtered select and that is why it's going to be faster and it requires us to denormalize in order to achieve that effect.

Jeremy: Yeah. And one of the things that I really like about denormalization is this is something I designed SQL databases for a very, very long time, built a lot of applications, a lot of eCommerce products on there and one of the things that always drove me nuts was you have a history of orders, maybe two years of a customer's order or customer orders and then they update their email address and then suddenly your join, now you have the email address that they currently have, not the email address they had when they placed the order because of the way that that works. And so unless you're denormalizing data which is eventually what I ended up doing anyways was to keep a record ...

Rick: Right. Yeah, you got to have the history of ... Some of the data is going to be immutable, even if the user changes it, right? You're going to want to know that it was ordered by so-and-so when they had this name before they were married.

Jeremy: And what their address was at the time.

Rick: What their address was at the time and what was their phone number when it happened, right? Yeah, absolutely and so when you normalize data, when you would eliminate that data from those records, then you're eliminating the ability of the system to keep track of it and as you pointed out, the only way to do that is to de normalize it, right? And so at that point, and this is again, actually it's a good point you bring up because it's one of the things we found at Amazon retail that we were denormalizing our data inside of our relational databases to deal with the scale of the system that we were trying to support, right? We couldn't calculate these common KPIs using queries anymore. Things like the counts on the downloads of the tracks for Amazon Music, right?

Jeremy: Mm-hmm.

Rick: I mean, for a while they were just select count from download table. Okay.

Jeremy: Oh geez.

Rick: Yeah, oh geez, is right. So after a while they're like, "Oh, well, let's create a roll up table and we're going to have a top level counter for downloads for the song and we'll just update that every now and then." Right?

I'm like, "Yeah, that makes a lot more sense, but what have you done? You've de-normalized the data, right?" And so at this point then, why am I not using a first-class NoSQL database? I'm trying to turn my relational database into a denormalized database, and then the next step you'll see people do, I see it all the time.

We have thousands of RDS customers that are doing sharded Postgre, sharded MySQL. I mean, hey, there's things out there, there's technologies people have built, pgRouter and stuff like this to be able to support and I'm telling you, as soon as you shard your relational database, man, you've gone down the road, let's go into NoSQL, right? I can't join in class instances anymore, right? So now, let's go back into a database that's built for that.

Jeremy: Yeah. No, and actually, one of the former startups I was at, we built an entire, MySQL cluster that had like a master master sort of directory service that would tell you which shard a particular customer was in. And then you were replicating relationships and things like that and you're like, "Wait a minute, why don't I just store this in one place that I can actually read this data from?"

Rick: That's exactly right.

Jeremy: But yeah, no, so I totally feel the pain there. So the audit trail piece of this I think is something that's really, really interesting and another thing you had mentioned in your 2019 re:Invent talk was this idea of sort of partial normalization where you might have, and the example you gave was this big insurance quote and you said that, "You don't want to store copies of the same quotes, especially if it's big. You don't want to store the same thing over and over and over again, but the immutable data, sure. Things that aren't going to change, their address probably isn't going to change, things like that."

But if you're updating maybe what the value of the quote is or something like that, you actually talked about breaking those up into smaller attributes.

Rick: Yeah, yeah. So in that particular example, it was an interesting use case. We had a customer very happy with the system. They were an insurance service, they had about 800 quotes per minute I think was their update rate. They were provisioned about thousand WCUs and their use case was pretty simple. Users came in, they create an insurance quote. They might edit that quote two or three times. Then they'll go ahead and execute the contract or drop the transaction and every ... The way they were kind of store the data is they would have the customer ID was the partition key.

The quote ID and version was the sort key and if they came in to get a particular quote, they'd say, "Okay, select star from customer ID." Where quote ID starts with or where sort key starts with quote ID and they would get the quote and all the versions of that quote and then the customer could go back and page through them. The thing was each one of those items was 50 kilobytes, 99% of the data in those items never changed. So every time they created a version of the quote, they're storing 50 kilobytes of data that what really existed in the last version, right?

So that's basically what I basically recommended to him was, "Hey, create the first version of the quote and then store deltas. Every time someone changes something, just store what changed." Now when you go use the same query where customer ID equals X starts with quote ID, but what you're getting is the top level quote and all the deltas and then the client side, you can just quickly apply the deltas and show them the current version and then whenever they need to see the previous versions, you just back the deltas off as they back through the various versions of the quote. So this caused a significant decrease in their WCU provisioning after they went from a thousand WCUs provision to 50.

Jeremy: That's amazing.

Rick: That's a 95% reduction. So that was a really good example of how understanding that ... Don't store data you don't need to store. Denormalization does not always mean copying data, right?

Jeremy: Yup.

Rick: And you've got, and this is a really good example of how you can look at what you're doing with your data, how is the data moving through the system, right? Because this is really oftentimes what we find is we're reading data we don't need to read, we're writing data we don't need to write.

One of the biggest problems we see in NoSQL and it's facilitated by the databases that support these really, really large objects, right? Things like I think MongoDB supports a 16 megabyte document, right? And the reality is that I don't know, very many access patterns, and again, I've worked with thousands of applications at this point that need to get 16 megabytes of data and it gives you in single request.

So oftentimes you'll see in MongoDB these really giant data BLOBs and users are going to say, "Get me the age of this user." [crosstalk] into this big giant data BLOB and they'll pull out a four byte in, right?

Jeremy: Yeah, and you have to read the whole thing. You have to read the whole thing in order to get it. Yeah.

Rick: ...to get this four byte in, right? So one of the things I do quite frequently is I work with a lot of customers on legacy NoSQL technologies like MongoDB or Cassandra or Couchbase and I'll actually get them to the correct modeling state of their application and so they'll come to me and they'll talk to me about migrating to DynamoDB and when I'm done, they ended up staying on MongoDB and I'll talk to them and get in a year when they actually do have to scale, but you know what I mean?

It's like it's nice because the design pattern's best practices and data modeling that we've built and that we've developed over the years working with the CDO and doing that large migration, it turns out that all of that stuff is directly translatable to every NoSQL technology and what it really exposed was how wrong the implementation philosophy is and how much of the industry is revolving around some really incorrect assumptions and things that they say. And it was eye opener to go through that. I was one of those people.

Jeremy: Well, I mean, I think that's what's really great is the ... To see the thinking evolve though and sort of get to that point. So I want to move on to something else so just quickly ... That's the quote thing. So one of the things you had mentioned too is this idea of pushing a lot of that complexity in terms of maybe reassembling the quote like pushing that down to the client. So is that something you ...

Rick: Yeah, absolutely. You know what? Those clients are 99.9%, "I had a loop man. Make them do some work." Right? I mean, we [crosstalk] a lot though. I've got some reservation use cases where I was talking to some customers and maybe they were creating items in the database in the table, an item for each one of the availabilities in the calendar or something like that and then they would come in and they would update that item with who booked it.

And I was like, "Don't do that." Just store the items that people booked and then on the client side, when they just say, "Here's the day that I want to book an appointment for, send them down the things in the book then let them figure out what slots are available."

Jeremy: Which one's are booked.

Rick: Otherwise, I've got to do a more complex query to kind of figure out which items are available and which items had been booked and I got more work to do with the application server.

I am a big fan of pushing whatever logic I can down to the end point, right? Give them a chunk of data and let them triage this, give them enough data to do the two or three things that I know they're about to do as soon as they make that request, right? I mean, it's a way better experience for the end user to have a responsive application, right?

Jeremy: Yeah.

Rick: Preload some of that data so that they know, I know 99% of the users that come in here when they ask for this, the next thing they hit is that, okay, great or the next thing they hear is one of these three things. Great, guess what? They're going to get all three of those things and it saves round trips to the server and what are we talking about? Most of the time we're talking about pushing down a couple of kilobytes of data, right? And [crosstalk 01:09:14]

Jeremy: Right, it's not a ton of data.

Rick: Yeah, it's not a lot. So let's get it down there. Yup.

Jeremy: And now with 5G on your mobile devices ...

Rick: I know and you've got one of the unlimited data plans and all that stuff.

Jeremy: Exactly. All right, so you mentioned MongoDB and you mentioned Cassandra and obviously, one of the new things that was announced was managed Cassandra. So I know the characteristics are very much very similar to DynamoDB, but other than somebody sort of already using Cassandra, why would you use the Managed Cassandra? What would be the reason for ... Would you start with that or would you just suggest people start with DynamoDB?

Rick: I think what ... The Managed Cassandra service is awesome. Okay, it's actually DynamoDB DNA. So when you're using the managed standard service, you're using the backing, a lot of the backing infrastructure from DynamoDB, but it's not DynamoDB. It's actually 100% full version of the opensource Cassandra. What we basically did was replace the DynamoDB request router with the Cassandra instance. It's fully managed in the back-end.

So it's actually a really neat piece of technology. I love how they did the implementation. However, it is more expensive for us to run those Cassandra front-ends than it is the DynamoDB head node. So as a result, the Cassandra cluster MCS is going to be ... Not significantly, but it will be noticeably more expensive than DynamoDB. So if I was looking at a brand new workload today, I'd go DynamoDB first. That's still our approach, it has always has been our approach.

I mean, we released DocumentDB for a subset of customers that have the need to have a fully managed MongoDB solution. They don't necessarily want to pay two vendors, right? If you go to Atlas, you're kind of paying MongoDB and paying us through MongoDB. They wanted an AWS native managed solution. So we did that for them. However, we are still running a DynamoDB first philosophy and for all the reasons that we've talked about in the past, right?

Cost efficiency, scale of the service, the robust nature of the system is unparalleled. It's unmatched. We get a lot of that with MCS. You get almost all of it. As a matter of fact, you do get all of it, but you're paying a premium for that managed Cassandra head mode.

Jeremy: Well, I'm just wondering too because I mean we hear all this stuff about people who want to be multi-cloud and they want to be some sort of vendor agnostic or something like that which again, if you were to choose Cassandra, you're still locking yourself into a vendor, but I wonder, I just wonder if it's something that would help maybe customers that don't consider DynamoDB first that this might be ...

Rick: Oh sure. I mean, if you are hung up on a cloud agnostic or a vendor agnostic. I guess like you said, vendor agnostic, what does that mean when I choose Cassandra or MongoDB, but whatever. I mean, I think what they're really worried about is cloud agnostic, right? They want to and I've seen ... This is a fun factor argument for legacy technology providers, right?

I mean, when you go to the cloud, you've got two choices, right? I can lift and shift my existing data center and deploy exactly as is and I'll never know the difference and I've done it. I've taken very complex enterprise IT infrastructures and recreated them 100% to the point where the IT admins have no idea that they're not working on their whatever, on prem facility, right?

It looks exactly the same, okay? Now, that's not a really great way to use the cloud, right? I mean, you're not going to maximize your benefit, right? Then you're probably going to see a slight cost benefit, maybe even not a cost benefit, right? Because you're, you're really not taking advantage of any of those cloud native services.

They're giving the elasticity and the consumption based pricing and all the things that you need in the cloud. So and with databases, if you think about this, what ... It's not the database that locks you in, it's the data, right? When I deploy 10 terabytes of data someplace, I'm locked in. It doesn't matter if I'm on MongoDB or Cassandra or whatever. If I want to go somewhere else, I've got to move the data. That takes a long time, right? And then the third factor to look at here is that once you boil down your NoSQL to the lowest common denominator which is the data model, none of those other fancy features matter.

As a matter of fact, those are the things you'd never want to use. Things like aggregation framework or nickel or any of these other query languages. Now CQL I'm not going to say that because CQL did the right thing. They actually said, "We're not going to try and implement join operators here. We're not going to give people the ability to modify the data. We'll just give them a nice familiar syntax to select their items."

I like that. Okay? I really like that, but if you look at like nickel or you look at MongoDB's aggregation framework, the bane of my existence at MongoDB was going around and dealing with customers and all those terrible aggregation queries and unwinding all that stuff and changing their data models to be more efficient and you really don't want to use that stuff.

Anyways, when you boil it down to that lowest common denominator, it doesn't really matter if I'm using Cassandra, MongoDB, DynamoDB or Cosmos DB, who cares? It's all ... The data model is the data model and it's all select star where X equals and they all do that just as well as each other.

Jeremy: Right, yeah. And I actually, so I mean my recommendation would always be just because I've been working with it for a while. I really do love DynamoDB, but I worked with Cassandra and I had to sort of peripherally manage the Cassandra ring and I can say using Cassandra was great, NoSQL was good, but managing it was not fun.

So at the very least, having a managed services is a nice alternative to somebody who is really hell bent on using Cassandra. So anyways, all right. So I actually got a couple of questions from a few people that I'd love to ask you and just sort of just give me a brief ... Just your, your two cents on some of these things.

Now, of them was about analytics, right? We talked earlier about the difference between relational databases and NoSQL. Obviously, you're not running a lot of analytics workloads on NoSQL. But we have DynamoDB streams, we have the ability to do scans and exporting data and some of that stuff. So just what are some of the best practices when it comes to taking that data and being able to analyze...

Rick: Absolutely. Yeah. You hit the good points there. Streams. For operational analytics, things that need real time aggregations. We're looking at like top end, last end counts, sums, averages, computed KPIs.

Streams and Lambda is your friend, right? I mean Streams is the running change log of DynamoDB. It's like a change data capture pipeline that's built into the process, built into the system and so when you update the DynamoDB table with either a write or an update or delete, any write operation, it's going to show up on the stream which causes a trigger to fire and that trigger can be picked up by a Lambda function and the Lambda function can process that change and update any operational metrics that are affected by that change.

So this is a really neat system because there's 100% SLA guarantee between the update to the table, the right to the stream and the fire of the Lambda function, it's going to process every single update at least once. So this is really useful for customers who are trying to maintain these operational analytics because you're guaranteed the process, right? If you try and process it all yourself, I mean, how processes die, right?

And if you managed to update the table, but not your analytics, then you know you're not going to be able to make sure that happens. We'll make sure that happens for you. So that's really neat. The other process you get is like you said, you can table scan, export, but one of the things I see people do a lot is they just actually snapshot the table, it's one of the nice things about table snapshots is they're fully consistent, right?

Rick: A table scan is not necessarily consistent, right? It starts at the first item and ends at the last item. If anything changed in between, right? I don't know. So if you do a snapshot, then you can that to a new table and then you can table scan that new table and it's like a point in time picture of your storage which is really what most people want when they're running these types of offline analytics, right?

And so you're going to snapshot this thing, you're going to restore it to this new table and then you can go ahead and export it to S3 as parquet files. You can run a Athena queries on top of this thing and do whatever you want. That's not the most efficient way to query the data, but it's highly effective and what you have is it a report that doesn't run with high frequency, that can be a really, really nice solution for you because you don't have to export it into a relational database or even putting into Redshift or anything.

Now, if you're running constant queries against this data, then maybe a regular process to using Streams, Lambda to export the data in realtime into a relational database to maintain kind of a synchronized view, so to speak as a normalized structure that you can run ad hoc queries against. I see that a lot too. All right, so depending on the nature of the system and the requirements of the analytics, we can handle it, but let's just make sure we do it the right way so that those ... We don't want to end up having to run a lot of random analytics queries on the NoSQL database. It's just not going to do it, right? It's not going to do it well.

Jeremy: Yeah, and one of the ... I mean, I guess the mindset that I follow is if I'm doing time series data or something that is immutable, it's just not going to change, it's writing the data in. I like to dump that into like Kinesis data fire hose and maybe S3 and then be able to query with Athena, but I absolutely love when I use operational type data where it's a lot of crud type stuff that's happening. Just copying that over into a SQL database so I have all that flexibility, but from an operational standpoint, that's my source of truth and I think is...

Rick: Yeah, no, absolutely, absolutely. Yeah, absolutely. As far as the time series data goes, that's a really good use case for Dynamo and we'll see a lot of people roll that time series data and use that Streams, Lambda processing to update those partitioned analytics, right?

Top end, last end, average and all that stuff and then they'll age out the actual item data, right? And do exactly what you said. So they'll TTL that data off the table, it will roll up into S3 go into parquet files and sit in S3 and then when they need to query it, they just select the top level roll-ups out of DynamoDB.

If they need to do some kind of ad hoc query, then they do exactly what you said, run the Athena queries and whatnot, but for the summary aggregations, they're still serving that up out of DynamoDB. They're just doing it and it works really well like you said because once those time-bound partitions are loaded and they're loaded, they're in chain.

So why calculate it every single time? Right, exactly. Yeah. Yeah.

Jeremy: So one of the things you mentioned about Streams too is that it guarantees at least one's processing. So you still have to think about item potency and some of those other things, if you're updating something else. I mean are there any best practices that you can think of for that or ...

Rick: I mean, I think it's just always a one-off, right? I mean it depends on what nature of the computation ... Some computations aren't even affected. If you process it multiple times, others you're going to want to make sure that you included this in the average already.

There's usually what I'll do is maintain track of items, events that are processed, what are the last end events or something like that so that if something processes twice, I'll see that it actually made the update.

Write the configuration data that normally would write from the Lambda function. You know these things are processing in order on a per item basis. If you have per item metrics, you can always record the last ID, right? So that when you come to update again, make sure that the ID isn't equal to my ID. If it is fail.

I mean, it's basically, it's going to be some trick somewhere, somehow. Oftentimes I end up tagging a UUID onto the items that I can use in exactly that way so that I know it's processed, it processed right? That way, balance those double-process things. Now again, it doesn't happen very often. I mean, like one in a million is going to and how you could run for months and never see it.

And of course, there's going to be that one random time where the lightning hit the data center and the container crashed and the thing was in the middle of your Lambda process, right?

Jeremy: It's always the customer who pours over their data that's going to find the issue.

Rick: Exactly. He's going to notice it. Right. Exactly.

Jeremy: All right, So another question that I got was the performance impact of transactions.

Rick: Okay. Yeah, yeah. So transactions are heavy, right? There's no doubt about it. It's basically a multi-phase commit across multiple items. So I don't know the exact numbers, but it's about three X the cost, right? To have the normal insert. So be aware of that when you're using the transact write API. If you have the need for that kind of strongly consistent update, then let's do that, but there's a couple of caveats to transactions people need to be aware of.

First-off is the isolation level is low. So this does not prevent you from seeing the changes on the table, right? You will see those changes appear across the table. Someone selects an item that's in the middle of a transaction, they will see it, nothing blocks the read. So if that's ... As long as that type of transactional functionality is right for you and that's really what we're doing.

We're giving you an ACID guarantee. It's an ACID guarantee with a low level of isolation. You also get essentially the same ACID guarantee from a GSI replication and that is guaranteed and it doesn't cost you any more than the cost of the right and the cost of the throughput. So again, as long as it's a consistency at the client, really that's what transactions gives you. You won't acknowledge the write to the client until all the updates have occurred. That is the only difference between a transact write API and a GSI replication.

So if that guarantee is something you absolutely have to have, then great. And there are plenty of use cases for that. I want to block until I know that all copies of this item have been updated. I don't want the to continue until X, right?Great, no problem, but it was transact write API. That should be a real subset of your use cases. I wouldn't use it by default.

Jeremy: All right, and then what are some of the biggest mistakes you see people make with modeling? And we only have a few more minutes.

Rick: Yeah, no problem. Look, the biggest mistake. Hands down. We use multiple tables, right? I mean and the bottom line is multi-table designs are never going to be efficient and NoSQL no matter what the scale. I mean, you can have the smallest application that you're working with, you can have the largest application you're working with, it's just going to get worse, right?

The small application you might not notice the cost that you're paying, but you will pay more and honestly, it's not easier to write data, write the application for a multi-table design, right? It's just not. I have to write multiple queries, I have to execute multiple requests. I have a lot more code that I have. I mean, if you want to compare code, it's going to explode the code to run multiple payables, right?

So you're going to be running with less code, less complexity, more efficiency with a single table design. Let's learn how to use that. I think that's really the biggest mistake I see people make. The bottom line is if you can't get over that, then stick with your relational database. You're going to be way better off, right? So I would definitely advocate that if you're saying that my app is too small and I don't need a single table, then okay, your app is too small, you don't need NoSQL, right?

Jeremy: So what about like a relational modeling? Is that something you see quite a few people doing?

Rick: Oh, in the relational modeling in NoSQL?

Jeremy: Yes. Yeah.

Rick: I'd probably say 90% of the applications I see for NoSQL like single items. There's not a lot of highly relational data. I would say that the majority of applications, it's probably about 50 to 60% of the applications that we migrated at at Amazon were ... Had a fairly significant relational model and that was because we were taking ... We were under an edict that basically every app that we had, all of our tier one applications, most of our tier two applications were moving to NoSQL and there was no choice.

So across customers, I'm seeing a larger number of apps these days. I'm seeing people starting to realize that, "Hey, you know what? We can manage this relational data. It's just it's not non-relational. It's de-normalized." Right? I think people are starting to understand that that non-relational term is a misnomer and that we're really just looking at a different way of modeling the data. So you're starting to see more complex relational data in these NoSQL databases which I'm really glad to see and I expect that trend will continue.

Jeremy: All right, and then the other thing I think you've mentioned in the past too is we talked about it a bit with that partial normalization on the insurance quote is storing really large objects.

Rick: Yeah, yeah, don't store the ... Don't use that large objects unless you need them, right? I mean, if the access pattern is I need all this data all the time, then great. Use the biggest objects you can. It's actually drives a lot of efficiency. But generally speaking, I don't see that, right? Most applications need little small chunks of data, right? This couple of rows from that table, this couple of rows from that table, they don't need the entire hierarchy of data that comprises this entity in the application space, right? And so if you're storing blobs of data that represent every single thing that could ever be related to X, then the chances are that you're actually working and using a lot more throughput and storage capacity than you need.

Jeremy: Yeah, and I think a thing to remember about that too is you're not paying for the number of items you read, you're paying for the amount of data that you read.

Rick: That's right. Yeah, that is the case with every NoSQL database.

Jeremy: Exactly. And so what's cool about that too is, I mean, if you go back to those GSIs, you may have certain attributes in a document, some of which you have to index a different way. You can use sparse indexes and only replicate some of those attributes or some of that data to another index.

Rick: That's correct. Absolutely. And you can pick and choose which pieces of that hierarchy end up getting projected onto those index is absolutely ...

Jeremy: Right. So optimize for the write, right? That's another sort of main thing to think about, the velocity of the workload and so forth. Those are the ...

Rick: It's a choice. It's optimized for the read or for the write, depends on the velocity of the workload. Depends on the nature of the access pattern. There could be times I want to do one versus the other.

Jeremy: Absolutely. Okay. All right. So just a couple more questions here, but I think it's really interesting and I had posted something about this a while back and you had mentioned all the applications that you moved over from amazon.com to to NoSQL, what are those numbers look like? Because I think a lot of people are like, "Well maybe I can't use DynamoDB." Or something like that. And I always say, "Well, if Amazon can fit 90% of their workloads into DynamoDB, you probably can too."

So just, do you have some of those numbers? What has the growth been like and how much data do you process on a given day or whatever?

Rick: Yeah, sure. So a raw number is, I mean, if you want to just think transactions per second, 2017 prime day, I think we peaked Amazon Tables peaked it about 12.9 million transactions per second. We thought that was pretty big. That was 2017. In 2019, Amazon CDO tables peaked at 54.5 million transactions per second. So it's in a [crosstalk 01:28:11]

Jeremy: So somebody else's application can probably be just fine using DynamoDB.

Rick: And this is the thing, so we get this question a lot. I mean, I get the question of is DynamoDB powerful enough for my app? Well, absolutely. As a matter of fact, it's the most scaled out NoSQL database in the world, nothing does anything like what DynamoDB has delivered. I know single tables delivering over 11 million WCUs. It's absolutely phenomenal and then the other question is is not DynamoDB too much overkill for the application that I'm building?

I think we can have great examples across the CDO of services. Not every one of our services is massively scaled out. Hell, I've got services out there, I've got five gigabytes of data and they're all using DynamoDB and the reason why I used to think that NoSQL was the domain at the large scaled out high performance application, but with cloud native NoSQL, when you look at the consumption based pricing and the pay per use and auto scaling and on demand, I just think you'd be crazy.

If you have an OLTP application, you'd be crazy to deploy on anything else because you're just going to pay a fraction of the cost. I mean, literally, whatever that EC2 instance cost you, I will charge you 10% to run the same workload on DynamoDB.

Jeremy: Yeah. And actually, I really like this idea too of using it in just even small applications as sort of a very powerful data store that yeah, if for some reason I happen to get 54 million transactions per second at some point ...

Rick: I can do it. Right. Yeah.

Jeremy: Most likely, I'm going to have 10 transactions per minute or something on some of these smaller things, but I always love the fact that DynamoDB tables are such ... They're so easy to spin up, right? And again, you can add GSIs, but if I think about building microservices, especially with serverless applications, I don't want to be spinning up a separate RDS cluster or Aurora serverless for every ...

Rick: Yeah, no. I mean ...

Jeremy: It's crazy and if most of what I can do, if my workload is fit, if my access pattern is fit and there are some that don't, but yeah, I love it as sort of this like go-to data store that you can do all these great things in and the other thing is is that if you do have some slightly complex queries that you need to run, but it is small scale, a couple of indexes don't cost a lot of money.

Rick: Yeah, processing data and memory doesn't cost much money. I mean, the best example is in what the community tells us. I got a tweet from a customer the other day said, he told me just deprecated as MongoDB cluster had three small instances. They were costing about $500 a month. It took him 24 hours to write the code, to migrate the data and migrate the data. He switched it all over to DynamoDB, he's paying $50 a month.

Jeremy: Oh geez.

Rick: I mean, it's just, it's amazing when you look at it. I mean, when you think about it, it's like, "That actually makes sense because the average data center utilization of a an enterprise application today is about 12%." Right? That means 88% of your money is getting burned into the vapor.

Jeremy: And you're paying those people to maintain the ... Paying those ops people to maintain that ...

Rick: On top of that, exactly. And that's the thing. So this guy is like, "Hey, I'm saving 90% of my base cost and that didn't even calculate his human cost of maintaining all those systems."

Jeremy: Which is likely much more expensive than ...

Rick: Than the 500 bucks a month, right. Yeah.

Jeremy: I will just tell you, I actually had a small data center. I had a co-location facility that I used when I had a hosting company when I was doing a web development company and I used to go there and swap out drives and do all that kind of stuff. I can tell you right now, Amazon or AWS can do it a lot better than [crosstalk 01:31:39]

Rick: We've been doing it for a while now.

Jeremy: Right. Exactly.

Rick: Petty much the scale that blows anybody else out of the water, right?

Jeremy: Exactly. Exactly. So definitely a wise choice to do that. All right, so let's move on to tools for development. So there's the NoSQL Workbench, I've been playing around with it, very cool.

Rick: Yeah, yeah. No, I love this thing. This was a tool. Again, and I said it and reinvented, built by the specialists, for the specialists. The North American Specialists SA team had been spending a significant amount of time plowing around in Excel, manually creating GSI views for customers and demonstrate to them how these things will work. It's an error prone process. It's a pain. To try and create those pivot tables, it's like, "Look, you're copying data out. Oh did I get the keys right? Is the sort order right?" All this kind of stuff, right? Whereas with this, what you do is that those NoSQL Workbench for DynamoDB, you just take a bunch of JSON data, you can load it into the tool, it gives you a nice view of what the aggregate looks like based on the partition key and sort keys that you configure in the tables.

It gives you all your GSI views. That's the best part about it. As you go from GSI to GSI, GSI, it pivots the data automatically. So you can put the sample data in and then you can visualize what happens when I translate that data across multiple indexes and then see what those sorts look like. It has a code generator. I don't know if you had a chance to play with that yet, but...

Jeremy: I did.

Rick: Oh, it's really nice because you know what I mean? One of the biggest problems in dealing with any database development is writing the queries, right? I mean, it's like we all know what conditions we need. Okay. Write the code, make sure that there's no errors, everything's correct. And the code generator for the Workbench basically lets you just set query conditions and hit the button and it generates all the code for your application. You can just cut and paste it. Actually, generates a runnable, executable.

Jeremy: It does. Yeah, that's right.

Rick: You can validate it from the console. Just go ahead and run it. Yup, it works. Okay, great. You can embed it in your application.

Jeremy: Well, and and the other thing that's great about it too is, and I think this is something that, because there's no interface to do this yet and you still have to add it to JSON manually is the facets. And I've been playing around with those a little bit. I'm sure you'll [crosstalk 01:33:49] to do it through an interface, you don't have to download and stuff, but the facets are actually really cool because this I think will be helpful to people who are thinking about entity-based type stuff.

Rick: Correct.

Jeremy: So we can sort of create this new entity for each different ... Facet for each type of entity, you go in, you can enter data for each facet which it maps or aliases, your PK and your SK. It just makes a lot more sense I think than people just entering everything into one big table.

Rick: Yeah. You hit the nail on the head there. Facets were intended to be the entities in your model. For each type of object that you have on the table, that's a facet, right?

Jeremy: Yes.

Rick: Now, I expect that over time, the functionality of facets is going to grow. It's going to start to align well with like data migration where customers want to move from a relational database into workbench and we're going to start to provide tooling to help them do this.

That will help them translate their normalized relational models into a single table and keep them in the context they're used to mentally because these items are related to those items. Here's a one to one, here's a one to many, here's a many to many. I'll create all my facets. One of the facets is my lookup table. Just to help people organize their data the way they like because we get that question a lot. "Well, I put all my data in this table, how do I visualize it?"

Jeremy: Right. Right.

Rick: There you go.

Jeremy: No, that's great. All right, so then other tools, there's the best practices guide. On the AWS site, there's some courses out there, the Linux Academy course.

Rick: Yup, Linux Academy, we just put out that new course out. That was published just last year, it incorporates all of our best practices and design patterns, it's completely updated content.

We've got, of course my content online. There's some great content from the community out there. Alex DeBrie has some really good content out there for people to pick up and we've got a couple of books in the queue here, so look...

Jeremy: Yeah, I heard. No, I know Alex DeBrie is writing a book and then I think you said I'm not going to be outdone by Alex DeBrie, but I'm also going to write a book. You have a book you're working on as well?

Rick: Yeah, we do as a matter of fact. We're leveraging the entire North American Specialist SA Team at AWS. They're all contributing content. I'm going to be editing the content. I'll probably write the forward and I look for that to come out sometime probably early next year I think is when that's going to hit.

Jeremy: Awesome. All right, so Rick, listen, I mean honestly, it's been absolutely awesome to have you here. I will say I don't think I would have ever discovered or found a love for NoSQL and DynamoDB if it wasn't for the presentations that you've done and the work that you've done.

So I really appreciate it. I know everyone that's listening and everybody in the sort of the DynamoDB community is a very appreciative of the work you did so again, thank you so much for being here. So if people do want to find out more about you, follow you or whatever, how do they do that?

Rick: Sure, you can hit me up on Twitter houlihan_rick and or hit me up on LinkedIn and we can connect there.

Jeremy: Awesome. Okay.

Rick: Thanks so much Jeremy. I really appreciate it. It's been great.

Jeremy: All right. Thanks Rick.

Rick: All right, thanks. Bye.

View Details

About Rick Houlihan:

Rick has 30+ years of software and IT expertise and holds nine patents in Cloud Virtualization, Complex Event Processing, Root Cause Analysis, Microprocessor Architecture, and NoSQL Database technology. He currently runs the NoSQL Blackbelt team at AWS and for the last 5 years have been responsible for consulting with and on boarding the largest and most strategic customers our business supports. His role spans technology sectors and as part of his engagements he routinely provide guidance on industry best practices, distributed systems implementation, cloud migration, and more. He led the architecture and design effort at Amazon for migrating thousands of relational workloads from Oracle to NoSQL and built the center of excellence team responsible for defining the best practices and design patterns used today by thousands of Amazon internal service teams and AWS customers. He currently work on the DynamoDB service team as a Principal Technologist focused on building the market for NoSQL services through design consultations, content creation, evangelism, and training.

  • Twitter: @houlihan_rick
  • LinkedIN: https://www.linkedin.com/in/rickhoulihan/
  • Best Practices for DynamoDB: https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/best-practices.html
  • 2017 re:Invent Talk: https://www.youtube.com/watch?v=jzeKPKpucS0
  • 2018 re:Invent Talk: https://www.youtube.com/watch?v=HaEPXoXVf2k
  • 2019 re:Invent Talk: https://www.youtube.com/watch?v=6yqfmXiZTlM

Transcript:

Jeremy: Hi everyone, I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Rick Houlihan. Hey Rick, thanks for joining me.

Rick: Hey Jeremy. Thanks for having me on.

Jeremy: So you are a principal technologist for NoSQL at AWS. So why don't you tell the listeners a little bit about yourself and what it is that you do at AWS?

Rick: Yeah, sure. So I've been at AWS almost six years now, I guess five and a half years. My primary focus in life since joining AWS has really been NoSQL technologies. Shortly after joining organization, I joined the specialist team and then for about two years, I spent a large amount of my time focused on the migration of Amazon's internal application services from a relational database, specifically Oracle to a NoSQL technology which was DynamoDB, of course.

So that was my mission in life and the last two years or so, I've been more focused externally taking the learnings that we gained from that exercise out to our customers and helping them solve similar problems.

Jeremy: Awesome, and I love the fact that you are actually working with customers and working with these big data problems, actually solving these problems as opposed to just sort of advocating for ...

Rick: ... thinking about them, right?

Jeremy: Exactly. Exactly.

Rick: We got called out. Yeah.

Jeremy: And that's, I mean, again, obviously this firsthand experience and all this work you do, you see all these different permutations and different ways that you can use DynamoDB and NoSQL. So I think if anybody is in the sort of AWS ecosystem, if they've ever sort of thought about using DynamoDB, your name has probably come up. You've become somewhat of a legend at AWS re:Invent with your NoSQL talks on data modeling in NoSQL.

So there are a million different things that we could talk about obviously, and I could probably talk to you for quite some time. I don't want to spend a lot of time rehashing the things that are in your presentations and I will put these in the show notes and people can go and spend some time looking at these things. But I actually, I want to be a little bit selfish here because I have all these questions that I've sort of come across and some people have asked me and now that I've got you here, I would love to sort of ask you those and see if we can dig a little bit deeper in them.

But what I do want to do is be fair to people who are not overly familiar with DynamoDB or NoSQL in general. So maybe we start quickly and just kind of explain or if you could explain to us what's the difference between NoSQL and relational databases or RDBMS.

Rick: Sure. Yeah, great place to start. So if you think about the relational database today, it's about normalized data. We're all very familiar with the idea of a normalized data model where you have multiple tables, we have all these relationships, parent child relationships and many to many relationships. And so we built these tables that contain this data and then we have this ad hoc query engine that we write called SQL. We write queries in SQL to return the data that our application needs.

So the server, the database actually restructures the data and reformats the data on the fly whenever we need it to satisfy a request. Well, NoSQL on the other hand eliminates that CPU overhead and that's really what the cost of the relational database is and the reason why it can't scale because it takes so much CPU to reformat that data.

So with NoSQL, what we're going to do is we're going to actually denormalize the data somewhat and we're going to tune it to what we call the access pattern, tune it to the access pattern to create an environment that allows the server to satisfy the request with simple queries.

So we don't actually have to join the data together. So we talk about the modeling and whatnot in my sessions, we can get into how do we do that, but the fundamental crux of the issue here is that the relational database burns a lot of CPU to join the data and produce these materialized views whereas the NoSQL database kind of stores the data that way and makes it easier for the application to use it.

Jeremy: All right, and then one of the things I think that you see all the time when people are sort of migrating or trying to figure out that NoSQL mindset is they think about access patterns and one of their access patterns is something like list all customers.

Rick: Right.

Jeremy: And it's one of those things where I think it's really hard for people to make that jump from just being able to say select star from customers and understand that data will get back and we can add limits and things like that. It's not quite the same with something like with NoSQL, so when do you suggest people not use NoSQL?

Rick: Okay, so that's actually a really good question. So NoSQL is really suited and as we talked about, we have to denormalize the data, right? Which does that means I have to structure it and tune it to the access pattern. So if I don't really understand those access patterns, if they're not really well-defined, then maybe what we're looking at is a different type of application that's not necessarily so well-suited for NoSQL, right?

And that's really what it comes down to. There's two types of applications out there. There's no OLTP or online transaction processing application which is really built using well-defined access patterns. You're going to have a limited number of queries that are going to execute against the data, they're going to execute very frequently and we're not going to expect to see any change or we're going to see limited change in this collection of queries over time.

And that's a really good application for NoSQL because as I said, we have to kind of tune the data to the access pattern. So if I only have a small number of access patterns, then it makes sense, but if the customer comes in and tells me, "I don't know what questions are going to be asked. This is maybe my trading analytics platform and who knows what the brokers are going to be interested in today or tomorrow and I look at the query logs of the server and there's a thousand different queries and some of them execute once or twice and never to be seen again and others execute dozens of times."

These are things that are indicative of an application workload that may be, is not so good for NoSQL because what we're going to want is a data model that's kind of agnostic to all those access patterns, right? And it has that ad hoc query engine that lets us reproduce those results. So lucky for us in the NoSQL world, that's actually a small subset of the applications, right? 90% of the applications we build have a very limited number of access patterns. They execute those queries regularly and repeatedly throughout day. So that's the area that we're going to focus on when we talk about NoSQL.

Jeremy: Yeah, and so with those access patterns and you talk about highly tuned access patterns, and if you think about an application that says maybe it has to bring back customer orders, right? Maybe a customer might have 10 orders, maybe they have a thousand orders. I mean, really, the amount of processing it takes to pull back either 10 or a thousand is pretty much the same, but there's other things that might come back with a customer as well.

Maybe you want to see their billing information, their shipping information or maybe they have a rewards program or something like that that they're attached to and one of the interesting things that I learned from you from seeing what you've done at your talks at re:Invent was the ability to put multiple entities if we want to call them that in the same table.

Rick: Sure.

Jeremy: And that was one of those optimizations where SQL we say, "Okay, select * from customers where customer equals one, two, three. So let's start from orders where this equals that and maybe join it on order items on ..." Things like that. And also, now I can make another query and I've got to bring back those rewards or bring back the shipping information. So you have all these different queries that you have to run, but when you put everything into a single table and you optimize that access pattern, you can make one query that brings back all of those entities in one round trip to the server.

So I'd love to get your take, I know you're a huge proponent of single table design, but I just like to get your take on why that's so powerful.

Rick: Why that's so powerful. Sure. It makes perfect sense. So if you think about the time complexity of a query, right? When I have to join through multiple tables, I will use the example you talked about. I've got customers, I've got orders and I have order items, right? Let's just use those three tables for example. And what I really want is I want all the customer's orders in the last 30 days, right? So I can select star from customer where customer ID equals X, inner join orders on customer ID equals customer ID, inner join order items on order ID equals order ID, right?

So I've gone through multiple tables and the time complexity, that is going to be significant because I've got a ... And let's just assume that the joins are all accurate or that the indexes are all accurate and we're sorted on the joint dimensions and so we're doing really nice, efficient nested loop joins, we're not doing any crazy things in the database and we're still looking at a time complexity. It's going to require a login search or an index scan of the primary table and in login complexity for the next table, right? I got to scan that table for each item that comes back off the parent table.

Now you might argue with me that the outer table has a single row return. So it's just who login, but then the next one comes, right? It's the order items table and now I have to scan through all the customer's orders and then I have to join in all the items from each order on that table and that's an in login scan.

So it just increases the complexity, right? It's the more tables you join in. Now, if we were to take the NoSQL approach where I treat the table like a big object collection, I'm just going to drop everything in there, just then you think of these as all the rows from all those tables, we just take all those rows, we're going to shove them into one table and then I'm going to go ahead and index these objects on maybe tag each one of those objects with the customer ID, including all the order items and everything.

And then I can index and I say, "Okay, well give me everything for the customer IDX." And it brings back the customer object, the order objects, the order item objects, or all those orders. Now, I might want to include some additional sort constructs to return a limited number of objects, but you get the idea. What I've really done is I create a collection of objects and I use indexes to join those objects together.

There is no joint operator, but I can still achieve the same result because a join is essentially a grouping of objects and so that's what we're really doing with the NoSQL table. So it's much, much more time efficient to do an index scan than it is to do nested loop joins and that's really what it comes down to is about cost and efficiency.

Jeremy: Yeah, definitely. Yeah. And I like thinking of them as entities in the sense of it's almost like a partition keys, and we're going to talk about this in just a minute, but I like to think of partition keys as almost like folders, the directory structure and you're grouping all of the things in those folders that are common to one another that you might want to bring back.

And you've got a few short features and you've got some ways to limit it like by date and things like that. So anyways, let's actually do that. We can maybe come back to the single table thing as we talk about some of this other stuff, but just for people who don't know and just to sort of hear it again so that we can all kind of be on the same page because I think we're going to get, we're going to start geeking out pretty heavily here in a minute and if you don't understand these basic concepts that you ... I don't want your eyes to gloss over too much if you're listening here.

So let's start with a couple of these key concepts. So I mentioned partition keys and we also have sort keys. So just give us a quick, what do those mean?

Rick: What are they?

Jeremy: Yeah, what are they?

Rick: So in a DynamoDB table or in any wide column database in any NoSQL database, you have to have some attributes that uniquely defines the item. In a wide column database like DynamoDB, that's called the partition key. So if I define a table in DynamoDB and I define it as a partition key only table, then each item I insert in the table must have a partition key attribute and that partition key attribute must contain a unique value and that supports kind of a key value access pattern, right?

Give me everything with this partition key value, brings back one item. If I add the sort key, what you said comes into effect now is now the partition turns into a folder, right? So the partition uniquely identifies the folder and the sort key uniquely identifies the item within the folder and now I can start to collect items together and I can use creative filtering conditions on the sort key value to return only the objects that matter.

And I really liked that folder analogy because if you think about it, when I store things, when I collect documents at home and I'm putting things in my file cabinet, well what am I doing? I'm putting related objects into a folder and I'm putting it in the file cabinet and because that's the easiest way for me to go access it in the end, right?

I say, "Here's everything associated with my mortgage." Right? I don't say, "Here's the home inspection." And put that in one place and then, "Oh, here's the title document, title report. I'm going to put that in another place." And then when I need to go get all those documents for my home loan, I don't go searching through 20 folders to pull together all those documents, review them, and then sort them back into their individual locations, right?

We put everything in one folder and it turns out that that's actually the easiest way to access data, not just for us, but for the computer too. So this is really what we're doing. We're using these partition keys to create folders, we're using sort keys to individually identify the objects within those folders and then we're using conditions on the sort key to restrict the objects that are being returned when I query that particular partition and that's kind of the fundamental construct that we're trying to create here.

Jeremy: Yeah, and so the ... So when you use those two together, when you use a partition key and the sort key, we call that a compound key and then the sort key itself, if we add extra data in there, you had mentioned we could filter on it for example and use something like a begins with query and things like that. So if we combine multiple pieces of information into that sort key, that's what we call a composite key?

Rick: That's correct. So a composite key typically takes multiple, like you said, multiple pieces of information to give me some interesting constructs to go query, right? So I might like say, "Well, a customer's use case is to give me all of his orders within the last 30 days." I might choose to prefix all the order objects in the customer's partition with an indicator that they are in order like O and then concatenate that with the date of the order.

And then I could start, I can say, "Give me everything that's between O - date one and O - date two." And it will return the order objects from that customer's partition. I might have many, many types of objects in that customer's partition, but that sort condition will allow me to filter out just the specific order objects in that date range that I'm interested in.

So that's kind of an example of the types of composite keys that we're going to create to produce queries. Those things can include things like states, dates, rankings, all kinds of different things that we're going to do to try and create conditions that we can query.

Jeremy: And automatically when ... No matter what you create for that sort key, that is ... They call it a sort key because it's what gets sorted on and you can ... But so the way that that sorting works is that if you had O# and then some date for example that anything, if you did a between query for example, you would have to include the entire sort key, the entire prefix in there as well.

I think that's something that trips people up. If you were trying to order ... Maybe you're trying to order a list of songs or something like that. If you had your sore key as one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, then it would sort as one, 11, 12, two, three, four. So you just need to be conscious of that to say it needs to be zero one, zero two, zero three, zero four, and so forth in order to get that sorted.

Rick: Yeah, zero padded or convert those to a maybe a four byte hexadecimal string or something that is string sortable. This is one of the kind of little caveats in using DynamoDB is that we really only give you one sort key attribute. So if we want to create these composites, you have to kind of create these string composites and stick it in that sort key attribute. If you look at other wide column databases like Cassandra, Cassandra gives you the ability to actually support true composite keys, right?

You can actually say, "I want to ... This attribute dot, that attribute dot, this attribute ..." And that would be the concatenation of those is your key and it will allow you to sort correctly between the data types. So that's a little bit of a difference between a Cassandra and a DynamoDB is that you have to maintain these composites kind of yourself and DynamoDB, but other than that, it's almost identical.

It's the same type of query conditions, right? In Cassandra, when I query the first attribute, it has to be an equality condition and the range condition can only apply for the last attribute that you query on the composite and the reason why is because all those range queries have to return a contiguous range of items from the CFP.

Jeremy: Right.

Rick: Otherwise, what I'm really doing is filtering the items from the partition.

Jeremy: Right. All right, so the other thing that all databases have is some sort of index, right? And obviously, what we're talking about with the partition key and the sort key, that's our primary index and we're going to talk about that a little bit more because there's some interesting things there, but one of the things that you can do is create other indexes LSIs, GSIs.

Again, we'll talk about all of that sort of, "Don't get lost if you're listening." But I do want to talk about this idea of overloaded indexes because this goes back to the single table design aspect of things where you have different values in different indexes. So that partition key for example, it might be a customer ID for a bunch of entities, but then it might be an order ID for other entities. It might be some other enumerated value to specify something else. So talk a little more about overloaded indexes.

Rick: Okay. So in DynamoDB, we've talked about the primary construct being this partition key and a sort key. The idea there is I'm creating groupings of objects and those object's grouping should be kind of related to the primary access patterns of your application, right?

In the customer example, we were talking about the partition key, it might be a customer ID, but there might be other objects on this table that I'm interested in tracking, right? Not all these objects might be customers, not all these objects might be related to customers and so let's say we had sales reps also. So I've got customer IDs and I've got a sales rep IDs.

So in order to kind of store customers and sales reps on the same table, when I define the table, I can't use a strongly typed attribute name, right? If I use customer ID as the partition key, then every object I insert into the table has to have a customer ID as a partition key, but sales reps aren't customers so they don't have a customer ID.

So when we define the table, we're going to use generic names. Things like PK for partition key, SK for sort key and this allows me to create multiple types of partitions. The object that I insert into the partition just needs to have a PK attribute. The value of the PK attribute depends on the type of the object that I'm inserting on the table and I don't have to worry about collisions here because when I query the system, I'm going to give that partition key quality condition.

It's going to have a value that's my application is aware of what it's trying to get, right? I'm going to query for a customer, I'm going to query for a sales rep, right? I know what I'm asking for. And so then what I'm going to do now is I'm actually going to say on the primary table that first access pattern is supported by the table, might be orders by customer, but I also might have a workflow that says at the end of the month, I need orders by sales rep or my customer, my orders aren't organized on the table by a sales rep, they're organized by a customer.

So what I'll do is I'll create an index, a GSI, and I'll use a ... I'll declare a new attribute on the GSI and I'm going to use that same naming construct, right? I don't want only sales rep items to show up on the GSI. As a matter of fact, I actually need customer objects. I need order objects to show up on the GSI and on the primary table, my orders are partitioned on the customer ID, not the sales rep ID.

So I'll create an attribute called GSIPK and all of the orders are going to have it and all of the sales rep items are going to have it and you know when we insert the sales rep item onto the table, it's going to have a partition key of the sales rep ID, it will have a sort key of the sales rep ID and then it will have this extended attribute which is also the sales rep ID. So you kind of notice we've duplicated that sales rep ID a few times, right? The reason why is because I'm going to index this item in multiple dimensions on the sales rep ID.

The first dimension that's actually indexed on is the table that's within the sales rep partition. Okay, while all those other order items are also going to have a customer ID as their PK, they're going to have a maybe an order date as there SK and then they're going to have a GSIPK, a sales rep ID. And so if I create a GSI on sales rep ... If I bring up an index with GSIPK as the partition key and the sort key, the existing sort key as the sort key for the GSI, what I've really done is I've re aggregated or regrouped the items on the table.

Now, the orders on the GSI are grouped by sales rep and if I query the GSI by sales rep ID, I'm going to get a copy of the sales rep's item and each order that he had and if I query it by sales rep ID greater than 30 days ago, I'll get only this stuff in the last 30 days, right? And so this is kind of what we're going to do. We're going to re aggregate, regroup the data in indexes to support secondary access patterns in the application and that's what we're going to use indexes for.

Jeremy: All right. So you mentioned GSIs in there and I want to talk about those a little bit more in detail, but let's start first with just this idea of accessing data via the primary index, right? So every table has the primary index. That is where you have your PK and your SK or whatever you want to call it.

Rick: Correct.

Jeremy: And one of the things that I see a lot of people do when ... Especially when they're asking me these questions and they're trying to format that is they often ... I think what they're trying to do is I guess what's the right word for this? They're trying to make DynamoDB work more like a SQL database and that they're trying to find a way to put ... Either create really, really large partitions and then use the sort key to trick the system into doing very large queries, but the thing that you need to remember or I think that the listeners need to remember, I know you know this is that we want to be careful about how much we're writing to the database when we're using some of these other indexes, right?

So for the primary key itself, or for the primary index itself, we want to, I guess, we want to make sure that we have all of our or as many of our primary access patterns accessible via that, right? When we want to be limited in terms of the amount of data we copy over to another index, right?

Rick: Yeah, I mean, if you think about it, every index I create is a copy of the data.

Jeremy: Right.

Rick: So if I can write the data onto the table one time and satisfy more than one access pattern, then that saves me copies of the data. Those copies of the data costs me money, they cost storage, they cost WCUs. So yes, we would like to store as many ... We would like to satisfy as many access patterns as possible with the primary table and then what you'll notice is as I start to create indexes, objects start to drop off because not all objects need to be indexed.

Not all objects need more than one access pattern. Some objects only have one and other objects might have three or four. So as you kind of go out across the number of index, as you're going to see fewer and fewer items transferring from the table to those extended indexes because they just don't need to be accessed on as many dimensions as as others. And so that's kind of what we're going to end up doing is looking at these individual items and say, "How do I need to group these?" Right? In the case of orders and I believe we've talked about so far, I need to group orders by customer, I need to group orders by sales rep, right?

But order items, I don't know, we haven't defined any other access pattern than group order items by order. But maybe there's another access pattern by order items, which might be the state of that order, right? Or of that item, is it back ordered, has it been shipped or whatnot. So we already have an index called our first GSI that we're indexing by sales rep and we're indexing order items by sales rep, but I can also use the same index, to index those order items.

If I decorate each one of those order items, it would say an order state that is also in GSIPK and then I'll just use the sort key of the item and I can say, "Okay, well which one of these things ... What orders were in this state at this date time?" Right?

Jeremy: Yeah.

Rick: And so those types of access patterns are we're going to do.

Jeremy: Yeah, and I think what I was ... I think the point I'm trying to get to is because I get this question a lot and I think you've explained it well in the past where when you ... So RCUs and this is something we didn't really talk about, but there's read capacity units and write capacity units for DynamoDB tables and the read capacity units, I think it's three...

Rick: 3,000. 3,000 RCUs.

Jeremy: Yeah, 3,000 versus 1,000 and you can read, it's the amount of data that you can read off of a WCU is ...

Rick: A WCU is four kilobytes, an RCU is one kilobyte. And I'm sorry, WCU is one kilobyte, an RCU is four kilobytes and I said that backwards.

Jeremy: Yeah, okay. So the point is is that you can read four times as much data as you can write in a single capacity unit and then obviously, a write capacity units cost a little bit or costs a lot more than the read ones do.

Jeremy: So I think what people tend to do, and this ... You say this all the time is that you should optimize for the write and I don't think people quite get what that means and that's, I think what I'm trying to get to here is that for every piece of data that you copy, you have to write that to another index. So if you're writing it to a GSI, you're paying for another WCU and it's also one kilobyte versus reading four kilobytes off of it. So if you are trying to write a lot of data, if you have a high write throughput, but you need to access that, putting another index to there just for the sake of maybe querying on one dimension might not always be the most cost effective thing.

Rick: Yeah, no. I mean, yeah, absolutely. And we see, I mean, it's a good point because it's important to understand that when you created a table for DynamoDB that you should use meaningful values for those partition keys, right? I mean, one of the mistakes we see people make is that they'll import their data from their relational database. They're going to go ahead and use those auto incrementing primary keys from their tables as their partition keys and they don't realize that the application is never really accessing the data using that value, right?

It's always saying, "Hey, get me all the objects for this customer log in." Or, "Get me the things that are related to this." And customers don't usually call up a help desk for example and say, "Hey, I'm customer UUID." So if your primary access pattern is get the information by customer, then probably using a, UUID as a partition key is not a good idea and using more something like an email or a login name or something like that that the customer is going to know is better because if I store the data on the table, that's one cost.

If I have to index it, again, that's another cost. Every index I create is an additional chunk of storage, an additional capacity, write capacity use is consumed, so make sure that every time I write the data to the table, that there's a meaningful access pattern associated to it. Otherwise, I've got to create an index to be able to read the data and that's just some dead data. Now, that's definitely a write optimization, but the other thing to consider when you're optimizing for the write versus the read is not just the structure of the data, but also the velocity access pattern, right?

If I'm not necessarily reading at a high frequency, then maybe an inefficient read is fine and I don't need an index, right? I can actually just write all these items on the table and maybe once a week I need to find the exception items. Okay, I'll table scan it once a week, but if I maintain an index or I have a nice efficient read, then I can make that nice efficient query once a week, but I'm going to pay every single time somebody updates to the table, I'm going to pay to maintain the index. Right?

I'm going to pay the constant storage cost of all those items on the index. So sometimes, the very inefficient read is actually the most cost effective read because it's the ongoing cost of maintaining the index. So I understand that there's a break point, right? When you're trying to find an index and that break point usually has a lot to do with the velocity, the access pattern, right? If it's not frequently accessed or used access pattern, then maybe optimizing for that read is not a good idea and instead, we should optimize for the cost and cost of the write.

Jeremy: Yeah, and I love that because this is one of those things that I think most people get wrong is just this idea of saying, "I need to be able to access this data in a number of different ways." And you may need to do, you have to understand when and why, how, how often you need to be able to do that.

So anyway, there are certainly cases where you can't just only do that, right? So we talked about GSIs briefly. So I do want to get into GSIs a little bit more and essentially, a GSI, you can basically choose any two attributes, right? And make one the PK and make one the SK.

Sometimes we see people do like we mentioned the index overloading. Well, we can kind of get back into that, but I want to start and I want to kind of reference your talk that you, you did at the last re:Invent 2019 and I've been looking at your table designs.

I posted something on Twitter. I've printed them out, I highlight or I make notes, totally geeking out over it because it really is fascinating to me and one of the progressions I've seen over the last couple of years, and I think this is maybe a change in your thinking too is that you started ... Rather than trying to reuse attributes, maybe do things like inverted indexes and some of that stuff, you've actually created new attributes that are specifically labeled GSI1PK or GSI2PK.

And you mentioned this earlier, you're copying data. So you're denormalizing data, not just across the table, but actually in the same records...

Rick: In the same item. Yeah, yeah, exactly.

Jeremy: ... Item in order to do that. So just what's the ...

Rick: So yeah, a lot of times we will do this because maybe I'm going to use ... Well, I actually want to create an index on the sort key dimension, right? But there's a lot items on the table and not all of those items on the table do I want to have that sore key index, right?

So simply flipping the PK and the SK is going to ... Will cause every single item on the table to translate to the GSI and so I think you're exactly correct. It's been a kind of an evolution in thinking. It's also an evolution in the complexity of the application services that we've been working with because that's becoming much more of a common case now that we've got lots and lots of types of partitions on these tables, some of which we want to have reverse indexed and others we don't.

And so in those cases, what we'll do is we'll take that sort key and we'll copy it to another attribute on the table and then we're going to create a GSIPK on that extended attribute and that kind of filters out all the items that we don't want if I just took the sort key and the partition key and flipped them around and so in some cases, it can save a significant amount of money to the customer.

Jeremy: Yeah, because you mentioned that if you just do the inverted index and you flip the PK and the SK for a GSI1 for another GSI. The problem is that every single SK value then gets indexed and then all of the attribute data depending on projections, we'll talk about that in a second. But on those all get ... So by not doing that, and this is something that I think is extremely powerful, sort of I guess technique maybe, but is this idea of sparse index, right?

Rick: Exactly.

Jeremy: Where yeah, where you just take a little bit of you take the data and you put it into the field and if there's ... Or in an attribute, and if there is no data in that attribute, whether it's the PK or the SK, it doesn't get copied over.

Rick: That's correct. Yeah. I mean, if the attribute exists on the item, then the item gets copied. If the item ... The attribute date does not exist on the item, then it just sits on the table and never goes to the GSI. So this is the technique that we'll use like you said when I ... When creating the just the inverted index with the key flip is going to cause a problem and that ... And why might that cause a problem? Well, because some of those items are going to have things like maybe a date stamp hash UUID and I ... That's a useless value on a...

Jeremy: You're never going to be able to use that...

Rick: So why do I want to pay to store it and replicate it and do all that? Right. So, and we've got a lot more of those complex use cases these days that are requiring that. So that pattern has always been around, you spend a lot less of a common case. Now it's becoming more of a common case. So we talk more to it.

Jeremy: Yeah, and I think the other thing that I sort of noticed about doing this is you had mentioned, you've always said this for quite some time that NoSQL doesn't mean flexible, right? It's not a flexible data model, but what you did say and your talk at re:Invent 2019 was that if you're using this, and I didn't quite reference it the same way or maybe I was just reading through the lines or I'm just too much of a geek and I picked up on it, but this idea of there is some flexibility here, especially if you do this copying of attributes within the same item.

You can go back and run some sort of batch query, redecorate the items, move some things around and if you're not relying on existing attributes in a sense, there is actually some flexibility there, right?

Rick: Oh, there absolutely is. I mean, we've now had to go back and extend, revisit, retool hundreds of of services that were built as part of the migration to Amazon as you can imagine. I mean, applications don't just stay static forever. So in many cases like I said, it's just simply add a new partition type, maybe add a new GSI or overload an existing GSI in a different way, right?

Rick: If I'm adding new types of items, I get to reuse all my existing GSIs again, right? Because those items need to be indexed and I have, let's say three indexes on the table. Every new item type I add can be indexed up to three times without having to create a new additional index.

Jeremy: Yeah.

Rick: So it gives you a lot of flexibility. If I need to regroup the existing items and the existing attributes need to maybe refine those access patterns somewhat, I can like you said, do a table scan, let's do a slight ETL, maybe we annotate those existing keys with an additional composite, or we change the hierarchy of the composite or something.

But yeah, it's not impossible. It's just simply a process and honestly, as I've gone more and more through this process, I don't see how it's different than relational. I just don't. We went through all this with relational databases, right? I got to add a new column. Oh boy, what's ... It can't have a default value, but it has to have a default value. Oh my gosh. It's a terabyte table. And so it will be three weeks while the database is offline, right?

Jeremy: Or adding an index. Even adding an index. It's like I'm going to go take a two week vacation on large tables and then come back.

Rick: Exactly. These are things that ... I mean it's just kind of like ... Look, altering the data model stinks. It's always been a pain. I don't know if it's any more of a pain in NoSQL. Now that I've done this dozens if not hundreds of times, I don't think it's any harder. I think the one thing that might be different is that the ETL portion of the process is more up to you as a developer in a relational database when I add an index and things like this, it just kind of happens in the background.

If I want to add a column with a default value, I don't have to worry about doing a table scan and write back, but I mean, other than that, it's not any real difference. So one thing I could say is that at least with NoSQL, when you're doing these types of things, the data is online [crosstalk] possible, right? Each one of these updates is kind of atomic in its own envelope whereas that relational database, you have that index, it's like come back tomorrow, right?

Rick: This is the kind of thing where ... And that's the other advantage of the cloud date of service, like DynamoDB is the ability to add a GSI and I want that GSI be available quickly so I'm going to give it a million WCUs for an hour or so that it will very quickly suck the data off of the table, populate itself and come online as fast as possible, right?

We can't do that in a relational database because you're stealing IOPS, you're stealing throughput and bandwidth from your workload when you add that index, right?

Jeremy: All of a sudden, every query slows down and you're starting to wonder what's happening ...

Rick: "What's going on?" Exactly. Exactly.

Jeremy: So another thing that has to do, I think, and this is important to me I think is cost optimization and that's why thinking about sort of these right patterns and optimizing for that, there's another way if you are copying data into a GSI that you can optimize what gets copied over and that's projecting data into those indexes.

And you can just copy the keys, you can copy some select keys or you can project all of that and so obviously, it depends on the pattern but just let's talk about projections for a second.

Rick: Cool. So projections are one of the most powerful features of DynamoDB. It's the ability as you said to restrict the data that hits the index, right? So in like document databases, like DocumentDB or MongoDB, when I create a compound index, I'm able to project additional attributes out of the index, but for each attribute I add to the index, it increases the complexity of the insert, right? Because those are essentially sorting dimensions, right? So if I say I want these three attributes on the index, I get to sort by those three dimensions every single insert that I make.

So that becomes a significant overhead on the system whereas DynamoDB, what you're going to do is you're going to specify those two keys, the partition key, the sort key, we don't store it on any other attributes, but you can choose what other attributes you want to project into that write and we'll go ahead and just take those attributes along, store them on the index.

Now, when I query the index, I don't have to ... I don't just get the items that matched, I get the items that matched plus the data that I projected. That saves me a round trip back to the table to go get the items, get the data from those items. So like if you look at MongoDB or DocumentDB, when you look at a query and you explain the query, it's going to tell you two numbers. It will say documents returned and documents scanned.

If documents returned is X and document scanned is zero, that means that the query was covered by the index. It means every attribute that I requested existed on the index definition. That is extremely rare. It almost never happens, right?

Jeremy: Right.

Rick: Usually what I'm doing when I query a document database it's saying, "Get me those documents." Or get me some subset of those documents. Get me something from those documents. So those are two things that happen in that situation. You're going to get a query explained back that says documents returned X, document scanned X. That means I found all the documents of the index and I get to go back to the collection and get all those documents and, "Oh, I have no choice, but I'm going to return the entire document because that's the only choice the document database gives you." Right?

Rick: So you're going to read the entire mess, right? So if you've got a bunch of 16 megabyte documents and what you're really requesting is just a couple of kilobytes of data, then you get to read all those 16 megabyte documents and pull that couple of kilobytes of data out and serve it up. So it's very inefficient way. One of the best things about DynamoDB is it just lets you choose which parts of the data that you need to tag along for that pattern. Maybe this item might be a couple of hundred kilobytes, but that pattern only needs five kilobytes of data. I'll only project the five kilobytes onto the index, right? So it's a lot more efficient.

Jeremy: Yeah, and I think that what you see is again, people just sort of copy over that index in there and I see this all the time. There's just like, "Oh, I'll just copy all the data over because it's just easier than I have it."

Rick: Right. Right. [crosstalk], project all.

Jeremy: Yeah. And I think that in some cases, maybe that makes sense, but I think in a lot of other cases, it doesn't make sense because again, every bit of data you're writing over, every attribute you copy is costing you and if you start going over that one kilobyte that write limit, then you're using more and more of these capacity units, these write capacity units to do that.

So the other bit of optimization around that though is that if you ... And again, this depends on the the access pattern, but if you were to write just the keys for example, maybe you have an infrequent access pattern. You have to think about and I'm sure you would agree with me on this, you'd have to think about it how often do you need all of that data? Is there a way that maybe you just need that and then it might be cheaper to go look that up on the primary index once you have [crosstalk]

Rick: Right. As a matter of fact, a great use case for this multiple service teams inside of Amazon look for exception cases, right? They don't expect exception cases and when they get those exception cases, they don't expect very many of them and so what we do is we maintain ... but they need to check quite frequently, right?

So in our order processing workflows for example, if orders are languishing, we've got to get those things out. I mean, we've got prime customers, they've got to get shipments out, if fulfillments and orders are languishing, we need to know and we need to know in minutes, right? So they're literally running these types of exception handlers every five to 10 minutes and all they're really doing is just looking for things that are in a state that is considered to be a languishing state or an air condition state that must be expedited.

Most of the time, those queries are coming back empty. Every now and then, they come back with a handful of items. So what are we really interested in? We're just interested in the keys, right? Oh, these items are an exception. Okay, go get those items and notify somebody and run the escalation workflow. But we don't have to store all those items on the table all the time because most of the time, we don't even get any items and when we do, it's not a big deal for us to go get them from the table. So save the throughput, right?

Jeremy: Absolutely. Yeah, and then one last thing on optimization here and again, I think ... I'm not trying to take money away from Amazon, I'm just trying to let people know that there's some good ways that you can think about this and one of those is the attribute name itself, right? The size of the attribute names, that affects the size of the item, right?

Rick: Yeah, sure. In all NoSQL databases, not just DynamoDB, right?

Jeremy: Well, of course, of course not, yeah. Of course.

Rick: As a matter of fact, one of my favorite stories is from my days at MongoDB, I was working with a university customer, I think 80% of the data they had in their system was attribute names and 64 terabytes data.

Jeremy: My goodness.

Rick: Yeah, so when you think about the impact of that, every item I write to the database carries a copy of those attribute name. So as a developer, I mean obviously, the simplest thing to do be do one, two, three, four, five or A, B, C, D, but that's obviously not very meaningful and you don't want to have to map all that, but if you use kind of meaningful abbreviations, right? Like Customer ID, CID, your developers are going to understand what CID means, these are going to be come kind of just parts of your vocabulary. Please do that. You'll save yourself a fortune in the long run, especially at scale and that's [crosstalk]...

Jeremy: Yeah, absolutely, absolutely. All right, so another thing that GSIs allow you to do, and this is ... You explained this brilliantly in your talks, but it allows you to actually do a lot of really cool relational modeling and there's things like z-index and there's ... You can do sort of graph or edge nodes and some of this other fancy stuff and there's a whole bunch of stuff on the Amazon website that kind of explains some of this. But I think one that is really, really interesting is this idea of adjacency lists. Could you just explain that quickly?

Rick: Yeah, sure. So now, adjacency lists had been around a long time, right? Adjacency list is really a graph. A graph is a nothing more than and it's at its core in many to many relationship, right? You have a bunch of nodes, nodes have relationships with each other, nodes can be related to many others and nodes can add many related to them, right? So what we're really doing when we build an adjacency list on a DynamoDB table is we're just creating a lot of partitions on the table that are related to each other, right?

Those partitions are kind of like nodes, right? We can have in the example we're talking about which was sales people and customers. We've got customer nodes and sales rep nodes and every time a customer buys something, our sales rep sales something to a customer, we can create a relationship between that customer and that sales rep and the way we do that is by sticking an item into the customer's partition that says, "Hey, I have an edge that points to this guy."

And so maybe the first item we stick inside of the customer's partition would be a customer item, right? It describes the customer. So we have, that ... The partition key is customer ID, the sort key is the customer ID and then we have some extended attributes that describe the customer name, login, email, all that stuff. Then let's say a sales rep comes along and sell something to the customer, we can put an item inside of that customer's partition that's sorted on the sales rep ID, right? And inside of that item, we can say, "Well, what did the customer sell on what date and how much?" Right?

Jeremy: Right.

Rick: And so that's how he's related to the customer, right? So you get the idea, what are we doing? We're building a graph, we're putting edges into the graph and these edges have properties to describe those relationships. So it's kind of the same thing as having two tables with a mapping table in between them, right? That many relationships. Now, inside of that structure, I've got a table with the edges inside of one partition.

So I have customer partitions, I've got sales rep partitions and the customer partition has all these edges, but the sales rep partition doesn't. So it's a directed graph, right?

Jeremy: Yeah.

Rick: It means that the customer knows what sales reps he's related to you, but the sales rep doesn't. So if I want to create an undirected graph where both sides know what they're connected to, then I'll create a GSI and in the GSI, what I'll do is I'll take the sort key and the partition key and I'll just flip them around.

And now, all the items, when I query by sales rep, I'm going to see all those edges for all the customers that he's related to, right? Because all those edges were sorted on the sales rep ID inside of the customer partitions. When I flipped the partitions around now, they're going to be sorted on the customer ID inside of the sales rep partition and that gives me the ability to query the other side of the many to many relationships. So this is one of the things about data modeling in NoSQL, it's an interesting exercise because it demonstrates that I do not have to denormalize to be able to maintain these many to many relationships.

Jeremy: Right. Yes.

Rick: And that is a very important learning so to speak because a lot of people will go ahead and create two documents. One for the customer, one for the sales rep and inside of each document, they'll update the other with the order and the details and all of this and that's how they're going to track those many amenities. But that is problematic, right? I mean, there's a lot of work you have to do with the application layer. What happens if I update one-half of the relationship and not the other? What happens [crosstalk] dies in the middle. I mean, all kinds of crazy things can happen, right?

So using this type of construct, which is I just make an insert onto the table, there's a 100% SLA guarantee in DynamoDB that that GSI replication will absolutely occur. You can count on the fact that both of those relationships are going to get updated and that's, that's actually very, very powerful.

Jeremy: Yeah, and there is like I said, there is a lot of information out there that the best practices, DynamoDB best practices on the AWS site explained some of this stuff as well, bu it is very, very cool and I will say, I had modeled quite a few tables using that and trying to do these overloaded indexes and reuse attributes and SKs and things. And then as soon as I change to doing it with these separate GSI attributes and so forth, it actually makes it makes it so much easier. It makes it so much easier to do.

So if you're confused with DynamoDB modeling, really take this approach of those extra attributes for GSIs. All right, so one of the things that you have never mentioned or at least I don't think I've ever seen you mention it, at least not in any of your talks for your modeling is local secondary indexes.

And I used to think, "Hey, this is great. They've got really strong guarantees and then it's sort of this great use case if you want to do a couple of different sorts." But LSIs are not quite ...

Rick: Not the panacea you might think they are.

Jeremy: Yes, correct.

ON THE NEXT EPISODE, I CONTINUE MY CHAT WITH RICK HOULIHAN...

View Details

About Yan Cui:
Yan is an experienced engineer who has run production workload at scale in AWS for 10 years. He has been an architect and principal engineer with a variety of industries ranging from banking, e-commerce, sports streaming to mobile gaming. He has worked extensively with AWS Lambda in production, and has been helping clients around the world adopt AWS and serverless as an independent consultant. Yan is an AWS Serverless Hero and a regular speaker at user groups and conferences internationally. He is also the author of several serverless courses.

  • Twitter: @theburningmonk
  • Blog, Courses, Workshops: theburningmonk.com
  • GitHub: github.com/theburningmonk

Transcript:

Jeremy: Hi, everyone, I'm Jeremy Daly and you are listening to Serverless Chats. This week I'm chatting with Yan Cui. Hi, Yan. Thanks for joining me.

Yan: Hi, Jeremy. Thanks for having me.

Jeremy: So you are a developer advocate at Lumigo. You are an AWS serverless hero, you are also an independent consultant and I think more people know you as the Burning Monk. But why don't you tell us a little bit yourself and what you've been up to lately?

Yan: Yeah. I'm all those things you just mentioned. I'm doing some work with Lumigo as a developer advocate where I'm focusing a lot on the open-source tooling and articles and in my sole capacity as an independent consultant I also work with a lot of clients directly. A lot of them are based in London where I used to be based. Nowadays I moved to Amsterdam. I still do a lot of open-source work. I just started a new video course focusing on Lambda best practices. Then I'm also doing some workshops around the world. In Europe and also now looking at U.S as well. So doing a lot of different things to keep myself busy.

Jeremy: Awesome. Listen, I can talk to you probably about anything. Anybody who knows or seen some of the work that you've done, it's quite expansive. It's very impressive. In 2019, I have some numbers here, you did 70 blog posts, something like 2200 students to your video courses. You spoke at 31 conferences in 17 cities. But more importantly, you helped 23 clients in 11 different cities. So you are on the front line here in seeing how companies are adopting serverless. And not from one perspective and I think that's what we get a lot from different companies is, there is one perspective of how they adopt serverless and how they are working with that.

You've obviously seen this from multiple perspectives, so just, I want to talk about adoption a little bit. We'll talk about a few other things, but just what are you seeing with companies now? The customers you're working with or the clients you're working with, what are they using serverless for?

Yan: They are using it for all kinds of different things. I think depending on, I guess the maturity of the company, the domain they're working in. I've got a lot of clients that they are either enterprises or a lot of small and medium sized enterprises, and even some stealth-mode to startup as well. And obviously your constraints are completely different. That's one of the things I really enjoy about being a consultant. Where I get to see a lot different perspectives and what may work for one company may be completely inappropriate because the constraints a different company would have. So in terms of the adoption patterns, you see a lot of the, I guess startups that are in that position where they can go all in on serverless.

They are your great serverless-first going to the game. But then at the same time, you also have lots of, I guess midsize companies and enterprises. They have so much existing intellectual properties that it wouldn't make sense for them to rewrite everything just so that they can run code on Lambda. For all those companies, you see a mix of Greenfield projects. They are serverless first and then at the same time there's some effort to migrate some of the existing projects to work on serverless at least to some degree, at least gradually. Of course, depending on a lot of constraints around how much of on-premises stuff you have.

Do you have to run everything in Java in which case it is the cold start performance that's a concern. So a lot of those limitations I guess affects how quickly and how much you are able to go in on this whole serverless first mindset that we like to have and I think that is probably one of the reason that serverless adoption hasn't been as fast and as many people expected a few years ago because, the fact that, you can't just lift and shift your weight anymore. It means that you always have to allow more thought process behind it and planning and also just risk involved if you make a big mistake and it's your flagship product and of course that's going to put you in a really difficult position. But we do see that companies of all sizes and all fields and all industries are adopting serverless for lots of different workloads, not just APIs but a lot of data processing, IoT, you name it.

Jeremy: That's actually one of the things I'm curious of too. You mentioned customers in all different industries, which is really interesting. Because we get to this point now where I think every company is a software company. Everybody is building some sort of software now. But so, what are the constraints that these companies are working in?

Yan: A lot of them, I guess again, it depends on the industry you're in. For finance companies you have to be very careful about a lot of the, I guess, regulated requirements. In terms of how you handle data and also in some cases you having a plan in case you have to move away from a database for example, that's where some of your vendor lock-in arguments start to kick in. And also for example, you have enterprises who have millions lines of Java code that has been accumulated over 10 years. It's not possible for them to move everything into Lambda if they're seeing one to three seconds of cold start time on those user facing APIs. So some of those constraints are being lifted.

At least they are now getting better with new features on the platform but still it's something that people have to be aware of and also have to understand the mitigation strategies, which a lot of times is where the constraint is, is a lack of knowledge and knowhow because you can even think of Lambda as the extension to a lot of AWS offers, then it means that, you can't just know is it to visualization, you have to know a lot of different services to take full advantage of serverless.

That's where I think a lot of companies are struggling, is that they just don't have the skillset available in-house. They're exposing developers to things that they've never had to think about before. And I think that's where you get a lot benefit from serverless from having autonomous teams that can be self sufficient and look after so many different things, but at the same time, a lot of developers are just not used to working that way. They're used to working in silos where they have very few responsibility, just write your code, someone else will manage running the code in production. They'll manage the infrastructure, but now more of that is your responsibility which can be a gift, but it can be a challenge to companies that are not used to working that way as well.

Jeremy: Yeah, I totally agree. And I think that, as you mentioned learning all these other services, I think we're at a point now where most of these use cases there is some sort of serverless equivalent or serverless alternative to doing it in a more traditional way. Obviously, we're still missing certain things like, I'd love to have some sort of serverless Elasticsearch for example which would be really nice. Are there certain applications that you see people are trying with serverless or are thinking about serverless and just say, "No, I can't do it." Because the throughput needs to be higher or there is too much latency or something like that?

Yan: Yes. You see cases where in say for example, one of my clients had a very complex microservices environment whereby they have so many API to API calls and the fact that you get cold starts on one function that may not be an issue, it may not affect your 90% or whatever SOA is set. But when they start to stack up, that becomes a massive issue. So having more control around the warm up process, provisioned concurrency should help with those things. But at the same time, that is a slow process. Having to get the teams educated on what these different features are, how to work. In fact, a lot of questions I get are fairly simple questions around, how do I even do a CICD? How do I do testing? It's not clear to a lot of newcomers how do you do these things?

A lot of what we've been taught has been tethered to, there's going to be a server, I can just run everything locally and I press F5 and I can just run a local HTTP server and now everything is running in the cloud. A lot of that mindset change, it needs to happen. Those kind of paradigm shift happens gradually because well, everyone learns in a different pace and you need to have some critical mass in the industry.

Jeremy: Yeah. I like the idea of provisioned concurrency actually, because I do think it does solve a problem for the right types of applications, especially when there's low latency requirements, that it helps. I think that AWS has been pretty good about addressing those problems. They've come out now with the RDS proxy, which is helping with connections to the to relational databases. But I always feel like when that happens, they have to add another service in order for you to make it work. It's not just Lambda functions. "Hey, we've solved the connection issue with Lambda functions." It's, "We've solved the connection issue with Lambda functions because we've added a new service that now you have to use." And I think those present a number of roadblocks. And you had mentioned education as being one of those. So what are some of those other roadblocks that you see companies running into?

Yan: Well, the biggest one I feel is by far is just education. Like I said, Lambda itself is getting more and more complicated because of all the different things you can do with it. Other roadblocks includes for example some organizations are still holding onto the way they are used to operating. With centralized operation teams, cloud teams. The feature teams don't necessarily have the autonomy they need to take full advantage of all these different tools that you get and all these power and agility you get with serverless, your team can build a new feature in a week, but it's going to take them three weeks to get anything they need done, provisions and to get assets to resources they need. Then again, you're not going to get the full benefit of serverless.

So a lot of that legacy thinking at the organization is still there and is still a prominent problem and roadblock for people to take full advantage of serverless. But in terms of actual adoptions, a lot of it is ... In terms of technical roadblocks, there's some, I think the last question you had was around some use cases that just doesn't fit so well. When you've got a really high throughput system, the cost of serverless can become pretty high. So imagine you've got something that's relatively simple, but how to scale it massively like your Dropbox, not a super complex system, but have to scale to massive extent. So for them it makes perfect sense to move off of S3 and start to build their own given hardware so that they can start to optimize for that cost.

For a lot of companies, they do have that concern as well. They may not have a very complicated system that requires a hundred different functions on this massive event driven architecture, maybe they just have five end points. But those five end points are running at 10,000 or 50,000 requests per second. So in those cases, the cost for using Lambda and API gateway would be excruciating and you'd be much better off paying a team to look after your community's cluster or your containers cluster than have them running them on Lambda.

But that's always a tricky balance. Because, oftentimes you can always get the reverse argument whereby, "Well, Lambda is expensive, so I'm going to just do this myself." But then you're hiring someone for $10,000 a month

Jeremy: Exactly.

Yan: ... to look after your infrastructure, and your Lambda bill is going to be, I don't know, $100.

Jeremy: And you're hiring more than one person too. Then you're still paying for the bandwidth and some of these other things and you're still paying for compute somewhere. So that's really interesting. You made a point a little bit earlier where you said, this idea of the paradigm shift or the mind shift of going from this traditional lift and shift and bringing things into serverless. And so obviously there is a ton that needs to change. We'd like to say, it's programming and you just need to figure out the glue that works there. But you really can't just lift and shift and get those benefits, right?

Yan: Yeah. It's a common pitfall whereby the teams try to lift and shift. And initially it looks like it might work and then later on, pretty soon they found out the hard way that, that approach doesn't scale, it doesn't work nicely. And you run into all kinds of different limitations. For example, one example from a client I've worked with that can be used to illustrate that point was, they had this API which used to do lots of different things, including doing some service rendering and some API endpoint, some penalties, some requests, and then they just moved the whole thing as one big fat Lambda function.

Because one of the end points have to access some VPC resource, so of course [crosstalk 00:12:59] and now you've got, every function have to ... when cold starts have to initialize React, which is not a very lightweight dependency. Even when you've got HTTP endpoint it doesn't need it. You have to initialize it, and then also you have to wait for the VPC initialization and all of that, and they were getting performance that was so bad that it's just not acceptable for anything that's going to run in production. And of course, unless you know that the reason why that's happening is because, we've got this Fat Lambda and how the whole initial decision process works.

Then you know to split your function up into one function per endpoint perhaps. Maybe at least some separation so that the resource intensive functions are separate from the other things. I like to find that you've got tools that allow you to take your express app and just run inside Lambda. They represent the easy path for people to get some of the benefits in terms of the infrastructure automation and improve their scalability and resilience with Lambda. But at the same time, unless there's a way for you to then later on do the idiomatic way of working with Lambda or having single responsibility functions, then it becomes a bit locked into the decision that the tool has made for you and it becomes harder for you then to migrate later.

Jeremy: I actually think that's one of the better arguments for moving away from Fat Lambda or the Lambdalith. I think a lot of people have a ton of success with that. I have used them in the past as well. There's been times where it just seems to make sense, but certainly the bootstrap process. If you're bootstrapping something as part of the warmup phase of a Lambda function that isn't used by 90% of the code, it's only used for it. It's a complete waste of time and memory to boot those things up. So, I totally agree with that. I think you and I were talking in the past too where, we just said, this adoption pattern here, this is just something that is going to take time. I know you're a big fan of functional programming. I'm actually a big fan of JavaScript functional programming, which I think people think is impossible, but it is. But anyways it's something that is probably just going to take a little bit more time for people to understand working in this different way.

Yan: If you look at where we are with functional programming, it's as old as OO, but when functional programming is going to hit the mainstream guys, TBD to be decided which is to some extent is frustrating because, for many use cases functional programming is probably better tool. I'm a big fan of F sharp and done a lot of things in the past with [inaudible 00:15:46] and stuff as well. I'm a big fan. But there is a big mind shift, change the you problem solve and it's not. Mind change doesn't happen overnight, and have you have to be patient, and you have to give people time to digest and internalize this change and really understand the benefits before they become advocates themselves or at least they become practitioners.

I do think it is happening slowly. Just judging by the amount of inches that community is showing and the number of serverless conferences around the world, their interest is definitely there. But we still are a long way from having enough people who are well equipped to succeed. There are definitely a lot of people. You Michael Hart, your Ben K. But we need a lot more of them.

Jeremy: I agree One other thing on functional programming, I tell people, "Listen, once you write a pure function, you'll never go back to writing something else." Anyway, one more question on the adoption side of things. Because one of the things I see quite a bit, I really love this use case for serverless, is sort of this peripheral enhancing the DevOps sort of stuff. Do you see a lot of that where companies are using it to either do auditing or doing DevOps automation, things like that?

Yan: Yeah, tons. There actually have been quite a few companies who, their main feature teams are not using serverless, but their infrastructure and their DevOps teams are using Lambda very, very heavily. Whereby before, there's just a lot of things they couldn't do, because there's no way to tap into what's happening in your diverse ecosystem. But now with Lambda, everything that's happening in your account gets captured by CloudTrail, you can use. Or Eventpattern or Eventbridge or CloudWatch events to trigger your functions to do all kinds of automated detection for changes that you don't expect to happen, to security checks and things like that. Or even just basic things like automating, doing some processes and resources that are no longer necessary.

There's tons of things that a lot of DevOps teams are doing now that they would have been really difficult to do in the past without Lambda. And I do see a lot of adoption in that particular space as well.

Jeremy: Awesome. All right. So I want to move a little bit past use cases, but I think maybe this ties into it. There are people who say, "Well, I can't use Lambda because it only runs for 15 minutes and I have ETL tasks that need to run longer jobs." Or, "I needed to do something like that." Or, "I have to have multiple jobs running together." Or something. And this new thing that seems to maybe have been sparked by Google Cloud Run, is this idea of serverless containers. I spoke with Brett McGowan about this and just the thinking behind that. And so obviously we have Fargate with AWS. So what are your thoughts on this idea of expanding the term of serverless to include things like Fargate and Cloud Run?

Yan: Well, listen, when I think about serverless, I don't think about specific technologies. I think in terms of their characteristics a technology has, in terms of the pay-per use pricing model, in terms of not having to worry about underlying structure and how it's going to scale and all of that. I think right now, Fargate is serverless in that, you don't have to worry about underlying infrastructure. There is two instances that your containers run on or the cluster, how to auto-scale them. But I guess what is missing right now is just the event programming model and the fact that this is not pay-per.

Jeremy: Pay-per use.

Yan: Pay-per use, yeah. But that's that. I think you will too get a lot of benefits that we enjoy from serverless technologies with Fargate already and it does eliminate a lot of the limitations that Lambda has. Also, I just don't think that Lambda is going to be ... we should not see Lambda as the silver bullets. Nothing ever used is going to be a silver bullet. So the fact that you've got something else that can allow you to run containerized loads very easily and minimize the amount of work that you have to do. Because, remember the whole thing about serverless is about removing undifferentiated heavy lifting. And a lot of that is around managing EC2 instances, configuring auto scaling groups and clusters and all of that. And the fact that you can get a lot of that away from my plate onto AWS with Fargate, and I think that is really good direction.

I'm not a purist in terms of the terms, all I care about is, what can I get from a technology? And from that particular standpoint, Fargate is quite close to what we get with other similar services. It Just would be nice if you can trigger Fargate with event triggers directly.

Jeremy: That's the big thing too. I think Tim Wagner has said this as well, where it's sort of like Lambda and Fargate are becoming closer and closer. For all intents and purposes, there's no reason why Linda can only run for 15 minutes other than it's a limit that AWS set. I mean they could run for an hour or 10 days if they are needed to. If they wanted to allow you to do that, they could add some sort of event triggering or some sort of event driven approach to Fargate. I mean you can start Fargate tasks now in a number of different ways. So there's a little bit of event driven, just not as clear as the Lambda stuff. As Lambda gets more of these server full type features and as Fargate gets more of these serverless features, is there maybe a point that they become the same thing?

Yan: Probably and hopefully. I think at that point, it'd be really confusing for people. But I think that that is ultimately where I hope we will get to. Whereby a a lot of the limitations that we currently have with Lambda is eliminated and a lot of benefits that we enjoy from Lambda but not available for as Fargate becomes available for Fargate. So it becomes more of a a choice in terms of, "Okay, what do I prefer working with? Do I have specific use cases that fits better with a containerized environment where I have more control of the infrastructure itself?" Then I use Fargate versus using Lambda. But in both cases, I can enjoy pretty much the same benefits. I think that would be a really good place to be.

Jeremy: Awesome. All right. So one of the things that I've been talking about a lot at the end of last year and it's something that I've been thinking about for awhile, is this thing that I call abstraction as a service. And it's probably an annoying term, but what I'm thinking of is, Lambda functions themselves are pretty easy. You created a Hello World one, fine, simple. You want to add an API gateway, you use the Serverless Framework or use SAM, it makes it very easy for you to get these simple examples up and running. But start adding in SQS queues, or EventBridge and Kinesis Streams and then understanding the sharding of Kinesis streams and how many concurrent Lambdas that you might need to have, and then the new ability for you to replicate the stream.

There's just a whole bunch of things that are happening there. And now suddenly your simple serverless upload a piece of code that is now completely dwarfed by the amount of configuration files you have to write and the understanding of all these different best practices. My sort of premise here or what I'm hoping to see, and I think this is something that serverless components are starting to do. And to a degree, the serverless application repository is starting to do is encapsulate these use cases or these outcomes and put them into something that's much more easily deployed.

Where you don't have to think about the 50 different components you might need to configure under the hood. You just say, "I want to build an API or a web hook that does this and that and whatever." And it's much easier to configure that with same default. So, we can talk about the serverless components thing. But really what I want to do is focus on the serverless application repository. Because you've done a bunch of apps. I think you've got 10 of them in there now. What are your thoughts on SAR?

Yan: I think SAR is a good idea. But the execution is still problematic right now. At least for my experience working with SAR both as a consumer and also as a publisher. So one of the things said that often stands out is that, with SAR, the integration path is not super clear. For example, as a consumer, to use SAR in my CloudFormation template is not just normal CloudFormation resource, this host servers application. It's not a native CloudFormation resource type, so you have to bring in the SAM Macro even if you're not using SAM. A few times when I had to do that with the server framework, it was just fine. I can bring in the SAR Macro, but it becomes a bit weird. And also AWS often talk about this idea of we should be doing lease privilege as a default, but then they want you to also just use a package, their profile, it's policies for your SAR applications.

Which means that, your application either have no enough permissions or have too much permissions. It's really hard to size and tailor your permissions to follow this privilege. But when you do the right thing, the discoveries in the console punishes you because someone had to take a box to find applications that are using custom IEMs which they're trying to do the right thing to give you lease privilege. Also I find a lot of the discoverability itself is also not that intuitive to use. When you trying to search something, it's giving you way more things than you're actually looking for.

If you look at some of the top applications in SAR right now, they're all Hello World or introduction to basic Alexa Skills example. There is a lot of example codes you can deploy to your account to have a look at how someone else is doing Alexa Skills as opposed to something that is actually truly useful. What that tells me is that, the AWS customer just don't really know what they can do with SAR.

Jeremy: Do you think it's a lack of incentive for people to publish those apps?

Yan: Part of that is that and Forrest wrote a blog post a while back where he argued that SAR being a marketplace of sorts, should be incentivizing companies and publishers in terms of putting out something that is not just a toy or example Codebase, but something that as a company, as an enterprise, I can actually have confidence. The point is being into my real production environment and know that it's been looked after. When there's issues, someone would actually be there to fix it and patch it rather than leaving me in a ditch. Because all I need is one experience like that to never want to touch anything in SAR ever again. So having some kind of a scheme where publishers can be financially rewarded by the resources that I will provision into my account, so AWS bill me for those resources so some of that learning can then be passed onto the publisher for the SAR application.

That way you hopefully would encourage more commercial companies to start to publish things that are commercially looked after, adhere to SOAs and guarantees your large enterprise customers will be willing and comfortable to actually deploy into their environment. The same way that when you look at AWS marketplace, where I'm buying some software that deploys EC2 instance, at least I have confidence that this is a commercial thing. It's not just someone's toy project that they may not look after when they find something more interesting to work on.

There's a bit of an image problem there for SAR in terms of what does it represent to the consumers. And if we want people to have faith in that, then we really need to do something about that. I think commercialization is one step towards that.

Jeremy: I wonder about that too, because I read Forrest Brazeal's post as well, and I thought that made a lot of sense. You have other open marketplaces or other open ecosystems. Just think about NPM for example. People use NPM packages all the time with probably no consideration as to how well some of them are maintained. So you already have people using those and running into certain problems like that. I think maybe because it's so specific to AWS and maybe it just doesn't seem as open source as something like NPM does in a sense. But I totally agree. I just don't know. Do people pay for some of these apps or are they paying more for the support of them?

Yan: I think that's an interesting point about NPM. But what I will say is that, the impact that a badly written SAR app can have on your organization is probably far greater than an NPM package. Because now you're talking about resources that are provision to AWS environment where they can ... If a malicious access for example, might be able to gain access to way more things, than say someone who's published a malicious NPM package of course can do a lot worse. We fear those dependencies too. Also a bad [inaudible 00:29:43] application can also just cost you a lot of money. Imagine you have someone deployed something to VPC with net gateway and start charging you 4 cents per gigabyte of data transferred, and then those-

Jeremy: Get expensive quickly.

Yan: ... can get very expensive really quickly. I think in terms of the impact it can be much greater. I will think twice about deploying something to SAR, whereas with NPM it's often just okay.

Jeremy: Maybe it was one of those things too, because I think you're right. You're deploying something that is actually going to cost you money directly. So you have some other costs of auditing and some of those things you might do with NPM packages, but certainly with this, you're deploying things into an account that could rack up serious bills. That might be one of those other things where SAR needs to go down this road of helping people understand exactly what types of resources they're provisioning and maybe cost estimates and things like that that could potentially help ease someone's mind. But I do agree. There needs to be more people flooding that marketplace with good tools that they can use, and without having some sort of backing I think that's kind of tough to achieve.

So speaking of other tools and other things that are available, the ecosystem that we have now for serverless frameworks and not serverless framework, but frameworks for serverless, I should say. Serverless framework being one of those, obviously SAM architect, Claudia.js. There is a lot of them now. There's ones for PHP, there's ones for Ruby on Rails, there's all kinds of these frameworks popping up. Pretty much every single one of them is doing the exact same thing.

It's taking some level of abstraction and compiling it down to CloudFormation or making a bunch of API calls to AWS. What are your thoughts? I know you're a big proponent of the serverless framework. You've done a ton of Serverless Framework plugins, and I know that you've done a lot of work with SAM as well. So just what are your thoughts on the overall ecosystem? What should people be using?

Yan: Personally I prefer Serverless Framework and I'm happy to go into details on why I think Serverless Framework does well compared to a lot of the other frameworks. I think that the Serverless Framework the biggest strength it has is the fact that it's got a great ecosystem of plugins that have support from the community. Pretty much anything that you run into, there's probably a plugin that can solve that problem for you or at least make your life a lot easier. Even when that's not a case, it's really easy for you to write a plugin yourself. I guess I'll complain about their documentation on how to write a plugin. I think the only two articles they have there is still from Anna from I think three years ago.

But once you learn what a plugin looks at like, it's fairly straightforward because you can do so much different things. You can make API course as part of the deployment cycle. You can transform the CloudFormation template. With SAM, it does a lot of things right out of the box, but the problem I have with SAM is that, when you don't agree with the decisions that SAM has taken, it's really hard for you to do anything about that. One time I was working with a client and we were using SAM and that's when SAM just introduced the IAM authentication for API gateway. But they were also changing how the IEM permission was passed through. So as a caller, I need to have the permission to evoke the function as well as the end point, which of course didn't make sense, it breaks obstructions and all of that, but there's no way for me to get out of that.

The only way I found was, I actually wrote a CloudFormation macro, deploy that and then change the template that SAM generate just to fit to that one tiny little thing. This is where having that flexibility gives you default like everybody else who is trying to do as well. But at the same time, give you a nice way out. I guess when it comes to framework, there's also this new CDK and [inaudible 00:33:59] which is a whole different paradigm where this lets you program with your favorite programming language. I have to say I'm not a fan of this approach. I think I can get the temptation that, "Oh, I like writing stuff in C Sharp, I like writing stuff in Java script and now I can use my favorite language to do everything."

But your preference to the language that you want to write, I don't think that should be very high in the list of criteria for choosing a good deployment framework. Things like, the fact that you can get the same infrastructure you have to reason in different ways, I think that is quite a dangerous thing. You can end up with arbitrary. The complex things that would have been a lot simpler if everyone just writes something JSON or YAML. That said, I do wish there's better tools for YAML. I see so many people struggle with just basic indentation problems. It happens all the time.

For me, I came from F Sharp and Python as well. [inaudible 00:35:07] methods. I kind of learned that, but most of what haven't. You have to be trained to look out for these kind of problems, and we do need better tooling support for YAML. That said, I still think YAML or something like that is a better way to define your resource graph compared to writing a code to do that. I remember before all these frameworks, I was writing bash scripts to provision everything and now I'm just substituting bash with C Sharp or a prettier looking language. And I don't think that's the right approach.

Jeremy: I actually agree with you on the CDK stuff. I know some people are huge fans of it and they like the idea that you can build these constructs and then you can build higher level constructs that wrap a bunch of things together. And it is kind of cool that you can encapsulate some of that stuff, but I do feel like there is that black box issue there, and maybe with Winston Churchill who said that, democracy is the worst form of government except for every other form of government or something like that. I would probably say YAML is the worst form of configuration language except for every other form or every other configuration language.

All right, so the serverless framework, they just came out with the Serverless Framework Pro. And I know you've kind of experimented a little bit with that, but what are your thoughts on that? Now that they've added things like monitoring and CICD and some of that other stuff?

Yan: I think it's a nice tool for someone who's new to serverless and just wants to have something that they can use. But it's certainly from the monitoring perspective, I don't think the Serverless Pro holds up compared to other more specialized solutions that offers monitoring and tracing that you can tell are done by people who are spinning this view for a very long time and understand this problem space. What I find with the Serverless Pro offering is that, it gives you a lot of basics, but it doesn't do much more beyond what you get with CloudWatch.

So as someone who's got a lot of experience with AWS and have used CloudWatch for many, many years, I don't see a lot of value add for me to invest into Serverless Pro. But at the same time, if I'm new to AWS and new to serverless having something that comes out of the box with the two that you need to use for deployment, I can definitely understand the temptation there. For a lot of applications, I've done where it gets complicated quickly. You've got some of the functions, lots of event triggers, lots of events flying everywhere. And I'm really interested in the tracing side of things and that's why I think a lot of tools that we have today, it's not quite there yet.

Everything seem to struggle for tracing. EventBridge for a moment and also X-Ray for example, it doesn't trace through SQS properly, it doesn't trace through Kinesis at all. We get all these fragments of our transaction, but you can see that this space has been evolving really, really quickly. You've looked at the work that Lumigo has done, AppScan has done and Thundra has done. Everyone has gone through a lot improvement over the last 12 months at least. And I do see this space are getting more mature and more of the, I guess traditional big monitoring companies getting into this space as well.

And also a shout out to Honeycomb as well. I think they also do a very good job with their product. It's quite a big mind shift for people who are not used to do event based monitoring. But once you have that, it's really powerful. Splunk has been there for a long time, but they kind of price everybody out.

Jeremy: Listen, I think you're right about the monitoring component of Serverless Pro. It is good. I played around with it and it does tell you about your invocations and things like that, but this idea of really understanding the tracing and some of that deeper stuff is a little bit more advanced. But I will say, Serverless Framework has been great at developer tooling. That's one thing that they've done really well. And I think the greatest feature of Serverless Pro, at least for me is the new CICD deployment stuff that they've released. They've got similar to what Seed.run did with being able to use the monorepo.

It's very hard to have multiple repos when you are building serverless apps especially if your services are relatively small. Sometimes that monorepo make sense and being able to just deploy changes from individual directories I think is a pretty interesting thing. But anyways. All right. What about your thoughts on this? Because this is another thing we hear all the time and it kind of drives me nuts as when we hear the term multi-cloud. And that people are trying to ... You actually mentioned it earlier where you were hedging your bets to say how easy is it for me to move from AWS to some other provider as that's something that we actually care about. Do you see using a framework like SAM as potentially locking you in even more to AWS or do you think that's a pointless argument?

Yan: I think it's more of a pointless argument. Even the tools that do support multi-cloud, they have different syntax in the same way that the Terraform have got different syntax for different cloud providers but give you a consistent tooling experience when you use it with different cloud providers, but you still have to learn the different syntax. You have to learn the cloud itself. What resources is available in AWS versus what resources is available in GCP or your [inaudible 00:40:55]. Serverless framework, it does support multiple clouds but at the same time I think it's not as valuable as the people probably make it out to be. Because again, how you work with different clouds is completely different syntax and different resource type.

It'd give you consistent tooling experience in terms of using SOS deploy if something happens, but it doesn't remove the fact that that's no way you're going to struggle with when you want to go from one car to another. There's been so many different blog posts on this, I've written a few of myself as well. And I think this whole multi-cloud thing is an argument about, for example, when I buy an iPhone, they take our insurance, but how much I pay for insurance versus the cost of the phone itself. If the insurance itself is going to cost way more than just getting a new phone, then why we're not wise to do that but at the same time you look at some of the vendor lock in arguments. Well, when you decide firstly it's not lock-in, you can still move things, just there's a cost to moving. They're coupling, so there's a cost of moving.

You either deal with that when that scenario comes up or you try to do a lot of work upfront. So essentially investing all the work that you will have to do later to this point when you don't even know what's going to happen in the future. And the worst thing is, you end up with a lot of complexity that you have to carry all the way and everything becomes more difficult if it becomes slower. Your developers have to work so much harder to do everything as opposed to just, make a decision, go with it and knowing the back of her head that if we need to move ever, this is other things that we need to think about and we need to do. I think Nick from Serverless Inc actually wrote a really good post about the fact that, moving compute is always easy. It's the data. Data is incentivized to stay where it is and accumulate as much as possible. So it doesn't matter how easy it is to move your APIs from one container to another in different clouds. Well, we're going to deal with the data because there's no exact replica of that in Dynamo DB.

There's no exact replica of the high replication data store. I think it's foolish to spend so much effort upfront to prevent something that is probably unlikely to happen. How much I spend on insurance should be proportional to the risk of my phone getting stolen, lost and also to the cost of the phone yourself as well. The same argument applies here, where a lot of this strategy is just insurance against the stolen phone.

Jeremy: I think I need to start hooking my guests up to a blood pressure monitor when I ask them the question about blocking. Actually it's very funny. I think people now, and I'm the same way when somebody asked me about this. I think I maybe do it just to get a rise out of people, but people get angry now about trying to defend this vendor lock-in thing. Because I think you're absolutely right. And the biggest concern that I have where people play this vendor lock-in argument. You're locked into everything. You're locked into your iPhone, you're locked into Microsoft Word or whatever if that's what you choose to devout your time.

Yan: Anything you use.

Jeremy: You're locked into these things. I look at it and I say, if people are using that as an argument to use the tools or picking the technology that's the lowest common denominator, then they're not choosing the best tool for the job. I think that is something that significantly impacts the ability for people to adopt serverless because they say, "Well, if I write a Lambda function, I can't just easily move that to Azure. I can't easily do that to GCP. I need to work with all of those constraints." But honestly, I think if anything, you're just adding more work for yourself. And you're right, you're insuring yourself against something that is very unlikely to happen. And in the off chance that it does happen, I still think you're going to go through a massive exercise in order to migrate something no matter how low the denominator was that you chose.

Yan: We went through all without with ORMs. There was a few years where there was ORM every single month.

Jeremy: I'm going on the record. I hate ORMs. I hate them.

Yan: Because when we do have to move to a database, it turns out ORM doesn't really help me. It's just another thing you got to deal with as part of the migration process.

Jeremy: Something new to learn.

Yan: And also it gets in your way from the start in terms of the complexity to start but also when you want to do anything, you've got to have to understand what happens under the hood but then also how to do it with ORM.

Jeremy: Exactly.

Yan: It's crazy.

Jeremy: And the optimization isn't there. You don't get the optimizations with an ORM. The biggest thing that drives me nuts about them is that, you write a query and then you have some ORM or whatever that has to run three separate queries in order to merge the data back in the application layer because that's how it was built, where you could have just written a native query and done a join or something like that, and it would have been a thousand times more efficient. But anyways, yeah. You mentioned Terraform as well when you were talking. What are your thoughts on Terraform for serverless deployments?

Yan: It's very laborious and painful. I remember on my previous jobs I was convinced by the teams to use a Serverless Framework, and all I had to do was show them a very simple API gateway endpoints with a Lambda function and there was three lines of code in the server framework. It was about 150 lines of Terraform scripts. And you can see the teams that are using the serverless framework, they just go in there and get it done. Get a feature shipped and test it and all that. Other teams would be spending next two weeks just writing Terraform script. I had engineers coming up to me to describe their job. We spend about 60% of our time just writing Terraform. When you are talking about serverless being don't do undifferentiated heavy lifting. Something is not quite right when most of your time you're just writing infrastructure.

Jeremy: That's the thing too with Terraform. Terraform is a very good product and there's all [crosstalk 00:47:28] Terraform the enterprise edition has a lot of great things like safeguards and some of those other things. I think it is a very good tool for cloud management but at the same time, I think you're right, not very productive for the serverless developer.

Yan: No. If I'm provisioning VPCs and networking and things like that, I'm very happy to use Terraform. It is a very good tool for that. But when I just wanted to write a few Lambda functions and hook up a few end points and have some event stores like SNS, SQS and so on, I really don't need Terraform, what I need is something that can give me good defaults and allow me to do what I need to do and get out and move on to the next thing rather than having to get bogged down with the detail of the specifics. That's just not productive. That's not useful. That's just undifferentiated heavy lifting.

Jeremy: You are preaching to the choir. All right, let's move on to, where is this going? Serverless in general. This is one of those things where I think you and I would agree that, and I think you mentioned it earlier, it's like we're making it more complicated. We're adding new features, the learning curve keeps getting steeper and steeper. There are still some use cases that are not necessarily perfect for it. AWS is making advancements in some of those things. Reducing VPC cold starts, adding things like RDS Proxy and provision concurrency and those sort of things. But are there other things that are holding serverless back? Does there need to be some other breakthrough before it goes mainstream?

Yan: I don't know about the major breakthrough, but I definitely think more education and more guidance, not just in terms of what these features do, but also when to use them and how to choose between different event triggers. That's a question I get all the time. "How do I decide when to use API gateway versus AOB? How do I choose between SNS, SQS, Kinesis, DynamoDB Streams, EventBridge, IoT Core. That's just six application integration services off the top of my head. There's just no guidance around any of that stuff and it's really difficult for someone new coming into this space to understand all the ins and outs and trade offs between SNS and SQS and Kinesis and so on.

Having more education around that, having more official guidance from AWS around that, that would be really useful. In terms of technology wise, I think I like the trajectory that AWS has been on. No flashy new things but rather continuously solving those day to day annoyances, the day to day problems that people run into. The whole cold start thing, again, often overplayed, often underplayed it's never as good as some people say, it's never as bad as some other people say. But having some solutions for people with real problems, where with clold starts we speak of various different reasons.

I really like what you've done with provision concurrency, even if I think the implementation is still, I guess it's a version one. So hopefully some of the kinks that they currently have would be solved. Other than that, I'd like to see them do more with some multi account management side of things. A control tower is great, but again, there's a lot of clicking stuff in the console to get anything set up, and it's also very easy to rack up a pretty big bill if you're not careful you can provision a lot.

NAT gateway for example and things like that. One of the companies I've been talking to recently as well, a Dutch bank, they are actually doing some really cool tool themselves to essentially give you infrastructure as codes. Think of it as a CloudFormation extension that allows you to capture your entire org. Imagine I have a resource type that's defines my org and the different accounts and then when they configure CloudTrail set up for multi-cloud to configure security guard and things like that all within my cell template, which looks just like CloudFormation. So some really amazing tool that those guys have built.

But having something like that from AWS would be pretty amazing as well. Because again, we've seen more and more people getting to the point where they have a very complex ecosystem of lots of different enterprise accounts, managing them and setting up the right things. The STPs and things like that. It's not easy and we certainly don't want people to be constantly going to the console and clicking things. And that's another annoyance I constantly have with AWS documentations is, they keep talking about infrastructure as codes, but every single documentation just tell us, go to this console, click this button.

Jeremy: That's how you do it in the console. Exactly.

Yan: What the hell?

Jeremy: Yeah, exactly. I guess one of the things that I try to tell people who ask me to get into the cloud or to start building stuff in serverless is sort to do a slow migration pattern. You can't just jump all in, you can't rewrite everything in serverless and do that. Often though that does require rewriting applications. Do you see a potential path where making it easier to move those applications into Lambda or into Fargate maybe like if there was an easier path to lift and shift, would that be something you think would make sense?

Yan: I think that would make sense. I guess I'll have to wait and see what kind of execution that comes from that. Because again, you're making a lot of assumptions about what people are using, what they're doing to be able to do that well. Of course if you do that, it's really easy to do them bad. I kind of think that would be great, but it really depends on the execution.

Jeremy: Awesome. All right, so any other missing pieces in serverless? I think you and I agree we need some sort of Elasticsearch utility.

Yan: Absolutely.

Jeremy: But anything else you can think of that's maybe missing?

Yan: Let's see. Nothing off the top of my head. But definitely some kind of serverless Elasticsearch that would be awesome.

Jeremy: Awesome. All right. So final question here. Because now that I have you and I think that with everything that you write with the courses that you do and you're doing a ton of in-person workshops and things like that and all of your talks, everything you do is very, very good advice. And I think you've been a serverless hero for quite some time. So just maybe we can capture, if people are interested in moving to serverless, what is your one sentence or it can be a little longer. How would you suggest people make that first step into serverless?

Yan: Subscribe to this newsletter that I heard it's because something like Off-by-none. It's a really good way to just get regular newsletters about all kinds of different content.

Jeremy: I did not pay you to say that. I just want to make sure that's clear.

Yan: But yeah, definitely. That's one of I think one of the dangers of having Lambda being deceptively simple is that, there's still a lot of things you have to learn. There's still a lot of things you have to understand too. You can make really bad mistakes. We keep reading on the web about horror stories, but a lot of that is because of the lack of research, and I think Joe Emerson said it really well that, if you spend two weeks researching and two days doing work, you're probably going to end up better off than if you do two days of research and two weeks of work.

Jeremy: Yes, I totally agree.

Yan: So that you don't make all these mistakes. But in terms of actual advice, I think we share to people in the community. People like you, me, Ben Kehoe, or others, we are all very happy to help and do some research and if you're stuck, just ask us questions. We're all very keen to see a world where human productivity is not wasted on setting up servers and managing them. We'd be very happy to help you. So we'll help you get started the right way.

Jeremy: Awesome. All right. Well, thank you again so much for joining me and sharing all of this serverless knowledge with everyone in the community and obviously the things that you continuously do to help people learn and educate people on serverless. If people want to find out more about you, how would they do that?

Yan: They can go to theburningmonk.com or follow me on Twitter as @theburningmonk.

Jeremy: And you've got a bunch of courses and open-source projects that you work on. Those are all available on theburningmonk.com.

Yan: Yeah, yeah. A bunch of courses you can find under the courses heading. There's also a bunch of in-person workshops I'm doing this year and also just lots and lots of blog posts.

Jeremy: Awesome. All right, well, we will get all that into the show notes. Thanks again.

Yan: Thank you. Thanks for having me.

View Details

About Ken Collins:

Ken is a Staff Engineer at Custom Ink focusing on DevOps & eCommerce architectures with an emphasis on emerging opportunities. Custom Ink is approaching its 20th year in business and is entering its second phase of Cloud adoption where he helps a growing engineering team succeed using AWS-first well-architected patterns. Ken lives near Norfolk, VA and organizes the area’s Ruby User Group.

  • Twitter: @metaskills
  • Custom Ink Tech on Twitter: @CustomInkTech
  • Blog: technology.customink.com
  • Lamby: lamby.custominktech.com
  • Full Stack to Functions and Back Again Talk:
    • Slides: https://speakerdeck.com/metaskills/full-stack-to-functions-and-back-again?slide=2
    • Video: https://www.youtube.com/watch?v=ktDXVn3EPfY
  • Migrate Your Rails App from Heroku to AWS Lambda: https://technology.customink.com/blog/2020/01/03/migrate-your-rails-app-from-heroku-to-aws-lambda/
  • ActiveRecord Adapter for Amazon Aurora Serverless: https://github.com/customink/activerecord-aurora-serverless-adapter

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with Ken Collins. Hi Ken. Thanks for joining me.

Ken: Hi Jeremy. Thanks so much for inviting me.

Jeremy: You are a Staff Engineer at Custom Ink. Why don't you tell the listeners a bit about yourself and what Custom Ink does?

Ken: Yeah, I think maybe first I'd like to say thank you for inviting me to the podcast, really great to be here. I definitely would like to think that it's not because we both have 38-inch, identical Dell Curved Monitors. A very exclusive club.

Jeremy: It is an exclusive club.

Ken: Sure. Custom Ink, let's see. I'm a Staff Engineer, I focus mainly on the eCommerce side. Custom Ink is about 20 years in business. And, I think we're probably only unique in the fact that we're a successful company that has a long history with Rails.

We've been going along for those 20 years, a lot of those years have been with Rails. And, we've sort of completed a lift and shift into the cloud since about 2017. And, we've had some interesting things to do with cloud adoption or adoption of serverless and all things basically AWS.

Jeremy: And, what about yourself? What's your background?

Ken: Well, let's see. I'm a self-taught programmer. Used to be a designer, used to be a marketing director. I think at one point in time I was the author for the act of record SQL server adapter. So, represented the Ruby community and Microsoft when they first started doing their transition to open-source.

And, I think I really love open-source. I'm focusing mainly on retooling my personal career and learning everything about AWS. And, that started about last year. And, doing everything I can at Custom Ink to sort of sell the serverless story, and to get more cloud adoption within the organization.

Jeremy: Very cool.

Ken: Thank you.

Jeremy: All right, so I wanted to have you on today because I've had a number of guests. And, we talk about serverless in theory all the time. And, we have all kinds of great ideas of architectures. And, I mean we get into some of the practical stuff. But, the hands on piece of it, and how companies like Custom Ink are actually getting their hands dirty. And, doing the work to figure out how to implement it. Every one of those stories is different, and I just really love the story that Custom Ink has. I think I saw a testimonial on the AWS site about sort of how you started with it.

I want to get into that because I think that's really interesting for people to hear how other companies get started with serverless and Lambda. And, how they start adding that. And again, you've gone through the experience of the lift and shift. I know you have some interesting microservices stories that I think would be great to hear about. But, let's start with that, let's just start ... you started moving to the cloud, you did the lift and shift thing. And, I'm assuming those were mostly monolithic applications, so what was the next step for Custom Ink?

Ken: Yeah, I think our story arc, maybe about 2014 was key sort of monolith. We had a very traditional big Rails frontend and a big Rails backend that sort of shielded us from a legacy Java backend. And, at that point in time, and we still didn't finish our cloud sort of migration until 2017. Which, is basically just a bunch of EC2 instances.

At some point in time in 2014, I think that's when Lambda came out. And, there was a lot of buzz around microservice architectures first. And, we even had a business unit that had started off in 2014. That we sort of gave them this majestic monolith, and the business unit decided to immediately retool that entire monolith into microservices first. It was just the hot thing to do in 2014.

Jeremy: Sure.

Ken: And, I believe that took about a couple of years to really fail miserably. In fact, everything that they went to engineer on, just breaking things apart. Eagerly because that was the architecture to do, versus the success of the company driving that microservice architecture all rolled back. All changed, eventually we got merged back into the core business line. And, that really sort of affected, I think a lot of the corporate memory about how we approached microservices.

Jeremy: Sure. Then after you sort of had this epic failure. And again, I think this is nothing against Custom Ink. Because, I think this happens to a lot of people, who try to do that second version syndrome and say, "we'll just take our existing application, break it up into smaller pieces and everything's going to be great." That typically doesn't work. Sometimes it does, but most of the time I don't think it does. Then you shifted to this idea of a kind of using Lambda functions and serverless in general really to start splitting off, you actually started more with DevOps tasks, right?

Ken: Yeah. I think there has always been a little bit of AWS Glue. I call it, when I look at Lambda I've sort of from my perspective today I put them into three buckets. AWS Glue is when you're just sort of glueing things together, maybe you're popping Kinesis Streams off of a DynamoDB. Or, you're just doing some little small tasks, maybe scheduling the shutdown of EC2 instances. Then there's the sort of microservice architecture of where I think a lot of people have their head space around Lambdas. And, then those are more sort of larger applications.

In, 2017 we had a brilliant engineer named Hunter. Who, took a key part of our design architecture and that really needed to come off of a Rails app, and ImageMagick. And, put it into a Node-based Lambda. And, that's what our customer testimonial on the Lambda product website is about. I think we had 90% in cost savings. Of course, we got the scaling up and the scaling down feature. And, I think starting in 2017 past that 2014 sort of failure story. We really got a good idea of what Lambda could mean for us, from a microservice perspective. And, it was all about the cost savings. And, basically breaking apart a key part of our infrastructure, around our design lab. And, pushing that just to Lambda. I think that's a traditional what they call sort of a North South client. It sits directly where we're exposing that microservice as a frontend, of graphic application to the design lab.

Jeremy: Yeah. And, I think that's a great strategy too. I mean first starting off getting that confidence with Lambda. And, being able to use it for things like you said, shutting down and spinning up those instances. And, using it for those peripheral sort of work cases or workloads, I think is really interesting. But then, that was I think the right approach. Where, and in the end I think this is why you had some success with this is; you took one specific piece of functionality or one sort of bounded context. Which was your graphic conversion system. And said, "we're going to build this as its own separate standalone thing." And then, how did that kind of integrate back into the rest of the application?

Ken: I think from an integration point of view, it really doesn't integrate at all. Other than what you're just calling this one service and you're offloading certain paths from one application. The way that it sort of fit in with our architecture is that, we would see more of these opportunities when breaking key parts of Rails applications outward. And, just converting very small parts of them into Lambda. Another good example that we did is our catalog application was a Rails application. And, it was still pulling S3 images right from the Rails app. It was also another application that was a key API for a lot of product information. Hitting that Rails application with s3 io eventually made the API less performing.

You found you had to scale a small part of it to make the whole better. And, we just sort of over the next year just sort of reapplied that pattern. What needs to come out of an application? What needs to be more performing? And, just pop it out. We let success define the lay lines of the architecture.

Jeremy: You go through the microservices exercise, you start building out these other components. How did you kind of go from lift and shift through to microservices. Then to using Lambda for these sort of side jobs or whatever. And, then moving into sort of this idea of full-stack serverless?

Ken: Yeah, that's a great question. I think full-stack serverless is a term that I sort of popped up on. I've totally co-opted it from Nader Dabit's Amplify work. And, the idea of sort of came to me when one of our engineers, one of our senior engineers. Looked at these different microservices that we typically had in Node. And, each one was very unique. Sure you would have like it would be on Lambda, but there was no sort of convention on how might your Pojos, your plain old JavaScript objects might look. How you might do routing, et cetera. And, he asked if there's any way we could sort of bring structure like we had in our Rails applications. To these microservices and or other things. And, that got me thinking, like how can we get Rails in? And, it never sort of came about, or sort of I thought about it before. But, it was only maybe about a few months after the official Ruby runtime release.

It was kind of about that time I was like, how can I get Rails into this Lambda? How can I make it work? How can we put structures into these Lambdas? Not only for microservices but also for maybe full-stack applications, and is that even possible? And, that started about last year with the project that we had called Lamby. Which, literally allows you to drop a Rails application into Lambda.

Jeremy: Yeah, that's pretty cool. Tell me a little bit more about Lamby.

Ken: Well, technically speaking it is a small rack adapter. Your Rails application would, in most other environments, whether you spin it up through a thing called Passenger with Apache. You would just start a process and then you would send it messages. Rails has a, or at least Ruby has a system for talking to HDB applications called Rack. And, if you look at the data structure for sending messages to a Rack application, it's basically an event.

And, all Lamby does technically is; it converts either Application Load Balancer or API gateway events in to that Rack event. So, that your Rails application knows how to deal with it. Your Rails app is no wiser of whether it's mounted in Apache, or Nginx or anything like that. All it knows is it's just getting Rack messages.

Jeremy: Very cool. All right, so let's move on and talk about the architecture of the first project that your team did. This was a little bit interesting. You made some, everyone makes architectural choices. Every architecture looks a little different, but you used Application Load Balancer instead of API gateway.

Ken: Yeah. Right. So, we used our API gateway for a lot of our microservices, that had just come out at the point that we just made the Lamby gem. That ALBs became a choice. That actually felt like a comfortable choice for me. I had no other reason to use it other than I knew I didn't need the features of API gateway. The only thing we used it for was just basically proxying events right down to the Lambda. Even on our microservices. It was just all about that HTTP event. So, ALBs made a great choice for that. And, the architecture of our first Lamby app really sort of afforded us a new way to sort of adopt and look at the cloud. One key aspect that we really liked, was sort of infrastructure as code. None of our lift and shift EC2 instances other than, we're basically just configuration management.

Lambda through the AWS Sam framework afforded us this infrastructure as code first dip into these architectures. The app was eventually, it's kind of like a Photoshop in the cloud application as a rendering services. We talked to DynamoDB and that we even deployed this application multi-region for availability and redundancy. Use a latency based routing with CloudFront, and even tied together a DynamoDB and S3 tables. Through cross-region replication and global table usage. Even did some optimization of another Lambda that spun off on the side, where we did image optimization. So, traditionally a Rails application might approach background jobs through something like a Redis database. And, adopt a processing system called sidekick. And, one of the things that I've been really advocating for is; when you think about these full stack server-side Rails applications.

And, Lambda, you don't try to lift and shift your thinking of what you might do with container based development. Either with a sort of Fargate, our Kubernetes. You really look at all the constraints that Lambda gives you, and you look at those as opportunities and ways to sort of feature develop your application in that new method.

Jeremy: Yeah, and you have a great slide that outlines the architecture here. And, it shows, the latency based routing. It shows the DynamoDB global tables, and the replication in here and some of that. So, I will make sure I put this in the show notes because I think this is super interesting.

Again, I love seeing how companies or how developers are building big apps with it. Not these little point of, or at least a proof of concept type app. So, this is really, really interesting. So, once this was built right, and obviously it looks like it's a solid system. But, then you have to put it into production. So, how does your team measure success with this?

Ken: Well, given that I'm sort of a person that likes to think about application usage from a customer's perspective. And, the value that things deliver either for our customer or the business. For me success is adoption. That's, the bar right there. Do we get it out? Is it getting used by other platforms? Are our marketing and other sort of social teams picking up this application, and using it for onsite personalization? That success to me. Nothing in CloudWatch tells me if it's success or not. It's, if it's not performing, then we can optimize that. We can break apart the application into two smaller bits. Optimize the rendering engine. Move from ImageMagic to Libvips, et cetera. So, success always for me is; did it reach the customer? Did it provide business value for someone?

Jeremy: Yeah. And, I think that's important. Like did it provide business value? And, part of that was, I mean you mentioned the 90% savings. But, what about like time to market? Like how long did it take your team to do this? How many engineers did you use to do this?

Ken: Well, I will admit that I totally shaved a yak here. I had to make Lamby in order to get this app out. So, my performance on this one is probably not good. What I'm interested in is if other people can sort of take the work that I've done. And, the Lamby gem and the work that we've done, and sort of co-op that for them. And, I think that's a key part of what we're going to be doing internally at Custom Ink. And, the idea is simply this; if you can take this, create a new application or an optimization. And, simply do a Rails new or follow the Lamby docs and get that spun up in hours, to days. Then you've fulfilled like one of our most important tenants at Custom Ink. And, that is ownership.

Can you control everything from the application of the creative side, the frontend, the backend, all the way to production. And, own that process all the way through, whether it be to an internal team or an external customer focus tool. And, make it happen. And, I think with Lamby you can do that in hours.

Jeremy: Yeah. And that, you know what, and what I really love about what your team did with Lamby is the fact that you went out. You built this for your own purpose, but you're all, everyone's trying to solve their own problems. That's one of the big things we do. We try to have this, I guess philosophy in serverless where we're not trying to reinvent the wheel. But, sometimes the wheel doesn't exist yet. And, so companies like Custom Ink, will build these things. And, there's other companies too. Like Nordstrom has done all kinds of amazing, they have a bunch of tools and services that they've released. And, that's just really great to contribute back to the community, so that the next group makes it easier. Since we are very, very early still in tooling and all this other stuff that has to do with serverless. So, that's awesome. So, I get that it took you a little bit of time to sort of get Lamby up and running. But, once you had it up and running, you're saying is pretty fast process to build these applications.

Ken: Yeah, absolutely. The hardest part is building what you need for your customer. And, I think that's what the focus should always be on, and the time should be spent. Not about like how do you get this infrastructure? How do you spend up EC2 instances? How do you configure for gate or God bless it, Kubernetes?

That I think the time that you spend in your app is going to be the time that you're delivering the features. The getting it up and running and getting your first [inaudible 00:16:29] in the cloud, can be hours. And, there's a story that we're telling here. We have more work to do with Lamby to make that story better.

Jeremy: Sure.

Ken: So, for example, we just finished the adoption of our Aurora Serverless gem. And, that opens up moving from DynamoDB to sort of your traditional relational database with the Aurora Serverless option. And, the next step would be something like with RDS Proxy. And, then after that we would really like to have something done to where we have this integrated with the serverless application repository. Where, someone can literally hit a button and start off with a Rail 6 app. And, a full CI/CD pipeline. And, just go from there.

Jeremy: That's awesome. All right, so let's go on to the next sort of phase here. So, phase one, you built this application out. And, you've had some success with, you had a lot of success with it. You had to build Lamby for it. So, where's Custom Ink going now? What's the next phase?

Ken: Yeah, I think internally the story is a lot bigger than a Lamby. So, we're working on trying to figure out how to get a lot of our EC2 instances into things like Fargates, or some sort of container based system. After that, our Lambda story, which I'm trying to spearhead is really broad. That's going to be, we have a lot of data teams that are doing traditional sort of ETL style data management.

We want to move to real true event at architectures that are afforded to us with EventBridge and Lambda Destinations. And, popping events off either from Aurora or DynamoDB. And, those will more than likely involve a lot of sort of smaller either Node, Ruby or Python base, Lambda glue and or Microservices to do that. My hope is that over the next year or so we're going to have teams during our hackathons and other sort of business units spin up these Rails applications.

Much like the one that we did for our initial Lamby deploy. And, just crack open new possibilities. We've got a lot of stores launching. Like physical stores that you could walk into. The possibility for reinventing small parts of the company, are sort of numerous. And, the good thing that I try to position my thinking around, is; I kind of don't care what that innovation is. If, I can just make the tooling so that smart people on the edges of the company, can sort of take that tooling and just run with it. That's where I succeed.

Jeremy: Awesome. All right. So, what about some of these other things that you're building though? So, you had mentioned this idea of kind of bringing back RDS into the mix. And, with the last project you did you heavily adopted DynamoDB, you got the replication from it. The global table, some of that stuff. So, I'm curious, why the move or why the focus back to RDS or RDBMS I should say?

Ken: Yeah, I will admit that this is a purely sort of adoption. I want to give people the option to look at Rails, and its entirety in Lambda. And, that from traditional Rails is a relational database. All the tooling around Rails is built around an open source adapter pattern, which can facilitate anything from SQL server, Oracle, Maria, MySQL, Postgre, et cetera. So, I think it's really easy for Rails developer. A lot of the process in Rails for running database migrations, it's just built around these relational databases. So, first and foremost, DynamoDB is great. It does have a steep learning curve. I'm almost certain my first implementation of the single table strategy that I did is okay. But, I would have rather started with a relational database, and sort of migrated to there. And, that's the story that I want people to adopt Rails and Lambda with.

So, I knew that wasn't possible when things first came out. You, I believe the written sort of a Node package that sort of manages these zombie connections. And, I took the stance where I just didn't want to play with that. I just waited for about a year. And, I waited to see if something would happen. And, certainly I think late 2019 is when Aurora serverless came out. And, that opens up a connection to the database through HTTP. And, it took me about a couple months to get through it, but I successfully wrote an active record adapter gem. That monkey patches the MySQL connection, from native protocols to HDDP. And, it took me, let's see, over the Christmas break here, probably about a week and a half to do it. And, it passes all 6,000 active record tests.

Jeremy: Yeah, that's awesome. That's awesome.

Ken: That's, that's the first story in the database adoption. But, I think it's not complete. It's only really good for say, these infrequent workloads or applications that need to just kind of sit there. Maybe pet projects that people have on Heroku or good candidates. And, we've certainly, at Custom Ink, we've migrated our internal innovation app. Which, sort of allows people to share the ideas around hackathons that we run twice a year. And, we just pour that right on over, right off of Heroku and Postgres and migrated it to Rails with serverless.

Jeremy: Yeah. And, that open source package, you're talking about serverless-MySQL. I didn't want to build it either. It was just one of those things where it's like you had to at the time. And, the data API is great. And, that's the HTTP one that you're talking about. I've actually found it to be very useful for a number of things, especially in the fact that you don't have to put it into VPC. Which is incredibly helpful if you don't want to have to get sort of that extra latency in some of the, set up a Nat gateway. And, some of these other things you have to do if you're running a Lambda function inside of VPC. And then-

Ken: ...you had a performance cost with that as well. For the startup time.

Jeremy: Yeah. I mean that's kind of gone away. I mean it's still there. I mean, but it's nowhere near as bad as it was before. And, I think that you still have a problem with relational databases. You're still going to have some sort of scaling problems.

So, even if you handle the zombie connections, and you handle the connection pooling, and some of that other stuff. So, you're not overwhelming your database. You're still going to face scale issues I think when you get to a very active system. So, do you have any plans to sort of deal with that, or is that just something that you kind of, is up to the implementer?

Ken: Yeah, I do. I think right now I'm putting my effort into looking into the newly released RDS Proxy, in preview mode. And, I think that'll help out. I don't know, I haven't looked too deeply at it yet. I don't know if it's going to work or not. But, from what I've looked at, apparently you can use the same MySQL protocol. And, then which means it should be transparent to the Rails application. And, that it will manage all this for you. Even in some of our larger EC2 applications with Rails, we have issues with connection management. There's already a need today for the RDS Proxy.

So, bless their hearts if people do want to put Rails, and sort of bring this true Lambdalith into reality. And, use a relational database. I want that tool to be there. And, I want to tell a story with that. I actually think RDS Proxy can help out with that. Yeah,

Jeremy: Yeah. And, I think that, I've had this discussion with James Beswick. And, basically, RDS or relational databases aren't going away. And, there's a ton of, there's absolutely a need for them. So, they are very useful for the right types of applications. So, I do think that, that's, the RDS Proxy is an interesting way to try to solve that problem. Again, even if it solves the connection problem, there's still potentially getting a really big RDS cluster. Is going to have its limitations at some point too. So, it depends on what you're doing with it. But, certainly for most use cases, I think that, that makes a lot of sense. So, what is Custom Ink doing to sort of help other engineers in your organization learn serverless, and learn sort of this cloud native stuff?

Ken: Yeah, we've been working really hard over the past year to up our game at cloud adoption. And, that's a broad story. So, we've done everything from a lot of internal workshops either around Fargate, Kubernetes, Lambda and stuff. We definitely have, I think we just, we do A Cloud Guru teams account. We've started encouraging certification at all levels, everything from management to the engineers.

I believe we have maybe 20% of our team now certified. And, I come from certification as a weird topic. So, I'm certified for the developer associate, I'm working on my developer pro and I think the architect pro. And, coming from sort of an old schooling system to where I never got grades in my initial school, for some reason my parents sent me to some hippy school and I never knew what grades were. So, I tend to just do things for the doing. And, achieve where I want to achieve. And, I believe certification in some way is learning how to pass a test. The real skill comes from how you apply that knowledge. So, we've taken the approach that certification is a way to open people's eyeballs up. Just so that they can have a good candid architecture discussions, when we're doing designing and architecting without, not from an implementation point of view, but just sort of from an understanding. So, if you simply knew that S3 batch operations was a thing. And, that got you to implementing some sort of Lambdas, and inventive hooks and data processing. That's success. We don't look at certification as a way of saying, "yeah, I know how to implement everything from VPCs to Lambdalith to microservices and all in between. It's knowing about where to look first, and then going and finding it afterwards.

Jeremy: Yeah. No, I think that's a really good point. And, that's one of those things with cloud provider like AWS. There are just so many services and so many, so many sub-services in a way. Different features of a particular service, that just getting exposed to know that something like that exists I think is great. So, you're doing some training and you're doing some hackathons and things like that. Is it something that you give engineers extra time to sort of focus on that learning piece?

Ken: Yeah, absolutely. And, we do it in two ways. So, one of the other ways that we sort of encourage all the cloud adoption, is we formed an internal Guild. That we've simply just call the AWS ambassadors. And, basically we take some of our senior engineers, and we put them into this group. And, we look to sort of build out this cluster of people that have that sort of cross functional knowledge.

So, there would be an ambassador that focus mainly on machine learning, Sage maker, etc. I'm sort of the ambassador that focuses on Lambda, and these Rails applications. And, we have many more that fill those gaps. Of, all those AWS services. And, then I think with regard to the teams, we have bi-yearly hackathons. Every other Friday, people get to work on what they want to. Whether that be sort of learning or training. Or, they can continue the sort of, we call it the imaginate day.

So, traditionally when you hold these hackathons. For multiple days, you're going to sort of be working with different business units. People that are not really in your sort of small fire team or group. And, learning to solve sort of problems outside of your normal purview of project work and roadmap stuff. What happens is; is if you don't care and feed that process throughout the year. Then those things just sort of, they gain traction and then they kind of die off. And, the idea with the imaginate days is that you have time and are totally encouraged, whatever it is you're doing. Cloud adoption, solving a business problem, to pick that stuff up and keep driving it every week or so and just bring it home. So, it's really up. It's really the individual's gumption on if they go into the stuff or not. We've done enough on our side to open that door and one of them is the cloud adoption and training.

Jeremy: Yeah, I love that. I mean that's one of those things without, there's just so much to learn and there's so much that's different. That if you don't have sort of a constant stream, and have some time to read the blog posts. And, go through the docs and experiment with some things and do that. It becomes really a, I think you can fall behind very quickly. And, that can be very frustrating, especially considering that I think serverless gets more and more complicated every day. Because, they keep adding new things to do. And, while it's got a whole bunch of benefits. The learning piece of it is still pretty tough.

Ken: And, that was the huge reason why I wanted to approach it from Lamby. I know when you just said you and Chris Munns talked, you all joked a lot about the Lambdalith. And, I want people adopting the cloud. And, one of the things, I attended Serverlessconf 2019 in New York. And, it almost seems cliche that everybody is coming up with their own, I call them like these serverless specific micro frameworks. This is how we do it. This is how we're sort of using Lambda with full stack applications and stuff. And, I found it kind of tiring. Like if you can almost dedicate for each one of those, and it could be that maybe in two years or three years time. That is going to not be around anymore, or Lambda is going to evolve in such a way.

The, idea that we've gotten with sort of Rails into Lambda is that your Rails app can live on no matter what. I can move it from Heroku, I can move into Lambda. If, I want to I can move my Rails app from Lambda to Fargate or EC2 or DigitalOcean to wherever I want to take it. It will work. And, it's been, Rails has been working for many, many years. And, I like that idea that you almost have some sort of bulletproof portable, for those people that want to go, I want to be cloud agnostic, workload essentially. That can be put everywhere and it helps me focus on the innovation at the Lambda level. Like the provision concurrency. Which I don't think we need, but I can now start focusing on things like destinations and other things and how AWS works. Versus these micro frameworks and the tooling around them. But, then again, people are different. So, anyway, you get into the cloud I think is a good story.

Jeremy: So, I think that's really interesting. Because I mean one of those things or one thing to think about when building services for Lambda and other managed services. Is that we want to kind of forget about how we used to build applications in the past. Because, things are so different. So, Taylor Otwell built his Vapor project that allows you to basically take a Laravel project and stick it into Lambda. And use, SQS and the database and some of those other things to just sort of make a Lambda, sorry to make a Laravel project work. And, you've sort of done the same thing with Rails. Yours is open source a little bit different. But, how restrictive might that be? When we think about building applications, don't we want to build them? When we build serverless applications, we want to build them using very specific technology. Making very specific technology choices, component choices, things like that. Do you have any worries that Lamby will kind of box people in and maybe not be able to take advantage of all the best practices?

Ken: Yeah, I don't think so. I think, if your app is the app. And, you need to move that up elsewhere, that's going to be portable. What's not portable is the implementation that you put around CICD, the IM permissions, the other things like that. That's always going to be the cloud locket, and I've never been anybody that's sort of like not do anything for vendor lock in. That to me is like a flood that you can easily just call out in a meeting, and then just stop things from moving.

Jeremy: Right.

Ken: So, Ruby code is Ruby code. It's going to run wherever you put it. So, if you have it run in a Lambda, doing popping events off of S3. Then your vendor lock in or your portability is just going to be what events you listened to. And, the Ruby's going to do what the Ruby does. And, that's true for Node or Python or any language. I always approach it from the language is portable. The application is portable. The way you integrate with the cloud, let's say if you're taking full advantage of API gateway and you're, proxying it right to DynamoDB and doing all those things. Those are definitely not portable.

Jeremy: Sure.

Ken: They're very valid. If you want to build a mobile application with amplify. And, they tell a really good compelling story. I would never say someone don't use amplify because you're going to get vendor locked in into the full suite of full stack services that AWS offers. Go all in, get customer value. That's the thing we should be doing.

Jeremy: All right. So, that brings up sort of this interesting topic too because Rails apps are typically Rails app. I, would see them as sort of their own little microservices. You can certainly build them out that way. So, how do you see serverless microservices? Because, I think people think of them differently. But, what's the organization of a serverless microservice look like to you?

Ken: I think that the ones that we've traditionally thought about are image processing. So, custom make, we allow people to personalize apparel. So, we have a flagship product of ours called The Design Lab. And, everything that we've pretty much used serverless for, and microservices. Are going to be some form of image serving an image optimization. So, for us it's really around taking an HDP event and giving a binary image back. And, that could be because I was a graphic artist, but I think it's going to be different for a lot of people. But, most of our microservices center around image manipulation and returning images back.

Jeremy: And, that's with multiple Lambda functions or are you building them mostly with these Lambdalith?

Ken: You actually could describe these as a Lambdalith. So, the one that does a lot of our proof generation. It has a, it does all this routing internally. It basically can do maybe up to 30 different types of routes, and 30 different types of images for any type of product and or design that you're putting on it. That is a Node, microservice that we call. Its package size is probably 35 maybe 40 megabytes compressed. So, it gets up to the size there. Because we're sort of bringing the layers in for our Libvips and things. So, it's very possible. Somebody might look at that, and even though it's just a small Node application that does one thing. Return product and images back with proofs on them. Somebody could say that, that's a Lambdalith.

Jeremy: That's interesting. Yeah. I just, and you're having success with it. And that's, and I mean the same thing. You know Michael Hart and Bustle, they use these Lambdaliths to do a lot of the processing. And it's, because there's less overhead in understanding how each individual function works. They have the same sort of scale behavior. So, it's not like you need to be able to scale up one particular piece. But, even how you're breaking it out. So, I think that's interesting. And, I think that the way people approach this, that's why I just love talking to people like you. Because, I love hearing all these, I love how people are actually implementing it. Because, we know what the best, practices are or supposedly are. But, those tend to break down when you put things into production. And, in terms of how people find more success and actually in implementing that.

All right, so before you go at one more question for you, because I think this is always super important. Especially for new companies or companies that are just starting to adopt serverless. What's your advice to other companies that are looking to get started with Lambda, or serverless, or just even maybe just starting migrating some things into the cloud?

Ken: Yeah, I think it's always to look at what your business is doing right now. And, where you need it to be sort of performing at first of right. So, always drives my success. I'm a very huge believer in DHH the sort of creator of Rails. That you do the majestic monolith first. You build an application out, and then you sort of look at where it needs to either be performing or broken apart. For Custom Ink. If, your story is anything like ours, it would basically be starting with the monolith. Looking to where sort of business units lie in that monolith and then breaking out into what we sort of call key domain services. So, we would extract the design lab from the monolith. We would extract the product catalog from the monolith, we would extract, a group order form and quoting systems and things like that from the monolith.

So, that to me is a really good way to sort of adopt the cloud. If you know you're going to be breaking up to this monolith into smaller parts. Some of them could be Lambdaliths, some of them could be say Fargate or EC2 instances, whatever. But, I think when you look at what's happening with your current application, let your success drive your architecture. And, I think Lambda is a good place for either moving apps, but it's also a really good place for, if you're in AWS. To question if that's an opportunity for you to look at doing data, and events and units of work in a different way.

Jeremy: So, step one is not just break up your entire app into a bunch of different microservices.

Ken: Oh yeah. My predecessor for Rails Atlanta did, I think it's a similar architecture. It was called a system called Jets. And, if I remember right, the architecture looks similar where in Rails. You have sort of this MVC controller and action pattern. So, you'd have a controller, a controller can have a number of sort of restful routes. And, the way that, that was thought about is; it's kind of like what we did in 2014. Every controller would be broken up and every action into its own Lambda. And, essentially you might have 30 or 40 Lambdas talking to each other. And, I can't even wrap. I think I'm a smart person, and I can't even wrap my head around how you'd managed the CICD pipeline. And, the maintainability of that.

Jeremy: It does get complicated. Well, anyways. All right. Listen, Ken, thank you so much for joining me and-

Ken: Thank you for having me.

Jeremy: Well thank you for sharing the story of Custom Ink and Lamby. And, what you're working on. So, if people want to find out more about you or get in touch with you, how do they do that?

Ken: Well, I've got a couple of places, so in Twitter you can follow us @CustomInkTech. Or, me personally, I'm metaskills. M-E-T-A-S-K-I-L-L-S on Twitter. We have a little product site that we made for Lamby and that's available at lamby.custommaketech.com and of course we blog a lot at Custom Ink. And, that is at technology.customink.com.

Jeremy: Awesome. All right, we will get all that into the show notes. Thanks again, Ken.

Ken: Thank you so much, Jeremy. I really appreciate it.

View Details

About Aleksandar Simovic
Aleksandar is an AWS Serverless Hero and an experienced senior software engineer at Science Exchange, a biotech company based in Palo Alto, California, that is helping scientists, research laboratories and big pharma companies get faster in experimentation and research. Co-author of “Serverless Applications with Node.js” book, published by Manning Publications. He is based in Belgrade and co-organizer of JS Belgrade, Map Meetup Belgrade and Serverless Belgrade. One of the core team members of Claudia.js, contributor to AWS SAM, AWS CDK, AWS Lambda Builders and many other open source libraries.

  • Twitter: @simalexan
  • Book: Serverless Applications with Node.js
  • GitHub: simalexan
  • Claudia.js: claudiajs.com
  • Blog: serverless.pub
  • The Computer/Jarvis Project: thecomputer.ai

Transcript
Jeremy: Hi, everyone. I'm Jeremy Daly, and you are listening to Serverless Chats. This week, I'm chatting with Aleksandar Simovic. Hi, Aleksandar. Thanks for joining me.

Aleksandar: Hi, Jeremy. Thank you for having me. It's awesome to be here.

Jeremy: You are a senior software engineer at Science Exchange, plus you're also an AWS Serverless hero. Why don't you explain to the listeners a little about yourself and what you've been doing at Science Exchange.

Aleksandar: Yup. You're right. I'm a senior software engineer at Science Exchange doing serverless a bit more than four years at the moment. Yeah, there's a lot of titles here, AWS Serverless Hero, where I work with other two serverless heroes, Gojko and Slobodan on Claudia.js, one of the first frameworks for serverless. Also co-authored a book, Serverless Applications with Node.js with Slobodan, running many meet-ups on JavaScript serverless Wardley Maps scene, Belgrade Serbia. My main focus is serverless, and business strategy, basically building product with serverless and Wardley Maps.

Jeremy: Awesome, all right, so I want to talk to you about something today that maybe is not going to seem like it's about serverless, but I think you and I will agree that it very much so is. That has to do with voice automation or the ability to use voice integration, I'm sorry, voice interface technology. I think that the ability to control something with your voice is absolutely the future of how pretty much most interactions are going to go. Maybe I'm a little bit crazy here, but I think you sort of agree with me?

Aleksandar: Yeah, this is something that there's a lot of heated discussion about, but I'm going to just tell you a story of this Christmas I saw my seven-year-old nephew, who basically doesn't ... He's Serbian. He doesn't know English. He doesn't know how to type properly. He doesn't know the Latin letters. I saw him using the phone in a very different way than we used to use it. He basically started ... He only uses the phone by using the Google Voice function, so he opens up the phone and he just presses the Google search function and he basically just says what he wants without even typing or anything.

For him, that was the most easy way to interact with technology. And that's something which blew my mind as I saw that the way we are interacting with technology has evolved so much that in our age we sort of ... We started tapping on the iPhones and everything, and now we have a new kind of age slowly creeping in using voice.

What's surprising is that for many humans that are not used to phones, are not used to the traditional ways of using technology, voice has become something as a normal thing, something very ordinary.

Jeremy: Yeah, and promised the listeners we're going to get to why serverless is important here, but I want to just quickly start with ... just sort of lay this out, like lay out the groundwork here and what we mean by voice interface technology. When we started with visual interfaces we were using desktops or computers, and then everything started shifting to mobile, and companies started thinking mobile first. Now there's this thing, sort of voice first, right?

Aleksandar: Yeah.

Jeremy: We've seen this with Alexa and Google Home and Siri and some of these other things. It started very simple, where we were saying like "Oh, Alexa play this song." Or, "Alexa set a time," or things like that, and I hope people aren't playing this over the speakers so that their Alexa devices are going crazy. I should say, "Alex, order a 100 rolls of toilet paper." But these sort of interfaces now have become much more sophisticated.

The technology's much more sophisticated, and now people can do very, very complex things. I want to get into that in a minute, but when I think about voice interaction or this idea of using your voice to control different systems, and of course this home automation and all this kind of stuff, this was sort of predictable right?

Aleksandar: Yeah, so voice ... As you can see, everything that we are ... in technology everything evolves, and everything evolves so fast, and how do we ... The main issue that we have is how do we anticipate change. How do we anticipate what's going to happen? Luckily, maybe around 15 years ago, I know something called Wardley Maps has appeared, some kind of strategic maps developed by Simon Wardley, one researcher at ... one amazing, actually, researcher, and a former CEO. He discovered this way of how can you actually anticipate change and have a situational awareness of how things are going and evolving. Many of your serverless listeners already heard about him, but he basically created this concept called Wardley Maps, which kind of represent the strategic maps of a business landscape.

Which are kind of represented in a form of a value chain of components, which evolved over time. Now that doesn't sound very, like, novel, for some people maybe, I don't know, but basically he created a very visual map, visual way of mapping business surrounding. Based on that, you're able to anticipate how things are going to evolve. For example, we know about the electricity, how electricity was something novel, new, unknown, coming to a point where it's commoditized, industrialized. I mean all of our common lives are kind of pointless without electricity at the moment.

Our technology and the things we do are pointless without it. And as these things, as Simon developed this amazing mapping technique, and basically a structure about our own strategy, he found out different things that were going on. For example, that as new things appear, as things become commoditized, sorry, they become ... you are able to build some things on top of them. We can see that with our electricity we got radio. We got television and we got computers, internet, and we came here to serverless.

So, basically, what happened, I mean what Simon saw 15 years ago, is that there's going to be a ... He actually even created the first serverless technology in his company, and he basically, 15 years ago, he said there's going to be something, such as AWS Lambda. There is going to be something where you're going to have a runtime, a runtime as a commodity, where you won't have to think about servers, where you won't have to think about infrastructure, so basically he developed a way how do you anticipate change and how do you anticipate where is it going and what's going to happen.

Of course, someone's going to say, "Well, it's not a crystal ball. You can't see that much." Well you can't see that far ahead, but you can see things that are in several layers on top of the current technology which we have, and this is where we came to voice.

Jeremy: Yeah, no, so I think you're right about this idea of once things become commoditized, then being able to build things on top of those is ... or adding more value on top of those is a huge thing. That's, like you said, where we come to voice here, so serverless, which is runtime, has basically been commoditized, and I should take that back, so that I don't get in trouble for saying serverless is only a runtime, but functions as a service, let's use that, functions as a service is a commoditized runtime.

Aleksandar: Sure.

Jeremy: It's not just AWS who's doing it. You've got Microsoft Azure. You've got GCP. You're got a ton of open source projects, everything that runs on top of native-

Aleksandar: Oracle.

Jeremy: ... and Oracle, yes, yes, the new Oracle functions, all kinds of crazy things like that. But now that you've commoditized the ability to process ... and not only that, it's not only just processing business logic, it's also the natural language processing and the parsing of the voices and all that kinds of... that's all been commoditized now through different providers, so now Amazon Web Services, and I guess it's Amazon in generally really, with their Alexa device, they have created a completely commoditized platform where it will do all of the voice recognition. It'll do all of the slot filling, and we can talk about that later. And then pass that off to a Lambda function, which is commoditized, a commoditized runtime, and then you can process your business logic off of that.

Aleksandar: Exactly, and what's interesting is now that's ... I don't know, maybe 20 years ago, when we saw the first Star Trek episodes and whenever we saw people talking to computers and interacting, we were thinking it's fantasy, but at the moment now it's the reality. We see that Alexa and other competitors are present in our everyday homes, and what's even more interesting that now you can ... I don't know, I think Amazon reported about 300% increase in shopping on Alexa just last year, so this is kind of ...

We're coming to a point where we have this new interface, new way of interacting with a software, and the new way of how do we interact with other entities, other companies or products. So, basically we're coming to a point where everybody is going to start booking more and things through voice. I mean we already see that, that for example, there's also been a report, there's a developer who actually started earning over $10,000 each month for just one Alexa skill, so things are going super crazily there, and this is still of course the kind of like the genesis where ... I mean where people are custom building their own skills, kind of, trying to build their ... trying to discover how can they voice and in which way.

If you're familiar with Wardley Maps, you can even anticipate that maybe in five to 10 years we're going to have this race of these intelligent agents that are going to appear everywhere. Where going to have like ... for example, I don't know if you remember 2014, 2015, you probably do remember the whole when Lambda came out, how could we use Lambda? There were these cases where people were thinking you know I had this [inaudible 00:10:58] that kind of works for me, but where could I put Lambda, and they started doing small chunks, pieces, a piece here, a piece there, maybe I'm going to do a small function as a converter and we get a PDF converter or something like that.

And now we're starting to see the same thing appearing in the voice space, where we have these voice skills, that are of course naturally, completely serverless, because that's the only ... that's like the most recommended way how do you direct with Alexa skill. Basically, here we see these small pieces where people are actually building small building blocks. You are, for example, you can order, I don't know ... Naturally, of course, you can order things from Amazon or whatever, but you're now slowly starting to see, for example, print, give me some report, or send this or send a message to someone else. And hidden and seen in the appearance of these [Eco 00:11:53], shows in the past two years, where you can even see when something is happening in front of you.

Jeremy: Right, it's like a visual interface that's on top of your voice commands?

Aleksandar: Exactly, and what my kind of prediction, and things that I'm working on, I'll talk about it later, is that we're going to come to a point where things are going to be automated using Alexa on many manual things that we're already doing right now, I don't know, office, office manager thing, office manager tasks, like I don't know, send somebody a reminder or schedule a meeting.

Actually we already can see that using Alexa for business, but all these small pieces are starting ... Like people are starting to discover how can they easily use it, but as serverless evolved, and now people are actually building huge applications on like enterprise scale applications on serverless, and we saw that on Reinvent, at this last year.

This is how things are going to evolve with voice as well, so we're going to see an explosion of higher order kind of software, like more complex software, that's ... You might have an intelligent agent or an Alexa skill that's going to be able to do some financial or maybe do your taxes, you don't know, you know?

Jeremy: Yeah.

Aleksandar: So, things are slowing building. These building blocks are appearing. We see AWS is building this whole serverless ecosystem around itself, where it's going to be a piece of cake actually combining these components and creating something out of the box.

And here we come to a point that ... so something which I've been building in the past, let's say, a year and a half or two years, something called The Computer, which is basically building software with voice. As much as it is, it's ... There's a video about it. I guess you'll put the link there.

Jeremy: Sure.

Aleksandar: Basically, I created these first prototypes of how you can build software using voice.

Jeremy: Yeah, and so this project, and I think you originally called it Jarvis, right?

Aleksandar: Yeah.

Jeremy: It was sort of a Tony Stark type thing.

Aleksandar: Yeah.

Jeremy: But I think that's absolutely fascinating, and before we get into that project through, because I do want to kind of talk about a little bit about that. The voice interface stuff that you deal with now, so we talked about things being commoditized, and obviously things get smarter every single day, right? So we know from the re:Mars conference that AWS had, that Alexa can now do emotion and things like that. There's some scary stuff happening there, but at the same time also some really interesting stuff.

And so as someone who has built a couple of skills, just sort of more playing around with it or whatever, it's very prescriptive, right? Like you have to really script out how these conversations flow. So, if you say Alexa open or Alexa use this particular skill, and then you say do this and then do that, you have to kind of outline how that conversation is going to work. I know Amazon recognized this. I know other companies have recognized that this is sort of a problem.

So these intelligent agents, and I think like Bixby is one, now Alexa Conversations, there's a few of these sort of tools and capabilities what are those about? Because that to me is really interesting?

Aleksandar: Yeah, well honestly, yeah, it can sound a bit scary. To be honest, this moment, the moment I heard that Alexa can recognize emotion, I was kind of initially scared, but then I discovered that actually this is an amazing feature where for the first time I could say in a ... maybe I'm wrong about it, but like in human history a machine's going to be able to understand whether this human is really angry at me or really sad or feeling upset from a voice of kind of perspective.

By the way, that's being by led by one friend I met at Amazon called Victor Rozich, which is a senior machine learning scientist at Amazon Alexa, and I was amazed when I saw that presentation, which his boss, I can't remember her name, sorry about it, and him, when I was amazing how is that possible. We see that every single piece, even this ... the whole Alexa device, and even the corresponding technology behind it is evolving as we speak, and things and patterns are emerging. For example, I guess all of you remember the first voice agent was Siri, right?

Siri was ... people were amazed by like I can tell the Siri this. I can tell the Siri that, but again, those were pretty basic things. It was a novel thing. People didn't understand what it was. I remember there was a podcast with Adam Cheyer on voice where he was ... because Adam Cheyer is actually one of the founders of Siri, where because, I don't know if you're familiar with this, but Siri actually was first on application. It was the first on app.

And then when they released it on the app store, if I recall correctly, Steve Jobs actually contacted Adam Cheyer, and he wanted actually Siri to be part of the ecosystem. I think Apple really lost a big advantage it had over everyone else. It could be like a real competitor to Alexa, and at the moment, we don't see it that much. We just see this kind of Siri shortcuts that appeared and so forth.

But anyway, the thing is, these developers are from Siri they went to Bixby, and they worked at Samsung, and which is interesting is they have actually discovered a better way of how do you handle skills and voice applications and voice agents, intelligent agents.

And they have actually created ... Bixby actually functions in a very different way than Alexa skill. You make those capsules where ... Actually, you teach Bixby on how to interact with a certain API or something like that, you know? You don't actually ask Bixby to create an app or to ask another application or a skill or to do something like that, because let's ask ourselves, how many applications, on our phones, do we know from the top of our heads, at max, 30 or 40. How many of those do we know that are residing in our Alexa device? Probably much less.

Jeremy: Yeah, exactly. I can never remember what the name of the skill is, and that's the worst.

Aleksandar: Yeah, exactly, and you're like, okay, what was the name of the skill? It doesn't work that way actually. So, the guys at VLabs, I think that's the name, that's actually the company there, so they have actually evolved this way of interacting. They have found out that it doesn't work to have like 500 applications, 500 Alexa skills. It doesn't work, but actually you have these capsule where you actually can teach ... that actually Samsung Bixby, you can teach it a certain, let's say, application or some method of interaction.

So, for example, to translate it to Alexa, you could say, "Alexa get me ... order me some cab or make an appointment with my ... I don't know, doctor," and actually it would actually invoke those skills you made, but you won't interact with those skills in that way. But you would say Alexa ask this to do that. You could just say Alexa do that. Which is more convenient, and people ... It's much more communicative for people, and we can actually see that with these Alexa conversations.

So, we see that Amazon has kind of discovered that. We have Alexa conversations where one skill can actually call another skill from another developer, so we see that this is evolving that way. I'm not the person who discovered this, and I'm just a reproducer in a way from Ben Basche, who is a product manager at MultiChoice I think in South Africa, so he actually talks a lot about the way how things ... how intelligent agents are evolving and what's the master agent and all of these other things, so yeah.

Basically, there's a lot of people now investing into voice. You can see that even voice by itself is evolving, and what is particular interesting for serverless, and from the aspect of voice, is that 90% or maybe even more than that, of voice skills are actually built on serverless apps. So, we actually ... what happened is that now we have, besides a web and a mobile interface, we also have a voice interface that we should start thinking about as people are ... as for people voice is one of the most natural ways of interacting with technology.

Jeremy: Yeah, all right, so I totally agree with this idea of these intelligent voice services or these intelligent agents, because that's something that is such a problem or I think a limitation, even when you're using something like Siri, you're ... You still have to remember what it is that you're trying to ask it to do.

Sometimes there's very specific ways that that needs to happen, so being able to just say Alexa order me an Uber, that's probably easy to say, order me an Uber, but if you said something like what's the weather for the next 10 days, or something, and rather than it accessing the default weather app that Siri's got, that Alexa's got built in, like if there was a customer skill that you wanted it to access... and then one thing that's very cool that I'm pretty sure Siri, and I know for a fact that Alexa does it now, is it actually does voice recognition, so not voice recognition in terms of understanding what you're saying, but understanding who's talking to it, which is very, very cool.

Because now if you set that up in Alexa, you can say, Alexa who am I, and it will tell you which user it thinks you are. It's very accurate, so that's kind of a cool thing. All right, so I do want to get more into some of the other business automation things that we might be able to do, because ... and I'm thinking of this, yes, there's all kinds of great things that you can do from a home automation standpoint, and yes, order me cab or order me an Uber or those sort of things. I think there's going to be some really powerful business use cases, but walk us through the Jarvis or the computer project that you did, and just explain basically what that process was and how that worked, because I think this was an interesting use of a voice interface.

Aleksandar: Yeah, well, I mean as I mentioned from the beginning, using this concept called Wardley Maps you can predict how things are going to happen. Basically, as we saw with serverless, at the beginning people were doing small scale cellular programs in serverless functions and slowly started to evolve, as I mentioned before.

We can see that that's going to happen with voice as well, so at the moment, voice can do some very simplistic things, but in the future, it will probably be able to read your email. You can probably type. You can use voice to send an email, just maybe even do CRM, like for example such as seeing the ... show me the last 100 orders that somebody did in my platform or something like that.

Of course, what I mentioned is some people are going to get scared, like am I going to lose my job if Alexa is going to be able to do that or something, but actually what's going to happen is there's going to be an explosion of even more engineers and even more people required to actually operate these things.

We'll need more and more people, and nobody should be afraid of losing their job, so in these small scales, as small sale services get integrated, like these small Lego building blocks, that AWS is providing us. We're actually going to have ... We're going to require more engineers to work on much more complex and larger problems in this space we haven't yet discovered.

There's actually a nice saying by a friend of mine, who was actually a software manager, software development manger in Alexa, who actually, [Ben Dimage 00:24:55], who said, "What is the problem you want to work on? What is the problem you want to solve? Do you want to solve the problem of like manually managing servers or whatever do want to solve business problems that are of high, or more customer value importance."

So, in a sense, it's the same thing with voice, as small things that get automated, we're going to be able to work on larger problems. As we see, when electricity was invented, more jobs were created with radio and everything.

Aleksandar: Anyway, but I believe that the next step after solving these more complex software tasks using voice, we're going to even be able to managed robots with IoT, maybe [Ben Ihoe 00:25:44] and other serverless heroes, who works at the iRobots, maybe we're going to be able to tell your iRobot to go clean the linen room and now clean the bathroom or whatever, and it's going to do it for you. Like I don't know if you remember the Rosie from Jetsons.

Jeremy: Yes, of course.

Aleksandar: So, basically Roomba will be the Rosie of our age, but we see this. I have to again come back to this Simon Wardley predicted this basically 10 years ago, speaking to everything in the future with voice, and everybody ... Nobody believed him.

Because it's kind of does sound weird, and kind of does sound like you're Nostradamus or something, but it's just understanding the way how do we interact at home with technology, and which problems are we solving along the way. And you probably ... you and I discussed before, maybe this Alexa for business, we have the meeting room scheduler, linking email calendars, to do with the reminders, but I have already seen that there's some skills and some people building assistance for bio and laboratory, retail industry, things are evolving super fast.

Jeremy: Yeah.

Aleksandar: I think in five years, it's going to be like ... not maybe five years, but 10 years max, we're going to see something happening with voice, what's happening with serverless right now, everybody jumping on the bandwagon and pushing in all directions.

Jeremy: Yeah, they should jump on it sooner, because ... So the point that I was trying to get to was, and you made a very good point about the sort of encapsulating these pieces of business logic into these building blocks, and maybe they're not only business logic. Sometimes it's a building block that might be an API interface or some other interface, but what you did with that computer project was ... and it's hard to say The Computer project, but what you did with the Jarvis/The Computer project, was you took some of these pre-configured blocks or these Lego blocks that were in AWS, and then you used voice to basically assemble them and launch an application with them.

And that application can do multiple things, and that's why ... What's cool about that to me is that shows a very extreme, in my opinion, sort of an extreme sort of no-code approach to building something with your voice. But there are things in between that, business processes, that would be a lot less complex, and fairly easy to implement, with the technology we have right now, and so I think we were talking about this at one point, where we saying that when you're a developer and you're building these backend systems, sometimes the front end piece of it might be the harder thing for you to build to visualize that or give people some way to interact with that.

And of course with things like the Alexa presentation, UI language or presentation language, you can visualize some of that without having to do any real design work, but I see this as something that could be like let's say you're just ... You're a manager at some company, and you want to see the most recent inventory numbers, so you say, "Alexa, show me the most recent inventory numbers," and either that shows right up on an Alexa show, or it sends it to a dashboard, a wallboard somewhere or it sends you an email with a PDF report in there or something like that. Those types of processes now, those are possible today.

Aleksandar: Yeah, so my opinion is that it's never going to be a voice only, like maybe rarely, like for example like I don't know if you remember you can say hey mom get me something over the phone or whatever, but when you interact with a voice module, my kind of prediction is that you're going to want to see something.

We, as humans, we're not just audio only or video only, which we're not like visually focused only, but we're going to actually want to interact. We want to say something and see the result in front. Like let's say maybe your Alexa skill or something or application got stuck or something, you want to see that something's going on. You don't want to stay in confused and be like, "Okay, what's going on here."

So, in a sense, that's kind of something which yeah, I'd be working on this project called The Computer, which basically is building small scale applications using voice. Building higher order workflows is a much more complex thing.

Jeremy: Sure. Definitely would need visual feedback to do something like that.

Aleksandar: Exactly, and that's something ... yeah, basically what I build is with this computer project is you can very easily explain to Alexa just by saying I want to create some service, I want to create some application.

We can very simply say add this element or add this other element, and anyone, who knows in way to explain their business to just basically just by using voice able to explain what they want from an application, like do they want to save, delete or manage in some way customers and just tell to the Jarvis to create the solution and while they're explaining it they can see on the UI, on the interface, how is this application ... how does it look like?

And then they can just say now deploy this solution and it's done. Naturally, of course, much more complex solutions require a lot more time, a lot more time and dedication to explain it, but it's basically able to do that. Basically, it's just a ... For some people it's more of an experiment, but my belief is that nobody's going to use also voice ... I mean how can you use voice, for example, like that, in your cubicle or [inaudible 00:32:11] software piece. You won't be able to do that.

Jeremy: That's a good point, yeah.

Aleksandar: Exactly, it's going to actually force you to sit down with your UI engineer, UI designer, UX, like the whole team altogether, and you're all going to try to work collaboratively, because you won't be able to use an Alexa device on say hey imagine like 300 people in an office space yelling at their Alexa.

Jeremy: Right.

Aleksandar: That just won't work, you know? So, what's going to happen is we're going to have this collaborative work, but several people are working in a meeting and discussing how should they build an app, and just using an Alexa as a support device for actually explaining what's actually going to happen.

We'll see at Amazon thinking in this kind of space, not really developing by voice, but these kind of no-code solutions I think as Forrest Brazeal, another serverless hero colleague, he even wrote, tweeted, like maybe nine months ago about that somewhere some were hidden. There's this no code kind of solution being in the works inside of Amazon.

Which means they have understand, like even Amazon has understood, that even though we have all these infrastructure services, that are extremely, extremely good and useful, we do not have these kind of UI interfaces, and you can see now a huge wave of no-code, that's going to ...

Like no-code is now the blockchain of 2017, basically. Everybody is no-code now. I mean it doesn't work that way, but anyway, it's ... for some people, it's a nice way to get some VC investors. Anyway, I'm sorry, I'm sorry to be...

Jeremy: That's a little advice from Aleksandar out there for anyone starting a startup company.

Aleksandar: Yeah, you want to get some VC money, just label no-code on it and it's done, you can even put a regular CRM, just say no-code CRM. There really is no code, you just click around it. Anyway, so coming back to this whole serverless, and voice thing, voice interaction, to be honest, we see everybody knows that there's going to be another wave on top of serverless, again, and things are evolving, so there's a high ... there's a big chance that voice might be that thing, maybe not in a direct way how we see it at the moment, but we see like the UI, for example, as you can see, Amazon is investing ... AWS actually is investing a lot into AWS Amplify, which is really an amazing solution, which I recommend to everyone. If you're building a product, don't start building a product ... I mean actually start building a product with Amplify first, and then try to separate the pieces using other serverless technologies.

But we can see that we're coming to a point that where now the UI's the next layer that AWS is building, which is something that we should also focus on, and then we are coming to a point where voice as well as just another side channel interacting, side interaction channel as well. So, we should think about these things as ... as I mentioned from the beginning of this episode, I mentioned this, like even seven-year-olds are capable, and they easily ... nobody showed my nephew how to do it. He just was pressing buttons and he discovered that and he was like, "Wow, this is easy."

Then he was like just first harassing the Google Voice a lot, and then he goes like, "Okay, this is useful." So, it's natural, and my belief is that we had first just visual interfaces, and actually we had the terminal at the beginning, and then we had more visual interfaces.

Jeremy: Right.

Aleksandar: We're coming to a point where we have web and mobile and now we have voice as well, so we have this Three Musketeers of human interaction, human competitor interaction, and yeah, I think it's going to go in that direction. I can even mention there's some experimental project we're working on called The Doctor, which kind of sounds like wait, the computer is building software, what does The Doctor do? Does it cure disease? Does he use it? No.

Actually, it's something ... both are going to be under the ... I mean I already have a domain, but, domains, but anyway, it's actually helping out like a researchers and doctors how to actually discover more important actually significant relationships in between the works they're doing or whatever. I'm going to be honest, for example, let's say you are searching for some heart disease or something, there's like a gazillion articles that you can read about.

Jeremy: Sure.

Aleksandar: If you mention certain keywords, and you talk in a certain way, it's again, it's something that it's not like I'm the inventor of the idea, but basically helping out people who want to quickly and naturally find out certain works and papers. They are going to be able to do it very easily and just by using voice, so yeah.

Jeremy: I think that idea, this idea of medical advice or medical feedback, in the moment, actually could be really, really powerful. If you think about in the emergency room, if a doctor or a nurse that's treating a patient could just say something like, "Does this patient have any allergies," right?

Aleksandar: Yeah.

Jeremy: And that was just kind of tied together and it'd say, oh, yeah, that whatever, or-

Aleksandar: Exactly.

Jeremy: What's the correct dosage of this or that.

Aleksandar: Exactly.

Jeremy: You could ask questions, and that could extend to things that were not quite as life saving as maybe the medical profession, but what if you were-

Aleksandar: Exactly.

Jeremy: Maybe this is a little ... I don't think this is farfetched, and I'm just curious, like let's say you were a plumber, and you're working on a sink, or you're doing something, and then you forget what the right, I don't know, fitting is or what the torque is supposed to be or something, and you could just ask that question, and a system could answer it for you, that would just make people more productive. I know that's one of those things, where I do it all the time, where it's like I stop for a second, and I'm thinking about like oh what's the function to write that piece of code, so I go and I Google it or whatever, and the smarter these devices become, and the more questions they can answer for us, in a way that they we expect them to answer those questions, would be really powerful. So, just a couple more things, and then I'll let you go.

Aleksandar: It's a pleasure actually.

Jeremy: I know it's the end of your day, so things like the accuracy, so you've obviously built a number of apps, or a number of skills, so you know that we use slots and intents, like we basically have to tell the system in what order or in what way to expect us or expect our users to speak into the system. Slots are very cool, because slots have ... They have like a type, so you can say that I'm expecting a number here or I'm expecting an actor's name and things like that, so they're very precise if you capture the right utterances, as they call them, to know what your user might ask it. But even as accurate as it is though, you still think that it's still a bit of a limitation, right?

Aleksandar: Yeah, I'm going to be honest, this is why I said it's going to take 10 years for Alexa to actually really be super, super powerful, maybe not 10 year, but seven or eight for sure. These slots are so, I mean from my experience, these slots are super limited. Even though it can understand emotion, my opinion is that it's on a level of a four-year-old. A four-year-old knows what you want. Do you want me to ... like the most basic stuff.

He doesn't know the answer to Einstein's formula or whatever, or if the answer is always 42 or something, but it knows that if you're angry at it, if you want it to bring something, is it okay, or find that information from a playbook or whatever. Even though it's very precise and accurate, it's still not on a level that we expect it to. I've tried from biological, for some biological states, medical conditions, it understands sort of things, but when it comes to like chemistry formulas or even more complex things, it's very easy to mix certain things and it's not able to really understand.

I've tried with some chemical compounds and solutions, and it really isn't able to understand anything, and not only that, but it's able to even mix certain words, so we're very far away from something super amazing, like an intelligent 12-year-old or something like that. It's still a four-year-old baby basically, which is able to understand many commands, but not a lot. It still needs to learn a lot, so yeah, even ... What you mentioned, it's able to understand voice supervision, to understand oh this user is Jeremy and the other is Alex or whoever, but it's still a four-year-old unfortunately.

But, we see that AWS is understanding ... Amazon is understanding how Alexa works and what are its kind of implications there and ... So, for example, we have this Alexa presentation language where if ... to come back to the point which I did with that actually people are not just audio only. You can use this Alexa presentation language to describe skill at the same time, I mean skills UI and voice basically at the same time. That's an amazing thing when you interact with an Alexa skill, and you say show me, like I want to get an airplane ticket from Belgrade to Boston, where you live, I'm going to be like, well it's going to say well there are five airplane routes that you can take, and you're going to be like, "Wait," and then it's going to repeat the five airplane routes, but if you don't see it visually-

Jeremy: Yeah, you're not going to remember it.

Aleksandar: You're not going to ... You're going to be like, "Wait, what was the..."

Jeremy: Even now, when you get those automated phone things, when you call in and they're like ... they list like six options, and like to repeat these options, you're like-

Aleksandar: Yeah, exactly.

Jeremy: ... because I don't remember, because I forgot the one thinking that it would be something better when it got through, so no, I totally agree with you. And it's funny, the emotion thing, if Alexa can detect emotion, then it's probably not going to like my kids very much, because they're always yelling at it. They always get very angry.

My wife hates it when she asks what the weather it and it has to give her the whole forecast and tell her to have a nice day, and that for some reason makes her angry, but-

Aleksandar: Yeah, but this is actually an evolution of machine learning by itself.

Jeremy: Of course.

Aleksandar: If we take a look at it, we have these ... At the moment, the majority of machine learning is actually in a broad way, like so we are learning from our customers, what do they really want, from a broad range of customers. The next step, which is actually even part of this ... There was a lecture by this friend of ours, Ann, who actually was speaking on the terms of Alexa like the next step is going to be this personal kind of machine learning.

So the next step is going to be where Alexa is going to understand that ... I remember even Gojko saying there's this ... I don't know, some musician that he likes, and he's like, "It's not able to understand." I mean if you say Alexa play this music, and it just puts the wrong artist, and say stop immediately, it should be able to learn that you don't want that.

Jeremy: Yeah, that's not what you meant, yeah.

Aleksandar: Yeah, that's really not what you meant. You want to actually for ... You want Alexa to learn on your habits, so this personalization is also going to be an important evolutionary step into voice.

Jeremy: Yeah, I totally thing too, just and again to tie this all back to serverless, it's been the commoditization of that runtime, that really made the accessibility of Alexa so much easier, and in building skills it's still kind of tough. I mean there's some things to do there, but there's the skill kit or the ask skill, SDK or something like that, that makes building it a little bit easier, but really serverless did enable, I think, what's going to be a mass adoption of some of this voice technology, certainly from the Alexa side of things, at least in my opinion.

But anyways, so listen, Aleksandar, thank you so much for joining me, and sharing all of this complex knowledge about voice interaction technology. Anyways, if our listeners want to get a hold of you, and find out more about what you're doing, how do they do that?

Aleksandar: Well, I mean first thank you for having me. That's super ... I'm super grateful for that, and I really watch and ... Watch, listen and read the show, because I basically, even sometimes when I'm busy and everything, I just skim in the transport or something. I read through who's here and whatever. It's really an honor for me to be here, but if somebody wants to contact me, so they can go on Twitter, @simalexan, or same for GitHub. So, there's three serverless heroes writing on serverless.pub, which is a place where we kind of write our discoveries and the things that we like, and things we want to share with the general ... like everybody who's interested in serverless or something like that.

And we also wrote a book, so if somebody's interested, they can read about it, like Serverless Applications with Node.js. But yeah, that's kind of basically it. They can even send us an email. They can find the email on GitHub or something like that.

Jeremy: Awesome, all right, I will get all that into the show notes. Thanks again.

Aleksandar: Thank you very much.

View Details

About James Beswick:
James Beswick is a Senior Developer Advocate for the AWS Serverless team. James works with AWS's developer customers to understand how serverless technologies can drastically change the way they think about building and running applications at massive scale with minimal administration overhead. He has previously worked as a Software Developer and Product Manager at various enterprises and startups, and has nearly a decade of experience building applications in the cloud.

  • Twitter: @jbesw
  • LinkedIn: https://www.linkedin.com/in/jamesbeswick/
  • Email: jbeswick@amazon.com

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week I'm chatting with James Beswick. Hey, James. Thanks for joining me.

James: Hey, Jeremy. Good to see you.

Jeremy: So you are a senior developer advocate at AWS. Why don't you tell the listeners a little bit about your background and what you've been doing on the AWS developer advocacy team.

James: Sure, so I've been working with serverless for about three years now. So I'm really a self-confessed serverless geek. I've used it to build quite a few applications, front to back using only serverless. And then in April last year, I joined AWS in the developer advocate team, and so this is truly the best job in the world because I like talking about serverless to people, so I get to go around doing conferences, blog posts, webinars, applications, and also some other things to show people how to build things. Since then I've just been going all over the place doing these things, but it's been pretty amazing just to see what customers are building all over the place with these tools.

Jeremy: Awesome. All right, so I was talking to Chris Munns when I was out at re:Invent, and I put together a podcast there, and we were talking about all these new things that AWS was launching. And I think what happens with serverless is that it's moving so fast that things are constantly changing. There's always new things being released. What serverless is is still up for debate, right? I mean, there's still a lot of questions around that.

So I wanted to talk to you because you and I talk as much as we can because I love talking to you. You have great insights when it comes to this stuff, and I wanted to talk to you about sort of what are we going to see with serverless in 2020, right? Because this is the year now where all of these pieces are starting to come together. We've got all of these tools, all of these things we've been complaining about like RDS Proxy, and we can't do this, and we can't do that. These problems are going away at a rapid clip. Maybe you can give me your take just on, I mean, what does 2020 look like for Serverless?

James: It's a great, great question. In the last five years, you know Lambda's really five years old, what's been happening is the space has been emerging and developing so quickly, we're simply seeing customers pick up the tools and build things and then find they need more features. So we've been building all these features as quickly as possible. And I think what's different this year is that this whole space is starting to mature very rapidly. And we're seeing customers, both startups and hug enterprises using all of these tools at scale. And starting to see the same patterns emerging from their use cases.

So what we're doing for the next 12 months is essentially looking at the entire list of requests that's coming right from customers where they want certain things and dedicating those resources to building out the features they want. So AWS is famous for listening to customers and building those features, but I'd say in serverless, I mean it really is the case their entire road map is coming back from these early adopters and these users and helping us to find what we now build.

Now in terms of actual concrete things, most of that comes down to improving performance all the time, always making sure we can make performance as good as possible but also improving tools and making sure that we integrate with developer tools that they're using all the time, and just making sure that all features, we sand off any rough edges that we have. So a lot of the time with AWS features, what we're doing is we deploying them out to customers as quickly as possible so that people get the first look at what we're building. And then when we get that feedback, then we build the additional bells and whistles to make sure it's exactly what people want.

Jeremy: Yeah, no that's great. And the other thing that I, I keep hoping for this, right? And maybe we're not there yet, and I ask everybody about this, but I really want serverless to go mainstream, right? Like it's just what you're doing. It's the way to build cloud applications, right? Because I think you have all of these use cases that are out there now, and from my newsletter I'm always trying to capture use cases to say oh someone's doing this with it or someone's doing that with it, and they have these interesting ways of solving those problems.

And like I said, these problems now have official solutions in many cases. What's your take on this idea of it really becoming mainstream and more customers starting to use it for, or just being the first choice of what to use when they're building something in the cloud.

James: So in my career, I've been one of these early adopt people where I was one of the first in the cloud. And I used mobile and got into mobile development very early. And one of the patterns I see over and over is that the tipping point of things becoming mainstream isn't always obvious. You go through this period where it seems like you're always walking uphill to convince people that this is something that's going to become the standard way.

And then magically at some point, it just does, and you didn't notice it happening. And I started to feel that's becoming the way of serverless because many of the groups I spoke to a couple of years ago, who didn't know what serverless was or they didn't think it was a good fit for their use case, are now starting to openly talk about serverless as an option at least and yet discuss how they could use it.

The great example is that last year I went to the DC Public Summit for AWS, where all of the government customers were there, and a lot of people were very interested in serverless. And I've seen the same thing at all the summits and events we go to that even people who haven't actually done anything yet are interested in what it can provide in terms of both agility and scalability for building their applications.

Jeremy: Awesome. So you mentioned tools and giving people tools in order to build stuff. One of my complaints from serverless, right from the beginning, is even though we are abstracting away all of this infrastructure, there's still a lot of configuration that has to happen. And with AWS that comes down to ultimately using either cloud formation or writing complex interactions with the APIs, which nobody wants to do.

So the CloudFormation side of things, there are extractions on top of that. We've got SAM, the Serverless Application Model that is, makes it a little easier. It's very similar in fell to the serverless framework. Then we have the CDK, which is relatively new that allows you to just write code, and that will generate an infrastructure for you. There's Amplify. I just talked to Nader Dabit the other day. We were going through this amazing tool that is Amplify and how it sets up all these things for you, does back end, does front end, and ultimately all of these things end up generating cloud formation.

So the question I have is, which one do you choose. If you're new to serverless or you've been using serverless. Maybe you're using something like Claudia.js or you're using Architect or you're using Serverless Framework and you want to use something that more AWS native, where do you start? Which one do you choose?

James: I think it depends on where you're coming from. If you're a startup and you're using things, that you're building greenfield applications where you can pick whatever you want, that's a very different situation to be in than if you're enterprise and you're migrating Legacy software into serverless. Most of the combinations of these things are designed for really different developers and different use cases.

So I'd say if you're in a greenfield space and you're doing some scratch, then using a framework like SAM or Serverless makes a lot of sense because you're starting at a point where it's going to build everything out the right way for you. Otherwise maybe if you're in an enterprise and you've got a certain set of tools, you might find that CEK is a more comfortable way to go. But all we're really trying to do is instead of saying to people this is one tool for everybody to go and learn, to really meet developers where they are and give them the tool they feel most comfortable with given their use case.

Jeremy: Yeah, I think that makes a lot of sense. I know for me that again sometimes you have to get into that cloud formation template and start doing things in there, and it can be a, I always complain about this, but again it's configuration, it's [inaudible 00:08:21] language. It makes sense. I mean, it's just as hard with Terraform or something like that, but you are, you know, there are a lot of configuration options there. So certainly as a developer that is new to infrastructure in a way, it is certainly a leap to learn some of that stuff. But again, those tools do make it easier.

James: Yeah, and if you look at something like Amplify, what's been built there is really interesting because when you've got cloud formation, it essentially gives you every knob and lever you have on the entire infrastructure, as you know, as YAML basically, and then when you look at something like Amplify, what it's doing is it's looking at the most common, sensible defaults for given use cases, and helping those developers in an opinionated way.

So if you're building those sorts of apps, that's a great fit. So when we hear customers saying to us that we want to have certain types of use case over and over and over, and they don't need to have all these controls, then we're happy to build tools that simply that.

Jeremy: Yeah, that's awesome. All right, so let's get into some of these tools and products, right? This for me, and of course your feedback on this or your insights on this is, I think will be probably more enlightening than mine. I think there are a few things that are, even things that haven't just recently launched, but tools that have existed. The way that we've been building serverless applications in the past, I think that there are a bunch of these tools, some of them are new, some of them are existing, but these are the ones that excite me the most, and I'd be interested to hear about this from you as well.

But for me I think one of my favorite AWS products right now is EventBridge. And when this first came out, which by the way, was back in July, right? July of last year. Warner announced it at the New York summit. When it first came out, there were a bunch of people in our space, you know we've got a very tight-knit group of serverless geeks that like to write about this stuff, but there was all this talk like, this is going to change the way we do serverless or this is the biggest thing since Lambda itself.

And I totally agree with that. There hasn't been a lot of fanfare around this. That sort of came out and then not much. There was no cloud formation support, which I think was part of the problem, but then you got all of this new stuff like the Event Schema Registry and some of these other new features that launched at AWS re:Invent this past year, so what are your thoughts on EventBridge and just how do you think people are going to be using it in 2020?

James: Yeah, I'm one of the people who are really super excited about EventBridge. I think it has a transformational possibility for the way that you build serverless apps. Because at the very least it can help decouple these applications. So if you built complicated serverless apps, you often find that you end up getting functions and services that become entangled with each other by accident. And by putting EventBridge in the middle, you can totally decouple the producers and consumers in the way where your application is so much simpler.

Now I think what's really interesting is in the last month or two, some of the features that have come out have really evolved the product in a very dramatic sort of way, so the schema registry and discovery features that you mentioned are, to me, just fantastic because I know from building event-driven applications before, one of the hardest problems is just keeping track of events, knowing what they look like, how they're shaped, and when services change versions, the events change.

Having a registry that's built for you just make like that much easier. Then the discovery feature where you essentially just pipe your events to this discovery service and it builds out the schemas is just amazing because it does all the work for you. You get 5 million events per month for free, and that should cover most use cases. And then once you've got a schema in place you can pull it into your IDE and then build applications directly off of that and use events as classes in your applications [inaudible 00:12:23] types. Those two features alone are just amazing.

And then recently we've introduced content filtering. So what that means before we just had rules and rules were just kind of this blunt object, things either match or they don't match. Content filtering now you can put much more dynamic rule sets in place in terms of ranges of values and things make it much more queryable. So the net effect of that is that you can push that logic back out of your code into a service so we're back into the business of less code, more serviceable applications. So all of these features have been coming out pretty fast.

You know EventBridge has a huge road map of things ahead of this, so I just keep watching in amazement. I'm super excited about it.

Jeremy: Yeah, so, one of the things I really like about EventBridge, and maybe this is even too geeky for this podcast, but so if we look at architectural pattern, right, and I'm a huge fan of architectural patterns, so we've got our monolithic applications and then we went to service-oriented architecture and then we went to microservices, and now people call things like nanoservices, which I'm not a fan of that terms, so if you think about the way that microservices, and that's how I like to think about serverless applications, building small services with clear boundaries, their own database to back them.

It might be four or five functions or a hundred functions that are part of one service, but essentially you are encapsulating all of that service logic in one cloud formation template. Or, you know, whatever, but basically you're breaking it down that way. With service-oriented architecture, that's when we introduce the message buses and things like that and being able to pass messages. But we were still sharing databases, and so again there's probably no comparison here.

But what I find really interesting about serverless applications, and certainly when you're thinking about serverless applications as microservices, and then introducing something like EventBridge, is it now what you're doing sort of using this enterprise service bus, if you want to call it that, that handles this communication asynchronously. So everything is completely decoupled. You're adding in the rules and the filtering into EventBridge, but then all of the configuration for it all tied back to the individual services that are subscribing to EventBridge, so you now have this sort of new type of architecture.

And I don't even know what you call it, but microservices maybe, but the way that we communicate with asynchronous just feels so much different. And honestly it feels a lot better to me. I don't know. What are your thoughts on that?

James: I think a lot of what we're building makes distributing computing just easier for developers, and when you think about the scale of lots of developers now have to face with their applications, even things like mobile apps, these are complicated problems to solve when you get spikey workloads and just huge numbers of transactions coming through. So a lot of these tools just make it that much easier.

But the mental hurdle is going from this synchronous model to this asynchronous model. And so if you're used to building synchronous APIs, initially it can seem a bit alien trying to figure out the different patterns that are being involved. But it seems like the natural evolution given the fact that you've got all these services in the middle that have to handle all of this traffic, and the timing issues involved, you know, start to evolve from where you are in the synchronous space, but I think what's been put in place is not too difficult to understand.

Once developers start using this, they find actually for many cases, it's the right way to go, but it's interesting to watch this because I know that just even 12 months ago people were talking about the API Gateway, this 29-second, 30-second limit problem, do all this stuff throughout your infrastructure. Or you heard about the Lambda limits of five minutes, then fifteen minutes because people were trying to work this way.

I think now we're going back to thinking about how do we break up these tasks. So it's shorter-lived tasks that run between services in an asynchronous fashion. So the whole model is really evolving.

Jeremy: Yeah, and so actually that is a good segue into talking about failing in the cloud, right. And so I'm doing a couple of talks at some serverless days this year, and the title of my talk is How to Fail with Serverless, and basically I should have a subheadline that like How to Fail with Serverless so that Your Serverless Applications Don't Fail, or something like that.

But basically what I'm talking about is, when you start doing things asynchronously, I just generate a job and now my Lambda function or whatever my service is, my client that generating that says, okay, here's a job or here's a request and then it says, okay I got the request and then it disconnects. So now this is somewhere out there in the ether. You have some thing and it's routed through EventBridge or it's in an [SQSQ 00:17:18] or something like that.

So you just have this thing out there, and at some point, you hope, that it will trigger something else to process that and do something with it. And those guarantees are very, very strong so you don't have to worry that it's not going to process it. What you do have to worry about is what happens if when I go to process it, something goes awry, and that the Lambda function fails to process it or there's some conversion issue and it can't insert that into the database because that was what it wanted to do.

And what I find really interesting about what AWS has done with the way that they've architected this stuff is to say, listen we know things are going to fail, as Werner says, "Everything fails all the time." I think I've said that about 12,000 times on this podcast, but I totally agree with it because I know. I watch things fail all the time. So when something fails, there are provisions in place that the cloud will handle those for you. There are ways to configure things to be handled for you.

So you have things like the DLQs, which used to be the primary way that you would take an asynchronous event that called the Lambda function if that Lambda function failed, it would put that in a DLQ. They just introduced Lambda Destinations, so now rather than using a DLQ and just getting the event itself, now you get all kinds of context along with that, why it failed the stack trace, things like that, which are super helpful.

You also have a success path, so if the Lambda function succeeds, I don't have to go ahead and put some code in my Lambda function and say, oh now do this with it. It will just automatically do that as part of the configuration. You have failures built into SQS. You added DLQ for S and S now. I'm sure there's all these different ways that we should let the cloud fail for us and not be capturing these events or swallowing these events with try-catches.

So I have a whole bunch of stuff that I'm working on to try to come out with that. But your thoughts on that, what is AWS's, if you can answer this, what's their philosophy on this because this is an essential part of distributing computing?

James: Werner's quote is the philosophy, that everything fails all the time, and so the question is how do you make your application resilient to survive those failures? Most of these new features are really just extensions of ideas we've had, they're in the infrastructure already in one way or another, but you know if you look at DLQs, they've been around for quite a while and Destinations is an extension of that.

Now it's not that we're telling everybody to go and replace DLQs with Destinations. It just becomes another way that you can handle a failure if you choose to, and so there are so many different ways of figuring out where failures work in your application. It depends on the sort of scale that you're working with as well. But I still meet lots of developers who take advantage of these features in the way that they should.

Although our infrastructure is very reliable, it's not 100% reliable. There's always a possibility that you have a service disruption in different services, so these features give you the ability to improve that reliability even further if you use them appropriately. But we're starting to see now with some of these new features with Destinations that it makes it easier to understand as a concept. So I think now developers are getting more comfortable with how you can build this into their serverless applications.

Jeremy: One of the things that I really like about Lambda Destinations is the success path allows you to, again, just write code. So traditionally what you would be doing is you'd say, okay I have to include, like maybe I want to write the information to SQS when it's, I do some processing, I do some transformation, I want to send that back into SQS to do something else, I would have to include the AWS SDK. I would have to make a call to the SQS service in order to post that event, or to post that message.

Now if that fails, I could retry it in my code or I could do some of these things, but there's a lot of logic I was building in there. So now I don't need to do that, right? I can send stuff to EventBridge, which again opens up a whole new possibility of what I can do with it, right? So I'm not just limited to SQS, S&S, a Lambda, or EventBridge. I basically have every service that EventBridge integrates with that I can utilize.

But my question is is that I think that some of that reliability and some of those retry mechanisms that people were trying to, and this is probably not the right way to say it but, sort of jimmy-rig them in a sense by using a Step Function. Because Step Functions are great, and you should totally use Step Functions if you have complex workflows. But even for some of those simpler workflows, it was just easier to say, hey this is supposed to do X and then send the data to SQS, if that fails, I want to retry it. I can encapsulate that in a Step Function workflow, and then that would kind of handle that retry for me.

But you don't need to do that anymore because of some of these new features that are added. But obviously Step Functions don't go away, but what are your thoughts on, you know, this obviously makes it easier, right?

James: I do get asked quite a lot by people, should I use destinations instead of Step Functions. The answer is usually no, but also it depends because if you've got a very, very simple process and it's really just a couple of steps, and you don't want to incur the cost of using Step Functions, perhaps this could be an alternative for you. But generally speaking, Step Functions provides a lot more functionality to that in terms of both length of the workflows and the complexity of things that you can build in. So in most cases you wouldn't want to go from Step Functions to this.

But it really is another option for developers to use when they're figuring out the right sequence of events for their Lambda functions. And the net effect of all of this is just less code because there's this boilerplate that you talk about, if you build it into one function, by the time you start building out these applications at 20, 30, 50 functions, you've got this duplicate code that you've got appearing everywhere. So if you can take that all out of those functions, it's a huge win in terms of shrinking your code base.

Jeremy: Absolutely. That is certainly something that I'm pushing in 2020 is, you know, do some research, figure out how to fail correctly because that is, there are so many features built in, and there's retries, there's throttling, there's all kinds of things that are automatically built in for you.

I think a lot of things a lot of people don't understand, I've probably mentioned this before, but if you call a function asynchronously, or you invoke a function asynchronously, and that function gets throttled, there's actually a built-in cue that will throttle that for you. So when you talk about using SQSQs and concurrency, function concurrency so that you can do some throttling, so you can reduce back pressure or pressure on downstream services.

But some of that is actually already built-in and if you don't need visibility into those cues, and there might just be temporary times where the concurrency spikes a little bit, some of that stuff is just automatically dealt with for you. So understanding some of that stuff and not arbitrarily putting another SQSQ in front of it when you most likely don't need to, I just think are really interesting things.

Of course you have to know that, which is part of the difficulties of serverless, is sort of understanding how some of these pieces work.

James: Yeah, so the team I'm on has grown from just Chris to seven people now, and so what we're trying to do is just surface some of these things in examples and things we've written. So my friend and colleague Ben Smith wrote a great piece on some of this. How you can figure out the LQs and retry mechanisms through applications. And so we're hoping, as the months go on, we'll start to build out more of these examples to people to make it more obvious.

Jeremy: Awesome. Let's talk about something else, and hopefully you won't get in trouble for answering this question, but why should you never use Provisioned Concurrency?

James: So Provisioned Concurrency, as a feature, is pretty interesting engineering. Behind the scene, how it's been built is pretty extraordinary. In the general serverless space, we like the idea of on-demand lambdas and mostly we focus our time on how do we improve the performance of those all the time. But there's definitely a subset of cases where you have this requirement for close to zero latency. And so you find that there are some of these cases where there's an enormous burst of traffic at a given time of day.

And the scale is so enormous that someone needs 15, 20,000 functions to run immediately. So really this is a great solution for that, and we've already seen since releasing it there are so many people who have used it for exactly that, and it just solves that problem because it solves both the cold-start problem of setting up the execution environment, but also the cold-start problem in your static initializer. So it's a really neat solution for that.

Now where it isn't designed for is just for the everyday lambda use case that generally people use. If you're using asynchronous flows, it's not something that's going to provide you any value, and in many use case it's not something that you'd necessarily want to add to your applications. It's an interesting feature just because where it's necessary it's absolutely necessary. I works really well. But you need to evaluate first whether it's right for your use case.

Jeremy: I totally agree, and I'm joking obviously about never using it, but I think that it's one of those things where we have, where cold starts, right from the beginning when everyone was like cold starts, cold starts, and I started noticing them very minimally when I was doing a lot of user facing stuff. And all of a sudden you get that, like through API Gateway you get a 10-second cold start, but then you take things out of VPCs and of course that problem has gone away too, and you start optimizing your code and you start tweaking some knobs here, and suddenly that cold start is two seconds or whatever.

I get a two-second delay sometimes when I go to load another website, so you know, those aren't cold starts, those are coming just from network latency and some of these other things that I think most people are pretty used to at this point. I mean, I watch my kids, if something's not loading, they just hit the refresh button about 7,000 times. So clearly that's not going to be the limiting factor.

If you were getting them all the time, it would be a huge problem, but I just found that for most applications, the cold-start piece is not a big of a deal. Certainly if you needed to do the pre-warming and some of these other things for having the amount of concurrency available to you, that's a different things, but certainly to solve the cold-start problem, I guess I'm not onboard 100% with that being a necessary solution, I guess.

James: I think with cold starts, it's a complicated problem because I would say 80% of the time when I meet people who are new to serverless and they find the cold-start program, it turns out they just haven't allocated enough memory to the Lambda function.

Jeremy: Right.

James: And so you meet people and they show you how something's taking 6 to 8 seconds, and I'll have a look at their function and you just change the memory and the problem goes away. But there's also lots of other reasons in terms of the code that's being implemented. So I think as people come onto the serverless way of doing things, they start to learn that time matters and the resources you use matter. They start to improve the code and the performance improves overall. And a lot of these issue start to go away by themselves.

Jeremy: Yeah. Oh and also the thing that's great about serverless is that you're not really touching or in control of a lot of that underlying compute. A lot of those optimizations, you know, Chris said this, you basically just make them and implement them and you just start seeing the benefits immediately.

But speaking of speed, so one of the things that has been a complaint for quite some time has been the latency that has been added and the complexity that has been added through API Gateway. And so the new HTTP APIs are out so why are these so much better?

James: This is a feature I really like as well because the API Gateway as a whole service has a lot of extra features that many people are often unaware of, but yeah because it provides all these extra features in terms of managing stages and API keys and DDoS protection, incognito integration, all these other things that, if you're using them all, it's fantastic because you basically pay a fixed price and you get all this extra feature set applied for you.

But in many cases, customers have told us they don't want all those features. They want to have a more vanilla API in front of the servers. And they want to have something that costs less. So this is really for that set of use cases, where you want something that's much more straightforward and slimmed down. And the nice thing about it is that typically we're seeing latency levels much more consistent and lower because it's a smaller service, and people have taken to this because obviously it's over two thirds cheaper. It's a dollar per million transactions and that drops to 90 cents at a certain volume.

So again, it's something where, if you're using all the features of API Gateway, you probably don't care, but if you only need this smaller feature set, this is great because you can use this and save quite a lot of money on your AWS bill along the way.

Jeremy: So what is the use case for HTTP API versus API Gateway because API Gateway has service integrations and, like you said, it has some of the quota management and some of these other things that are there, so what can I do with HTTP API, what can I do with it that I could with API Gateway?

James: For the basic API management set that you would expect, the things that pass through to Lambda functions and interact with your application and you don't need anything beyond that. That is essentially what this is designed for. And you see actually a lot of applications that people have written where, especially for internal applications or things that are just small scale, this is absolutely fine.

But it's really, I still think there's this combination where some people use API Gateway, the original version, with all of these features in place, but many use it without knowing those features are there, so the conversation I frequently have with people is they end up building that functionality inside the Lambda functions, then you show them something like VTL or I could just do it all on API Gateway, so I think it's a good opportunity to look at what feature set each service has and see what fits your application.

You tend to find that one or the other is a very strong fit rather than being a toss-up between the two.

Jeremy: Yeah. I've been using WebSockets quite a bit, which obviously is an API Gateway feature and not a HTPAPI, but certainly there are a lot of use cases where the more straightforward I just need low latency and that routing of course the Lambda Proxy integration, that's sort of how that all works and so it's a very, very cool service.

All right, so another thing I'm going to ask you about is, RDS Proxy came out, so the question is should we just forget about DynamoDB and use RDS again?

James: People have this reaction often where you use one or the other. I'm a huge fan of DynamoDB. I think it's an absolutely essential database for serverless development. The incredibly low latency, the massive scale it provides, and it's just the pure simplicity from a serverless point of view. And I've used it for awhile, so I've learned a lot of the ways you have to construct your application around it. So I'm still very much in DynamoDB camp.

But at the same time, everybody's got a [inaudible 00:33:11] database somewhere, and you speak to enterprise customers who have all their data in RDS, so they have to RDS as a data source for their application. So I think for customers in that position, it makes a lot of sense to use proxy because it just takes a lot of the headache out of managing this connection problem where when your lambda function scales up, you can drain the resources of your database, and we've made some other improvements there in too as well as through security and also fail over speed and other things.

But I think from a conceptual level, this DynamoDB-or-RDS question is one that you and I will be talking about for a while because I like some of the alternative ideas where you use both. You know, why not have your DynamoDB database as your operational database for your app and then use streams to push the data to RDS for analytics? And so I think there are lots of interesting other ways of doing things. But for certainly for people who just need to use RDS and not worry about it, the proxy's a great answer for them.

Jeremy: I think the other thing about the RDS proxy, which I really love the idea of because I do agree with you. There are people who have analytic purposes or analytics workloads, I guess, and you need to use something like RDS. You need to use MySQL or [inaudible 00:34:27] or whatever. What I don't like about some of these things, and this is just opinionated on my side, is what I liked about the fact that it was kind of tough to use RDS with serverless was the fact that it kind of forced you into using something that was a little bit more cloud scale.

And so if you were using RDS with a lambda function, and you had a low workload, right? Say you're doing some ETL task with it or you're doing administrative APIs or something like that where it's low interaction, low concurrency, it was never really a problem. Those zombie connections eventually cleaned themselves up. I built that serverless MySQL package that sort of worked really well for those sorts of use cases even if you get to a point where you were using close to your limit for connection.

But now with this what I'm sort of afraid of is that people will be like, oh I can just use my relational database now with lambda functions, and that kind of goes away. But I do agree with you that this hybrid approach is probably the best way to go about it. And I have this in a ton of applications now. I have DynamoDB as the operational database. You can pound against that thing. It will handle as much traffic as you want to throw at it. It will handle the right speed, the right [inaudible 00:35:53] you put on, it's amazing.

And then you just have a DynamoDB stream setup, and then you just take that data and you push it into RDS. And what's interesting is what I've been doing lately is using the data API, which again I think people forget exists maybe. But what's great about the data API is it doesn't require the VPC. It doesn't require you lambda function to be in a VPC. So you can have a VPC running and you have your RDS database in that VPC, obviously it has to be aurora serverless to use the data API.

But you just take that data off that stream and then you just use the data API from a function that is not in a VPC, right, so you don't have to worry about configuring that or whatever, and then you just push that data over there. So I really like that combination of things. Now granted, I see your point. There are many people who are on RDS and need to use relational databases to do it.

I still think that even though RDS Proxy is going to handle the connection issue, still think you're going to run into scale problems at some point.

James: There's a couple of things that I was thinking about recently. One is it comes down to choosing the right database for the right reason.

Jeremy: Also true.

James: It's been so easy to spin up MySQL for database for so long. It's becomes almost a habit that you lean on the database's capabilities because it handles so much for you in terms of multi-threading and developing large applications. And so now we have all these other tools available. I think you have to reevaluate. Are you using RDS for the right reason? And it opens up this broader question of where should your data live in a serverless application because now we have all of these solutions. You know, the S3-based ones, RDS, MySQL, and everything in between.

And so as developers we actually have a more complicated choice now about where the data goes and how I should manage it. But I think overall if you make the right choices, it gives you more resilient applications.

Jeremy: Yeah, and I think you make a good point about the source of truth, like where do you want that source of truth to be? Obviously in something like MySQL, you can export that, you can move that to other places fairly easily. It's not quite as easy to export data out of DynamoDB. I mean you can just run scan operations and you can do that, but I really like the idea of having that data in DynamoDB as that sort of source of truth.

But I agree with you on that. Choosing a purposeful database, or purpose-built database is certainly a smart move. But anyways, we could probably talk about the DynamoDB versus RDS for quite sometime. I mean obviously I'm also solidly in the DynamoDB camp, but I do greatly appreciate and have always loved the ability to write queries in SQL as it's quite easy.

Although it is quite easy to write them in Athena and so if you're pushing data into S3 and now with the new DynamoDB connector for Athena, you can actually query DynamoDB directly using Athena, which is just using regular SQL syntax as well.

James: With DynamoDB I think it's one of the greatest things I did as a developer before I joined AWS was taking the time to learn how it worked. A couple of times I almost gave up because the model is so completely different. But in recent months you've seen all of these new tools coming with DynamoDB like the Workbench, there's a lot of materials coming through that show how to use things. So I think it's easier to pick up now.

There's a Rick Houlihan video that's become legendary at this point for training on DynamoDB, but once you click and you realize how it works and you see it work at scale through applications, it really is just an amazing service to have in your toolbox.

Jeremy: Yes. And speaking of toolboxes, I do have the DynamoDB toolbox that I'm working on, which again just makes writing data, right now it's mostly focused on writing data to DynamoDB, but that is actually one of the complex things that you have to deal with is the fact that you have a different type of query syntax in order to pull data from it and also in order to do these complex updates. You know, sort of the put items is simple, but then when you do the update item, there's a bunch of syntax things there, so that's actually what my project does around that.

All right, so there have been some changes with some of the run times. We went to No. 10 then we went to No. 12. Those all seem to be pretty stable. Everything sort of worked out there. I've been pretty happy with the performance around 10 and I started using 12, and that's great. But there are some things changing with some of the SDKs and there's something changing with the Python SDK that's sort of important, right?

James: Yeah, so what's happening is that we've changed the way that boto core works so that the request module is no longer part of that. And we've unvendored it, enables some additional flexibility in the way that boto core operates. Now from the point of view of using Python and lambda, what this means is is you're already bundling your version of the SDK into your function, which is the best practice and keep doing that please, you don't need to do anything. That works just fine.

If you're not doing that, if you're relying upon the included version execution environment, when that changes, you will have to make some changes too. So what we've done is we've published some layers you can use that just give you the option of continuing to use that request module that you want to use within your function.

I've just written a blog post about this that went out on the AWS Compute blog that gives you step-by-step instructions, but we just wanted to make sure everybody who's using Python in lambda who's relying upon the request module is aware of these changes that are coming up.

Jeremy: Awesome. All right, so, last thing, then. 2020 serverless, what are your general thoughts? Is this going to be the year?

James: Yeah, it's really snowballing in terms of popularity and certainly seeing just the sheer number of people from all these different companies. You have startups and enterprises and so many different types of industry all starting to pick up serverless tools. And a lot of things that we talked about just a year ago, that really seem an incredibly long time ago now, the conversations that don't really necessarily matter that much anymore.

There was a discussion about what is serverless and all these sorts of things. And now we're starting to talk about architectural patterns, and starting to talk how it's not just lambda anymore. Serverless is this concept of taking different services from different providers and combining them. So I think, you know, we see people building things where you connect API Gateway, DynamoDB, S3, but also with services like Stripe or with [Orsero 00:42:41] and then lambda is just connecting things in the middle.

There's just a lot changing in the way people are building very sophisticated applications at scale, and I think it's finally gotten to that tipping point where it's becoming generally adopted.

Jeremy: Yeah, that's awesome. All right, so, James, thank you so much for being here. If people want to get ahold of you, how do they do that?

James: So I'm available on Twitter at @jbesw or I'm on LinkedIn, people often send me questions on there. If you look up my name James Beswick. And I'm also available through email at jbeswick that's B-E-S-W-I-C-K @amazon.com, and I'm on Slack and everything in between, but essentially anytime you send me a message, I'll do my best to get back to you as quickly as possible.

Jeremy: Awesome. All right well I will get all that into the show notes. Thanks again, James.

James: Great. Thanks so much, Jeremy. Take care.

View Details

Please visit our EPISODES page for links to the full episodes.

View Details

About Nader Dabit:

Nader Dabit is a Developer Advocate at AWS Mobile working with projects like AWS AppSync and AWS Amplify. He is also the author of React Native in Action, & the editor of React Native Training & OpenGraphQL.

  • Twitter: @dabit3
  • Twitter: @AWSAmplify
  • AWS Amplify: aws.amazon.com/amplify
  • Blog: dev.to/dabit3
  • Github: github.com/dabit3

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week I'm chatting with, Nader Dabit. Hi, Nader, thanks for joining me.

Nader: Hey, thanks for having me.

Jeremy: You are a Senior Developer Advocate at Amazon Web Services. Why don't you tell the listeners a bit about yourself and your background, and what you do as a Senior Developer Advocate?

Nader: Yeah, sure. Before I joined AWS, I was basically a front-end engineer, mainly a mobile engineer for the last, I guess four or five years before joining AWS. I kind of come from a traditionally front-end background, but the team that I work on is the mobile team, but we cover Amplify, we cover AppSync, we also cover Device Farm and the Amplify Console. And yeah, we have a couple of developer advocates, I'm one of them. And our role is very kind of lenient in the sense that we don't really have a traditional role as someone might think of maybe a developer evangelist or something.

I think it's really team dependent on what that role actually means. But to our manager, it's a way for us to have a lot of leeway in what we do, so we can write code. Most of the stuff we do is open source so we can contribute to the open source, we can speak, we can write docs, we can write blog posts. Whatever we feel is going to contribute the most to moving everything that we're working on forward, we're able to attack that and work with that.

Jeremy: Awesome. So, speaking of things that you're trying to move forward, you mentioned AWS Amplify. Which is this really cool project that Amazon is working on. Why don't you give the listeners a 30,000 foot overview of what exactly that is?

Nader: Sure. The Amplify was first, I guess, introduced as a client SDK for web and for React Native that basically allowed you to interact with things like API Gateway, things like AWS AppSync, Cognito, much easier I guess, than some of the old way. Before you were using probably the AWS JavaScript SDK, we just added improvements that were really meant for interacting with these services from client apps. The Client was first introduced, that started game gaining steam pretty quickly.

We then introduced the CLI at the... I think the next reinvents. I think it was actually, I'm not sure exactly when the CLI was released, but it was really after to the Client. And the CLI is something that basically allows you to create AWS resources in a similar fashion as you would do with something like CloudFormation or SAM or even something like the Serverless Framework.

But it gives just a different approach, so instead of having to maybe do it in the way that you're used to doing it, maybe writing some CloudFormation or maybe writing some templates with JSON or YMAL, you can just go to the command line and create an update categories versus kind of having to know what's on with AWS.

If you're coming to AWS as a newcomer, it makes a little more sense based on the feedback that we've gotten to use Amplify, because they can say, "Hey, I want an API," and in the background we'll spin up an API Gateway and point with some configuration around a proxy to pass the event into a Lambda and we'll also generate the Lambda. It's kind of an easy entry point for people, but it also is a very helpful way to generate a couple of things at once that kind of tie together, so the CLI is another part of it.

Then there's the Console, which is something that was introduced, that re:Invent 2018, and the Console is a hosting and CI/CD platform that allows you to just kind of connect to a get-repo, and then we do the build and we deploy to CloudFront with S3. It's a really nice way to deploy your web apps. We also have a lot of stuff that's been added over the last year to improve that. I would say that's the main focus, those three things, the Command Line Interface, the hosting platform and the client libraries.

Jeremy: The purpose though of these three tools sort of working together is to build mobile applications, or web applications using something like React Native or just React, or actually view an angular, like there's plugins for all that. The point though is that rather than you having to go out and use something like SAM or build all these things out individually with CloudFormation, that this is just sort of giving you a unified way to create both the front-end and the backend, right?

Nader: Yeah. I mean, we have a lot of people that are actually... there's two ways to kind of look at it, I guess. You can use the client only and still use CloudFormation and SAM and Serverless Framework and we have a ton of people actually doing that. Or you can kind of buy into the whole framework and then use the CLI as your resource creation platform of choice. You can kind of go at it both ways.

But yeah, the idea though is to build these full stack applications and to kind of, we have a couple of main focus points. One of them is around developer experience and the other is just around developer velocity. And I guess that could, maybe tied to developer experience. We want to be able to allow you to create things, and configure things, and deploy and try things out, experiment quicker than maybe you would have been able to in the past.

Jeremy: I mean, I get that you can sort of break these things up and you can use them individually. But I mean, really what you're building is sort of a philosophy around a full stack development, right?

Nader: Right, right. We think that what we're doing is a little different than anything that's kind of been out there before, I think. And we don't really have something to compare it to, but we talk about it in a couple of different ways. One of the things that we talk about is this idea of a full stack serverless development, where you're a developer, or you're a team, or you're a startup, or you're a company and you want to be able to enable a developer or a team of developers to build the front-end and the back end, versus having the traditional maybe engineering team where you have a backend developer and then you have a front-end developer.

We're looking at it like, what if a developer could just be looked at as a full stack developer like we've seen forever. But instead of the traditional full stack developer where the backend developer might be in charge of creating servers and creating a database and patching and dealing with all of the different backend resources, we could take the serverless philosophy, use that and then apply the front-end developers and merge that together and enable a single developer to build out these full stack apps, or a team.

Jeremy: Yeah. And I definitely love that idea because I do think there's a huge growth in front-end developers that need the capabilities of the backend, whether it's simple crowd apps or something like that, or something more complex. But serverless has always to me, seemed like a really easy entry point into that backend or full stack experience. But it's still hard, right? There's a lot of stuff that you have to do.

I mean, especially if you're doing CloudFormation, even if you're using something like the Serverless Framework, which makes these things really easy, there's still some barriers there and some things that you need to learn. What I really like about the way that the framework works is that you use CLI, you can create these resources, you create the front-end, create the backend. Those things are sort of tied together, but you definitely take a very opinionated approach to it.

Nader: Right, right. We definitely do. I mean, for us to be doing the amount of things that we are doing and for us to be enabling some of this stuff, we do take an opinionated approach. But when you're building with Amplify on the CLI, all we're really doing is generating CloudFormation. All that CloudFormation is available in this Amplify folder and we keep two versions of your cloud configuration. We have kind of a development version and the deployed version, that's kind of what your current backend looks like in the cloud, what it's looking like in your account.

We have the backend and we have the current cloud backend. In the backend folder you're doing your development, you're deploying. And then once the deployment is successful, we're tying that into the current cloud backend folder. You a dev and a deployed version, but really all we're doing in those folders is generating that CloudFormation. We're not really, I guess doing anything new, we're not creating any new services as far as the CLI is concerned. We're just kind of using the existing things that are out there. Just a new abstraction on top of them.

Jeremy: And I think when you're doing the level of abstraction that you are doing, you have to take an opinionated approach. But there's ways to eject. You said you're just, let's say I go through this process, I'm a new dev or maybe I'm an experienced dev and I just want to get something up running or something up and running very, very quickly. I can use the Amplify framework in the CLI, use the Console, it all works together get something up and running. But then if I have to go above and beyond that, if there's something that the framework doesn't do, how do I keep what I've done but then keep building on top of that?

Nader: I guess there's a couple of ways to do that. Of course, Amplify being just a abstraction on top of an existing abstraction is going to be either... because also we're fairly new, we're going to be a little bit behind or a lot bond depending on what you're trying to build the current feature set of something like confirmation. Once you start building with Amplify, you might run into a situation where you need something that we just don't offer.

You can either create your own CloudFormation within the Amplify project, then you can run Amplify Push, and this will kind of allow you to take what the CLI offers you and add to that maybe your existing knowledge of CloudFormation. Or you can just eject completely out of the Amplify workflow, and if you have something and you need to, I don't know, maybe take this to the next level or whatever you see Amplify limiting you for whatever reason, you can completely just take that and move that to whatever other framework that you'd like to use.

Jeremy: Going back to the opinionated approach too, I mean, in terms of what the roadmap looks like for this, because obviously you're adding a lot of features and we can talk a little bit more about that later. But is this something where you see, just keep adding feature after feature or is it something where you feel there is a target market for this, and that as long as you cover a certain percentage of use cases that probably is going to be the philosophy that you stick with?

Nader: Definitely target market, there is a target market for this. We're growing super quickly, I think we had a 400% increase in community contributions around content this year, 2019. Super big increase in downloads and usage and all that stuff. And we're seeing a couple of different types of developers, where we're seeing the traditional AWS developer, someone that's been using and building with AWS for a while that is building maybe these full stack apps.

They're not just building the infrastructure, they're actually building a mobile app or a web app. To them this is the easiest or the fastest way to do certain things, so they're taking that approach versus... or maybe they're just adding it to their toolbox. This is just another option they have. And you look at the app that you're going to build and you decide, "Should I use Serverless Framework?" These are all great options, I think.

The other developer that we're actually really excited about, it's kind of this new generation of cloud developers that we are seeing enter the space via Amplify. A front-end developer or even a developer that's using something that you could kind of put in the bucket of competitor for Amplify that was scared to use AWS is coming to Amplify. Seeing that we have this great user experience and seeing how easy it is to add authentication and API and kind of deploy this on this scalable infrastructure is something they've never been able to do before. They get super excited about it, and then they tell their friends.

We're seeing this new generation of traditionally front-end developers kind of moving towards the framework. We're seeing the traditional AWS developers add it to their toolbox. And then startups are really excited about it. We just did a big week of events and in New York, a three day Amplify week. Busier, we had more people show up there's standing room only. A lot of startups where there, they're really looking for the most efficient way to build in this day and age of expensive developers and also the average.

If you're a startup and you want to go and hire a mobile developer and then you want to hire a cloud developer, you might want to hire a dev ops. You start adding all these things, starts adding up. What would we have today if you kind of choose the right abstractions? You have things like React Native on the front-end where you can write a single code base and deploy across iOS, Android and maybe even web. And then you look at something like Amplify, and if that front-end developer or that mobile developer can learn how to use Amplify, that might be enough for them to actually ship Version 1 without having to spend a million bucks, right?

Jeremy: Right.

Nader: The startups are super excited about it because of the efficiency and the cost savings of what they're able to build, and they're building on the same infrastructure that Netflix is building on, so that excites them as well. Because if they're like, "Oh, I don't have to just build a flimsy V1 or V0 that we have to rewrite, we're actually building on something that can scale. I think those are the three biggest buckets of developers that are pretty excited about Amplify.

Jeremy: And I love the sort of serverless for startups concept too because it's just one of those things where you're like you build that first version, you can always add onto it, but you just immediately have that scale. You have that reliability back there. And then again, not to hammer on the opinionated thing, but there are a lot of decisions that you need to make even with serverless when you're planning your backends and figuring out exactly how they're going to work. And you still have to think about scaling and think about how your systems are going to behave underload even if the backend scales for you.

There's still some decisions you need to make around there and Amplify basically makes most of those decisions for you. Which is I think really handy, especially for people who are not cloud experts, because that in and of itself might need to be a PhD just to figure out how some of these things need to work. I think that's great. You mentioned this idea of sort of there's on ramp to building these tools for production that they can get that view on out there. Are you seeing a lot of people using it to build PLCs or are they building Version 1 that actually goes into production and is out there as a full on production application?

Nader: Oh, totally both. I mean, we definitely see both. I would say the people that would be, if you're a massive corporation or you're a big company and you have a big team of established engineers, you might use Amplify to build that prototype and then move to something that is along the lines of you're basically existing development process, I guess. But if you're any company I guess, and you kind of do want to build something, there's really no reason to I think at this point, eject out of Amplify for any reason unless you really find the current way to get around some of the things that we don't support too cumbersome or whatever.

But I mean, going back to this weekend in New York, we had about a half dozen companies that had already shipped their Version 1, and they were there to talk about what their next steps are and then maybe just show us their apps. Like a really, really cool dating app that was just... they have a dozen people or so at that company and they're using Amplify and they've already shipped V1. It was was really cool to see to that type of stuff. We're definitely seeing a little bit of both.

Jeremy: That's amazing. All right, so let's get into the framework itself. All these things work together. You mentioned the framework and the CLI and the Console itself. But just the framework, the development experience. I'd like to talk about what that workflow is sort of the tool chains that you have in place. Maybe we start with sort of the basic thing here. You go to the Console, you type in Amplify, right? Just to get started, right? How do you get started with this?

Nader: There's a couple of different ways. If you do go to the AWS Console and you type in Amplify, we have a feature, I'm sorry, we have a service I guess in the Console and you'll have two different links that you can click there. One is going to take you to the Amplify Console, which is the service, but the other link is actually going to take you to our open source framework landing page, which has all of this stuff around the open source projects, which are the Client and the CLI. To get started you could go to the AWS Console, you'll find us there. You could just google it and the first thing that pops up is going to be the framework landing page for the open source.

If someone's just getting started, typically they'll find us through whatever means of a blog post, or maybe a reference, or just searching I guess on Google. And typically they'll then decide which platform that they want to see the information for their docs. And we really have four main buckets I guess, you could say of developers. We have as far as the platform on the front-end is concerned, we have native iOS, native Android. We just released new clients Amplify, native Amplify for iOS and Android on the Client, was just released actually in re:Invent this year. We're excited about that.

There's a native iOS, native Android, there's cross-platform, which is kind of React Native and we're also investigating other frameworks like Flutter. And then we have just web, and the web is growing crazy too because our team is AWS Mobile, but when we introduced this framework with JavaScript support, we thought React Native was going to be the fastest growing segment there, but web has actually been the fastest growing. And React Native is growing too, but just the sheer number of web developers just outnumbers that number of React Native developers, so it's kind of crazy.

React view, Angular, Ionic, all of those are the different frameworks that we support. We're looking at spells hopefully in the next, in the near future. Typically, you would kind of choose which framework you're interested in. And then we have a tutorial to get started where you create the Amplify project within whatever framework that you're working with, say for instance, React Native. Then you can add authentication and we have this two-liner that allows you to add a real production ready authentication flow in your app and two lines of code.

And that usually is really the light bulb moment where people are like, "Wow, this is awesome. We're excited about it, let's learn more." And then they'll go from there to maybe a Lambda function and an API Gateway. Maybe they'll look at GraphQL without sync or maybe a S3 with storage, just kind of playing around from there. Or they'll investigate based on the app that they want to build. Like, do we have the feature set that they want, or not?

Jeremy: Right. That was a lot of information. Basically, really the best way to get started is download the CLI, is what you're saying, right? That's really what you want to do the most work with?

Nader: I guess there's two types of, if you want to get the whole experience, yeah, do the CLI. If you already have some AWS stuff you want to connect to, just look at the Client docs.

Jeremy: All right. Then you've got the CLI, you said you pick your sort of the front-end of choice, whether that be React Native, or Web or something like that. And by the way, just in terms of, I can see Web being very popular because even a bank that I use got rid of their mobile app in favor of just using a mobile site. And I think it's just because they didn't want to maintain all those different versions. I do think that, especially with iPhones and Android and things like that, that the Web, the mobile web versions-

Nader: The web platforms

Jeremy: Yeah, exactly, will be very popular. All right, you start that, you get that sort of going there, you mentioned the components, right? Let's talk about components for a second. Because this is kind of cool. You and I talked about this in the past. Some of these components just add sort of features to the front-end that allow you to connect to backend resources. But some of these actually create, like you add them to your project and it actually will create front-end and backend resources, right?

Nader: Well it's kind of, we have this category or we have this set of components that are UI components, user interface components. And for those we have framework support. We have just the raw JavaScript library, which means you can just use this with any JavaScript project. And then in addition to that, you can install additional framework components. If you're using React or Vue, then you can NPM install these additional components. And yeah, they're kind of a hybrid between front and backend in the sense that they are going to scaffold out some UI for you. We're going to scaffold out for the authentication for example, a sign in, sign out, sign up flow along with things around multifactor authentication and all that stuff.

Nader: And from there you configure that, but it is opinionated in the sense that we're just assuming when you're using this component that we're going to be communicating with the Cognito backend that you've configured locally in your project. But yeah, we have those framework specific components for things like authentication, chatbots, images, feeder galleries and stuff. And they're a good entry point for people just getting started because you can just write a couple of lines of code and see something that looks really nice and that is ready to roll.

Jeremy: All right. Then you have sort of a tool chain here as well where you add those category components and then you just type ad hosting and it will just add it to the Console, or add it to the in fly Console.

Nader: Yeah. The tool chain is its own category within the CLI. We talk about the CLI either as a CLI or as a "CLI tool." And yeah, it does quite a few things. I mean, for instance, with GraphQL, if you've ever written GraphQL and you've called a GraphQL API from your Client application or from anywhere, you have to have the GraphQL operation definitions. And with the rest you call a rest endpoint, but with GraphQL, you pass in this AST, which is the GraphQL AST which contains the GraphQL operation.

You typically would have to just write all that code yourself. But the CLI tool chain, one of the things that it does, it'll look inside of your GraphQL schema and then it'll just generate all that code for you and then we'll create a folder and then drop it in. That's one of the things we call that GraphQL cogeneration. We also do Lambda function generation. If you want to have a Lambda trigger for S3 or for Cognito or whatever, we can give you a boilerplate to start with.

We also have a lot of other Lambda boilerplate generation for things like an express API running on Lambda or maybe even a DynamoDB CRUD app. If you want to get started with that, there's just boilerplate that you typically have to write over and over. And we're kind of giving you a starting point and then from there you can update that. And that's mainly, I would say what the CLI tool chain is there for.

Jeremy: Awesome. And then one of the other library components that was added recently was sort of this AI/ML that does sentiment analysis?

Nader: Right, right. We added the support both on the backend on the CLI, I guess, and on the Client for this new category that we call it Predictions. And Predictions is pretty cool because it basically just is taking advantage of all of the powerful stuff that's already there in AWS. And it wasn't a lot of work to get it working for us because all we were basically doing is kind of, if you have an identity pool within your project or if you want to create one, of course you could just run Amplify at auth and create one.

All we're doing is granting permissions to interact with those different AI and ML services within that identity pool. That part was pretty simple. And then on the Client we have added a couple of different methods I guess you could say around this new Predictions category. And from there you can work with recognition, you can work with Polly, you can work with Transcribe, you can work with many of the AWS or Amazon, forgot, depending on which service it is that the Amazon or AWS part, but yeah, all of these different ML and AI categories.

And I demoed this for the first time at an event in New York, actually this week again back to that event. And people were just blown away and it's been one of our fastest growing categories too. Because, it's just really simple to do and a lot of people are interested in this stuff. People are building apps around medical data and stuff with a lot of the text transcription and reading text off of a documents. People are pretty excited about this entry point because it's kind of easier in my opinion, to get up and running with than if you walk through the AWS Console and try to figure a way to actually integrate this into your app.

Jeremy: Well, I mean, it just goes back to this idea of making it super simple and easy to do this integration because, again, it's not hard necessarily to set up these things via the Console. But if you want to put the infrastructure as code and get the right CloudFormation in place and then connect everything together, and then you've got the UI piece of it that ties in, I mean, it's just a much better experience.

Let's talk a little bit more about some of these other features that are being added because your team, or the mobile team there is adding features to Amplify at an incredible pace. And one of the coolest things that came out lately was at re:Invent you announced the Amplify DataStore. I spoke with Chris Burns about this and I think, I forget his exact quote, but I think he said it was just awesome, but that I should talk to somebody else to give me the details on it. I'm talking to you, so tell us about the Amplified DataStore because this is pretty cool.

Nader: It's the culmination of almost two years of work from our team. And Richard Threlkeld, and Michael Herbenick, and Manuel Iglesias, a bunch of the people on the team have been either thinking about or talking about or working on this. Those are the ones that actually do the actual real work, by the way. And if you're on Twitter, it'd be worth to follow if you're interested in this stuff because they talk about this stuff.

DataStore is kind of, when we first launched the GraphQL support for AppSync on the Client, we had a library out there that allowed you to interact with the AppSync. And we ended up having two available options actually. We had one from Amplify and then we had one from a separate SDK that was kind of a fork of the Apollo SDK, which is a popular GraphQL library.

One of the things that Apollo had built in was this cache. And the Apollo Cache allowed you to work with some of that data in memory locally. Instead of having to query from the backend, you could query from that cache. And then we enhanced that by adding offline support. But the limitations that we started running into with that was that we had a lot of developers that started treating the cache as a store, and they wanted to do more complex queries on the store. They wanted to do things like query-based on some type of predicate. They wanted to maybe have more complex stuff than we could basically support, because we were working with someone else's implementation that we didn't really have control over.

Also, a couple of times there were breaking changes and stuff that they would implement that we would have to kind of go back and figure out what was going on, and it just wasn't the best experience. The Amplify DataStore is the next generation or the next version of that, but it has just a bunch of new features that we didn't have before. And I think the idea around it is, it's basically a single source of truth. It's a store that you write to locally from your front-end app. If you're on React, you would just write to the store and you don't have to actually from then worry about writing to your backend, the backend and the Client all sync together from this DataStore.

The way it works, I guess, if you talking about from the very start, if you're a user the way it would look like this you, you come online with your mobile or web app and from there it's kind of, imagine a new build or a new view of the app. The user has no data right on their app, so we make that initial fetch to data store via DataStore to whatever database you're working with. We get that initial set of data. Maybe you get like a thousand rows of data from DynamoDB. We keep the last sync time stamp on the Client at that point. We know when you fetch that data from there, we then have an observer that's set up that will then get any new data that's sent to that database and send it back to the Client.

You have that new data coming in. It's kind of like if you've ever worked with GraphQL, a GraphQL subscription running in the background, you could think of that. Or maybe if you've worked with Pub/Sub or something like that, WebSocket. But that's kind of all set up for you. Whenever anyone makes updates, that data comes back and it's stored for you locally. Also if you go offline and then come back online, we use that last sync time stamp to then fetch from what we kind of call it the delta table or the change table. And instead of fetching the entire data set, we actually just fetch the death between when you went off long last and went and when you came on last.

When we fetched that small set of data, you're getting a more efficient update to your local device. But also when you're offline you can actually write to your store locally and everything works. We make the optimistic update in the DataStore, you can have all of that updated data. Like if you've ever used Twitter or Facebook or even Instagram offline, you can still like stuff and all of that. It feels right, it feels like it's working, it's not going to break. And then when you come back online we send that new data back to the data database via the DataStore and you have the conflict resolution and the conflicts detection also built in, and we do a lot of that stuff.

We have a couple of different ways of doing conflict resolution. That you have the option of implementing, we have auto merge is the default and that's kind of the new conflict detection that we built in. And we kind of are taking the GraphQL schema and then the types that are whatever data that you're sending up. And we have a really enhanced version of conflict resolution that will make sure that even if two people write the same piece of data to the data database, maybe there's a field or there's a type that has three different fields and I change one field and someone else changed another field, instead of kind of like taking the last rider wins approach, we'll kind of merge those two objects. And that way the data is still up to date with both of our changes. But the user doesn't have to do anything and you actually as a developer don't have to write any code to make that work, it just works.

Jeremy: And I think what is really cool about this and also just so it doesn't get lost, is the fact that sort of traditional approach was you would sync your local cache or sync your local store with the database, and even if there's some pushes and things happening. But there was sort of a management locally on the Client where you said, "Okay, well this is in the cache, this is on the web, or this is online.

Nader: Right, right. We keep up with two sets of data, right? The local data and then the data that you have-

Jeremy: Right, right. And what this does is you just interact with the DataStore, right? You don't even care about the web. You just basically say, "Here's the DataStore, here's my source of truth. I'll interact with this and then Amplify will take care of syncing all that data to the cloud. All of my updates will get pushed. All of the updates that'll get pushed to me, that all just automatically happens. I think it's such a huge increase or a much better user experience, only have to deal with that one particular thing.

Nader: Right. That's exactly right. You're totally on point with that. You're writing to the local DataStore. You're no longer having to deal with interacting with the API yourself. But you still can, for whatever reason if you want to do something on the server or if you still want to do something on the client, you can. And then another thing that's kind of a big part of this is that we've added this query language to the data store itself.

You can call queries based on different predicate. You could say, "I would like to query the data based on, if this field is as this, or if this field is that." And you can actually chain different arguments into that query. You could say, "I want to query all of the to dos that are done, that are created by this person or something like that. And now you're just querying from that local DataStore and you're not actually having to send multiple requests to your backend. You're actually maybe, saying, of course the latency is going to be much faster. But even you might even save some money because you're sending less requests.

Jeremy: All right. If somebody took a look at Amplify maybe a year ago, there was a lot of cool features in there. I think the project had a ton of promise back then, but over the course of the last year or so, over the course of probably just the last few months so much has been added to it beyond just the DataStore. And these were things like the ability to do local mocking, more authorization support. You added the ability to do or to connect to Aurora Serverless, simplified all apps, that whole Delta deploys.

There's just been a ton of different things, instant cache validation, and I think that all has to do with really there being a good merger between, or a unification between the framework and the Console, which I think was a little bit disconnected a year ago or so. And that has kind of come together. Are there any of these updates that have happened over the last couple of months or so that you think really if people looked at it in the past are things that would be really game changers for them to look at the framework again?

Nader: Totally, the multiple authorization types, the local mocking and testing, the backend visitability and the Console, the machine learning stuff that we've added. A lot of these things are the things that I hear a lot of good feedback from our customers. But our roadmap is pretty much set by what people say that they want. We're not saying, "Oh, this would be cool to add or that would be cool to add." We're actually being like, "Okay, let's look today what people ask for, and tomorrow." And we're taking all of this feedback and we're prioritizing our roadmap based on what people ask for.

We're customer driven. And I know you always hear about AWS and Amazon being customer focused or whatever. I mean our team, that's all we really are. That's the main thing that care about, it's what people are asking for. Of course, we're always investigating other improvements that we feel our engineering team as engineers would think would be additional enhancements, but everything that we've basically done is based on that feedback.

A lot of the stuff that we do release is super, super popular and we get a lot of good feedback. Just because we've had people say that they wanted that stuff anyway, we just go and build it. We have a pretty transparent process around it, which is pretty different than... I think I've actually seen some parts of AWS now doing this. But we have our process open in the GitHub repo. You can look for RFCs or you can create your own RFC. And there you can actually talk about what you want and you can be in the discussion around what we're actually about to add. If you see something or if you have an idea, you can actually maybe get it implemented. It's pretty cool and a lot of people will get excited when they're in that discussion.

Jeremy: That's awesome. All right, I want to move on to a slightly different topic because you do a lot of work with GraphQL, obviously. And I'm a huge fan of the one table design strategy or single table design strategy for DynamoDB. And you wrote a post recently where you took, I think it was the example that was on the best practices article or whatever it is on the AWS site. And there was actually 17 different access patterns that you basically showed people how to do that in GraphQL, straight GraphQL, right? Like Lambda resolvers obviously I think people like to fall back on that. Sometimes they might think it's a little bit easier, but you set this whole thing up all through GraphQL. Do you want to tell us a little bit about that?

Nader: Yeah. And this is another feature that we added. That post was enabled by another couple of new features that we just added this year. Via the GraphQL Transform library, which is part of the Amplify CLI, you could say, you have the ability to use custom indexes and define custom keys and indexes using the graph Guild Transform library. And I think one of the things that a lot of people coming to Amplify have as a question or concern is, when they're done with their hello world project or when they're actually building their production app, they're stuck because they don't know how to take what they have and add a lot of this more sophisticated data access patterns that they might want.

And that post was just a direct response in addressing of some of those questions, some of those asked. And we put that also in the docs, so it's in the docs now, you can kind of go there and see that. If you have a business idea and you're wondering, "How can I model my database, and how can I leverage Amplify at the same time? How can I actually build this? Going and reading through that might give you some good ideas. Even if you're not using the exact the data set or whatever that we're kind of talking about there, of course, it's probably going to be something different, but you can use the ideas there and probably get some good ideas.

Jeremy: That's awesome. All right, well listen, Nader. It was great to have you here. Thank you so much for joining me. If people want to find out more about you and Amplify, how do they do that?

Nader: Yes, thank you so much for having me. Actually, I'm a huge fan of this podcast and I've looked up to pretty much all the people that you've had on here so far. It's been pretty cool to be here. If you want to find me on the social media, @dabit3, D-A-B-I-T and the number three. And then if you want to follow Amplify, we're also on Twitter @AWSAmplify. Definitely check us out if you're interested in this stuff.

Jeremy: And the Amplify site on AWS is aws.amazon.com/amplify, right?

Nader: Right.

Jeremy: And you have your blog, we mentioned your blog. What's your blog if people want to read all those articles you do?

Nader: Right, right. I used to be on Medium and I still have articles there, dabit3, but I've actually moved my blog to dev.to, so you can get a D-E-V dot T-O. That's dev.to/dabit3. And I'm pretty much dabit3, across all of the different social media and stuff. And also on GitHub, I have 270 repos on GitHub and I'm guessing that maybe there are 70 to a 100 of them are Amplify stuff. So dabit3 on GitHub also.

Jeremy: Awesome. All right, well, I will get all that into the show notes. Thanks again.

Nader: Thanks for having me.

View Details

About Ant Stanley:

Ant is a consultant and community organizer. He founded and currently runs the Serverless User Group in London, is part of the ServerlessDays London organizing team and the global ServerlessDays leadership team. Previously Ant was a co-founder of A Cloud Guru, and was responsible for organizing the first ServerlessConf event in New York in May 2016. Living in London since 2009, Ant's background before Serverless is primarily as a Solution Architect at various organisations, from managed service providers to Tier 1 telecommunications providers. He started his career in 1999 doing Y2K upgrades in his native South Africa, and then spent 5 years being paid to write VB6. His current focus is Serverless, GraphQL and Node.js.

  • Twitter: @IamStan
  • ServerlessDays: serverlessdays.io
  • For organizer information: organise@serverlessdays.io

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Ant Stanley. Hi, Ant. Thanks for joining me.

Ant: Hey Jeremy. It's a pleasure to be here.

Jeremy: You're the co-founder of ServerlessDays Global. Why don't you tell the listeners a little bit about yourself and what ServerlessDays is all about.

Ant: Yes, I helped co-found ServerlessDays in 2017. I've been an early member of the Serverless community. I originally was one of the co-founders of A Cloud Guru and helped get Serverless [Consults 00:00:30] off the ground. After leaving Cloud Guru, I took a year off, worked on a few side projects, then joined up with a few folks here in London, and we decided to get a community-based Serverless conference going. It was supposed to be one conference called JeffConf. Then, it took off and became a thing of its own due to the amazing community. That's pretty much, not quite how we got there, but it's the start of how we got to where we are.

Jeremy: All right. I actually want to talk to you about ServerlessDays. So I helped co-organize ServerlessDays Boston, a crazy event. I went to one in New York, and I've seen, basically, these ones all over the place now. I went to one in Milan. This is becoming a pretty big thing. So, there's all kinds of ways people can get involved. There's some really, really great speakers at these events, but I just want to talk about, really, how this got started. Let's go way back to the beginning, understand what the motivation was behind it. Then, let's talk about some of the events that are happening around the world and, then maybe, how people can get involved. Why don't we start with that? What's the history of this whole thing?

Ant: The history, it goes back to April, May 2017. There was due to be a Serverless conference in Amsterdam, run by the then organizers of the Serverless user group in Amsterdam, and it, kind of, fell apart. I think end of April, beginning of May, it got canceled. I don't think they could raise enough sponsorship funds. I think they were trying to go too big, and, at the time, that was going to be the only Serverless conference in Europe that year. So at the time, I ran the... Well, I still do... run the Serverless user group in London, which is the largest Serverless user group in the world at this point in time. I had a conversation with Paul Johnston. He used to work for AWS and he's one of the early the early Serverless bloggers or contributors, and James Thomas is a Developer Advocate for IBM, on their OpenWhisk functions platform. He's also London-based.

The three of us had a conversation via a Slack channel. I'm saying, "Well, there isn't anything happening in Europe this year. Why don't we try and organize something?" What became an idea, started to become reality, and Paul popped up, and he said, "I might have a venue that's really cheap." So I said, "Well, I've got a user group with a whole bunch of users, and we don't have anything planned in the summer because that's normally an awful time to run a user group cleanup. So, I said, "Well, let's try and run an event." We decided to call it JeffConf, based on a very bad joke, because of the name Serverless. The in-joke, at the time, was we could've called Serverless anything. We might as well have called it Jeff. So, as a joke, we decided to call this thing JeffConf.

We organized it in six weeks from the point of saying, "Yes, let's do this," to actually running the event. It was a six-week window. We didn't run a CFP. We ran on an absolute shoestring. We spoke to whoever we could. Companies jumped in to sponsor. So the first tweet we put up about it, Chris Munns, from AWS jumped all over it and said, "Hey, can we sponsor?" IBM got involved. A few other companies, local London agencies, also got involved.

Yeah, We managed to get off the ground. We had some great speakers. We had Simon [Woodley 00:04:01], that I've been trying to get into my user group for ages. I managed to convince him to come to London for the day, and he gave our opening keynote. We basically managed to cobble together a great among of local, predominately, London/UK-based speakers to come speak, and it worked. We had about 170 people attend, which wasn't bad for such a short period of time. We somehow made a tiny profit on it because we managed to get a venue, which was the St John's Church in Hoxton, 196-year-old church, where Paul was friends with the pastor who runs it. So we ran it in a 196-year-old church in the middle of Hoxton Shoreditch area, which is the heart of London's tech scene. We had some great speakers and basically had a beautiful day, and it was a great day out.

That was the first JeffConf, and we didn't really think we would go further than that, at that point in time.

Jeremy: But it is sort of fitting that the first ServerlessDays was in a church because Serverless is, kind of, a religion, if you think about it, to some people.

Ant: Oh, yeah. Yeah, it definitely is a religion for a lot of people. Yeah, it is, kind of, fitting, and we've made jokes about Serverless dogma and religion, but it is fitting. Ironically, part of the reason it did take off is, when we announced that two Italians got on a plane and helped us out. Alex Calaboni and .... They came over and helped us out. Then, Alex, at the end of it, said, "Hey, I want to run this in Milan." So, it's like, "Well, we didn't have any plans beyond running once, so yeah, you can go ahead. If you want to run it in Milan, go for it."

Two and a half months later, it was the end of August. So first ServerlessDays was in first week of July, first Serverless JeffConf. Then, it was in September of 2017, Alex ran JeffConf from Milan, copied by my awful, awful website that I'd designed. It was a point where I thought rolling my own single-page app framework was a good idea. It's being used for sum total of three websites, which is more than it ever needed to be. So, he put that together. I think he had about 150, or so, attendees the first one. He had a whole month extra to organize it and that was a great event.

Then, Soenke from Hamburg, was one of the speakers of that ServerlessDays at that JeffConf. He approached Justin and says, "Hey, he wants to run this in Hamburg." So, at that point, he said, "Well, if you want to do it, and you want to put the effort in, we'll help you." So, Soenke decided, with some of his colleagues, at the company he'd just co-founded, to run a JeffConf in Hamburg. That turned out to be the last JeffConf because, in the process of organizing this, we all stopped, myself, Paul, James, Alex, Soenke. We all said, "Maybe there's something in this.

Just organically, without trying, we managed to get three of these events in a six-month period. That's when we decided to rebrand, and we spent a lot of time trying to think about what the name should be, and how we should rebrand, and that. Yeah, we announced ServerlessDays as the last talk of JeffConf Hamburg. So, JeffConf Hamburg was the end of JeffConf, and the start of ServerlessDays, as we know it.

Jeremy: So, why ServerlessDays? What was the reasoning behind that?

Ant: We thought about 101 names, because we knew the JeffConf name wouldn't expand. It was an in-joke, and it was too open to misinterpretation. One of the core tenets of every single ServerlessDays is it needs to be representative of the community. JeffConf is very representative of people named Jeff. So we needed a name that didn't exclude a large proportion of the population. Hence, we decided to rename it. Ironically, the actual name... I'm debating... It didn't come to me. It came up when I was having lunch with James Governor, the Monkchips from Twitter. He's a friend of mine. I have lunch with him couple of times a month, and it was debating, "What should we call this thing?" James was like, "Stop mucking it about. Just call it what it is, and call it ServerlessDays.

So, we decided to stop trying to be too clever. Stop trying to cover up with a funny, clever name that some people market. Call it what it is, and that's exactly what it is. It's ServerlessDays. It a day to learn about serverless technology and to engage an expansion on the [inaudible 00:09:20]. That's where the name came from. It was basically telling us to stop being clever and just call it what it is.

Jeremy: Well, yeah. JSConf and DevOpsDays, they're very descriptive of what they are all about. So, yeah, I think that works well.

Ant: Yeah, exactly. Exactly.

Jeremy: All right. So, now you've go three of these in the books. You changed the name to ServerlessDays. Then, it started taking off even more than that.

Ant: Yeah, massively. Our MVP for ServiceDays, so to speak, is another badly-cobbled-together website, which still has not changed, since I put it together. serverlessdays@io. I think I actually pushed it live about two minutes before I went onstage at JeffConf Hamburg and announced it. So, the MVP was the website, and on the website was a link to an email address. It said, basically, "Email us us if you want to run a ServerlessDays." That was, pretty much, it. We have a Slack Journal, where we have a bunch of people who can support and help, and we've got a bunch of documentation from all the various ServerlessDays that we share with new ServerlessDays organizers, sponsored templates, sponsored contract templates, artwork. We basically, through the website, we've got a whole bunch of characters designed to try and give it a bit more of a feel to it and put an email address up. That was enough to get it going and take it international, take it beyond something that got set up by people who had been to a previous one. That's how we got it outside of Europe, essentially.

Jeremy: Yeah, because there's been Portland. There's been Austin. There's been Atlanta, Boston, New York. There's one coming up in Nashville, and across the world, there's two in Japan, this year, I think.

Ant: Yeah. Tokyo was the 22nd of October, and Fukuoka is coming up on the 14th. Yeah, it's after that.

Jeremy: I think you're right. I think it's...

Ant: Yeah, 14th.

Jeremy: Yeah.

Ant: So, this year, there's going to be 19 ServerlessDays. So, we've gone from two in 2017 to 19 in 2019. Conservatively, I think we'll go over 30 ServerlessDays in 2020. Don't hold me to that, but I think we will go to 30. With the amount of inquiries we get, I think we should get to 30.

Jeremy: Well, I think there's 10 in the first quarter, or something like. Right?

Ant: Yeah, there's 10 in the first quarter. I have about approximately a six- seven-month view of what else is being organized. Typically, the runway to organize a server if there's some people starting to organize the actual running is about four months. We don't recommend doing six weeks. We say a minimum four months, ideally, six months. We've had teams that have been working on this for over a year. We do expect a bit of a ramp up.

Then, you've got the teams in Japan that just absolutely hit the ground running. They started to speak to us in July. They ran their first one 22nd of October. They had 450 people added to that.

Jeremy: That's crazy.

Ant: Yeah, so yeah, I think 30 could be achievable this year, which is good growth.

Jeremy: So, a couple of things. Let's talk about just the main goal of ServerlessDays. Obviously, it's to get people to learn about serverless technologies. It's not necessarily any one specific cloud provider. We always have a lot of diversity at these events, from different providers, as well as really trying to have a diversity of speakers and a diverse audience, as well.

Ant: Yeah, that's the core aim of ServerlessDays, is to grow the serverless community. It's to create a community, grow it, and nurture it. It's not an opportunity for vendors to pitch. We actually almost have to coach some of the vendors in terms of what talks they submit to CFP. It's not pay for play, and if you sponsor, you're not guaranteed at all. You still have to go through the CFP. I did have a conversation, once, with a senior individual at a certain card vendor to explain why none of their talks got accepted, and it wasn't their fault. It wasn't. Yeah, the main aim is it's about growing the community and it's not just about growing a community for one sector of the population. We want these... as I said before, one of the core [inaudible 00:14:12] needs to be representative of the community that exists, beyond tech.

So, basically, being blunt, having a room of white guys with a bunch of white guys talking to him doesn't further the aims of ServerlessDays, and it's definitely not what we want. I think we're getting there and achieving our aim, but we can always do more.

Jeremy: Absolutely, and I know that your team has done some work as well, reaching out to other people to try to get things like diversity and inclusion and really pushing those.

Ant: Yeah, I think success varies from region to region. We do try and push it onto the organizers and some of the organizers absolutely take it and run with it. London, we're very lucky in that diversity and inclusion is slowly becoming embedded in the tech community, there, and there are multiple groups that we can work with that are being very supportive. So, we've almost got it easy, compared to other regions where the diversity inclusion efforts are not as mature. But, it's always about improvement. You're never going to be perfect. You're never going to achieve all of our goals, as long we get better year and year. Then, maybe one day, we will get there, but just that constant improvement is what we're aiming for.

Jeremy: Right. So, speaking about the future of where this thing is going, you've recently asked me, and I have agreed because, for some reason, I can't say no to these things, to help with having a United States-based entity along with [Farah Campbell 00:15:49], to help out organize ServerlessDays in the US, and to make it easier for people to start organizing. I know, for me and for the team that I worked with when we did Boston, it wasn't easy getting started. We had to form a legal entity. We actually had one of our board members or one of the organizing members was also going to be one of the sponsors. So they fronted us some money, so that we could pay for an attorney and do some of that stuff. We were very lucky, I think same as you.

With the event space that we found, we were able to do it at the Microsoft NERD Center. Because it was a community event, they donated that space for free to us, which was very, very helpful. There were other costs. We recorded all the videos, so we had to pay for a videographer. The bill for the catering was the biggest we had, I think. We had banners printed. We had the happy hour afterwards, and things like that. So there's a lot of costs that are involved there, and I think it can be very daunting, especially for people who are busy professionals, trying to do this on the side, and help out, that they really can run into a lot of roadblocks. So, what are the plans, here, to make it easier for new organizers to come in and run these events successfully?

Ant: Yeah, I think we definitely do need to make it easier. I think, early on, we didn't really have a framework. We had a bunch of templates from previous events that we could share, a bunch of documentation we could share. When we do these onboarding calls, the typical process was myself or Alex [Castleburny 00:17:32], who runs ServerlessDays Milan, would get an email. One of us would pick it up. One of us would jump on a call with a potential organizer and one of the first things we'd always highlight immediately is, "You need to get a budget." So, give them a budget template. Fill up that budget template. "You need to figure out how much this thing's going to cost you, at a high level.

You need to go put money up to book a venue, and you need to have a legal entity that you can do the legal contracts and all the financial transactions through." Those three things created a barrier to adoption. I'd say only about 40% of people who contact us end up learning ServerlessDays. So, 60% of people don't really even get past that barrier. In some respects, it's good because you test someone's commitment to actually doing it, if they're willing to go through all of that, but the other hand, it should be we do lose events because of that, because not everyone is in a position that they, as a company, can help with cash flow up front. Not everyone is in that position to make these things happen.

So, a big element of what we're going to be doing in 2020 is creating a US entity, which enables us to essentially get a lot of the sponsor money. Particularly, the major sponsor's all US-based. Amazon, Google, Microsoft, Cloudflare, IBM are all US-based and enables us to give them one company that they can pay for sponsorship for the year, and they can give us both sponsorship. And we can just handle that up front. Then, what that will enable us to do is that then gives us bootstrap funds for new ServerlessDays.

So, if you want to organize a ServerlessDays, we can help front some of that money for you, because we've already got sponsors onboard, and we can basically help give you an easier life and take that stress away because, honestly, the financial stress of planning these things is probably the biggest element of it. It's always interesting seeing the emotional journey organizers go through as they see themselves signing up to very large costs with the promise of money, without the money in the bank account. So, hopefully, we can ease that emotional journey a little bit and help them focus on the main elements.

We had a ServerlessDays that ran last year, where they spent so much time focusing on sponsorship and focusing on getting paid, they didn't spend enough time on promoting the event, and they had a mad rush in the last two weeks to try and get it [inaudible 00:20:18], and they sold a hundred tickets in the last and they didn't achieve what they wanted to, but they got to a point that it was a successful event, and what we need to do is let the organizing team focus on promoting the event. Let them focus on curating and running the CFP, and creating an event that's unique to that area, that builds a community and take away some of that financial stress, really. So that's the big part of that.

Jeremy: No, I love that because I think, for the Boston team, just the procurement process... We had to fill out of these forms on people's websites and do all this kind of stuff and then you don't get paid right away. Big companies pay when they pay, and we love sponsorships, and these companies, the ones you mentioned, have been excellent sponsors, but certainly, I think, putting something into place that makes that A, easy for someone like AWS or Google to put that into their marketing budget at the beginning of the year.

We found, with a couple of sponsors we reached out to, we said, "Oh, our event is in April or March," last year, and they had said, "Oh, well, we already did our budget planning for next year." So, it's not there. Some of these events can pop up and can run fairly quickly. Four months, like you said, at a minimum, if somebody does that, then you might be out of cycle for some of these sponsors. So, being able to get some of the bigger sponsors and know that those main things are covered, I think, is really, really great. Besides financial support for organizers, you mentioned some documents and things like that. What else is the global organization doing to help organizers?

Ant: One big thing we want to do is looking at creating an organizer's guide, a one-one-one guide on how to run a ServerlessDays. Up to this stage, we've been sharing tips on that in Slack, and someone jumps in Slack and asks the question, there'll be someone who can answer it. We also point people to the organizing guides for JeffConf and DevOpsDays and [inaudible 00:22:31], and those. But those conferences all have different [inaudible 00:22:35] and different things that make them unique. So, what we want to do is actually create a organizing guide that's specific to ServerlessDays. I do remember, last year, basically, we created a URL, guide.serverlessdays.io. It's got no documentation on it, but it's there. There's an outline. So we need to populate that, and I do remember putting that URL on the ServerlessDays to organize this channel, last year. You popped out, I think it was just after the Boston one, you popped up when you say the names. It was like, "Hey, that would've been useful." So, apparently, it would be useful.

Jeremy: It would be nice to just have the checklist. That was one of the biggest things, for us, was we had a bunch of people on our committee. We had some really great people, and they had run some conferences before, or have been part of these organizing things before. I know Matt Williams had done quite a bit with DevOpsDays, and Erik Peterson, Christina Wong, a couple of others. If you had this master list, okay, you've got to call the venue. Call the venue, get the contract, sign the... Maybe it doesn't have to be that detailed, but certainly, all the incidentals like, "Oh, do you want to organize open spaces?" Well, then, you go talk to some speakers and try to get that worked out. What do you need for food, at least that kind of stuff, and videography. Things that you might not think of, like the little incidental things, t-shirts. You're printing t-shirts. You're doing conference badges, or you're doing whatever it is and just having some of those best practices in place and some checklists, I think, would be really helpful.

Ant: Yeah, that's one of our aims. Let's be a little bit more prescriptive. Let's help these teams. A lot of ServerlessDays success comes from building on the shoulders of previous communities. So, like DevOpsDays, where there's a healthy overlap of DevOpsDays organizers and ServerlessDays organizers. Also, DevOpsDays again. It's the 10-year anniversary in October, and I saw organizers from six other ServerlessDays there. So, we want to learn from them.

Bridget Kromhout, who runs DevOpsDays, came up and did a tour about how DevOpsDays grew. The first five years of DevOpsDays, they'd only grown to 15 events, and what happened during year five is they decided to create documentation on how to run DevOpsDays. There was an exponential growth after that. So, I think we're lucky. Our growth has been quite rapid, and a lot of that is because we've had people who've organized other events in the community already, and people with experience. But, we can't rely on organizers being able to pop into Slack and have someone answer their question, to help us grow.

We really need to have to standardize this and be a little bit more prescriptive and have a clear guideline on how to run these things and, sure. What do we do about food? How do we handle sponsorship? How do we handle covering travel and accommodations for speakers? What are the policies on that? Have that all covered, basically, just to make it easier.

As I said earlier, it's just about the greatest value our organizers bring is organizers are all practitioners. They're all members of the community. We don't have marketing teams running these things. It's developers, engineers, running these events. So, let's, then, focus on putting together a great event with great content, and focus on getting the right speakers in the room, and let them focus on building the community, and make all that other stuff, that has to be done to have a great event, and make that as easy as possible.

Jeremy: Yeah. Even the finances, just understanding the financing stuff. We had Christina Wong, who was on our team, and she was the Treasurer for us. That, in and of itself, is just a huge undertaking. Then, just like space coordination, like I said, we were very lucky with Microsoft NERD Center and Simona Cotin, from Microsoft, actually got involved with our group. She's on your side of the Pond, there, but she was able to reach out and help out with our team and get us space. So we just had a lot of community support, basically, in order for us to make these things work. I think that codifying that and making it very accessible to people who want to do it, I think, would be hugely advantageous to new organizers, and existing organizers, like people who have run this in the past, and thinking, "Okay, we're going to do this again. Do I have to go through this whole process?" Maybe get their feedback and, like you said, incorporate that into the overall documentation, there.

Ant: Yeah, that's exactly what we're looking to do and potentially do some sort of documentation sprints in the New Year, and get a bunch of the organizers together and, basically, do one big data post of everything we've gone through, and what we think should happen and shouldn't happen, and get that documentation up. The [inaudible 00:27:54] is there. It's just about putting content on the [inaudible 00:27:58], which would help everyone.

The guideline thing, it's being a victim of the success of ServerlessDays because everyone who runs ServerlessDays is part-time. For everyone, it's a side project. So, because it's a part-time, side project, one's actually had the time to actually write these things. So we actually have to find a little bit of extra time to save ourselves time, later down the line, to get it done.

Jeremy: All right. Coming up, this year, we know there's a bunch of ServerlessDays that are happening. You've got Belfast. You've got Cardiff. Boston is happening again. We're just waiting on the final date for that. There is Nashville. There's a whole bunch of... Hamburg is happening again, which must be... Is this their fourth one, now, that they're running, third one or fourth one?

Ant: Yeah, this will be their fourth.

Jeremy: Their fourth one. Okay, and then Helsinki is running one. So is there any other-

Ant: I think that'll be their third.

Jeremy: Their third one. Okay.

Ant: It'll be their third.

Jeremy: Are there any places in the US, simply now because you're asking me to help with this, are there any places in the US, where we're not seeing any of these pop up? Are there some places we want to target?

Ant: The US is the trend, because coming from outside of the US, we always think of the US as one country, but the reality is it's a very large country, and it's a very diverse country and not everything is evenly distributed. So, I think, areas that are potentially underserved, like ServerlessDays Chicago, we had an organizing team there, and they basically, to the team members, moved out of Chicago. So, she stepped away. There's one person, there, who's still very keen to get it going, but needs a team to support them. So, if you're interested in getting involved in ServerlessDays Chicago, just ping us on organize@serverlessdays.io, and we can introduce you.

The Twin Cities in Minneapolis and Minnesota, there's a great team there that approached us last year, saying, "Hey, we want to run a... I said last year. It was this year. They approached early this year, saying, "Hey, we want to run a ServerlessDays, but we don't know what the community's like, and there wasn't a Serverless user group. So the recommendation, there, was start a Serverless user group, and they've done that, and they've been running that for six months. They're, now, starting to plan their own ServerlessDays. So it should, hopefully, happen in 2020. So, there are a couple of these.

Ant: I think the key thing is if you can get an organizing team together, it doesn't matter how big the center you live in. ServerlessDays can be 100 people. It can be 400 people, but get a team of minimum of three people together, because it's way too much to take on unless there's three or more of you. Ideally, if you're a bigger center, you want more. Just start it. It doesn't have to be big. The first one in London was 170 and, like I said, we've had smaller, and you'd be amazed, too. We'll come and speak in your area. Actually, sometimes a smaller area is actually get better attendance, because there isn't competition with other events.

Jeremy: Yeah, not the bigger events. Sure.

Ant: Yeah. There's other good ones, I think. ServerlessDays Phoenix is starting to get planned. Yeah, there's a few others in the US. Then, the other interesting one is, we might to into China, this year, in 2020. There's potential for three ServerlessDays events in China, which would be huge. We've been speaking to organizers in South America.A couple have come close, but never run one yet. So, we want to see that. I'd love to see a ServerlessDays in Brazil or Argentina, or Chile or wherever. Columbia is another country we've had conversations with, before. So, we'd love to see-

Jeremy: There was just a JSConf in Columbia.

Ant: Yeah, exactly. There's development communities there that would love to have a ServerlessDays. So, it's just about finding individuals who're willing to put the effort in. Obviously, India is another one where I'd like to see a a lot of growth. We've had a lot of conversations in India, and, hopefully, 2020 is when we, hopefully, will see quite a few events there, or 2021. But, there's a lot of growth to be had and I think the key to a ServerlessDays is start small and grow from their. Don't think you have to run a four or 500-person event. 100-person is fine. Keep your budget low. Minimize any risk. We'll have an organization behind you that can, in 2020, financially back you up and support you. Hopefully, it'll grow significantly.

Jeremy: I just want to go on the record, saying that I am willing to help and attend a ServerlessDays Hawaii. If anybody is interested, I'd be more than happy to lend a hand, there. All right, great. So, while I have you here... Part of the reason why you started ServerlessDays was because you are a serverless fanatic, like many other rabid serverless fans in the serverless community. So, I'd love to get your take... I know you were just at a reinvent. I saw you out there. What's your take on where Serverless is right now? What do you see? I don't want to ask you the, "Where do you see Serverless in five years?" thing, but how do you feel about the adoption and where it's going?

Ant: I think the adoption is good. It's still strong. One of the challenges it has is for better or worse, AWS is the dominant player in the serverless world. So, if you live in the AWS world, you see serverless everywhere, but if you're using Google Cloud or Azure, or potentially you're a web developer, Serverless isn't as prevalent. Google Cloud and Microsoft Azure, both, seem to have a greater preference towards containers, though they are getting more and more serverless with the container offerings, [inaudible 00:34:19] is a good example from Google Cloud. I think Serverless is in a good place. I don't think it's got great marketing behind it. It doesn't have an opensource foundation, and the core value of Serverless will probably never have an open source foundation, because it's all about the platform, and the platform benefits.

So, it's slow and steady. People are adopting it because they see the value in it. The ecosystem's building out. One of the great things I love about it that there's a solid community. It was great being at re:Invent, because got to see all the people I see every year at ServerlessDays. Now, we're all in one place. I think it's great. I actually like what the Microsoft teams are doing. They've got a 25 days of Serverless, at the moment, where they're doing those serverless challenges every day, leading up to Christmas. That's great.

Microsoft has a small team, and they're very passionate about their platform. They don't have the investment of Amazon. I know the Google teams who work on their service platforms are also very passionate about it. I think there's a lot of passion, a lot of commitment to what Serverless is and the value it can bring, and I can only see that growing because, one of the last things I do like about it is doesn't have the marketing dollars of other open-source foundations. Maybe in 2019, 300 thousand, 400 thousand is probably the total marketing spent on every single ServerlessDays, total. That's about the marketing spins of two headline slots at one CubeCon.

The marketing's been done on other cloud native technologies is magnitudes of what Serverless is, but Serverless still goes slow and stead and keeps going. It keeps going because it's got such a strong value proposition. It's so much easier than other options. There's no clusters to manage. There's none of that configuration. There's trade-offs. There's other things to manage, but it is significantly easy and you get great value, really quickly. I was seeing more of that, and that's great, and it's a lot of where that value, I see, is in the talks at ServerlessDays.

When you see Lego Group get up and tell you about how they started with one function, three years ago, and where they are now, or Liberty Mutual, very large, Boston-based insurer with a good presence in Belfast, how they talked about their improved [inaudible 00:36:58] market, their reduced costs, their better availability, and all these positive metrics when they've moved things to a serverless platform. That's where I see where that growth is. It's not just one company that's got one genius engineer who can make this thing work. It's multiple companies who are seeing that value, and it's just happening slowly and slowly. I think slow and steady is going to carry on growing, and it's definitely... It's on its own cycle. I don't think... It's not marketing-driven hop, for me, personally.

Jeremy: Yeah, one of the things I think we saw, this year, which was very encouraging, and I think is great for pushing adoption is we started to move away. We started to move away from a lot of these toy projects that you've seen people build with them and, actually, into things, now, that are production grade. Liberty Mutual is... That's a big one obviously, what Capital One has been doing, what [LEGO's 00:38:02] doing. There's so many companies, now, that are starting to do that, and they're building really interesting... They're interesting patterns around microservices with functions and managed services and things like that. So, I'm really encouraged. I think 2020 is going to be the year where you've got a few of these vendors that are really smart, know what they're trying to do, in terms of getting this out there, and I think you've gotten that incremental adoption but, at some point, I think, someone's just going to step on the accelerator and you're going to see this explode. So, I'm excited for that.

I don't know if 2020 is the year for that, but I think we're getting closer.

Ant: Yeah, I agree. I'm not 100% sure if 2020 is the year. It's been this growth under covers but, I do think we are going to hit that. I think it'll hit when there are more and more of these companies that come out and start publicly talking about the value of it. It's been a bit of a weird one, because we know some of the very early service adopters who get such value out of it, and they see it a differentiator to what they do, where they don't actually talk about the numbers. The might go up onstage and talk about their technology, how they built this integration or how they monitored this, but they don't talk about the numbers, and they don't talk about the business value, because they don't want their competitors to find out. They don't want to understand that this thing is a bit of their secret sauce. I know a few companies who've said, "We love it. It's amazing. We'll talk about the technology. We won't talk about the numbers because we don't want people to find out." I think that, as we get more and more adopters, we'll get more and more companies, getting onstage, doing case studies, and talking about the amazing value they can get from this when they do it properly. Yeah, I think it's going to come. I don't think we're far away, whether 2020 or 2021 or 2022, but it is going to come.

Jeremy: Awesome, and I hope that anybody who's listening, if you have not been to a ServerlessDays, you should definitely go check one out because, like you said, there's a great pool of speakers and sometimes you get those absolutely amazing talks that go into the detail of how some of these companies are actually using it, as opposed to more theory, which is what my talks tend to be bout, more about patterns and things like that, but, hopefully, you can use those to build your own stuff.

Anyway, so listen. Thank you, Ant, for being here. This has been great. I'm really excited about the future of ServerlessDays. I think we've got some cool things there. If people want to get in touch with you, or find out more about ServerlessDays, how do they do that?

Ant: You can find me on Twitter @IamStan. Just a warning, most of my posts are about my dogs, because they're way prettier than me. Otherwise, just drop an email to organize@serverlessdays.io, and that would be the English spelling of organize@serverlessdays.io. Maybe we should set up an alias for our American friend.

Jeremy: Probably. Switch out the S and the Z. Right?

Ant: Yeah, absolutely [crosstalk] for that.

Jeremy: And, if you want to check out all of the upcoming ServerlessDays events, serverlessdays.io, it's all listed there.

Ant: Yes, serverlessdays.io/events. They are currently not on the website, right now. I think another three will go live, by the end of January. So, yeah, go find the ServerlessDays, go look at it, look at attending. They're typically dirt cheap, ranging from 20 pounds in Cardiff to, I think, Boston, you're at $50?

Jeremy: 50, yeah.

Ant: Yeah, the price is nuts. It's so cheap. So we get good attendance and it's a day off. It's one day. Go attend a conference. Go. Spend a day learning and seeing how the companies are doing this.

Jeremy: Awesome. All right. I will get all of that in the show notes. Thanks, again, Ant.

Ant: Cool. Awesome. Been great to chat.

View Details

About Chris Munns:

Chris Munns is the Senior Manager of Developer Advocacy for Serverless Applications at Amazon Web Services based in New York City. Chris works with AWS's developer customers to understand how serverless technologies can drastically change the way they think about building and running applications at potentially massive scale with minimal administration overhead. Prior to this role, Chris was the global Business Development Manager for DevOps at AWS, spent a few years as a Solutions Architect at AWS, and has held senior operations engineering posts at Etsy, Meetup, and other NYC based startups. Chris has a Bachelor of Science in Applied Networking and System Administration from the Rochester Institute of Technology.

  • Twitter: @chrismunns
  • Email: munns@amazon.com
  • AWS Compute Blog: https://aws.amazon.com/blogs/compute/

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Chris Munns. Hey Chris, thanks for being here.

Chris: Hey, Jeremy. Thanks for having me.

Jeremy: You are the Senior Manager of Developer Advocacy for Serverless at AWS cloud. Why don't you tell the listeners a little bit about your background and what you do in that role?

Chris: For sure. Definitely. Going back to the earlier parts of my career, I started as what I guess, I would have considered a sysadmin. Maybe these days, you would call it a DevOps engineer or an SRV or something like that. I took care of servers and infrastructure, a jack of all trades across the stack below the application. Then, just a little over eight years ago, just about eight years ago, I first joined AWS solutions architect, did that for a couple years, actually went back out to a startup and then came back again. Then for the last three years, I have been a developer advocate for Serverless at AWS.

Then, just in the last year or so I've actually built out a team of people that are all over the globe. What we do as a team is we create a lot of content, we deliver a lot of content, we do a lot of interacting with our customers, trying to share the good word about Serverless and get people over the challenges and things that they are understanding the various aspects of our platform, I would say. You'll see a lot of our stuff show up in webinars, and Twitch and blog posts and in conferences, and in social media and all that stuff. I would say the next biggest part of what we do is act as a voice of the customer back to the product teams. We are embedded in the product organization, we have influence over and what product is built and to a degree, how it's built. We want to make sure that our customers, concerns, the things they're trying to solve the challenges that they have are being properly represented back to the product organization.

Jeremy: Great. All right. We are live actually in Las Vegas, we're at the Big Show, as AWS fans, I guess would call it. We're at re:Invent 2019 there have been ton of announcements so far this week and I think we're pretty much done, we've hit the max on cognitive load for the number of serverless announcements that have come out. There are a whole bunch of them that I want to talk about, and we can get into some of these in detail. There were some really great ones that I think solve a lot of customers pain points. What do you think are the biggest announcements that came out so far? Maybe not just that re:Invent, but also in the last couple of weeks? Because the last few weeks, there's been a ton of announcements as well. What are your thoughts on that?

Chris: Yeah, it's been a really hectic period for us in the serverless organization at AWS, in the last two weeks a whole bunch of things. Really, I like to boil it down to four key big things that we've launched in last couple months, announced in the last say three months that I think take on some of the biggest challenges that our customers have. The first was back in September, we announced that we were going to be changing the way that VPC networking worked for your lambda functions. We announced this new concept of what we call a VPC to VPC net, it's built on in a data based technology called Hyperplane, it's part of the advanced part of our networking stack.

As of last week, the week here before re:Invent, we got Thanksgiving here in United States, we actually finished the rollout all the public regions that we have across the globe. It's taken some time to get this rolled out. It's actually a really huge infrastructure shift, but basically what this did was it drastically lowered the overhead of having your functions attached to a VPC for cold start, we had examples where it was shaving 8, 9, 10 seconds off of that initial cold start pain. It also reduces the total number of them and so really huge one. That's the biggest one out in all public regions today globally and customers are just seeing the benefits of that.

The next is on Tuesday of this week, we announced a capability in Lambda called Provisioned Concurrency. You and I have some fun history in this that it was almost two years ago at a startup event in Boston, maybe it was? Where I talked a little bit about, some of the pre-warming hacks and then you and I just went through back and forth on it for a while you launched your-

Jeremy: Lambda Warmer.

Chris: Lambda warmer project, which has become the de facto standard. We're full circle here, you and I have this like, I don't know, two years later almost?

Jeremy: Right. It's actually funny because I wrote a blog post like an open letter to the lambda team that was asking for provision concurrency. You and I had this conversation way back when, and you said, "Well, we really don't want to do that. Because, we want to improve the cold starts and get those down." I'm actually really glad that the team at AWS did that because I think if you would have gone with provisioned concurrency before then the need to make those improvements wouldn't have been there. It pushed you and your team to and the engineers there to get those cold starts down and work on that. But now provisioned currency adds a whole new level, which again, I think is great. Do you want to talk about that now?

Chris: Yeah. We'll riff on that.

Jeremy: Okay.

Chris: I saw some commentary that people thought that this meant that we were giving up on continue to improve cold starts. It's like, we want to be really clear that, that's not the case. This is a knob or lever that you can turn that yes, it changes the way that functions are essentially, it's difficult to use the term pre warm, but effectively, they are pre baked up through the init phase of the function lifecycle. What we've done throughout the year, and I've made a couple of tweets about this in the first half of the year of places where we've shaved tens of milliseconds off of some part of the overhead of the platform, or we've lowered jitter on various aspects of it. There's a lot of that stuff that continues to happen and actually, I talked about this at Serverlessconf, New York City back in October.

I basically said one of the key benefits of Serverless is that it just keeps getting better for you. There's basically three ways that happens, one is all the stuff we do behind the scenes that you just never see. The second is the stuff that we launch, that we tell you about but it's just automatic, there's no like opt in, there's no option you need to take to enable that. That's basically what the VPC improvement look like. Then the third is where we give you an op where we say, "Hey, for certain things, you're going to want to make this conscientious decision or not to turn this on." Provisioned concurrency is an example of that third one. We see it primarily being for interactive or synchronous space workloads, primarily API, chat bots, things like that.

Again, what it does is it provides, I think, a more trusted solution than some of the things that you and I had even talked about that when you were like, "We can do this. It's a little hacky, you run this fourth logic and your handlers and you do this other thing." This is going to give folks a just a much more consistent method for doing this and then the outcome of that method is greater consistency, lower latency, and potentially even lower costs. That's one of the interesting aspects about this, we're looking at pricing. We didn't want to make this be a penalty for performance. We consider it a premium feature, for sure. That's mostly because we don't think that everyone needs it. There's only certain folks that are really, really can have this extremely low latency requirement.

By and large, and this is another thing we've talked about in the past, cold starts impact very, very, very few in locations by the biggest misunderstanding about the cold start, and how it impacts a workload. Is that really below the 98th percentile or something of traffic, you just don't see it.

Jeremy: You're not going to see it.

Chris: Where that last two percentile is something that people typically care of, so that some people really care about a lot, I should say. Provision concurrency help solve that for them.

Jeremy: Yeah. I actually, I liken this and maybe this is not the best way to think about it, but I think a good mental model is this like buying reserved instances of lambda functions, in a sense that's where you refer. You pay a little bit upfront, but you get a cheaper execution or the upper execution is a little bit cheaper. I look at it that way, but I totally agree with you on the cold start thing, where it's for almost 99% of what you're doing cold starts will never come into play and it's not that big of a deal. I can see there being workloads that are fairly consistent, that are fairly heavy that you might want to... I think the pattern here might be to slightly under your provision concurrency, so that you're almost always using 100% of that provision capacity. Then, the other thing that I thought was interesting is some people have pointed out, they're like, "Well, why not just do the lambda warmer or the cold the CloudWatch ping technique?

Chris: Yeah.

Jeremy: What's different is every time like my project has to call your function. You're using one of those. You're using the concurrent connection in order to do it and the system would allow you to do multiple concurrent connections and the way that it did that was by actually running a wait timer to keep the other ones open. Then there was enough to though you could actually build up that concurrency. The problem with that is that ties up that concurrency.

Chris: It's blocking.

Jeremy: It's doing blocking, right? Where this provision concurrency doesn't do blocking, you always have those available.

Chris: Again, I think that was one of the things where we never formalized the warming hack model as we consider it in any real way. There was also this unwritten what I call the 5, 15 rule, where we said, "We keep functions or execution environments, warm five minutes outside of VPC and 15 minutes inside of VPC." That was before the VPC networking improvement. That was before, a bunch of other things we might have coming out that might even make those times dynamic. The 5, 15 rule might go completely away and then people might have to even more creative and do all this other hacking, and do all this other stuff. Yes, the warming bottle that you and I have talked about it was blocking, effectively could be detrimental to customer requests and of itself.

Jeremy: Absolutely. Yes.

Chris: Again PC, provision concurrency called PC for short now, it gives you an official solution to this problem. Not everyone by far is going to need to solve this problem this way. For folks who do, this is the way to do it for now. That said, we just announced this thing two days ago, we're digesting everyone's feedback. Obviously, we've been talking to customers for quite some time in private conversations about it, but what my team is doing at this point is aggregating feedback, we're going to boil it down, we're going to keep looking at it. If we find out that this isn't the right model, then we're going to work to find the right model for our customers.

Jeremy: Actually, one of the things I really like about this provision concurrency too, is that yes, there are probably those use cases where you want to keep 500 lambda functions warm and that may be for enterprise, whatever, but even for smaller things. I think about a bunch of administrative APIs that I have where they're synchronous APIs. You maybe have, 10 20, 30 people using it at any given time, maybe even less than that, if you wanted to keep that warm, and you wanted to keep some important endpoints warm, so that admin users didn't run into a cold start, which again, would probably be fairly minimal anyways, at this point. That's something where, provision five or something like that it barely cost you anything, it's still going to be very inexpensive, but it would take away that pain that the occasional user would think.

Chris: One other big thing about for provision concurrency is that it's integrated with auto scaling. It is auto scaling pervasive across so many different parts of the platform. People are very familiar with its mechanisms and how it works. There's some little tweaks here that happened with provision concurrency. You can follow your traffic, you can schedule the ability to rise and fall at certain times the day. I think we're going to see people who say, "Yeah, I have a highly latency sensitive application, I'm going to over provision because I need my TP99, TP100 to be really, really low and consistent with the TP50 that I have." If you're not familiar with that term, we're talking about, essentially, performance across the graph of performance over how customers are experiencing your application.

I think we're going to see people doing a lot of different things with it. I think we're going to see people use it and be like, "I probably don't need this." Now it's a tool that's out there, it is going to solve some big customer challenges, that's going to unblock a lot of people who want to build serverless applications.

Jeremy: I want to move on to other things, but one more point that I want to make on provision concurrency, because I was just talking about this actually with Slobodan and Alexander, who are two other Serverless heroes. We were saying, the pattern here or the best practice, leading practice is single purpose lambda function. We're building a lot of similar purpose lambda functions that are endpoints on APIs. Now you say, I want to build an API that has 50 different endpoints, maybe I want to keep those warm for some reason. Now the pricing model here is I'm paying for warming all these different lambda functions. One of the things I don't want to see happen is to slide back into the monolithic lambda function, because now I can keep that warm. Just a thought.

Chris: About that. In my talk, one of my talks that I have this week, the code for it is SVS-343, building microservices, AWS lambda. I actually have some... I talked about this. I talked about the per function, per API action model. I talked about the lambda length, the lambda based monolith that's out there. Internally, with our internal field experts and our subject matter experts internally, we've had a bunch of heated discussions about this. One thing that we see is that there are a number of developers that have frameworks that they love and that framework says I want to route logic inside of my function. Number of API frameworks that are out there that will do this. It ends up working for people, they end up being happy with it, they end up being successful with it.

Now, I would say that I think that you, Alexander, Slobodan and myself, were purists and wanting to see people use the platform the way that the maker intended to a degree.

Jeremy: You want to fail the mechanisms, you want all that resiliency and all of that stuff. I mean, I always tell people don't put dry catches in your lambda functions, let the function fail and let the cloud handle that.

Chris: We can talk about the new stuff that helps make that even better. I think at the end of the day the lambda lift also enables potentially some submitter code portability, people really care about that. There are definitely points where it starts to become problematic, where you reach the balance of what you can do in that lambda lift. Then as soon as you reach that tipping point, you're in like, "No, how do I break apart this thing?" Then you run into all the challenges that you would have with any monolithic application.

I think yes, we will see people who say, "Well, I've got this provision concurrency thing and upload my coder to one function. I don't have to prison that much or something." Maybe to a degree, they'll still end up having to provision some amount towards the cumulative amount of requests that they would have gotten across all those individual function. Maybe being too smart, despite yourself to certain degree, we'll have to see, we'll have to see what patterns end up really coming out from this.

Jeremy: I mean, that might be something interesting that could happen is where it's not about provisioning a specific lambda function. It's just maybe provisioning a certain amount of concurrency across a group of lambda functions or something like that and I'm sure there're things.

Chris: I mean, one of the big things that that provision concurrency does, we could definitely change topics. It gets you all the way up through in it. All the way in the lambda life cycle up through we execute your pre handler. What we've actually seen is that that pre handler code is the more expensive part of the equation most times, then the platform overhead. You've got people importing packages, they're talking out to other API services. Maybe they're getting stuff secrets manager from parameter store, it's crypto mass, we've got decrypted all these things. Provision concurrency getting you through in it, but stopping at your handlers actually, that's a big part of it.

We're going to have to see what patterns emerge, we're going to have to see what feedback we get and how to tweak it. I don't think this is a one and done feature now. I think you'll see some stuff in the future that tweaks in one way or the other.

Jeremy: Great. All right, so that was two.

Chris: That was two.

Jeremy: Next one.

Chris: Man, I'm trying. I got two more here and I'm trying to think of what order I want to talk about them because they're both so important. I'll do it in chronological order from when they launched. On Tuesday, we also basically co announced with the RDS relational database service, something called RDS Proxy. Again, those folks who have been building serverless applications for some time now, working with relational databases has been a challenge, you have to deal with the fact that your functions may need to establish new connections. Then when the function is idle, it doesn't necessarily tear that down. That establishing of new connection overhead could be expensive of both the database and your function. There's the scale challenge of you potentially consuming lots of connections on your database.

Jeremy: Zombie connections.

Chris: Yes. That would lead to people out sizing the database to deal just connections, which is an unfortunate thing to do. Long story short, this is going to help remove that problem. With RDS proxy, we basically are creating a database proxy. There are lots of other solutions for this in the industry. There's PG bouncer for Post grads. There's MySQL proxy that existed, open source packages, a couple others that exist that are out there. This one is built by us for the cloud, scales for the cloud. It's going to do connection management shared connections, helps you with things like failover for multi AZ, databases for you. It's only in preview right now, it supports just MySQL. We're obviously hearing a lot of feedback about Post grads. I'll just say seat tight folks.

Jeremy: I've heard that too.

Chris: Seat tight. We'll see what happens by the time we get to GA or soon after, maybe. This is a really big one. Now between the performance improvements and VPC, allowing people to put more functions in a VPC where RDS is typically running. Then this, you now basically get to the point where you can do these clicking relational database, have it really efficient and effective, not eat up a lot of resources, have to be really fast. This is a huge one. I've had some people say, "Wow, this is the biggest one of the weeks for me."

Jeremy: This is the other thing is funny. On top of the lambda, or the provision concurrency, which I had that package on lambda one more package. I also have a package called Serverless-MySQL which essentially does connection management for you. What it does use the process list cleans up zombie connections, it'll say if you set like 70% capacity, then it will automatically kill connections up to a certain point. AWS this week has killed two of my open source projects, but I actually love it because I don't want to do that. Right? I love the fact that you build workarounds, I think that's where we are with Serverless right now anyways, is that there are certain things that you think you can do or the easy to do with other things and connection pooling should be the simplest of things. When you have a femoral compute, and each one has to compete for those connections, like you said, you need some way to manage it.

Managing it with an open source package worked really well I get a lot of people using that package and things like that, I think people will still use it because they're it does some good transaction handling and or just got a better workflow for transactions and stuff. As I promote my own stuff. Seriously, now I think it's great because this is this was a huge missing piece where people are just not willing to let go of relationships of the human basis,

Chris: For very good reasons, right?

Jeremy: Yes.

Chris: It's where their data is today, it's what they understand they've maybe been trained in it or just have so many years of experience with it. I don't think we were ever comfortable with the idea of telling people like, "Up too bad. Time is relational." Big one. It's a block a lot of workloads, especially in the enterprise. Even for developers that are just in general, more comfortable relational databases is going to open up lambda for them.

Jeremy: All right, so just a couple of questions on this to clarify. The RDS proxy you still have to run inside of VPC, right?

Chris: Correct.

Jeremy: Then in terms of how that connects, there's some secret's manager, stuff that you need to do at the proxy layer itself?

Chris: Yeah. The RDS proxy ends up using secrets manager to handle the secret management between your database and the proxy itself. What's great then is this tie in up through your lambda function, so that you're not hard coding usernames and passwords in places. You can use either the IMF formication methods that they have with RDS today, or you can still use username and password however, managed by secret's manager. You just have to give your function access to that data inside of secrets manager, and it's pulling on the fly and all that cool stuff.

Jeremy: Right now in the preview, it only supports RDS Aurora MySQL, it doesn't support Serverless Aurora yet, right?

Chris: It supports RDS MySQL or Aurora MySQL but not yet Serverless MySQL. Correct.

Jeremy: All right, next one.

Chris: Yeah. Then last, but definitely, definitely not least is we also just announced yesterday, basically a new model for Amazon API gateway. For Amazon API gateway, the first effectively API model that we launched with was what we called it REST APIs. It was very much meant to be almost kind of buy the book to the purest vision of REST APIs. Last year, we announced web socket support, which was one of the biggest things that we were asked for from our customers. I think if we look at API gateway, it's almost like a misunderstood product. It's an incredibly sophisticated, powerful product that just can give you so many knobs, levers and so many things. To a degree customers are just like, "We just want something really simple, really easy, like bit more basic."

Given that, given feedback on performance and cost, and what people perceived what they were spending their money on, which is all valid feedback. We take it, we track it, we quantify it, all of that, like that's what we do at AWS. Yesterday, we announced something that we're calling HTTP based APIs. I'm trying to make sure I don't get the numbers mixed up here, it's 70% less cost than the REST API model.

Jeremy: It's $1 per million locations as opposed to $3.50 per million.

Chris: $3.50. Yes. I believe it's a 60% reduction in the overhead that API gateway used to add. I don't think it was pretty quick, but it did have a whole lot. We still had people saying, "I want even faster." This is going to make it so that you can build, really simple basic APIs up through still fairly complex APIs. We've got a bunch of new authorization capabilities. We've got just a bunch of other things that you could do with it that make it a little more flexible in some ways. There are some things that were in the REST APIs that are not in HTTP yet. It's also a service that's just in preview right now. Again, I think for Serverless customers, in particular, with the way that how a lot of our Serverless developer customers are building APIs, this is just going to be just a solid win for them. It is effectively a new product inside of the product name. You do have to relaunch your applications to support it. It's not just like a toggle. You can do a number of things to export an API that imported in. Or you can just redeploy your application up with a new version of it.

Jeremy: I think that this is a very cool new product and the new way we're going to do, it's definitely going to reduce costs. The model for this is, is that the lambda proxy model, where it's that full pass through into the into lambda function?

Chris: Correct. Today, it's definitely meant to be a much simpler, easier experience.

Jeremy: No VTL templates and that stuff.

Chris: No, no, you shouldn't do any of that right now. Definitely, not. We've got a bunch of new easier capabilities around Corps, talking about the authorization of authorized capabilities. Now we have JWT authorized tickets for open ID, which was a big one that people have been requesting.

Jeremy: Announced Apple log in for Cognito, right?

Chris: Cognito, now it's Apple login. That's going to help folks that are in the Apple iOS, OSX world of things. Which is not a small community of developers. A bunch of new stuff that you'll be able to do with this. Again, I think the lower costs, the better performance and some of the easier configuration capabilities spec.

Jeremy: Great. All right, so that was four. Those are the four big ones, I know been talking for a while, but there are a few other ones that I do want to get to, EventBridge schema registry. I'm a huge fan of EventBridge and now using it in every project I'm building. Again, I talked to some of the team about EventBridge, and I think there're some amazing things coming down the road for that. Tell us about the schema registry.

Chris: Backing up for those who don't really know what EventBridge is at the end of the day. EventBridge has a concept called Buses. This is built upon what we were doing previously with CloudWatch events, so takes the same underlying technology for that. Essentially, what you do with EventBridge is you have the ability to connect some source service and event source. This could be a number of AWS services, it could be third party SAS products, or something custom that you want. Then what EventBridge does, its bus models then allow you to pass that event through set really fine grained rules against the actual individual attributes of the event. It's JSON structure we support today. Then that can then be targeted out to I think it's like 17 different services. It's lambda, its SQS, SNS, it's step function.

Jeremy: Step functions, which you can't do with SNS.

Chris: It's Kinesis, it's Fargate. There's a bunch of different places that you can pass the events. I table this, event will how to structure, you can see that schema and what we're now giving you the ability to do is to basically for your applications, for those third party application, for AWS surface applications. Track the schema of those, register it, you'll be able to determine like a type on it. Then the coolest, coolest thing I think is that we give you the ability to generate what we're calling code binding, which is basically code that will allow you to pull out the individual attributes of that event.

I've heard this from a number of folks over the years, Mike and john from Sinfonia. They always talk about like, "Just give us some code to pull apart the events." Well, to a degree we did here. Code bindings is a super big part of it. I think this is still some early capabilities is still in preview, but being able to track the types of events, which when you have that schema registry there, and stepping back to bigger picture with EventBridge. One of the ideas that we see is when you have all these events flowing through your infrastructure, people will just say, "How do I discover when an event is?" Service registry helps with that. You could see an event, you could see the schema of it, you can see the attributes inside of it, you can decide how you want to filter or have a rule set configure for that.

As you're in, especially for larger organizations, as the number of services that you have expands that you want to do new things consuming different services, the registry is going to make it really easy for you to discover what types of events are flowing through your system.

Jeremy: I think for me, I use EventBridge so much and I have so many events flying through it right now, with different microservices. That's using that as that main bus, that just the number of events is staggering, there're different types of events. Obviously, having the registry for the AWS events, that's easy downloading those code bindings, auto complete in your VS code or whatever you're using. That's a super handy feature, but when you start generating your own events across teams, being able to discover interesting events, categorize them and put them into that registry now I think will be really, really helpful. I posted something on Twitter about this, where I said, "The registry would be great getting people to put stuff in the registry is a challenge in and of itself." With the schema discovery piece of it, I think that is going to be a huge step where people will really find it useful. I'm very much so excited about that.

All right, so then another thing that came out and we can... this is a quick thing, but Express workflows.

Chris: Express workflows, this is four step functions. Again, for folks who are not familiar with step functions, step functions is a managed orchestration service, if you will, that allows you to take all of the workflow logic that you would otherwise be writing code for. Things like decision tree logic, how to chain functions or chain capabilities together, parallelization, failure handling, things like back off and retry. This is all stuff that developers always have written code for, that you end up pulling in some random module off of NPM or hit package or something like, "This is exponential back off retry and it's made by code blaster 317, how fast is this?"

Step functions helps you take that logic out, put it up to a managed platform for you, so that you don't have to think about how to do that. Step functions has been out now for a couple years and what we did, basically, or what we heard from customers was a couple of things. One, the way the service was built, the throughput of it maybe couldn't handle so the most extreme workloads. There was a cost aspect to it that also made some folks, it didn't work for them, basically. What Express workflows do is they have a really, really massive scale difference. The default limits here, and I just had to pull this up from my notes for standard workflows was over 2000 per second. This supports over 100,000 per second, there's a magnitude difference there.

There's some trade offs, I would say, are just the differences between these with standard workflows. You could have a workflow that could run for a year. That's an unusual thing, but it's something that we've seen people need. Because workflows aren't just for lambda functions, you could have human actions, you could have batch processing, you can have things that go away and come back again, at some point.

Jeremy: They're called back pattern and some of those other things.

Chris: Yeah. This actually gives you a maximum time, five minutes. Where you have a very discreet scope workflow, where you still have the same logic that you want to capture, you still want to do the same type of failure handling, but it isn't one of these much longer types of runs that might happen. This is pretty key for that. One of the other things is a giant cost difference for this. This works out to be about $1 per million in locations where I believe the standard workflows was about $25 per million state transitions.

Jeremy: It was 2 cents per 1000 states transition or something.

Chris: I may actually be wrong on the math on that one, but basically, still it is a magnitude difference in cost. Again, I think this is another thing where it's just going to unblock people for being more comfortable with saying, "Yeah, you know what, I don't have this really long running workflow execution, I want to take some data and I want to pass it through a couple different services real quick. Then a couple different lambda functions, or whatever it is that may be trying to get the end result of that be done." This is just going to enable that to happen a much greater scale.

Jeremy: Yeah. Even just from the pricing standpoint is huge for a lot of those workflows, even the ones that I've been doing. We have an article system at the company I work at, and we pulled out articles from the internet, and we run them through a series of national language processing, and we do some extraction and then we do some algorithms. Now all of that happens within 30 seconds, but I have several pieces of that logic that I like to reuse in different ways. What I ended up doing is either stitching those functions together with a synchronous call or something like that, which has some problems that you can get, or you're just basically putting all of that code into one lambda lift as you said and I don't like doing that.

This is one of those things where now, step functions will actually become very, very useful to me and affordable for the volume that we're doing. Because even so it just gets out of control. I think that is a very important one, and it'll make function composition for the right types of things. Then just having those guarantees and the back offs and the retries and the error handling taken care of for you. I think that one is pretty cool. All right, there was the Amplified data store that was launched, I think people should go and look at that if you're in the mobile space. That's a little bit outside of your scope.

Chris: My knowledge and ability to be hands on with it's limited to this point, but my understanding is it's really here to help you with offline data synchronization, storing data on the device. For mobile developers, it's super powerful.

Jeremy: I think that is certainly going to fit in, especially with amplify and all these other things that are happening. Amplify is launching a lot of stuff into the Serverless space and so there's some overlap there, but definitely more mobile. I'll have to get like maybe Nader Dabit on and we talk about that.

Chris: You should. Yeah.

Jeremy: Those are the main things, I think that were launched this week. As we said, there are some things that were launched leading up to that. One of those things, which I think is a huge game changer is lambda destinations.

Chris: Yeah. Yeah, absolutely. This is a big one. It's one for people who are new to Serverless application design, they are like... they scratch their head a little bit trying to think about it. You mentioned this before, of why. What does lambda destinations allow you to do? Basically for asynchronous invocations it allows you to capture either the success or failure outcome of that function execution. We have for many years now had a concept called dead letter queues, which would allow you to capture in failure scenarios the event requests that went and that failed. You could take that message, the dead letter queue, reprocess it somewhere else, pull it back up later. Basically, it would help you with capturing, and then being able to retry failed events.

We had a lot of situations where customers were creating and writing a lot of code for the success path as well. You had times where, let's say you're processing data out of s3. People are uploading images, uploading data files, whatever it might be, s3 calls lambda, lambda executes, well, what happened? For a lot of folks, they'd be doing a lot of log writing. Maybe they're creating almost like an inventory system in Dynamo DB or something like that to track actions.

Jeremy: You're including the SDK and then having a call from the SDK, and from the lambda function to another service.

Chris: Yeah. It was overhead in a couple different ways that people didn't want to deal with. What lambda destination does completely out of band from your code, to write any code to handle this, it's now just built into lambda. You have the ability to take both the success and the failure of a lambda execution, and pop it off to one of a number different places. You can either send it to a different lambda function. You could also send it to an Amazon SNS or SQS target, as well as EventBridge. This will allow for some interesting chaining of functions, it would allow you to take at least in the success cases and say, "Hey, okay, so we did complete this stuff in this lambda function is this action in my workflow. Now let's send it elsewhere for something else, or some other service or some other space."

I could see where this plus EventBridge could be super powerful. EventBridge is just such an awesome product because of all things that could do. It's like, "Yes, you could send these directly to SS and SQS, you decided to EventBridge, we could also send it to their and then also a lot of other places."

Jeremy: With the event registry, with the schema registry, you basically could have other teams saying, "By the way, when that s3 file gets uploaded, the file gets uploaded a lambda function processes it." Then after that's done, it sends an event to EventBridge that says, "Hey, this was processed with some information around it." Now you may have a service from some other team that says, "Hey, would really like to go when somebody has uploaded that image?" You don't have to do anything anymore.

Chris: Exactly.

Jeremy: I mean, in that situation, s3 to lambda to EventBridge, you don't even need to include the Amazon SDK or the AWS SDK and write any of that code. You just process it and payload back and it handles everything for you.

Chris: One of the things I talked about in actually two my talks this week. The two talks that I have is about this tendency for people to always want to build synchronous applications. They want to glue part A to part B to part C. This is one of the things that over the last couple years, we keep having to reinforce this aspect that building more asynchronous distributed applications is the only way to build them. We still see a lot of people doing these types of motions, reading a lot of this glue code, doing all this extra work. This is going to enable it being even easier to build these distributed workflows to build, but not just build resilient lots.

Jeremy: Yes.

Chris: You capture both the success, the failure, we give you both the request and potentially the response. You can take either of those pieces of data as you might need, and then do something else more beyond that. The failure side of things now, you could take this and send it to EventBridge and say, "You know what, let's send it to..." Here's maybe like an interesting one, "The work that you were trying to do was too big for your lambda function." You could basically build a system here that captures the failure of the event and attempts to reprocess it and Fargate and UCS and some other place. You can get away from like, "The limitations of lambda block me from doing this occasionally." One of the hundred requests is out of the bounds of the limits of lambda, I can't use lambda anymore. It's like this basically changes it so that you don't have to have a lot of crazy logical front, you don't have to do anything super creative behind the scenes, you can pass it back through the rest of the ecosystem of things that exist and solve the problem. It's a minor, little thing, but it's going to be super powerful for how people architect distributed applications for Serverless.

Jeremy: You mentioned that people really wanting that synchronous experience and I totally agree, I think that, that's something where people are like, "Well, I need to know that this happened before I can maybe move on to the next step." I think where people are maybe not thinking about this is you want to do that with long polling, like you want to do an HTTP request and wait for all this stuff to happen and come back. You don't want to do that. I mean, there you could use web sockets with API gateway if you wanted to, make connection calls API that bounces around a bunch of different services do it if you need an incentive back to, or even app sync, for example, has this or two way push subscriptions built into it.

Once you make some change and make some calls, and when that data is updated, then that will go ahead and push back to you and do it. There are different ways to build that model and get a synchronous feel, but still be using the benefits and the resiliency and not have service six, the six service in line fails, and then everything fails. Then you have to reprocess the whole thing. That's interesting. Those are the success pass stuff, you mentioned a little bit about the DL queues stuff. This is the failure mode on this really should replace DL queues now on a synchronous and that's because you now get the context you get the payload like you used to with the DL queues, but now you actually get the context of the error when it fails.

Chris: Yes. I mean, if you have your DL queues today, great. Keep using them. This supersedes DL queues. This is better DL queues, it doesn't cost you anything extra, it doesn't change your code, you could plug your DL queue, consuming or retry model back into this same thing. This just gives you better options for it. If you've got DL queues today, again, this is just going to give you more information. It's a minor configuration change.

Jeremy: I think the pattern here is probably on our send it to EventBridge, EventBridge maybe puts it back into SQS, if you want to do some replay or something like that. Now the shape of the data is going to look a little bit different, but here's a trick I haven't tried yet. I think this will work, do your failure into EventBridge, and then use the subscription to SQS to actually transform the data to just put the original payload back into SQS. I think you can do that.

Chris: I think you should be able to do that. I mean, you should be able to do that. The is question do you need to? I don't know.

Jeremy: Do you need to? I mean, that's right. Whatever your retry mechanism or your replay mechanism could certainly do that. I just like to think of weird things like sort of play around. All right, so then more on DL queues actually on another announcement was SNS DL queues.

Chris: Distributed systems have some unique qualities to them, and actually-

Jeremy: Everything fails all the time.

Chris: Yeah. Verner has its line, rid of rebel CTO of Amazon has his line about, everything fails all the time. We get a lot of people to say things like, "I want to guarantee that nothing's ever going to fail." I'm like, "If I had that for you, I would be gambling right now and not talking to you because I would be some mega genius or some fortune teller." There's always the potential that something could go boom in an application. There's no way to avoid this, if you think you've avoided it somehow in your on prem server somewhere. You're wrong. There's even a reason why NASA does things in threes and fours and fives and stuff like that.

One of the things that we have with services like SNS, it gets an event that comes in for the sources. Then it wants to send it to something like lambda or to another target like an HTTP endpoint, or to SQS. If for some reason, it can't reach that endpoint. It can't deliver that message. SNS by default does do retries different targets. If it continues to fail over a period of time, you could have hypothetically in the past lost that message. Straightforward just gives that DL queue or the dead letter queue mechanism to SNS, so if you have a failure off of it, that you can capture it and retry it, or do whatever you might need to do with it. For those rare situations, where you have that type of an issue, again, especially with if you do DL queues to something like SQL, then I'm not paying for anything, unless you have the problem. If you the problem, you have safety net, that's to pay for your use model that you can pull stuff out of it as you need to.

Jeremy: It's entirely worth it.

Chris: Yeah. I've been preaching around lambda for a while now for general asynchronous workflows to just enable DL queues. I probably wish that I could say that we would just like enable it for default for you, but it is a mechanism I think a little bit about. If you have SNS go enable DL queues, just go ahead and do it. Set it up to go to a queue, setup CloudWatch monitor to look at queued up, and maybe you don't even know how to pull stuff out of it right now. Set up that alarm when it does, you could say, "Okay, well, I can queue stuff up for pretty long period of time couple of days, and then find a way to process it, pull it out."

Jeremy: Another pattern that this opens up, which I think is really, really interesting. This is in combination with the lambda destinations is let's say you're running some piece of processing logic and a lambda function. That has to now make a call to an API, a third party API, you take your lambda function, you process it and the payload uses success that sends it to SNS. SNS has a subscription, an HTTP subscription, and then a DL queue on that. If that fails, there's nothing, you're not writing any code here other than here's what it should be that goes into my endpoint.

Chris: Exactly.

Jeremy: I think, that just gets rid of so many headaches. Because if you're trying to do that HTTP call from your lambda function, you have a problem. By the way, you don't need a nap anymore, if you're running it in a VPC because now you can just send it out. Let SNS make that HTTP call for you and now you've saved $30 a month for having an ad or whatever that is there. Right?

Chris: Yeah. It's synchronizing and it's calling.

Jeremy: Really interested things it's all this does.

Chris: When you use a service like SQS or SNS or EventBridge, or Kinesis for asynchronous communication, you get with it the benefits of the persistence and durability capabilities. It's not like, "I have this data in my single execution environment and if that goes, boom, I lose it." You put it into there and you have these durability capabilities and these persistence capabilities that again and of themselves are super powerful.

Jeremy: All right, one more, and then I'll let you get back to the expo floor. I know you have another talk later on today. SQS FIFO support for lambda functions.

Chris: This is a big one, we first announced SQL support for lambda the ability for lambda to directly consume off of SQS queues. SR queues, I should say, back in the summer of 2018. It was one of these use cases that is just so lambda-y I should say. People wanted it for so long, and we got around to it. We had to build some things before we can make that happen. We launched that everybody was like, "Great, now FIFO support." FIFO support or First-In, First-Out, basically gives you order data inside of a queue. There are a lot of situations where people care about order of records coming in, whether it be for transactional things, whether it be for sensor data, IoT workloads, tracking of all sorts of things. Clickstream tracking, basically any other options that have Kinesis and you can just throw things into and pull it out when you want.

Chris: FIFO support, again, now supported in lambda has the ability for you to pull out batches of records aligned based on attributes-

Jeremy: Message group ID.

Chris: They're message group ID. It could still actually support some really massive throughput and be able to again allow you to buffer up a bunch of information and pull it out all out ordered if you need to, you don't have to run any extra code, you don't have to do any pulling yourself. You just have your lambda functions consume records and do the thing.

Jeremy: What I love about the order nature of this as you think about Kinesis and that's been the go to when you wanted ordered records, and then have to do shards, so and then shards would scale and I know you can added some other cool things like the parallel stuff.

Chris: I had a similar one too.

Jeremy: We don't have enough time to talk about that.

Chris: Yeah, next time.

Jeremy: With the FIFO thing that's really cool is, like you said on message group ID essentially, that creates, essentially a shard ID, if you think about it, then you can have... if you want 10 separate lambda functions to be processing things in parallel, you can pull stuff off of that stream using that group ID, you'd have 10 different group IDs that could do that. You can use that group ID then or the message group ID to segment maybe different customers or whatever. It's only parts of the stream that you need ordered or in group by a certain things. I think there're some patterns in there to where you might be able to use it for priority and some of those things you could do some interesting stuff.

Chris: Absolutely.

Jeremy: All right, great. Let's close with this because I love talking to the people that AWS everybody all the PMs, all the engineers.

Chris: Thank you.

Jeremy: Everyone is just so excited about the future. I think you got a lot of people like me and there's a whole bunch of we hang on the What's New blog like, "What's coming out next? What can I build with it?" I think, you've got 65,000 people here who share that enthusiasm but that is infectious because of the way that your team and the PMs that are building these products. I think you're more in the role of advocacy, obviously and you've got a great team of people, you've added some great people. Then I know like, there're some evangelists as well who are doing a similar, a little bit different. What's that future of advocacy at AWS? What are your plans for Serverless around this and getting that getting people on board with, "Hey, Serverless is the way?"

Chris: For sure. In a broader sense, I and maybe it's a little boring to like our strategy, I think of customers across like three different life cycles. There's the completely net new green customer who is like, "What is this stuff? How does it work?" They're the folks that are further along in their journey. They're building applications for production. They're they're doing real life, real world things with it. They're like, "You know what, I'm running into some rough edges, I'm looking for some best practices guidance. How do I scale this right? How we do the best patterns and stuff like that?" Then there are the folks like yourself, like Bengio from IRobot. Many of our other heroes, many of our big customers where you're like, pushing on the bounds of what we can do and how we do it.

Across all those different areas, we want to look to be able to tell stories, share advice, give guidance, gather feedback, continue to grow the space, grow the workloads, we use this hashtag a lot all of us, which is #ServerlessForEveryone. I said big picture, the view that we have is that we want to be in a position where customers could say, "We're going to be Serverless first." Those customers could be in any industry, in any vertical and any size, building any type of an application. Then we want them to say, "Okay, we're going to be serverless first, and we're going to knock it down based on a roadblock or a limit or a challenge." Then part of what my team has to do is, "Okay, great. Tell me more about it. Tell me more about that story. Let me take that feedback and deliver it back to the PMs, so that we can start thinking about what is the potential solution that we built for it?"

I think in 2020, you're going to see my team and some of the technical evangelist, other folks inside AWS all over the place talking about Serverless. You're going to see a lot more blog posts, you're going to see maybe some more instructive guidance around certain topics. We're going to keep doing tech talks, twitch and in a lot of conferences and a lot of places. If you're lucky, you'll get to see Eric Johnson up on stage and Munns full of energy and excitement. You get to read the incredible stuff that the team is writing as well. Every year for the last five years Serverless is getting bigger and bigger and bigger, bigger. Lambda is typically one of the top 1, 2, 3 topics that are summits or even this year as well. The talks are super packed. This space just keeps growing, the customers keep doing incredible things. It's changed the way they build applications, and we just want to continue to magnify and grow that I'd say.

Jeremy: That's awesome. I know they just announced the Builder's Library and I'm sure that you'll probably contribute to some of that.

Chris: I'm not smart enough for that stuff. That's for the real experts.

Jeremy: Well, hopefully we get some good stuff there. Then also, I know, Heitor Lessa is working on a revamp of the Well Architected Framework with Serverless Lens.

Chris: Yes. Very cool stuff on there.

Jeremy: I think that'll be out soon, if it's not already. There's just a lot of good stuff, a lot of good information and like you said, and there's a bunch of great conferences now and all those Serverless state conferences, all around the globe, really interesting people speaking. Not just from Amazon either, I think it's really great to see what Azure is doing and see how GCP is doing and where they're pushing the boundaries where they're going. Because I think that, if it's solving customer problems, that's what AWS is focused on and that's great.

Chris: My view is that the deeper and richer the ecosystem, the better it is for everybody.

Jeremy: You can't be in an echo chamber.

Chris: Exactly.

Jeremy: Awesome. All right. Well, Chris, thank you so much for being here and taking the time to do this.

Chris: Of course.

Jeremy: If people want to get in touch with you and find out more about Serverless or what AWS does, how they do that?

Chris: You could find me on Twitter @chrismunns. You could also if you ever need to reach out about something deeper, I can give me my last name, Munns, munns@amazon.com. Real quick, I would say people often ask me, "Where can I find out about the latest greatest launches and things like that?" We post almost all of our content on AWS compute blog, and that's where Serverless does a lot of its announcements, how to post updates, things like that. You can go and search for AWS blog, you'll find it right away and see the most recent posts that we have.

Jeremy: Perfect. I will get all that into the show notes. Thanks again, Chris.

Chris: Cool. Thanks for having me. Take care.

View Details

About Farrah Campbell:
After 10 years of working in healthcare management, a serendipitous 20-minute car ride with Kara Swisher inspired Farrah to make the jump into technology. She has worked at multiple startups in many different capacities, eventually working her way to being the Ecosystems Director for Stackery in Portland, Oregon. As the Stackery Ecosystems Director, Farrah has managed the Stackery relationship with AWS including Stackery as an Advanced Technology Partner, achieving the AWS DevOps Competency, a launch partner for Lambda Layers. Farrah has cultivated the serverless community as an organizer of Portland Serverless Days, the Portland Serverless Meetup, along with numerous serverless workshops and the Portland tech community events from Techfest to bringing multiple luminaries to Portland. She's also an AWS Serverless Hero.

  • Twitter: @FarrahC32
  • Email: farrah@stackery.io

About Danielle Heberling:
Danielle Heberling is a software engineer with a background that includes being a musician and teaching at a K-8 public school. She’s passionate about building things that make the world a better place, whether that be through social change or a good laugh. When she’s not coding, you can often find her reaching back to her teaching roots by mentoring folks from underrepresented groups that would like to make a career switch into tech.

  • Twitter: @deeheber
  • Email: danielle@stackery.io

Notes:

  • Translator GitHub Repo: https://github.com/stackery/language-translator
  • Translator App: www.serverlessing.io
  • Stackery Blog: https://www.stackery.io/blog/
  • re:Invent Session: https://www.portal.reinvent.awsevents.com/connect/search.ww?searchPhrase=DVC16

Transcript:

Jeremy: Hi everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with Farrah Campbell and Danielle Heberling. Hi Farrah and Danielle, thanks for joining me.

Farrah: Hey Jeremy, thanks for having me.

Danielle: Thanks for having me.

Jeremy: So you both work at Stackery and we can talk a little bit more about what Stackery does in a bit, but I want to start with you Farrah because you are the ecosystems director there and I think it's a really interesting role. Can you tell us what that's all about?

Farrah: Sure. Well, essentially the way I look at it is my job is to connect with people across AWS and other technical partners along with the serverless ecosystem so that we can increase serverless adoption.

Jeremy: Awesome. And Danielle, you are a software engineer at Stackery, and I'm curious what that role looks like when you are building serverless applications to help people build serverless applications.

Danielle: Yeah, it's pretty meta actually. Well, at Stackery we're a small startup. There's only six software engineers total. So I guess you could say we're all technically full stack, so I just jump in anywhere in the stack where I'm needed. And sometimes do customer support too.

Jeremy: Very cool. So I saw the two of you give a talk at Serverlessconf, New York, called Leveling Up Serverless. And you talk about this app that you built and we will get into that in a minute, but what I really loved about your talk was the story behind it. And as you both explained, you have very different backgrounds. Neither of you started in tech, but somehow you sort of serendipitously came across serverless, started participating in the serverless community. And that's what inspired you and enabled you in a way to actually build this application. And I think your story is inspiring, especially to people who are getting into tech or thinking about getting into tech. So I'd love to just talk about your experiences today and we can go through that. So Farrah, let's start with you. How did you get into tech?

Farrah: Well, my intro to tech wasn't like many people's, I hear. In fact, it wasn't until after high school that I really actually explored the internet. My mom had met a new man that she married who owned a computer company called MicroAge and he started a new startup in the back of that where he had multiple engineers working.

And I talked him into letting me work for them to research the horizontal and vertical markets. Multiple years later being a single mom needing to pay the bills, I found a job in health insurance but always really still had that love for working with developers and would always find ways to try to work with the engineering and IT departments, which was awesome because when I moved to Portland there was all these tech companies and there's this very vibrant growing community.

And I wanted so badly to be a part of it, but it took me... Well, I applied for jobs for about a year without any... I didn't get any interviews scheduled, and then started working or volunteering at a conference where I was able to meet a number of people. One who introduced me to a small startup in Portland where I was able to start working for four hours a week and then quickly turned that into a full-time office management position.

And then multiple startups along the way ended up leading me to Stackery. The cool thing about it though is with each one of those startup jobs, I was always trying to really understand the tech and trying to find ways to work and be a part of the engineering teams. They even set up a way for me to update documentation so I would actually have contributions to our source control. And it's been pretty cool to be able to do something more with that at Stackery.

Jeremy: Awesome. So Danielle, I think you and I probably had a very similar start. I went to college with the intent of being a music teacher actually, and instead I ended up switching into technology. But you started with music and you kind of went from there.

Danielle: Yeah, definitely. So I grew up in a small town and in order to keep out of trouble, I just became very active in music throughout middle school, high school, took piano lessons starting in kindergarten even. I just really loved creating and sharing music, and it got to the point where I had to pick a college major to go to college and I really had no idea what I wanted to do.

So I kind of looked at what I enjoyed doing activity wise and gravitated towards music. And I really like helping people out, teaching people. So I decided to go for music education. So I was a music teacher in a public school system for a few years and it took me to a different state, a completely different mindset in that area. I definitely enjoyed my time there, but longterm teaching just wasn't for me.

So I went back in a whirlwind of not knowing what I wanted to do for a career, but I did need to pay my bills. So I just put out a lot of job applications and ended up getting hired in tech support. I really love tech support because every day was a completely new challenge. A customer would write in with a question about how to use the product and it would be like something that I never even thought someone would want to use the product for.

So it really kind of gave me a great lesson in empathy and kind of role playing in how other people think. It got to a point where I was supporting the developer version of the product. So the customer's writing in were people writing custom code to add to our platform offering that we had. And it was my job to determine is this a bug in our system or is this a bug based off the code that the person added. So I did some free online tutorials in order to be able to have more educated conversations with the people writing it.

And it goes back to being a musician, just creating things. And I build a few silly websites and just really loved doing that and sharing what I created with other people. It was just a different medium. So music to code, and as far as customer support goes, I just wanted to be the person to provide the better experience for the customer. So that's kind of what led me to attending a code school and becoming a software engineer.

Jeremy: Yeah. And I actually, I think that a lot of software engineers don't have enough experience at the product level and so they don't have that customer empathy that you were talking about. And I think your experience dealing with that actually it makes a better software developer, because I think we too often get sort of bogged down with the tech and we don't think about that overall customer experience.

And I totally agree with you on the parallels to music. I mean writing software is definitely more of an art in my opinion than it is a science. And it is really, really rewarding to get to a working product and get that out there, and knowing you've created something. So anyways, so let's talk about how you both got started with serverless. Farrah, Stackery was your first foray into serverless. So how did you get connected?

Farrah: The last startup that I was at, it was called Reflect, was being acquired by Puppet and that fit wasn't going to be for me. And I had put a few feelers out within the Portland community and actually got introduced to the CEO of Stackery at the time who created this role for me. Something that was completely out of anything I had ever done, but I'm always willing to tackle a new problem and I always like a new challenge.

Jeremy: Awesome. And Danielle, you are a container convert, correct?

Danielle: I am, yes. Fresh out of code school, my first job was ECS. We did all containers. Luckily it was microservices, so it wasn't too much of a mind shift set on that front. But ECS is how I first got my feet wet with AWS and cloud services.

Jeremy: Right. So then you both got involved in the Portland tech community and attending meetups and things like that. And that's actually how the two of you got connected.

Danielle: Yeah, so I'm not sure how it is in other cities, but in Portland pretty much every single tech meetup has their own corresponding Slack channel. And Farrah I just happened to be hanging out in one and she mentioned Stackery, talked about how they were hiring software engineers. And the one thing that really appealed to me was she said that half of the software engineering team was women.

So on top of that I was also starting to get my feet wet with serverless. So that's kind of what prompted me to message her just to learn more. I wasn't even looking for a job, but I asked her out to coffee, just kind of more of an informational interview and somehow I ended up applying and now I'm there.

Jeremy: That's awesome. And that's the power of meetups and conferences, right? Where you go and you meet new people, you get new ideas, you expand your list of contacts and just really, really good things can happen. And if you have the ability, if you get the chance to be able to go to a conference or meetups, I definitely suggest you do things like that.

So you had mentioned that half of the developers at Stackery or on the Stackery development team were women. And so this is obviously something we don't see often enough in tech. We're not quite there yet. So is there something about serverless and the serverless community that is more welcoming to people you think?

Farrah: I definitely do. Serverless is a new approach. It's inventing a whole new way to build software applications and it's sticking. And this community, I feel like everybody has a lot of work to do, a lot of big things to accomplish and everybody's at a starting point. Everybody's willing to have open conversations without putting others down or explain to you why you're wrong about something. Everybody's at the same starting point and just trying to learn from one another.

Jeremy: Yeah, I totally agree. And I think one of the things that's great about serverless is the fact that we are still very, very early and we're still figuring things out. I mean we're still developing the tech, let alone the tools and the best practices that go along with that stuff. So when you come from a more traditional sort of computer science background, because things are so different in serverless or some things are different, you need to unlearn quite a few things to be able to do this. And I think for most people this is ingrained in their heads and they probably resist the change. So coming from a fresh perspective is probably a benefit.

Danielle: Yeah, I agree. And I think just because everything's so new and serverless is a lot more than managed services, but that is a big aspect of it. So there's always eyes on the new thing, AWS or GCP is releasing and people are talking about it and how they're using it, planning on using it. So that's what makes it fun too.

Jeremy: Yeah, definitely. All right. So let's talk about this app that YouTube built, and we'll get into the details in a minute, but first of all, like what was it that made you say, "Hey, I want to go build a serverless app?"

Farrah: I actually went to Danielle because there was, I'd say there was a lot going on in their lives at the time and sometimes when that happens you want to take steps forward and try to grow. And so we knew we wanted to work on something together but weren't quite sure what it was. And as I started doing tutorials through Stackery's website, I was able to build my own web hook. I was able to follow along the Wild Rydes tutorial and build my unicorn app.

And knowing that Danielle has been a teacher before, I went to her and asked her if she might be willing to help me to build a serverless application. I then would get the experience of learning how to actually build a serverless app, and at the same time Danielle would have, I think, more experience trying to teach somebody and build something from scratch.

Danielle: And I also think it's a lot of fun to work with someone who's brand new to something because they have that beginner's mindset, that endless optimism. So it can really be infectious in a good way.

Jeremy: Right. The anything's possible until you run up to deploying a cloud front distribution and it takes 45 minutes. But anyways-

Farrah: I was always worried that it would be terrible working with a beginner, because I seem to just run into problem after problem after problem after problem, and Danielle was always so incredibly patient and understanding, which made things so much easier.

Jeremy: Yeah. But that's just how we code. I mean there are probably some people who can write perfect code the first time through, but for the vast majority of us it is a lot of trial and error. We have to test our code and we make a lot of mistakes. We have to test it and test it to get it working properly. So if you are new to development or you are a experienced programmer, that is a perfectly fine way to do it. So certainly don't feel bad about that. So anyways, so you built a language translator app. So what was the motivation behind building this?

Farrah: So the motivation for building the app is we knew we wanted to do something. We knew we wanted to have some sort of positive theme or to be beneficial to our community in some way. And after a trip to Australia, I got to attend the AWS community day in Melbourne and got to meet an incredibly talented group of people.

But in these conversations, and we were talking about what are the barriers that every single one of us has had to developing or gaining new skills. And they started talking about documentation being one of them, and I had never ever thought about that being a barrier before. And so I brought the idea back to Danielle and said, "Do you think that maybe we could build a language translating application or something?" And she was stoked about it.

Jeremy: Awesome. So then how did you get started? How did you plan what you were going to build?

Danielle: So when it comes to planning the app, I just wanted to do something as simple as possible just because I didn't want Farrah to get overwhelmed and frustrated and not want to do this ever again. So a simplest thing possible that I could think of was just kind of a small backend pipeline that could be plugged in elsewhere with serverlessing.io, I added a front end onto it, but I use like that core architecture of, it's a bucket function bucket. So the first bucket takes in a plain text file of what someone wants translated, the function talks to the AWS translate service and then gets the translated text, and then it puts it in the second bucket.

Jeremy: So does it just do text right now or does it do like HTML files as well?

Danielle: That's actually something we'd want to improve in the future, because we were playing around specifically with markdown files because the idea of translating documentation, and a AWS translate was giving us some unexpected results in regards to spacing and things like that. So we actually have an open GitHub issue for that, if anyone wants to help out.

Jeremy: That's actually something you find with working with a lot of these managed services is that either there are limitations or there's these nuances, and sometimes you have to dig to find them and you kind of end up banging your head against the wall to try to figure these things out.

Danielle: Yeah, definitely.

Jeremy: So anyways. So that's really interesting. So then how do you-

Farrah: That was actually... Can I talk about that just for a little bit?

Jeremy: Yeah, go ahead.

Farrah: Because this is actually, I think, one of the hardest parts as a beginner to development, is understanding what the error code messages are, understanding where should I start to debug, when should I research? Is this a limitation on the actual service? We were getting all kinds of error messages and couldn't figure out what was wrong, and we came to find out that AWS translate only processes 5,000 bytes, which is about two paragraphs, which is not something that I would have initially thought and looked for. And so it was definitely a huge learning experience for me to understand how to understand what these error messages are, where to start to look. And again that that just basically is part of the development life cycle.

Jeremy: Yeah, another great point. Another great point. All right, so this exists though, right? It's up and running now, we can go and use this if we wanted to?

Danielle: It is, yeah. If you don't want to add code to it, you can just upload a plain text file at serverlessing.io, and if it's a longer file it might take a bit to translate for you, but be patient and it'll spit out the translated text in a language of your choice.

Jeremy: And you open sourced this, right? So that people can download it if they want and play around with it.

Danielle: Yeah, feel free to download it, play around with it. We also have some open GitHub issues. Feel free to jump in on one of those if you want to work with Farrah or I or both of us on it collaboratively, we're always into working with new people too.

Jeremy: Great. All right, and so what are you going to do with this application next? Got big plans?

Farrah: That's a very good question. That's a very good question. I mean there's definitely some plans. I think the next thing that we want to do is to be able to process more than just plain text files. And then also my next is like adding Polly. So we'd have talk to text or text to talk.

Danielle: Text to speech.

Farrah: Text to speech.

Jeremy: Yes. So you are dealing with a bunch of cloud resources here, S3 buckets and Lambda functions and things like that. So you used infrastructure as code to build this, correct?

Danielle: We did, yeah. We started with the drag and drop interface of Stackery, but we also ran into the classic circular dependency issue with the S3 bucket. So we did have to manually edit the CloudFormation template a little bit.

Jeremy: And how did you find editing a CloudFormation template to be?

Danielle: A lot of fun?

Jeremy: It is. It gets-

Farrah: I thought it was fun. I actually did think it was fun. I actually felt pretty cool at that moment because we were in the CLI and it was really cool just to... if you start to understand the lingo that's being used, like we're navigating through the JSON object to drill into the function. Like those are all big learning moments for me. So I thought it was fun.

Jeremy: That's awesome. And CloudFormation is one of those things where if you stare at it for long enough your eyes will go cross. But once you figure out what it's doing, it's extremely powerful. And the Stackery interface, and actually let's talk about that for a second. So you can drag and drop these resources. You can wire them together, and it's not just the drag and drop either.

Like you can edit the information, and edit those configuration options. There's a bunch of sane defaults in there, and if you've ever had to write a CloudFormation template by hand, then you know you go to the documentation pages and you find certain attributes and then under each attribute there's a complex type. And then you've got to go and look at the details there, and then there's a complex type under that and it just keeps going and going. And so this is something that the Stackery, the visual builder, helps you do really easily and just abstracts all that away, correct?

Danielle: Correct. Yeah. And actually when I first got hired at Stackery, it was just drag and drop. We didn't have the CloudFormation template surfaced in the UI yet. So that was actually one of my first big projects, and it took a while for me to realize how huge it was going to be. But it really is, especially with our new VS code editor where you can have a tab of the template and then a tab of the Stackery visual UI, and you can physically see it change right in front of your eyes. It's really cool.

Jeremy: Yeah, definitely. And it's all local, right? So I mean there's no uploading or anything like that. It's just a very, very cool and simple way to build a CloudFormation. So if you are struggling to write CloudFormation templates, definitely take a look at stackery.io and check out the tool that they have available there.

So anyways, so this episode is actually airing during re:Invent and the two of you are giving a talk called The Power of Serverless for Transforming Careers and Communities. And that is Thursday at noon in the Dev Lounge in the Venetian. So if you're at re:Invent, you're listening to this and it's before noon on Thursday, definitely go and check out Danielle and Farrah giving this talk. And so this is going to be about the app that you built and again, the motivation and how you came to build it, right?

Danielle: Correct. Yes. It will start with our backgrounds and then the inspiration for the app, and just kind of that whole storyline behind every time you build an application, the roadblocks that you go through and how you get yourself unstuck.

Jeremy: All right, so you'll get into a little bit more about the debugging and some of that stuff, right?

Danielle: Correct. Yeah, we'll get more into the storyline of that.

Jeremy: All right, cool. So Farrah, anything else going on at re:Invent that people should know about?

Farrah: Well, there's a lot of things going on at re:Invent. Oh my gosh, I'm still trying to keep up with all the recent announcements that keep happening even up until today.

Jeremy: I know, it's crazy.

Farrah: Well, we're having a party, we're hosting a party Wednesday night. I'm hoping to see many friendly faces there. Please come say hi if we have not met before. I know I've met a ton of people on Twitter and I love to put a face with the name. And yes, please come to our talk, let us know how we did and let us know what we could do better.

Jeremy: Awesome. All right. Yeah, so Wednesday night is the ServerlessForEveryone party. There are so many people signed up for it. I think we're going to have to take over re:Play next year and that'll just be the serverless party because this is getting bigger and bigger, and definitely I will be there. Come talk to me, talk to Farrah, talk to Danielle. The serverless community, honestly, everyone is so approachable. There's just so many great people. It is such an honor to be part of all this. Oh, and by the way, Farrah, you were recently named a serverless hero.

Farrah: I was. I still am taken aback by this entire experience. I was reading the, the RedMonk article about it last night and just tears were running down my face. It's such an honor and such an achievement, and I really just appreciate everybody at AWS for designating me with that. To join people like yourself, Jeremy. It's pretty awesome.

Jeremy: You definitely deserve it. I mean with everything you've done with like the Portland meetups and Portland ServerlessDays and all the outreach that you've done, I mean you're just a huge part of this community. I mean you both are. So you know, thank you for everything that you two do. So if people want to get in touch with either of you, Farrah, how would they go about doing that?

Farrah: Twitter is by far the easiest way and the fastest way to get a hold of me. I'm just at @farrahc32, and then also you can always reach me on LinkedIn or by email. And I think my email is listed on the AWS heroes website.

Jeremy: Okay. And Danielle, how do we get in touch with you?

Danielle: The best way to get in touch with me would probably be a DM on Twitter @dheber, you could also send me an email, danielle@stackery.io.

Jeremy: Okay. And then the Stackery blog is stackery.io/blog and then the translation app that you built is available at serverlessing.io.

Farrah: Right. And you can also reach us through GitHub too, through the GitHub link, which is https://github.com/stackery/language-translator.

Danielle: Yes. Yeah.

Jeremy: Perfect. All right. I will get all of that information into the show notes along with your contact information and where you're speaking on Thursday. This was great. Thank you so much again for being here.

Farrah: Awesome. Thank you.

Danielle: Thanks for having us, Jeremy.

View Details

This is PART 2 of my conversation with Ory Segal. View PART 1.

About Ory Segal:
Ory Segal is a world-renowned expert in application security, with 20 years of experience in the field. Ory is the CTO and co-founder of PureSec (acquired by Palo Alto Networks), a start-up that enables organizations to build and maintain secure and reliable serverless applications. Prior to PureSec, Ory was Sr. Director of Threat Research at Akamai, were he led a team of top web security & big data researchers. Prior to Akamai, Ory worked at IBM as the Security Products Architect and Product Manager for the market leading application security solution IBM Security AppScan. Ory authored 20 patents in the field of application security, static analysis, dynamic analysis, threat reputation systems, etc. Ory is serving as an officer of the Web Application Security Consortium (WASC), he is a member of the W3C WebAppSec working group, and was an OWASP Israel board member.

  • Twitter: @orysegal
  • Prisma by Palo Alto Networks: https://www.paloaltonetworks.com/prisma
  • The 12 Most Critical Risks for Serverless Applications: https://www.puresec.io/serverless-security-top-12-csa-puresec

Transcript:

Jeremy: All right. So let's move on to number four. So number four is over privileged function, permissions and roles. This is one of my favorites because I feel like this is something that people do wrong all the time because it's just easy to put a star permission.

Ory: Yeah. And this is an issue that I've been thinking about a lot of, why is it like from a psychological perspective that developers put a wild card there? So, obviously, we've talked about the very granular and very powerful IAM model in public clouds, and that's very relevant to serverless. You break your App down into functions, you assign each function you need to assign to each function, the permissions that it actually needs in order to do its task and nothing more than that, and that's the point here. How do you make sure that if somebody exploits the function, if somebody finds a problem in the function, they are not able to manipulate that function to maybe do some lateral movement inside your Cloud account, move to other data stores, etc?

So that's very important and we see that developers have a tendency, and this is one of the most common issues, to just use a wild card and allow the function to perform all of the actions on certain resources. And that, as I said, this is something that I've asked a lot of developers, why are they doing that? And I'm hearing different answers. Some are just lazy, I have to admit, I do that from time to time as well. It's much easier than actually having to go to the documentation and figure out the name, the exact name of the permission that I need. The other set of developers talked about future proofing the function. So they said, okay, now the function only puts items into database, but maybe next week I'll need it to read, which by the way violates the principle of single responsibility, but let's put that aside. And so they did just put maybe crude permissions or they put everything.

And then there are those who just either don't care, or don't know, or are not aware that this is a problem. So those are the three types of developers or answers that I've heard. But this is by far the most common and I've seen frameworks that automatically generates wildcards as well, which is also bad. And I've seen some bad examples as well in tutorials, which is the worst thing this can happen because we're trying to teach people how to write [crosstalk 00:48:21].

Jeremy: To go the other way. Yeah. Well, so the example that the document uses is the Dynamo DB star permission. And I love this example because you would think, okay, put items, get items, query items, delete items, that seems like that's what I'm giving it permission to. But, no, when you give Dynamo DB star permission, you're giving it the ability to delete tables, or change provision capacity, and you can do a lot of really bad stuff there. And obviously, this is all predicated on someone actually being able to get into your function, but that is something that is possible ... again, it's limited in how you can do that but it certainly is possible. We'll get into that more.

Just one point about about the permissions per function. One of the things that I like to do is I try to give each function the permissions that I think it needs, then I publish it to the cloud, and they try to run it. And then actually, AWS does a great job of giving you the error saying, this function doesn't have Dynamo: put item permission or something like that.

Ory: You're basically using debug branding to figure out the right permissions, right.

Jeremy: It works.

Ory: It's not nice. But they do have a service I think called, access advisor, and I think Google came out with a much better automated solution for that. Eventually, I think the access advisor would look at historical logs for, I don't know, a few days or a few executions and will tell you, it looks like you have too much permissions, you should probably reduce them. But this is something that we've done in Pure Sec with the there's the OpenSource list privilege plugin that we wrote that you can use for serverless which basically statically analyzes your code, extracts all the API calls and then maps them to the list required privileges, and will actually generate a list privilege role for you. So this is, and we'll talk about the future, but I think this is something that will eventually will have to be solved somehow or Cloud providers will probably produce better tools around that.

Jeremy: Well I think some frameworks are doing work with guard rails and stuff like that too that help a little bit but it's not quite there yet.

Ory: Mostly around the asterisks around the wild card, but actually-

Jeremy: What you actually need.

Ory: Exactly. Yeah.

Jeremy: All right, so let's move on to number five. And this is probably tied to number 10 in a way. So number five is inadequate function monitoring and logging. And I think this extends a little bit to number 10, which is improper exception handling and verbose error messaging. Whereas logging is a good thing [crosstalk 00:51:04] but can be a bad thing, right. So let's talk about inadequate function monitoring and logging first.

Ory: Well, I think looking at this issue now in hindsight and seeing that there's an entire industry of serverless monitoring vendors and solutions, I think, we already see that this is a real need. And it becomes more critical for security, not talking about performance and tracing and things like that, but being able to properly monitor your functions to log the right thing is critical. And if you look at this, for example, if somebody runs a sequel injection attack and triggers some exceptions, where would you even see that? I'm not even sure you would see that in the default logging facilities that the cloud providers give you.

So it goes back to the fact that developers have to worry about application security specific logging. And, yeah, so you have to write more into Cloud watch and if it's related to IAM and things like that, then you would see that in cloud trail, I'm talking AWS of course. But yeah, without that you're pretty much blind to the attacks that you're experiencing. Yeah.

Jeremy: And then you have the whole issue too, where is if it doesn't necessarily trigger an error and you're not capturing what the original input was, then how do you inspect that input? If you're not logging that input somewhere, it's not just getting logged automatically for you like an access logs. I mean, obviously, if you have API gateway setup you can enable access logs and you can see some of that. But even the post data and some of these other things are definitely hard to see.

Ory: That was never available. By the way, if you looked historically at Apache logs never logged the post bodies and for good reason. And I think from a PCI and data privacy perspective, you don't want them logged all the time, maybe just when there's a security exception.

Jeremy: Yeah.

Ory: It's the same for for serverless. I don't think you can actually enable full event logging end to end...

Jeremy: I think some of the observability tools let you though.

Ory: Yeah.

Jeremy: Which can be somewhat dangerous if it contains information that it probably shouldn't contain. But anyways. So, yeah, I totally agree though. I think that logging is important to make sure that you have visibility and certainly the monitoring tools are helping with this. All right. Okay. So let's move on to the next one because this one I think is probably the biggest security hole in any type of application, this is not specific to serverless, but certainly something that if it was taken advantage of, and again, this may be theoretical, it could cause pretty bad or pretty dangerous side effects. And that is number six which is the insecure third party dependencies.

Ory: Yeah, so I think I read somewhere recently that third party libraries make up, I think, 75% of the code we right today, is actually coming from external third party non-trustworthy resources. And that's, as you said, that's really not specific to serverless and this is also something that I mentioned when I give my server security talk is, this one is not specific to serverless. I think the only difference that I like thinking about when talking about the third party depends, how do you monitor those dependencies? Because you don't see them running, it's not running in your environment. If for some reason it's leaking data or sending your credentials or your API keys to some third party, you have no perimeter to block that from happening and you can't really monitor their behavior. You have to somehow run them locally and monitor their behavior. So the problem is, I think it becomes a bit harder to locate that you have infected third party dependencies.

I think another maybe thing to think about is that in serverless, you always reduce the amount of code in a function to minimum, as we talked about the principle of single responsibility. And so a lot of time you rely a lot on third party libraries to do some of the heavy lifting for you. I think you can't even write a serverless function without at least do one import, you have to import JSON for example, if you're talking about Python, right. So you start by already importing a library and then, frankly, nobody is monitoring these functions. So, I have a severe trust issues with open source projects. Nobody is keeping keeping us safe and if we're talking about all the the snakes and the white sources, they're very good in listing non vulnerabilities like CVE type vulnerabilities, but I think all of the all of the malicious packages or the malicious code that was injected into open source packages lately, was only discovered, I think I did some research, more than three weeks after it was injected. So there's nobody really monitoring this and there's no real solution to tell you that somebody injected malicious code in there. And that's an issue.

Again, not a serverless issue, but it becomes worse when you don't have any tools to control the environment?

Jeremy: Yeah. And I think this is the, we're going to talk about remote code execution, people can't just hack into your function, right, that's not a thing. It's not like a traditional server where you're going to see that happen. But remote code execution, I would say, third party dependencies is probably the number one overwhelming possibility or exploit where people could use remote code execution. And if you look at a lot of tools that people are writing, again, like you said, this applies to every type of service you're building. But certainly, you look at a lot of the tools that have been built, a lot of open source tools, they run sequel queries, right. So how would you even detect that this runs a sequel query? Maybe it's supposed to be running that delete, maybe it's supposed to be running that select scan on the Dynamo DB table. You'd have to really understand what each one of these dependencies is doing to know whether or not it's doing something it's not supposed to.

And obviously posting that data to some third party service is probably the the easiest way for these RCEs to get the data somewhere else, and that's something that, again, the cloud provider does provide some mitigation to if you're running into VPC, then you can disable outgoing HTTP calls. I know Pure Sec, did some work around that as well and that's part of the system there. But yeah, I just think this is one of those things where people say, oh, there's a service that does this or a plugin that does this, I'm just going to use that. And what you don't realize is that you may be installing hundreds of dependencies, each one of those maintained by different people of questionable intent or of questionable reputation. And we just do it because it's easy, but yeah, this one scares me the most and I think it's a very good practice for you to be very conscious of the third party dependencies in you use.

Ory: And if you remember, I think we talked earlier about my wife's WordPress [crosstalk 00:59:09], right, blog. Then WordPress itself, I think the amount of vulnerabilities found in WordPress is rather low but it's usually those third party plugins that you add that God knows who wrote them, and what's their background in application security. And so, the majority of holes in CVE's and exploits you see around WordPress, is because of those plugins, and I think the same applies to application security for modern applications and for serverless in particular where you can write the perfect application, and do threat modeling, and pen testing, and everything properly an input validation, but then you're using some de serialization package which includes a remote code execution, and that's it. That's the weakest link in the chain. So yeah...

Jeremy: And you may have followed every other best practice and it doesn't matter. Which I think ties into this idea which, besides the third party just being able to somehow execute a little bit of snippet of code, part of this is, again, for the purpose of maybe stealing application secrets, the tokens and the session tokens and things like that. So number seven is this idea of insecure application secret storage. So what's that about?

Ory: Well, all applications, most applications, use secrets, API keys, passwords, I don't know, whatever. And well, the security of those secrets really depends on where you store them. And I think we've seen again, and this started because of some bad examples and tutorials where people stored obviously secrets hard coded, which is the worst, and then they push the project to Git or something and it leaks out. And then later on people use the environment variables because that was the best or maybe the only way to store in the stateless world of serverless applications and people store them there, and those have a tendency to leak pretty quickly. And I have a good example of that in the presentation that I usually give.

And at some point, I think cloud providers decided to solve this and now they all I think offer secret storage that you should be using with some KMS, some key management, and you encrypt the secrets. That's all terrific and you should be using that. Keep in mind though that if somebody manages to run code, if you end up with an RCE with remote code execution, then those secrets are not secrets anymore. So that problem hasn't been solved yet. So if somebody can run code on behalf of your function through some RCE, then it's game over. I haven't seen a solution yet for that.

Jeremy: But I do think that environment variables are the more dangerous place to store those because it's very easy for a third party service to just look at the environment variables, that's a standard place to look. If you did store them, if you used accessing them at runtime, even cashing them in the function, but pulling them and storing them in global variables that weren't necessarily accessible through the environment variables that the attacker would have to do something a little bit more tricky, they'd actually have to redo some static file analysis to find where these variables are being used and then try to load them that way. So it might be harder if they're not stored in an environment variables anyways.

Ory: I would love to see cloud vendors make use of technologies like the Intel SGX secure enclaves technologies where you can store things in a very secure manner, and even if somebody has remote code execution, they won't be able to read them. We're not there yet but the technologies exist, I think at some point, they will make use of it. Yeah, I think we even said too much on this.

Jeremy: Yeah probably. But speaking of securing or having the cloud vendor do more, Microsoft Azure Functions actually just released something where now their secrets are available, but you don't have to do any code or have to add any code to access them. So, I don't know what that's all about. I have to look more into it, but that seems like an interesting approach. All right, so anyway, let's move on SAS eight. This is the denial of service and financial resource exhaustion.

Ory: So I don't think we need to spend time talking about denial of service specifically, I think there's some interesting aspect of denial of service especially in serverless environments where you can cause that financial resource exhaustion. So if you find the function that, with some input, will work harder and longer, you can definitely inflict some financial pain if you'd like. I think on the denial of service side, the only thing I want to say is that the auto-scaling and whatever infinite scaling or whatever people claim serverless platforms to be, you have to pay attention to that. You have to design with scalability in mind, it's not all for free. So you have to think about how you design the application, use the right services, think about what is invoking what, what are your timeouts, how your application is going to handle these things. Just because you're using Kinesis in the serverless function doesn't mean that you're infinitely scalable as an example.

Jeremy: Right. Yeah. And certainly flooding queues and things like that are other ways to raise the costs or to slow down the execution of other important messages. And that's things to where just good practices there are validating the shape of the data when it's coming in through something like API gateway or making sure that you're not flooding certain downstream services. But I agree, that's something that I think is well written about so there's lots of ways to find information there. Probably same thing here with SAS nine which is serverless function execution flow manipulation. So thoughts on that.

Ory: I think we changed the name later on to serverless business logic manipulation.

Jeremy: You did. Yes. Serverless Business Logic Manipulation.

Ory: But it's same thing. I think, again, this is not specific to serverless but rather to more service oriented architecture, web services, service mashes, etc. where, again, we're back to the fact that we broke the application down into many, many tiny laser focus services, and now they only interact with each other and who knows if something is even enforcing the flow that you expect it to do. And if you think about an application where you have an API request coming in, triggering a Lambda function that then put some data inside a bucket and that triggers an event that triggers another function that stores something inside IQ, which eventually triggers another function. At least from what I know, there is no way for me to enforce the order, the order in which the ... unless I'm using step functions and then it's a whole different game.

But if I'm building at the classic traditional way, if you can say traditional in serverless applications, but if I build it the normal way, nothing is actually promising me that the services are invoked in the right order. And that one service can, with 100% certainty, say that whoever invoked it was in fact the service that the it expects to be invoked.

So, as an example, you have the API triggering the Lambda function stores in a S3 bucket, and then that triggers the rest of the chain. What if some insider or a developer stores or throws the file into that bucket and starts the chain from the middle? Is something even validating that? And the default, when you create those applications with some framework or SAM or serverless framework, it's not that the default is deny all and then you loosen up the security permissions a bit. So, anyone with execution permissions can run any function in the account, and anyone with access to the bucket can drop files into the bucket. And so you have to think about that. And I think that in the future, at least, this will change and you will be able to create applications that the default is deny everything and then you say, okay, to this bucket, only that function can write and anyone else is blocked. And that's how I would want to see things evolve.

Jeremy: Yeah. Totally agree. All right, let's move on to number 10. So this is improper exception handling and verbose error messaging. And the reason why I compare this to number five, which was the inadequate logging, is because you have a tendency to log too little, but then you also potentially have the tendency to log too much sometimes.

Ory: Yeah, just like how you debug IAM permissions. We all do that. So we all do debug prints and then sometimes we leave them there, sometimes we don't catch the exceptions properly and that spills out just looking things like show down or Google you can do Google Dorking and find a lot of very juicy verbose error messages. And in serverless, I think because at least at the time of writing of the documents, the debugging capabilities that you had weren't on par with traditional applications where you're writing an IDE and you can debug line by line. There is a tendency that we see that people write verbose error messages and then leave them there. So that's how this is related to serverless specifically. But yeah, this is a classical error in any type of application.

Jeremy: Absolutely. AlL right, so this is actually somewhat specific more towards I guess cloud native, and this is number 11. This is legacy/unused functions in cloud resources.

Ory: Yeah, that's a good one actually. Just I think a couple of days ago, I was giving a presentation at the London serverless conference, and we had a booth and I was showing the serverless radar that we have, and somebody asked me, what there's more than, I think, few dozens of functions just lying there on the radar not related to the demo I'm giving? And I said, okay, I wrote these functions, I left them, some of them are two years old. I am not even deleting them. Why would I delete old functions? And they're just lying there with the IAM permissions waiting for somebody to access them, and invoke them, and exploit them.

So, I saw I think somebody published a serverless pruning plugin for serverless framework lately.

Jeremy: Yes.

Ory: Yeah. So I think in any Cloud account that I have, there are hundreds of functions just lying there, and I think last I checked was in most of my accounts, I had like 600 rolls that [inaudible 01:11:38]. Lambda execution role, one Lambda execution role, two and three, and nobody's pruning them and until you hit the limit, the account limit and then it will take you a few years before you start getting rid of unused resources and there's really no reason for them to stay there. Again, not serverless specific, serverless does bring a new set of resources that you can leave behind, so mostly functions.

Jeremy: Yeah.

Ory: Yeah.

Jeremy: I think that's one of those things I am the same exact way. I go in I look and I'm like, I don't even know what this function does, I don't know how it got published, I probably deleted the folder that actually published the actual app, locally that published it somehow. And I'm like, I don't even know how do you go and then trace everything? Maybe it wasn't even unpublished through cloud formation so now because it's like a test you are doing or something, so yeah, that is definitely something you want to clean up.

Ory: It took me I think three days to find a function. I had a function that scans the entire set of S3 buckets that we have publicly open and sends me an email every day. That was like an experiment that I have done I think more than a year ago, and I couldn't figure out even What account, what region, and what's the name of the function that was sending me the emails. And it took us I think three days or even more to eventually manage to find where it is by looking at the email hitters, figure out which AWS zone, not zone, but region it is, yes. And then, as you said, I completely deleted the project so it's not like I had something to delete the serverless framework project that I was using, yeah.

Jeremy: That's crazy. All right, so the next one is SAS 12. That's cross execution data persistence. And I think this is a really important one. I mean, obviously, in traditional applications, we have global variables that can get manipulated and are reused across each execution. But with serverless we think about, we're used to this single execution model, but we do save data outside of a handler which gives us the ability to reuse that on warm invocation. So what's the issue with that?

Ory: I think you explained it very well. People don't think about the fact that the same container is being reused. So it's frozen between executions but then it's revived, or defrosted, or whatever they call it, and the environment stays the same. So if you had anything stored in /temp, which is basically the only place where you can store things locally, or if you're storing things in the environment variables like session variables, it stays there. And in fact if the next execution belongs to a different user who's malicious, if again, you screwed up in one of the other top 12, the other 11 of the top 12, there is a chance that this information will leak.

And I don't think there is a way for you to automatically flush everything between executions. You have to code that. So delete or destroy what's in /TMP, throw away environment variables, and start from fresh, if you want. Again, that really depends on the application itself. But this is something that you have to keep in mind that stays there and we've seen a few examples, some demos where people store stuff in /campaign and then somebody comes in and grabs that data. Yeah, just something to keep in mind that ... and actually, this goes back to the theoretical versus practical. This is a classic one where we haven't seen this getting exploited, this is purely theoretical, we added it to the 2019 because we believe that this is something ... This is almost obvious that this is some kind of vector that attackers will target because that's pretty much what's left behind with Linux executions. So it only makes sense that somebody will eventually use that somehow.

Jeremy: Yeah. And actually, I mean, not even putting this in the security context but just putting it in the things you can do to step on your own toes from a security standpoint is, if you're not fully aware of what data you're saving in global variables from execution to execution, it's very possible that user A accesses this function and then user B gets that thought version of the function again on the next invocation. And if you've saved data about that user or some other information about that user that you might then share back, maybe you're a pending rose to an array or something like that. So these are things you can certainly basically make your own application and secure by just not knowing that this exists. So I do you think it's important. Yes, it's theoretical maybe from an exploitation standpoint, but certainly possible to do on your own if you're not paying attention.

All right, so I think that's a ... I mean, we went through I think in more depth than I was planning to. So we've been talking for quite some time and I hope that what people are taking away from this is just, again, not all of this stuff is specific to serverless, but it certainly is things just to pay attention to.

So, I want to wrap up though before I let you go. I do want to talk about the future of serverless security because I do see things changing quite a bit just from two years ago or a year ago in terms of how people are starting to address these things. And one of the things that I read recently was, this was a study, I don't know how valid it was but something like 63% of containers, so containers, not serverless, but containers run for less than 10 minutes and then they go away. So I think this nature of a femoral compute, even if it's containers, and it's Kubernetes, or Fargate or any of these other orchestration management systems that the time that containers are actually running to the time they get recycled is very, very low. So a ephemeral compute seems like, whether you're using containers or serverless, is the future of the cloud. So just what are your thoughts on that, and what does security look like next year or in the future?

Ory: Interesting. Wow.

Jeremy: Sorry, I packed a lot in there.

Ory: A lot of different ... No, that's good. That's good. It's a very interesting topic. Let's start with containers, and ephemeral, and server lists, and all of that. So, a container is a technology of how you package applications, existing applications that we had previous technologies like web apps and databases, you just package them as containers so it's easier for you to redeploy them. And so you can think of containers as a technology that actually came from maybe the operations side of the world and DevOps. That's a technology that comes to help them to more easily deploy and maintain infrastructure.

I think, on the other hand, serverless is something that is more geared toward developers. It's a technology, or an approach, or a platform, whatever you want to call it, that comes to help developers with getting rid of the need to actually maintain infrastructure and rely on IT teams, and it's much easier to now deploy applications and go into a very fast agile CICD deployment cycles where you push changes, you don't have to ask anyone, there is no gate, nothing is slowing you down here.

So, I don't think that looking at the fact that most containers run for less than 10 minutes means that we should all be migrating to serverless, that's not the right reason to do that. You might still run containers and feel comfortable with the way you package existing technologies, web servers, databases, etc. You might want to consider using services like Fargate where you don't own the infrastructure and you don't care about the underlying container orchestration. But the packaging is still more comfortable or easy for your teams to do through containers.

So, again, the running time, I don't think that's the right reason to go serverless. Again, if you have containers, they run for 30 seconds and then you destroy them. Maybe it's still easier for your ... you need some things that require you to package them as a contain. So that's just a comment about containers versus survivalists and which one is going to take over the world and that's the whole thing. About security and where this is all going, if I remember correctly, the discussion.

Jeremy: My convoluted question, yes. I mean, with the ephemeral compute nature of this, how does security apply differently to ephemeral compute versus what we've traditionally seen with long running resources?

Ory: Yeah, that's a discussion I actually don't like to do. I remember, in the early days of serverless, people talked about the fact that it's ephemeral, it's not staying there, you can't infect it, malware is irrelevant, all of those claims. You can't store anything there. I think we've already demonstrated that, that's not the case. You can infect serverless functions. Again, it depends on the vulnerability and the way you exploit it. But the fact that it's ephemeral, you can overcome that by reinfecting. So let's say you have a remote code execution through an HTTP API call, I can reinfect the function, it might not be the exact same instance, but again, it depends on what I'm trying to achieve.

On the other end, the likelihood of me using that to get into your network, that's pretty much done. So, yeah, the fact that those platforms are ephemeral, I think has its benefits and its drawbacks both from the attacker and the defender side. At this point, I don't think this is something that's worth paying too much attention on.

Jeremy: But I just I'm curious about the visibility of it, right. So I mean, I think that especially in being able to inspect log files and some of that other stuff, just the ephemeral nature of some of these compute models now just seems like it's probably harder for the defender to put all this information together and make sure that their application is secure.

Ory: Right. Actually, there's two sides to this issue. If you think about it, if you do logging properly, you're probably logging to cloud watch, which means that if you did IAM permissions properly, the function will be able to write but it won't be able to read from the logs and the ability of an attacker to destroy logs and cover their traces becomes less probable. Something that again, if you found a remote code execution in a web app, you can probably destroy the logs and nobody will be able to trace the actual attack events back. So in that sense, I think if you follow the book, if you write the logs properly, you do IAM properly, You're much better off.

Jeremy: Right. Now do you think this is something that cloud providers are going to put more emphasis on? I mean, this idea of application security, it seems a little bit outside their purview, but when they're managing so much of the underlying resources, it seems like they would want to do more in this space.

Ory: The simple answer is, no. I think cloud providers are now trying to get us all hooked on their cloud platforms. And so, they will build more features that are related to enabling more use cases. Obviously, they are doing some efforts around security but that's just enough so that people won't be scared of adopting these new technologies. And I think it's correct to leave the security aspects to the security vendors who are experts in this field and have the right personnel and experience, and it contributes to an ecosystem of vendors. You want your cloud to have a nice rich ecosystems of vendors and choices for customers to use different tools. So I think at the end of the day, they will do the minimum required, at least in the next few years and it's correct.

They now need to put an emphasis on how to get everybody on board. We want to see people adopt these cloud native environments. And so more tooling, the more debugging, more tracing, more logging, things like that, and not pay attention to runtime protection and things like that where obviously there are vendors that already have experience with that.

Jeremy: Sure. All right, so I've got one more question for you and this is around the containers versus serverless debate and it's not really about the debate because I think both of those things live in harmony and will continue to be used. And maybe I'm harping too much on the ephemeral aspect of this. But in terms of container security and serverless security. I mean, serverless security essentially is just going to be a container running somewhere that you probably don't know about, right, and maybe it's firecracker or something a little bit different or some of the more lightweight ones like the V8 in the cloud flare workers and some of that stuff. But is there going to be a big difference between how container security and serverless security are handled or eventually do you think there's going to be a merger of the two paradigms?

Ory: Interesting. Okay. So, first of all, regarding containers, I see in general two types of platforms that obviously the ones you manage yourself, like the traditional containers and the ones that are fully managed more like Fargates, where you run them in the cloud native environment. Actually, by the way, I don't consider containers to be cloud native. I have no idea what started this whole ... I know who coupled these. I think serverless is cloud native containers. If I run the container on my own host inside my network, it's not cloud native in any way. But I think serverless, and the Fargates of the world, and the fully managed public cloud container services, those are more similar in the sense that the majority of the backend heavy lifting work is being done by consuming cloud provider services like buckets, and databases, and etc. And there the security is different.

First of all, obviously, we talked about that, but you can't deploy anything other than a serverless security platform or whatever, cloud native security platform. You can't deploy agents and things like that. And you don't control the network. But more importantly, the network disappears in those cases. You no longer deal with network like layer three, four networking, and so firewalls and things like that are less relevant and everything is done through API's. So these cloud services, they offer you API's and then the control over who can access those API's is being dictated by IAM permissions. And so that's why people say, and I love that, that IAM is the new perimeter.

And basically the network layer security is pretty much dead. Yes, there are people who are running VPCs and things like that, but that's not the case. And if you look at the other types of containers, like the traditional where you package things as containers and you run the container, you still control the network and you can put a web application firewall, and you can put a next gen firewall there, and you can do network access controls and things like that, that's a completely different story. So, yes, some runtime behavioral protection logic can probably apply to both but the way you deploy it, the way you do it is completely different.

Jeremy: Awesome. Well, I think that certainly demonstrates why you are a senior distinguished research engineer with Serverless because, I mean, honestly, Ory thank you so much for having this conversation with me. I think, again, not to go back to the FUD thing, but I think this information is super important to get out. Should it scare people? No, I think it should just make people aware that serverless or security doesn't go away with serverless, right. It's not magical like it out of the box, it's way more secure than probably anything that more traditional that you would launch but there are still things you have to pay attention to, and many of these things now fall on the developer, and you don't have the benefits of some of those higher level Ops, Dev Ops people taking care of some of that protection and security for you, or the SecOps people.

So anyways, again, thank you so much for being here. If people want to find out more about you and actually more about Palo Alto Networks, the Prisma brand there, how do they do that?

Ory: Well, you can just type Prisma Security or Palo Alto Prisma in Google and you get to that page. I don't want to do any pitch here.

Jeremy: Feel free.

Ory: So just go into the website and see what we offer. It's like an end to end cloud security platform. I do want to say one more comment, and I think you summarize it well but the last thing we want is to get people scared. We want people to adopt serverless. I definitely think that serverless is what people imagine the cloud to be, which is why I love serverless, I would never think about using today and that would be my first choice for almost anything I build. There are just security things that you have to keep in mind, and there are some nuances that you have to keep in mind, which is why we published that document and why we do those presentations and conferences. But don't be scared. I think that's the one thing that we want people to do is to adopt more and more serverless.

Jeremy: Absolutely. All right. If people want to find you on Twitter.

Ory: At Ory, O-R-Y, Segal S-E-G-A-L, or just look for Mr. Serverless security, I'm just kidding.

Jeremy: That's going to be your new title. Awesome. All right. I will get all of that into the show notes. Thanks again Ory.

Ory: Thank you. Thank you very much for having me. It was awesome.

View Details

About Ory Segal:
Ory Segal is a world-renowned expert in application security, with 20 years of experience in the field. Ory is the CTO and co-founder of PureSec (acquired by Palo Alto Networks), a start-up that enables organizations to build and maintain secure and reliable serverless applications. Prior to PureSec, Ory was Sr. Director of Threat Research at Akamai, were he led a team of top web security & big data researchers. Prior to Akamai, Ory worked at IBM as the Security Products Architect and Product Manager for the market leading application security solution IBM Security AppScan. Ory authored 20 patents in the field of application security, static analysis, dynamic analysis, threat reputation systems, etc. Ory is serving as an officer of the Web Application Security Consortium (WASC), he is a member of the W3C WebAppSec working group, and was an OWASP Israel board member.

  • Twitter: @orysegal
  • Prisma by Palo Alto Networks: https://www.paloaltonetworks.com/prisma
  • The 12 Most Critical Risks for Serverless Applications: https://www.puresec.io/serverless-security-top-12-csa-puresec

Transcript:

Jeremy: Hi everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with Ory Segal. Hi, Ory, thanks for joining me.

Ory: My pleasure.

Jeremy: So you are a senior distinguished research engineer at Palo Alto Networks. So, why don't you tell the listeners a bit about your background and what you're doing at Palo Alto Networks?

Ory: Sure. First of all, congratulations for managing to actually say it, it's a mouthful. So yeah, actually I got this title after a PureSec, the company that I co-founded and was the CTO of got acquired in June of 2018, by Palo Alto Networks. So, as I said, I used to be the CTO and co-founder of PureSec a small vendor, actually the first vendor to offer a serverless security platform. And my current role at Palo Alto is mainly to oversee the research for the security algorithms and the product features for serverless security within the Prisma brand, which is the cloud security brand in Palo Alto.

Jeremy: Awesome. All right, so I want to talk to you about what you've been working on for, I don't know, how many years now it seems like, but serverless application security. And I want to start by discussing what's different about traditional security and why serverless security is a bit different.

Ory: First of all, I think it's important to get some background. I've been doing application security for I guess, over 20 years since the end of the 90s. Starting with Sanctum, which was the first company that that built the world's first web application firewall and later on Apps Scan, which was the first DAST scanner, which was later acquired by IBM. And after doing that for a while, I worked at Akamai for about five years leading the threat research for the cone and cloud security product. And at some point, somebody approached me and started talking to me about severless security. I can already tell you, that was one of the other co-founders. And the story or the technology behind severless sounded very interesting both from an innovative aspect but also from security. Everything I knew about application security seemed, at least from a protections perspective, seemed to be sort of irrelevant or not exactly fit the serverless model.

So obviously, and we'll talk about that later, you still need to do input validation, business logic enforcement and all of those things, but the form factor and the way you deploy serverless applications made it very challenging to the point that it was mind boggling and interested me very much and I started thinking about, okay, how can we apply runtime protection to serverless applications? And that, I guess, got me interested and eventually I left Akamai to join Pure Sec.

So, and back to your question, serverless security, should actually refer to this as serverless application security is indeed application security. The same old application security that we know and some of us love from other places like mobile and web apps. So input validation and configuring the platform and hardening and all of those things, but it has some twists, some very interesting twists that you definitely have to keep in mind when you're building those applications. It's a different way of performing threat modeling and different methods of input validation that you need to think about, where inputs are coming from.

Obviously, configuring the platform is very Different, we're talking about cloud native environments, usually public cloud. And again, we'll get back to that a bit later. So that twist is what I think makes it more interesting and obviously more challenging.That's a high level overview.

Jeremy: So let's get into a little bit more of the details there. So I think one of the things that changes quite dramatically, and I know you've written about this, is that shared responsibility model that the cloud gives us, right. So what changes with that shared responsibility model?

Ory: That's actually one of the topics that I really love talking and just discussing this offline, not always in conferences because this is something that I usually bring up when I talk about serverless security. So in every public cloud scenario, there's a shared responsibility model between the customer and/or the App owner and the cloud provider. And there's a line at some points, and really that line or where the line is drawn really depends on the type of cloud model or public cloud model that you're using. And so we start to think about infrastructure as a serverless, then the cloud provider is responsible for the physical infrastructure but any anything above that is your responsibility. So the VM, the host, the hardening of the operating system, and the users and everything, that's the responsibility of the cloud provider.

And in serverless that line reaches new heights, which is something very interesting, because for the first time you're really not responsible for the majority of security requirements or demands. If you look at PCI compliance requirements and you compare, and I have an article about that as well, between infrastructure as a serverless and functions or serverless, you see that your role is reduced or your responsibility is reduced to even less than half. Which brings me to the next point that's, theoretically speaking, serverless applications actually are a terrific enabler for application security. Takes away a lot of the things that we usually miss or we usually screw. So patching that we all know is a very tedious task that you have to constantly be on top of.

So in serverless your starting point, from a security perspective, is actually much better off. Somebody else is responsible for almost everything except for the application itself, which is, I think, the future of what I was hoping for application security to see all those things patching, and OS updates, and physical infrastructure taken care of by somebody else and leaving you to deal with the things you actually understand about, which is your core business and the business logic that you own.

Jeremy: Right.

Ory: I recently heard a very cool analogy about serverless. Somebody was comparing it to transportation or automobile industry where, when you own servers, it's basically like you own your own car. And then Infrastructure as a service is more like you rent or you lease a car, and then serverless is more like Uber, where you just drive the car when you need it, or you don't drive actually, somebody drives you to where you need to go, and that's your only responsibility basically. And from a security perspective, I think that's brilliant if you think about that. Do application security and leave the rest to somebody else. You just use the infrastructure and then dump it, which is very cool.

Jeremy: Also it's the pet versus cattle analogy as well.

Ory: Yeah.

Jeremy: Yeah. So I think that, that's a super important point, right, this idea of you having to worry less about the security, getting all of this perimeter security out of the box obviously with the cloud providers which gives you a greater security posture right off the bat, which is awesome. But then you can go even further, right, because you can take IAM roles and you can assign those to individual functions, which gives you this really, really fine grained security.

Ory: Yeah. And I think the IAM topic, which used to be mostly relevant for AWS by the way until recently, but other cloud providers are now closing the gaps there, is a very interesting topic because for the first time, as you mentioned, you can get very granular with access controls to the point where, and that's that's almost unheard of in the world of server full or I don't know how you want to call those traditional applications, where you can dictate that a specific function can only do a very specific action on, let's say, a database table. So think about allowing a function to only read. And you could never do that, and you can see that when somebody used to hack into a system and take over an application, usually it would end, it's a game over, like a remote code execution or sequel injection, that's pretty much game over because you can elevate privileges very quickly and do lateral movement inside the network.

In serverless if you do IAM like Identity and Access controls properly, you can get to a point that somebody completely exploits a single function, but is left unable to do pretty much anything other than what that function is allowed to do. So think about a function that, let's say, registered somebody's user, it's like a user creation function, the only thing that attacker will be able to do is create users, which is, okay, it's not good, but it's not the worst case that you can think of. That wasn't the case. If you had a sequel injection or some kind of injection in an App, you would probably dump the entire database in seconds, and that's something that if you do IAM properly, you are reducing the blast radius. And that could be by the way, very frustrating for it from an attacker perspective. I'm pretty sure you remember the article about Lambda Shell and if not, go check it out.

So from an attacker perspective, that could actually be the thing that will block any lateral movement further on. So, again, kudos to whoever thought about that very, very granular IAM model and using that obviously, in cognitive environments.

Jeremy: Yeah, because I think you make an awesome point about that where in the traditional server full application or you've got some application spread across servers that the entire code base basically has access to your entire database, right. So dropping tables or deleting rows, and not that you just can issue a standard delete or drop table or something like that in Dynamo DB environment, but the fact that you can just say, look, you can create as many users as you want to but you can't access the table, you can't select records from that table. And if You really want to get fine grained with sequel as well in serverless, and I think not enough people take advantage of this, but if you have Lambda calling sequel, you can create separate users that have limited access as well. So Lambda A, has access to maybe the write role, and Lambda B has access to the read role, and then lambda C has access to just delete. So there are certainly ways that you can add that granularity even if it goes somewhat beyond just the IAM configurations as well. So, very cool stuff.

Ory: Yeah. If I take you know, I'll give you a real world example of the exact opposite. So my wife runs a blog, it's running on WordPress deployment that she has somewhere in some hosting service provider. And I don't know, she's using some WordPress plugin which is obviously vulnerable and every month or two, we have to completely destroy the entire installation and install everything from scratch because somebody manages to find some vulnerability inside that plugin specifically, which was written very poorly, and take over the entire infrastructure. I'm not talking about the WordPress installation, but also the the OS, they're destroying everything, destroying files, rewriting database tables, and all of that because of a plugin. So, if that plugin had the minimal permissions it actually needs to run, which would probably be nothing, there's no way that would happen. So I think the the new IAM model is a blessing.

Jeremy: Yeah. And we'll talk about insecure third party modules or dependencies in a minute. But before we move on to that though, just quickly maybe, what are some of the security practices that I mean, you've been doing for the last 20 plus years? What are some of those security practices that we're missing the tools for in serverless now?

Ory: Okay. Where do we begin? Let's start with the simplest one, static analysis. I have yet to see an adequate static analysis solution that can handle serverless applications. There's just so much complexity in severless applications. There's a lot of logic that's not inside your code, there's a lot of logic that spreads across functions, there's a lot of glue between functions that, when people use cloud services like queues, message queues, and Kinesis and things like that, we're simply statically scanning the code of the functions and maybe even the configuration is not going to help you. It's going to be impossible to actually locate some data flow related issues like injection attacks and things like that. So static analysis is extremely limited today. These vendors are going to have to do a lot of research and improve the technology to support, I guess, cloud native environments. So that's one.

I'm not saying that they're completely useless. So, if you have vulnerabilities inside a specific function, then obviously if they support the runtime language, that should be okay. But more complex vulnerabilities, they're just not going to find.

Dynamic analysis tools. Again, if it's a web based application and serverless is just the backend, then you might get lucky, and you'll be able to instrument the API's and then fuzz insane attacks. I'm not sure regarding the automated validation. So, you don't always get a response directly. It could be something like asynchronous where you send an API or some data and then it shows up somewhere entirely different, not in the HTTP response. And so being able to validate that the injection succeeded is going to be a problem. And obviously, IAST, which is I think interactive or integrated application security testing, which is you deploy an agent, and then you run test, and then the agent hooks into the different syncs inside the application. I haven't seen anything being offered yet by any vendor. So, security testing is currently a big drawback or automated security testing, I should say, regardless if it's dynamic or static.

With regards to protection, I think I mentioned that in one of the ... I might have not blogged about it because it was after the acquisition already. But if you look at the range of inputs that serverless applications consume, and I'm basing this on I think Chris Month's a slide from a while ago or a tweets about a slide, then web triggers or API gateways, just some percentage, I think less than 20%. I have the number of somewhere I don't want to say the exact number. But not a lot of the triggers going to Lambda functions are actually coming from API gateway, which means that applying WAF, even if it's a cloud WAF, it's not very efficient from a coverage perspective. You have a lot of functions triggering from asynchronous events, from message queues, from S3 buckets. Where would you place a web application firewall there? And even if you did, we're talking about non web traffic, so eventually it is Restful API calls but it's not the classic standard HTTP message parsing that WAFs are used to.

So, WAFs were irrelevant for serverless or for large chunks of serverless applications, which is, by the way, one of the reasons why we came up with the serverless security platform in Pure Sec, given that, that's not giving you a lot of coverage.

What else, so host based solutions and network based solutions like IPS ideas, host based intrusion detection, Endpoint Protection solutions, even outbound traffic inspection like web security gateways that can help you to avoid server side request forgery and remote file inclusion attacks and things like that. Again, not relevant, there's no place for you to deploy them, which is a problem.

Jeremy: All right. Well, I'm glad I asked that question. So I think actually mentioning the things that you mentioned, brings up a really, really good point and I've spoken about this before, and I've even been criticized in some way in the past for this idea of FUD, right, this fear, uncertainty, and doubt especially when it comes to serverless security. And I have always been a practitioner of good security policies or at least I like to think I have been and I tried to be hyper vigilant and maybe I'm a little paranoid, but I ran a web development company that hosted servers and things like that. And so, I have a lot of scars to prove why some of these things are more important or need to be worried about. And so, I want to talk about the CSA top 12 that you spearheaded and did quite a bit of work with. And this is this list of the most critical serverless or the most critical risks for serverless applications. It was inspired by the AWS top 10 a lot of similarities there, although because serverless goes well beyond just web application security, there's a lot more happening behind the scenes so this is a little bit of a broader list. And so, I know there's been criticism again, because I've gotten some of it, that maybe this is instilling a lot of fear saying oh, now here are all these other things you have to worry about and serverless isn't secure and and I don't know how many times I've said this and I know you said it, serverless right out of the box is more secure probably than any other programming paradigm or application paradigm that exists. So maybe just give your thoughts on that about, why this list is important. And I know we've seen a lot of things with miss configurations lately causing lots of problems. But why is this list so important to you or important to everybody?

Ory: Well, you brought the topic of FUD and I think it's worth spending maybe two minutes on that and my views of it. So, I think FUD is not necessarily something negative. It is used in a negative context, especially when people want to call out FUD in a very critical way. However, well, there are two types of FUD. The negative one, which I think is the one that people are usually referring to is the vendor FUD, where you scare people that the end of the world is coming and you have to buy my product or else you're doomed, that's I agree, not the best approach for security marketing. However, my Security vendors, by the way, take that approach. And for the reason why I think FUD actually has some positive aspects to it.

To convince developers and stakeholders that they need to take security seriously, the toolbox that you have in your disposal is not a very rich, I guess. It can say, I don't know, maybe people won't see you as a good developer or you're a lousy Product Manager or things like that, or you can threat them that there's a liability issue here. And if somebody will find a vulnerability and will exploit that, it pretty much means that they will lose their job and the company's going to suffer the consequences.

So, I think that telling or explaining what is the worst case outcome of a security issue is not necessarily a bad thing. So, yes, sometimes fear, I'm not sure about uncertainty and doubt but instilling fear in people that they need to think about security and it's critical and not because ... I can't even find positive things to say of why you need security. You need security because-

Jeremy: You need security.

Ory: Yeah, because people will hack, and exploit, and steal, and ex filtrate your data, and you need to worry about that, it's a risk, you need to fear the risk. So I think it's not all negative. Now, back to the, and by the way, this might be a scoop that a security person is actually admitting that FUD is not necessarily bad and I'm accepting the fact that I'm sometimes spreading fear, I think, especially in presentations when you give talks in conferences, showing sexy attacks and spreading some fear generates more interest than when you speak in a very monotonous way and talk about I didn't know the fact that you need to apply, I don't know, strict IAM permissions. If you don't give the scary examples, people don't go away with anything. So, using drama is always, I think, good in this case.

You have to make sure you don't overdo it, of course, and that actually takes me to the top 10 and later on the top 12 that we published. You have to keep in mind that without the top 12 documents around, developers and architects wouldn't have any materials about serverless security to learn from. And this document talks about potential risks that we prioritize based on what we've seen with customers and prospects. And we collected the data from other evangelist and the industry experts. And prior to this effort specifically, if you take a look back to two and a half years ago, before the original top 10 came out, the majority of materials if you looked for serverless security dealt with IAM permissions and third party library vulnerabilities. And there's a reason why that was the case because either the cloud providers that's what they allowed you to control or the other vendors had legacy security vendors had to offer.

But nobody touched about the actual risks that were we're going to cover in a second, and so I think, at the end of the day, this document provides architects and developers with a good starting points, a good reading material that they can rely on and turn to, to understand what they need to worry about. It's in no way very dramatic, it goes through the list, it offers you the right remediation depending on the cloud provider, it shows examples. I don't think in any point in the document, there's anything scary or FUD like in that negative sense.

Now, there were some people that mentioned that the majority of the document deals with security of functions, I think, that the document exclusively looks at what you can do to secure functions and neglecting other aspects of the cloud native environments in which those functions run. And that's, I think, simply incorrect. If you look at the list itself, it talks about authentication issues, and cloud configurations, and permissions, and monitoring, and how to handle secrets, application secrets, and to prune obsolete resources and things like that. There's a lot more than just the security of the functions. Truth be told, there is an emphasis on functions, but if you think about it in serverless, today's serverless at least, the point or the location in which a developer can actually control input and control business logic is the function. It's where your custom code lives. And so I think it's almost trivial to say that in serverless, at least in serverless that's functions oriented or centric, application security is going to be probably mostly or some of it will be applied inside the functions in. And so I think it's not entirely wrong to pay attention to functions.

In the future where, I don't know, serverless platforms will not necessarily mandate you to write functions like code lists on whatever, you take some pieces Lego pieces and you glue them together. Even today, if you have an application that all only uses API gateway and then goes from there to some S3 buckets or some queue and there's no function logic then obviously, the only type of application security you will be able to do is configuration based.

Jeremy: No. And I totally agree with you. And I think that there are criticisms of some of these things. One of them is well, a lot of this is just plain application security. Well, good, right. That's a good thing. You should know that. I mean, the first one we're going to talk about is basically based around sequel injection or this idea of function injection, and that's one of the things I wrote a post about this where you can upload an S3 key that has sequel in it, right. And so if you aren't practicing good stripping out things or using the right way to parametrize your sequel, if you're using that inside of function that processes that, that has no WAF in front of it or whatever, these are just good security practices.

So I certainly agree with you that application security is really at the root of most of what this does. I also think there's a lot of overlap with what this points out versus what maybe also applies to other types of micro-serverless architectures or event driven architectures certainly. But that's not the point, right. The point of this document is to take these top 12 things that when You're writing a serverless application, whether some of it spills over into configuration, some of it is more traditional application security, some of its just good practices and maybe not specific to serverless, I do really like this document because it is a way you can put in front of a developer and just say, hey, be aware of these things because they can cause problems down the road. But that's my opinion so I certainly agree with you on that.

Ory: There's another comment, I think, and that's important in this phase of the lifecycle of this document is the question of, how much is this practical versus how much of this is theoretical? Because at the end of the day, we haven't seen a lot of attacks and a lot of vulnerabilities in that sense. So...

Jeremy: But how would you see those attacks?

Ory: Exactly. And actually, I'll get back to that point in a second. But you have to remember that since the early, I don't know, the dawn of this internet age, security researchers usually dealt with theoretical issues. If you look at some of the things that I was a part of the effort to discover them things like HTTP response bleeding, and sequel injection, and X path injection, and LDAP injection, cross-site scripting, when we publish those advisories 20 years ago, nobody was exploiting them, it was entirely theoretical. You could say we are to blame that people later on...

Jeremy: You made people aware of it, that was the problem.

Ory: Exactly. But you have to remember that as a security practitioners and specifically as researchers, we are trying to flag potential future risks. If we were to only look at what's being used and exploited today, we will always be in a dog chase with attackers. So, I think it's very good that security experts and security researchers look for the next attacks in a new technology and finding it before it's being exploited. And so you can then teach developers how to avoid these and hopefully, reduce the attack surface.

So I think it's not necessarily bad that we're pointing out things that haven't been exploited yet. And as you mentioned, regarding the evidence, usually attackers there aren't web forums where attackers share war stories of how they hacked into a system. So obviously, attackers don't publish anything about their techniques. And especially if you talk to the application owners and companies, most of them also don't like to share information. In fact, usually when they give the server a security conference talk, at the end, there's that five minutes that you save for questions. Usually, I know that nobody's going to actually ask a serious technical question because they are embarrassed. It's like something that you don't want to talk about around other people from maybe competing organizations.

So there's no resource to go and look at and see how people are exploiting and what are the vulnerabilities. We collected information from customers and prospects, I've reviewed dozens if not hundreds of serverless Apps at this point, and we collected this information to see what are the most repeated risks that people do.

Jeremy: And I think the proactive versus reactive approach is exactly what security people should be doing because again, it's no fun to go and clean up security breaches, it's much easier to stop them right away. So anyways, all right, so let's get into this top 12 list. There is an entire document on this and I will put the link to it in the show notes because it is certainly something people should go and download and take a look at. But just for the benefit of people listening, why don't we go through these and then just give me a quick minute or so on each one. And just to make people aware of them and then, like I said, definitely dive into these in more detail. So the first one is this idea of function event data injection, what's that all about?

Ory: So here's one that's interesting actually when people ask about what's the difference between serverless and I don't know, maybe web apps. In web applications that used to be called, originally, historically, parameter tampering, I think later on it was just called injection attacks. And that's simply when a malicious actor or a user can control some of the data fields that your application relies on, and manipulate those fields to inject some kind of attack payload. So think about sequel injection cross-site scripting, path reversals, command injection, all those injection based attacks, that function event that injection basically encompasses them. The main difference here is, I guess, the rich set of events that you can consume in a serverless function. And that's the main difference.

Again, if you talk to web developer and API developer, they know which fields they need to rigorously inspect. They know about parameters body and query. They know about headers and cookies. Maybe the passing foot path of the URL, everything they're really used to it and there's a lot of frameworks that help you to actually validate that input. But when we are talking about serverless applications, I'm not sure everybody knows which fields they should be inspecting. So think about when you get an S3 bucket event. Yeah, I'm not even sure you know all the fields that actually arrived from that events to the function, obviously it includes the file name, the change, but other fields as well. So which fields do I have to inspect? Obviously the ones I rely to, but do other fields, can an attacker even manipulate them? How is that going to affect my application? I don't necessarily know.

So the problem is the same problem, its input validation, it has been input validation since we wrote Coble applications for mainframe, and on mobile, on web apps and now serverless. But the way or what you inspect, how do you inspect, what are the environments that the value then goes to, requires some different attention than what we are used to?

Jeremy: Yeah. And I don't think this is specific to serverless. Any application you're building now that is getting events, certainly from SMS, or SQS, or any of the other AWS services or other cloud provider services, I don't think this is specific to it, but certainly, yes, if somebody can dump poisonous data into one of these, into one of these hoses, then your system does need to be responsible for parsing that. And I actually think this is somewhere or this is one of the applications where the CNCFs events project that they're working on is standardizing these events and what these payloads might look like, could be really interesting for a vendor to come in or some sort of open source to come in and build a WAF to some degree that could inspect these events.

But it goes back to trust too, right. I mean, that's the S3 example with the sequel in the key, what fields do you trust to right too? So, I mean, you might say, I don't trust user input but the name of the file might be one of those things you could overlook. So, just certainly something to definitely be aware of.

Ory: There are even more basic things that people haven't thought about I think yet. How do you actually trust the event? How do you know that the event actually came from the source you think? I can easily spoof an S3 event, send it to the function and invoke it, and claim that I'm S3. Is there any way for the function to know that the event actually came from the service that claims to be that service? What about schema validation, and you mentioned the CNCFs events. The schema for these events is not even standardized inside a single, if you look at AWS or Google, different types of events have completely different formats, with different fields, some of them with different formats. So schema validation, which used to be ... if you look at like XML security gateways, web service security gateways, schema validation is the bread and butter of how you protect API's, but how do you do schema validation something that who's schema you can only guess based on some examples you see on some documentation on the web?

So other than the input validation there's also, as you mentioned, the issue of trust and well formedness of the event itself, which is interesting.

Jeremy: Yeah. All right, so number two, broken authentication. So again, this is something that can be a problem for anything, but how does this apply to serverless?

Ory: It's the same old broken authentication that we know from any other type of applications, like you said. I think the main difference is, we as serverless practitioners, tried to preach for reduction in I guess, focus. So each function should have very laser focused tasks that it should be doing, at least in our [crosstalk 00:39:54]. Yeah, exactly. Exactly. The principle of single responsibility. You want a function to do one very specific thing. You don't want people to write monolithic functions. And so we are pushing people to break their application into dozens or more functions. And suddenly you have a lot of, I guess, input vectors into the attack or into the application, and you have to think about how you authenticate and authorize invokers.

So it's not a different problem, it's just it increases the scale of the problem. If you didn't do authentication well in a standard web app, and now you're breaking the web app into dozens of functions and each one can be invoked by God knows who, then the authentication issue becomes, I guess, more complex. It's not different, it's not worse, it's just there's more of it. [crosstalk 00:41:01]

Jeremy: It's typically right you're putting some authentication logic in middleware or something that every part of your application has to pass through but now you have the ability where you write a function that has no authentication at all, and your authentication to that function is from a higher service like API gateway, that is controlling that access, which again, it's just different and just something to be aware of. All right, so then what about insecure, this is number three, insecure serverless deployment configurations?

Ory: Yeah, if you drop the serverless and maybe turn it to Cloud and the Cloud Native, I think it becomes more obvious. So configuration and we hear about that a lot, and we see a lot of examples almost every day of people not doing their cloud configuration properly, and it's the same problem. In this case serverless is just because it's in the cloud, in public cloud, your applications use buckets, use databases I think I just read a while ago that Adobe left some cloud database open and people ex filtrated data. So, yeah, I think without configuring ... This is not different than, okay, you could have misconfigured your Apache server and leave directory [crosstalk 00:42:28].

Jeremy: You probably didn't misconfigure your Apache or [crosstalk 00:42:32].

Ory: Exactly. And your your PHP file included some nasty configurations. And yeah, that's not something new but you have to remember that, at least in serverless with the lack of other types of runtime protections, and firewalls, then the cloud configuration or the cloud and the IAM are your new perimeter. So this is your way of securing your account and your application, and so that is probably one of the most important things that you have to make sure that you cover. And there's a good reason why I think cloud security posture management vendors are very successful. So, that's something that everybody understand that there's a need to have. So you need to be aware of all your cloud assets, where they are deployed, how they are deployed, whether they have configuration issues and improper permissions, that's going to be the new command and control for CISOs those cloud security posture management tools.

Jeremy: Yeah, and I think this is one of the points in here that goes well beyond just this idea of serverless being only functions, right. And I mean, it doesn't matter what application you're building that's using these, it's just that when you're building a serverless application and you're using Lambda as the glue, or using API gateway to do some of these serverless integrations, things like that, that the configuration of these managed services is very important. I mean, we've seen people leave Elasticsearch wide open, right, and just be able to query Elasticsearch, just knowing the domain, and obviously the S3 buckets stuff in the Capital One breach some of those things, those configurations certainly there are many issues that can happen when you don't configure these things correctly, broader topic well beyond just the serverless aspect of stuff. But bringing it back to the serverless ... Sorry, go ahead.

Ory: No, I wanted to say and I think you're just about to take it back to serverless and functions. Functions also have configurations that are very important and have security consequences to them. Even the timeout settings, people think about the serverless applications because they're based on functions which auto-scale and support a lot of concurrent executions, people tend to think that these applications are automatically resilient to denial of service attacks but, and we've proven that's not the case, you really have to tweak and configure your functions, and the memory, and the timeouts, and the dead letter queues and whatever properly if you want to actually enjoy these benefits. So it's not automatically secure.

Jeremy: Right. And that actually is a lot of that configuration falls back on the developer, which you're probably not used to doing those things.

Ory: Yeah.

Jeremy: All right. So let's move on to number four. So number four is over privileged function, permissions and roles. This is one of my favorites because I feel like this is something that people do wrong all the time because it's just easy to put a star permission.

Ory: Yeah...

ON THE NEXT EPISODE, I CONTINUE MY CHAT WITH ORY SEGAL...

View Details

Show Notes:

On event-driven architectures...

Mike Deck: (Episode #5) I think that it's probably easiest to understand it when contrasted against kind of a command-driven architecture, which I think is what we're mostly sort of used to. So this idea that I've got some set of APIs that I go out and call and I kind of issue commands there, right? So I maybe have an order service. I'm calling create order or I've got downstream from that. There's some invoicing service now, and so the order service goes out and calls that and says, "Create the invoice, please." So that's kind of the standard command-oriented model that you typically see with API-driven architectures. An event-driven architecture is kind of, instead of creating specific, directed commands, you're simply publishing these events that talk about facts that have happened, you know these are signals that state has changed within the application. So the order service may publish an event that says, "hey, an order was created." And now it's up to the other downstream services to, they can observe that event and then do the piece of the process that they're responsible for at that point. So it's kind of a subtle difference, but it's really powerful once you really start kind of taking this further down the road in terms of the ability to decouple your services from one another, right? So when you've got a lot of services that need to interact with a number of other ones, you end up kind of with a lot of knowledge about all of those downstream services getting consolidated into each one of your other kind of microservices, and that leads to more coupling; it makes it more brittle. There's more friction as you're trying to change those things, so that's a huge kind of benefit that you get from moving to this event-driven kind of architecture. And then in terms of kind of the relationship to serverless, obviously with services like AWS Lambda, you know, that is a fundamentally event-driven service. It's about being able to run code in response to events. So when you move to more of this model of hey, I'm just going to kind of publish information about what happened, then it's super easy to now add on additional kind of custom business logic with Lambda functions that can subscribe to those various different events and kind of provide you with this ability to build serverless applications really easily.

On understanding the connectivity of microservices...

Ran Ribenzaft: (Episode #8) We broke them from being a big monolith, a big single monolith, to multiples of microservices, you can call it microservices, service, nanoservices, but the fact that there was one giant thing that broke into 10 or hundreds of resources, suddenly presents a different problem. A problem where you need to understand what is the interconnectivity between these resources, that you need to keep track of messages that [are] going from one service to another, and once something bad happened, you want to see the root cause analysis. This is like a repetitive thing that you can hear over and over. This root cause analysis, so the ability to jump from the error - the error can be like a performance issue or like exception in the code - all the way to the beginning. The beginning can be the user that clicks on a button on your business website that caused this chain of events. So these are the kinds of things that you want to see where, in traditional APMs, in traditional monitoring solutions, you don't have it. And in the future, once you'll find it more and more like that.

On monitoring interconnectivity...

Emrah Şamdan: (Episode #12) In serverless, on the other hand, it is like you have different piles of logs, which it comes out of box from CloudWatch, from the resource that Cloud vendor propose. But these are actually separate, and these are not actually giving the full picture of what happened in the distributed serverless environment. And what you need here is that the problems are different. In a normal environment, the problem, most of the time, was actually about scalability and you were responding to that by giving more resources, by just increasing the power of your system. But with serverless, the problem is about like some problem occurs in any kind of a system in a distributed network and you need some more than log files. You need like all three pillars of observability, which is called traces. In our case, it is distributed traces, which shows the interaction between Lambda functions and the managed APIs and the managed resources and third-party APIs, and the local traces, which shows what happens in the Lambda function, and the metrics and the logs.

On the purpose of AWS X-Ray...

Nitzan Shapira: (Episode #2) You can do it to some extent. X-Ray will integrate pretty well with the AWS APIs inside the Lambda function, for example, and will tell you what kind of API calls you did. It's mostly for performance measurements, so you can understand how much time the DynamoDB putItem operation took or something of that sort. However, it doesn't try to go into the application layer and the data layer. So if information is passed from one function to another via an SNS message queue and then going into an S3, triggering another function - all this data layer is something that X-Ray doesn't look at because it's meant to measure performance. That's why it would not be able to connect asynchronous events going through multiple functions. Because again, this is not the tool's purpose. The purpose is to, again, measure performance and improve the performance of certain specific Lambda functions that you wanna optimize, for example.

On instrumentation...

Ran Ribenzaft: (Episode #8) Instrumentation is the way or a technique which allows a developer to, let's call it hijack, or add something to every request that he wants to instrument. For example, if I'm making a calls using Axios to a REST API for my own code to an external or third-party API. I want to be able to capture each and every request and response that is coming in and out from that resource, from that Axios request. Why would I like to do that? Because I want to capture vital information that I'll be able to ask questions about later on. For example, if my Axios is calling Stripe to make a purchase or to send an invoice to my customer, I wouldn't know how long it takes, because I don't want my customer to wait on this purchase page or wait for his invoice to get into his email. I want to make sure of how long it takes so I can measure that, put that as a metric in CloudWatch metrics or in any other service. And then I'll be able to ask, "Well, was there any operation against Stripe that took more than 100 milliseconds?" If so, it's bad, and this is only accomplished using instrumentation. I mean, the other way around is just to wrap my own codes every time that I'm calling Stripe or every time that I'm calling any other service. But with the amount of annotations that you'll have to add to your code, it's almost unlimited, so you won't get out with it without a proper instrumentation in your code.

On the problems with manual instrumentation...

Nitzan Shapira: (Episode #2) It's not just the fact that you can forget. It's also just going to take you a certain amount of time - always - that you're going to basically waste instead of writing your own business software. Even if you do remember to do it every time, it's still going to take you some time. Some ways that can work [are] in embedded in your standard libraries that you work with. If you have a library that is commonly used to communicate between services, you want to embed that tracing information or extra information there, so it will always be there. This will kind of automate a lot of the work for you. That's just a matter of what type of tool do you use. If you use X-Ray you're still going have to do some kind of manual work. And it's fine, at first. The problem is when you suddenly grow from 100 functions to 1000 functions — that's where you're going to be probably a little bit annoyed or even lost, because it's going to be just a lot of work and doesn't seem like something that really scales. Anything manual doesn't really scale. This is why you use serverless, because you don't want to scale service manually.

On adding distributed tracing early...

Sheen Brisals: (Episode #20) So we do sort of a structure logging that kind of evolved from a simple log messages. So we have sort of a decent level of logging, you know, if you look at the logs, we are now able to trace things through. Then at one point we started, we have a monitoring system in place, so we kind of stream the logs to ElasticSearch as well as to the monitoring system, so that the structure logging, with ElasticSearch, we are able to, you know, go through and try to identify any issues, and engineers work with that. But one area that we didn't focus or we didn’t put in place was the distributed tracing side of things. So that's why I think I once stated that if you're starting your serverless journey, please you know, start with the distributed tracing and manage. I mean, you can start with XRay, or bring in a third-party tool. So that's a really cool thing that gives lots of confidence to the team.

On the best way to prepare for an incident...

Emrah Şamdan: (Episode #12) So the best way to get prepared for an incident is actually to experience it before. But no one wants to experience something bad over and over again, right? And the nice thing that we can do with chaos engineering is that you can just get yourself prepared by actually simulating that these kind of problems. So you can ask yourself, what if this third-party API that I'm using starts to respond slower? What if the DynamoDB that I'm just leaning on completely starts not to respond. So you can you can run such kind of chaos engineering experiments, and in this case, you should be knowing what will happen. And you should be knowing that not just because of, not from the perspective of what to do, but how to inform the customers, how to inform the upper management, how to have the, let's say, the retro. You can understand how we can respond to these kinds of situations from many different perspectives.

On chaos engineering...

Gunnar Grosch: (Episode #9) Well, the background is that we know that sooner or later, almost all complex systems will fail. So it's not a question about if it's rather a question about when. So we need to build more resilient systems and to do that, we need to have experience in failure. So chaos engineering is about creating experiments that are designed to reveal the weakness in a system. So what we actually do is that we inject failure intentionally in a controlled fashion to be able to gain confidence so that we get confidence in that our systems can deal with these failures. So chaos engineering is not about breaking things. I think that's really important. We do break things, but that's not the goal. The goal is to build a more resilient system.

On the difference between resiliency and reliability...

Gunnar Grosch: (Episode #9) Resiliency isn't only about having systems that don't fail at all. We know that failure happens, so we need to have a way of maintaining an acceptable level of operations or service. So when things fail that the service is good enough for the end users or the customers.

On the purpose of these experiments...

Gunnar Grosch: (Episode #9) So we do the experiments to be able to find out how both the system behaves and also how the organization, their operations teams, for example, how they behave when failures occur.

On the differences when testing serverless applications...

Slobodan Stojanović: (Episode #10) There are a few different things but, in general, testing is still the same. You want to check if your application works and the way that you want it to work. But some of the things are not your responsibility anymore. For example, infrastructure is, like, the responsibility of your vendor, such as AWS or Microsoft or someone else. So there's no point in really testing that part because that's not really something that, they have their own testing, things like that. But you still need to be sure that your business logic is working in a way that it works. And also all serverless applications are basically microservices, that they're working together. Most of the time, you don't have one monolithic application that is just uploaded to AWS Lambda or something like that. Most of the time, you have, like many different functions. For example, in Vacation Tracker we have, I think more than 80 functions now that they're working together. So it's really important to be sure that all those small services are working together the way we want them to work together, and that our end users have a decent experience and that they can use our application.

On how software development has changed...

Chase Douglas: (Episode #2) Yeah. So the way that we've always developed software up until very recently was it would, in the end, be running on servers, whether it's in a data center or in the cloud. But these servers were monolithic, compute resource. That meant that typical architectures might be a LAMP style stack. You've got a Linux server, and you've got a MySQL database off to the side somewhere, maybe on the same machine, maybe on a different machine. But mostly as a developer, you're focused on that one server, and that means that you can run that same application on your laptop. So were we become very comfortable. We built up tooling around the idea of being able to run an entire application on our laptop, on our desktop in the past, that faithfully replicated what happens when that gets shipped into production in a data center or in the cloud. With serverless, everything is kind of a little works differently. You don't have a monolithic architecture with a single server somewhere or a cluster of servers, all running the same application code. You start to break everything down into architectural components. So you have an API proxy layer. You have a compute layer that oftentimes is made up of Lambda, though it can include other things like AWS Fargate, which is a docker-based, serverless, in the sense that you don't manage the underlying servers approach. So you've got some compute resource, if you need to do queuing instead of spinning up your own cluster of Kafka machines, you might take something off the shelf, whether it's SQS from AWS or their own Kafka service or Kinesis streams. There's a whole host of services that are available to be used off the shelf. And so your style of building applications is around how to piece those pieces together rather than figuring out how to put those and merge those all into a single monolithic application.

On the tools for serverless...

Efi Merdler-Kravitz: (Episode #13) There are a lot of tools today that enable you to package, to upload, to deploy your code. You have tools today that help you to monitor, and debug. Use them. Don't write something on your own. Don't waste your time on it. And I think one of the first things that you need to learn is to learn tools like AWS CloudFormation or Terraform. These are the tools that enable you in the end, that these are the basic tools, these are the building blocks that enable any serverless packaging technology to deploy your code, to deploy your various sources. So no matter what serverless framework you choose, either the Serverless framework, or Chalice, in the end, behind the scenes, everyone is using either CloudFormation or Terraform. So I think it's very important to learn the best building blocks, and I think you need to learn how to automate your tools, automate your testing. So use good testing libraries like Pytest or Jest and there are many others that are very good. And also use serverless plugins to test some of your flows locally, like DynamoDB or API Gateway.

On testing locally...

Efi Merdler-Kravitz: (Episode #13) I think it's a painful point right now in serverless, in serverless testing. And I think the only thing that I can say right now is that testing locally just as you said won't give you the quality that you're expecting. In the end, local testing will give you a certain amount of validation on your code. But I think that the best way to increase your testing velocity is to give your developers that build it to run their code easily and fast in the cloud. That's the only way to actually test and make sure that the code that you wrote is working.

On the challenges of testing locally...

Chase Douglas: (Episode #2) You start with some code, and that code for a Lambda function has this handler that gets invoked. One of the things that I did early on when I was starting to play with this to try and speed up this this iteration workflow is "well, I could write a little wrapper script that invokes that handler code function with some test data just to get it running locally without having to deploy it all out." And, there came along some tools that kind of helped facilitate this mechanism. AWS Sam, their tooling has SAM local invoke where it will take your function code, and it will actually spin it into a docker container and run it as though it's in a proper lambda environment. Ah, the Serverless Framework has a similar thing. But even there you have a challenge where the permissions that your function has is based on the permissions that you have locally on your laptop. Now, a lot of developers, they have permissions on their laptop, but they have sort of administrator permissions. They can, if they wanted to interact with any resources inside of their AWS account. Whereas the function that you're building it's tied to a very specific set of permissions where you don't normally give it full administrator access. So you have to sort of a lot of times you get your code working. And then as a second step, you have to figure out is the code still working when I deploy it to the cloud and I've got the permission set the right way. And then lastly, you've got the challenge of that service discovery piece where if I'm running on my laptop, how does my function know which DynamoDB table it should be interacting with, which SQS queue it should be sending messages to. So you've got to solve these problems through some mechanism, and a lot of people come up with their own little test scripts on the side that help here and there. But there's a real challenge there around having a workflow that a whole team within an organization can uniformly use and provides them with that sense that they're bringing the cloud to their laptop locally.

On serverless security...

Hillel Solow: (Episode #11) One of the things that you see in the move to serverless - and again I don’t think it's a serverless-only thing, but I think it's in serverless more than anywhere else thing - is that sort of divide between security people and developers. It's not really tenable and, you know, in serverless, a lot of the security controls that security used to own are now security controls that developers control. Like configuring IAM roles and setting up VPCs and things like that. And so in a lot of ways, we've actually put more responsibility on developers, but we haven't necessarily empowered them in real ways to make security decisions, and at the same time, we haven't given security people a way to meaningfully understand and audit some of those things when they don't necessarily understand what the application does or what the code wants to do. So I think that's been a big change, and I think that's true across a lot of cloud applications. But it's just truer in serverless applications. You're kind of forced to reconcile that. The other thing about serverless applications specifically that we like to talk about is the fact that developers have gone from an application that comprises 10 containers to one that comprises 150 functions, you know, could create all sorts of nightmares in testing and monitoring and, you know, deployments and things like that. But for security it’s an interesting win there where you get to apply security policy, IAM roles, runtime protection, at a very fine-grain level — you know, at kind of a zero trust, small perimeter level. And that's if you can do it right, if you can do it at scale and automatically, that could potentially be a huge win, really for, you know, mainly least privilege, reducing attack surface and reducing blast radius. You know, something goes wrong; my developer left a back door accidentally into a function. But now that function really can only do right to one particular table, as supposed to, you know, in the old world, where that gave an attacker a lot more capability. So I think that's an opportunity that is on the table. It is challenging to capitalize on that. Like you said, there's less time. There's less gates between developers and running production code. And that means that, you know, how do we automate and capitalize on a lot of that value without trying to slow everybody down? That's the big challenge.

On building fat functions...

Brian Leroux: (Episode #17) I think it's totally appropriate to build out your first versions with just a few fat functions. But as time goes on, you're gonna want that single responsibility principle and the isolation that it brings. There's one last small interesting advantage to this technique is that the security posture is just better. You have less blast radius. If your functions are locked down to their least privilege and their single responsibility, you're just going to have a way better risk profile for security.

On developers understanding cloud costs...

Erik Peterson: (Episode #6) If you're a SaaS vendor, you know your value delivery chains is built on top of cloud. That's your cost of goods. That's your gross margin. You need to understand that if you're going to deliver a profitable product to the market and and you want that conversation to be part of your entire organization because, I mean, the reality is is that the buying decision is being made by your engineering team now, right? They choose: am I going to use this type of instance or that type of instance? Am I going to implement this kind of code or that kind of code? They make a buying decision every moment of every day. Essentially, every time, every line of code that they write, they're making a buying decision, and and so you have to think about that. And then it gets even more complicated, though, because there are so many intertwined, and particularly in the serverless world, which is so, I think, honestly I'm sure our listeners here will appreciate our point of view, is that we think, we believe serverless is the future of all computing. But, you know, it's even more powerful because you create these very interesting applications that are composed of lots of different services. It's not just Lambda compute. It's I have Lambda connected to SNS passing to SQS, DynamoDB, Kinesis ⁠— all these things flowing together. And I'm not just going to the cheat sheet on Amazon and saying, "well, how much does it cost for one hour of compute?" to try to estimate my costs. No, I now have to think through that whole story, and I think it's kind of a shame that, actually, for most organizations, they consider the state of the art there to be well, let's just try it and see what happens. And a lot of times they try it and test and they go, "Oh, looks like it was gonna cost a couple bucks. Great. Let's ship it." And once it gets into production, it's a much different story, and they just don't ⁠- organizations really struggle with this. It's unfortunate.

On making sure your developers are aware of costs...

Efi Merdler-Kravitz: (Episode #13) I think that people, you know, people that come to serverless for the first time sometimes forget how easy it is to scale serverless. So in a matter of minutes, you can easily get a hundreds and thousands of Lambdas running simultaneously. Millions of requests to the DynamoDB. And in the end of the day, you suddenly see a bill of a couple of hundreds of dollars, and you ask yourself, “What?” I as a manager, I check the cost on a daily basis, and I'm trying to understand the trends. I use the Cost Explorer in AWS quite a lot. In addition, we also use our own tools. We have our own monitoring tools which also gave us a cost breakdown, and I think again, part of the code review is part of the checklist that I mentioned earlier. We ask the developers, why did they choose, for example, this amount of memory for this specific Lambda. Or why did they add another index to DynamoDB? Each index costs more money because you are duplicating the data. And for example, while they are using Kinesis and not Firehose. So there are many questions we ask along the way when doing the code review. Again, it's not something that can be done automatically, something that people need to see the code and understand what's going on. But you ask the questions in order to make sure that developers understand the trade-off, in order to understand that it costs a lot of money. And you know, especially for startups, where money's always tight, suddenly paying thousands of dollars per month, it's dangerous, can be really dangerous. So it's not only “Oh no, we'll use the corporate credit card.” It can be really dangerous for the startup. So you need to pay attention to it.

On premature optimization...

Alex DeBrie: (Episode #1) In terms of purity versus practicality there, you need to think about your use cases and what matters to you. If you're not gonna have a user-facing application, I wouldn't worry that much about optimizing it, or if your bill's not that high right now, don't worry about optimizing it. Most importantly - I think this is true of serverless or non-serverless, but I think it's been a focus in the serverless community - focus on building a product that brings value. If speed is something that brings value to your customers because they want a quick, responsive app, then maybe focus on speed. Otherwise, focus on building those features in that core experience that your users are really going to care about. Focus on that first rather than some of the optimization techniques.

On best practices...

Michael Hart: (Episode #19) Don't try and prematurely optimize. Because I'll tell you, we at Bustle, we have very few, very large Lambdas, and we do billions of invocations a month. You know, we do many, many, many, many page views and our latencies are very low, and it's not perhaps as bad as you think. There are some tricks you might need to do, like we webpack all of our JavaScript into one single file, right so that there's no file system calls being made whenever it's required, and we minified — so there might be people that go, “Oh, that's kind of a gross hack, but well, alright. For us, that's fine.” You know, we've got plenty of developers that know how to do that and that are comfortable doing that and would be less comfortable managing 50 or 100 tiny functions, maybe, and dealing with the ops of that because it's not free. You know, a function isn't a zero cost piece of infrastructure. You still need to monitor them, you still need to maybe tune them. There's a whole bunch of things that every function you have, you need to think about a little bit and monitor and that sort of thing. So yes, so even things like that. I think there's a spectrum for best practices, and I would say, try things out first. Maybe be aware that it's a lever that you can pull, but try them at first and don't stress too much about having the perfect — there's no single way to do these things basically.

View Details

About Sheen Brisals

Sheen is an experienced software engineer, solutions architect and a team coach, currently at LEGO architecting serverless solutions. Previously as a principal engineer, tech lead and development manager with leading organizations such as Oracle Corporation, Hewlett-Packard, Omron, TATA, BAe and others. Sheen is a regular speaker at tech gatherings. He is a keen participant at Serverless meetups, AWS conferences, Serverless Days and others.

  • Twitter: @sheenbrisals
  • Blog: https://medium.com/@sbrisals

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with Sheen Brisals. Hi Sheen. Thanks for joining me.

Sheen: Hey, Jeremy. Pleasure to be here. Thanks for having me.

Jeremy: So you are a Senior Application Engineer at the Lego Group. So why don't you tell the listeners a bit about yourself and what you do at Lego.

Sheen: Right. So, yes, I am a Senior Engineer at Lego. I joined Lego three years ago. And as part of my role with Lego, I I act as a team lead; I act as an architect and also I coach fellow engineers on their career progression. So in terms of my career, I started way back in the nineties as a software engineer. I’ve been through quite a few organizations - both big and small - been involved with number off software development projects all the way through. So I joined Lego at a juncture where when they were thinking of moving to microservices so that's why I came on board with Lego, around three, three and a half years ago.

Jeremy: Awesome. Alright, so you are like jet-setting around the world telling the story of how Lego.com went serverless and I've seen you speak at a number of conferences and I think it's a huge service that you are doing for the serverless community sharing this because it is important, I think, for teams to see how other companies are doing it and how other people are implementing these things because it's very new. Serverless is very new, and even though it's been around for five years at this point, there are still a lot — there’s a long way to go. So I want to talk to you about this idea of the serverless journey at Lego.com. So let's start sort of where are you now? Where is Lego now with serverless?

Sheen: So within Lego, there are different teams or departments embracing serverless. So I am with the team that focuses on shopper technology, shopper engagement technology, that includes Lego’s ecommerce platform and all the toolings around that. So we are progressing well with the serverless approach, so as you know that we migrated the ecommerce platform onto serverless. So we didn't stop here, but our journey continues beyond that because now there are new issues, new developments coming up because we see the benefit of serverless that gives in terms of the speed or the, you know, velocity with which we can bring out new features to customers.

Jeremy: Great. Alright, so let's go back to the beginning, because this is one of those interesting things where I think companies get to this point where they're either starting their cloud adoption or they're running on legacy hardware, and they say, “Okay, we need to now make this move.” And a lot of people go down that container route, but it was a little bit different, so let's start with where you guys were. Where were you a couple of years ago when you came in there? What was the technology?

Sheen: So that time, all on-prem. And so we had Oracle ATG, an old version of Oracle ATG ecommerce platform hosted on-prem and we had an Oracle database talking to, the platform itself talking to a bunch of other services within Lego. And at some point, there was an initiative that happened to make it more API-based and even that time they first APIs around. But still, it was on-prem and a monolith, but the front-end moved on from being a JSP onto a JavaScript based, hosted on Elastic Beanstalk from AWS. That's pretty much what we had two years ago out in that time frame. So we had, you know, they did a number of issues associated with the typical monolith platform maintaining or releasing our new features and fixes and things like that. So that was pretty much the landscape we had at that time.

Jeremy: And then you had sort of a “come to Jesus” moment on Black Friday, right?

Sheen: Yes. So yes. So there are different things. So though I focus on that particular incident, there were one or two other thoughts that were going on. So, first one is the ecommerce platform itself was aging and very old. So we had to move on and then Lego, as a company, wanted to reach out to many children around the world. So that means they need the platform to go out and launch the shop and the availability of the bricks and everything to children around the world. So that was — from the business side of things — that was a need. So we need a platform that can provide us that capability. At the same time, we wanted to migrate from the old platform, so we were not thinking of serverless. So a typical microservice, put everything in a Docker, or, you know, containerize an instance-base — those were the different ideas floating around. Then came this Black Friday. And we had this catastrophic failure of the platform so that triggered some of the conversations internally to get to a point that we must break up things. We can't have everything together as a one-piece. So we wanted to make sure that, you know, we don't fail just like that. If something fails there are stuff the platform should be able to carry on. That's where things got ignited I would say and started the business and the engineering teams started to discuss and come up with proposals and ideas. So serverless was very much new at that time, and no one had any experience or exposure within the team I belonged. So I went out to office while looking at AWS and I'm talking to people, attending different conferences and meet-ups and things like that to gather the idea. Still, not to say where to take the leap to. So that was, at that time, when we had that Black Friday failure. So from then on, based on the technology improvements and the Cloud initiate deals and the organization-wide that the need for our digital transformation initiate too. So all those things came together for us to, you know, take on this serverless journey.

Jeremy: Did you already decide to move to the cloud? Like, basically, you said on-prem is not working for us or we need to be able to scale more. So you decided to move to the cloud and you said you looked at containers. But what was the thing about serverless though, that you said no, we definitely have to go this route?

Sheen: Okay, so there was another factor. So as part of the platform migration or the upgrades, there were a bunch of things that were looked at so either have a platform similar to what we had and everything available within a part of the platform or go the other way around. Look for simple, headless, API-based platform and then we can put our logic around it so we have the freedom to innovate, scale, bring in new features and the all other capabilities from our side. So that was kind of the discussion point. Then he chose that, okay, we need the flexibility because we don't want to be constrained by some of the commerce platforms out there. So we wanted the freedom to innovate, freedom to bring on the features the way that we wanted to bring out to the customers. That was a main shift. That's when we started looking at cloud as a, well, you know, enabler for us. So and, then, with all that, you know, the availability, HA and the scalability and all the things and came along with, you know, the Cloud thinking. So that's sort of the position point where we started to focus on AWS Cloud and then came the serverless mindset.

Jeremy: Nice. Alright, so now that you've made this decision. You’ve said, “Okay, we're moving to the cloud. We're going to go ahead and take the serverless-first approach.” So what was the process for getting started? Because obviously, moving to the cloud is a big step for a lot of companies. There's security. There's understanding, you know, just the ecosystem and what's available. So what did that process look like for Lego?

Sheen: Okay, so at that point, so you know, the Black Friday failure happened and that gave us the opportunity to try out something simple when we wanted to decouple small part of our system. And that's where we introduced serverless, because not knowing anything about serverless, we didn't want to take a huge risk. So we said, “Okay, let's try this out. If this works, we can take it further. If not, we can change course and something else.” And for us, it worked, and from there what happened was that we started to realize the potential that serverless can provide us. So slowly the different managed services became familiar to us. So the S3, DynamoDB and SQS and all the different managed services that become part of the serverless ecosystem kind of became familiar. So then we were able to see the opportunities that that would bring us going forward. So that's kind of the, you know, the time we started. So based on that, then we put together the sort of the guard rails, if I may say, so that when the team starts focusing on serverless, we won’t kind of debate and deviate from where we wanted to go. So that's when we said “Okay, let's go serverless and let’s use the managed services where available so that we don't need to bring anything else and reinvent the wheel.” And then, “Let’s use AWS Cloud, because lots of people use AWS Cloud and all the, you know, the services and features that we looked at, they’re all there and all that. You know, so there was a great community around it, especially and then there are plenty of meet-ups, and the talks, and the user groups are awesome. So that’s sort of the initial principles that we put together to move us forward to the next stage.

Jeremy: So before we move on I want to go back to what you said earlier about sort of taking a small chunk of the application or biting off a small piece and testing that. I think that is an awesome way to get started with serverless, especially for large organizations, especially when there is potential — disagreement is the wrong word — but maybe there's not quite the amount of confidence from upper management. So is that something that the engineering group or the management of the engineering group saw or you were able to do these small experiments and have successes and that was sort of what built that confidence for you to go entirely serverless or to really push the serverless mindset there?

Sheen: Yes, that's a very, very good and important point because when we often look at a monolith, we often get confused. Where do we make a start? Because it looks everything big. But the thing is, we need to start looking more closely part-by-part. So then we will be able to identify some small entry point into the system that will give us the comfort to, you know, try out something new. And also, when you have, when you work in the organizations, you need to prove or showcase these things to stay stakeholders, you know, to get their buy-in. So for that, it's important that we identify a part of the system that is not complicated, small enough that we can experiment with the new ideas and show them, show them the proof that it's working and that’s feasible for us to go forward to everyone around — not just the engineering team. Bring the business stakeholders, everyone together and yeah, so that's very crucial when, especially when we start this monolith to microservices serverless journey.

Jeremy: Yeah, I love that point because again, it's such a good way to do it, you know, and whether it's just peeling off the email sending component of your system or something like that. It is such a useful way, and then it's easy to have those early successes and prove out, especially if you're trying to get adoption. Alright, so you talked about some principles or some guardrails that you sort of formed around serverless. And so what were those? But were those sort of about this idea of using services or managed service when they were there? Or was it also sort of just coding styles, and you know, fat Lambdas versus, you know, single purpose Lambdas — things like that? What were those guardrails you put in place?

Sheen: So the initial guardrails were all kind of high level. So managed services I talked about, and the tooling, language preference, testing, those kind of things because those, I mean, if you think of the lean or the fat Lambdas, well that was too early to think of all those things. Those things come as part of experience. So initially when we start, we can set up these sort of toolings and all the different frameworks and etc., because often you know, when you really have a bunch of engineers around, there are always different preferences and choices. So it's important that you respect your whatever is happening outside. What do we have, say, for example, in our case, we could have argued between Golang and JavaScript forever. But when we looked at the skillset we already had, JavaScript was the obvious choice. Then they started looking at okay, so where is the, you know, tradeoff between these two or their different language or tools. So then you realize, okay, it isn't that bad. So you're not making a — you know, [you’re not] completely out of choice or anything. You're still within the preferred toolings. Similarly, you know, the framework we choose, we want it to get going faster. So what gives us that capability to move forward faster? So these were the initial sort of guardrails and the principles that we set in place. Yeah.

Jeremy: So once you had that, then you started sort of forming some teams around these different components you needed to build, right?

Sheen: Yeah. So we went away from this typical full-stack setup at that point, simply because we didn't have enough AWS or serverless knowledge that we can distribute across different squads. So that was an important position being made which paid off really well for us because we pulled the AWS skills and serverless skills in one team. We started out with two or three engineers with some of the knowledge, then slowly grew the team or the squad. So that was one of the best things that we did at that stage. Because then that particular team was able to come up with the required serverless services implementation-wise, and then the other teams are they’re looking at the GraphQL layer or the front-end. So they were able to consume these services, and then, you know, what among them to take this forward to the next level. So that was one of the, you know, the best things that we did at that stage and then when we kind of migrated the shop, he had, like 10 or 12 engineers focusing purely on AWS around technologies.

Jeremy: So how did you make that shift sort of from a DevOps standpoint, right? So you had a lot of stuff on-prem. You're migrating things into the cloud, obviously with serverless where, you know, the engineers are probably putting the IAM permissions, you know, as part of the deployment there, you want to automate some of that stuff. And I know you did quite a bit of automation right from the beginning, which was, I think is the advice that I would give to everybody, like just take a few, take a few steps or a few ticks and and and figure out how to do the automation piece because you don't want to configure anything manually when you're launching something to production. But so sort of what was that process for Lego, you know, when your team was thinking about automation and DevOps? How did you make that sort of responsibility shift?

Sheen: Okay, so that's partially from the previous experience. So we had in the traditional way we have it, we have a different team looking at all the infra and doing all the infra instead of coding and everything and engineering team had nothing to do with that. But with this one, when we started this journey, we wanted to have that infra team as part of the engineering team. So that happened. But still we were in a dilemma how to split the responsibility, where the engineering team is going to be responsible for this thing or infra team. Now initially we started out with the infra team, then soon that became a bottleneck for us because the engineering team was able to come up with the service implementation faster, whereas it relied on an engineer to set up all the, you know, scripting and everything else for the department. So we took some time off and we discussed and said, “Okay, let's kind of merge these responsibilities.” So an infra engineer is a specialist, they will still have whole certain areas that they will work on whereas the day-to-day, the service implementation and delivery kind of things, the engineering team will start incorporating the scripting so that we don't have this sort of, the blockage progressing further in terms of the delivery of the services. So these things won't happen overnight. So if you don't have this previous serverless experience so this will be part of the new learning. So that's fine. So that's how the teams learn and move forward. As long as we are flexible enough to understand and take the approach appropriately, that should be fine. There is no reason why we should have, you know, we can't do that. We have to have this team doing that. And the engineering — no. If we have that mindset, then won’t, you know, progress fast enough.

Jeremy: I definitely agree. I think that when you're building out serverless teams, if you have a few people that are working on sort of that cross-functional team where you've got people who understand security, people who understand the development, and depending on how big the team gets, you know, if you have people who are dedicated to DynamoDB or Kinesis or any of these other things or can focus some of their skills on that, I think that is a huge advantage when you when you kind of put that together. Alright, so the other cool thing you mentioned and maybe you could explain this because I think it's probably a relatively sort of common-type practice, but you call it “solution detailing.” Can you talk about that a little bit?

Sheen: Yeah, sure. Yeah. So as I mentioned when we started, we didn’t have enough AWS or serverless knowledge. So I was at AWS at that time. I had maybe one or two engineers coming on with AWS skills. So in order for us to gain the momentum in terms of engineering the services faster. Someone had to provide the details to the engineers, because imagine that engineer coming on board with the JavaScript knowledge with no or limited AWS skills. So that means we need some way of enabling them to use their, you know, the programming skills, coupled with the serverless, AWS skills. So that's where the solution detailing helps. So people like me would take a small part of a system that we're going to architect. Then we'll put the architectural picture or diagrams in there, and will go on detailing each service level. Say, for example, I would put okay, so this isn't an S3 bucket or put in the diagram. This package is going to be, you know, doing these things. We’ll have these sort of human triggers for these Lambdas and it will have this life cycle policies, so we don't need to keep it beyond certain duration or days. Those sort of things will go as part of the solution detail. So what that what does is, for the engineers, they can go through that, especially when they are new, they know exactly what needs to be done then that will raise discussions or questions. So obviously that would be a collaborative effort between me or someone else doing the solution detailing with the engineer or engineering team. So that way, the knowledge slowly gets transferred to them as well. So they get familiar with these terms and technologies, especially, you talked about the IAM permissions, that's an important point, because when someone starts new on AWS, they wouldn't have a clue of what we're talking about, the IAM permissions, it's important that the senior engineers or the architects explain [to] them: this is the reason, and show them these other ways you need to do and why it’s a best practice and all other things. So for those reasons, this solution detailing helped a lot. One additional thing I do is that when I finish the solution detailing, it is like a Confluence space that I give to the engineering team or engineer who will be working on and say that, “Now you own it. If you make any changes, you don't need to talk to me unless you get change the entire architecture,” for example. That probably won't happen. So small changes you carry on. That's your document you keep up-to-date, and then you know, go along with it. So that then becomes kind of a reference for teams when they do the testing and, you know, the quality QA teams, they keep up that. You know, exactly, you know what’s expected and what is there so.

Jeremy: Yeah, and I really like that approach. I know for me, you know, I've been living and breathing serverless for I don't know what, four years now, or something like that. And every time, I mean because things move quickly, new services become available. There's new ways to do things; there’s new leading practices that emerge and you'll be there and you’ll be working on something. And maybe you have it in your head. You know what you want to do. You're doing some TTLs in Dynamo or you're going to make this update. Whatever it is, I know I do that — I'm in the middle of writing something and then I question whether or not this is the right way to do it. And is there a better way? And so I think putting together sort of a proof of concept or, you know, the idea of having this solution detail or whatever and being able to share that too, and kind of iterate on it a little bit before you kind of commit because that is one thing that I really like about serverless. I find myself spending 80-90% of my time thinking about how I want to solve the problem and about 10-20% of my time actually writing the code that does the solution. So it is really good to think those things through first and, as you said, for new engineers, if you think that stuff through, you know, writing code is writing code and if you've outlined that underlying architecture, I think that's a great approach so that will certainly speed up the learning process for people new to serverless.

Sheen: And also just one more thing to add. Also, it kind of, sometimes to some ideas, new ideas say, for example, recently a few weeks ago, I did a solution detailing for a particular feature, and in there, I opted for the on-demand Dynamo backup. So I said, “We need to have a scheduler, we need to back up, you know, this many times a day,” because that was the need. There's no need for the automated backup for that particular data. Then the engineer started implementing and one of his colleagues, he kind of got into the discussion, said, “Why can't we use an existing thing? Where can we use automated backup?” So I explained that, you know, there is no need for the automated backup. You don’t need to, you know, keep 30 or 35 days backup or anything? This is a simple service and the data isn’t that important. Then the sort of discussion continued. Why don't we have anything available as part of AWS or something? So that triggered and I realized that, of course, there is AWS backup that supports DynamoDB. So we don't need this sort of, you know, our own scheduler doing this —AWS backup does it. So that's sort of, you know, ideas just born out of, started off from the solution detailing, then the discussion followed up, then led to a better approach.

Jeremy: Yeah, definitely. That's a great idea. And I think the distinction, too, to make — this isn't like waterfall development, right? We're not designing out the entire application and specifying exactly how everything works. We're just basically pinning down or you're outlining a good architecture, but the implementation details would still be up to your engineers.

Sheen: Yeah.

Jeremy: It's great. Okay, So if I remember correctly from your talk or one of your talks, you had a bit of an awkward start with CICD, right?

Sheen: Yes. That is, I wouldn’t say it’s perfect. It's still a work-in-progress because, you know, the time when we started, we were in a hurry to, you know, finish something by a certain deadline. You know, though we were doing agile, we had set up, you know, a target to beat. So this is one of the areas where we had to do some sacrifice, because we didn't have enough engineers or info engineers to carry on all these things. So this is still bit off manual process and work for us. We don't have the end-to-end automated deployment pipeline. So when a PR merged, it goes through the pipeline and gets to the QA environment. So from there on to the acceptance and production is still manual. A bunch of what's happening as I speak to get it more fluent and take it, you know, all the way through. But it's not going to happen anytime soon because there are a number of different, you know, stages we need to get through and even improving the CI pipeline itself. So it is still a work-in-progress, but we are in a much, much better position than when we started last year.

Jeremy: Well, yeah, I mean, the other thing is, is that CICD and serverless just hard. It's not as straightforward as, you know, just putting something up there and then and then having your servers, you know, deploying it to the individual servers. I mean, there's so much, so much more with the infrastructure deployment and the infrastructure as code and things like that that have to happen and then, of course, you know, if you're deploying into different accounts and things like that, I had a whole episode with Forrest Brazeal about this and yeah, it's complex. So don't feel bad about it, because it's certainly not the easiest problem to solve.

Sheen: yeah, exactly. That's why I'm, you know, in my talks, and when I talk to other people in other conferences, I'm open and say that you know, it's not perfect progress. We’re still improving. Yeah, that's, you know, that's the best thing. So you should be able to continuously improve as you move forward.

Jeremy: Awesome, alright? So a lot of best practices or sort of, I guess, lessons learned came out of this journey for you and the team at Lego.com. So you want to talk about a couple of those best practices and sort of lessons learned?

Sheen: Yes, sure. So there are plenty of lessons obviously. One area is — I think I blogged about this — is about lean versus, you know, fat Lambda, how we make the choice, you know, And also, the other one is like news, the service integrations, where available and where possible so that you don't need to always write Lambda functions for everything. And so then obviously there are a number of other areas, well, say, even when you have storage, there are ways to make a decision on where you want to keep the data, or you want to remove the data. To remove the data, there are a number of ways you can do that and you don't need to, you know, clog the data there and pay for your storage. I mean, DynamoDB has, TTL has policies, and CloudWatch has it, so it's very, you know, duration. So a number of these things, I mean, these are all the lessons learned because early on, we didn’t have certain things set because no one knew these things were there, but then started to learn about, don't know about these things, so then they become best practices for us. So obviously you learn something, and then you put that into practice as one of the things So if I talk about the lean versus the fat Lambda, that again came out of our experience, especially dealing with the Lambda functions behind the checkout flow where everything needs to be fast and crisp and quick because customers otherwise will laugh at the customer experience. So we tried a couple of options. We tried, we thought of using step functions and we thought of splitting into different Lambda functions, but those approaches didn't give us the fast response that we were looking for. So that's one of the reasons, right? That, okay, we need to be open, so there is no right or wrong way. You choose the approach required for that particular situation and someone I remember, someone commented about their approach because they have functions they complete in a few milliseconds. So if they had to split in the number of different Lambdas, they will obviously pay for, you know, 100 milliseconds for every Lambda that as a single Lambda for them is less than 100 milliseconds. So there is a cost implication as well. So a number of these different things, as you become more familiar with serverless, as you gain more experience, you start to learn and then put into practice.

Jeremy: So what about what about security? What lessons came out of that

Sheen: Security is a big area. There are different ways of looking at security. So, you know, alright, we talked about the IAM permissions. So initially we had a serverless written with sort of, you know, a lifeguard? Yes. Then as part of the code review, RR review, these things will be got out and will be put, you know, corrected and put in right. Then we have the APIs. So we have a bunch of APIs, API Gateway endpoints and how do we secure them. So there was an issue, because we didn't have the time or the expertise to put together the client side authentication mechanism in place, so for a while, we had to go with the API key-based approach, even though that's not a recommended thing. But then we have sort of a locking down mechanism implemented in Lego so that you know, the access gets white-listed, otherwise. So that area, that is still improving. So especially the new services that we are now working with. So we have planned authentication put in place from the beginning. So we work with an incognito user pool and the scope-based authentication mechanism. So that is coming in slowly for all the new services and also will then get applied to the existing APIs as well.

Jeremy: Alright. And how did you deal with logging and tracing and monitoring?

Sheen: Yeah. So we do sort of a structure logging that kind of evolved from a simple log messages. So we have sort of a decent level of logging, you know, if you look at the logs, we are now able to trace things through. Then at one point we started, we have a monitoring system in place, so we kind of stream the logs to ElasticSearch as well as to the monitoring system, so that the structure logging, with ElasticSearch, we are able to, you know, go through and try to identify any issues, and engineers work with that. But one area that we didn't focus or we didn’t put in place was the distributed tracing side of things. So that's why I think I once stated that if you're starting your serverless journey, please you know, start with the distributed tracing and manage. I mean, you can start with XRay, or bring in a third-party tool. So that's a really cool thing that gives lots of confidence to the team. So that's kind of what we have in terms of logging. But again, this is an area that is always improving, including the tool that is the monitoring tool that we use. They also going to come up with the enhancements that will provide us some sort of capabilities. So that's also coming. And also we had, we spoke to a bunch of the distributed tooling providers as well, so they sort of are still open. So that’s an area in which we’re constantly improving

Jeremy: It sounds like one of the one of the biggest lessons you learned is you can always keep improving right. You don't have to, doesn't have to be perfect the first time around, as long as you get some of the, you know, most of the rough edges are at least from a security standpoint, and some of those things. You know, there's a lot of improvements you can make, and it's really fast to iterate, too.

Sheen: Exactly. The point is, I mean, we say serverless and Lambda functions, but as they start growing, you are then soon, a bunch of the things immediately, you know, are swamped, because you don't have the time or the liberty to sit down and make everything perfect. So you identify the most important things that will take us to deliver your most valued features. But then you slowly, you know, start improving on the areas to bring everything.

Jeremy: Absolutely. And I think that probably ties into a little bit, ties into your concept of Set Piece architecture, right? Can you explain that a little bit?

Sheen: Yes. So that term I don't know if you follow the football, it's not your football, but the English football…

Jeremy: Soccer!

Sheen: So that is this sort off play, they say often set piece play. What they mean is that they sort of, they don't play end to end to score a goal. So they send someone coming in front of the goal post on the corner kick and things like that. So then they score a goal so that they can offer for us a set piece play. And I thought. OK, so even an architecture or even building and Lego model. The concept is there. So we don't build something from one piece to the end, we kind of build things, group them into smaller modules and smaller pieces together. So I found that way of looking at the end there architectural and focusing on areas that we work with, as you know, good progress way off dealing with the complexity. So that means between the cloud and serverless, we get this sort of opportunity to practice that thing. So that means so we can come together with the quick solution implementation. Then we can kind of run it through. You don't need to wait to get to production to do that. Because all environments are the same, you have your DEV, or TEST, or PROD. They are just AWS environments, so you can basically try it a number of times in a way off in sports, say, for example, in rehearsal or you know, that sort of thing. So you you practice that again and again and make sure that that particular solution is now performing and solid production ready. All you need is just take it and deploy it to production. So that means, so I always say that we need to have the vision off the entire architecture, but we need to focus closely on a particular piece at the time. So this is how we kind of, you know, build the different areas and, you know, brought them together. Because obviously the services these days communicate with events or messages, and if API calls, so that means we don't need to tightly integrate many things together. So we have that loosely coupled or a decoupled upwards, so that we can focus and make one bit perfect before we, you know, look at the other thing. I'm slowly then bring together one by one into their architectural walls.

Jeremy: Yeah, and speaking of, sort of decoupled applications or decoupled services, so event driven and events streaming, that's something you and the team embraced quite a bit as well, right?

Sheen: Yes. Yeah. So from the beginning way, we started using the event-driven approach. We have a bunch of SQS, SNS, all sorts of things. But then the team are now looking into the amazing EventBridge. As you know, that kind of changing the landscape. So that's one of the things, as soon as they announced, we realized the value in it and the flexibly is gives so we gave it to an engineer to come up with some sort of POC sort of solution. So I worked with an engineer and asked her to, you know, do these things and explain the benefits that we gain. So here is how we filter the message and set up the routing rule and things like that and the new service, as me move forward, we are having the EventBridge as a core component. Event buses, a core component, at start up, because, as you know, that the filtering capabilities are far, far better compared to... Yeah, yes.

Jeremy: And you can trigger more services with it.

Sheen: Exactly, exactly. So, yeah, so in my patterns talk, we talk about the EventBridge as well, because I see that, you know, many I already spoke about, it's a cool, cool service that we have.

Jeremy: Yeah, yeah, I love EventBridge. I just started using it in a real project that my latest project that I've been working on for the last couple of months incorporates it and of course they just added CloudFormation support for creating custom buses, but which is good because this project isn't live yet. So I could go back and fix that which I had to do some work arounds in order to get that to work. But so that's there now. All right, Awesome. So, listen, this is this is, I think, been super educational for for anybody who is thinking about kind of bringing serverless into their organization. But while I have you, I do want to talk to you about a post that you just put up. Called “Don't Wait for functionless, write less functions instead.” So just take a couple minutes. I'd love to hear your perspective on this. This concept of functionless.

Sheen: Functionless as a term I heard, I think, probably at one of the conferences. It's probably in Helsinki, they had a panel discussion at the end, and someone was asking what’s the next for serverless? And someone said, “Oh, it's functionless.” Obviously, at that state functionless has been used by a few. So then I started to think about it around that time I was speaking to my friend and his organization moving to serverless, and he was probably explaining a simple piece of architecture he put in place. And there was this Lambda function kind of between an SQS and an API Gateway. And he said he set up the data transporter and I asked him, what exactly it performs. He said, “Oh, no, it just takes the request payload and puts it into the queue.” And I said, “You don't need a Lambda there,” I mean, he became uncomfortable, but then I realized that some off the situation many engineers go through, especially when they come new to serverless, start writing Lambda functions. So that kind of thought process evolved. Then I realized that that a number of ways, we can reduce the use off Lambda functions. So API Gateway being one, and then obviously, we talked about the approach of fat Lambda versus slim Lambdas. And there are a number of other areas and now with EventBridge where we can avoid having custom Lambda functions written. So, you know, when we explain this to people they may say that’s just another Lambda function, why we making so much noise? But thing is, if you look down, you avoid so much of a hassle like poor maintenance, you know, security issues and everything all the integration points and pains and obviously of course, depending on the usage of that function, there is cost implication as well. So that's sort of roughly the how. I came up with this sort of idea and I thought, “Well, okay, let me collectively put everything out there, you know, a nice you know, light-hearted reading so that I'm not hitting so hard, but just conveying the message so that someone reading it will have something to think about it?

Jeremy: Sure. You know, and I love this idea, cause I think things like API Gateway service integrations, they're they're not a silver bullet. I mean, obviously, there are a lot of things you still need Lambda functions for. If you have business logic, then yes, Lambda functions are needed. But sometimes you're just moving date around or doing some transformations even, there are ways to do that without even touching Lambda. And it reduces a ton of complexity, right? It makes your architectural a little bit simpler. I do find that when you start doing some of these integrations or service integrations, that it becomes a little bit more black box and the observability isn't quite there, you know, so that can get a little bit nerve racking. So I know for me I love using Lambda because I can log everything. I know exactly what's happening, but at the same time, there's a lot of reliability and retries and resiliency built into the cloud for you already, you know, and the more you lean on that, the less sort of technical debt and overhead that you have to worry about. So...

Sheen: Yeah, exactly. So that's two things. So one is, uh so you don't need to worry about say cold starts, which is a good thing. But then you don't know exactly what happens inside. And that's an area there that AWS needs improvement. Because velocity template, you know, the struggle getting that right and making that work, that’s what needs, you know, improvement drastically. So yeah, so that that's the other side. But you feel just shifting data, and simple things from, you know, one place to another. And if there is a way not using Lambda functions, then please don’t.

Jeremy: Awesome. All right, well, I'm going to end this with something that Forrest Brazeal said at Serverlessconf, which I thought was quite clever. He said, “Most people have been saying that Serverless is Lego, and now Lego is serverless,” which I think is brilliant so good on him for that. But again, Sheen, thank you so much for being here and, you know, going all over the place telling this story and just sharing your knowledge with the serverless community. If people want to find out more about you and what's going on with LEGO.com, how do they do that?

Sheen: So obviously, I'm on Twitter. So @sheenbrisals. Also, we have the LEGO engineering blog channel on Medium, where I'm trying to encourage engineers to put more out there. That's another place and then also, I share the journey expedience around, and also there are other engineers now picking up and spreading the knowledge around so that's sort of the ways we can communicate. And, obviously, you know the tech community is growing, so there are multiple ways we can get in touch and learn from each other's experience.

Jeremy: Awesome. And I'm sure you'll be speaking at a conference near someone soon. So, again awesome, I will get all this into the show notes. Thanks again, Sheen.

Sheen: Thanks a lot, Jeremy. Pleasure to talk to you.

View Details

This is PART 2 of my conversation with Michael Hart. View PART 1.

About Michael Hart

Michael has been fascinated with serverless, and managed services more generally, since the early days of AWS because he’s passionate about eliminating developer pain. He loves the power that serverless gives developers by reducing the number of moving parts they need to know and think about. He has written libraries like dynalite and kinesalite to help developers test by replicating AWS services locally. He enjoys pushing AWS Lambda to its limits. He wrote a continuous integration service that runs entirely on Lambda and docker-lambda, which he maintains and updates regularly, and has gone on to become the underpinning of AWS SAM Local (now AWS SAM CLI).

  • Twitter: @hichaelmart
  • Github: github.com/mhart
  • Medium: medium.com/@hichaelmart

Transcript:

Jeremy: Alright, so now we're going to go to the next level stuff, right? So if you’ve been...

Michael: That's not next level enough for you.

Jeremy: Well, that's what I’m saying. If you made it this far, I hate to tell you what we just talked about was kid's stuff, right? We're going to the next level. Alright. So you have been working on a new project called Yumda, right? Tell us about this. Because this thing — this blows my mind.

Michael: Right. So this is basically what was born out of the realization that people have struggled traditionally to get things compiled — native binaries or anything like that compiled for Lambda. For example, if you do want to write a CI system like lambci, then you will need some sort of git binary or a git library. But I would suggest using the git binary because libgit is just not there with all the features, But, you know, you'll need to get binary running on your Lambda so you can do a git clone of the repo that you're then going to do your CI test on, and getting that on Amazon Linux 1 was kind of hard enough. Getting it on Amazon Linux 2 is much harder, because there are so many fewer dependencies that exist there. I think on Amazon Linux 1 already had — it has curl on it. You know, if you're in Node.js 8, you could just shell out to curl, so a git has curl as a dependency. So if you're compiling git for, you know, the older runtimes you didn't need to worry about a curl or anything like that. You just need to worry about git. On Amazon Linux 2, you don't have curl, you don't have some really, really basic system libraries. So if you want to get git running on Amazon Linux 2, you need to pull in a lot of stuff yourself. And I got to thinking, well, what would be the best way to provide, you know, a bunch of pre-built packages out of the box? Yes, you could use layers. And I think layers are a great idea for very high-level packages, very, very large binaries that have a huge tree of dependencies or certain utilities. But it's impractical to be creating a layer for every single dependency that your native binary's going to use. You don't want to be creating one layer for libcurl, and another layer for libssh and another layer for this. Firstly, you're only limited to five layers that you can currently use in your Lambda, so you'd need to be squashing them together anyway. And secondly, it's just layers, certainly as they stand at the moment, they're not — if there's no particularly good discovery around them. It's nothing like doing an npm install or a yum install or something like that.

Jeremy: Well, and I also think that many of those layers, that if you install five layers, that a lot of those might be sharing dependencies under the hood as well, like they might have shared dependencies and then you might be installing those twice or three times. I don't know if they would...

Michael: Right, right, could they be clashing.

Jeremy: Yeah.

Michael: No, no.

Jeremy: But anyway, sorry.

Michael: No, no. So that's another consideration. So I thought, well, ideally, what people want to do, and this is certainly what people do in the container world, if you're writing a Docker container, you know, and you need native dependencies, one of the first steps you'll do in your Docker file is you'll do yum install whatever dependency I need. And that'll go down and pull all the sub-dependencies and then that will be installed in your Docker container. And then you can, you know, run your app from there knowing that this stuff exists. We don't have anything like that for Lambda, so I thought, well, I want to run YUM install essentially and have all those packages — all those Amazon Linux 2 packages that are there — you know why couldn't I just get them and install them for Lambda? And the reason that you can't do that is when you run a yum install, it installs in the system directories; it installs software in /usr/bin, /usr/lib64 if it's a dynamic library. And you can’t install that to those places on Lambda. You can only install to, if you're using layers, /opt so /opt/bin is in the path and /opt/lib is in the lb library path, which is where dynamic libraries get loaded from. So you need to make sure that your binaries in your dynamic library sit in those path. That's where they'll be unzipped to essentially when your layer is mounted or /var/task if you've bundled them up with your Lambda function. So you need to make sure that the binaries that you're shipping and the dynamic libraries that you're shipping are okay living in those paths and there's a lot of binaries and libraries out there that aren’t. You can't just copy them from /usr/bin to /opt/bin because something's being compiled into that binary that is assumed that's living in /usr/bin. There are a bunch that you can just move around and that is a good first test. You may as well try it out, see if you can move a library from here to there or see if you can move a binary from here to there. But there might just be something down the track while you're using it, where it’s suddenly like, hey, I can't find this file or maybe it's depending on the configuration file in a path that's been hard coded as well, and you can't get your configuration file to that path because it's not writeable by you. So what I did was I took the Amazon Linux RPMs, and you can get, you know, all these RPMs are open source. You can get the source RPMs. RPMs is this sort of Red Hat package manager format for what a native package looks like on Red Hat Linux and all of its various children, including Amazon Linux, which, which sort of stemmed from Red Hat. So RPMs are what yum what YUM Install will use to install. So I pulled all these RPMs down there, and then I just re-compiled them instead of instead of /user being the path they were compiled for, compiled them for /opt so then you know, I had all these packages that I had re-compiled on, and then I created just a little a little Docker container that has YUM on it that is configured to install these RPMs in the right place because you also — the way that you, if you ever do want to YUM install the package in a non-system director, you have to provide a bunch of configuration to it to let it know that you're doing that. So I sort of pre-configured all that and to talk to the YUM repo that I had set up and that sort of thing. So basically, I've, you know, created a little Docker container where you can just do YUM install git and it will pull down git and all of its dependencies, everything that's been compiled for a /opt environment. And it'll install it all in a directory of your choosing, which you could then zip up and create a layer from, basically. Or you could also bundle it into your Lambda if you wanted to as well. But typically, I think people will want to create layers.

Jeremy: But the idea would be is that if you wanted git and SSH and a couple of these other things you compile, you can compile all those or combine all those into one layer, right? So you just have one layer and you have some limitations there, you know, like package sizes. So obviously you couldn't install the moon, right? You need to be a little bit…

Michael: Yeah, yeah, you're still limited, at least currently, I think the limit is 250 MB or something, I mean, as the total package size that you can have. So, yes, you're still limited to that.

Jeremy: So basically, there's those two sides. So you have your own sort of YUM repo that you’ve built around this. And then...

Michael: Yeah, yeah no, you're absolutely right. So it's a YUM repo, which is literally just it's where all the packages live. You know, they're up on the web, in S3 somewhere. So there's that part of it, all the re-compiled packages that live up there and then the other part is okay, that's fine. They live up there. But how do you tell YUM to A) use that as the repo to pull the packages down and B) to make sure that they know to install into /opt, which in Docker, you mount a local directory into a directory in the container, and that's how you would do it. You know, I'm tossing with the idea of turning this into a CLI, so that you wouldn't need to worry about the Docker run command. But it's a pretty basic command. It’s basically just Docker run Yumda, YUM install package, you know, and then packages, and it'll discover all the dependencies that it needs and make sure that they're all there, and all installed in the right places as well.

Jeremy: And sort of the long term goal here would be to open up that repo, right? So other people could compile and put it in there. But you've got a really good start. How many packages do you have already?

Michael: Yeah, so I've compiled 868 packages so far because the process is relatively easy to convert. Basically usually, the way the RPM packages are compiled, they use a spec file, which I guess is akin to like a Docker file or something like that if you're in that land. A spec file that says, okay, how’s this package going to be compiled and it tries to use variables for things like, what's the top level directory, you know, and things like that. And a lot of spec files, the way they've been written is pretty good, and they don't have any hard coded paths or anything like that. So as long as you re-compile and you pass it in the right variables, it'll just work without even needing to modify any of the source basically — any of the source code. There are a couple of packages, which you know have just assumed — made these hard-coded assumptions — that it's going to be installed in /user because that's where everyone installs everything and so they, you know, I needed to modify something's there. But I think the long term goal would be okay yeah, release this repo and the packages and then give instructions about how people could create their own repo. I think — so one thing about because I tossed around with the idea of: is there another package manager that would be better for this? Would it make sense to NPM install, you know, native dependencies and you certainly could do that, but it would require rewriting a bunch of stuff. There would be some advantages in that, you know, other people could then just NPM publish and that sort of thing. Yum, you know, and RPM, they're much older systems. It's not quite as easy as that, to just publish your own package, you need to kind of host your own Yum repo. But I think, yeah so I'd provide instructions about how you'd want to do that and then obviously, try to accept as many sort of pull requests and ideas for hey, I want this package and that sort of thing. I think that's the reason I haven't released this year. It’s because I keep thinking about do I really want this to be on my plate? What's the best way to sort of tell people, “Look, I'm happy maintaining a certain, you know, number of packages and obviously some core things that you might need,” but I don't know if I want to be brew or you know, homebrew or something like that.

Jeremy: But listen, the thing, though, about this is that if people don't realize how powerful this is, I mean, just think about the ability to quickly and easily compile a layer that has git or SSH and a lot of the use cases, the packages that you've done, you know, they play into, like, a use case like CI/CD. But at the same time, you were telling me about some of these ones you have. I mean, you did like GraphicsMagick for image manipulation...

Michael: Yeah, exactly. So there’s image manipulation. There’s sound conversion that you might want to do. There’s video conversion. All of these sorts of things are going to require native binaries, basically, because they’re difficult things to do. PDF rendering, things like that and Gojko, who you mentioned earlier, he's done a lot of exploration on this, but, you know, I remember on Twitter watching him bang his head against the wall about just how hard it is to figure out how to compile some of these things for that restricted environment. So this would really allow you to, you know, pull in a lot of these pieces that you might need.

Jeremy: But you've got runtimes compiled. You've got, like, Apache and MySQL. So you could literally use a serverless...

Michael: I’ve got MySQL, postgres...

Jeremy: ...So you could use a Lambda function to spin up and test. You do integration testing on your LAMP stack project. It’s kind of wild.

Michael: Yeah. I’ve got PHP and Python compiled and a whole bunch of things. Yeah you could run MySQL locally on Lambda, and I actually think it wouldn't be as crazy as it sounds. You know, a lot of these systems that, you know, a lot of them have options for hey, just start with an in-memory database or something like that. Start up because people do do integration testings with these servers.

Jeremy: Crazy. Alright, so listen, we've been talking for a very, very long time, but there's another thing that is even more — I don't know if this is higher level than what we just talked about. You wrote an article earlier this year called Massively Parallel Hyperparameter Optimization on AWS Lambda. And you were using this ASHA technique or some paper that you read. And now I’m going to tell you — this is above my head, too. So everything you're saying — I was reading the article, glassy-eyed, like not sure where I was, you know. But this is just really, really cool. And so maybe for my benefit and maybe for some of the listeners, you could explain what you did as if I was a five year old.

Michael: Sure. So, basically, there's a machine learning tool out there called FastText, which is for categorizing text. You know, spam is a classic example of this. This text is spam. This text isn’t spam. This text is spam. This text isn’t spam. We do a lot of that, that sort of thing at Bustle. You know, we might use it to aid us in saying, “Okay, this article belongs in this particular category of vertical. This article belongs in this one. And you typically, you often have a bunch of training data. So here’s data where we've had humans come along and manually label this stuff and they might have done, you know, if you're lucky, a few 1000 articles like this, but then we've got 300,000 articles that we need classified. It's just, you know, it would be incredibly tedious to get people to try and classify them all. Machine learning's great at this so let's do that. But you need to sort of train a machine learning model to do that and machine learning models are quite finicky in terms of what works on one data set might not work on another. You might have to tune certain parameters about the way that the machine learning algorithm runs. The way in this case, there’s a binary that runs the machine learning algorithm. The way that it runs is a bunch of parameters that you pass to it that affect how good it's going to be on a particular data set. So you essentially need to tune it. That process is called hyperparameter tuning. It's the idea of okay, I want to adjust the parameters to this algorithm so that it's sort of suits out our data set the best. And there are a number of ways of doing this. You can kind of try and do an exhaustive search of all of the combinations of parameters hopefully that you can find. But in practice, that ends up being a pretty bad approach. Yes, if you wait long enough, you know, all of the combinations to have been tried then you'll have a good idea of what was a good combination. But it could just take a really long time. And often with these things, it’s a little bit like programming. When you start a job running like that, that might run for hours and hours and hours, it might only be five hours in or a few days in that you realize, “Oh, hang on. I messed something up. Ah God.”

Jeremy: Everybody knows that feeling.

Michael: You know, or maybe you have this lightbulb go off and you go, “Wait. I could have included this extra data in there, and I think that would give it extra accuracy. Okay, scrap everything that I’ve done. I'm going to rerun the experiment.” So there's that sort of thing as well. So there are techniques out there, and you can use SageMaker to do this as well. It will sort of try and autotune your hyperparameters for you. But it typically takes hours, if not days, to run these sorts of jobs because, you know, they're speeding up big instances and they often using techniques that very fairly serial in nature. So they need to wait for a couple of different parameters to have been tried before they before they say, “Okay, maybe if I move in this direction, it will be a better set of parameters,” and it could just take a very long time. But there are some algorithms out there, the simplest, of course, which would be a random search. So just randomize all the parameters. Try that and see how you go and then randomize them again. Try that. Now that's a perfect case for where you could do a parallel search because you could just start up thousands of searchers, all with completely different random parameters, and then that'll, you know, maybe take roughly the same time to finish doing these sort of training jobs, and then you come back and you just pick the one that had the best accuracy. That actually works surprisingly well. And in the world of hyperparameter tuning, there's a lot of algorithms that really struggle to beat random search as a baseline. A bit like in the drug world, beating the placebo is really hard. It's similar to that with hyperparameter tuning, but this particular algorithm uses random search but does it in a few phases. It'll you know, the first phase it'll spin up thousands with different random parameters. Then the next phase, it might, you know, cut that in half, do a bit of a binary search and say, “Okay, half of the ones that were the best, let's test them again and change them slightly.” That's a very simple way of thinking of the algorithm. And I just thought it was a perfect use case for Lambda because I was like, you know, I'm sitting here completely in the dark, you know, testing on my local laptop, all these different sorts of parameters. And, you know, I feel like I'm just in the dark. I'd love to just be able to to do this 10,000 times, and then it comes back to me and say, “Hey, here's the best combination parameters.” So I got FastText, which is a native binary. I mean, this machine learning binary built compiling on Lambda, you know, using my Docker Lambda. That was pretty easy. And then it's just create a Lambda function that called out to the binary. There are some limitations. Obviously you only have 500 MB of disc space, so if you're needing to train on data that was bigger than that, you just couldn't do it at the moment — or at least you couldn't train on all of the data in a single Lambda. You’d need to come up with a clever technique to split that up. But the data I was training on, 500 MB is quite a lot for text. I think I was training on 50,000 articles or something like that. So that was fine. Wasn't going to run into any limitations there on and so I could just do a test run and each Lambda I would invoke with a different set of parameters. And then, you know, you use sort of just a coordination process to then, once the first batch had finished, use this ASHA algorithm to figure out which set of parameters do I keep and then try and manipulate. But, you know, I was getting results even within the first five or 10 seconds, because I was speeding up 3000 Lambdas in parallel, you know, and that's 3000 experiments being running parallel. You're going to get quite good results, unless the hyperparameter space you’re trying to search, you know, the number of parameters is incredibly large. 3000 covers of fair bit of that space, so already within the first launch, you're getting very good results. And then, you know, if you're successively sort of narrowing down that search space, yeah within 30 seconds, I got basically a state of the art result, because I then went back and benchmarked it on some of the data sets that they use in the papers. And, yeah, I was getting state of the art results within 30 seconds, whereas I tried the same thing on SageMaker, and it was, you know, within half an hour it still hadn't returned a result that was any way near as good as what I got on Lambda. So I think for things like this — and there are plenty of other examples that you could imagine in the AI ML worlds. Reinforcement learning is another perfect example. Like game playing — anything like this where you're trying to tune and algorithm and spin up many, many instances and run many sort of games in parallel or many environments in parallel. Anything like this. I think Lambda is a great use case for it and and there are some limitations at the moment. But, yeah, I'm hopeful that AWS will, you know, just bump them up a little bit, increase them a little bit and then Lambda will become a supercomputer.

Jeremy: I was just going to say though. That's the use case, right? I mean, that's where we're getting to where that promise of Lambda and that parallel computing being that supercomputer is there. And I'll just say I mean, I've spent a lot of time working on some NLP stuff and then using the output of NLP to do some like multi-class classification stuff with Baseyian and, like I mean, but honestly, like it's all — it works well. It works really, really well, but the stuff you're talking about is just it's insane and the fact that you're pushing the limits is pretty crazy. So alright, so again, we've been talking for a very long time. I’ve got one more thing to ask you, because you and I agree on this, and I want you to share your real world use case here because people ask this question all the time. Some people think it’s an anti-pattern. Lambda’s calling Lambdas. Reasons why you would do this? And you have some very good reasons for that, right?

Michael: It's true. So I'll start off with, I think, why most people suggest it's not a best practice to call a Lambda from another Lambda. And I think that's A) they’re specifically talking about a synchronous use case. So when you're using the request response mode of invoking a Lambda. So you're invoking the Lambda. You're waiting for it to finish, and then you're using the output of it to then do something yourself. So, you know, there are some obvious caveats with that. One is, well, if the timeout of the Lambda you're calling is greater than your timeout and it runs for a really long time, you might timeout before the other Lambda comes back. So that's one thing to think about. And then the other thing to think about, of course, is, if you are doing things at massive concurrency, then you might be hitting, getting close to your limits and if you don't have sort of per function concurrency set up or anything like that, you might be getting yourself in a situation where, you know, you're chewing through your concurrency basically because you're changing your Lambda and you're calling them like that. And actually Joe Anderson just pointed out on Twitter, another why people say this is because they perhaps they think that people have split up their functionality into lots of different functions, and they're trying to compose them in a way that would really be better just composed in a single Lambda itself.

Jeremy: Exactly.

Michael: I 100% agree with that. If you're literally just trying to like call a method from another class, don't turn that class into another function unless this is an incredibly good reason for it to be living, you know, like it's managed by another team or something right about then that I think that's a good use case. So I think that's why people say that. I think, of course, the simple comeback is: well, hang on, but Lambda is just another API. Are you saying that I shouldn't call any API from my Lambda? And people might go “Well, no, you can call other APIs. Maybe just don't call Lambda from Lambda. It's an anti-pattern.” And you go “Well, hang on. If I call other APIs, what if that API’s backed by Lambda?” You know? What's the logic? You know, you need to give people a reason, I think, for why you're saying this is a bad practice. There are certainly use cases, I think where asynchronous patterns work well and I think this is true for any microservices, regardless of if you're using Lambda or not. If you're needing to wait 10, 20 seconds for an API to get back to you and this is a request that a user’s waiting on, there's probably a better way to architect your app, you know. And that's where SNS or SQS, or, you know, using some sort of messaging system is probably a good idea and having an architecture in place that you can return early to the user and then go back and poll. But you can invoke Lambdas asynchronously, you know, like the idea of putting a message on an SNS queue and then when that SNS message gets picked up by a Lamba, it does something. I would say, well, think about what you're actually getting from that on top of just the Lambda asynchronously invoking the other Lambda because you can do that. You could, you know, just say you call this using the event invocation style, and that API call will return within milliseconds. It'll invoke the other Lambda. It'll go and do its thing, and it will return within milliseconds. Now...

Jeremy: It bypasses all that API Gateway stuff.

Michael: Right. It bypasses all the API Gateway stuff. It's incredibly low latency. There's no, you know, it's very fast assuming that you haven't run into a cold start or something like that. But yeah, it's probably as fast as calling SNS. It is, you know, in a lot of cases. So I think it's perfectly valid for that. I think I'm actually the sort of person that prefers to have much less architecture. I think queues and notification systems and things like that can be really useful. And especially, well, if you get a dead letter queue for free and it’s something that you can go revisit later, that's a good use case I think, you know, for having queues, if you need to throttle things and there are plenty of reasons, obviously, for having these things in place. I just I prefer to start without them. And I think you'd be surprised at how far along you can get without needing some sort of intermediary and it probably saves you a lot of headache. And look, maybe if it fails, you log it, and then you haven't alert on your logs, like, you know, it's not as though, there aren't patents for dealing with this.

Jeremy: Yeah, and I mean, so just my quick two cents on this is: I do this all the time where you usually don't want to wait for a synchronous request from a customer, goes through API Gateway, hits a Lambda function, and then that Lambda function has to call another Lambda function to get something and then return it back. Although, I will say when you build other services, you might have a user service, you might have an article service, whatever it is, that one of those services does need to grab some data in order to denormalize into its own service, that it's pretty fast. And if you set low timeouts, you know, so you say, “Listen, this doesn't respond within three seconds, then I'm just going to go back. I’m going to send it to an SQS queue or I'm going to log it somehow and fail back to the customer.” Or say to the customer, “Hey, we got it. Alright, we'll deal with this. We’ll deal with it later.” I think that's a perfectly good use case because you're just calling an HTTP connection when you’re calling Stripe or calling…

Michael: Right. There's nothing special about Lambda in this respect.

Jeremy: Exactly.

Michael: It's like, this is just sort of best practices if you were calling any API or if you're writing any API that if you're waiting for many, many, many, many seconds, then you might want to deal with that. And those are the sorts of use cases where I think, okay, fine, that's perhaps not a good practice. You actually, you asked me. We use this at Bustle. So we have a Lambda that renders our frontend HTML code. It's a preact app. It does service side rendering of the HTML, but it delivers to the, you know, to the browser via API Gateway and a CDN and things like that. But it calls our other Lambda directly, which is a GraphQL backend. It calls that to pull in the data that it needs to render the HTML page. Now in the browser, it also will call that GraphQL backend , but it will do it via API Gateway. Because it's coming from the browser, so it needs to make some an authenticated HTTP request into the function. But when you're in the Lambda world, well, that Lambda can just call that Lambda directly, and call the GraphQL Lambda, and that goes to Redis and Elasticsearch wherever it needs to pull the data and send it back. And we just make sure we have the timeouts tuned such that, you know, I mean it responds within milliseconds. It’s not even a thing we would really run into.

Jeremy: But if it didn't, I mean, you build in that resiliency, right? You just figure out what do I do if it does fail? You know, the happy path on these things, which is 99.9999% of the time, or whatever it is, usually is going to give you a low enough latency that these things don't matter. And then the other thing I would say too is, oftentimes, when you send a request that is asynchronous, that asynchronous Lambda function that's running, you have a little bit more flexibility. You can wait a little bit longer if you have to make some synchronous calls from — You know, once it's disconnected from the frontend, who cares if it takes a little bit of extra time? I mean, there's some tuning you might want to do there from a cost standpoint, but if you need to pull data from four different APIs in order to compile some, you know, some object that gets saved so then it's denormalized and accessible by a customer with a single call to DynamoDB? Do that. It's a good way to do it, in my opinion anyways, and I do it all the time and never had problems.

Michael: 100% agree, and I think this is actually the thing — it's a pet peeve I have with a lot of the sort of best practices that, you know, that you see a lot of the thought leaders in our space out there talking about. Best practices are like a spectrum. They’re not binary. And there are things that people tell you you shouldn't do from a Lambda or whatever. But it's like if you know what you're doing, of course, do it, but just know what the sort of failure cases are. And if it's like, oh, okay, if I do this bad practice and my function is going to fail like half a percent of the time and I know exactly how to deal with it when it does fail, then you forget about it. And I think this is true for things like, oh, don't make TCP connections from your Lambda and things like that. It's like, Well, why? Let's break down why that's considered a best practice, a bad practice or something like that, because you're making TCP connections every time you call HTTP. So and little things that this — even like having large Lambdas and, you know, that sort of thing. It's like, well measure it first, see if it's really a problem for you. Don't try and yeah, prematurely optimize. Because I'll tell you, we at Bustle, we have very few, very large Lambdas, and we do billions of invocations a month. You know, we do many, many, many, many page views and our latencies are very low, and it's not perhaps as bad as you think. There are some tricks you might need to do, like we webpack all of our JavaScript into one single file, right so that there's no file system calls being made whenever it's required, and we minified — so there might be people that go, “Oh, that's kind of a gross hack, but well, alright. For us, that's fine.” You know, we've got plenty of developers that know how to do that and that are comfortable doing that and would be less comfortable managing 50 or 100 tiny functions, maybe, and dealing with the ops of that because it's not free. You know, a function isn't a zero cost piece of infrastructure. You still need to monitor them, you still need to maybe tune them. There's a whole bunch of things that every function you have, you need to think about a little bit and monitor and that sort of thing. So yes, so even things like that. I think there's a spectrum for best practices, and I would say, try things out first. Maybe be aware that it's a lever that you can pull, but try them at first and don't stress too much about having the perfect — there's no single way to do these things basically.

Jeremy: And this is the last thing I'll drop in here is the fact that even if it's not cost-optimized, I mean, unless you go from like zero to a billion invocations on day one, the cost is going to be low enough that you can experiment with this stuff, and eventually get to a point where you know you'll get it tuned the way you need to.

Michael: Right. I agree. And typically the sort of cost tuning that you're doing is going to be nothing compared with the time cost or the people costs or whatever, you know, it is that you would have spent in the extra development trying to do something in a particular way. Probably.

Jeremy: Like you could, for a $180,000-a-year engineer spends five weeks to figure out how you can save $50 on a Lambda function.

Michael: Yeah, nope. You end up doing these calculations in your head all the time as a CTO, and it's like, yeah, forget it. I mean, this the CI argument as well. I'm like, CI seems to be this last bastion of where we're still actually willing to accept multiple minutes of something happening. That's like — what? No. This stuff should be immediate. We shouldn’t be accepting anything less.

Jeremy: Alright, Well, listen, this has been awesome. Honestly, this should be a 500-level course at re:Invent. They don't even have 500-level courses I don't think, but seriously, this was great. And the stuff that you're doing, obviously, AWS Serverless Hero and all these cool things you're doing this Yumda thing. I think it's going to be a huge game changer. It's going to open up a whole bunch of use cases for people that, you know, right now might be compiling these things to Docker or using an EC2 instance. So how can people find out more about you so they can stay up to date with all the stuff you're working on?

Michael: Right, so I am @hichaelmart on Twitter, which is just Michael Hart, but with the first two letters swapped. I’m the same on Medium, I am MHart on GitHub.

Jeremy: And that's H-A-R-T.

Michael: Correct.

Jeremy: Awesome. Alright, I'm going to get all that into the show notes. Thanks again, Michael.

Michael: Thanks, Jeremy. It was great, as always.

View Details

About Michael Hart

Michael has been fascinated with serverless, and managed services more generally, since the early days of AWS because he’s passionate about eliminating developer pain. He loves the power that serverless gives developers by reducing the number of moving parts they need to know and think about. He has written libraries like dynalite and kinesalite to help developers test by replicating AWS services locally. He enjoys pushing AWS Lambda to its limits. He wrote a continuous integration service that runs entirely on Lambda and docker-lambda, which he maintains and updates regularly, and has gone on to become the underpinning of AWS SAM Local (now AWS SAM CLI).

  • Twitter: @hichaelmart
  • Github: github.com/mhart
  • Medium: medium.com/@hichaelmart

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Michael Hart. Hey, Michael. Thanks for joining me.

Michael: G’day, Jeremy, mate. How’s it going? You having a good day? Is everything going alright so far?

Jeremy: I love the Australian. I love the Australian accent. You don't actually talk like that, but that was...

Michael: I don't know what you're talking about. I talk like this all the time. Yeah, I do. I do wonder. I feel like if I did speak like that all the time, people might find me charming, but I don't think they'd have a clue what I was saying.

Jeremy: Exactly. Yeah. No, I actually thought Australians spoke English until I met a bunch of Australians, and I said I don't know if that's English, but anyway, so it's awesome to have you here. You’re the VP of Research Engineering at Bustle. You're also an AWS serverless hero. Why don't you tell the listeners a little about your background, what you do, and what's going on at Bustle?

Michael: Sure. So a little bit of my background: I have started a couple of companies, co-founded a couple of companies, been CTO before, in Australia. This was moved to New York, did a bit of consulting and then joined Bustle as the VP of Research Engineering. So I, you know, do a bunch of interesting research things there. Bustle is a digital media company. We have a bunch of sites, mainly targeted at sort of millennial women, although we've recently been expanding that market.

Jeremy: Awesome.

Michael: And we have, just in the last year or two, I think I think we've sort of acquired or started about nine other sites, so yeah, growing.

Jeremy: And you guys are using serverless up and down your entire stack.

Michael: Yes, serverless across the board. Yeah. Been pretty early on that, yeah.

Jeremy: Awesome. Alright, so I have had a number of conversations with you. We were out in Seattle. We were out in New York City the other day. We've had a ton of conversations about serverless and Lambda and all these things that it can do. I would have recorded the conversations, but usually we're in a bar drinking Old Fashioneds or just being, you know, whatever, and the audio quality wouldn't be that good. So anyways, I want to talk to you about all these cool things that you do with Lambda functions because I have talked to tons of people and I capture use cases in my newsletter every single week and, you know, they're interesting things, but I don't think I've met anybody who has pushed Lambda to the limits like you have. And, I mean, not just like one thing, like multiple things. So I want to get into all of that stuff. But just maybe we could start by talking about, you know, in case people don't know, what is it— The Lambda function itself, it's actually an execution environment. There's an Amazon Linux runtime underneath there, or operating system underneath there. So you know this inside and out. And this will become abundantly clear that you know probably more about this than some of the AWS engineers as we go through this, but just let's start with that. What is a Lambda function? What is it made up of?

Michael: Sure. Yeah, so you're absolutely right. It is the environment that your function is running in is sitting on Amazon Linux, until very recently until the Node.js 10 runtime. That was all Amazon Linux 1, which is getting pretty old now. And then the ruttimes themselves would would sit on top of that just in a directory in the operating system, and each runtime would have, you know, whether it's running on Python, then the Python binary and all the libraries. And if it’s Node, then the Node binary and other libraries so that it'd sort of be the only difference between those two runtimes; the underlying operating system’s still the same. I mean, and these are launched very quickly. Now it's on Firecracker, which is Amazon’s sort of new VM-type technology that sort of provides isolation. But essentially, you know, these isolated environments spin up very quickly and they're running an operating system that runs a runtime that then invokes your function, which is also sitting on the file system.

Jeremy: Alright. So let's get into a little bit more than the details though, I mean, in terms of things that are installed and ready to go. I mean, it's more than just the runtime, right? I mean, there's other libraries, and other things...

Michael: No, you're absolutely right. So, unlike — I'm trying a bit of a good example — unlike if anyone's played with Cloudflare Workers or something like that, that's just running JavaScript. There's no sort of file system that you have access to or anything like that. It's just sort of a JavaScript environment with all of the things that you would have access to if you were in the browser, for example, those so you know most of, or a lot of those sorts of APIs. In Lambda, there is a file system in there. It's a Linux operating system running a process and and then including your Javascript file or Python file and running that. So as well as your file and and the runtime itself, there are, you know, a bunch of base operating system binaries sitting there that you can access as well and Amazon’s pretty cagey about what guarantees they give you about what binaries will be there. They say, “You know if you want to compile something native that your Lambda uses, compile it for an Amazon Linux environment, you know, and that's pretty vague because obviously, you could have an Amazon Linux environment with a whole bunch of binaries or dynamic libraries installed or you could have an incredibly stripped-back operating system that has nothing installed. So in that case, you'd need to sort of bring those binaries yourself into your Lambda. So a good example is if your Lambda function wanted to call out to Bash, do you just assume that Bash is there on the operating system? That's probably a pretty fair assumption. Bash is on most Linux operating systems or certainly the larger ones. So that might be a fine assumption. But then another example might be Perl. You know, maybe you need — maybe your function does something a little bit exotic. Maybe it's doing some cool image manipulation or video manipulation that it needs to call out to a Perl script. Do you assume that Perl is in the Lambda environment or do you bundle yourself or include it in the layer or something like that. So, yeah, those are the sort of questions I think that you need to think about, but because a bunch of binaries do exist in the operating system and you can kind of see them there in your Lambda, it's quite tempting to use them.

Jeremy: Yeah, and so I think that's actually something that's really, really interesting, because when I first started using serverless and Lambda functions, it was basically, the idea was okay, great. I got a stipid of code. I upload it, and it can access the database or it's going to, you know, call an API, right? It's going to use some of these basic functions that are built into the runtimes. But then, you know, every once in a while you have that image manipulation thing. You know, so you have some sort of resizing an image or, you know, and there's basic binaries out there, and this used to be really, really hard, to sort of compile those, and then you have to package them. And it was such a pain to do it. Obviously, there’s Lambda layers now, and we can talk about those in a few minutes. But so let's talk about — before we get into that stuff — let's talk about the Node 10.x runtime stuff, right?

Michael: Right, right.

Jeremy: So this was sort of a huge — this was kind of a big leap when they when they implemented this.

Michael: It was and they didn't announce it as such, not from memory, anyway. They were basically, just like, “Okay, we're now supporting Node 10.” But then when I looked it up, I was like, oh, hang on a sec[ond]. This is actually quite different to all of the other runtimes up until this point, even the custom runtimes that they had introduced used the same sort of base operating system, whereas Node.js 10 was running on Amazon Linux too. So all the other runtimes run on Amazon Linux 1 and Node.js runs on Amazon Linux 2, which, you know, is a lot more modern and has a lot more modern, a more modern kernel, and more modern binaries available for it and, you know, if anyone's using EC2 and using Amazon Linux, that's exactly what they will have been using for quite a while now. So I was running Amazon Linux 2 and it was running an incredibly slimmed down — or it runs an incredibly slimmed down version of it. The entire OS is something sort of like around 90 MB or something which, whereas Amazon Linux 1 — I'm gonna pull a number completely out of nowhere, but it's something I want to say. It's 500 or 600 MB or something like that. What you'd consider maybe a fairly standard OS installation. Whereas Node.js 10 is running on OS that doesn't even have the “find” command. It doesn't even have — there's some very, very basic commands that they've just removed completely, which, when you think about it, kind of makes sense because these Lambdas, you want them to spin up as quickly as possible. And if you are only running, your .js code or your Python code, and you're not using any native dependencies or anything like that, you're not relying on any binaries, or you've compiled your own go binary or something like that. You don't want there to be any extra fluff and you want the little VM instance that it’s running on to spin up as quickly as possible. So it makes sense that you'd want the smallest OS possible. And this is true for anyone who's using Docker, you know, they know this, that the smaller your image is, the smaller your container is, the quicker it starts, so it completely makes sense. But it's certainly a little bit of a departure from what they had done before. Thankfully, you know, they didn't they didn't sort of do it under the feet of anyone on the existing runtimes. It was a completely new runtime, so no one's code was going to break. But if you did want, there are definitely a bunch of functions out there that you can't just move from, say, Node.js 8 to Node.js 10 to move from one runtime to another, if you were relying on a certain behaviors or, you know, expecting certain binaries or dynamic libraries to exist there.

Jeremy: Yeah. I mean, there's been a bunch of changes with sort of how callbacks work in some of that other....

Michael: Right. So that's actually on top of the OS itself. Yeah, so that was another departure - you're right - as well as the underlying operating system and how tiny it was. The runtime that sits on top of, which is really only just a JavaScript file or two that gets executed, then loads your code and runs it — that had been completely rewritten. The runtime that exists on Node.js8 and Node.js 6 and Node.js 4 was all pretty — the code’s basically the same. It's the same Node.js code that loads your handler function, and then runs it. The code that was on Node.js 10 in the runtime was quite different. Firstly, it's using the same mechanism that custom runtimes use, which is sort of an HTTP client method, as opposed to the other runtimes that were using some sort of native binding to sort of speak to the Lambda supervisor. So it was written in a completely different way.

Jeremy: And in terms of some of the other breaking changes — so I know there was a bunch of complaints about the logging stuff, right?

Michael: Right. Yeah. So, as I do, I kind of love to look under the hood of these sorts of things. So I had a look at the code that was written for the Node.js run-time. AndIt wasn't not good, I would say. Just as a Node.js developer who's been developing for many years, it was code, but you know, you get an intuitive feel for whether someone kind of knows what they're doing or not. And it looked as though the code had been written by people who had never written Node code before. They were doing things where they were completely ignoring call backs. They were completely — it was not how you would write asynchronous code in Node. I did a blog post, made some complaints about it, and other people had seen changes as well that, you know, weren't necessarily directly as a result of the quality of the code, but was certainly as a result of some of the decisions they made. Errors were being swallowed at the top level and not being printed out as they were on other runtimes. And yet, as you say, the logging had changed. They had added another field in the logging. So any of your existing logging as is, or anything like that wouldn't work anymore. They — and then this is still true — strip out any new lines in your logs and replace them with carriage returns. And that could be lossy potentially, depending on how complex the thing that you're logging is, but it meant that, you know, a log that previously would be multiple lines is now all on the one line. And these sorts of things I think would be fine if they had done them from the very start, but because they hadn't and they change them in the new runtime, I think people upgrading to that runtime, you know, have had to maybe just deal with things a little bit differently.

Jeremy: Well, I will say anytime you see a Node script that uses double quotes instead of single quotes, you know that’s suspect, right?

Michael: Yeah, single quotes for life.

Jeremy: Exactly. Alright, but right now, I mean the thing was, is that you're right. When they first came out, there were a lot of things you wrote about. You had a great blog post about that whole thing. But right now, I mean, I've been using it now and in my latest project, and I mean, I've been pretty happy with it. It is fast and stable. I think it's a huge improvement over what was there before.

Michael: 100% agree. They they went in. They made a bunch of changes. I mean, the runtimes actually change every now and then. I noticed this, but because you know, I have this Docker Lambda project that we might chat about later. So I'm constantly sort of checking for changes that will have happened. And they actually changed the runtimes every now and then, sometimes it’s security patches. But sometimes it's just that they'll add something in a little bit extra or that they're clearly they've added something in because there's been an edge case or a bug. You can kind of see exactly what it was with the code that's changed. With a Node.js 10 runtime, yeah, they completely rewrote all of the code, and it's much better now, and I agree with you. I think I would recommend anyone use it. You know, be aware if you're upgrading from from the 8 runtime to the 10 runtime. There are some minor changes with logging. But aside from that, I think everything's pretty rock solid.

Jeremy: So one of things you mentioned with the new runtime is that the Amazon Linux 2 strips out all kinds of these basic commands, and we generally don't need them. We might not need to do like a disk usage command line within a Lambda function, so we don't need that. We don't we don't need “find,” but we do need some things, and maybe we need a different runtime. So custom runtimes and Lambda layers are two things that were introduced back, you know, almost a year ago now, which is kind of crazy that that much time has gone by.

Michael: Yeah, wow.

Jeremy: But so I still feel like the custom runtime stuff, you know, there's a lot of blog posts about this too, like, why would you use custom runtimes? And I was thinking about it myself, and I'm kind of like, yeah. I mean, you know, why would I use my own Node custom runtime? And I know you have version 12 custom runtime that you've built for it. But I think you have a really good perspective on this because it sort of changed the way I look at it. What do you think about custom runtimes?

Michael: Yeah, so I actually agree with the general advice that, you know, one of the reasons that you choose Lambda is because so many things are managed for you, and you don't need to worry about so many things and there is an issue if you choose a custom runtime that is an extra thing that you now need to think about in terms of patching — at least the runtime anyway. The OS under the runtime will still continue to get security patches and kernel updates and all that sort of thing without you knowing about it. You know it'll happen one day, and suddenly your new Lambdas, whenever they cold start, will be running on a, you know, a patched OS. So you don't need to worry about that with custom runtimes, but you do need to worry about, of course, upgrading the language that the runtime's running on itself and if that happens to be a dynamic language. I mean, I think plenty of people [are] using custom runtimes for languages that just aren't supported on Lambda, and that's an excellent use case. You know, if you want to be using Swift or something like that, or Rust or whatever it is, then I think that's a very valid use case for a custom runtime because you have no other choice other than to use an officially supported language. But yeah, like, why would you want to choose a custom runtime if you're running Node, when there are Node runtimes available? And I think there are a few reasons for this — not many, but a few. One is if you want to include some things that, across your organization, you know that you're gonna need and that everyone's gonna need. Then you can sort of bundle them in the runtime itself. I don't know. You might have some wrapper around your function that every function needs to have and that every function needs to invoke when it first starts up. And you don't want the authors of these functions to worry about that or to also have that extra step of adding a layer because arguably a layer could do that. But having the custom runtime there means that you're in control of that boot-up process. You might need to authenticate with something before the function runs or there might just be a bunch of things that you want to manage across your organization, and providing a custom runtime might be the easiest way to do that. And then there's just, I don't know, if you're the sort of person that likes staying up -to-date I mean, I know that I can release my Node custom runtimes quicker than Amazon can. So when I'm using my custom runtimes, I know I'm actually less out of date than everyone else who's using Node 10 because I haven't checked, but the last time I checked Node 10 was still a few versions, you know, a few patch versions, if not a few a couple of minor versions out from the from the most current Node. So, yeah, if you're the sort of person that is happy managing your own runtime — because that's not quite the same ask as managing an entire operating system or…

Jeremy: Installing Kubernetes.

Michael: Yeah, yeah, yeah, yeah. It's not quite that intense. it's literally just okay, yeah, I know how to get the latest Node binary and turn it into a...

Jeremy: But I actually thought, though, that the — and we can talk about layers too, because actually — well, let's talk about layers and we'll go back to your custom runtime use case. Because in my most recent project I started, I'm like, I'm gonna use layers, right, because everybody says I should use layers and so forth. And then I ran into this problem of alright, well, where do I deploy my layer to? Where do I package that? Is that in a separate service. I'm sharing that across multiple services. What's the service discovery on that? There's no semantic versioning on it, and everything's just adding a number. And then I found out hey, I could put it in a SAR app and I can version that and whatever. And then I realized I can just put “GitHub:” and then the path in my NPM, I mean in my package.json file and I could just install it right from my own, with semantic versioning, right? I can put tags and things like that, and I was just thinking like, you know what, for the purpose of this, this is just going to be easier. And the reason why I was even doing that was because I have something that works for specifically formatting logging a certain way. I have an authentication component that will take in the authentication headers or the custom authorizer headers. And it sort of transforms that into a common object that all of my scripts can deal with. I've got an EventBridge emitter function that allows me to easily send notifications to EventBridge and things like that. And rather than writing that service for every single, you know, service that — or writing those scripts for every service that have, I just created those, have those as shared services, and I pull those in when they're installed. There's no dependencies basically, and your use case or what you were talking about with the custom runtimes, was actually kind of interesting, because I'm thinking to myself: if all of my services are using Node and I have five or six different sort of common packages that I want, like, why not have those sort of built into the runtime as opposed to having to make sure that I installed them as Node dependencies and have access to things like that. And that way I could just say if there's a major update to these individual dependencies, I could just update the runtime and use the next version of the runtime. And to me, that seems like it's a lot easier to manage that one thing. Now again, you're still managing your runtime, but for a larger organization that wanted to enforce certain security policies and certain logging requirements or whatever it was, I think it’s probably a really slick way to do it.

Michael: Yeah, 100% agree. I think a classic example of that would also be, you know, the AWS SDK.

Jeremy: Yeah, absolutely.

Michael: You know, so many people use that. Of course, you can use the one that exists in the Lambda itself. But even AWS recommends…

Jeremy: Suggests you don’t.

Michael: ...Recommends against that. It's very convenient that is there because that means you can write a one-line Lambda that does amazing things, but...

Jeremy: Especially if you're forced to write code through the console for some of those…

Michael: Under duress.

Jeremy: Under duress on a live video stream. Yes.

Michael: Yeah. I don't know why you’d want to do that or put yourself in that situation. But yes, yes, for things like that is very handy to have the Node runtime there. But of course, yes. managing especially, you know, a new service comes out and you want to use it from your Lambda. Well, guess what? Hey, it's not supported by the AWS SDK version that's installing Lambda just yet, so you're gonna have to wait or you bundle it yourself. And yeah, a custom runtime would would be one example of doing that. Again custom runtime, it does mean that you need to — that you would need to understand how custom runtimes work. A soon as you understand that, I would say there's no extra onus on you to keep it up to date or anything like that. There's nothing particularly fancy that you need to do. And of course, there are a number of custom runtimes out there that you could just extend yourself and add your own packages to. But yeah, the layer thing. I agree with you on that. I think it's a real pity it doesn't have semantic versioning. And it was really interesting to find out that, yeah, you could do it via SAR if you want, but it's almost like a hack, in a way — a hack of SAR, of like, creating a whole application that is really just a layer. You get some anti-versioning and things like that, and the layer can be across all regions, which is kind of nice and a little bit of a pain if you manage your own sort of open source or a commercial layer that you want other people to use. It can be a pain to replicate that across all regions. But one thing that you should consider, with not using layers, is deploy time. Now, your dependencies might be nice and small, which is great, but, you know, if they start getting up into the multiple MB, if you have something in a layer, you don't actually need to deploy it each time. And yeah, I think it was maybe even Brian Leroux who made that lightbulb switch in my head, because I was, when layers first came out, I was like, “Okay, this is kind of neat,” but, you know, assuming you know how to manage dependencies, what's the point? There's no real point to this because I think I was hoping when they were announced that I was like, “Oh, great, this is gonna give you extra package size.”

Jeremy: Which it didn't.

Michael: Give you extra space. It's maybe going to make cold starts quicker because, oh, they're gonna be able to bake the layer in with the base image and have that all ready to go. That could still be an optimization that they might make, but my understanding is they don't you know, I mean, that would, I guess, I mean, they'd have to bake as it were a lot of many, many, many different combinations of peoples runtimes. But yeah, there is this advantage — so you don't get any other advantages, really, but there is this advantage of deployment time, which is where you don't need to deploy the layer again and all the code that sits in it. And if you have dozens of MBs, that could make your deployment a lot quicker.

Jeremy: Yeah, I definitely think — I mean, I think the public layer stuff, especially some of that stuff that Gojko has done with FFmpeg and some of those others, or NumPy or some of those other things, those are great, because nobody wants to compile those things down. Although you might have an alternative to that that we'll talk about in a minute. But, you know, that's just easier for someone to say why I don't want to have to run a local Docker and compile it down and then do that all myself and package my own layer and then manage that layer and so forth. And I think some of those sort of things are relatively safe. If I need to pull in, you know, FFmpeg or something like that, pulling that in via a layer is so much easier than trying to manage that myself.

Michael: Right, and I think that's a good point is that perhaps for people who are new to Lambda, they might not actually realize that oh, your Lambda is going to be running in a different environment probably than what you're developing on.

Jeremy: Yes.

Michael: You're probably not developing on Amazon Linux. You might be developing on macOS or something like that, and you're used to being able to just sort of, you know, maybe brew install something, or uninstall something, and have it ready to go for your application. And I think this trips up people a lot is when they first go to start using Lambdas is they don't realize that the package that they've installed, you know, that's sitting in their Node modules that they then zip up and fire off to Lambda, well, that was natively composite for MacOS which is what they’re developing on. That's why it works locally, but they've bundled it up. They've sent it to their Lambda, and hey, it doesn't work anymore. Because it was a native binary component for MacOS and I think that so there's that step of going “Oh, hang on. Okay, I messed up. I have to do something about this. What can I do?” And then answer, really from Lambda - it's not a great one - you know, it's like, “Oh, well, you need to make sure that your binary is compiled for an Amazon Linux environment.” And it’s like, did you just tell me to go “F” myself? Like what does that mean? Does that just mean I can find a Linux binary somewhere and use that? Maybe, you know, sometimes you can. But what does a Linux binary mean? What is it depending on? What is it assuming about the environment that it's running in? And I think people, especially newcomers, they end up going having to go down a rabbit hole that they really didn't want to go down and they were just like, “I just wanted to install this dependency.”

Jeremy: Alright, so that's a perfect segue, because now we should talk about Docker Lambda, which is essentially a Docker image of that Lambda runtime environment that you can use to test things locally, compile binaries. You can do all kinds of things in there. Why don’t you tell everybody a little bit about that?

Michael: Yeah, 100% right. It's basically a Docker image or a set of Docker images that are trying to be the most faithful reproduction of the live Lambda production environment as possible and they have all the exact same files, or at least the ones that you have access to. There are some obviously that only accessible by root that I can't replicate. But aside from that, it's the exact same file system and as much as possible, the same sort of commissions because in the Lambda environment, you only have write access to /tmp. You don't actually have write access to the directory that your code is in and that can trip people up as well. So I created it because I was trying to do some some pretty fancy stuff in Lambda, and I was just getting frustrated with this. You know, the idea of oh, I have to spin up in an EC2 instance with Amazon Linux and then compile my stuff there and then copy that over what, to my local machine and then zip that up or, you know, or deploy it from EC2. It was getting very painful when, you know, you could maybe do it once, but then the cycle of development is really slow. So I was like, well, I wonder how hard it would be to at least try and, you know, get a similar sort of environment to Lambda just running locally using Docker. It's a sort of perfect use case for Docker, you know, because Lambda is kind of like containers. And the way that I ended up doing it was essentially by running a Lambda, which tars, you know, creates a tarball of the entire file system and and copy that over to S3, and then I pull that down. And you know, there's a Docker command that you can use to create an image from a tarball of an entire file system — so sort of as much of the file system as I can grab. I grabbed that, put it in a tarball and then create a Docker image from that. And then during that process, as I was doing that, I realized, “Oh, hang on. All of these runtimes, they're actually running the same operating system. It’s just one directory that's different — you know, /var/runtime. That's where the runtime specific code lives. That's where the Node.js 8 code or the Python 3.6 code or whatever lives. The rest of the operating system’s exactly the same. So that's kind of cool. I can create a base image and then each runtime can just — I just need to dump the /var/runtime directory. And there's a couple of others now, there’s /var/lang, /var/rapid I think, is one of the new ones. There's a couple of directories that change within each runtime, but everything else is the same, so that makes it easy to create separate images for each runtime then as well. And then what I did on top of that was — so that's great. That gives you an image that's a replication of the environment, and you can use that to then compile stuff, and that sort of thing. But I thought, well, hang on. I could also use this to mark out running a Lambda, and I could actually use the runtime code itself because it's sitting there so I could use the exact same code that requires your index.js and looks for the handler and then executes the handler with all that. And, you know, on Node, for example, it overrides console.log and adds a bunch of fields so that whenever you console that log, it actually outputs a bunch of extra fields, and it’s the runtime that's doing that patching. So I was like, well, instead of trying to replicate that, I'll just let that code do that, you know? And then the only thing I do need to mark out is obviously, whenever it tries to talk to the Lambda parent process, you know, talk back to the supervisor and say, “Hey, here. Sends these logs off to CloudWatch Logs. I'll just intercept that and we'll output them to the console and that sort of thing, so, yeah, thus Docker Lambda was born, and I think a lot of people started using it. They found it really useful. It's a little bit if you are testing your Lambda locally and you don't have any native dependencies, and it is just a simple function. It's not really necessary to go to the degree of this level of replication. You can very easily write unit tests that just test your functionality without requiring a whole Docker environment. But if you are doing anything like writing to the filesystem or anything like that, it sort of gives you this extra level of parity with the real environment. Yes. So I wrote that. I remember chatting with, at the very first ServerlessConf in 2016, with Tim Wagner about it who is the sort of father of Lambda. He created Lambda. He was at AWS At the time. I was chatting with him about it and a bunch of other things, like oh will Lambda ever be able to run Docker and things like that. I said, “So I've been experimenting with this idea of creating this thing, and, you know, it'll allow people to test.” He was like, “Okay, yeah. It sounds OK, but I don't know. I think people would probably just want to test in the cloud really? Like I think that's where we'll put our focuses on, and we'll just make it easy for people to test in the cloud. I don't think they are really going to want to run Docker on their local machine as part of a testing environment or anything like that.” Or all more just, you know, he was like, “No, that's probably just not an effort that I think we want to pursue at Amazon or anything like that.”

Jeremy: And then...

Michael: Which I think, you know, i’s okay. That's a valid response, and I could buy it if you could literally deploy everything, you know, deploy your entire cloud formations, stacking milliseconds and have a really fast cycle and also be able to, like, have free development accounts and things like that. Then I could buy that you could use the cloud as a testing environment. I still think that's something that Amazon could aim for is to make that process much easier, because what better way than, you know, they're testing in the environment where your production's actually going to run. but But anyways, I created t and then yeah, a good yea later and AWS Sam’s CLI team reached out to me, and they were like, “Hey, we noticed your Docker Lambda project, and we're thinking about writing a local testing utility in Sam’s CLI. Maybe we could use that.” And I was like, “Yeah, great. Go for it.”

Jeremy: So when you said people are using it, you meant a billion-dollar company was like, “Right, we’re going to borrow this, if you don't mind.”

Michael: A trillio- dollar company

Jeremy: I’m sorry. Did I say billion? Trillion-dollar company, yes.

Michael: So now if anyone uses Sam locally. Long story short, if anyone’s using Sam locally and then they’re spinning up the Docker Lambda containers basically.

Jeremy: Awesome. Alright, so that was one cool thing that you did, which was sort of just duplicating the environment — not “just” duplicating — but you did all this work, right? But knowing all those internals, you then took this further and you started doing these other crazy things with it. And you launch this thing called lambci, that basically built an entire serverless build system for CI, for continuous integration. So tell us about that.

Michael: Right. So actually, lambci, preceded Docker Lambda. I wrote Docker Lambda as a utility as I was writing lambci. I was like, well, hang on with Lambda we've got access to, you know, you can suddenly spin up all these instances incredibly quickly. What's one of the most painful parts of sort of development and sitting there twiddling your thumbs? And that's waiting for a CI build to finish. And I was like, “Well, hang on, if we could get CIs running in Lambda, that’d be incredibly fast.” Not only that, but at the time — because this is back in 2016, maybe even 2015 when I first writing it — there was no sort of pay-per usage CI systems. You basically paid per month. You know, you'd be like okay, yes, I want access to one build server or maybe four build server or something like that. You’d pay per month. There was, as far as I knew anyway, there was no way to get a per request pricing or per build pricing or something like that, which, you know, especially in the serverless world, whether that thinking maybe wasn't as common — it seems obvious now — but it wasn't as common back then that we'd only want to pay for the resources that you use. And I think people would do things like they might be running Jenkins, and then at night, they'd shut down their build cluster, and in the morning, they’d start it back up again. You know, you could do cost saving things like that. But I was like, well, hang on, we can just spin up a build in milliseconds and have it shut down again. You're only charged for the time that it’s building. So that's one advantage. And then the other advantages well, you’re not just limited to maybe four concurrents because, you know, I was CTO of a company that had not many developers, but 16 developers. And once you've got 16 developers, if you're really doing CI/CD and you're pushing stuff out all the time and you want, you know, to be able to everyone have their own staging environment and all this sort of thing, you're pushing the CI service pretty hard. And the worst thing is for people to be sitting there stuck in a queue waiting for someone else’s build to finish so that they can then access the CI server. With Lambda, obviously. I mean, you could get to that point, but you'd have to be doing thousands and thousands of concurrent builds basically before you started stepping on someone else's toes. So, yeah, I thought, look, this would be great. It's running Linux. Yes, there are some limitations, but I'm sure there's a whole bunch of things that I could get running. So I started exploring what you could do, and it's very straightforward to sort of do very basic sort of Node testing you because it's just running Node files. But then I started pushing it more. I’m like what if you wanted to do an NPM install and you had native dependencies. How are those native dependencies going to build in Lambda? I was like, okay, I need to get GCC compiled on Lambda, you know, and that was quite an adventure, but I did it and so you know, then it was like well, how you could do MPM install and have native dependencies. And so because GCC’s running on Lambda, it's compiling your files. And initially I think the timeout limit was five minutes and then they expanded to 15 minutes, and that opened up, because people were using lambci and running into this limitation. But then it became 15 minutes and that actually, there are people, obviously, that have builds that run for longer than 15 minutes, but not many, whereas there are a lot of people that have builds that run longer than five minutes. So that opened up a big bunch of use cases there. There was some, I think, restrictions early on about TCP sockets and things like that that stops people from opening local servers that they might want to do integration tests on. You know, you create a local express server that you then test all your HTTP routes against and that sort of thing. You can now do that very easily in Lambda. So as time's gone on, there have been more use cases that you can kind of do, and that you can then use lambci for basically.

Jeremy: So the other thing that I’m just thinking about, too like, you can do other tests. Like, can you run like selenium tests? And like headless browsers?

Michael: You can do a headless, yeah. So, you know, some awesome people have experimented, I guess, just as I have, with getting some crazy things compiled and running on Lambda, including TensorFlow, but also including Chromium. So you can run a headless Chromium in Lambda. You can you can do a whole bunch of UI testing. You can run Lighthouse. You can actually use it for all sorts of interesting things. So you could be doing that as well.

Jeremy: And part of the reason why this is so cool is because, as you mentioned, you have a Jenkins build server or something like that, it only has so much processing capacity, and it runs a lot of these tests serially. I didn’t know if I was going to be able to say that word. I’ve got to channel Peter Sbarski here and say “parallelized.” So you can parallelize these jobs with lambci and you had a post the other day on Twitter where you said you took like build times down from seven minutes down to, like, 15 seconds or something like that?

Michael: Yeah, yeah, yeah, that's right. So we use this or at least, have been experimenting with using this, at Bustle for a while now. And you know, I came on two years ago. There was already an existing CI system in place and we do some pretty hairy things. We talked to Elasticsearch and we talked to Redis and we do that in our tests as well. So we need the CI system to be running Elasticsearch and running Redis. And you could do this in Travis and Circle CI. You can have these services set up, you know, a local Elasticsearch running, and a local Redis. But you could do it in Lambda as well. You can, as part of your build process, download Elasticsearch, certainly in the Amazon Linux 1 runtimes. Java is sitting there. So you can download Elasticsearch and you can run it on and have it exposed locally. Redis, even easier. You know, this is a very small binary and you can run that. Just expose local ports and then your unit/integration tests can talk to that locally. So we do that at Bustle, and that's back-end testing, front-end testing, similar sort of thing. But you know, you're often running tests, unit tests, but you’re also running linting and you’re running formatting. And you know there's a whole bunch of things that you can do in parallel. People might be parallelizing these already in their CI systems if they've got a couple of concurrent servers available to them. But with Lambda, you can really go crazy parallelizing it so I've got it running so that it tests every nth test. You can, with most test runners these days, you can pass the list of files to it instead of just passing a directory, you can pass a list of files so you can just use the find utility and an awk command or whatever to go and grab every fifth file or every tenth file, or every fiftieth file, pass that to your test runner, and each Lambda can be doing every fiftieth file. And you can run 50 Lambdas in parallel and boom. Your tests suddenly run 50 times faster. Any job like that that you can do in parallel that you don't need some sort of serial bunch of steps for, yeah, is a perfect use case for — I mean is a perfect use case for any parallel system. It’s just that Lambda is incredibly good at that, and it's gonna be a lot cheaper to do on Lambda than it is to be buying 50 concurrent servers from, you know, renting them from a traditional CI.

Jeremy: That's awesome. Alright, so now we're going to go to the next level stuff.

Michael: Right.

Jeremy: So if you've been...

Michael: That's not next level enough for you.

Jeremy: Well, that's what I’m saying. If you made it this far, I hate to tell you what we just talked about was kid's stuff. Alright, we're going to the next level. Alright, so you have been working on a new project called Yumda.

Michael: Right.

Jeremy: Tell us about this because this blows my mind...

ON THE NEXT EPISODE, I CONTINUE MY CHAT WITH MICHAEL HART...

View Details

About Brian Leroux

Brian LeRoux is currently building a continuous delivery vehicle for cloud functions called begin.com on an open source foundation called arc.codes. Previously he worked at Adobe on PhoneGap and Apache Cordova. Brian believes the future will be writ as functions, seamlessly running in the cloud, agnostic of vendors, on an open source platform and it will be stewarded by hackers like you.

  • Architect Framework: arc.codes
  • Twitter: @brianleroux
  • Begin: begin.com
  • GitHub: https://github.com/brianleroux

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Brian LeRoux. Hey, Brian. Thanks for joining me.

Brian: Thanks for having me. Excited to be here.

Jeremy: So you are the Co-Founder and CTO at Begin. So why don’t you tell the listeners a little bit about yourself and what Begin does.

Brian: Yeah. Cool. So, I'm Brian. I’m a Webby hacker. I guess you could say I've been building software for a really, really long time now, and my focus has been, in the last few years, the cloud. In a previous life, I used to work in the mobile space quite a bit and Begin.com is continuous integration and delivery for serverless applications — modern serverless applications.

Jeremy: Awesome. Alright, so I wanted to have you on today to talk about the Architect Framework. So this has been out there for quite some time, but just in case people aren't familiar with what it is, maybe you could just tell listeners what the Architect Framework is all about and what you can build with it.

Brian: Yeah, Architect is a serverless framework. It papers over some of the more complex bits of getting up and running with a serverless application. It's sometimes accused of being pretty opinionated, but I think maybe we'll dig into that a bit in this episode and how maybe it's not so much opinionated, just, you know, makes some choices up front for you and saves you time. It's much more convention than it is configuration. And it's really targeted at building super fast web apps.

Jeremy: So let's get into that. So first of all, why did you build this? Because there's Claudia.js and there’s Sam, and there’s Serverless Framework, and there’s — you name it, there's a framework out there that helps you build serverless applications. So what was the reasoning behind it?

Brian: We never actually set out to build a framework. In 2014, I left my role at Adobe. I used to work on the PhoneGap in Cordova projects. And a big part of that was this thing called PhoneGap Build, which is a hosted service that’s part of the Creative Cloud where you can upload HTML, CSS, and JavaScript, and we spit out native iOS and Android apps. And building for Creative Cloud and building a big load-balanced rails application taught me a lot of lessons. One of those lessons was: I'd never want to do that again. And when I came out the other side, Ryan Block and I started the first iteration of Begin, which was a Slackbot, and we didn't really know a whole lot about what we were going to be building or how we were going to be building it. But we knew what we didn't want to do. And what we didn't want to do was take on a traditional load-balanced, monolithic architecture. And this new serverless thing was around and it looked like this was the future. It was 2014 at this time. Lambda was new, but there was no way to do an HTTP call. And then that August, API Gateway was released and I was floored personally. I saw it, and I was like, that's it. It's just got to be the future. So we built the first iteration of Begin, which was a Slackbot, and it had really strong real-time requirements, because it was a bot and had a web app component to it. And at that time, there was only one framework. It was called JAWS and it really was early days. So we built our own thing, but we built a product. We didn't build Architect per se. That project didn't work out, but we extracted Architect afterwards to build our second thing, and we knew we were onto something because, very similar to the Basecamp story, we didn't try and create a framework. We built a thing, and in the process of building that thing, we came up with something that looks a little bit different. When you look at Architect, it's not the same as other frameworks because we weren't trying to build a framework. We're trying to build a product.

Jeremy: And I think that's actually kind of typical of building serverless applications. I know as I started building serverless, as I was working on my first few serverless applications, I built the Lambda API web framework, right, because I just needed a better way to process API Gateway calls using the Lambda proxy integrations. And then,I built the Serverless-MySQL package to deal with the max connections issue. And you start building these components to help you build the products that you want to build, and eventually, you get some really cool tools that come out of it. And that's basically what you did with Architect, and you’ve made that open source, right?

Brian: Yeah. We donated it to the, at that time it was called the JS Foundation. But the JS Foundation, and the Node.js Foundation merged and became the OpenJS Foundation. I've got a pretty long history in open source, and I really believe in foundation-backed governance for projects. And so it was important to me to pull that IP out of a privately-held, venture-backed startup and put it into a place where the commons could contribute back to it. And just as a note, so the listeners understand, I'm not a zero sum thinker. There's going to be more than one framework. Technology tends to be additive, and there's going to be different things that we could learn from each other and build on from each other. So by all means, check out Architect, but, you know, if you're a Python hacker, you would be remiss not to check out Chalice. And if you're deep into AWS, you're probably already using Sam, and I think that's just fine. That makes sense to me.

Jeremy: Awesome. Alright, so let's talk about this opinionated thing because when I first saw this — it is kind of funny — but when I first saw the Architect Framework, I was looking through it, and I'm like, okay, this seems like it is built for solving a very specific problem. And again, I know it can be extended and you've got other things we can talk about, like some of your macros and things that you can build on top of it, but it just seemed very opinionated to me, in the sense that you were enforcing small file sizes, single purpose Lambda functions, that kind of stuff. So what are your thoughts on that? Because you maybe don't think it's so opinionated, right?

Brian: I didn't and Yan Cui from the Burning Monk wrote an awesome blog post where he threw Architect way up in the opinionated corner. And when we first saw it were like, “Oh, weird. Okay.” So we do look different and we do look like we have opinions, but I think most people will share those opinions. So one of our opinions is that we need to be really fast. And we need to be faster both author time and at runtime. So by author time speed, I mean, we need deploy iterations and lead time to production to be really quick. Monolithic apps have pretty poor characteristics for this. They tend to be deployed in minutes to hours, if not days or weeks, whereas serverless applications, because we break them apart into small constituent pieces, or we can, we can deploy those artifacts in parallel. We get a lot faster deployment speeds as a result. So that's really nice. And I like that, and I like small functions for that particular aspect. There's a single responsibility principle as well. When you have bigger functions, they're going to be harder to debug by default. If you have small single responsibility functions, your discovery and maintenance gets a lot easier. If there's a, for example, if there's a bug on the get about page, and there's one Lambda function serving the get about page, well, you know exactly where the problem is. It's on the get about page function. It's not somewhere in this ball of code that you have uploaded. The kind of final piece to this with the small functions bit was really driven by practicality. So in our first version of our bot for Begin, we did the same thing everyone else did. We put an Express web server inside of a Lambda function, and it worked — it worked really well until we started building something bigger than Hello World, and then it started to not work so well. And in particular, the thing that didn't work well was the cold start. And you hear about cold starts all the time. And it sort of irks me because I feel like this problem has been solved for a long time, too. We measured cold start with every different runtime with varying payload sizes, and we determined that it was correlated to payload size. So we did the thing where we just made small functions. So we found that in the earlier versions of Lambda, there was a 5MB kind of magic number. If you were over 5MB, you would be over a second cold start. If you were under 5MB, you would usually be sub-second cold start, which was totally a suitable performance profile for a bot. So we just set up our CI to fail build after functions got bigger than 5MB, and we started dividing up our app in the single responsibility principle.

Jeremy: And by payload size, you mean the package size of the artifact, right?

Brian: Yeah, the zipped package size, actually. Yeah, and I've been talking to other devs about this, and there's been a lot of movement in the last four years on Lambda, and it's gotten a lot better at cold starts. And I imagine these numbers are different today, but just because you can, doesn't mean you should. I think it's totally appropriate to build out your first versions with just a few fat functions. But as time goes on, you're gonna want that single responsibility principle and the isolation that it brings. There's one last small interesting advantage to this technique is that the security posture is just better. You have less blast radius. If your functions are locked down to their least privilege and their single responsibility, you're just going to have a way better risk profile for security.

Jeremy: Right. So the other thing again, maybe this is why Yan and maybe some others, including myself, thought it was an opinionated framework, and that’s because it's — I don't think limited is the right word, but it's specifically curated, I guess, with just a few core components that you can use to build applications. But with that small set of services, you can quite accurately replicate the execution environment on your local machine, right?

Brian: Yeah. This was another really important thing for us, and I think it was, frankly, probably more of a coping mechanism than anything else. When you open up that AWS console for the first time, it's a pretty intimidating experience. There's over 300 services there, you know, and you don't know where to look. You don't know how they integrated with each other and they've got different UIs for each service. And we sat back and really looked at the requirements of our application and distilled it down to its constituent parts, what protocols we needed to support and how. And we realized we only needed eight services. We didn't need 300. We just needed a subset of them to build a CRUD-y web app. So with that knowledge, we really built our abstractions on top of those eight services. We don't hide the other services from you. We just paved the path for those eight and make it really smooth and easy to get on board with them. The canonical example of difficulty for configuration would be probably API gateway. And most people would agree it's a bit of a beast and that it’s a powerful beast, but it's a scary, powerful beast. And so most people just want to give it some URLs and say, “please return values from a function when these URLs get invoked.” They don't want to get into the depths of velocity templates and the rest of it. So we we paper over API Gateway. We make that part look really, really simple, even though it's really complicated under the hood, and then we add some sugar — some terseness I should say — to the configuration file format for Dynamo, SQS, SNS, and a handful of other services that really aren't that user facing. And that's kind of it. We have a macro primitive as well that lets you reach into the cloud formation that gets generated. So you can access anything in AWS that you want. It's just our contention that, most of the time, for a large portion of applications you won't even need to.

Jeremy: And so that's another thing that I think is interesting about the Architect Framework, is the fact that you have those primitives for, like, SQS and SNS. And I know they're named after AWS services, but if you think of like the Serverless Framework, for example, they sort of abstract away the connection to things like SQS queues. Like it'll create certain resources for you, but it's generally in the context of a connection to a function. But if you just wanted to create your own SQS queue, you'd have to write the CloudFormation and put that in the resources section of your serverless.yml file in order to build that SQS Queue, and the same with SNS and DynamoDB is a great example. So by having these primitives — and we can talk more about that in a minute — but is that something that you're looking ahead for, something like cloud portability?

Brian: I think that there is a possibility of portability and a longer run future. I’m less interested in a disintermediation, and I’m more interested in velocity on the de facto cloud. And this isn't sucking up to Amazon, this is just being straight up with where the state of the industry is. If I had better options, I would take them for sure. But AWS has a really big lead in this world, and the other players are, frankly, just catching up. Azure has a concept called Azure Resource Manager, which is — ARM is the acronym — that is their equivalent to CloudFormation. It's very new. Can't really build a full app with it yet. Google doesn't even have an answer to this idea of infra as code yet. So I guess Terraform would be — Terraform and Kubernetes, YAML files would be the answer. So portable where? Would be my question. And I just don't think…I think they'll catch up. I absolutely believe that there's going to be more than one cloud. I just don't know who it is yet. So what I do know, though, is that AWS is amazing. It gives me the characteristics that I want, and they've been a wonderful partner to work with too, so we can build on there with a high degree of confidence that they are probably the de facto standard. And we'll see where the other players get to. But like ARM looks a whole lot like CloudFormation to me, so can we translate CloudFormaton into ARM one day? I don't know. Maybe. Do I want to do that? Not super badly. To be honest with you, it all comes down to the data store and DynamoDB is real hard to beat. I'm sure Cosmos is going to try, but I am an extremely happy locked-in Dynamo user right now, and I don't see why I would adopt more latency to use it. So that's kind of my perspective there. But, you know, is it possible? Sure. Do you want to? Probably not. Not right now.

Jeremy: Yeah, at least not right now. Definitely. Well, I want to talk about DynamoDB. I have a whole bunch of questions around DynamoDB that I want to ask you, but maybe let's go back to the Framework and talk about, you know, how does it help you build sort of these modern serverless applications?

Brian: Yeah. So it gets you off the ground running really fast, and, you know, everyone makes that claim that I'd like to quantify it. So within 10 seconds, we should have a local development environment that fully replicates exactly what you would have in the cloud. Folks would say, you know, “This is impossible.” I heard the same thing about mobile emulators 10 years ago. It's not impossible. It's very possible. AWS doesn't change your APIs all the time. In fact, they change them extremely rarely and again, we're subsetting so like a very small number of services. So we have a pretty kick-ass local development environment. It's a couple years old now, and it's been matured in the open source world, and it works real fast, so you don't need to deploy or even to have AWS creds to get started. And that's a really big advantage for understanding how this thing fits together. We even have DynamoDB running locally using Michael Hart’s amazing DynaLite project. And then the next step beyond that is like, okay, cool. I want to get these bits up in the cloud. And once you've credentialed yourself with Amazon correctly, it's one command and you're there. There's no configuration to speak of, other than the arc file, and everything else gets generated by convention. And this leads to a pretty slick development experience, you know, kind of what we're used to. CRUD apps back in the day where you could generate routes, and see them deployed and see them respond to HTTP events and, you know, have states shared between them but in a stateless execution environment. And that's really the class of app I'm building on AWS today. There's obviously a whole lot of other workloads that are possible, but we're really tuned for that use case, in building a web app as fast as you can.

Jeremy: Right and that's another thing that I really like about Architect, is that you're not trying to be all things to all people, right? And it’s great, that bootstrapping locally or just getting you up and running right away without even connecting to the cloud, I think is sort of an interesting approach to doing that. But in order to deploy, you end up just generating SAM templates under the hood though, right?

Brian: Yeah, but we didn't initially. Actually, this is a fairly recent addition. We did the Sam, Architect 6 is all SAM and CloudFormation-based. We were able to delete roughly 20,000 lines of code, which I'm going to get into in my Serverlessconf talk. So just going to CloudFormation, for the listeners, by the way, like screw whether or not you use Architect. That's cool if you do. But even if you don't, learn from us and and the learning is, we did an SDK-based framework for the first few years of its life and the net result of us moving to CloudFormation was a massive code deletion. And we gained features in the process. We gained a ton of features in the process. CloudFormation is absolutely where the puck is going to be. And to me, it's the de facto standard. I'm certain a lot of people would cringe at that one, because it's not, you know, consortium-based spec standard, but it's the right way to build an AWS application. It's less code. It offers a huge amount of determinism, and yeah, we've been really happy moving to it as the baseline. And so ARC really is a file format and that file format is extremely terse. It's like if YAML chilled out and just didn't and so that the file format, a lot of people get a little bit bent out of shape about, but we had good reasons for doing this. So one, we wanted comments and JSON doesn't have comments. Two, we didn't want deeply nested structures, and YAML really, really encourages deeply nested structures. And so we didn't want to use JSON. We didn't want to use YAML, and so we were using ini like files for a little bit. And then we kind of end up creating our own syntax along the way. And it's really terse. It's extremely readable. You can also write it, and this is a big benefit to Architect. You can look at the manifest file, and within a few lines of code, you will understand what that application does. Nobody could look at a SAM or a CloudFormation document and know what that application does. It doesn't tell you anything. It just shows you a lot of stuff. So we translate that ARC file into a SAM document for you, and we dump it in the root of your directory so you can see the delta for yourself. But usually it's a 60 to 80x reduction in configuration, which is a very huge productivity improvement that is quantifiable. You know, you can see it for yourself, just run it and boom. You just generated a shitload of CloudFormation.

Jeremy: I have spent days working on CloudFormation files, so yes, I know full well that it would be nice. So you had mentioned being able to access those files as well. So the ARC file is your configuration. Listeners have to go check this out, because, like you said, it’s ridiculously terse. I mean, it's so short what you need to do in order to generate probably, what? 500, a thousand lines of CloudFormation at the end of the day or something like that?

Brian: Yeah, and so we've really distilled a lot of the what we feel are the best practices. Maybe, we've distilled a lot of the opinions that are out there on how to do this. So because we know all the resources up front, we can give you a least privilege role and attach it to those.

Jeremy: Oh right. Yeah, Yeah, that’s awesome.

Brian: So you don't have to do any of that. The other big, tricky thing in CloudFormation is getting service discovery, right? So, ideally, when you're building out your serverless application in CloudFormation, none of your resources have a human-readable name. You can give them logical IDs, which are human-readable names. But you want that generated stuff to effectively be GUIDs, and you don't want that to have any significance or meaning because we want to treat our services like cattle. We don't want to treat them like pets. So want to be able to wipe those out and recreate them and what have you. And so at runtime, this becomes a pain in the ass. If your database table is a GUID, it's really hard to find. So Architect generates a service discovery scheme for you by using SSM parameters, which are free tier. There are other ways to do this. Some people like to use environment variables, but those get out of hand with structured data, and other people like using Cloud Map, which I think has probably got a good future. But Cloud Map is also extremely expensive. It's 10 cents per resource per month, and you could rack up a lot of resources. We have thousands for Begin, so it just wasn't realistic for us. So we use SSM, which is a free key value story and has great throughput. And we can cache the results of those lookups and so you can interact with your DynamoDB table at runtime as though it had a real name. But under the hood, it does not. So those are just some of the things that we can do with that CloudFormation for you without you having to think about it. There's a ton more of minutia, especially around IAM roles and the security aspects.

Jeremy: So then if you were to build out something, you generate all this CloudFormation or the SAM template, you said you can add custom — well, there's two things you could do. One, you can add custom resources through your macros. But the other thing you can do is if you just say, you know, I don't want to use Architect anymore, I just want to take my bootstrapped SAM file, I can just eject and go my separate way.

Brian: You can bail and we actually have a playground on the website. If you go to arc.code/playground, we've got one of those two-up things on the left side. You write Architect and on the right side, we show you the generated CloudFormation template. And yeah, there was an eject path in this that we really wanted to have. I feel that increasing the CloudFormation is the standard way to build AWS applications. If the cloud had a file format, it would probably be CloudFormation. So in order for us to have interoperability with that ecosystem and also the, you know, the portability into things like SAR, we really wanted have that ability to eject and not hide it behind a leaky abstraction.

Jeremy: And so speaking of this idea of putting things into SAM, when you were using SDKs, all of the services at AWS, you can control with their APIs, right? That's like the first thing they release. CloudFormation, not so much, so are you handicapped at all by using SAM now, in some cases, or just have you not run into those limitations yet?

Brian: So one of the beauties of the ARC macro system is that it can run an after deploy. And we haven't exposed all of this yet in documentation but it's in the code and we do it ourselves. And so you can run a patch after you deploy, or you can do whatever the heck you want with those generated resources. It's kind of like CloudFormation custom resources, except for it runs locally on your machine and uses SDK calls. This seems impure and dangerous — and it is. But we had to do it. So believe it or not, AWS has bugs sometimes. And when we did this move to CloudFormation, we found some of those bugs and we were not stoked on them. We found bugs in particular with API Gateway that were pretty deal-breaking around binary content encoding. So we ended up having to write a patch that ran after deploy and did another AWS or did another API Gateway deploy, which is a bit of — it felt dirty. But it worked.

Jeremy: It does. It does. I'm right there with you on EventBridge, because that's the other thing. My latest project is using EventBridge, and you can’t create custom buses. And you can’t create, or you can’t add rules to custom buses through CloudFormation...

Brian: See, this makes me angry now. At this point, this makes me angry as a customer. I really feel Amazon is dropping the ball on this one. So CloudFormation is not a nice-to-have. The service, the team’s got to get together and have this on day zero for every one of their releases. If they expect us to be following the best practices, they publish.

Jeremy: That would be nice.

Brian: Yeah, and I feel like it's not really a disadvantage because we can drop down to the SDK at any time. And there are actually times when you maybe do want to do that. So Architect has another semi, not well-known feature called Dirty Deploys. So we'll deploy using CloudFormation by default. And by default, we deploy to a staging stack, and we have a production stack, which is an extra step to get to. You could do your own arbitrary stacks if you want, but we bake in staging/production because we feel that's essential complexity. The other thing we can do is just deploy functions using update function code calls, and this is, the syntax for invoking it is “arc deploy dirty” and it is dirty. But what we'll do is we'll literally zip all your functions and will replace the ones in staging and a Dirty Deploy usually runs for, say, 10 functions within two seconds. So your iteration speed is incredibly good on staging. When you do these Dirty Deploys better than CloudFormation even, although CloudFormation’s gotten pretty fast. It's not that slow anymore. It used to be really slow. Yeah, so sometimes you need to drop into that SDK, get a little bit dirty, and I think that's okay. But if Amazon's listening, they really got to get CloudFormation support day zero for every new product. That's table stakes these days.

Jeremy: Alright, so let's talk about microservices It sounds like we're building a single application, right? And I know that what I do with the Serverless Framework or with SAM, is I will build multiple services with separate CloudFormation stacks as microservices and so forth. Is that something we can do with Architect as well?

Brian: Yeah, it's not super well documented, but we have a way for broadcasting SNS events between stacks. The service discovery allows them to talk to each other, and this has worked really well for Begin.com. I haven't concluded exactly the best way to do this yet. So Pub/Sub is is great. And we have a lot of tools for doing it. We've got SNS, SQS, and EventBridge. It seems to me that the kind of sweet spot is actually combining these things into like where you would maybe broadcast an event, but it hits a queue, so you know you don't lose it, because the availability and processing guarantees between these things are a little bit different, and you sort of want the — you want the guarantees of SQS probably. Like that message got delivered. Maybe more than once. And yeah, so there are ways. And generally, I would say those ways are Pub/Sub. And I would also say that this is an interesting new ground to figure out. Once it gets really sophisticated, you might want to get into proto buffers or something like that, but I don't know that devs really want to get into that. You know, we could get a lot done just JSON payloads over Pub/Sub.

Jeremy: Right. Yeah, that's actually, that's why I've been big into EventBridge lately. Because I do think that, I mean, I think they're ramping it up, and I know they get billions of messages that go through there every day. So obviously it's a very reliable service, and it's something that if we can make that part of our application, especially for cross-boundary communications, I think it would be really interesting. Of course, you do have the issue with service discovery, but...

Brian: Yeah, it's related. And I feel these are like, really the bleeding edge problems of the cloud right now, are service discovery, inter-app communication. Maybe these are just always problems too.

Jeremy: Function composition. I just talked to Rowal Udell about Step Functions for function composition, where it's great for certain asynchronous workflows, but, how much coupling do you want to create across service boundaries, and are Step Functions the right choice there, and what if you need something synchronous? But I don't know, there's just — there's a lot around that I think that causes confusion, especially when you go to that single responsibility principle.

Brian: Absolutely. You know, I think it's also speaks to AWS’ maturity in the space. You know, you look at it, and it looks like they got a lot of Pub/Sub and a lot of databases, for some reason. Shit, they even have two ways to invoke HTTP events through API Gateway. But these things have different availability guarantees and different service guarantees and different limits. And you really have to, you can't just, like, abdicate thinking about it. You really got to dig in and understand: okay, is the characteristics of SQS appropriate for this use case? Do I want, you know, this thing to retry forever? Or do I want to fail at some point, like, that kind of thing. Or, like database is another really good one. I mean, there just is not going to be a database that fits all workloads. And so Amazon has a lot of database products.

Jeremy: Although Oracle says their database can handle all of those use cases. You know, I’ve heard that lately.

Brian: Yeah. I don't pay too close attention to Oracle and the cloud.

Jeremy: Yeah, that’s probably not worth it. Alright, so just one last thing on building apps. So another thing we're seeing a lot with modern applications, is a lot of front end developers hosting static sites, right? So what's the Architect solution for that?

Brian: Yeah, this one actually came at us a little bit of a blindside. So we be built out Begin in — the earliest version of it — around late 2014, 2015, 2016. And I kind of missed the rise of the static app. I was a part of the PhoneGap team, so I was close to the rise of the single page app. But I did not participate in the rise of the static app. So when Architect was first released, a lot of people struggled using it to build things because we put a function at get / and we make it greedy. And that's a lambda function that's returning HTML. Oh, my God, why would you do that? It works really great if you server render, which you probably do want to do. But if you're content is inert and it's not changing, then you know, static sites make a lot of sense. It especially makes sense for your landing page and that kind of thing. So we re-tooled Architect in version 6 so that the root of your application is a public folder, which we greedily proxy. So anything in that folder that you compile to with, you know, Gatsby or React or whatever will be available at the root of your application. And then any functions that you add will be mounted at sub URLs so you could have post GraphQL, for example, would call a Lambda function. But you know, all your static assets could just live in public. And this seems to be the architecture that people really want these days. We don't do it ourselves. For what it's worth, we server render things through Lambda functions, which at first sounds disturbing and slow to people. But you have to remember that we put these things behind the CDN. So your Lambda function’s not getting invoked a lot. It's getting invoked maybe once a day or something like that.

Jeremy: Yeah, because you can send — because actually, that's one of the things I think maybe people don't know is that API Gateway has a CloudFront distribution in front of it, and so if you send back the right caching headers, then it will cache that for you and not hit the backend.

Brian: Yep, and CloudFront’s a great solution. You know, if we actually do the regional API thing and then we put a CloudFront just in front of the, you know, the blanket URL and get rid of that ugly /staging or /production that you get. Yeah, it works great. CloudFront takes standard caching headers. Even better if you're putting your stuff in S3 and you upload your content using our deploy. We’ll set all of the headers and content type for you on the S3 buckets so the etags comeback correctly, and you get a cache for free basically, and you know, you're going to see sub-10 millisecond responses on the majority of your content. You know, once in a while, you'll get a cold start and it'll be 200 milliseconds. It’s like, fine.

Jeremy: Still pretty good. Alright, so let's talk about some of these primitives, cause this is again, this is one of things about the Architect Framework that I really, really like. That you have, I think, what, 12 different primitive services and 12 different primitives, maybe? Maybe you can explain it better than I can. You built it.

Brian: I'm one of the people that built it. I didn't build it exclusively. There's actually quite a few people hacking on it nowadays. And I have to go to the website just to remember myself. So when we first started, you know, we were building CRUD-y web apps, and I remember we actually had a piece of graph paper out, and we were figuring out, you know, what we needed to build. And then what services that Amazon facilitated those. And so we've ballooned to 12 services, but we only used to be eight. The services are Lambda, API Gateway, S3, SNS, SQS, DynamoDB and CloudWatch events. And then there's some sort of supporting cast, services that we use that you don't really see and don't really come into play very often, but CloudFormation, obviously, Route 53, CloudFront, Parameter Store, and IAM kind of set up the supporting side, and that's it. That is the core of Architect. We just paper over those services and make it really easy to build a web app. We kind of felt that those, you know, facilitated CRUD. They facilitate background tasks that could be long running. You can put a CDN in front of this stuff. You can host static assets. You know, this is basically everything that you would need to build the majority of a web-based app today.

Jeremy: Yeah, that makes sense.

Brian: It isn't to say that you don't have those other use cases, and that's why we dump out the CloudFormation for you and let you modify it, because you will have other use cases. We've got a few macros floating around out there already. The most complex one is for uploading directly to S3. So directly uploading to S3 sounds like it should be an easy thing to do, but it turns out it takes a fair amount of CloudFormation to pull off. And so we've wrapped all that up and made it a single line directive inside of ARC. And it's a good example of not something that we built it for initially, but we were able to extend it into.

Jeremy: Nice. All right, so what about limits in AWS? Because that's one of things about AWS, which is great, is they do publish all their limits and people are like, “Oh, well this has a limit.” Yeah, well, everything has limits, but it's nice to know what those are. So does Arc sort of deal with those gracefully?

Brian: Yeah, and I feel this is worth a shout-out because this is something the other clouds don't do well. They sort of claim that they'll handle it all, and they leave it up to you to discover where their failure rates happen, when you're gonna get throttled, when you're going to overwhelm them and that kind of thing. And that's a bad experience. I want to know what the service boundaries are, so I can design my application for them. Though, the big one that you run into when you start working with CloudFormation is the resource limit of 200. We do some tricky nesting to get around that limit. Otherwise, I kind of don't feel the limits are the limiting factor anymore. There was a time when it felt like, you know, Lambda needed a lot more memory, and Lambda is going to be better once we get X. But that time has passed. We have tons of memory. The execution limits are pretty generous. And I haven't run into problems with that. A while back, I actually heard someone concerned about DynamoDB limits, which I thought was pretty laughable because DynamoDB has an extremely large amount of potential theoretical throughput, but potentially infinite storage, uh, guaranteeing single digit millisecond latency queries, like these are characteristics I have never seen in a database. I don't think there is another database with these characteristics, so yeah, it's got limits, and they publish them. But I view this as a positive and not a negative. And it helps you build a better app. You know what you're in for.

Jeremy: All right. So let's talk about DynamoDB. Because I know you mentioned that earlier. I think you gave a — I think I heard you give a talk about DynamoDB one time. So DynamoDB is sort of woven into Architect. It’s sort of the database of choice or database, you know the default database I guess you would say within the Architect Framework. Sowhy DynamoDB? Maybe let's start with that.

Brian: Yeah, I mean, it's a decision making process. And it's one that a lot of people are aren't comfortable with. It’s a managed database, which is a nice way of saying that it's a proprietary database. It's owned and run privately by Amazon. And, you know, after, our history has a, or our industry has a long history of being gun-shy of these databases because of Oracle, frankly. And I don't blame anyone for painting Amazon with that brush. “Oh, my database. That's my data. I don't want them to have that. I want to control it.” The only people that say that, by the way, are people that have never sharded a database, you've sharded a database once, you are happy to let someone else manage that for you. You are more than happy. How much does it cost? Fine. Less than a DBA. So that's going to be a good deal for me. So once you get over that initial concern, which isn't a real concern, by the way, that free tier is extremely generous. You could run a local instance of this thing yourself headlessly if you want for testing and building out locally, so you don't have this requirement of the cloud. And the free tier’s insane. I think you get something like 20GB in the free tiers. So, like you could build a lot of app with 20GB. A lot of app. You could put images in there, you don't want to, but you could. Yeah, it's a great DB. I guess the other thing that people get a little tripped up on is the syntaxes. It’s a bit strange. It’s coming out from a different world. I don't think it actually is that strange, for what it's worth. I'm pretty sure if you'd never seen SQL before and I showed it to you, you’d be like, Well, that's strange. I think just what you're used to. It's a sadly verbose query language. It takes a lot of directives in JSON form to make it do pretty trivial things. We've written a few higher level wrappers for it to make it a bit nicer to work with, but it's all about the semantics. Single digit millisecond latencies for up to a MB at a time querying, no matter how many rows I have? That's unreal. But we've never had a database that can do that. And I'm happy to pay for that capability.

Jeremy: And I think you mentioned a good point about the query language and more so about how you go about getting data in and out of DynamoDB. Because getting data in is fairly simple. You just kind of put items. But it's not always just put item, right? Sometimes it's update item, and sometimes you need to use an update item to put an item, depending on what you're doing and you want to minimize lookups. Upserting exactly. And you have to use these tricks like if_not_exists on the created date, if you don’t want that to get updated if you're overwriting a particular item and you know, there is, there's some interesting syntax, obviously the “begins_with” to select a part of the sort key. You know, there's a lot of really cool things you can do, but it’s certainly not very straightforward, in my opinion. I mean, I come from a SQL background. I’ve done SQL for 20-some-odd years. So for me, SQL is super easy, right? With DynamoDB I’m always thinking about a different way to access it. What I have to be worried about overwriting different records and then how efficient you need to be when you're grabbing data and that you really can only grab data with the primary key or a GSI, and then add composite keys to efficiently filter on that sort key, so you don't have to add filters after that. So I do think there's a big learning curve there. I've talked about this on the podcast a million times, but at the end of the day, I don't think I would want to go back to SQL, especially for most of the use cases that I have, I can use a DynamoDB table to handle that workload.

Brian: Yeah. Me too. And I'm sure there are, and actually, I want to give some props to Erica Windisch, about opening my mind on this one a while back. I was sort of kvetching about the cost of Dynamo, and we had a lot of rows for not very changing data. And they correctly were like, “Well, why don't you just dump that into S3 and use S3 Select?” I was like, lightbulb went off. I was like, “Wait a second. Can I do that?” Yeah, you can, and you can get crazy good querying speeds out of S3 and S3 Select. And this is sort of the beauty of the serverless managed database cloud in that, it's no longer a tradeoff. We don't have to put it just in Dynamo, or just in S3 or just in RDS. Why not all three? We can use Dynamo streams to pump all that data into Redshift. If we really want to use SQL to query it. If the data is historic and not changing very often, maybe just dump it in S3 and read it with batch and Select.

Jeremy: Or Athena.

Brian: Yeah!

Jeremy: It’s great for data that doesn't change. Any time series data, Athena is amazing.

Brian: Yeah. So I kind of think where we're at now is less about making a tradeoff choice and more about what are we opting into, for what characteristics, and when? So if the — you're absolutely right. I don't want to trivialize querying in Dynamo. It's a real — it takes a minute, and it's not the easiest to model, and you're probably going to get it wrong the first time. And that's totally okay, because your iteration speed is already 100 times better than it was before, So you're going to be able to fix it. And my recommendation to people is to just dive in, you know, maybe model it a little bit relational and feel that pain and start learning about wide column design. A neat thing about this key value store thing is that your skills with Dynamo are transferrable to Mongo and Cassandra and these other key value stores. They all model roughly the same way, where you start with the query and build out your columns from the query. So yeah, I get it. I talk to friends about this all the time, and they're like, “No, I want to use RDS,” and I'm like, “Yeah, I know.”

Jeremy: A lot of people still do. I mean, that's the other thing, too, is that a lot of companies still do the whole sharding thing, right? And certain companies — sharding is not that difficult, if you have a very good key that you can use to shard on, right? So like Slack, that's easy. It's a work space or whatever it is that they can easily just shard on. But when you build applications that there's a lot of intercommunication between them, you're going to be doing a bunch of denormalization in relational databases. So why not stick into Dynamo, do the denormalization and get the speed benefits without that overhead of doing all that sharding?

Brian: Yeah, or do both, you know, if you really do need it. I think there is a data gravity thing here too. There's a lot of not just skill investment, but actual literal rows in a database sitting there right now. And if you're in BIG CO and you've been around forever and you've got, you know, GB of data in SQL, you're not going stream that into a Dynamo table. This is going to have to be a different story for how that thing gets migrated and/or how those apps evolved into the serverless world. It's not a zero sum game. But I think if you're using Lambda and you want to play on easy mode and get the best performance characteristics, Dynamo's a no brainer.

Jeremy: So what about lock-in, right? I don't know if we mentioned this earlier, but just this idea of lockin it’s talked about all the time. Obviously, you said your skills are transferrable, but not all of your data or code might be as transferrable with Dynamo. So what's your thought on that?

Brian: Yeah, the lock in discussion. I like to dig into it when people bring it up, so, you know, there are concerns with lock and one of the concert primary concerns should be price. This is the lock-in a lot of people suffered with Oracle, where they squeeze you as the years go on and your data becomes harder to move. Amazon doesn't really raise prices. I haven't seen or heard of an instance where they do that historically last 10 years. So maybe that'll happen, but I'm not betting on it. It's a pretty competitive market, and they're really interested in margins. We know that Jeff Bezos always says your margin’s my opportunity. So I don't see database getting more expensive because I do feel this is the main anchor differentiation between clouds and right now, Dynamo’s in a really good position, so it's a little bit expensive. But as spanner and cosmos get better, where they're going to start competing on price, which Amazon is more than happy to do, so I expect price to go down. That's not really a lock-in concern. Another lock-in concern is they shut the service down. Well, Amazon’s still running SimpleDB.

Jeremy: Right.

Brian: So, if there's anyone on that, they don't shut things down. That's what Google does. So I'm not worried about Amazon shutting it down, so the next lock-in concern would be breaking changes. To be honest with you, I kind of wish Amazon would do some breaking changes once in a whil3, but they don't. And if you want evidence of that, go look at the S3 API. They literally have API methods that have V2 in the name of the method. Amazon only does additive change, so you're not going to suffer a breaking change. You're not going to suffer a service shutdown. You're not going to suffer price pumping, so I don't know what the objection is to lock in. Sure there's got to be another one. I'm sure someone's going to cook one up, but it's just not a rigorous argument. And for my time, the danger is picking the non-Amazon that goes away. So if my solution to lock in is to use a venture-backed third-party vendor that's privately held, then I have done some very poor risk analysis because we all know how that story goes. A privately-held venture-backed company is looking for an exit.

Jeremy: And probably and hiring somebody specifically to deal with that different type of technology or whatever, right? That's that DBA concern. It's just, you know, even if they, Amazon, did raise their prices, it’s probably going to be a lot cheaper than paying a bunch of DBAs to keep re-balancing the MongoDB cluster or the Cassandra rings or whatever, right? I mean it’s just, anyway, I totally agree with you on that. Alright, so we've been talking for a very long time. We probably could keep talking for a while, but I do want to get to Begin.com, because this is a super interesting thing.

Brian: We should do that.

Jeremy: It has been in private beta for quite some time, but you've got some news, right?

Brian: I have some news. I'm stoked to announce that Begin is now publicly available. Anyone can try it out. We have a free tier where you can deploy in an app serverlessly to AWS in 30 seconds or less. Usually takes around 10 seconds, but we have an internal benchmark of 30 seconds. It’s CI/CD, but serverlessly. So CI/CD is not news.They've been around for — it's been around forever, and there's tons of people that do it, but most of them are for traditional architectures and haven't really taken advantage of this serverless world. And so Begin is a fully ground-up cloud native serverless deployment service in CI/CD service. And as a result, our build times usually, at most, expand into a minute. Usually they're around 30 seconds. That's great lead time to production. Lead time to production is the main metric by which companies live and die and we give you an extreme advantage to that. And it's all just Architect. So you can eject any time and run on your own end of AWS. Our paid tier will target your AWS. Our free tier is running on our AWS, because we found a lot of devs and I think this is an interesting thing. But there’s a huge Amazon community, obviously, it was a huge amount of people building for the cloud. But I talked to a lot of newer devs and they're really intimidated by AWS. They don't know how to get started and they don't know where to get started, and they find it to be just so overwhelming. And it's just way easier to get started with with someone else and Begin is seeking to fix that problem. They shouldn't have to get started with someone else. They should be able to deploy straight to their Amazon within 30 seconds. That's our goal.

Jeremy: That's awesome. Alright, well, listen, Brian, thank you so much for taking the time to talk to me, sharing all of your knowledge with the community, all the open source stuff that you've done. So how can listeners find out more about you, Architect, Begin, all that stuff?

Brian: Yeah, Begin.com. It's open. So go log in with your GitHub. If you want to find me, @brianleroux on Twitter and GitHub and I usually respond pretty quickly. And if you want to learn more about Architect, you can go to arc.codes.

Jeremy: Awesome. Alright, I will get all of that into the show notes. Thanks again, Brian.

Brian: Thanks, man.

View Details

About Rowan Udell:

Rowan Udell is Cloud Practice Director at Versent, an AWS Premier Consulting Partner in the Asia Pacific region. Working with customer and internal Versent teams, he helps them deliver change at scale and speed using serverless and AWS native services. He co-authored the AWS Administration Cookbook and has published video courses on AWS.

  • Twitter: @elrowan
  • Blog: blog.rowanudell.com
  • AWS APN Ambassador: https://aws.amazon.com/partners/ambassadors/ambassador-apac/
  • AWS Administration Cookbook: https://www.packtpub.com/virtualization-and-cloud/aws-administration-cookbook (2nd edition coming out soon)

Transcript:
Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Rowan Udell. Hi, Rowan. Thanks for joining me.

Rowan: Hey, Jeremy. Thanks for having me.

Jeremy: So you are the technical director at Versent which is in Sydney, Australia. And you are also an AWS APN Ambassador. So why don't you tell the listeners a little bit about yourself, what Versent does, and actually, I'm kind of interested in this AWS APN ambassador thing, if you could tell us about that as well.

Rowan: Yeah, sure. So Versent is a premier consulting partner here in Australia and, you know, we work with a lot of enterprise customers, really helping them do cloud the right way. You know, if I was to describe us to a wider audience, we kind of want to see ourselves as the heart specialists for AWS. You know, when you have a heart problem, you go to see a specialist. You don't just go to any old doctor and we want to be that for AWS. In my role as technical director, I work with a lot of teams that are using Step Functions and other serverless technologies, building out applications on AWS. I'm part of the APN Ambassador Network, which is a new program that AWS has started up for consulting partners. APN Stands for the AWS Partner Network, and what they've done is get together a group of like-minded partners in one room so that we can kind of give feedback to AWS but also help them get new features, services, technologies out into a wider audience, you know. And so a big part of what we do is things like coming on this podcast, but also doing blogs, doing meet ups, giving speaking events and things like that. And so they really kind of encourage us and enable us to get the word out there and try and make AWS easier to use for everybody, not just consulting partners.

Jeremy: Great. And what about your background? Where do you come from?

Rowan: Yeah, So, I mean, I've been working in IT for longer than I'd like to admit now, you know, mainly with AWS, especially over my last four years here at Versent, just working purely with AWS. Before that, I was leading developer teams, working at some startups, things like that. And, you know, when the cloud came along, I really kind of jumped on board because I was sick of administering servers in the first place. And that's probably another reason why you see me online talking a lot about serverless stuff.

Jeremy: Awesome. Alright, so I have had a number of episodes where we've gone down sort of the technical path of some subjects. Some of them we've talked more about the business things, but I know that you do a lot of work with Step Functions, both as obviously as your role as a consultant, but also just I think just what you're doing out there working with teams and what you're seeing. So I want to talk about Step Functions today, AWS Step Functions. I think this is a really fascinating service that they offer. I think it's under utilized by a lot of people. I mean, it certainly has good use cases and use cases that may not be the best for it. But maybe so if people don't know, what are AWS Step Functions?

Rowan: Yes, so AWS Step Functions is the name of Amazon's state machine as a service offering that they have. State machines are sometimes called finite state machines, and this just makes them much more easily consumed and also integrated with your serverless applications.

Jeremy: All right, so let's let's start high level here. What are state machines?

Rowan: Yeah, so a lot of people get turned off by some of the terminology. But I mean, really, it's just a lingo or, you know, some vocabulary that's actually been around for quite a while. State machines are nothing new. State machine’s really just a mathematical way of modeling an application. Most people kind of understand what that means conceptually, but what it means in reality is you can describe your application as a mixture of your inputs, the states that your application can be in, and the transitions between those states. And what this forces an application designer or really a developer to do is to think really concisely and clearly about what their application does and what it could do next, and it kind of forces you to do that upfront, which I think is something that's really valuable and often overlooked.

Jeremy: Yeah, and one of things that I really like about state machines and specifically Step Functions in the serverless world is that ability to do function composition. Because I think that's one of the things that many people are confused about. Like how do I have function X talk to function Y, right? So state machines are the glue in between those, right?

Rowan: Yeah, definitely. And you bring up a really good point because we see a lot of people out there discussing on forums and in Slacks about like, “Should I have one function call another function directly?” and usually someone will jump in and say, “Oh no, you should never do that.” But obviously, then that's going to complicate things a little bit. And in some ways, if you're using Step Functions that problem goes away because you're able to link two functions together, you know, which allows each one of those to do one thing well, and you don't have to worry about calling them directly and coupling those two functions to each other. At the same time, you don't have to introduce things like SNS topics or SQS queues in between those functions. You know, in my opinion, it's kind of the best of both worlds.

Jeremy: Yeah, I mean, and that's where we're talking about orchestration versus choreography, right? So and that's one of the things that if you're using SNS or you’re using EventBridge or using some other communication channel or messaging bus, you can decouple the applications — state functions are a way of sort of creating coupling. But it's a different type of coupling, right? Because a function can run on its own outside of Step Functions or it can be part of several Step Functions, right? Different steps within that could reuse that same piece of logic. It's just a way of kind of gluing all that stuff together, like you said.

Rowan: Yeah, it is a form of coupling, but it's a loose coupling. You know, the function that is calling or being called doesn't know that it's being called by another function, to your point where it might be part of a complicated workflow or it might not. It really doesn't matter to the implementation function.

Jeremy: So what about the visualization of these things, too? That's another sort of important piece, right?

Rowan: Yeah, when it comes to the visualization, I think that's another thing that Step Functions provides that's really valuable, especially for developers getting started with state machines that maybe don't have a lot of experience with it, is it has a really nice way of rendering your state machines so that you can clearly understand how states are connected, what transitions are in play. And you know, this really makes things like troubleshooting and designing a lot easier. And it's this kind of thing, which, if you're trying to use state machines in your applications as I did, years ago, you had to do that yourself. There was no easy way. So generally what you found yourself doing was okay. I'll go over here and I'll design my state machine in isolation. I might use a diagramming tool, and I'll draw the circles and the lines that connect them, and then I'll go and implement that in code. But there's a little bit of a leap there. Well, I might think I've implemented it in code, but maybe I didn't. Maybe I got something wrong, and that's going to make my life later on a lot harder, because what I think it is and what it actually is is two completely different things. You know, there's a quote I love, which is, at the end of the day, state machines are a way of modeling your application, and we often say all models are broken, but some models are useless. Step Functions make it easier to get that model more, right. I wouldn't say correct, but it gets it a lot closer.

Jeremy: Yeah, and those are certainly some of the benefits that you see — just that ability to troubleshoot faster, like you said, and sort of encapsulating that application logic. Are there any other benefits though that sort of come to mind?

Rowan: Yeah. Look, I think with Step Functions, it really forces the developers to break up their applications into discrete steps — things that most developers already do in their head, it forces them to articulate it and write it down. And that obviously makes it easier for them to communicate that with another developer, and that other developer might be them in six months time. And so, by forcing them to break these things down, you can see where there's a lot of parallels with serverless applications in general. We talked about having functions do one thing and one thing well, and at the end of day, I guess what this is forcing developers to do is make all of their implicit models that they have in their head and really make them explicit and define them and say, “Well, you can only do this thing after that thing” or something like that.

Jeremy: And so the other thing, too, is I think people look at Step Functions and they think that they're extremely complex. But really, they're not actually that complex. I mean, they're powerful, but from a complexity standpoint, they're quite easy to implement actually.

Rowan: Yeah, if you look at most Step Functions in the console and you can see that visualization of them, they are really quite simple. There's a built-in limitation to them that I think makes them quite powerful because it really forces you to do what you said you're going to do and not kind of make things up on the fly because it just won't work. And I think a lot of people are turned off this perceived complexity, partly because of the vocabulary that's used. Like I said, traditionally, they've been very academic and they've been around for a long time. So people kind of look at that and go, “It's computer science. I don't have time for that,” or something like that. And really, it's not that bad. And like most things, if you start off simple, you have some simple use cases in mind, you can then iterate on that and build up to high levels of complexity, and there's been some nice features released that can really help you with that. At the end of the day, I wouldn't say Step Functions are required for every serverless application out there, but they really do suit themselves well to complicated workflows things that need a higher degree of traceability, or auditability. For some of those really important workloads, they can really help you, give you confidence in the system.

Jeremy: Right. Alright. So let's get into a little bit more details here. So you mentioned input states, transitions, that sort of stuff. So let's talk about sort of the state's tasks, activities, those sort of things as part of the, well, I guess what we would call maybe the core functionality of Step Functions. So why don't we start with that?

Rowan: Yeah, sure. Yeah.

Jeremy: So if you give us an overview of that, that'd be great.

Rowan: Yeah. So, look, most of what I'm going to tell you about today is defined in the state’s language spec. I think it's states-language.net/spec.html. This is a document produced by AWS, but if you actually read it, there's nothing specific to AWS in there other than the name of it. And really, when you talk about the core functionality of Step Functions, you're going to talk about an execution of your state machine and an execution has an input, and that's just a JSON payload that comes into it. And then once you're inside that execution, your state machine is going to have a number of different states, and those states can transition to other states as you've defined in your state machine definition, which at the end of the day is just a JSON object. The most interesting part of the state machine is going to be those states that you can have and in Step Functions terms, we talk about these different kinds of states that you can have that you can define using the language. The main one is called a task state, and as the name suggests, this is where you go off and do something. And then there's a few other kinds of states that you use to kind of control how your Step Function flows. So you have a choice states that can direct you down to different following states.

Jeremy: And those choice states can all be done — those are all based off of the output from the previous state.

Rowan: Correct, correct, because every state has an input on its own that it uses to then do something with and again, this is kind of the attraction for me with state machines, is that you know what it's gonna do based on its inputs, and it forces you to kind of define that up front.

Jeremy: Right.

Rowan: Some of the other states that you can use — that you will have to use — is the end states as we talked about, and these are special because you can't go to another state after them. And you have two kinds, as you might if you think it through. You’ll realize you have a fail state or a succeed state, and this is just about reporting the status of that execution back up to Step Functions so you can see it in the console, wherever you're monitoring it.

Jeremy: Okay, makes sense.

Rowan: Another state that you can have is what's called a pass state. And this is a relatively simple state that you use to either inject certain values into your input and output payloads, or often you'll use it for debugging just to try and, you know, kind of explain, okay, why are we transitioning from this state to this state and what is that being done for? Another state that's often used is called the wait state, and this is where you might build in a delay into your system and again, because you're forced to declare everything, you actually have to tell whoever's looking at your execution, “Yep. I'm just waiting for something now.” You know, it's not kind of hidden away inside the code.

Jeremy: But actually, the wait state is is incredibly powerful because there have been several people, I think Paul Swail and Yan Cui and a couple others, have used that wait state as a way to set a like a timer — like a dynamic timer, basically — you know, so that if you want to execute, I know some email or you’re scheduling emails or something like that, you could create all of these different wait states in order to schedule a job later in the future, and it's much more reliable than something like a DynamoDB TTL or something like that.

Rowan: Totally. Yeah, and you know, this is a really good example of how a state machine is made up of these relatively simple states, but when you combine them together in this way, you can come up with some really complicated workflows that are really powerful, yet still easy to understand.

Jeremy: And I actually think that's one of — speaking of combining things — the other really cool thing is obviously parallel branches, right? So what can you do with that?

Rowan: Yeah, so look, this is another one that I find myself using a lot where, just as you would if, you know, you as a human were performing a task, you might have a couple of different things going at once and again, this is where you can really leverage some of the benefits of a serverless platform and the serverless approach, what makes it really easy to parallelize your workloads and really get some time and efficiency, time savings and some efficiency, out of it.

Jeremy: And what about this new thing “dynamic parallelism” that has just come out. Any thoughts on that?

Rowan: Yeah. Look it's a slightly more advanced feature. You know, it's been out for a couple of weeks now and we're using it. We've run into a few issues using it at scale, so it is still relatively new, but it does kind of fit nicely with how I think a lot of users really wanted to use Step Functions. I think I saw a tweet the other day saying how this was like the most requested feature for Step Functions. So it's really cool to see them adding that to the language and really kind of putting these additional advanced features in.

Jeremy: Cool. Alright, so let's talk about tasks because I think of all these other things — choice states and pass states and all those other things that you can use — really where the work gets done is in these individual tasks, right? So probably, I mean, I even when I think of it, I think 99% of the tasks I use when I use Step Functions are Lambda functions, right? Basically, I'm calling some discrete piece of business logic and doing that. So that's obviously the most common. But there are a lot of other services you can integrate with, right?

Rowan: Yeah, definitely. You know, I think you're right. Lambda is by far the most used. Obviously you can do anything in there within a 15 minute time limit, but there's a lot of other ones now that are really useful and especially some of these more advanced callback patterns coming in, they're seeing more and more use. The ones I use a lot are things like the SNS task, which allows you to send off notifications, and this could be great for things like Slack integration, kind of letting people know whereabouts in the state machine the execution is. Another good example is using Fargate or ECS to do some longer-running container base tasks. And probably my favorite one, which has just been announced relatively recently, is actually calling out to other Step Functions. This is a really nice model where it allows you to kind of encapsulate some complicated workflows in another state machine and from the parent state machine, you don't need to worry about those details. Before this feature was released, you used to kind of have these nested state machines, and they did look pretty complicated when you kind of opened them up. Now you can kind of collapse that all down to just a single state and just say, “You know what? Go off and do that job. Don't worry about telling me the details. Just let me know when you're finished.”

Jeremy: Yeah, that's a very powerful, and, of course, recursion in code, especially if you can write a recursive function, they can be very, very powerful, sometimes very complex, complicated, difficult to understand, but when you make it work, I know I always get excited when I'm like, “Oh, my goodness, it actually works.”

Rowan: Yeah, not without its challenges.

Jeremy: Exactly. So what about some of the long-running activities? You mentioned the callback pattern stuff, but that's relatively new. So what are we talking about with just long-running activities?

Rowan: Yeah. So in the past, you used to have to use something that Step Functions referred to as “activities,” and this was a way to kind of farm out that long-running work to workers that were usually going to be based on an EC2 instance, or maybe even in the external system. And, you know that involved a lot of polling. You know, those workers would poll for work and eventually work would appear in that queue. And I think this pattern is less relevant now than it used to be. You know, it used to be the only way to do this. But now with this callback pattern where you send off the work and you supply a task token and you say, “Hey, when you're done, let me know that this task is finished and then I'll continue on the execution.” And that doesn't involve any kind of polling or any kind of long running work. You can really get it down to the point where your workers only work when they need to, not just in case.

Jeremy: Yeah. I mean, and that's the other thing, too. Ben Kehoe had a great post about the task tokens because that's just one of those things where it was so inefficient to be polling every 30 seconds or every 10 seconds. And if it was an important job, you might be pulling every one second in order to check for something to be done. So anyway, I really do like that — that new callback pattern. It obviously is much more efficient. And what is it, like a year or something like that? How long can it wait?

Rowan: Yeah, so executions can sit there for up to a year, you know? And they don't cost you anything while you're doing that. Their pricing model doesn't care how long it's been running for. So yeah, really powerful for some of those really long-running workloads.

Jeremy: Yeah. I hope you don't have a task that takes a year to run, but if you did, you'd be okay. Maybe that sounds a little complex as we kind of talked through those things. I think, though, if you look at it and you go into the Step Functions console even and you just use the little visual builder and you can build some Step Functions yourself, I think these sort of makes sense. But let's go to the advanced side of things, because one of things we know, and I think I've said this 1000 times on this podcast already, but everything fails all the time, right? So something is going to break. It's not going to be your fault. It’s going to be a network issue. It's going to be can't connect a third-party API. It's going to be SQS hiccups or something like that, and it doesn't submit the job, and then something fails. So one of the really, really cool things, especially about complex workflows that Step Functions handle is the ability to do error handling and it’s sort of built-in for you. So let's talk about that a bit.

Rowan: Yeah. So the really cool thing about when you're configuring your tasks, you can configure things like, okay, how will this behave on a failure? And obviously the simplest thing is just fail and kind of cancel out the execution. But we're seeing more and more teams, if the response to a failure is something that's kind of predetermined, like oh, you know, I should try again. You can configure these things in your state machines really easily now, and it's a lot easier than if you had to write all that code yourself. Because if you've ever tried to do kind of elaborate error handling inside a Lambda function, it can get a little bit tricky as to what happens when and it often ends up being a lot larger than the actual code that's doing the work. And so this allows you to kind of take that out. And you can even use things like the choice state that we talked about before to say “OK, on a particular kind of failure, I'm going to go off and do this other stream of work and kind of address that problem.” And again, you don't have to put that in the Lambda task that's actually receiving the error. You can kind of hide that away from there, and that enables you to keep your functions nice and small.

Jeremy: Yes, so also, you have this idea of things like the saga pattern where you have, you know, maybe a job completes or some step complete successfully, and then it goes on to the next step. And then that step fails, right? Like you can't charge a credit card or something happens, and it has to go in reverse other states and that's all stuff that you can build in with Step Functions as well, and it gets complex and it gets complicated. But that is possible.

Rowan: Yeah, definitely. And as much as I always prefer a simpler solution, at the end of the day, sometimes you have to do those complicated ones because that's where the value is and, you know, being able to very clearly map out, okay, I've done these steps and it's failed, so now I know which steps I need to undo because I've just defined them in my state machine. So it becomes really clear and easy to troubleshoot. And again, if you have any failures around there, there’s some nice integrations in the console, for example, to see okay, if you have a Lambda task and it's failed, it will actually link you directly to the function that failed. It'll actually pull the exception out of the Lambda execution result and show that to you in your Step Functions console, and it'll even give you a link to the Lambda logs for that particular execution. So, as much as I don't want to be using the console regularly, if I am troubleshooting, it is relatively streamlined process compared to, like you said in those more complicated workflows, trying to correlate across different Lambda functions exactly what went wrong, where and why. It could really help for those kind of things.

Jeremy: Yeah, definitely. Alright. So we talked about the parallel state a little bit, and we didn't really get into too much details in terms of what you would do with parallelism. So what's an example of that? Because if you're thinking about Step Functions that this step, then that step, then that step and you’re passing data through those different steps. But what's really cool about when you run parallel jobs is that one, obviously, you can use it for things like fan out and some more complex sort of work flows like that, but Step Functions also aggregate all those results for you, right?

Rowan: Yeah, definitely, and it can work really well if you have different jobs that can be performed in parallel, but they might take different amounts of time to finish, and you can bring all those results back and then make some decision based on the aggregate results of those. And you might even say, “Hey, I can accept some failures in this process and I don't have to fail the whole thing.” And when you combine that with this new dynamic parallelism, it can, with the ability to nest Step Functions, you’re going to end up in a situation where you can have a really simple, kind of overreaching Step Function that can call out to lots of other Step Functions in parallel. And the potential is there for processing data, and things like that are going to be really, really cool once people get comfortable with these. You know, I didn't mention it before, but there's also integrations for tasks to call out to things like SageMaker and Glue and some of these really data-intensive services from AWS. And so being able to pull together these complicated ETL workflows is actually going to be quite simple to implement and maintain going forward, again, once people get comfortable with things like Step Functions.

Jeremy: Yeah, so then another sort of advanced thing, and maybe it's not that advanced, but I think it's a pretty cool is that when you are sending data into a particular Lambda function or into a task for whatever, you can actually manipulate the shape of that data. So let's talk a little bit about how that works.

Rowan: Yeah, sure. So this is another one of those features which, when you get started out with Step Functions, you totally don't need to worry about, you could just have input from into your state, then become some kind of output, whatever your land of function returns, and then that becomes the input for your next state. And that's a really simple model. And it works really well for most workflows. You do have to worry about the kind of total size. You don't want to put large objects in there. You know, that's where you might put a reference to something in DynamoDB or S3. But once you get to the stages where you may be doing things in a more complicated fashion, you can pick and choose what parts of the returned payload from your Lambda function or whatever service you're calling in your task you actually want to keep, and then you might only pass a subset of that on to the next state. Whatever it needs to do its job, and that way you can, using this kind of JSON path syntax, you start at a dollar sign that represents the root of your JSON object, you can just say “Yep. I want this particular property that will get sent to the next state.” Or you can even do it on the input passing as well to say, “Well, I'm going to give you a large object. But, you know, since I'm running a whole lot of things in parallel, you don't need the whole object. I'm just gonna give you a subset,” and then you can do what you need to do on that. And as you said, you can then reaggregate that at the end of those particular states.

Jeremy: Yeah, and this is super important when you're integrating with some of the other services, right? So if you want to send something to DynamoDB, you really want to be able to control the shape of that, or you need to interface with Glue or something like that. Step Functions will make the API call for you, but you still are responsible for the shape of that data.

Rowan: Yeah, another good example is when you're doing those notifications via the SNS integration, I only want to pass particular subsets of my input object to, say, the message field or the body of that notification message, so they don't need to see all the internals. I can actually send them what is relevant just to them. And that's really nice, because again, all of that complexity is hidden from the service that you're actually calling. And it's really in the place that needs to be, which is the state machine, which knows about all these things. That is kind of the core of your application.

Jeremy: And the other thing, too, so combining the data afterwards, so when you execute maybe a couple of parallel tasks, maybe you shape the data differently in terms of what goes into each one of those tasks, when that data comes back, you not only have access to what those return, but you also have access to the entire state of the application, right?

Rowan: Yeah, yeah, it gets pretty complicated, you know, when to trying to bring these bits and pieces all back together into a coherent message. But the reality is that complexity was already there. You just were kind of ignoring it and hoping for the best. This kind of forces you to think about it right at the beginning and go, “Okay, Well, what will I do with the various returned values?” And really kind of makes you do it now rather than after It all goes horribly wrong. And you have to, you know, pick the pieces up.

Jeremy: Definitely. All right, All right. So let's move on to some recommendations here. So you work with Step Functions all the time. So what are some of your best recommendations for people using Step Functions?

Rowan: Yeah, sure. One of the things I recommend to the teams I work with that a using Step Functions is to take the time at the initial state of their execution and really setting up their payload, you know? So this might be where you pull in variables from parameter store or something from databases, and then use that in your payload for all of the states that follow in that particular Step Function. And the reason why you do this is that it stops all of those other states making their own calls to those configuration services, all those data services. Now you know you can't obviously put things like secrets in there, but you can put references to where these things are so that you kind of minimize the touch points or coupling for those states that follow afterwards and everything you need to know. It's kind of like dependency definitions. You know, you do very much of the start of your execution, and that way it's there for all the other states to use. The other thing I really like about that approach is that kind of mimics, you know what you do in things like functional programming where you say, “Okay, the behavior of my code is determined by the inputs.” And so you're really calling out these are my inputs, and I find that makes it easier to test the various states in your Step Function, you know? So if you have a Lambda function in there and it's always gonna be getting the same inputs and everything it needs is in those inputs rather than in things like, you know, environment variables or external data stores, then it makes testing those functions in isolation a lot easier. So, you know they're gonna work in the context of the state machine that it will eventually exist in. As I mentioned earlier, you know, I think it's really important to bubble up errors into the state machine as much as possible. So in the case of Lambda, this is where you know you don't want to put too much complex error handling into your Lambda. Obviously, some make sense. But at the end of the day, if the function has a problem, it should just let the state machine know that “Hey, I've got a problem.” And that's gonna make trouble shooting quicker and easier because you gotta find the call problem in a short amount of time. The other thing, which I know I don't do enough off and I see people kind of learning the hard way is setting sensible timeouts on your states, you know, so particularly for these kind of callback based tasks, if there's only a reasonable amount of time that it should take to run and it could be hours, days, whatever, then you should call that out in the actual state definition. You know what I generally sees it in development, it's fine because you're watching every single execution. And if something goes wrong, you see it and you fix it. Once you've deployed this and into production and you're not looking at every single execution or maybe you can't even see it every single execution, that's where timeouts will really save you. You know and kind of let people know that, “Hey, there's something going wrong here.” Another thing that's really cool about breaking your application up into these discrete states is especially in the context of Lambda tasks, it allows you to set an IAM role just for that specific task. So rather than saying, “Hey, this is a role that my application needs to run.” You can say, “well, this particular part of my application only requires these kinds of actions, and resources on these resources”, and it allows you to define that they're rather than giving, you know, that kind of a common set of permissions to your entire application. So this is really kind of the best possible least privilege scenario that I think you can get.

Jeremy: Yeah, I'm a huge fan of least privilege. That's my my mantra. That's why I live by.

Rowan: Yeah, look, and it's definitely something that, you know is the responsibility of developers these days because they're the ones that are writing the code. And so anything I think, which makes it easier for them, is a good thing.

Jeremy: So what about what about metrics, though? Getting metrics out, right? Because the observability is one of those things where it's possible to do it. But Cloudwatch is a good place, you think, to get those metrics?

Rowan: Yeah, definitely. Like most AWS service is you get a whole bunch of default metrics that will come out of there in terms of number of executions, how long they took, the various states they were in. And so I think that's a really good place to start with your monitoring of of Step Functions. You know, obviously, CloudWatch dashboards has its limitations. But in lieu of any other kind of monitoring solution, it's a really good way to get some basic visibility, cause that'll let you identify any kind of anomalies. You know, if most state machines you have run in a matter of minutes, and you've got one that's going for days, should probably have a look at that.

Jeremy: Definitely. So now that EventBridge is out too, you have the ability to trigger Step Functions from events. I mean, you could technically do it with CloudWatch events as well. So is that something you recommend using that as a way to invoke them, as opposed to maybe invoking them directly from a Lambda function or something like that?

Rowan: Yeah, definitely. You know, I think there's a lot of kind of scheduled tasks that will suit themselves really well to being defined in Step Functions and then executed on a regular basis. Whether that's using scheduled events or I see, you know, using Event Bridge as a really good way to decouple the request for an execution from the actual triggering of that execution. So the thing that wants it could just talk to EventBridge and say, “Yep, I need this job done”, and then that can trigger the Step Function that can then later on, kind of return its status. Whether it's you know, via the various other kind of integration service integrations and you've really decoupled, the kind of the request from the performing of the work, and that's gonna work better in a serverless and an asynchronous workflow in the long run.

Jeremy: Now, what are your thoughts on just this idea of service boundaries and where Step Functions might fit in? Because for me, I always look at it to say if I'm building a user service and I'm building a billing service, and those are two separate things. I may have workflows within each of those services that require several steps, and I need to make sure they all complete. And I would definitely use Step Functions within those services, but what about coordinating across multiple services, right? So would you have a Step Function that says, “all right, I'm going to run a task that processes the inventory for this product, and then I'm gonna have another that's gonna go to another task that charges the credit card or whatever those different steps are.” Just what you thought of thoughts about using Step Functions across these service boundaries?

Rowan: Yeah. Look, I think it's definitely something that can make sense. But the caveat being that it needs to be done in a kind of a simpler way as possible. And so that's where this new nested Step Functions approach could do that kind of processing, I think the key thing that anyone developing a serverless application needs to remember is you do need to plan for failure, even when using things like, Step Functions. You know? Definitely. As as you said earlier, everything fails all the time, so as long as your system can handle, retries across those multiple systems, yeah, I think it can definitely work. And I really hope that as developers get more used to these tools, they'll be able to find new and interesting ways to apply them in hopefully ways that make everyone's lives better.

Jeremy: Yeah, definitely. Um, so just maybe one more question about long running tasks. So you don't recommend using activities anymore? Those even still available? I haven't used them, I haven't used them in a while, and I actually haven't tried the callback stuff yet, but so what are your recommendations on that?

Rowan: Yeah. Look, activities aren't gonna go away you know, AWS doesn't get rid of services very often, they just kind of stop mentioning them in their updates. Look, there might still be occasions where activities make sense, especially if you already have those long lived workers, and you really just want to control them and orchestrate them. That's where I can definitely see it making sense. Perhaps if you also have a worker that is not as well integrated with AWS and can't necessarily return that task token itself, although it's gonna have to be talking to AWS to get the jobs in the first place. But, you know, there may be some limitations there if your worker can't handle that. Other than that, you know, I think that the waitfortask token, the callback pattern is really just a much simpler and more elegant approach that's hard to beat.

Jeremy: Yeah, all right. So I actually have one more question, and that's on pricing, because I think pricing is something that turns a lot of people off to Step Functions because they're not cheap, right? So I mean something like two cents per 1000 or 2 and 1/2 cents per 1000 transitions. If you have nested Step Functions and you have all of these different workflows running through them, that can add up pretty quickly. What are your thoughts on pricing?

Rowan: Yeah, look, so it's, you know, if you look at the raw numbers, it's definitely not the cheapest. But what I found in practice is that it is pretty forgiving. You do have to be aware of it. And, you know, if you had a really high volume workload, maybe you deliberately try and not put Step Functions in the mix there and instead maybe use it at a higher level to, say, coordinate some of that batch processing rather than having a state or transition for every single processing activity that you do. I know for myself, you know, I definitely have learned the hard way. I misconfigured how my Step Function was being triggered and came in the next day and found that it had been triggering itself consistently for the last 20 hours and run a couple of thousdnd executions. And luckily, you know, the free tier is pretty generous. So I think it cost me a couple of dollars, but, you know, there's, I think, you're probably gonna have a lot more issue with costs around service is like API gateway and some of those things than you are Step Functions, unless you're putting it in your kind of high volume processing. And maybe don't do that.

Jeremy: Yeah, I mean, I think that's the biggest key is sort of like, if you're using it to process like clickstreams, that's gonna get expensive very, very fast.

Rowan: Yeah, definitely. You know, and there are other tools that are probably more ideally suited to that. I see Step Functions fitting into a much more kind of generic kind of higher level workflow rather than something quite as low level as clicks.

Jeremy: Yes, definitely. All right, well, so, I guess let's wrap this up. What would be your advice? You know, for people that have not used Step Functions yet? What's the best way to get started with Step Functions?

Rowan: Yes, I think going to the console, there's a lot of sample projects there. In particular, they have the callback pattern listed there. You know, I think most people, if they sit and think about that, they can probably come up with one or two workflows that they’d like to have automated. And you just have a try at representing that as a state machine. You know, the JSON language that's used to describe them in the language spec is really simple. You do have to get used to a few of the terms. You know, some of these various states and tasks that we've talked about here today, but it gets familiar pretty quickly, you know? So don't be turned off by the language.

Jeremy: Awesome. All right, well, listen, Rowan, thank you so much for being here, sharing all your knowledge with the serverless community. How can listeners find out more about you?

Rowan: Yeah, I do most of my work, probably through my blog, which is just rowanudell.com. I'm on Twitter occasionally as well. Spend most of my time talking about AWS stuff there. Other than that, you know, I've written a book on AWS, The AWS Administration Cookbook, which is actually about to get its second edition. I'm not writing the second edition, but you know that should be out in the coming weeks. Yeah, that's me.

Jeremy: Awesome. All right. Well I will make sure that we get all of that into the show notes. Thank you so much.

Rowan: Thanks, Jeremy

View Details

About Gillian Armstrong and Mark McCann

Gillian works as a Solutions Architect at Liberty Information Technologies. Her team is focused on thinking about big problems, and working out how to solve them using innovative technology in interesting new ways. At the moment she is working on Artificial Intelligence, with a particular focus on Conversational AI design and development. She has more than a decade’s worth of experience in many technologies across the full stack, and loves being a software engineer as it allows her not just to think up big ideas, but also to make them a reality.

Mark is an Architect at Liberty Information Technology that has been developing software and solutions for Liberty Mutual for nearly 20 years. He is currently working on making "Business idea to production in minutes" a reality. Mark holds several AWS Cloud Certifications and has a vast amount of experience with microservices, event-driven architecture, Docker, AWS, and other emerging cloud technologies.

  • Gillian Twitter: @virtualgill
  • Gillian Web: virtualgill.io
  • Mark Twitter: @markmccann
  • Liberty IT: liberty-it.co.uk

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Gillian Armstrong and Mark McCann. Hi, Gillian and Mark. Thanks for being here.

Gillian: Hi. Thanks for having us.

Mark: Hello.

Jeremy: So both of you work on the team at Liberty Information Technology, which is a part of Liberty Mutual Group. So Liberty Mutual, if people don't know 100-year-old insurance company, one of the largest here in the US - you have, what - 30 countries you work with, 50,000 employees, something crazy like that. So let's start with Gillian. You’re a Solutions Architect there. Why don't you give us your background, a little bit more about what you do?

Gillian: Sure. So I have worked across a lot of the areas in Liberty, including our emerging tech space, where I was first able to work on some completely serverless-first projects. This year, I’ve been working with the teams in our digital ecommerce space in Boston, looking at serverless, looking at AI, and I've just moved to our Data and Analytics unit. I definitely have a big focus on driving the serverless mindset in the company and also trying to get involved in the serverless community as well. And I'm also looking at AI sort of from an engineering and serverless perspective. So how far can we get using the managed services? How do we bring it into a large enterprise systems? Because, as you said, Liberty Mutual is a huge company.

Jeremy: Great. All right, Mark, what about you? You're an Architect there. Why don't you tell us about yourself?

Mark: Yeah, I’m an Architect with Liberty and similar to Gillian, I have worked across multiple different teams and areas in my 19 years working here with Liberty for everything from C++ mainframe development, the introduction to JavaScript, the introduction to Java and Spring and moving into this sort of adoption cycle. And then now, more recently, moving into microservices, all the good DevOps practices, and ultimately, where we're heading there with this big push to the cloud and serverless adoption. So been through the entire journey from you know, from mainframe to serverless.

Gillian: Yes, so Mark, and I work in very different areas, but we try to be really collaborative across the company and sort of some of these bigger things like serverless.

Jeremy: Well, Mark, it's good to be talking to another developer old-timer like myself. I appreciate that. Alright, so let's start, because again, Liberty Mutual is huge and you are actually part of Liberty Information Technology. So these are separate companies and I'm fascinated by how large organizations work and how all things are distributed and stuff like that. So maybe one of you can explain to me and to the listeners, what's the relationship between Liberty Information Technology and Liberty Mutual?

Gillian: Yes. So we did say Liberty Mutual had about 50,000 employees. About 4000 of those are in IT. And the company we work for Liberty IT is a wholly-owned subsidiary of Liberty Mutual. We have about 600 software engineers based between Belfast in Northern Ireland and Dublin in Ireland. And they are all fully focused on delivering world-class software and solutions for Liberty Mutual.

Mark: Yeah, so we're very much a software house, focus on high performance engineering and really delivering those world-class solutions that Gillian mentioned. So we're slightly different from the rest of the Liberty Mutual sort of area where they may have a mixture of developers and business. We're very focused on software engineering.

Jeremy: Awesome. Alright, so that makes a lot of sense. Thank you. Alright, so I want to talk to you today because you both mentioned a lot about serverless. Liberty Mutual is obviously embracing serverless. So let's talk about that. How is or how Liberty Mutual is embracing serverless? And maybe let's start just by sort of how did the team kind of discover serverless like, what was the point where you said, “Hey, let's start looking into this new technology?”

Mark: Yeah, I think that also goes back to where, in 2014, where we started our public cloud journey, I guess. So the form the public cloud came, and they started opening up the access to AWS and seeing how this whole new cloud thing would work within the big enterprise. And so a lot of that start in 2014 into 2015 was really just dipping our toes in the water and exploring cloud capabilities. Then we had very little workload in there. Coming into 2015, it was starting to get the initial learning, starting to set up the pathways to get to the cloud. Try to build in those capabilities that a big enterprise like ours needs. So real focus on security. Real focus on how do the development teams actually get access to this stuff. You know, what are good practices? And again working with AWS and partnering with them to figure out what that looks like. So our public cloud team did a really great job in starting to explore the space and open up for the enterprise this new awesome capability that’s the cloud. And then into 2015 or 2016, we were into the, you know, starting to really think about what apps we could we migrate to the cloud. What modernization can we do? What sort of approaches could we take so that we can break down our big monolithic applications and break them into things that will actually fit in the cloud. And then all the way through to 2016, we have maybe 10.5 percent of our workload’s in the cloud. 2016-17 we’re in the 12.5%. Then 2018 we’re at 20%, and now we’re up to about 30% of our workloads, plus, in the cloud, and all through that time it’s been around developing the capabilities, developing the expertise and partnering with AWS, and really learning what the cloud capabilities are there. A lot of this was traditional through EC2, RDS-type work, moving the workloads into that. And then we get into containers, we get into the whole DevOps practices, and now ultimately, we're starting to really get after serverless.

Jeremy: Okay, Awesome. Alright, so then let's talk about just your strategy for adoption, right? And we can probably go a little bit deeper into the timeline as we get further into the conversation. But let's talk about the strategy. What was your strategy for adoption?

Mark: Yes. So our strategy, really at the minute is to create an environment for serverless to succeed. So we mentioned the work that our public cloud teams has done and that’s really been about building the CI/CD pipelines, building up the developer access, making sure that as a new developer joining a team, there's nothing blocking me from experimenting with the cloud, delivering some innovative solutions in the sandbox environment and then ultimately promoting that though all the different stages to production. So a lot of the work has been done by the people have come before us — Gillian Armstrong, Delia McCann in particular, have really helped accelerate a lot of the public cloud capabilities that we have. A lot of that is now in place. So what we've been trying to do now is evangelize what a serverless-first approach looks like. Educating our leaders, educating our developers so that whenever they're looking at any potential problem that we usually have, our customer has, they are trying to adopt that serverless-first approach, seeing if a serverless-first approach works for that particular problem and seeing where it’s a good fit.

Jeremy: So is that something now where — you said the serverless-first sort of thing — so everything you build now, you just do serverless?

Mark: Not necessarily…

Gillian: Yes.

Mark: ...as we much as we would like to say yes…

Gillian: I wish.

Mark: We want them to have a serverless-first mindset and a serverless-first approach, but we know that it's not a good fit for all contexts just yet. So we have a number of fallback options on that serverless spectrum that we want our teams to fall back to, but ultimately, we want them to at least try to see if the serverless-first approach works. Does it work in full-on Lambdas with managed services Is that a good fit? Does it give you the solution you need? If it doesn't, then let’s fall back onto some sort of container solution potentially or maybe some sort of past solution. And ultimately, if you fall back far enough, you might end up back on prem. But then we hope that you're really trying to aim for that full-on serverless experience.

Jeremy: It sounds like it sounds like Gillian doesn't agree with you 100%. So Gillian, get your take on this.

Gillian: I think we find that anything we build fully greenfield, we really have been able to go completely serverless. Even with huge enterprise systems, I know we’ll talk more about that. I think people feel like serverless is just a little thing. It's great for playing about small applications, but we've really been able to prove that we can build very sophisticated enterprise systems. The challenges do come when you are dealing with legacy applications in your 100-year-old company. You do tend to have a lot of stuff left about from from back in the day. You know, I I remember when you if you needed a server, someone maybe had to drive it on the truck to the data center and you had to wait until that happened.

Jeremy: Those were the days.

Mark: No, it’s really because we're evangelizing this serverless-first approach. There are some challenges to that way of delivering solutions. So we're trying as a strategy to identify what those challenges are ahead of time. Trying to make sure that the managed services that we want our team to adopt are actually available to them. They have been through our security assessments. They have been added to our allow list of services that we can use. So if you want to use something like AppSync, it's been pre-approved, that's been enabled in our accounts. It's available for the development teams all the way to production. So not only re-pushing the team's adopt this serverless-first approach, we really want to make sure that the serverless capabilities we would like them to use are actually available to them. And so we've been working very closely with our public cloud team, our DevOps teams, but also security, legal, privacy teams to make sure that the services we are asking our teams to at least evaluate are actually secure, cost-optimized, any sort of legal or privacy concerns have been mitigated. So we've been trying to put a lot of those things in place to make sure that it's not just your architects telling you to go serverless-first and then not enabling any of the teams to actually do it. We're trying to make sure that those blockers have been removed on the pathway to serverless.

Jeremy: The labyrinth that is the giant enterprise, right? So how did you get going with with serverless. I know you talked a lot about the cloud, you have a cloud team and things like that. But is this something where it was sort of like a ground-up approach where the developers are bringing it in? Or is this something that came down from the top?

Gillian: Yeah. So we do have a cloud team, but they manage our cloud accounts. They, as Mark said, there's a lot of enablement needs to be put in place to make sure that as developers move from on-prem to the public cloud that we’re given good governance and we're not going in and opening all the ports to the whole world. So they very much put those structures in place. But then it was the teams that came along and actually built applications that started to prove that it really could be used. So I was on one of those teams with Mark's wife, Gillian McCann actually, And we built out an employee digital assistant, and it was something that definitely brought a lot of learnings. Where were the problems with security? Where were the problems with how do you observe? How do we audit? What's going on? And then really from that, and from some other projects that were going on, we were able to start to really learn about can we use this in-house? What are the blockers? What are the barriers? And go back to the public cloud team and get them to support us, get them to help us, start building out CI/C pipelines. I know Mark’s been doing a lot with developer enablements and helping other developers. And everybody doesn't have to sell the same problem over and over again because serverless is supposed to make it easier.

Jeremy: Supposed to.

Gillian: Supposed to, yes.

Jeremy: Yeah. So what about leadership? Because especially in the multi tiered organization, where does the influence come? Are you working with the leadership to try to get them to kind of integrate that into their vision or what's happening there?

Mark: Yeah, absolutely. And part of the job is to articulate the benefits of serverless and the serverless advantages that come with it. And talking to your leaders, talking to your senior your leaders and making sure that they understand the value proposition of a serverless approach. You know, it really is the ultimate cloud capability that we want to pursue. We want to focus our developer efforts on delivering business values, so really trying to articulate that we do want to deal with this undifferentiated, heavy lifting. We want to focus on really solving real customers’ problems. And because we are a big insurance company that means real people are ultimately the users of our software capabilities. So we're very acutely aware that we don’t want to waste time, money or effort on stuff that actually doesn't have an impact to our customers. Ultimately, that's our mission here as an insurance companies, is to really actually help people.

Gillian: It's been great because we know our CIOs have publicly stood up and told the company that this is a new approach that they recommend. This is the go forward. So it is great to see the message coming down from the top as well as coming from the grassroots and then the developers as well.

Mark: And recently, with the senior leaders in some of the spaces adopting a serverless-first approach is now in some of their OKRs and some of the objectives that we're rolling out across the thousands of developers that we have in some of our divisions. So it's a nice stated objective off our teams that we need to evolve to meet this serverless challenge.

Jeremy: That's awesome. Alright, so now you mentioned AppSync is one of the things you kind of approved. So you must be using a lot of managed services though, right?

Mark: Yeah, absolutely. Across the board, we're using pretty much everything. So depending on the team and the problem at hand, you know, we have everything from the typical Lambda, DynamoDB, the Kinesis stuff...

Gillian: API Gateway, SQS, SNS.

Mark: Yeah, AppSync is more recently, as we’re starting to get more of the GraphQL and evolving from full-on RESTful services to much more GraphQL-type services. And yeah, we're in CloudFront. Everything really. We have access to the full portfolio of serverless capabilities pretty much across the board. There are some enterprise decisions that were made that have Route 53 for example, we don't have access to. We have our own DNS solution. So that becomes some of the challenges that we have when we start evangelizing serverless. Developers go on the databases or they go on to some of these talks to people like yourself or some of the patterns that are right there and they see these things solutions, and they try to bring them in-house. So maybe they don't work quite the way that you would expect, because some of the capabilities aren't allowed for us for good reasons. And so we need to, unless something we're working on a lot is to take those patterns and make them work for our ecosystem in our context, so instead of this DNS provider, we’re using this and you update the pattern or update the templates or update the CloudFormation to work within the ecosystem that we have. So a lot of that’s around identifying blockers and barriers and making sure that we've got a solution in place for them.

Jeremy: So okay, I don't know if this is true, but my understanding is that if you're a technology company in the UK, you have to use Wardley mapping now. Is that true with you guys as well?

Mark: 100%. We use it a lot for our strategy and most recently have been using it to talk to my teams that I'm sort of responsible for to really understand what the team purpose is. What their, who their users are, what they're delivering to their users, but also to get a real understanding of their tech stack and I'm using the Wardley mapping to see how they could evolve their tech stacks to meet these serverless builds that we have as a company. So it really helps me to talk to the team. It helps me to actually show the teams this is where we're heading and it also gives us fast feedback on any sort of blockers that that may have to serverless adoption.

Jeremy: Awesome. Alright, so let's go back to the timeline. So you mentioned 2014, public cloud. Somewhere around, I think you said 2016, you started breaking down some of the monoliths. So maybe we start after that — after you broke down the monoliths and you started moving into actually bringing things to the cloud. Let's start there.

Mark: Yeah, so, we started the digital assistant work sort of kicked off around 2016 as well — the work that Gillian and Gillian McCann were doing. We had a lot of the security teams were starting to build out serverless capabilities there as well. Some of the auto remediation stuff and we can talk about those because they’ve talked publicly at re:Invent and other areas about some of these capabilities, but they built in a lot of really awesome security capabilities in the serverless way. It provides guardrails for the development teams and makes it so that we can do things that are against to our company policy. But it frees us up to go off and try and experiment with stuff. It'll give you a nice feedback on don't open these ports or you can't use this particular capability or everything has to be encrypted at rest and encrypted in transit. So again part of that developer enablement is enabling them, but putting guardrails around some of the things so that we're not exposing ourselves to risk.

Jeremy: So you started — a lot of the serverless stuff you were doing originally was internal, like, sort of compliance, that sort of stuff, right?

Gillian: Yeah. So some of the first things were utilities for the public Cloud team themselves. They were definitely the experts at the time. And then we really moved from that to other internal systems. So we do a lot of employee-focused software and when we build internally for ourselves, we do have the opportunity to use a lot more emerging technologies, to take a little bit more risk, experiment on our employees for employees. So we have an internal productivity tool that we're actually selling externally now to other customers. We called it MyHub. On top of that, we built the digital assistant. It was a chatbot, and we built it fully serverless. And alongside that we had a few other small tools internally were being built-out and that let us work with those security teams, that sort of public cloud team that exposed where we were having issues, let us negotiate. So you know, we have to go forward. This is the future. So how do we this? Not, you know, if we sit down with the legal team and security team, it's not can we do this? It's like, we have to get to the cloud. We have to be able to use these tools. This is the future. So let's work together to work out how we can do it and alongside that, working out patterns, working out sort of best practices, if there is such a thing as a best practice. Creating resources for other people, here's the CI/CD pipeline for how you can deploy this. Here is a custom resource you can use yourself , and then as we moved on and hit the customer-facing applications, they were able to move really, really quickly that a lot of barriers had been removed for them.

Jeremy: So before you did that, though, in terms of the actual like customer-facing workloads were you — I think you mentioned something about containers — were you still sort of going down the container route initially?

Mark: Yeah, absolutely. And I think 2015, 2016 into 2017, we were very much still moving down that move to microservices so we did a lot of work talking to the teams, doing things like event storming to really break down the monoliths and the more bounded context and sort of microservices. But building out the containerized solutions, whether it's a Pivotal Cloud Foundry or Docker data center we had, you know, these cloud-ready, cloud-native, sort of containerized environments. And so we were doing a lot of work there to break down the monoliths and the microservices and moving them onto containers and getting them right onto the cloud. So a lot of the work we did around 2016 and 2017 was really around sort of moving that stuff into the sort of workloads as well as also the work that Gillian and others were doing around the serverless stuff. So a large bulk of the work was really moving that to that containerized world. And that gave us a huge amount of advantages and really accelerated a lot of our cloud adoption and accelerated our time to market. And like Gillian mentioned, we were doing a lot of education, building a lot of patterns and building a lot of the good practices around into what does an event-driven architecture look like? What do microservices look like? What does that mean for our teams? What ways do we educate our team? How do we train our teams? How do we help guide them on some of this stuff? And then, ultimately, we've got to look, we've had a lot of success with that, and that's helped really pushed a lot of our workload onto the cloud. Now we’re into the next phase of evolution, how can we then just move on to the serverless-first approach. Containers will still play a large part of our future, I think, for a while, but ultimately we want to evolve and move everything to that serverless ecosystem for all the advantages that serverless brings.

Gillian: I managed to somehow miss containers and skip past them. No one’s going to drag me, back. I think it is interesting. So as Mark said, obviously, we have a lot of applications we couldn't get out quickly to the cloud if we if we didn't use containers. So in that way, that is a step between, you know, going cloud native, going serverless. But realistically, you know, containers does not get you closer to serverless. It just gets you onto the cloud. So I always say to teams, “Look, if you've not gone to containers yet, see if you can just skip past them. Don't spend the time putting on containers. If it is possible to go to serverless.”

Jeremy: Yeah, I totally agree. I mean, the problem that, you know, I always see is that you can't lift and shift to serverless. It has to be a rewrite and just a lot of teams don't have that time to do that. So I have nothing against containers. I think containers are great, especially when you're breaking things in the micro services. But I totally agree. If you can skip them and go right to serverless, that is my preferred approach as well.

Mark: Yeah, 100% agree. I think having those fallback options as well is good and making sure that people know that the problem they're trying to solve and the context that they're within, and that’s the whole rationale behind the serverless-first sort of mindset we're trying to push. Try to do this in a serverless-first way and we think it will solve, it'll be suitable for a lot of the problems we're trying to solve, but if it's not a good fit, we have plenty of other options.

Jeremy: How much of your workload is in the cloud now?

Mark: So I think we're due from 2017, 12.5% and 2018 we were 20%-ish. Now, we’re probably high 38%. But that's always increasing. So there's a lot of acceleration there in the cloud. So we're really starting to accelerate that, that move and moving everything into the cloud properly.

Gillian: Yeah, there's tens of thousands of servers, you know, sitting in our internal data center. So it's amazing that we've already got to almost 40%. Hopefully, we're going to keep going, keep accelerating but I know that the last ones in there will be, there’ll be dragons in there.

Mark: But I get you anything that you do, we’re trying to make it so that they build serverless-first, and all the pathways and the blockers for serverless adoption have been removed.

Jeremy: I can imagine with an enterprise like Liberty Mutual, you probably still have some servers that are actually writing onto stone in order to save information. So I've talked to a couple of development teams and a few development team leaders — relatively small teams, maybe five to 10 developers that are trying to get their teams to move to serverless. You have 600 developers, you said, that work for Liberty IT. So let's talk about how you get the word out, right? How you train people, how you evangelize serverless internally, how you enable those developers. Let's kind of start there. What do you guys do to get the word out?

Gillian: Yeah, so the first thing — we say not everybody's a cloud developer yet. But so the very first thing we needed to do is really educate people about the cloud and the benefits of the cloud. You can't get to serverless if you don't know how to build on the cloud. So there has been, for several years now, a lot of support put in place to let people learn about the cloud, to get their certifications, to get time to do that, and as we came a couple of years ago to start thinking about serverless, initially, internally, we were trying to share how we have internal systems where you can blog. We have internal tech talks. So we spent a lot of time talking about what we were doing around different groups. But ultimately, the big tipping point came when myself and Gillian McCann got a spot at re:Invent and went and talked about a solution we were building. Because sometimes being able to go externally and showing that your expertise is on par with other people in the industry is actually a really great way to get people internally to listen to what you're doing, to pay attention, to actually find out about what's happening. Since then, we have had a lot of people out at Serverless Days, Serverlessconf, QCon, and sometimes that's actually more effective because people will go and listen to the recorded talk. Then they'll maybe come and talk to you internally. We've also been doing a lot of things where we get AWS, Google to come in and run workshops or give talks which gets people away from their desk for a day doing something. And then also we're getting teams to do informal hackathons, engineering days, letting them try things out, which really sort of leads into that developer enablement because that's part of, you know, enabling them.

Mark: Yeah, And even on top of that, we've had cloud native open spaces where we get the whole company together just to talk in a very open way about cloud adoption and the challenges and the successes that teams have hard and just really get developers to talk to each other because we're a large company, so lots of people have had a good experience or challenges with certain things and just getting developers to talk to one another has been a great empowerer for this adoption over we're trying to get there. So we're very collaborative, engaging sort of company, and that's been in our culture. So we really want to our engineers to share that and really try to help each other.

Jeremy: So how do you sort of, you mentioned things like repeatable patterns and different models that people can use. How do you share those? Do you just have a sort of internal Wikis and things like that? That that's where that stuff goes?

Mark: Yeah, we've had a number of big initiatives around their enabling these cloud native approaches. So within the GRM in which is another customer-facing insurance area, we've had a DNA, which is digitally native architecture approach, so we built a team that took all these patterns and made them sort of code executable for developers so they could go to dashboard, click a few buttons, and they would have a fully cloud native solution deployed with CI/CD pipelines, with all the security checks and balances, with all the quality baked in and that would enable you to deliver a new cloud native API all the way to production in a few minutes, and similarly, in the area I am in now, we're doing the exact same thing, but with the serverless-first pattern. So we were baking the patterns into templates that are then executable by developers and really sort of make sure that all of our good practices and security standards and quality are baked into those patterns. So with a few clicks of a button, the developers have a full-on serverless solution with the pathway to production all automated and ready to go. So that's our trend to sort of capture all the good stuff that Gillian, Gillian McCann, others within the company have done. We want to bake those good practices into these repeatable templates that are really executable. And that's why there's something like CDK coming out. It has really piqued our interest. We want to try and capture some of our infrastructure good practices and make them CDK constructs so that again, we can accelerate that developer enablement so that these well-proved, hardened, secured, complaint capabilities are available for all our development teams.

Gillian: So we have an “inner source” as well, sort of like open source. But internally, we have repositories where people can share things, GitBooks are always a big fan of people. There’s Slack channels where people share different links to different things, either externally or internally about serverless.

Mark: But ultimately it's that serverless-first approach. We want to remove any undifferentiated heavy lifting, so we want to package up and automate as much as possible. So that if a developer needs a full-on serverless stack, they shouldn't have to go and craft it themselves. There should be good stuff available to them, to accelerate, so that they can truly focus on business value.

Jeremy: So what about giving people time to sort of experiment? Because even whenever I come across a new service or service that I might have been using, but now I want to do something, codify it with CloudFormation and automate some of the processes and things like that. I mean, that just takes time. I could spend an entire day just, you know, messing around with a CloudFormation template sometimes. Do you give teams the ability to do, you know, sort of experimentation on their own?

Mark: Absolutely so, because we are software engineering sort of company within LIT, and we have a culture of engineering excellence, we need to give teams capacity to learn, to explore, to play with some of these new capabilities. So, a lot of our teams have innovation time, to dedicate an innovation time that allows them to explore and experiment with some of this stuff and it doesn't even necessarily need to be a lane to your business goal. Can we just explore this new technology because it’s cool. But we have encouraged all of our teams to have 20% time pretty much at least to explore new capabilities and learn and read and do the right thing for their teams. So, yeah, with the pace of change and the pace of new capabilities coming out, if we don't have that capacity for teams to explore new technology, you quickly get left behind.

Gillian: Yeah, I think also we try to educate product donors and other groups outside of IT that when they're bringing in new technologies, they're actually using for features on their projects that may take a little bit longer. But ultimately, you know, they will go faster and rather than just demanding that they keep accruing tech debt, you know, repeating everything that's ever been done, but they let them, you know, have a little bit of time to bring in new things.

Jeremy: And is that something, I mean, I can imagine if you've got 600 people or, you know, however many people are working on different serverless projects, they've got to be discovering new things, better ways to do it, so even if you've codified a pattern or whatever and they say “You know what, I found a better way to do that” — sounds like very grassroots, like it just kind of works its way back up through the system?

Mark: Yeah, yeah, pretty much. And we have a lot of vehicles for people to share. There's internal Wikis; there’s internal collaboration platforms that we have. There's pretty much all the teams would have some sort of tech talks or a regular sort of schedules, sort of show-and-tell-type time so that they really, if one team has found a new way to do something, or say a new capabilitie been released on AWS, or on-demand for Dynamos, you know, on-demand sort of capabilities will really reduce the cost of your Dynamo instances. If the team's already turned it on and figured out how the CloudFormation gets updated to enable that, they may talk to another team from the other side of the building and say, “We did this crazy thing and saved us hundreds of dollars a day, you might want to do that.” So we’ve a very collaborative culture within LIT and so if somebody steals something cool, we very quickly hear about it.

Gillian: Very. The absolute joy of a serverless architecture is that it's evolving. Because it isn't this big upfront design when an architect goes and creates a massive big architecture diagram and then prints it out and puts it on the wall, and that's it for the rest of time. Because it is evolving; because it is easy to take pieces in and out and change them. That lets your team, you know, contribute to the architecture. It lets individuals maybe have a bit more expertise and come along and say well, we should change this. If new patterns are emerging, you can go back and update the architecture with a lot less overhead than it would have in the past. Maybe not with none, but certainly the joy is that it is a much more flexible and changeable way of building something. And we definitely in the serverless world, in the serverless community, there are definitely debates about the best way to do things. And I think people are coming up with new and better ways as new tools are coming out. As new functionality is coming out and I know I've been to a lot of conferences and listened to a talk, and it's just completely blown my mind. And I've gone, “oh, I've just been doing this thing completely wrong.” So I think it's pretty exciting. One of the joys of working in any emerging technology is you don't have to worry so much about decisions because you definitely have made the wrong decision, and you just have to accept that you're just gonna have to keep changing and evolving.

Jeremy: We'll always find too where it's something like, you build something really cool, you find a great work around to do something, and then, like two months later, Amazon releases a way to do it in one line of code, which is great. I mean, I that's actually one of the great things about just where we are in serverless, and the serverless space right now is as we start to find use cases and people actually start to use it, that's when you start butting up against those limitations and AWS and Azure and Google, they're all working to get rid of those, which again just is amazing.

Gillian: I honestly think that they lie in wait, and will wait and wait. And maybe a new feature will come out and then we'll spend, you know, four or five weeks building it ourselves, cause which along. And then the day you check in, it's done. You'll like wake up the next morning, you’ll look a Twitter and there'll be a blog post dropped, saying “Hey! That feature’s there.” And you’ll say, “No!”

Mark: And again, I think our teams are well aware that any sort of custom sort of workarounds that they’re building they may have evolve to use whatever the managed service delivers, because, you know, just be prepare to throw this stuff away. Because, you know, sooner or later, somebody will bake it into the ecosystem you're working within. And then you know that custom built thing you need is is no longer relevant.

Jeremy: Still worth it, though. No, still worth it building those things. Alright, so you mentioned a couple internal tools and stuff like that. But you guys have built a ton with serverless. So let's talk about some of the success stories. I think that'll be interesting to people, especially in a large organization. So why don't we start going with some of those internal tools again, like the employee digital assistant? What's that about?

Gillian: Yes, so the employee digital assistant. It's a chatbot. Very exciting. This is a little bit where I get to dabble in AI, and applied AI in that sort of serverless mindset where we pick up those managed services. And it was really based around making things easier for employees. So nobody wants to spend all their time searching for things on SharePoint or emailing someone for an answer, doesn't get back to you for a week. So what we did is we hooked up a whole pile of internal functions. We hooked up our internal help desk, finance, HR, and got them to put a lot of different things. You just go and ask the chatbot whatever you want, you know. So, “Something's going wrong with my payslip this month. What do I do?” Or everybody's favorite one, “What's on the cafe menu today?” Which was the top search from the intranet. So because we're building it completely from scratch, it was this amazing opportunity to build it completely serverless to try things out, and we were able to experiment a lot. And one of the most exciting things is, from that, we now have a whole spinoff company called WorkGrid that is a startup that has spun out from the Liberty Mutual that has taken that, rebuilt it as a SaaS solution, and is now selling it to other companies.

Jeremy: Awesome. And you also have something called Radar. That's your cloud adoption one?

Mark: Yeah. So we can talk about this because we've talked about it at re:Invent, so it’s cool. Our public cloud team created the security or auto remediation capability tool that really helps to prevent sort of any sort of bad behavior from developers and creating their resources or playing in the cloud, doing stuff that we shouldn't be doing. So it has a number of your security policies as code that will prevent us from creating new reserves with open ports to the world or not encrypting at rest or not encrypting it in transit. And it gives you nice reports and it gives you feedback, you know, where you're going wrong, where you're not following company policy but it also auto-remediates as well so it’s also triggered off CloudWatch events and will actually rectify and stop you in some cases. But really, it's an awesome enabling constraint because it means that our good enterprise practices are baked into an automated policy. So that security policy as code really helps keep us on the straight-and-narrow and so I think it really, it's not about the security, the Department of “No”. It enables us to go faster, but go faster with a good practice and good security.

Gillian: You're going to you every single developer in here who has the very first time they started to build something out has found their resources disappeared because they didn't tag them properly.

Jeremy: Yeah, well, that's good though. Alright, so what else? Any other interesting internal projects?

Mark: Yeah, so there's a lot of good work in our sort of financial central services space where we’re processing hundreds of thousands of records a minute in serverless and that’s all through step functions, Lambda, SNS, SQS, Kinesis, DynamoDB. But that's huge volumes, you know. That's really starting to stretch your serverless and the managed service capabilities. And so that's ongoing at the minute. But it’s a massive sort of success, and we’ve had _____ talking about it at Serverless Days in Dublin around some of the cool stuff that's going on there. So it really though, you know, serverless is not just for your simple APIs and simple, you know, utility-type libraries for the core sort of financial processing engine of the company. So it’s almost like shows that serverless is ready for pretty much every problem you can throw at it.

Jeremy: That's amazing. Alright, so what about customer-facing stuff? Have you put anything out there now that your actual customers are using that’s built in serverless?

Gillian: Yeah. So one of our biggest ones out there is the virtual agent, which is in our call centers. They've deployed a virtual agent that's answering some of the calls, so it's taking some initial stuff or answering really simple stuff. And they were actually able to use some of things from the digital system, some custom resources, some of the patterns we put in place for the use of Amazon Lex, which is Amazon's natural language understanding service. And that let them move really quickly. They were picking up with Amazon Connect, which is their call center as a service as well, which was a brand new tool and a lot of learnings. But it's massive. I mean, the cost is so low, because it's all managed services, and they're able to, you know, bring it out, trial it with users for a little bit, make sure the whole system works, and then just scale seamlessly up. And they're just adding in more and more functionality all the time. Ah, and in fact, they presented that last year at re:Invent so that talk is out there too.

Jeremy: Awesome.

Mark: I think one of the early really customer-facing ones was around the document generation, policy generation sort of capabilities and within our sort of Liberty Mutual benefits space, and a lot of the underlying document management, document generation stuff is very heavily built on serverless and around the same sort of time as the digital assistant, so there was two sort of teams really help pioneer a lot of serverless capabilities within the company and helped to prove that this is something that is going to be a game changer for us as a company going forward.

Jeremy: Awesome. Alright. So let me ask you this question because obviously you two are deep in the enterprise world. You're on the forefront of bringing serverless in there. So what would be your advice to other enterprises looking to adopt serverless?

Mark: I think getting access is number one. Creating that access to some sort of sandbox environment where your developers can explore an experiment risk-free whether getting chewed out by some manager for what are you doing opening up in AWS account in your credit card. So I think that was key. Also, we had a sandbox environment that we could explore new features and play a little bit and actually even access the console. Beyond that, having a clear pathway to production is critical and our public cloud team has done a fantastic job really automating all of the security compliance or legal issues that we may have and making sure that they are part of that automated pathway to production so literally a developer could create something today, have it in production this afternoon. And that's how you automate it and how fast, we can deliver capabilities with that clear pathway to production is a big enabler for us because a company. I think you can't compromise on security, so there's a real enabler there around don't do anything risky, so make sure you know what your security profile is and make sure that you have on approach for dealing with security in the cloud space. We've spent a lot of time working on threat modeling and making sure that we work with our security architects and security teams to to really make it easy for developers to show the risks that they may have and show how to mitigate so I think that zero compromise on security is number one for us. And ultimately for us, if you're starting out, testing is a big thing, you know, really focus on those good testing practices. Making sure you have testing, unit testing, integration testing for it’s different in a serverless environment. But that allows you then to go safely — quickly, but with safety — so have a real testing approach and then… Observability is probably the big one for us. So monitoring observability, making sure you know where stuff is and you're getting appropriately alerted, alarms and everything's things happened is a big one for us. Gillian, do you have any?

Gillian: So I tell people, especially in big enterprises, the same for both serverless and AI, which is: start now. You’re already behind. If you haven't started, you need to start now. It takes a while to learn all the things that Mark’s just said — a lot of things. It takes a while to move your mindset from highly architected things before to serverless. Serverless is very different, even the microservices. So even if you're familiar with microservices. This is still a different paradigm. So it just takes a little while to learn. It takes a little while to move all your existing practices and thinking about how you build your systems. So you need to start. You need to find places that are sort of safe-to-fail places where you can try things out and then gradually scale up. I think the big thing is, if you run the company, do you create time for people to learn. Do let them have that space and, you know, find your people who are really, really passionate about it and then let them loose.

Mark: I think that was key for us. We had Gillian Armstrong, we had Gillian McCann, Laura MacFarland, Chris Gormley. We had a number of real serverless pioneers in the space. We really do blaze the trail and opened up a lot of number of doors for everybody else to come from behind. So kudos to those guys.

Jeremy: So I think that we can sum up your advice, Gillian, maybe by using that “old planting a tree” proverb. The best time to start with serverless was five years ago. And the second best time is now. Right?

Gillian: Exactly.

Mark: I think one of the big things that we've certainly noticed is, you know, getting certified has actually been very useful for us. Certifications for certifications sake are about pointless. But we find that really has helped accelerate the knowledge of our development teams. It gives them something to aim for, but also maybe it's just the way that the AWS searchers are set up. You know, they've been very applicable to the technologies and the approaches and the patterns that we are preaching our teams to embrace. So the AWS certification journey has definitely been a worthwhile one for our company. And I think we have reached the tipping point for that. We're over for 10% certified now, so it's really helped accelerate so that whenever we talk to our teams, they know what we're talking about, which is usually a good first step.

Jeremy: Definitely. Awesome. Alright, well, listen. Thank you, Gillian and Mark. This has been an awesome conversation. How can we find out more about you? Let's start with Gillian.

Gillian: Yep. So Twitter, I am @virtualgill on Twitter. And Twitter's definitely where I hang out most of the time, and always happy to have conversations with people. My website is virtualgill.io, if you want to look at some of my talks or read some of my blog posts where I rant about various things. Um, but definitely Twitter’s the best place if you want to have a chat.

Jeremy: And that's @virtualgill with the “G.

Gillian: With a “G,” yes.

Jeremy: Okay. And then Mark, what about you?

Mark: Twitter, I’m probably on their Twittering all the things serverless, so you can get me at @markmccann on Twitter.

Jeremy: And then if people want to learn more about Liberty IT, the website for that is just liberty-it.co.uk, right?

Mark: That's correct. Yeah.

Jeremy: Awesome. Alright. Well, listen, I'm going to put all this into the show notes. Thank you guys so much.

Gillian: Oh, there is one more thing, Jeremy.

Jeremy: Oh, there's one more thing.

Gillian: One more thing, right. One more thing. Yeah. There's going to be a Serverless Days in Belfast.

Jeremy: Awesome.

Gillian: It’s going to be November. So I'm helping to organize it. So if people in podcast want to come, then they should definitely, follow me on Twitter for when we tweet the formal announcement.

Jeremy: Awesome. Alright, Thanks again.

Gillian: Thank you

Mark: Bye.

View Details

About Forrest Brazeal

Forrest Brazeal is an AWS Serverless Hero and a Senior Cloud Architect at Trek10, where he hosts the Think FaaS serverless podcast and contributes to their open source efforts. He understands the challenges faced by enterprises moving to the cloud and loves building solutions that provide maximum business value for minimal cost. Outside of his day job, he interviews the top names in cloud for his “Serverless Superheroes” series at A Cloud Guru and also creates the FaaS and Furious webcomic. He is the co-chair of ServerlessConf and regularly speaks at workshops and other events in the serverless community.

  • Twitter: @forrestbrazeal
  • LinkedIn: linkedin.com/in/forrestbrazeal
  • Blog: forrestbrazeal.com
  • Trek10 Blog: trek10.com/blog
  • Think FaaS Podcast: Think FaaS with Trek10 Podcast
  • ServerlessConf in NYC: nyc2019.serverlessconf.io

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Forrest Brazeal. Hey, Forrest, Thanks for being here.

Forrest: Hey, Jeremy. Thanks for having me. It's great to be on the show.

Jeremy: So you are an AWS Serverless Hero as well as a Senior Cloud Architect at Trek10. And I think many people in the serverless community, as well as the underground technology rap scene, know who you are. But why don't you explain to listeners a little bit about your background and what Trek10 does?

Forrest: Sure thing. Well, and I think the underground tech rap community is basically just me. So that's a small pond there, but yeah, so my name is Forrest Brazil. I'm a Senior Cloud Architect at Trek10. Trek10 is an AWS advanced consulting partner. That means we work with AWS Technologies. We focus primarily on the cloud native side of things. So think serverless, IoT, basically anything you can build while using the best practices that AWS wants you to use. I spend a lot of time helping clients put things like that together. I also spend a fair amount of time being community-facing. I love to educate and advocate for serverless technologies wherever I can. That's why I do the serverless hero thing. And it's great to be here with you now.

Jeremy: Awesome. And I think I could probably talk to you about — I think I have talked to you about pretty much everything that serverless has. But I think what I want to do today is focused more on the sort of CI/CD process for enterprises. And maybe I know you've done a lot of stuff in that space, you and Jared Short have, and maybe you can start by telling us, sort of what's the difference between sort of your typical CI/CD process and one that involves serverless now?

Forrest: Sure, Ultimately, there's not necessarily that much difference. I mean, the underlying goal is the same. I've got code. I want to get it off my laptop. I want to get it into production and serving my users as safely and quickly and efficiently as I can. And that goal doesn't change because you're suddenly using different infrastructure, or less infrastructure hopefully, but I think where it gets a little bit interesting for serverless is you all of a sudden start having the cloud become part of your software development lifecycle a whole lot earlier than you may have been used to it in the past. So you don't write this code locally, and then you throw it over the wall and it gets deployed on some infrastructure that you're not thinking about. So much of your development process now actually is tied up in that configuration and figuring out permissions, thinking about latency at the cloud level, right? You're simply not getting a very realistic picture of what your application's going to look like if you're trying to mock all that and develop it locally. So we see people trying to figure out how can I get this app into the cloud tested earlier on in my lifecycle? How can I do that collaboratively? Or how can I do that without stepping on other members of my team who may be working on different features, perhaps in a shared account? That's where we try to put some best practices together that make things easier for serverless developers.

Jeremy: Right. Great. Okay. So why don't we start talking about it specifically with enterprises then. So you've worked with a lot of enterprises. I've worked with a lot of enterprises. We know that there are certainly challenges facing enterprises when moving, just to the cloud, let alone you know, just moving to serverless. But just from a CI/CD perspective, what are some of these challenges that enterprises face?

Forrest: I think any time you're in a large organization where you're contending with multiple teams, teams that have different priorities, different technology stacks that they're comfortable with, the biggest challenge you face is just that. Despite the temptation, there's really not a one-size-fits-all solution that you can impose. So we see a lot of these central cloud teams now, and a lot of times they have a cute name, like the Cumulonimbus team or something like that. You can always spot them a mile away. You see these teams come in and they say, “Well, we gotta solve the CI/CD problem.” They've been given that problem to solve and understandably, it's a problem. Right? Code is not getting into production fast enough or it is, and there's all this shadow IT happening, right? So people want to find a way to get around that. They give this task to a central IT team. They come up with some great solution that involves CodePipeline or something like that, and they put it in production, and they say, “Hey, everyone, come use this.” And they're rather shocked to discover when no one uses it. And the reason for that is, as I said, is that you've got teams that all want to do things their own way and they view the centrally imposed CI/CD standards is being just another thing that are stopping them from getting code to production quickly, which, of course, is exactly the opposite of what CI/CD is supposed to do. And so what you have happened then is the shadow IT problem just gets worse and your development teams continue to work around your central team, and you, frankly, wind up with a worse problem than he had in the first place. So when you're working with an enterprise then, the CI/CD challenge is how do you put something together that solves the centralization problems you have, the governance problems where you really do need some amount of rigor around who can deploy to production, when it can happen and what those deployments look like. And then you marry that to the also very real necessity for dev teams to be able to accomplish their work the way they want to. So that's the problem that we try to solve.

Jeremy: Yeah, and I think you get this other problem where nobody knows who owns the process, right, especially if you have certain teams working on different components, you know that throw-it-over-the-wall sort of DevOps mentality in a fast-paced serverless environment, anyways doesn't really work anymore, right?

Forrest: That's exactly right. And it's not that it can't work. It's that it doesn't work the way people expect that it will, because they tend to think about these problems at the level of the team. Once you move into a large enterprise again, the moving pieces rapidly increase in number, and you've got to find a way to coordinate those people together.

Jeremy: So what is the typical, or maybe let's say, a basic CI/CD pipeline look like and maybe we can call this the “hello world” version, which I think most people read a blog post, they see this laid out, they know what that typical, or they maybe feel comfortable with that typical deployment process. But why doesn't that work in enterprises?

Forrest: Sure, the “hello world” pipeline, as you call it, would involve some combination of a few different steps. Obviously, you want to start with a source control repository. You want to push your code to source control. After that happens, you want to run some number of build tests on it. You want to run your unit tests. You want to do your linting things that happens statically on the code, perhaps static analysis if you're doing some security stuff. And then you move on from there to a deployment phase where you want to actually pick up that code and you want to push it out into some sort of non-production environment where you can test it and people have all kinds of different environments they like to involve there, whether it's dev, whether it's staging or QA, non-prod, whatever you want to call it and then at some point, that code will be validated. Ideally, you'll run some kind of smoke tests or blackbox tests, end-to-end tests on the deployed environment. Once that's out there, then you will want to deploy to production. And a lot of people like to have some kind of a gate, some kind of approval there, where someone has to manually give a thumbs up and say, “Yes, I agree. The automated tests look good and this is ready to go out.” That's the “hello world” idea I think people see that laid out, and they say this is going to be great. The challenges then that you run into with that in the enterprise, of course, you've got, as I said, different teams, but you've also got different branching strategies. You've got these problems where oh, I had a problem that I need to hot fix to production immediately and my boss is telling I simply don't have time to go through all this normal process. So how can I just circumvent all that, because the business reality is I need to push it out now. That's you know, all of the hypotheticals that you've put down on paper sound great until you are at three in the morning and you really need that IAM permission to be corrected all at once.

Jeremy: And I don't think that's a common case where somebody has to rush a hot fix out to production, is it? Come on now.

Forrest: No. Although, you know, I've seen cases where people put these branching strategies together where you go to master and then you have, like, staging branches and production branches above that and one of the challenges people have there, of course, is they want to go to hot fixes and they hot fixed a production. They've got to figure out a way to merge those branches back down onto each other, and you very quickly end up with a true version of your code, which is deployed in prod, and then several different half in-progress releases in different other environments, and it rapidly collapses in on itself, and no one knows of the state of any of the environments. And that, of course, is worse than what people had at the start.

Jeremy: Definitely. Alright, so let's talk about repositories for a second, because this is another thing. There's the age-old debate between mono repo versus multiple repos and so forth. And I've seen this a lot, and I'm sure you've seen this a lot, where sort of the strategy they choose, that the enterprise chooses, for their repositories in terms of, you know, multi versus mono repo, that can also be a challenge when developing a CI/CD pipeline as well.

Forrest: I think that's true. I don't see a lot of organizations going with a true mono repo. I mean, nobody's Google, and I think people understand that. So you don't see an entire organization that's putting literally everything in the same repository. So when people talk realistically about mono repo versus multi repo or micro repo, or whatever you want to say, they're typically thinking about it in the context of some smaller problem that they have their heads around. So for this particular application or service that I'm designing, should I have a single repository or should I be splitting this up somehow by pipeline or by microservice or component of my app? That's really what the question is, and I think maybe that terminology gets confusing. What I've come down on, I guess, and I've spent a lot of time trying to get my head around this, multi repos are great and it a lot of times makes sense to have a repo per pipeline if you have multiple pipelines out of a single repository that can get confusing and hard to manage. But where people struggle is they try to break that down too much, and what they've done is they've created dependencies that span multiple repositories. So if you repeatedly find yourself having to encompass two or three different repositories or more in a single feature that you're pushing out, you're probably creating problems for yourself that you could solve if you had all that stuff in one repository to begin with. So I see people doing things like trying to create these orchestration systems on top of multiple repositories, and on top of multiple pipelines, just to try to control when dependencies go out and that just, it's adding an extra level of complexity that I don't think needs to be there. You know, don't be afraid of git-merge, and merge conflicts. That's actually a protection the repository is providing for you, and when you get away from that, you have to start reinventing wheels that I don't think are necessary.

Jeremy: Yeah, I like that idea too about being very careful about having dependencies or too much coupling between multiple repos, because that's always that problem where you're launching into a new environment or into a staging environment or something and you need to launch each deployment, or each pipeline needs to run an order. Otherwise, the other ones won't be able to access bits of information from it. Alright, so what about feature branches? Because that's another thing that I think helps, especially when you go to a smaller repos — maybe not the micro repose, but packaging individual services — when you're launching different features to have those separate. What are you thoughts on feature branches?

Forrest: I'm a fan of feature branches. I think that in my experience, it's the pattern that is best suited for enabling teams at your standard development shop as opposed to something like trunk-based development. So when we talk about feature branches, we're talking about a flow that would look something like: I make a branch off of master or off of my stable branch. I check that out. I commit some amount of code that solves a problem for me. I test it. And then when I'm ready, I pull request that branch to master, goes through code review from the team. And then once that code is merged into master, at that point, it's off to the races and it's being sent through whatever staging QA accounts, and it's on its way to production, and that process hopefully happens very quickly. It's mostly automated, and it's not like it's sitting in limbo off of master for days, but it's not necessarily an environment where you're creating a lot of bespoke release branches off of master. It's very much once I merge this feature, and then I'm ready to go to prod, or I should be ready to go to prod if the tests pass. I’ve seen that work well for teams because it lets different developers collaborate and it also lets you, if you need to, to make those hot fixes directly to master without a lot of additional fooling around and without worrying about having you know do they also need to be merged into existing releases that I'm adding on to. It just seems to work well for a lot of especially small-to-medium sized teams when they understand their code base pretty well.

Jeremy: Yeah, and that's also, you know, sort of one of those processes where when you're developing using feature branches, it's good to break up, have very modular code in there so that you're not always overwriting the index.js file or some some main handler file where you've got some layers of code in there where things can be manipulated without constantly having those conflicts, like you said when you’re merging multiple branches. So then, before we move on to the next thing, I just wanted to talk quickly about some of the common pitfalls with CI/CD, because I know I mean, I've seen this before, especially when you have multiple teams that are merging into some sort of centralized repository or centralized, you know, process that you get these QA teams that are constantly getting updates and is there a way to sort of mitigate against that? So as you have all these different teams that the QA team should kind of stay on top of what they're testing, as opposed to always getting new updates.

Forrest: Yeah, so you're talking about, like release cadences and things like that, and every team is different here and their process is different, and a lot depends on the maturity of where they are and what they already have from a testing perspective. I think that it it does make sense to have a single staging environment where things come together before prod. But it's important to not mistake that environment for production. No matter how similar you make it, and how automated it is, you won't really know how the app’s performing until you get it to prod and can test it there. I'm a big fan of, you know, blue-green releases for that reason, and rolling things out that aren't live yet, but that are in the production environment so that you can run tests on the infrastructure that's going to be serving your users before they've actually had a chance to see it. And you can get obviously more complicated with that, with feature flags and and canary releases as well. But that's a whole discipline that I think a lot of teams are just not ready for yet. It's not something I'm seeing in most places. There's a lot of work to do before we get there, where teams are comfortable with that.

Jeremy: Alright, Cool. Alright, so speaking about all this work to do, you wrote a very good post earlier this year called Enterprise CI/CD on AWS: a pragmatic approach and I really liked it. Thought it was really interesting how you sort of split these things up. So let's talk about that. Let's talk about this pragmatic approach to CI/CD. And in there you mentioned, you know, you run into this Conway's law problem where every organization is different. Every organization does things differently. So maybe we can talk about that. What is your pragmatic approach?

Forrest: Sure. So at some point, designing a process like this for a large enough enterprise is almost entirely an organizational problem and a human problem rather than a technical problem. The technical challenge is relatively easy to solve with the tooling we have today. I can set up a pipeline and I can set up approval gates, no problem, and I can figure out a way to pass artifacts back and forth. Big deal. It's been done a million times. But when I have a whole bunch of different teams that have different technology stacks, as I said early on in the show, that all have their own processes for how they like to get to production. But I also have some need for central governance and a need to control who actually has access to deploy things in prod, it gets more challenging. And the way that seems to work for a lot of folks to tackle this is to split build and release. So when I say split build-and-release, I mean you're taking the build piece of your pipeline, which a lot of times maps onto the CI that continuous integration half of CI/CD. That would be something that typically is very tightly controlled by an individual dev team. So they would be using a source control repository of their choice. Could be GitHub. Could be CodeCommit — probably not CodeCommit. Let's be honest, but I suppose it could be. And, you know, whatever they're choosing there, that'll be hooked into a build pipeline of their choice, which could be Jenkins-based, could be CodePipeline, GitLab, who knows and then that will run build tests that they write, unit tests they write. It'll run their code coverage stuff. They'll probably have to hit certain standards that are set by the organization, in terms of hey, you have to have this percent of code coverage. You have to have this sort of security scanning that's running on your code, but they get to choose how they how they implement that and what kind of pipeline they put it in. And I find that most dev teams now you know, have that capacity or want to have that capacity, to be able to put something like that together. The output of that build pipeline, that build process, is, quite simply put, one or more deployment artifacts. It's a deployable chunk that can be put out into production into other accounts, and that's something that could go into Artifactory, could go into S3, it just needs to go somewhere where you're release pipeline can access it. The release pipeline then is the piece that makes more sense to be owned by a central team and the central team then would take care of actually doing the code promotion across accounts. So their pipeline will be the one that spans, you know, not just the development account, but staging production, QA, whatever. They would have the permissions to deploy into production. They would have the approval gates on there, they would have the timing in terms of whether there needs to be some kind of a maintenance window before something gets rolled out. They can own all of that. Is this you know, the healthiest possible paradigm in a perfect world, right, where you have build-and-release split that way instead of the dev team owning things themselves all the way to production? Probably not. But this is actually the way most large enterprises function and will function for the foreseeable future is that there's got to be a little bit more governance around who could do what. And it also takes some of that need to juggle all those accounts away from the individual dev team. They can focus on what they do best. Your SRE team can focus on what they do best, and this seems to work fairly well.

Jeremy: So then, for something like feature pipelines would that be on the build side? Would you deploy your own feature pipelines there? Or is that something that also goes to the release side?

Forrest: Yeah, that's a great question. And that would definitely be on the build side. You would want them to be able to spin up and test their own infrastructure before the code gets merged to the master branch, again envisioning this feature-to-master kind of flow for your source control. So you can have your central team help thereby providing some scaffolding, possibly even something that gets vended out at account creation time. So having really good process around how you create another AWS accounts and how you see them with tooling and processes for people is a really big part of this. But yeah, having that ability for people to push code and have it spin up a feature branch should be in control of the dev team.

Jeremy: All right. And you mentioned security like some static analysis and things like that. But really, where does the security fall? Does it sort of fall on both sides here?

Forrest: Yes, absolutely. And that's where I mentioned, even if the dev team has some control over what tooling they're running there, and, of course, the dev team, they're writing the code. They know better than anybody, you know, what they should be looking for when it comes to certain security things. There do need to be standards across the org and there needs to be communication in terms of what's getting looked at. So from an AWS perspective, you know things like stars and IAM policies, there needs to be policies around “hey, eyeballs are on this” and you know, that if we're seeing the show up in a code review, there needs to be a conversation happening. And that's just something that you work with the organization over time to develop that excellence.

Jeremy: Alright, so what about testing? So you mentioned unit testing and linting on the build side. But is there more testing that would be done sort of on the release side as well?

Forrest: Absolutely. Big fan of any kind of automated smoke testing, blackbox, end-to-end testing you can do. In fact, I think you can't go to production without it. So when that infrastructure is deployed, whether it's QA, staging, prod — even, and especially in prod — you want to have tests that run as part of your release pipeline immediately following that that are doing deep health checks, that are, you know, provisioning test records, making sure you can get all the way back to the database and back, things like that. And you know that's just a no-brainer. The challenge there, too, is those tests, you want them to be run and deployed by your release team in these various accounts, but you actually want them to be written by your dev team because they're the only ones who actually know how the application’s supposed to behave. So I built things in the past where you'll have a standard, an interface that's set up between your build-and-release team, where the dev team will, for example, let's say you're working in a CodeBuild CodePipeline type of an environment, your dev team will go ahead and write some scripts that launch their end-to-end tests. So it could be, for example, Cyprus if you're using that tool for end-to-end testing, and they will go ahead and create build specs that call their Cyprus test, and those will be stored in a known location in the repository. So by the time your release artifacts get out to release, the release pipelines, which are standardized, will know to go look in that location, find whatever test files have been defined, and run those end-to-end tests.

Jeremy: So you mentioned approval gates in there, and this is something — so again, I think we are talking about continuous delivery, not continuous deployment, when we talk about CD in this context anyway. So I don't know how many organizations are confident enough to just automatically, once they pass the automated test, to go ahead and have those deployed to production. So obviously, we're putting in approval gates here, and that's a human function that needs to happen. So what about multiple gates though? That seems to be something that people look for.

Forrest: I think more people think they want that than actually want that and what you're talking about I think is something like a nuclear key approval, you know. So envision the two guys in the nuclear command center, each of whom has their hand on one key and both keys out returned for the missile to launch. People sometimes think they want that for a software deployment process where you need tWo individual people who probably belonged to separate IAM approval groups or whatever to be able to approve something — and you can certainly implement that in CodePipeline. I think I've got a blog post out there somewhere that explains how you can do it — but over time, that seems to be a little bit frustrating for folks. It actually does slow them down a bit. And what seems to be better in most of those cases is to have a trusted group of people who can approve, but then also just to maintain really good auditing. So if something does go wrong, you know what happened and you can go back and address it.

Jeremy: So I also see that sort of being helpful for escalating a problem, because we would only fail a build or we would only not approve something if there was something seriously wrong with it, right? So is escalation another sort of reason to have these approval gates?

Forrest: I think so. Escalation is another big topic, and we could probably spend a whole podcast talking about that. In general, and especially in these big org, where you just have huge multiplications of AWS accounts now being vended out through AWS organizations, or there's even people who've outgrown the capability of AWS organizations, how do you keep track of all these accounts? How do you figure out who can be escalated to access different functionality inside of them, you know? And is there such a thing as a central team that has to have eyes on all of that? Or can you delegate that permission to trusted people inside of the individual dev teams? That's a question that more and more folks are starting to answer. There absolutely are ways you could do that with escalation. We actually do that here at Trek10. We've got a nice little Slackbot that lets us temporarily attach policies to our roles, to be able to deploy things in production for various accounts. But I think that's a little bit of a separate topic.

Jeremy: Okay. Alright, well, so let's move on to tooling, right. So you mentioned, you know, obviously the CodeBuild and different code repositories, GitHub, Bitbucket, things like that. I mean, obviously that's part of it, having that. But then CodeBuild and all these other tools that do it. So first, let's start with this. Can we use standard build tools to deploy our serverless applications, like the Travis and Circle CI and Jenkins and those sort of things?

Forrest: Absolutely can, and people do it every day. However, if you are really getting into the serverless mindset, which is, I don't want to maintain infrastructure that's not directly relevant to me in my business, it would seem that a massive piece of infrastructure that will be nice to not have to worry about anymore is your build server. I think all of us have spent time troubleshooting a giant Jenkins instance that's been running at a disk space because of one rogue job that's affecting everything else, right? That's not a problem we should have to have in a serverless world. We want to be using tools like CodeBuild, which reliably gives me an ephemeral build job, and it doesn't affect any other jobs that are running at the same time. And if it goes away, I just spin up another one and I only pay for what I consume, right? Those are all great serverless fundamentals that we want to be working with. So, yeah, you can provision any old build server and launch your serverless apps. But I think it's much better if you can, to try to find a way to deploy your serverless app serverlessly. And there's more and more tools now that are letting you do that.

Jeremy: And so what about for people that are starting to move to serverless or are starting to migrate parts of their application? You know, would you suggest that they rebuild their existing CI/CD pipelines using these tools? I mean, I would think that some of these legacy tools would sort of be slowing down the process.

Forrest: They can. And I do recommend that wherever possible, it's ultimately just a really long process though usually. It's interesting. We see a lot of organizations now that are a few years into their cloud journey. But for them that means that they have five-year-old tooling that in some respects is sort of legacy now. So they’re that first wave of cloud folks and perhaps they've developed this large multifarious Jenkins instance that’s serving an awful lot of cloud teams and they're outgrowing that and they're running into all kinds of problems with it. But at the same time, it took a lot of years to get people actually onboard and using that and using probably shared tools and libraries and plugins that are accessible to everybody. So it's no joke. You can't just unwind that overnight, but it can be done and part of the way that can be done, as I was saying earlier, is trying to hand some of that control back to the individual teams, so you don't make a central team a bottleneck for that in the first place.

Jeremy: Yeah, so what about some of these tools that are coming out that are specific for serverless? So there's seed.run, you know, something like ZEIT or Architect, and I think Begin is part of the Architect framework, and serverless.com now has their own. I mean, I like these tools because I think they're great for small jobs, but what are your thoughts? Do they fit into the enterprise world?

Forrest: So you know, a lot of these tools are early on in the process, and I don't want to speak individually to them right now. Some of them I've seen do great things. I'm excited about some of what Begin seems to be offering. I don't think it's quite released yet. I think the important lesson to take away from those tools right now because they're so young and so early — so who knows where exactly they'll go on an individual basis — but I think the important thing to take away is that this is obviously a big problem that has not been adequately solved yet, at least not at the 80% level and that's why some of these tools are starting to proliferate because people very rightly are saying, “Oh, it's challenging. It's difficult to be able to put together a reasonable CI/CD pipeline that works for my developers and works for serverless without a lot of frustration.” So they're trying to automate some of that away, and we'll see how that goes over the next few years if more of those features are just available natively in the cloud providers or if one of these tools really takes off, but I think we would be foolish to dismiss the obvious gap that there currently exists for developers to be able to do this more easily,

Jeremy: Right, and so you mean basically, I don't think any organization, no matter what we build or what someone builds, there's gonna be a turnkey solution for enterprises.

Forrest: No, I mean, what do they say about enterprise software? It's solving the problems of today by creating the problems of tomorrow. I think that's probably true for any kind of CI/CD system as well, whether you build it yourself or whether you buy it from somebody, but it's just because it's an organizational thing at that level, as I said, it's always going to be growing and changing. There will always be interesting new challenges to solve.

Jeremy: Alright, so another tool that is out there was built by you and Trek10 which is this Quick Start CI/CD for AWS. So I'm not sure this is the best thing to walk through step-by-step on a podcast if someone's listening in their car, or mowing their lawn. This might not be the easiest thing to digest, but maybe we could just go over it quickly because I do like how it's set up. I do like what it does, all the different features that provides. So do you want to just kind of walk us through it at a high level?

Forrest: Yeah, let me give you the elevator pitch for it. Maybe that will help. So this grew out of a couple of different projects. Well, really, I guess multiple projects we've been doing for various clients at Trek10. And we were doing exactly what I've been describing over the course of this podcast, which is going at enterprises and helping them figure out how to set up that feature branch-based workflow and how to get serverless apps deployed using these ephemeral services like CodeBuild and get them pushed out across multiple AWS accounts. And we were solving some of the same problems over and over again and had all this CloudFormation that we were using. And so we said, “Hey, let's let's spruce this up and let's make it available for everybody.” And that's what the Quick Start is. Some of the things that the Quick Start does I think probably won't necessarily be needed long term as AWS continues to roll out more features. But right now a lot of this you can't do natively today, so you do need some additional CloudFormation to come along and help you with that. The biggest thing there is we really wanted to enable people to dynamically create feature branch deployments. So as I was, I think saying earlier, when I push code to a feature branch as a developer with a serverless app on AWS, I want that feature branch to automatically spin up for me a deployment of my app that's namespaced according to my branch; it doesn't conflict with anybody else. So if I need to, I could have 20 developers all on the same account, all deploying. Nobody steps on anybody else. Harder to do in the serverless world than you would think, but at the same time, it's also very cost effective because you're not paying for that infrastructure unless you're actively using it. So we built some Lambda, and some SNS that spawns those dynamic pipelines for you and you can check that all out. It's open source. There is also the ability for you to promote those things across accounts and you know it's basic. It's got the full end-to-end life cycle in there. No doubt as you work through this, if you decide to work through this, you're going to find things you'll want to tweak to your use case. There's the concept in there of a shared services account, so we actually take the pipeline infrastructure. We segregate that off into its own AWS environment, and then it reaches out and passes roles cross account to be able to deploy in dev, and staging and prod. And that's a very best practice thing to do. Some people want to actually create multiple shared services accounts, or sometimes they call them tools accounts, one that has production access and one that does not. But ultimately, you know, you can play with that. The goal is you want to make it easy for your devs to deploy and spawning those pipelines on their behalf seems to be a good way to do that.

Jeremy: Now, when you're building like a development version or something that's launched to a different environment, obviously with serverless, we usually have some naming conventions in there for functions and for DynamoDB tables and things like that. So are you generating, are you rebuilding the artifacts at every step here, or is there some way that you're creating artifacts that can then be promoted?

Forrest: In this case, and this is me casting my mind back a few months, so apologies for any inaccuracies here, but I think in this case, we did end up with a separate set of artifacts for the feature branches. And then we rebuilt the artifacts when you push to master but then that same set of artifacts is used across all the accounts you promote once you merge code into the master branch. And definitely, if you split up your tools accounts between dev and your prod workflow, you would have two sets of artifacts that were generated. I don't know that that's necessarily a big deal, but if it is, of course, you'd want to go with a different approach.

Jeremy: Okay, great. Alright. So maybe we can wrap this up by, you know, giving some advice. So for the enterprise that is looking, or is having some trouble with CI/CD or looking to go serverless — I'm asking you multiple questions here, but let's say for the enterprise looking to go serverless, has some question about their CI/CD process, besides calling Trek10, what would be your best advice?

Forrest: That's a great question. So obviously call Trek10, but even before you do that, check out the Quick Start. You can play with that yourself. It'll give you a good feel, I think, for how far you can go on your own. I definitely would recommend trying to use some of the AWS native services if you are in AWS. If you get to where you need something that is a little bit more flexible or something that enables you to work with additional tools that are not natively integrated with the AWS services yet, at that point I would suggest looking into that build release split that I discussed earlier and that you can find in that pragmatic enterprise CI/CD post.

Jeremy: Awesome. Alright, well, I will get that stuff into the show notes, but anyways, listen Thank you. This was awesome. This is certainly a complex topic that we probably could go into detail about and spend hours talking about. But anyways, why don't you tell the listeners how they can find out more about you?

Forrest: Sure. Well, I'm pretty easy to find. You can find me on Twitter @forrestbrazeal. Also on LinkedIn. And of course, I do write a number of things for the Trek10 blog. I have a podcast there called Think FaaS — functions as a service. We've been a little bit lax about putting out episodes of that recently, but we are doing a bunch of live episodes of Think FaaS next month at ServerlessConf in New York City. Jeremy, I think you're speaking as well.

Jeremy: I am, yes.

Forrest: Yeah, you got a full size talk. They're not a Think FaaS talk, but yeah, we'll be putting those out on the podcast here in the coming months. I am co-chairing the conference this year. So if you have not signed up for a ServerlessConf yet, please do that. I would love to see you there. I think It's going to be a great couple days and you can hear Jeremy speak on — what are you speaking on, Jeremy?

Jeremy: I am speaking on building resilient serverless systems with non-serverless components.

Forrest: There you go. You don't want to miss that.

Jeremy: Alright, Awesome. Alright. Thanks again, Forrest.

Forrest: Absolutely. Thanks, Jeremy. It was a blast.

View Details

About Efi Merdler-Kravitz

Efi is a software expert and currently the R&D director at Lumigo. Over the last 12 years, he has been working as a developer, team leader, group manager and director in the healthcare, mobile, security and agriculture industries. Recently, Efi has been working on developing serverless applications and building tools to make serverless easier.

  • Twitter: @tserverless
  • Lumigo: https://lumigo.io/
  • Email: efi@lumigo.io
  • Lumigo Twitter: @lumigo

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Efi Merdler-Kravitz. Hey, Efi. Thanks for joining me.

Efi: Hey Jeremy. Thanks for having me.

Jeremy: So you are the R&D director at Lumigo. So why don't you tell the listeners a little bit about yourself and background and then what Lumigo is up to?

Efi: So as you said, I'm leading the R&D of Lumigo and I've been working on pure serverless applications for the last 2.5-3 years. And on a personal note, I think it's the best technology decision that I ever made. So a couple of words about Lumigo. Lumigo is a SaaS platform for serverless monitoring and troubleshooting. Basically, Lumigo connects to your AWS accounts and alerts when things go wrong and then tells you the entire story of the requests that lead to that issue so you can quickly get the root cause. Let me elaborate a little bit. When you break your application to small pieces following the microservice architecture, it becomes very hard to debug your application, by the way, both in production and in your local dev environment, and it becomes especially hard when using async components like SNS, SQS, Kinesis, etc. And as someone who worked with serverless extensively before, we understood the challenge here at Lumigo, and therefore we developed the platform to help developers like us to understand the environment quickly when something goes wrong.

Jeremy: Awesome. Alright, so we've had a couple of shows so far where we've talked about observability, and we've kind of gotten into all of that sort of stuff that Lumigo, I think, as a product does, which is really interesting. But I actually want to talk to you about this idea of managing a serverless engineering team because one thing that's kind of unique, I think about Lumigo is even other companies that are working on serverless products, they're not entirely serverless themselves. And Lumigo is. You pretty much manage an entire serverless engineering team, right?

Efi: Yeah, we are 100% serverless from deployment, packaging, monitoring. Everything is serverless. We don't use any physical or virtual servers in our back.

Jeremy: Awesome. That's so cool. So alright. So you’re a manager. You've been doing this for a very long time. You've been managing engineering teams. And so I really want to get into this idea of what’s sort of different about managing a serverless engineering team versus managing a traditional engineering team. And I know maybe some people are thinking, “Well, you know what's the difference?” But I think there is. And I think you think there are some differences. So maybe we start first by what we have to do to sort of move our team to serverless. So if we've got, for an established organization, you know, it is great to be a greenfield startup and be exploring new things. But most companies are not. Most companies are established companies with legacy systems and so forth. So how do we go from you know this idea of taking a team that's used to working with all these different services and EC2 and containers or whatever, and moving them to something that's a lot more serverless. So maybe we start with that. What's the first step to getting people to move teams to serverless?

Efi: Great question, Jeremy. So I think that the best advice to any new beginning, especially in the technology world, is to start small. Try to taste the technology before jumping headfirst. So what does it mean in our case? First of all, I think you should ignore buzzwords. You hear a lot about new cool services. For example, AWS have dozens of ways to save data, where you have the ability to run machine learning on it. Try to use simple, trusted, and well-documented services I think services like API Gateway, DynamoDB, S3, Lambda, of course. These are the services that are the building blocks of any serverless application in AWS, and there's a good chance that you use at least one of them in the final solution. Try to use the simple version of services. What do I mean by simple? For example, in AWS, you have six or seven services that provide queuing capabilities. There's a good chance that choosing SQS, at least in the beginning, when you start is good enough. Avoid more complicated services like Kinesis. And you know, in the end, what they say, no one was ever fired for choosing SQS. So I think it's a good choice. And read and learn. Many people think that serverless identical to previous technologies that they use. And sometimes they forget that serverless is not only a new technology but also a different way to approach development. There are many great blog posts, newsletters, and, of course, for specific AWS services, read the AWS documentation. They are a great source of information. Choose a good framework that will help you with the transition. There are many good ones like AWS SAM, the serverless framework, Chalice or Zappa if you work for Python, and don't forget other practices that you used in the past. If you used .Node and Express the past, then AWS released, I think a couple of years ago, a framework called AWS Serverless Express. It's a framework that was released by AWS to help you in the transition. So if you are familiar with Node and Express, it will help you in moving to the serverless world. For example, in the Python world, if you use Django or Flask, then you can use Zappa, which I think is a great choice for the transition. So you don't need to change your methodologies, at least in the beginning because eventually I think, that these frameworks, for example, the framework AWS provides, and the framework that Zappa provides in the end, don't provide the code that should be running in the serverless environment. But at least in the beginning, they remove a lot of overhead from your head.

Jeremy: Right? And so what about the transition? I mean, do you just transition the whole team right away? Or do you pick like a point person to do something like that?

Efi: I think that especially serverless, which is something that is just very new, I think you should lead the way as a manager. As I said before, serverless is not only technology. It also shapes the way you develop software. Therefore, you as the leader has the service responsibility of how your software team works and delivers. You need to master the tools of the trade. So before giving it to anyone in your team, make sure you sit down and learn it on your own. Again, as I said earlier, read the blogs, do the tutorials, read as much information as possible before moving the entire team or transitioning the entire team.

Jeremy: I think also, it's one of those benefits of being an engineering manager where you get to try out some of that cool stuff first, before you let anybody else use it.

Efi: Exactly, exactly. You know, most of the time you do the boring stuff, and now you have a chance to start something new.

Jeremy: Exactly. So how do you then introduce it to the team? So if you as a manager are going through and learning some of the basics here, I mean, obviously you can lead the way, but this is why we have teams, right? Because teams can really go deep and and start implementing those things which you might not be able to do with your busy meeting schedule. Right? As an engineering manager. So what's the best way? How do we introduce that to your team once you sort of feel comfortable with it?

Efi: I think that the best thing about service is that it’s a cool technology, it’s a new technology, and I think the moment you introduce it that way to your team, it will make your life a lot easier. So make sure it looks cool. Show that it saves time in deployment. Show that it's very easy to configure new components and show them that you can easily and quickly deliver new code. And in the end, you know, developers hate configuration and the moment you save them the hassle in configuring new Dockers and configuring new services, but all they have to do is just write the code and let it run magically. The moment they'll see it, you know they get hooked up immediately.

Jeremy: How do you impart that knowledge on to your team? Do you do formal training or is this something that you just encourage they explore on their own, for example?

Efi: Yeah, it's a good question, and I think it's a general question. I don't think it's related only to serverless. So what's the best way to learn new technology? So in the end, it's a personal taste, and I think each of the managers and developers choose their own path. So my own personal taste is to do it together, planning together. First of all, I think it's great for bonding. You know, developers usually walk alone, knowing their own environment on their own laptop, putting headphones on the ears, listen to music and most of the time they don't interact with each other. So it's a good way to bond, you know, to talk with each other. It makes it very easy to ask questions. And the best of all, I don't think you're interfering to each other because you're learning together. So it's not like you are learning right now, then someone in the middle of you know, off his own task, you ask him a question, and you need to stop him on what he is doing right now. I think that two good resources that I recommend on learning. So, first of all, there is the official AWS tutorial on serverless. I can share the link at the end of the recording. And I think there's a very good tutorial by serverless-stack.com, a tutorial on how to use serverless code-wise and how to use the various tools. And each of my developers, you and the team, the day they arrive, must actually go over these two tutorials before they actually start code for the first time.

Jeremy: So when you do that, it sounds like it's sort of a combination of both, you know, that you would let them sort of go through some of these tutorials. But is that something you do kind of get everybody together and do like a formal training though?

Efi: Yeah, and again, I think it depends when a new developer arrives or when you as a team begins to learn serverless. So I think when you as a team, as a whole team starts to begin to learn serverless, then I think it's better to make it formal, see together and learn together. But after the team gathers up enough information, enough knowledge and now a new developer joins the team, then you won’t gather the entire team again. And then you know you have a set of links that each developer can go over alone.

Jeremy: Alright? And in terms of especially like with the new developer coming on, or just I guess, the team learning in general you know, sharing information between the team. Like how do you recommend doing that? I mean, do you still put stuff in Wikis? Or should you be using something more real-time?

Efi: Actually, that’s a good question. By the way, I just want to add a lot of people forget that many developers jump from on-premise development to serverless development, so there's also another gap that they need to learn or to jump over. This is not only serverless, but also cloud computing. So to many of the developers who have joined my team, one of the first things that they do is also play with the AWS environments on its own. Create new resources, you know, create an EC2 instance, many of the developers see it for the first time. So remember that. So jumping straight to serverless, sometimes it's not the best way. Make sure to move along the path that will allow developers to easily transition to the cloud computing. Now, regarding your question, I think that sharing is very important and always share, you know, we here at Lumigo, we use Slack or in your case, any other tools that you prefer, and we have a weekly meeting where all the developers gather together and we define the agenda before the meeting. And each developer shows a new serverless material that he’s learned in this week. So serverless, it’s a great — there are a lot of new things to learn in serverless on a weekly basis. There's a very low chance that we'll repeat the knowledge that you’ve learned from previous weeks.

Jeremy: Right, so you would repeat a lot of that stuff. But what about sort of capturing that? Do you do internal documentation for that? Like using a Wiki or something?

Efi: No, we have only Wiki online to public material on AWS or to any other good tutorials that we find. But we don't have any Wiki on, you know, on the new stuff that the new developer learned last week, mainly because serverless is such a dynamic field, a dynamic technology. So writing something in a Wiki will make it stale in a matter of a couple of months.

Jeremy: Are you telling me that my Wiki page on how to install SQL Server 2000 is out of date?

Efi: Yeah, unfortunately.

Jeremy: So what about best practices, then? I mean, so if you're not, I mean, it's kind of hard with serverless because you throw the term “best practices” out there, and I like to think about it to say I'm not sure if they're are best practices, just anti-patterns right now. Like things we try to avoid. You know, so codifying those, I think, is somewhat important, at least between the team, so everybody's doing things the same way. So how do you do that? Do you establish best practices for a team?

Efi: Yeah, it's a great question, and I think you can also divide it into best practices in actual code. What's the best practice in using DynamoDB, or S3, and the best practices to behave as a team, as a serverless team. So I think serverless promotes end-to-end process, which means that developers takes full responsibility from the design phase after the production phase. So, for example, in Lumigo, developers are responsible to the product. You know, we are eating our own dog food and you know what bothers us the most, so we're the best candidates to develop our own platform. And developers are responsible to the quality, making sure everything works as expected, and we're putting a lot of emphasis on automatic testing in the deployment. One of the, I think, one of the benefits of serverless is agility, but you need a set of tools to help you get the most from serverless. Just using DynamoDB or just using Lambda is not enough. You need all the tools behind the scenes that will help you make the most out of it. And I think that developers, in the end, are also responsible to production and making sure that what they develop works as expected and is being used by our customers.

Jeremy: Right, yeah. And I actually really like that philosophy, too, where I feel like when you’re developer or when you were a developer the past, you'd write a snippet of code, you know, maybe you would write the test for it. Maybe there's somebody else writing a test for it, but then it would go to QA. Someone would test it. You throw it over the wall, some Ops person maybe puts it into production, and then maybe at some point it comes back to you. But I like that full ownership of it because one, you can see things in production very quickly. And as long as you follow good practices for security and things like that I think it's a really, really good way to keep developers motivated too to come to sort of see and own, you know, that entire process. So I do like that. Alright, so let's say we've done this now. We've introduced our team to it. We’ve got a good knowledge base going. We’re sharing information, you know, we're writing serverless applications. So now how do we run this team on a regular day-to-day basis? And maybe let's start, actually with growing the team, because this is sort of those funny things you see where someone's looking for a serverless developer with 10 years experience, which is kind of hard because service hasn't been around that long. So where do we start? Who do we look for when hiring for serverless teams?

Efi: Yeah, yeah, it's a good question, and as you know, as a developer, if you see someone post the job with 10 years as a requirement, that means it is not serious. So that's a great question. And I think it's very hard to find someone with a lot of serverless experience, and it really resembles the early days of mobile app development. Now, I believe that in the future a lot of developers will have this experience, but right now they don't. But I think that in the end, a good developer is a good developer. It doesn't matter what technology they use. But I think that experience in the technology, in the methodologies, that enable us to use serverless to serverless are a good plus. So I think things like experience in the cloud that you're using, either it’s AWS or Azure or Google Cloud. I think that experience using serverless components like S3 or DynamoDB. I know many developers that I don't know, wrote code that runs on EC2 but broke quite a lot with S3. So I think it's a very good experience. I think knowing agile processes, like a continuous integration continuous delivery because serverless is very agile and promotes end-to-end the ownership, the knowing how the entire process works is very important. And I think automated testing. A developer in you know, in our era needs to know how to write tests, and needs to know how to write them good.

Jeremy: Right. Yeah. And I think you know, the other thing is that beyond just knowing how S3 works, for me anyways, I've always been looking for people who understand distributed systems or at least have some knowledge around you know what happens when something fails and things like that. What are your thoughts on that?

Efi: I think it's a good point. I think you know, that's one of the ways that I personally test developers that come to the team. I think that a good exercise is to let them to design distributed system. For example, try to design a system that mimics the mechanism of Lambda. So say you're on the Lambda team, how would you design and build something that's scales indefinitely? Or a system that runs multiple process, that tries to deliver a message from one place to the other? I think that serverless is about scaling. So you also need a developer to think about how to scale the design, and how to scale the system that they built.

Jeremy: Right, and one of the things that I found interviewing, especially people graduating from college, is that they don't have a lot of distributed systems experience. And one of the questions I tend to ask them is, “Alright, what happens when your database isn't big enough anymore?” And usually the common answer is “We'll just get a bigger database,” right? Like, but eventually, if you can get them to answer the question, well, maybe I could put some of my data on this database and some of my data on this other database and spread the data, right? If they can start thinking about sharding and things like that, I think that's really interesting. And then also a lot of people — and this is really scary, in a sense. I mean, I'm not sure what they're teaching it in some of these colleges — but you know, this idea of like even connecting to an API, it's always happy path, I think, is what you get from a lot of developers. And so if you say to them well, what happens if the API doesn't respond, right? And if they say, “Well, try it again.” Okay, that's the first step, and “Try it again?” Alright, what happens when you can't keep trying again? What's that failover? How do you build in that resiliency or at least thinking about? And I think that's always a good sign if people can come back to you, give you answers like that, then I think that's really interesting,

Efi: Actually, that that's a very good point. And I think that in the end, developers don't need to mention the right terms. So even if they don't say “shards” but they mean, okay, like you said, “Hey, split the data between various databases,” it shows thinking in the right the direction.

Jeremy: Right, and curious people, I think make the best programmers — people that are willing to just, they want to learn something new. So all right, let’s move on a little bit past now we've hired. Let's say we can hire some good people. There’s probably some training in there and so forth. We kind of talked about that. But what about you know, this day-to-day work in serverless? Is there, let's say we're working with cross-functional teams. That's something that's very popular in agile environments now. You have a product manager. Do things like the granularity of user stories change? I mean, now that we're building much smaller components, do we need to get that detailed and say we need five functions as opposed to we just need to solve this user story. I mean, is there a difference there?

Efi: No, no, I think it's the same. I think the moment you move to microservice architecture, I think the way you think of all the products stays the same. Nothing changes.

Jeremy: Alright, and that I think that's probably music to product managers’ ears, right? So that they don't have to learn anything too new.

Efi: Job security.

Jeremy: Right. There you go. Alright. So what about tools? Because I mean Lumigo is a tool, obviously. But that's sort of a monitoring, you know, debugging sort of after, in-production tool type thing. But in terms of tools to help you build the services - you mentioned frameworks - like how do those come in?

Efi: Yeah, it's a good question, and I think, you know, first of all, when they teach us in college about encryption, they always tell us not to write our own encryption algorithm. Use something that is ready. So I think that the same thing applies to serverless tools. There are a lot of tools today that enable you to package, to upload, to deploy your code. You have tools today that help you to monitor, and debug. Use them. Don't write something on your own. Don't waste your time on it. And I think one of the first things that you need to learn is to learn tools like AWS CloudFormation or Terraform. These are the tools that enable you in the end, that these are the basic tools, these are the building blocks that enable any serverless packaging technology to deploy your code, to deploy your various sources. So no matter what serverless framework you choose, either the Serverless framework, or Chalice, in the end, behind the scenes, everyone is using either CloudFormation or Terraform. So I think it's very important to learn the best building blocks, and I think you need to learn how to automate your tools, automate your testing. So use good testing libraries like Pytest or Jest and there are many others that are very good. And also use serverless plugins to test some of your flows locally, like DynamoDB or API Gateway. I think that Bash scripts, or scripting, depends on the OS that you use...

Jeremy: So let me… Sorry to interrupt you, but I want to go back to the testing locally thing, and we can talk more about that. But the mocking libraries and so forth, and maybe we disagree on this, and that would be great. But I'm thinking, you can emulate DynamoDB locally and certainly the Serverless offline plugin that allows you to run, you know, the end points, I think is great locally because that we don't have to publish them, and you can make changes quickly, and that works really well. But I think interacting with some of the cloud native resources like a DynamoDB or an SNS and SQS, from a local development standpoint, I feel like it's better to interact with the cloud services of those, and I get unit tests, right? Doing stubs, you know, or some sort of mocking, maybe for unit tests, I think makes a lot of sense. But, I mean, you know, how far would you go with these local mocking libraries?

Efi: Yeah, that's a very good point. I think it's a painful point right now in serverless, in serverless testing. And I think the only thing that I can say right now is that testing locally just as you said won't give you the quality that you're expecting. In the end, local testing will give you a certain amount of validation on your code. But I think that the best way to increase your testing velocity is to give your developers that build it to run their code easily and fast in the cloud. That's the only way to actually test and make sure that the code that you wrote is working.

Jeremy: And are you a fan of giving each developer their own environment?

Efi: Yeah, in the end, that's what we're doing here in Lumigo. And again, I think it really depends on the size of the team because although serverless is supposed to be a pay-as-you-go. But there are various components in serverless, like Kinesis, that even if you didn't use them, you’re still paying them. So if it's a large thing, you need to think of a better way to control your costs. And I think you know, I think it's a different discussion, a discussion about costs. But in a smaller team, I think you can give to each team member its own environment and give each developer an ability to easily deploy the code to his environment.

Jeremy: Yeah, I love that model because then you just have, they're not messing up anything. They're not even messing up the dev environment, right? They’ve just their own sandbox that they can play around with. Alright, so let me go back for a second to the frameworks, right, because you said CloudFormation and TerraForm is usually what happens under the hood. And I totally agree with you on that. I think that's a good point, because what you have, even with the Serverless Framework or with SAM, which are the ones that I'm most familiar with, you know, if you want to create an SNS topic or a DynamoDB table or something like that, and include it in the SAM template.yaml or the serverless.yaml file, you're still writing just straight CloudFormation in order to make that happen. So even if you know how to do a function, you know, the functions in the events and some of those other things that are super handy and easy to do, you know, they've got short hands in those different frameworks. If you want to get a little bit more advanced, yeah, knowing that other stuff is absolutely imperative. And again, I'm monopolizing most of your time here, I know I'm doing a lot of talking, but the Bash script stuff is another thing where I know it's a really low level, or it seems low level, but there's so many things you can do with Bash scripts that a little tiny script here, especially as part of your CICD process or your testing process, you know, makes a huge difference to know that. So I'm totally in agreement with you on that. So did you have anything else on Bash scripts or?

Efi: I think that again, Bach scripts, or scripts in general, you know, the group. So you need to learn how to group things together. And the only way to do it is to learn scripting.

Jeremy: So, yeah, so then the other thing too in terms of tools — and I know there are some tools that do this now — but I guess this goes back to certainly CloudFormation, this would be very much so specific to AWS. I guess other services as well. But understanding the IAM permissions and roles and things like that because, you know, I just had a conversation with Hillel Solow about serverless security, and we were talking about the least privilege principle and things like that to make sure that we're not opening things up too much. But that's a tough thing to do to one, learn all of this stuff, but then to enforce it. So do you have a way or do you use tools to enforce IAM permissions?

Efi: Yeah, that's a good point. We don't have any automated tools to enforce it, but it's something that is very important to remember. Security is important. It doesn't matter if you use serverless again or any other development paradigm or any other technology, and I think it should be part of the development cycle. And when working with AWS, understanding how IAM roles work, I think it's crucial because otherwise the security will be partial at best. And I think that you know, the term that you mentioned, the least privileges, that it's something very important — an issue developer should learn as part of his welcome to the company. But again, just to show you that security is not only IAM, for example, in AWS in serverless when using S3, so making sure that your S3 buckets are not public, making sure that they are private. So again security is not only IAM. In Lumigo, the security phrase is being done through the code review. So during the code review, we have a checklist that needs to be passed. So one of the checklists is security. So we do ask the developer, “Why did you add this kind of IAM permission? Why are you using it? Can you reduce it to something lower?” So the developers need to answer “no” to answer these questions before they can actually take the code and merge it to production. But by the way, I think also it's a good time. I want to go with you over our flow here at Lumigo just to understand the development flow that we do here at Lumigo and how everything is working together. So we have our task in JIRA and the moment that developers set a task, he opens a branch in GitHub. We're using a GitHub flow, which means that the end, each branch that is merged to master actually is being deployed to production. And so the developer creates a branch, write the code, test the code either automatically overriding automation. He has to test the code on its own, on their own AWS environment. They do a pull request. They do a code review. There are a lot of automatic gating that we do. Again, the word automatic is very important here. And so things like linting, unit tests, integrated tests, static analysis, and if the code review passes then it merges to master and we have an automatic continuous integration service, specifically we use CircleCI and we push it to our monitor environment and then to production. And in the end, the developer self-monitors it through the Lumigo platform. And again, pay attention. It’s very important that the developer is responsible to the entire cycle, you know from the product, from writing the code, writing the testing and monitoring and production.

Jeremy: Awesome. And I really like the code review process that, you know, you have this sort of checklist of things that you have to do. Now, I know some companies that are very, very good with this. I know some startups that are not so good with this. So I know you guys are still a relatively small team, and that's great to enforce those those policies earlier because I do see those possibly breaking down because you still have that human element. And maybe in the future we'll have some better tools that will automate all of it for us. Okay, great. So that's awesome. And I think that the outline of what your process is is really, really helpful. So let's get you know, so now if we arc this story here, we moved our team to serverless. We've been running serverless for some time now. Now what about things like roles and specializations in serverless teams? Because some of the AWS services or services in Azure or Google Cloud, they could require a specialist in and of themselves, right? DynamoDB, designing DynamoDB tables, understanding Kinesis and some of how those things work. Athena. QuickSight. All of these tools that are very complex in and of themselves. Is that something that you want to start doing, is shifting some of the responsibility to individual users so that they can sort of go deep on one service and sort of broaden the knowledge for the entire team?

Efi: Yeah, I think it's a tough question. In the end, I think it depends on the size of the team. As a rule of thumb, I think that for small team, I don't know less than 10 developers — it's a ballpark doesn't have to be, it can be also 15 but it's on your personal feeling — everyone should know everything. I think that it's when the team is small, it’s a chance for the team to learn serverless together. You know, when the team is very big or when the team grows, it becomes very difficult to share knowledge, to gather together and to talk together about the various problems. So when the team is small, it's a good point in the life of the startup or in your team to create a core knowledge of serverless. So as the team gets bigger, so I think you need to start to specialize. And always remember, I think you need to always remember to have redundancy. You know, making sure that at least you have two developers that know how to use resources.

Jeremy: That's a really good point. Yeah, the redundancy aspect of it is huge. But sorry. Go ahead.

Efi: Yeah, and I think that some services just like you mentioned — DynamoDB, Kinesis — are difficult, difficult to master. And they change all the time. And I think you should share knowledge on these services with a couple of developers, even on small teams. So in small teams, most of the developers or all of the developers know how to use DynamoDB, let’s say, in the general level, how to write a DynamoDB, how to read from it, but, for example, how to design an index, and how many shards should I have for Kinesis? I think this is kind of speciality that, I don't know, maybe one or two people in your team should know as deep as possible and the others should ask them questions.

Jeremy: Yeah, yeah, that makes a lot of sense. Alright, so let's talk about maybe the day in the life of an engineering team. I mean, you mentioned your workflow, and you kind of gave us that, but does anything change in sort of the methodology, the way that we approach software development?

Efi: Yeah, I think it's, in the end, it’s very similar to teams that use microservices. Again, it's full ownership, product, code, testing, deployment, production. I think there's a very major change from, you know, from using serverless to using other microservice technologies, the cost. Think we started talking about it earlier. So people need to understand costs. Developers need to understand cost. It's part of their development site.

Jeremy: How important is cost, right? I mean, you take larger organizations. I know there's a lot of jokes about you know, my serverless infrastructure cost us $30 a month or something like that, but that applies if you're maybe using, you know, just Lambda and API gateway or something. But add Kinesis in there. Add DynamoDB. Start adding some of these other services and get some scale, right, and then all of a sudden cost is an important factor. And if you're calling the KMS API too many times or the Secrets Manager API, you know, things start to add up. So how much time and energy should developers be thinking about costs, or be spending thinking about costs?

Efi: I think that people, you know, people that come to serverless for the first time sometimes forget how easy it is to scale serverless. So in a matter of minutes, you can easily get a hundreds and thousands of Lambdas running simultaneously. Millions of requests to the DynamoDB. And in the end of the day, you suddenly see a bill of a couple of hundreds of dollars, and you ask yourself, “What?” So I think it's very important what you just mentioned. So, for example, in Lumigo, we have cost alerts in each of our environments. Both our dev environments, have cost alerts so developers know if they use their resources too much. I as a manager, I check the cost on a daily basis, and I'm trying to understand the trends. I use the Cost Explorer in AWS quite a lot. In addition, we also use our own tools. We have our own monitoring tools which also gave us a cost breakdown, and I think again, part of the code review is part of the checklist that I mentioned earlier. We ask the developers, why did they choose, for example, this amount of memory for this specific Lambda. Or why did they add another index to DynamoDB? Each index costs more money because you are duplicating the data. And for example, while they are using Kinesis and not Firehose. So there are many questions we ask along the way when doing the code review. Again, it's not something that can be done automatically, something that people need to see the code and understand what's going on. But you ask the questions in order to make sure that developers understand the trade-off, in order to understand that it costs a lot of money. And you know, especially for startups, where money's always tight, suddenly paying thousands of dollars per month, it's dangerous, can be really dangerous. So it's not only “Oh no, we'll use the corporate credit card.” It can be really dangerous for the startup. So you need to pay attention to it.

Jeremy: Yeah, and I think too that's one of the things that I know I always did. I mean, I've worked in many, many startups. So even before serverless, cost was always a factor in optimizing the solution and I think that is something you can do. Like you said, Kinesis versus Kinesis Data Firehose. Some of those things have variable costs as opposed to fixed costs for shards and some of these other things. So I think that's something you build in early. You don't need to worry about maybe premature optimization, but I do think that if you say well, look, this is going to be $5000 a month, and this is going to be $500 a month. If there's a way for you to see that up front which, in most cases, I think there are, you can do some good estimations. Yeah, that's definitely definitely something you should be paying attention to.

Efi: I agree.

Jeremy: Alright, so what about the overall responsibility of the team, right? So smaller teams, we talked about this a little bit more in the beginning, you know, Ops teams, security teams, all kinds of teams that do things in the modern cloud, and do things in modern organizations. Where does the overall responsibility change for serverless development teams?

Efi: Yeah, you know, I have never talked about security teams because I think that security again, you know, especially in today's world, with security so prominent, I don't think that developers can specialize in security. They need to know security. They need to know how they write the code by the things that — security needs someone who specializes in security. Now it depends on whether you're in a big corporation or in a startup, whether you want to hire someone who specializes in security or doing some kind of training for developers to be security specialists. So that's a different question. But I think that security is a different role. But you've mentioned also DevOps, and I’ve mentioned also on a previous note, the QA. And I think that DevOps and QA, today in the serverless, are actually one role. It's developer, DevOps and QA is the same person, is the same developer who is doing everything. And I think that in the end it produces a better product because it's a developer. A developer knows how to test his code. He knows how to write the testing in order to think about all the various edge cases that might appear, either the developer or doing the code review with the other developers. But I mean the developers themselves, and not someone who is external to the development process. The same thing about operations. I think that again, because serverless gives you the ability to deploy your code very easily, especially the tools today, I don't think there's any need to have a separate role for it. The developers can do it, and with the monitoring tools that you have and the monitoring that AWS provides, I think that developers can do it. They don't need someone to do it for them. Of course, I'm not talking about customer support and things like that, that probably will require a different world. But I think that the day-to-day in monitoring and making sure that everything ticks as expected, I think that developers can do it.

Jeremy: Yeah. And I think I think like you said, the idea of owning all the way through QA into production is really interesting. And, I mean, maybe this even goes back to the cost optimization thing when I'm writing something in serverless now, you know, I might be building a couple of Lambda functions. I'm interacting with Dynamo. I'm doing some of these other things. I spend a lot more time thinking about the design of the application and how it should be built, than I do actually writing code. Like a day that I write only a few lines of code, but I've launched something that is production ready, you know, is a pretty good day, so we certainly don't measure — you know, the less lines of code, the better, in my opinion. I think that's shared amongst quite a few people. Alright. Great. So let's kind of wrap this up maybe, because we've been talking for quite some time, but I'm fascinated by this conversation, so I think I could probably talk to you all day about it. But maybe you could just give us some general advice based on your experience, some general advice for engineering managers that are starting to manage serverless teams.

Efi: Yeah, I think use the serverless benefits. So move fast. Test serverless ideas. Add new features quickly without getting bogged down with provisioning problems. Again, it's something that serverless, the cloud provider, gives you. I think if your code, if your application is monolith, again, we start with an existing thing. Yes, according legacy code, your application is monolith, I think you should start breaking it into smaller components and you don't need to break everything. You can start with breaking the peripheral stuff like report generation, email services, Slack alerts, and all kinds of services that are not the core, and slowly but surely start, you know, eating parts of your monolith code and responding to serverless. Again, I am returning back to the original question that you've asked me. Don't use the buzzwords. Use services everybody uses and move slowly.

Jeremy: Right, yeah. And I think that that's a really good point. I mean, it's starting simple.

Efi: Exactly.

Jeremy: There's no reason to launch something with a very complex, you know, multi-connected EventBridge with Kinesis in there. And yeah, or SageMaker. Like these things get complicated or can get complicated pretty quickly, so alright. Well, that's great advice. Listen, Efi, thank you so much. This has been absolutely awesome. I think I've learned a ton. Hopefully, the listeners have learned a ton. So maybe you can tell people how they can find out more about you and about Lumigo.

Efi: Yeah, sure. So I'm on Twitter @TServerless and you can find me also on Lumigo.io. I write blog posts over there. And of course, you can contact me by email efi@lumigo.io. I always like and love to help others, especially in the serverless world.

Jeremy: Awesome. And then Lumigo’s Twitter handle is just @Lumigo.

Efi: Yeah. Yep.

Jeremy: Perfect. Okay, awesome. I'll get all that into the show notes. Thanks again.

Efi: Thanks.

View Details

About Emrah Şamdan

Emrah Şamdan is the VP of Product at Thundra, a tool to provide serverless observability for AWS Lambda environments. With the development team, Emrah is obsessed with helping the serverless community with their debugging and monitoring effort both in production and during development. He is responsible for making trouble for the Thundra engineering team while finding solutions to ease the life of serverless teams.

  • Twitter: @emrahsamdan
  • Thundra: Thundra.io
  • Blog: blog.thundra.io.
  • Demo: demo.thundra.io

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to serverless chats this week. I'm chatting with Emrah Sandam. Hi, Emrah. Thanks for joining me.

Emrah: Hey, Jeremy. Thanks a lot for having me today.

Jeremy: So you're the VP of product at Thundra. So why don't you tell the listeners a little bit about yourself, your background and what Thundra is up to.

Emrah: Yeah, sure. So I'm, you know me as a product manager for Thundra. I started as a product manager at Thundra while it was a start project — it was an insider project in OpsGenie. We were some engineers, me and some designers that began an internal product for OpsGenie engineers. Then it turned out to be a product and company, and now we are serving serverless developers for observability. So in 2017 Serkan, our CEO and the CTO actually acting, and he was developing some modules of OpsGenie with AWS Lambda. And he had some problems with the observability and he couldn't find any solution that fits the purposes. And he said, hey, I can write general libraries because they were writing in Java at that time which can give me some ideas about like how my Lambda functions are performing. And he developed this as like an extracurricular activity for OpsGenie, and he made this available, and it was sending data to Elastic at that time. They were seeing some Thundra produce data. And they thought that, even before I joined OpsGenie, they thought that why don't we make it as a separate product? And why don't we make it as a separate company and I joined and they hired me as a product manager for that. In October last year in 2018, we decided to spin off it as a separate company because, you know, OpsGenie was sold to Atlassian and Thundra will continue as a separate company. And we are helping serverless developers with observability by aggregating traces, metrics and logs.

Jeremy: Very cool. All right, so I wanted to talk to you today about reducing MTTR in serverless environments, because I think when we think about meantime to repair, normally we have a lot of control. Like if we're running our applications on-prem, then we likely have access to the physical servers and the hardware components, and even if we're running our applications on something like EC2, we still have access to the operating systems, the VM instance sizes, the attached storage, and the same is typically true with containers as well, right? So we have a lot of ways in which we can affect the time it takes to repair some of these hardware or even scale issues. But If you are in a serverless environment, then it’s quite a bit different, especially if you're using a lot of managed services from the cloud provider, you really don’t have access to the underlying operating systems or hardware anymore. And I know some people have changed the “R” in MTTR to mean “recovery” or “resolution” since it’s really less about actually repairing hardware. But maybe we can start there, maybe you can give us your thoughts on what's different with how we respond to incidents in serverless versus how we would respond to incidents with more traditional applications.

Emrah: Definitely. So in traditional applications, as you say, there are some resources that we can easily gather the information when we see some problems, some incidents in our system. But in serverless, on the other hand, it is like you have different piles of logs, which it comes out of box from CloudWatch, from the resource that Cloud vendor propose. But these are actually separate, and these are not actually giving the full picture of what happened in the distributed serverless environment. And what you need here is that the problems are different. In a normal environment, the problem, most of the time, was actually about scalability and you were responding to that by giving more resources, by just increasing the power of your system. But with serverless, the problem is about like some problem occurs in any kind of a system in a distributed network and you need some more than log files. You need like all three pillars of observability, which is called traces. In our case, it is distributed traces, which shows the interaction between Lambda functions and the managed APIs and the managed resources and third-party APIs, and the local traces, which shows what happens in the Lambda function, and the metrics and the logs.

Jeremy: Yeah, right. And I think you’re right that the distributed nature of serverless is something that might be relatively new to people as well, so just figuring out where the problem is, or what component is causing the issue, is a challenge in and of itself. So the point about metrics is interesting too, because as you just said, the scalability is handled by the cloud provider for you with most of these services. So we’re likely not as worried about low level metrics like CPU usage anymore. So what are the signals of failure, like, how do we know that something is broken? What are the things that tell us something might be wrong in our application that we might want to address?

Emrah: Yeah, sure. So, like, if the scalability is not the problem, so what might be the problem? So what might be the good metrics that we should look at? In this case, there are some metrics which are actually very predictable by everyone listening here. But they are actually saving our system’s availability a lot. Say first, and the most important is actually latency, because of our aim to not receive timeout alerts, right? So the latency metric, the duration of invocations metric, is something that we should keep our eye on. So you need to keep an eye on how long do our functions take? So you need to see that if the duration is approaching to timeout, if there is, we should check what might be the reason we should check that? Is there something? Is there a problem with the third-party APIs? Is the problem with the resources that we are using? And this gives you like, when you see a latency, you should be approaching it very, very carefully because you don't want to be in the storm of timeout errors. So the second metric, in my opinion, is that memory usage. So you know, we are provisioning a memory to Lambda function. And this is the only thing that we control in serverless. Most of the time, developers are giving the memory more than it actually requires just in order to increase, speed up the IO and throughput and let’s say the CPU. But, in this case, we are having problems with the cost. You know, because whenever your function gets triggered, it runs for a time, and our billing is decided by GBs per second. So in this case, if you allocate more than necessary memory, we may be losing some money for Lambda. This might be very negligible if you don't use Lambda excessively. But if you're using Lambda in production and you're using Lambda and serverless mostly, you'll have some problems with the costs. In order to do that, you should tune your memory accordingly, and I love what Alex Casalboni does about this. You should see how your function is performing best with which memory configuration. Even after that you should keep track of memory usage and see if there's a jump that you don't expect, and you can again tune it again. So you should tune it again because there might be some changes in the managed resources they're using. There might be some changes in the inputs that you are processing. So memory’s just something that you should keep your eye on even after you successfully and carefully tune it. And the other metric is actually not related with Lambda itself, but the resources that we are using. So let's say, for example, they are using Kinesis, you are using SQS, so we need to keep an eye on how our Lambda function is performing in terms of these managed resources. So we should keep an eye on the Kinesis iterator age. We should take a careful look at the SQS queue size in order to not overflow the messages and not to have data losses.

Jeremy: Right. And so then, all this stuff is telling us when we see the increased number of timeout errors, for example, then we could assume that it’s potentially a third party service or one of the managed service we’re using is taking longer than it needs to. And if we see things, like you said, the iterator age for Kinesis or the SQS queue size growing, that those sort of things give us indications that either our application isn't performing correctly or maybe some sort of downstream service isn't performing correctly. But distributed systems have failures all the time, right? Random connectivity or network latency issues. So should we be worried about occasional timeout errors here and there, or is it more in the aggregate that we want to look at?

Emrah: Yeah. So when a timeout error happens, it depends on your use case with Lambda. It might not be like, the single timeout message might not be the end of the world for you. But you should, again it depends on your use case. When you have something that's critical, one timeout error can be something significant. But sometimes when you are using data processing and you can just retry it from Kinesis, let's say, and it doesn't mean that much and the signals can also come from the latency again. Latency when you are processing your message, it may not be that problematic when the invocation increases to some extent because that might be a steady state of your functions when there is lots of load on your system, the agency can increase to some extent and it may not be something that problematic. So what I'm trying to say is that there might be a range that your function is performing daily, weekly, maybe monthly, and there might be some spikes because of the traffic, because of the third-party API slowdowns and these may not always signal failure, but something normal in this steady state and you should know about it by constantly observing your system.

Jeremy: And what about things, though, that you maybe could address? So I talked to Hillel Solow the other day and we were talking about flooding kinesis streams with junk data and then having trouble draining those. And that was more on the security side, but I can see things like bad messages in an SQS queue without the proper redrive policy causing lots of problems too. So how do we know those sort of things are happening?

Emrah: So like in this case, you should be able to have that observability tool that helps you to see what is the request and response that's coming through your Lambda function. So let's say that you have this function, gets to you by Kinesis, and it gets some message and it doesn't throw any timeout, but it just throws some input. In this case, you should set an alert for this condition, and when I see this error type excessively, more than 10 times in a row, I should take an alert, and I should just go to SQS or go to Kinesis and take out the poisonous message.

Jeremy: Okay, so I feel like the SQS situation, where maybe you have a poison message in there, and it's just causing the function to fail over and over again, that those are fairly obvious when you start getting those exceptions. But what about something like, so for example. So this happens to me on one of my projects. I get an error that says something like, and it includes a stack trace, but it's something like a mapping error. And it happens maybe once or twice a week, you know, it's something like 0.001% of my invocations. So, very, very small. It's this sort of edge case, and I know it's just something that I need to address, but I see other errors like that too where something will just kind of pop up and they feel like they're anomalies. Maybe, or maybe just outliers, but, how worried should we be about things like that?

Emrah: So again, it depends. So the first thing that you should be aware of is that outliers happen. So these kinds of situations happen, but as I said, this can be, the outlier even can be in this steady state, but you should be able to understand the reasons of outliers. I talk with many customers from many Lambda users, and they have this information. They’’ll tell you “My 99 percentile of my functioning invocation duration is like, let's say, one second. And, normally it’s like 200 milliseconds.” And when I asked them, “What is the reason?” Do you know that most of the time, [they say,] “I don't know. Maybe this third-party API that we are using. I see it’s slow sometimes.” But in some of the times, it’s not this third-party, but let's say, the DynamoDB table that they use. So when your function is having an abnormal invocation duration, you should know like not exactly at the same time, but maybe later, as a retro, you should know what caused them? So in Thundra we provide kind of a heat map to our customers and they see what are the outliers in the heat map. And when they focus on that, they see what is the normal usage of, let’s say, Dynamo and what's the value of Dynamo in this outlier region. It helps them to see if what is jumping during the outliers. After that, they can also have a closer look to outliers. So what are the inputs? What are the outputs like? What made my Dynamo run slower than normal? And it might be because of the input, it might be because of your bad coding practice, and it might be just because of Dynamo itself. So we made an experiment with Yan Cui, actually. So maybe two months before and it was DynamoDB Keep Alive Connection, and he was seeing some spikes in DynamoDB duration, even if he’s using keep-alive. And we see with Thundra that it's getting, it's actually renewing the connection even its keep-alive. So we called how DynamoDB actually renews the connection, and it causes an outlier. Maybe, if we trust DynamoDB keep-alive itself, it may not be sometimes — it may be causing problems. But detecting this and knowing about it, actually makes us prepared for such kinds of frustration.

Jeremy: Yeah, so that's really interesting about the keep-alive thing, because AWS doesn't put it in the AWS-SDK, but a lot of people use it because it does speed up subsequent HTTP calls. But that's interesting that it could potentially cause a problem if it needs to be reset. And I would assume that some of these outliers will be things like cold starts too, which I know most observability tools will say, “this is a cold start” or “this is a warm invocation.” Yeah, alright, so let's move on to talking about what exactly a “failure” is. You mentioned earlier about not getting an alert every time there's a timeout error or something like that, and I know failure conditions are sort of different based on different customer preferences. But maybe, we could just talk about when does an error or group of errors constitute a failure. And maybe a better way to put it would be, in the serverless world, what are the signals of failure that we can actually address. Like, when should we take action?

Emrah: Yeah, this is something that we are thinking [of], and we [have been] trying to find answers for a long while. You know, most of the tools that, both with CloudWatch and with the other monitoring tools, did the alerts are just for a single error. So you're just having a one Lambda invocation, then an error happens, and most of the time they're paging an alert. But we thought is this something that is actually wanted by people. Is this something that prevents people from alert fatigue? You know, we are coming from OpsGenie, and that's why we are very, very careful about not putting people into alert fatigue. So we ask people, “What is the definition of failure for you?” Like we asked tens of people,”What do you think? When do you think that this serverless architecture has failed?” And the response is that not a single error, like most of the time, it’s not an error. So when I call something an incident, when it causes something cascading failures. So I have a problem with Lambda function, and this Lambda function should should have triggered another Lambda function through SNS, and this triggers another Lambda function to, let's say, SQS. I'm just throwing out a scenario here. So this first Lambda function fails and the others couldn't even start. So in this case, we can understand that we are in very big trouble, that we lose some transaction there. That’s a failure for most of our people that we talk with and the other stuff is that for, at least, especially for upper management, the invocation duration, invocation count, any kind of abnormality about these metrics, are not very important. And they are seeing cost as a signal of failures. So let's say when they want to allocate $10 per day in to the serverless architecture and let's say, one dollar per a function. In such cases, they want to get alerted when the cost exceeds this threshold. They are not interested in if the function is running more than expected because of a third-party API. They're not interested in if the problem happens because of an input error. They just wanted to see if the cost is exceeding something, some threshold. Because all off their motivation was, when joining to Lambda, to save cost. And they don't want to read that with a problematic situation.

Jeremy: Right, yeah. I think cost is always a good indication even, you know, just to see if things were taking longer to run than expected. But I always like this idea. I mean, you mentioned tuning for the memory side of things, but I think tuning for the timeout is also important, because I think it's easy to say, “well, let's just set the timeout to 30 seconds and then that way if something's running slow, then, you know, we'll just absorb it in the 30 seconds,” and maybe that's right for certain situations. But I really like this idea of failing very fast, especially to make sure that the user experiences is good. And speaking about user experience, you mentioned alert fatigue. And I think that actually is really, really important because I know I've worked for several organizations that have had all different types of alerts for their applications and infrastructure. And there's always that one error that you just get every day, like 10 times a day, something random that’s completely benign because you just know, like, “Oh, yeah, every once in a while, this runs for a little bit longer,” or whatever. And maybe you go in and you tweak your alerting system to say, “don't show me these alerts,” but in my experience, most people, almost never do. So then you get to the point where you've got an email folder that's flooded with alerts and you just start ignoring all alerts and that gets very, very dangerous. So how do you fight against that? Is it something like anomaly detection? Like, what can the providers do? And this doesn't have to specifically be about Thundra, but what can you do, or the tools that you use, what can they do to prevent alert fatigue, because I think that is it is a serious problem in many organizations.

Emrah: Definitely. This is a problem that most of our customers are facing. And even if you are using Thundra or not, you just know that what type of errors that you're facing and when you face one of them, how many more is coming for you. So when you see an alert in a time, so when you see a specific error type, it means that after a while, like for a while, you will have thousands of them, tens of them, like hundreds of them, the kind what we call alert storm. In this case, so you should either actually configure [an] alerting mechanism that don't raise me an alert until the error type, the errors with this type exceeds, let's say, 100, 10, whatever about your case, or you can say that, after I create, after I got the first alert about this error type, I don't want to get alerted for, let's say two hours for, let's say, 10 minutes. In this case, you don't see like, lots of messages in your mailbox if you're using OpsGenie or some other tools that you don't need to page an alert for every single error and you can just get the first alert and understand the severity, and you can just go and focus on the problem and solve it. So in this case you can do it yourself, but if you are shipping your logs to, let’s say, some other stuff from CloudWatch, and you can do it with Thundra by writing your own query. Let's say that I like to get alerted, then there are more than 10 errors with this specific type, and after I get the first error, I want to throttle this alert. I don't want to get more than one alert for two hours. So this is again your SLAs, your rules, but you should be able to understand what's happening. And for now, what we do with Thundra, just as I said, we are giving people a flexible querying system, which lets them configure the others as flexibly as possible. I can tell. But we're also now working on some learning capabilities on how the problems occur in their system, and most probably we will be giving up this option to them by the end of year about, like, we will be actually understanding when to raise an alert for you. And this will be more flexible and even they want me to do this, but, you know, then they want to keep the control. We will continue to let them use the queries.

Jeremy: Yeah, and I think it’s hard to build tools that are self-learning like that because our architectures do change quite a bit, so I imagine keeping up with the changes wouldn’t be easy. All right, so what about failures outside, I don't want to say outside of our control, although they probably are outside of our control? But things like SQS errors and DynamoDB errors, like the inability to connect for some reason. Or I get SQS errors all the time that just say “Internal Error” or something like that when you try to put a message to SQS. And it happens every once a while and I know that there are retries in there that happened automatically for you. But when you get messages like that, I don't know if I want to be alerted on those because they just seem to be like normal distributed system errors that I really can't do anything about other than have a good retry mechanism in there. So what are your thoughts on those types of errors like, should we be alerting on those, or should we not?

Emrah: So this is something that is very controversial with our customers as well. So some of our customers wants to know in any case, like they want to understand what even if there's nothing that they can do. They want to report the situation today, to the directors, their team leads. And some of our people say that “Hey, there's nothing to do with this. Let's grab a beer. This is something SQS errors and retry will fix it anyway.” But in such situations what you need to do is this, like when you first received this error, and even if you get alerted or not, you should just check from a distributed tracing application like us and see if there's a problem, there's something abnormal from your side in order to verify that it is something caused by service providers, something caused by Dynamo, something caused by SQS. In this case, you just check the message that you sent, you just see if there's a problem into operation name, if there's a problem with the timing of these and you should see what happened in your code, maybe line-by-line, maybe in a method level. And how did you prepare this message for them? And then you are sure that there is nothing wrong with the message that you sent, you can then blame the others that, “Hey, this is not related with us” and you can just sit and wait until it is fixed actually.

Jeremy: All right, so I think this is a good segue , maybe to the next topic, so I mentioned actionable alerts. So we know bad things happen, right? We know that something can't communicate and APIs slow down and managed service might become unavailable for a period of time, or there's some bad code in your app somewhere, so we can detect some of that stuff, or you can detect some of that stuff with observability tools. You can alert users, but I kind of want to talk about how you respond to those things. Because first of all, there's this responsibility shift, maybe, right where we used to have operations teams that would address the scalability and the hardware issues. And, you know, and as we moved to sort of DevOps, that sort of translated into the developer and OPs working together to try to figure out, if it is a code problem or configuration, or things like that. So we're obviously not using as many operations people in serverless, especially for smaller customers. I don't think you're gonna have operations teams at all, startups and things like that that are using serverless are going to be almost entirely serverless or mostly serverless. So now you get a developer who gets this alert. So when a developer now gets an alert that is actionable, that they can actually do something about, what’s the first step? What does the response process look like?

Emrah: Yeah, as you say, like in previous systems, which like non-serverless system, so, like most of the time, it was a scalability issue, there were Ops people who [are] actually on top of that, and they were scaling systems and they were solving issues even without asking the developers. And they had these issues. Now with serverless, like, as you say, there are teams that don't even have any kind of an operations people inside. And even front-end engineers are building up very nice products with serverless these days. But there are still some problems which actually impact our systems, so which actually impacts our end-user experience. So the problem, most of the time, is actually the errors and the latencies. So we don't have the scalability issues now, but you have these latencies and errors. So the operation people actually tune themselves as the people who teach the colleagues about how to respond to such kinds of errors too. So in this case, like when you take an alert, and I also talked about this: what makes an alert actionable? So an actionable alert, which shows the stack trace of this alert, and this shows how actually it is cascading between functions. So we provide such kinds of alert for our customers when that happens. They can show what's the stake trace. They can see what happens. What was the request in response to this function? And in this case, they can actually understand if the problem is occuring because of them or because of the message that they just received. And is it because of that piece of code? Is it because of a misconfiguration? And they can understand from that perspective. And there are those alerts because of latency. So like what makes an alert, what makes a latency alert actionable? In this case, whenever you receive another, let's say that you set up an alert, which says that I want to get alerted when the dysfunction or all transactions as a change invocations of several functions exceeds one seconds. So alert me when this condition happens more than five times in the last 10 minutes. So let's say you have such kind of an alert. When you receive this alert, the first thing that you will actually want to see is that what caused the latency jump? So in the alert body itself, there should be some[thing] somewhere saying that “Hey, your function is performing worse than before, worse than normal and the cause is that you are reaching out to this API on like some URL and this URL started to respond slower. Like normally, it was like minutes 50 milliseconds. Now it is 500 milliseconds. In this case, you can say that, “Hey, this URL is performing slowly. Is there something that I can do?” This is something [where] you can take action actually.

Jeremy: Yeah. So what I'm thinking is that when you get alerts, and that’s one of the big things that observability tools certainly help with, is this idea of identifying what type of error it is. Because certainly, if it's things like IAM permission errors, which can happen, right? Somebody on your team changes an IAM permission and then all of a sudden one of your functions that doesn't often use the DeleteItem call in DynamoDB now needs to delete an item, and you start getting errors around that. Obviously, things like the timeout errors and the memory errors those are all things that certainly would kind of get you to start looking at stuff. But maybe that's where this, this ounce of prevention is worth a pound of cure saying comes in, right? So basically if you prepare for this sort of stuff, if you're in better shape for these types of failures, because you can sort of anticipate them or build in the resiliency to protect against these things, then you’re way ahead of the game. And I had a whole episode where we talked to Gunnar Grosch about chaos engineering, and he actually mentioned Thundra as being one of the tools that can sort of automatically inject latency and errors into your application to help plan for this kind of stuff. So I don't want to talk too much about chaos engineering, but I think it is a really, really interesting topic. So from the standpoint of mean time to repair and the ability for you to recover quickly if there's an outage, it might be that some of these things could almost be automated for you, in a sense, especially if you did things like graceful degradation, right? So if a third party API isn't responding, there’s better ways to deal with that. So what are some of the other ways that chaos engineering could help us improve our MTTR?

Emrah: Yeah, so, first of all, I listened to the episode with Gunnar and it was very, very nice and and you talked about the basics of chaos engineering and how it is important for serverless a lot. But I can tell about, like how it is actually useful for responding to incidents. So the best way to get prepared for an incident is actually to experience it before. But no one wants to experience something bad over and over again, right? And the nice thing that we can do with chaos engineering is that you can just get yourself prepared by actually simulating that these kind of problems. So you can ask yourself, what if this third-party API that I'm using starts to respond slower? What if the DynamoDB that I'm just leaning on completely starts not to respond. So you can you can run such kind of chaos engineering experiments, and in this case, you should be knowing what will happen. And you should be knowing that not just because of, not from the perspective of what to do, but how to inform the customers, how to inform the upper management, how to have the, let's say, the retro. You can understand how we can respond to these kinds of situations from many different perspectives. So let's say you can set up a HTTP latency chaos experiments and you can see how what are you going to do then when this kind of situation happens. So you can you can set up a graceful delegation policy there, and in this case you can still serve your customers, even if this third-party API is slow. Let's say that what you can just stimulate an idea with Thundra, you can just inject an error to DynamoDB, like illegal access exception, IAM permission exception, and you can't reach out to Dynamo. Maybe Dynamo can experience some throttles and maybe DynamDB — I don't know. I hope it never happens — but DynamoDB becomes unreachable for most of the people in the region. So you should have seen what will happen in your system. And you should actually implement some some exit points for that. So this is all that chaos engineering is about actually: getting prepared for the incidents. So you should maybe, for example, for a DynamoDB issue, you can put a circuit breaker. You can put like another version off your Lambda function with the circuit breakers to Dynamo and you can just upload this. You can just deploy this Lambda function and in this case, you can return some default responses to your customers. But in any case, you won't collapse. You will be able to still answer your customers with some dummy answers, but customers won't be able to see an error message and won't be able to see white screen. In this case, you'll be able to understand what's happening. And the last thing that I'd like to talk about chaos engineering is that you get yourself prepared about like how actually you will respond to the incidents as a team. So when you receive an alert from Thundra or whatever, this will be most [likely] the first time you receive this. And with chaos engineering, it won't be the first time, and you get prepared. You will be have a run book then when you have this error.

Jeremy: So that's actually a question that I have because again, we might be putting circuit breakers in, and gracefully degrading the experience for the user because the latency is too high for something. But when do we want to know about that? So this is something I spoke with Gunnar about, where we said, even if we can't process a credit card transaction through Stripe right now, that's fine. We just buffer those requests in SQS, and then process them later when the service comes back up. So if we do that, if we build in these circuit breakers and we build in these things where you know we have a short timeout on purpose. Maybe we say, if we can't reach Stripe within five seconds, then we're just gonna fail it and we'll buffer the event and we'll just replay it later. Should we still get alerts on these? And maybe these are the non-actionable ones like we talked about, where it's okay. You know, Stripe is down again, although, I think that doesn’t happen very often. But let's say it's some small third party API, and we say, “Oh, yeah, that's down again. It's fine. It'll come back up. We know everything is being buffered.” But can we get alerts on those sort of things to let us know an incident is happening? Just so you're aware of it, but you don't need to take any action?

Emrah: Yeah, this is actually a very [educational] for the team, so you should be knowing the resources that you are using are not performing well from time to time. And you should be knowing this specific statistically that what happens with third-party API? Like Stripe doesn't fail that much, of course, but like this third-party API, which is something not very manageable. Such kinds of statistics give you an idea that if you should still continue to use this, if you should still continue to make a request to these third-party APIs or you can start searching [for] some alternatives if you can. So one of our customers did this with with some APIs that they're using. Actually, because of us, they switched to another product because they see that the latency with this third-party API started to increase continuously, and this, they see that it's not getting any better. And they switched to some alternatives, some competitors of this third-party.

Jeremy: Awesome. Okay, well, listen, Emrah, I really appreciate you being here. Why don't you tell everybody how they can find out more about you and more about Thundra.

Emrah: Yeah, so thanks a lot for having me, first of all. You know I'm on Twitter. I'm with my name and surname @emrahsamdan, and my DMs are open. And I'm I'm talking with many people about serverless and observability. I actually love to speak with the serverless thought-leaders like you, and like many others. Thundra is reachable by Thundra.io and we're releasing continuously some blogs about, like how-to articles, some of our feature updates on blog.thundra.io. We also have an open demo environment in which you can see what Thundra does like using a sample data that we produce continuously. It's reachable by demo.thundra.io, and this is more or less what I can tell about Thundra. If you need something like distributed tracing combined with local tracing, Thundra is, for now, the only solution that you can find. We can just talk about it on Twitter, on our Slack if you join, whenever you want.

Jeremy: Awesome. All right. We'll get all that information to the show notes. Thanks again.

Emrah: Thanks. Thanks. Thanks a lot, Jeremy.

View Details

About Hillel Solow

Hillel is passionate about security innovation, and is driving product innovation and security at Protego. Prior to co-founding Protego, he was CTO in Cisco’s IoT Security Group, where he worked on innovative security solutions for new technology markets.

  • Twitter: @hsolow
  • Blog: protego.io/blog
  • Protego: protego.io
  • Twitter: @ProtegoLabs
  • LinkedIn: https://il.linkedin.com/in/hillelsolow

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Hillel Solow. Hi Hillel! Thanks for joining me.

Hillel: Hi, Jeremy. Thanks so much. It’s a real honor to be here.

Jeremy: So you're the co-founder and CTO at Protego. So why don't you tell all of our listeners a little bit about your background and what Protego is up to?

Hillel: Sure. Thanks. So Protego is a security company focused on serverless security? We've been around for a couple of years. Prior to that, I had spent about 20 years in security at companies like Cisco and various other companies. And we really started Protego because we saw that serverless and cloud native was going to really usher in a wave of changes in how we deploy applications and build applications. And that was really going to upend a lot of what we do in security. And so we really focused on trying to ground up understand what is it about serverless and cloud native applications that changes? What's the best way to secure them? What do people worry about? And how do we help them solve those problems?

Jeremy: Awesome. So I wanted to talk to you about serverless security in the real world, and by that I mean the things we are actually seeing. Because I think that there's a lot of misinformation that is out there. And I know there's a lot of security companies starting to focus on serverless and cloud native. And every once in awhile we here about these security breaches in the news, so I think this is just a good opportunity for us to talk about what we really have to worry about. I mean, obviously want to have a good security posture for whatever we do in the cloud. But maybe we could start by discussing a recent, sort of, high profile, or highly publicized, successful attack like Capital One, for example. So I know this wasn't serverless related, but what are your overall thoughts on that attack? Does that scare people when they see something like the Capital One thing?

Hillel: Yeah, it is interesting because I think Capital One has done a really great job of leaning into the cloud and taking advantage not just from a development and deployment perspective, but from a security perspective of everything that cloud can offer. So it's a bit unfortunate now that they're going to get hit on the head here. I don't think it's a result of them moving to the cloud. To a large degree, this kind of attack that we’re looking at, it's kind of similar to the other kinds of Equifax attacks in some ways. You know, it's some misconfiguration and some access to an EC2 machine machine that then had access to some S3 buckets that shouldn't have had access. So those kinds of things, you know, obviously they can happen across any kind of infrastructure. The fact the Capital One is leveraging, you know, Amazon to do a lot of the securing of the infrastructure below what they're doing is great. It does highlight the fact that at the end of the day, though, we're all responsible for our own applications. And Amazon says that you know, day and night. And so for us to focus on the things that you know, that we deploy our business logic, that's really important. It’s important, obviously for Capital One, and I think you know, they do a great job of it for the most part here, and obviously they're going to have to improve. But I think for all of us, it's a lesson in how careful we need to be about applications security and about how we're using the cloud. Because just because Amazon is securing the underlying platform might lead us to believe that we don't have to deal with security. And it’s obviously not true.

Jeremy: Yeah, definitely. All right, so let's talk about the first aspect of this, because like I said earlier, I think there’s misinformation out there about what it means to be serverless and what your security posture becomes once you go serverless or even just move to the cloud in general. So there's this concept of FUD, right? This fear, uncertainty and doubt that you tend to see a lot of people and companies using to maybe “exaggerate” the risks. And I know your team is great at sort of shutting down the FUD, right, just giving people real, honest answers. Which is really refreshing. So maybe we can jump into that, and just give me your thoughts on how you feel about — you know, this idea of people scaring people, by spreading misinformation about the security of serverless.

Hillel: Yeah, look, I don't want to discount the value of fear. You know, I think if you're a security company, it's nice to be selling a product that solves the problem people are really worried about, and that's obviously important. But I think this notion of us becoming hysterical about things that aren't really issues is something we need to avoid. And specifically for us, as we’ve looked at serverless and how it changes security, I think one thing is really clear. Serverless is not less secure than other things. I think, you know, in a lot of ways, serverless applications stand to be the most secure applications that organizations deploy for a bunch of interesting reasons. They do raise some interesting challenges in terms of where do I put the stuff that I used to run on machines or where do I put things that don't scale the way that I want them to scale in the serverless world and things like that. And obviously they do create different types of opportunities for attackers. They do change some of ways attackers are moving, but overall I mean, my strong belief is if you're making the move to serverless, you're going to get a net win on security. You just need to take advantage of a lot of what's out there. And for us, you know, I'll talk a little bit later about what we do, but in particular, a lot of what we focused on is: hey, what happens when you move to serverless and cloud native? What new opportunities are there? And how do we leverage those for security in a way that maybe in the past was challenging?

Jeremy: Yeah, and I think the other piece of that, too, is that you have developers that are now much closer to the stack. And I've said this a million times, but this always makes me a little bit nervous because there are some new things that a developer might be responsible for when deploying and securing your application code in serverless. And like you said, the infrastructure security provided by the cloud providers already gives you this great foundation. But, if you don't have those skillsets or you're just not used to implementing IAM policies because maybe they were handled by Ops people or there were tools like WAFs and things like that, that gets a little bit scary for me anyways, when I see what some of the younger or junior developers do. And certainly that’s part of their cloud learning experience, but without proper controls in place, it does open up risks. So let’s talk a little bit more about what's different with serverless security versus more traditional security systems. And one of those things would be, speaking of IAM roles, this move to very, very small fully managed compute units, as opposed to the security of an entire machine or maybe a container where you have full access to the execution environment. So what's the difference there?

Hillel: Yeah, sure. So I think first I’ll state to your earlier comment, you know and again, I have nothing against young developers. I would like to think I'm still young developer in some ways, although I don't think anybody else thinks that. I think a couple things have happened over the years. I think we spent a bit of time, you know, 10 years ago and 15 years ago, focusing on getting developers to be better about security. I think we took our foot off the gas little bit on that, and over the past 10 years or so, I think a lot of what we've done in security is focus on trying to wrap up developers in an environment where they're kind of sandboxed from evil so developers can write any stupid things they want. We've got scanners and WAFs and agents and all sorts of things to try to secure them. And I think one of the things that you see in the move to serverless - and again I don’t think it's a serverless-only thing, but I think it's in serverless more than anywhere else thing - is that sort of divide between security people and developers. It's not really tenable and, you know, in serverless, a lot of the security controls that security used to own are now security controls that developers control. Like configuring IAM roles and setting up VPCs and things like that. And so in a lot of ways, we've actually put more responsibility on developers, but we haven't necessarily empowered them in real ways to make security decisions, and at the same time, we haven't given security people a way to meaningfully understand and audit some of those things when they don't necessarily understand what the application does or what the code wants to do. So I think that's been a big change, and I think that's true across a lot of cloud applications. But it's just truer in serverless applications. You're kind of forced to reconcile that. The other thing about serverless applications specifically that we like to talk about is the fact that developers have gone from an application that comprises 10 containers to one that comprises 150 functions, you know, could create all sorts of nightmares in testing and monitoring and, you know, deployments and things like that. But for security it’s an interesting win there where you get to apply security policy, IAM roles, runtime protection, at a very fine-grain level — you know, at kind of a zero trust, small perimeter level. And that's if you can do it right, if you can do it at scale and automatically, that could potentially be a huge win, really for, you know, mainly least privilege, reducing attack surface and reducing blast radius. You know, something goes wrong; my developer left a back door accidentally into a function. But now that function really can only do right to one particular table, as supposed to, you know, in the old world, where that gave an attacker a lot more capability. So I think that's an opportunity that is on the table. It is challenging to capitalize on that. Like you said, there's less time. There's less gates between developers and running production code. And that means that, you know, how do we automate and capitalize on a lot of that value without trying to slow everybody down? That's the big challenge.

Jeremy: Yeah. You mentioned this idea of developers running their code in sandboxes. And when you put traditional tools like WAFs in front of incoming web traffic, obviously it inspects that, and we hopefully take care of basic things like SQL injection and cross-site scripting attacks and things like that. But what about all these other events? Because that's one of the things that I always like to address, and I don’t bring up to scare people, but serverless certainly promotes event-driven architectures, and AWS Lambda has something like 90+ event sources, plus custom events, and really only two of those, I think, would even look like traditional web traffic. WAFs aren’t designed to understand these other types of events. So what are the WAF equivalents for other event, or is this just all about writing good code?

Hillel: Yeah, absolutely. And I think it's interesting to see the evolution because I think if we go back about a year and a half, you know, there were, I think, 16 event types and not as many as there were today. But it was still interesting to see that I would say 90% of what was going on was kind of API gateway and then CloudWatch timers. But over the past year and a half, I've seen a lot of evolution in terms of how applications are built, and people are really learning how to use some of these triggers that are out there to build more interesting applications — applications that are more efficient that bypass a lot of the bottlenecks that existed in the past by using, you know, things like AWS IoT or AppSync or we're using, you know, Kinesis to directly upload data. So a lot of those triggers are now out there, and you're right. WAFs typically are somewhere between difficult and impossible to put in front of most those things. Again, that's true in non-serverless applications as well. It's just that the norm now is to build some business logic that's triggered by some data coming in from one of these APIs that I can't necessarily put a WAF in front of. So, yeah, we need to reconcile where are we putting security so that it's more agnostic to where things are coming in, where are we putting security, so it's not assuming that all the bad guys are outside this big perimeter that has one big front door and all the good guys on the inside. And again, I think we should be assuming that for the rest of the cloud as well, but we just don't have a choice when it comes to serverless.

Jeremy: Yeah, and the other thing, too, and I guess this would apply to the entire cloud infrastructure, or infrastructure as code really — is where it's very easy for you to spin up all these dev and staging environments, and you see developers and companies creating lots of functions and lots of versions of functions, API Gateways, and things like that. But those are all out there, and often never get cleaned up. Is that a security risk, in your assessment?

Hillel: It's interesting because the rest of the world tries to decrease friction, right? That's what we're focused on. How do we do things with fewer barriers? Security kind of likes their friction. We really enjoy a little friction. The fact that there's friction means it's a little harder for a developer to do something. And maybe there's more things along the way that might prevent him from doing something silly. So the fact that we can now spin up machines instantly, we can deploy functions instantly, that's obviously great for productivity in a lot of ways, but it does make security a bit of a nightmare. And yeah, I think the move to serverless as a mindset, but also as a technology. I mean the fact that you could literally go into the console in Lambda, hit, create function and write some code, hit save, and that thing's in production? That's incredibly powerful, but it's a lot more scary than it is powerful in the real world. And so what we see a lot is we see sprawl really quickly. You know, in a lot of ways, the last protection from kind of resources you didn’t need that were hanging out there in the cloud world was your CFO with somebody at the end of the month going “Our Amazon bill is what? Oh my God, let's do an audit of that. You know, what's that machine called Hillel Test 1 running in Australia. Do we need that?” And I’ll go, “Yeah, probably not. It's called ‘test.’” So thank God we paid $200 a month for it because from a security perspective, somebody wants to care about it. Now you can throw functions and stacks and versions, old versions of functions and API gateways instantly. You pay virtually nothing for them unless they're being called. So from a developers perspective, okay, I put some tests back up. What do you want? I put it up in US West-2, which is where we don't put our production applications. So it's easy to realize it's not production. And I tried something out and I didn't delete it. But from a security perspective, well, that's all callable. That's accessible. It's probably got very little error-handling because you were just testing something out, and it's open to the outside world, and it's probably not going to go away unless we’re really mindful about it. So yeah, I think you know, we see people go from zero Lambda to hundreds of Lambda where they really only need 13 or 14 Lambda functions in production, you know, within the course of months and a lot of what you need to focus on — whether it's through tools or honestly, whether it's just through good hygiene — is what's out there, what's deployed. Why is it out there? Do I need it? Is it part of something important? Can I prune? Prune everything. You know, it's really gotta live in a Spartan way in the cloud.

Jeremy: Yeah, and I think, too, that when you start piling up all these old versions of functions and you have all these old API gateways that are just for testing things, as you mentioned, I’ve seen this quite a bit, where people will publish an API gateway to a test function that does something like dumps a whole bunch of information that you probably shouldn’t be dumping. And security by obscurity only works for so long, so I can definitely see how all these shadow APIs can certainly open up some security holes. Plus the other thing you mentioned too, about your “Hillel test” server in Australia, I get the higher bill can be somewhat beneficial from a security perspective, but even small resources can start to add up. So I see this too where people will create something like a DynamoDB table, but not use the on-demand pricing and then all of a sudden, they've got a bunch of read and write units that they're getting charged for, and you don't realize it because it’s often spread across lots of services. So I find that to be an interesting side effect, as well, and pruning would certainly help there, too.

Hillel: Yeah. Yeah, for sure. Kinesis is the killer. That's the one that really gets you.

Jeremy: Yes. When you're paying for those shards. Sure. Okay, so let's talk about this idea we mentioned regarding real-world serverless or serverless security in the real world. And so you've obviously seen quite a bit of security stuff. You’re a security guy; you run a security company. So maybe you can push past all of these anecdotes, and discuss some of the things you're actually seeing happening out there.

Hillel: Yeah, it's great. It's really an opportunity, before I even say anything else, to say it's worth remembering that the people out there who are trying to attack our system, they're the same people who are out there before, whether they're, you know, personal people, state actors, you know, criminal actors. Whatever it is, it's the same people. They may or may not even be aware that we've decided to build our application based on Lambda or Azure functions or something. So they want the same things. By and large, they're going to use a lot of the same technologies. They are starting to adapt a little bit to the fact that it's a serverless environment. They are starting to understand how serverless environment scale and how they could benefit from that, how they can do more brute-forcing, perhaps in certain cases, and they might have been able to in the past without being detected. But by and large, we're not seeing a huge shift in what are attackers doing. I think from a defense perspective, from the things I care about perspective, we're definitely seeing that there are new things you need to focus on — you know, the whole debate between denial of service, and then I look wallet or, as do I leverage the cloud to scale horizontally tremendously for me so I don't worry so much about being knocked over by a denial service account. But then maybe I'll pay a tremendous amount money at the end of the month. Or do I try to figure out, you know, what's my minimum concurrency I could set up and still service my customers and try to avoid paying through the nose? Those are interesting discussions that maybe weren't as relevant in the past and are relevant now. But mostly I think from the outside things are quite the same. I think denial of service in general, this is a category of attack that, you know, broadly speaking, we need to focus on a lot with applications. Denial of service not so much in the “what if I get hit by a million requests at the same time?” You know, that's what Cloudflare is there for, or AWS Shield will help you with things like that. But denial of service more in the what happens when people exploit the business logic of my application to just make me consume resources. And you know, you can hit a login endpoint and consume a lot of Dynamo resources. And as you said, you know, in a lot of these cases where there isn't on-demand billing, we've got to set up a capacity. So we, you know, until Dynamo had on-demand, we needed to decide how many read units we set up for Dynamo, and if we were storing our hashed passwords in Dynamo for some reason, and looking them up during login then we needed to figure out, you know, what's the right capacity so we didn’t get knocked over by that. But we've seen more complicated cases, and a lot of this has to do with architecture. I mean, we recently saw someone who got attacked, but that attack didn't just manifest in, “Hey, my stuff gets really busy.” It manifested in “My stuff got really busy. My Kinesis pipelines and analytics pipelines applications got full of all sorts of spurious data that were part of the attack.” That data wasn't getting consumed or drained from those pipelines in the way that maybe they should have been or with the speed they should have been, because of the way I had structured my application. And this customer was down for a couple of days until they could figure out how to properly, you know, drain these attacks out of their system and then figure how to mitigate those things. So a lot of the ways that people architect around these managed resources, and these more complicated, more distributed architectures could be really important, not just for performance, but also for security.

Jeremy: Yes. So the point about the design of the architecture, I think that is right-on, because I see this problem quite a bit, especially with things like cascading failures. So what are those trade-offs that we have to make? What do you feel is appropriate? Do we want to continue to scale up to handle all these extra requests, if we do get some sort of a flood? And like you said, Cloudflare or something like AWS Shield will knock down the DDoS type attacks. But what about the more common ones, the ones where somebody's maybe trying to brute force your password form or flood a webhook endpoint or something like that — where do we draw that line between, you know, letting it scale up and shutting people down?

Hillel: Yeah, it's always the eternal questioning in architecture in general, especially in cloud architecture, and when you factor in security, it becomes a little more complicated. I think, obviously, the earlier you can know something is wrong, and something should be ignored or discarded, the better. Right? And so one of the one of the mistakes we often make with distributed architectures is we figure out that the logical place for something to be checked is somewhere down the line. You know, Lambda function puts something into SQS, gets pulled out by another Lambda function. It does the processing on it, then spins off three other Lambda functions and one of those third Lambda functions that’s all the way down the ranks, they're the ones who are going to go, “Wait a minute this request isn’t signed properly,” or “This request doesn't match the account ID it came from,” or something like that, because, logically, they have access to the data at that point. And that's great. Except that now you think about that from a security perspective, it took a long time and it consumed a lot of resources before you figured out that that request was something you'd actually like to ignore. So in some cases, figuring out what's the earliest place that I can filter out bad data from good data, if I can recognize bad data, how quickly can I do that? How can I avoid looking things up in a database before I know something's a problem? Those they're going to be helpful. And obviously on the architecture side, trying to make sure you're using resources in the cloud that scale really well, resources that let you handle edge cases, you know, of crazy overflows or lots of extra data inside your pipeline. How do you deal? How do you deal with an SQS pipe queue that's full of requests you don’t want to handle? How do you drain those? How are you going to handle that? So some of those are architecture decisions you have to make to make sure you're not scaling out too rapidly; you're not consuming all your resources. But at the same time, you can handle that sort of massive scale that you want to handle. The whole point is to say, “I build this application.” Just as an anecdote, we have a customer who started with us last year. They had about 100 million requests, you know, per month invocations per month on their Lambda infrastructure. They’re at close to a billion now. They haven't changed their operations team. They haven't changed their code all that much. So they've really capitalized on the fact that you can use serverless to build things that just scale magically without having to worry about it. You want all that, but at the same time, you will make sure your security concerns are mitigated that, you know, in that same system.

Jeremy: Yeah, and that point you mentioned about using a lot of resources because you're just passing data through without inspecting it, that is another very interesting trend that I think can be quite a problem. But it’s also not impossible to protect against, either. If you're streaming something into Kinesis or you're putting something into an SQS queue, there are ways for you to validate or verify, and it might go beyond just validating the signature too, but looking more specifically at fields like phone numbers or whatever, and make sure that they meet a certain format. And that rather than you pushing that into the Kinesis stream or into SQS or EventBridge or whatever you're using now, to verify that before, just do some even basic checks on it before you start flooding downstream systems. I think that's a really, really smart point. So what else do you see developers struggling with? Because we now have this thing where, like you said, our production systems aren’t running in traditional sandboxes anymore. And a lot of the burden of security is now on us as developers to make sure that our applications are secure. So what are the things that you see developer struggling with?

Hillel: Yes, I mean, I think the number one thing we see a developer struggling with is IAM roles and permissions. This is something that developers have recently inherited. I wrote last year about a company we had talked with, and are working with, where the developers all got together and said they were going to quit unless the security team gave them ownership of IAM roles. And the developers made an interesting point, which was we can't do our jobs if we have to go to security every time we change IAM role, because we're now talking about IAM that governs 5000 API calls we could make, you know. In a world where we're expected to move really rapidly, we have to own that. And I think that makes a lot of sense. Developers aren't necessarily equipped to make those decisions. And when you look at how developers work and they say “Okay, now I'm reading from a database. So I'm going to have to put Dynamo something there. I don’t know. Is that query or scan or both? Let's just throw a wildcard at it because that's the safest thing to do. And then, which resource? I mean, the table name depends on whether I'm in staging or production. So let's just put a wildcard on resource because that's the thing that's going to make sure I don't break,” right? To a large degree, I understand developers. It's the same reason why the second most common configuration for the duration of the Lambda function is the maximum. The first one being the default, right? Because if I'm going to change it, I'm going to change it to the maximum. That way I don't have to worry about a timeout; that's great. So I think the same way we see people struggling to deal with security configuration in a way that really, you know, meets least privilege and minimizing risk, I'm not sure I blame them. I mean, as a developer, I also I want to move rapidly, and I want to get things done quickly, and I'm now being, you know, pushed even ever harder to write code faster, deploy code faster, less testing, more automation, more deployment. So that's the thing where I think people struggle the most with and it has a huge impact, right? I mean, negative and positive. As I said earlier, properly configured IAM roles on lots of little Lambda functions will give you a tremendous amount of security joy. It will melt away huge swaths of your attack surface. And then at the same time, when you discover that your entire account is using a single roll, by the way, sometimes because security people demanded that they control that one roll and it's got IAM wildcard Cognito wildcard, Dynamo wildcard, S3 wildcard on it. That's going to put you at a huge amount of risk. So I think that's the place where we're saying, “Hey, developers go do this,” but we're not empowering them to make an easy, good decision and as security people, we don't necessarily even know what the right decision is, and that's a real struggle.

Jeremy: And I think that this idea of the star permission too — and I don't know if it's just an education thing, but it almost seems like it's sort of becoming a joke, like everybody's like “don't use stars,” but, obviously I totally agree with that — but when you do use star permissions, what are the risks of doing that? Because again, if you were running, let's say you're running, a EC2 instance, whatever permission that EC2 instance has — which is probably a lot because it has to interact with a bunch of different services and connect to the databases and do all these other things, plus, it probably has its own role that has access to different parts of the network, you know — why is that different than opening up permissions on, say, a Lambda function?

Hillel: Sure. So fundamentally obviously it's not really different. And a star here and a star there are equally bad and I've seen organizations that have, you know, policies in place to prevent stars from being deployed, and that's all great. At the same time, we need to recognize that least privilege is always important. But if you asked me whether I would spend 20%, 30%, 40% of my security effort on getting to least privilege on my EC2 roles, I'd probably say no, I want to spend those people on other things, because at the end of the day, those EC2s are still going to have pretty big roles. So I would take a quick stab at, you know, carving out stupidity from those roles, and I'd probably accept the fact that having extra privileges there, it's probably not terrible compared to what I could use those other people to do configuring security groups or doing audits or configuring my WAF. But when you go to a Lambda function, which probably does one thing, accesses the same resource is every time, probably just reads from that one table, and writes the other table, the difference between “*” and ListTags is incredible. Right? You know, we've seen functions that literally only need to list the tags on resources and they get a star because nobody wants to figure out what the right permission was for list tags, which, by the way, is ListTags. So it's not that hard to figure out in this case. But that star means that that function can not only list tags on, say, IAM or, you know, on an EC2 instance say, but it can delete it or created or spin 10 up or spin 1,000 up, right? And so that's, in the world where this function literally needed to do one or two things, and you gave it everything, that's really challenging. The other thing to remember is that wildcards are not categorized by risk, right? So it's not like someone said, “Hey, there's really risky star and a little bit risky star and pretty benign star and totally benign star,” and you could just say, “I just want totally benign star,” right? There's a little bit of read and write stuff you could do, but for the most part, a star somewhere gives you access to something really benign along with tremendous risk. And so you really have to be mindful about that. And again, yes, you should be mindful about that in EC2 as well. You shouldn't use stars in EC2 as well, but your mileage will vary it when you do it on small resource is like Lambda functions or Fargate containers, it's going to really, really do a lot more good for you.

Jeremy: And especially the Lambda function, right? Because they're ephemeral, and they can be used over and over again. So if there is an exploit there, that’s sometimes hard to see what somebody may have done to exploit it. And if you give it a star permission, and like you said, I mean, DynamoDB is the example I like to use, if you give DynamoDB star permissions, people are like, “Oh, I can get items. I can write items and oh, I can delete items,” but no, you can delete tables. You can create new tables that you can change provisioning capacity. There's a lot of things that you can do and so IAM is so incredibly powerful. But again, those stars just make it a little bit dangerous and I’m not trying to scare people. It's just one of those things where it's like if you could tell anybody you know the one thing that's probably the most important security piece that you have full control over is IAM permissions. So, learn them and use them correctly. Alright, so what about things like public buckets? Are people still doing that? Because S3 now makes them private by default.

Hillel: Yeah, I think one of the things that we see around the world is that stupidity doesn't go away that quickly and you know, part of this is some things are just easy to do wrong. And I think one of the things we focus a lot on in security is the best security solutions are the ones that make the easiest thing the right thing. That's not always so easy to set up — maybe sometimes impossible. But if you could make the easiest thing to do the right thing, that'll be great. And to a large degree, the cloud providers have done better recently by trying to make some of the defaults better, making the process of making a bucket public harder to do, more friction. You know, again, like we mentioned before. That's great. Yeah, you know, it's easier in the cloud to often configure things. And it's not just public buckets, it’s API gateways, it can be, you know, DynamoDB or AppSync resources that can more easily be set up to be visible from the outside world. So that's something that people have to be mindful of. But I think the most important lesson about public buckets is just because you remember this was a problem five years ago does not mean that problem has gone away. And I think in some cases, you know, we talked a lot about other problems that are large resurgence in, like SQL Injection or other things like that. There’re some interesting reasons why what I call “millennial stupidity” is back. You know, things that we did 15 to 20 years ago, and we thought we eradicated. you know, the malaria of coding is SQL injection, and I think we thought it was all done. And now we're discovering no, it's back. It's back with a vengeance. Maybe because we're writing in languages like JavaScript and Python more than we're writing in languages like C Sharp and C++, but also because we're demanding people to deliver things faster, and the fastest thing you could do is concatenate strings. And so you know, that's going to be the quickest way to your database, and we're not focusing enough on “do it slower, but do it secure.” And we’re making that harder. The hardest thing to do in those languages is to write some complex query that's going to protect you from SQL Injection. The easiest thing to do is concatenate strings. That's obviously a software challenge. So I think a lot of those things, they're not going away so quickly. We need to get better about them. We need to detect them. We need to have a process. Part of problem with public buckets, for example, is sure, I've got public buckets. They’re part of my web app or they're part of my single page, you know, application or something, and they should be public. How do I make sure the ones that are public are the only ones that should be public and nothing else is public? And I think a lot of that has to do with hygiene and posture. You could automate a lot of that. But you also have to just be diligent in mindful about a lot of that.

Jeremy: Right. And I think that the mistakes that people make on a regular basis are probably the ones that, like you said, we thought we solved 15 years ago with things like WAFs. And honestly, nobody ever taught me about SQL injection, I actually learned it the hard way very early on in my career. So I like the term “stupidity”, because people still tend to do some pretty stupid things like check credentials into GitHub, but a lot of it is probably just ignorace and inexperience as well. For example, there's been a recent movement away from using environment variables to store secrets in Lambda functions. And I think things like this are all smart things that we should be doing to minimize risks, but there are still plenty of posts (probably even some of my older ones) that say that’s okay. You mentioned this idea of hygiene and posture, and I think that probably brings us to this idea of mitigation, maybe? So we have tools. There are tools out there that we can use that can help us mitigate some of these things. So what are your thoughts on how we do this? Can we just put tools in place to try to block some of the stuff and save developers from themselves?

Hillel: My number one cliche would be: let's focus less on mitigation and more on prevention. I know that's super cliche in security, but I think here, one of things that we see a lot is that you get a lot more mileage out of trying to make sure that the things you're deploying are deployed with least risk than you do at trying to chase after attacks. And again, that's not to say we don't need to do both. We will forever need to do both. No amount of proper configuration hygiene is going to prevent every type of injection attack or cross-site scripting attack, or whatever is on our infrastructure, right? We need to mitigate all of those things. But in cloud applications and particularly in serverless cloud applications, the value of hygiene and posture is much greater than it was in the past. You know, whether it's the things we talked about earlier, like just leaving around old stuff that could put you at risk but you don't need, or it's getting IAM configured properly, or it's things like setting timeouts to their minimum threshold, if you can. All those things are going to give you a tremendous amount of value in making it hard for attackers to do what they want to do on your system, before you even started looking for a SQL Injection, right? You still need to look for a SQL injection. We still need to run the tools that we’re going to run. But before we get there, before you start worrying about blocking and mitigating kind of runtime attacks, spend a significant amount of energy on: What do I have? Where is it? Do I need it? Is it configured in a way that gives me the least risk? Have I isolated the things I can isolate? Am I doing all that continually? That would be my number one focus.

Jeremy: It sounds like a lot of work, though.

Hillel: Well, yeah. I mean, that's why we get employed, right? Security people need to get paid?

Jeremy: So maybe that is a good segue to this question. Whose job is this? Right, so if the application developer is now the one fighting for IAM roles, they want control over these IAM roles, and you've got different types of events coming in that could have SQL Injection, where again, just good coding practices, we know that takes care of some of those things. But then things like configuring what the timeout should be. That's on the developer now. Things like how many RCUs and WCUs do we need for DynamoDB, or how many Kinesis shards do we need? A lot of those decisions now fall on developers. So whose job is this? And maybe it changes with different organizations, but is it Dev? Is it DevOps? Is it AppSec? Is it Ops? Do we need SREs to come in? Who owns this now?

Hillel: Yeah, I I think in a lot of ways it's all of the above. And I know we've said that for a lot of years. I think the difference is, to your point you made in the question, there's just a lot of stuff that those people in those layers own now that they need to be responsible for. Now, I'm a big fan of the idea that security owns overall responsible for security. And if you don't have a security organization in your business, then you're going to find out sooner or later that that was a mistake. You need somebody whose job it is to care, right? But that person can no longer imagine they can solve the problem on their own. You know, I think we were able to imagine for a while that we could throw up a WAF in front of application and then feel good about ourselves. I think we're recognizing that's not really going to be enough. And so, yeah, we need developers to be empowered mainly to do the right thing in the easiest possible way. We need DevOps to help us automate the process of making sure those things happen. So kind of a trust-and-verify model. Sure, you own IAM. Sure, you go ahead. But at the same time, there's stuff in the pipeline that's going to make sure that if you go too far left or right, there's guardrails in place to help put you back on track. Then security needs to know that even though they put some of that stuff in the DevOps pipeline, and that's supposed to give them a lot of good hygiene and posture, stuff will still make it into the cloud in a way that's not ideal, whether it's because it looked okay when it was deployed, and then later on, we discovered it had a third-party vulnerability we didn't know about; or because it was grandfathered in; it got some waiver; it bypassed something; it's been there, etc. We still need to worry about okay, what's actually happening at runtime? Can we know where our risk is and can we go deal with that? And can we still, yeah, look for SQL Injection. Look for code injection. All those things still have to happen just in a way that lives it well in a serverless applications, scales with the serverless application, doesn't get in the way of a serverless application. So yeah, we need to layer all those things on, empower everybody to own their layer, and make sure somebody else is responsible for verifying everybody else's job.

Jeremy: Yeah, I think that a good way to approach it. So let's talk about tools just for a minute. And I know there are a whole bunch out there. Protego obviously is one of the security tools. What does Protego do? How does it mitigate different parts of this?

Hillel: Sure. So first, let me mention that they're obviously lots of tools out there. There are a bunch of tools that you absolutely should be using that come from the cloud providers. So Azure, AWS and Google all have very robust suites of security tools that are available to you if you're in the cloud and you definitely should be taking advantage of those tools. They're going to help a lot in a lot of the areas that you're in. So, you know, in AWS World, you should be looking at using things like WAF and Shield in places that it makes sense. You should be looking at using CloudTrail for auditing certain things and GuardDuty will give you some visibility to anomalies and certain types of parts of your account. So those things are all really great. Where we come in is we really come in trying to bridge the gap between security and dev and infrastructure and code. So we’re kind of saying, “Okay, in a world where the key asset you deploy or in code and API and where you're going to put security around code and API, not at the infrastructure operating system level, how can we come in and try to drive security in a way that's meaningful?” Where we can say “Okay, developers can do what they want to do. Security people can have their policies. Security people can know that they're driving towards least privilege and optimal security configuration and runtime policy without having to close their eyes blindly, and you know, hope. Just throw a WAF in front, and not configure it.” I think the single hardest problem to solve in the AppSec world is how do I configure my WAF and there're usually, you know, three answers. One is, don't bother. It's not going to work, just run it and kind of DDoS and basic attack mode. The second is, yeah, put a lot of people on it, have them constantly evaluate what the application does and try to get the WAF continuously configured. And obviously that gets more and more challenging as the application changes more and more rapidly. And the third is kind of these learning mode WAFs that exist where it's like, let me learn the application, figure out the right configuration, and the challenge there in serverless, and cloud native is things don't run statically for long enough for that to work. You know, by the time three or four weeks have gone by and you've learned what the API should be doing and not doing, those functions have been redeployed seven, eight, 10, 12 times, changing their functionality so you kind of have to keep continually trying to relearn before you can block anything. So to a large degree, what Protego is focused on is, on the one hand, automating least privilege, automating risk mitigation and automating is key. We talked earlier about how much work it is. You mentioned that, right? it's true; it’s a lot of work, but really, the way to deal with security in the modern world is to automate 99% of workload. Set up the right policies, set up the right tools, set up the right processes and let those do most of what your job is. And you spend all of your incredibly taxed time trying to deal with the 1% of things you haven't figured out how to automate yet. So we really try to automate the process of IAM role, IAM role generation vulnerability scanning, looking for keys in the code, and all sorts of other things that people are doing more and more. But mostly I'm trying to make sure that before it hits the cloud, it's configured optimally. After it hits the cloud, we continue to monitor it and make sure it's configured properly. And in the second half of what we focus on is what is runtime application security in a serverless world in a way that takes advantage of serverless. That scales with serverless. You know, we really tried not to build a WAF for a function. I think that's the easiest cliche to fall into. And it's not to say that part of the solution is not a WAF for a function. So I'm kind of contradicting myself. But it's really more to us about focusing on “hey, what is the best way to get security?” And so yeah, part of securing an application in serverless is securing each function and part of securing each function is making sure the inputs and outputs, you know, are validated and make sense. But a big part of it is really focusing on application behavior. And that's really where we're very powerful. Because one of the things we can really do is we can build kind of, I call it, “layer 8 microsegmentation.” It's microsegmentation at the level of API calls, process creation, interaction with third-party resources, things like that, where we can automate the process of figuring out by understanding what the code does, what the right configuration is, and we can really automatically build the proper microsegmentation around each function of your application simply by having an algorithm understand what the code wants to be able to do, understanding what the code is doing at runtime and saying, “Okay, I could build a whitelist for every one of those behaviors, and I can literally lock the function in a cage that says you can do exactly what you need to do. Not a drop more. Not a drop less.” And then above and beyond that, we let security come in and say, “You know what? I don't care what a developer wrote. I don't want anybody to be saying. IAM create user. IAM delete role. Those are things that need special exceptions from me. So I want to apply a policy above that.” So Protego lets you kind of automate all that, build your policies, deploy them, and really kind of sleep at night for the most part.

Jeremy: Awesome. Alright, well, listen, Hillel. Thank you so much for joining me. If people want to find out more about you and Protego, how do they do that?

Hillel: Yes. So a bunch of places. Obviously you can go to our website, and there's quite a bit information about what we do and how we do it. And, you know, you can see a demo. You can actually sign up to use the product. A bunch of clicks, you could start playing with the product and see what it does in your environment. I try to post on Twitter both at @hsolow and @ProtegoLabs, and as well as, you know, on LinkedIn. We've got a blog on Protego on the Protego website that you can take a look at. We try to really ah impart our findings and our wisdom both from a security perspective and a serverless perspective and just try to keep people interested. We've got a podcast, which you're going to reciprocate, and come on real soon, where we try to focus on some of the things that are new with the security angle in the serverless world and in the cloud native world. So there's a link on our website to that as well. And I think hopefully that'll be interesting for people as well. I mean, I see the ecosystem growing. I see more people interested, and I think that's it's reflective of the fact people are figuring out how to use this stuff to their benefit and hopefully, how to do security as well there.

Jeremy: Right, and I think the education piece of it is huge. And the more content we can put out there, that hopefully gets people thinking about this stuff is great. Alright, awesome. I will get all that in the show notes. Thank you again.

Hillel: Really appreciate. Jeremy. Thanks for the opportunity. See you soon.

View Details

About Slobodan Stojanović

Slobodan Stojanović is CTO of Cloud Horizon, a software development studio based in Montreal Canada, and the CTO of Vacation Tracker. He is based in Belgrade and is the JS Belgrade meetup co-organizer. Slobodan is an AWS Serverless Hero, Claudia.js core team member, and co-author of "Serverless Applications with Node.js" book, published by Manning Publications.

  • Twitter: @slobodan_
  • Vacation Tracker: vacationtracker.io
  • Cloud Horizon: cloudhorizon.com
  • Claudia.js: claudiajs.com
  • Serverless Applications with Node.js: manning.com
  • Blog: serverless.pub

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Slobodan Stojanović. Hey, Slobodan. Thanks for joining me.

Slobodan: Thanks for having me.

Jeremy: So you're the CTO at Cloud Horizon and Vacation Tracker. So why don't you tell the listeners a bit about yourself and what these two companies do.

Slobodan: Yeah, so for the last seven years, I’m a partner and CTO at Cloud Horizon. We basically do services for companies, and we build web applications for them. We worked with start ups and some enterprise companies and things like that. For a long time, we thought about, like, building a product. So last year, we finally started doing that. And we built a small tool that will help us to, so whenever our... we have, like, almost 30 people in the company and it's really hard for us now to track who is on a leave and who will be on a leave at some point. So we built a tool to help us out, to track just that. We don't need the full HR system and things like that. So we used serverless to build Vacation Tracker, our product, which is basically a slackbot and now web application that will help you to manage leaves for your company in just a few clicks, and people can request leaves through Slack and things like that. Besides that, I'm writing a lot about serverless. Not a lot in the last couple of weeks. But before that, I wrote a book about serverless called “Serverless Applications with Node.js” with my friend Aleksandar Simović and I have a couple of Medium posts and a few other articles that are explaining mostly testing architecture of serverless apps. So, that’s it, basically.

Jeremy: Awesome. All right, so I wanted to talk about testing in serverless applications. And so maybe for people who are either new to development or, you know, are maybe used to different, ways of testing. I mean, why is testing so important? Let’s maybe start with that.

Slobodan: Um, probably the best example I saw so far is the story, one of the stories from the previous book from Gojko Adzic called Humans vs Computers. So there was a guy somewhere in, I think, US that wanted to have custom plates for his car and he tried to fill out the form and he was into sales and boats and things like that. So, he had three choices in that form. First one was like ‘boat’, second one was ‘sailing’ and he didn't want the third choice, so he tried to leave it empty. He wasn't able to do that, so he typed ‘no plates’ or something like that. The first one was occupied, the second one was occupied, so he got ‘no plates’ plates. That was fun so he kept them and after a month, he started receiving a lot of tickets for parking because, you know, in the software for the guys that were filling the tickets for parking, no one predicted that there will be a guy with no custom plates or without plates. And whenever they don't get the plates, they just typed something. And most of time, they typed ‘no plates’ plates. So with testing, they would probably have handled that thing much before it hit the production and everything. So our applications are not perfect. There are so many things that when people start using, that can do in our applications. And when we start testing first we do some analytics of our application and think about the end users and the way that they will test our application. And on the other side, when our application grows really big it's easier for us to, like, be sure that we didn't break something. Unless we wanted to break it, of course.

Jeremy: Right. And when people are building applications, too. I mean, this is something where I mean, I'm a big fan of test-driven development. Where you you actually write your tests first and then you write code to make the tests pass, right? Because then you know what the expected outcomes are, as opposed to, you know, kind of going back after the fact and trying to make some changes there. So, let's talk about testing with serverless, right? Let's get a little bit specific. So is there, or are there, different things that you need to do or, sort of, what's different about testing serverless versus maybe testing a traditional monolithic application?

Slobodan: There are a few different things but, in general, testing is still the same. You want to check if your application works and the way that you want it to work. But some of the things are not your responsibility anymore. For example, infrastructure is, like, the responsibility of your vendor, such as AWS or Microsoft or someone else. So there's no point in really testing that part because that's not really something that, they have their own testing, things like that. But you still need to be sure that your business logic is working in a way that it works. And also all serverless applications are basically microservices, that they're working together. Most of the time, you don't have one monolithic application that is just uploaded to AWS Lambda or something like that. Most of the time, you have, like many different functions. For example, in Vacation Tracker we have, I think more than 80 functions now that they're working together. So it's really important to be sure that all those small services are working together the way we want them to work together, and that our end users have a decent experience and that they can use our application.

Jeremy: Right. And so that the types of tests that you would run I mean, you're still gonna do unit testing, right?

Slobodan: Yeah. So, basically, these types of tests are not that different. We still want to have unit tests because they are still the fastest. But, we also want to have integration tests, that are more important than ever, because we want to check if all these things work in integration, not just our code with the database, but maybe two different Lambda functions that are talking through some SNS topic or something like that. And of course, you wanted to have some kind of end-to-end tests. And maybe UI tests if your application is heavily using UI and things like that. So you still want to keep all these different types of testing that we had in previous non-serverless applications.

Jeremy: So, the other thing I think that's important about UI tests and integration tests is with serverless, they're not quite as expensive as they were before, right?

Slobodan: So yeah, one of the things that I used to kind of show that these testing pyramids. So, testing pyramid was defined by, I think Mike Cohn in his book Succeeding with Agile, a couple of years ago. Probably much more than that, actually, 10 years ago or something like that. So, he tried to explain with the pyramid which tests are the most important for your application, with traditional non-serverless applications. They're probably not traditional, but whatever. Let's call them traditional just to illustrate the point. So you should have a lot of unit tests because they're fast. If you want to test your code, you don't need to speed up the database. You don't need to have even the server or anything. You can just run them on your local machine, and that's it. They're the fastest one, and, of course, the cheapest one because you don't need the infrastructure for them. Then you need the service layer or integration test that will just test your code against the real database and some other things. And these tests are much lower because they require you to have the infrastructure, they have latency and many other things, and they're more expensive because you need to have that infrastructure. In the past, infrastructure was really expensive for some applications, especially for big applications. And, many times so servers that they're working just during the day for test environments and dev environments and they're shut down during the night and things like that. And it's even more expensive and slow when you want to run end-to-end or UI tests because you need to have everything in some environment, and then you need to simulate clicks and many other things. So you don't want to have, those are the things on top of pyramid that are the most expensive and the slowest , so you don't want to have a lot of them.

And in serverless, I like to call that serverless, the testing pyramid. It's basically still a pyramid, but more like Mayan pyramid without, it doesn't look like a triangle anymore because unit tests are still the fastest one and the cheapest one. But then we have integration tests that are cheaper than ever and faster than ever, because we can run them in parallel and speeding up a new Lambda function or DynamoDB table don't take like minutes, it takes a few seconds or even less. And, you pay just for execution, so it's really cheap for you. Most of the time, it will be free. So you can have more of them, and you need to check like integration between your services, that's why you want to have more of them. And finally, for end-to-end, again, tt's cheaper and faster than ever, because it's easy for you to speed up a new environment that looks exactly like production, where you want to test everything. But you can also do some crazy things, like putting your browser inside your Lambda function or something like that. And then you can, instead of running just one track off your UI test, you can split them into, like, 10 different tracks and running 10 different Lambda functions and the cost will be the same because you pay per requests, so it doesn't really, so 10 requests that will be sequential will cost you the same as ten parallel requests. So it's easier and faster.

Jeremy: Yeah, It happens, it happens much faster. So you had mentioned too, this idea of multiple environments, right? So that's actually another really cool thing that you can do in serverless where each developer can have their own environment. And if you wanted to just spin something up and do some manual tests or some even some integration tests in multiple environments, it's a lot easier to do that with cloud then it was to do traditionally.

Slobodan: Exactly. So, especially when you're using serverless applications, and when you have some kind of infrastructure as a code, like CloudFormation or something else, Terraform and things like that. But you can, most of the time, It's just one command that will speed up your new environment, and it will take you like 2, 3 minutes, maybe 10. But it's not like hours or days anymore. So it makes sense, and you don't pay for environment that no one is using, so it makes sense for you to have different environments for each developer, but also different environments for your manual testers. So every person in your QA team can have their own environment, but also, if you have a big feature that you want to test for a few weeks or something like that before you release it, you don't block your test environment. Instead, you can spin up a new environment, and you can still have your test environment for some other smaller futures that you want to ship in the meantime.

Jeremy: So, are there certain things in serverless that you see I need to be tested differently? Or maybe the question is, what is it that we're testing? You mentioned integration because obviously you're hitting up, maybe against a DynamoDB table or different SQS queues or different SNS topics. Are there specific things that you would test in a serverless application that you wouldn't test maybe in a normal application?

Slobodan: So it depends. You're testing things that you think can be risky. For example, when I'm in Claudia.js core team and, when my friend Gojko started Claudia.js, those early days of serverless, it was really hard to start with serverless. So for serverless applications at that moment that you had, like, 10 lines of NodeJS code that will do something and then, like if you want to automate upload, it was like 200 lines of bash script that will do that. It was obvious that risk park is not just in your code anymore, it's also in the process of set up because you need to upload the code, which is not that hard. But you need to set up all the permissions and do many different things to connect everything and to be sure that everything works together as it should. So, yeah, risk is shifting from your code to different parts. I don't think you should build your own deployment library anymore because we have so many awesome libraries that are well tested. But you need to use something that is really well tested to deploy your application. So you need to be sure that your integrations are working fine, that your permissions are working fine and many other things, many other problems that you didn't have before. Like microservices in serverless. And there are certain risks, of course, that you want to cover.

Jeremy: Yes, so let's, I want to talk about those risks. Actually, maybe this gets us into how do we actually test these applications? But maybe we start first, I mean, you mentioned deployment services or deployment frameworks. I mean, obviously, Claudia.js is one, Serverless Framework, SAM for AWS, and there's a whole bunch of them that are out there now. Architect framework. There's a lot of them. And again, the preference for which one you choose is based on a number of factors. But, maybe additional tools for actually doing the testing, right? We just use traditional Jest and Mocha and things like that, right?

Slobodan: Yeah, for our applications we mostly use Jest for writing tests. Before that, we used Jasmine, but Mocha or anything else will work for JavaScript. The same for Python, you can just use, or any other language, you can just use the tools that you used before. And of course, if you want to run end-to-end tests and things like that, there are so many cool things that you can use such Cyprus, for example. It's really cool because before that we're telling you, that's basically another like, I don’t know, another set of skills that you need to have to run your end-to-end tests and things like that. Right now, you can do like, testing, so end-to-end testing by just writing JavaScript, which is awesome for us, because we, our full stack is basically JavaScript now with Node.js, some Typescript and many other things. So, yeah, basically, there are some new tools that can help you such as, I know there were a few. I don't know the names, but there were a few, like testing and CI libraries emerging with Serverless. But most of the time, you can just use your preferred tools and it will just work.

Jeremy: Awesome. All right, so let's get into these risks, right? Because that's one of the things I think when you, when you're talking about testing or testing strategy, we need to know what to test and that's an important question. And you kind of use the word ‘risks’ here, which I think is perfect. It's a good word, because again, you have liabilities, right? Every time we write a line of code, we have introduced some sort of technical debt or some sort of liability that we are now responsible for. If that ever changes, that could break a whole bunch of things. So maybe outline some of these risks, so as you see them.

Slobodan: So, yeah, in our book, my friend Aleksandar come up with four risks, and I really like that explanation on and, like, list of risks that that you need to cover. First one was configuration risk, which is basically if you have the right permissions to do something and if your application is configured to talk to the right database and things like that. Then you have technical workflow risk, which is basically, if you’re handling your errors correctly, or if you're returning the correct response to API gateway or some other tool that you're using or service that you're using. Then you have business logic risks, which is basically related to your code. If your code is working the way it should and things like that. And finally you have some integration risks. With configuration you configured the right services, but integration risks will tell you and will cover the part of these services working together, so if your Lambda function is writing the right way and can talk to your dynamodb in the right way and things like that. So, in order to build a serverless application, in my opinion, you need to consider all those four risks and to make sure that you're testing all of them. And some of them can be tested with like unit tests like part of business logic and, of course, things such as technical workflow risks and things like that. But for some of these things, such as configuration risks and integration risks, you need to have integration tests, of course, but sometimes you even need to have, like end-to-end tests to be sure that everything is working together fine and that your users will be able to use the application.

Jeremy: Right. And actually, I mean, I think that's a really important point where you know, when we write traditional applications and again, I don't know if that's the right word to call them ‘traditional applications’, like you said. But you know, you usually have your databases set up for you. So you know what the URL of the database is and you know the username and password and so forth. But with infrastructure as code, we're spinning up, you know, a separate database or a separate DynamoDB table for, you know, just for this Dev environment or for Developer A's environment or whatever. And so you run that risk of you spin something up, and you might be able to access it locally because you have a profile that you're running on with admin permissions or something. But then once the code starts to run, all of a sudden you can't access the database, right? Or you can't, or when an error happens, you're not catching that error the same way in the cloud that you might be catching it locally or whatever. So you have all of these different risks. So it is super important that you do spin up an environment to actually test it in the cloud. Otherwise, you could run into all kinds of problems like that.

Slobodan: Yeah, of course. And I saw that so many times, sometimes you try to cut the corners and just deploy something to your test environment and just see if everything works. And then you end up with, like, now digging through cloudwatch logs and many other things to be able to figure out what’s, I don’t know, what’s failing and things like that, right. But yeah, you don't want to waste your time that way. You just want to be sure that things are working together in a way that they should.

Jeremy: Yeah. And so when you're using things like DynamoDB or SNS, you know, and you're writing tests against those, what's your strategy for doing that? Are we using real, do we want to make sure we're always using real cloud versions of those? Do we want to mock those locally? Do we want to run local versions? How's your, how do you deal with that stuff?

Slobodan: So in the early days, when I started working in serverless, I tried to just do the things that I did with non-serverless applications. So my first try was like to install DynamoDB locally and use it as a local database and test against it and things I got. But then there were so many different services such as like Cognito or SNS or many other things that they can, cannot just install locally. So there are some things that can simulate them. But these are simulations. You don't want to run your tests against simulations because they will not give you the right results. And then the second, my second try was basically to mock these things. So I tried to find some complex mocking libraries and things like that that will mock everything and return realistic results and things like that. And even that is like leading to so many errors. And some things are not mocked. You have some special things in your code or we're using the old library, so you need to send some pull requests and who knows what. So these things become more and more complex, and in the end we just decided not to do that. Instead, we want to run our tests locally. Unit test mostly. And whenever we want to test, to run integration tests, I want to test my code against real DynamoDB, which is on AWS or real SNS or real services that are on AWS and that they're working in the cloud. So basically, yeah.

Jeremy: And so what I do is, in my local applications or my unit tests, I do like to run mocks, but I hate those services that do the mocking for you. I mean, I think there's one, like AWS mock service or something like that. I actually just try to capture the actual response from the AWS event, and then I usually save that and use that as a way to quickly run unit tests. But then when you, which is great for just running tests locally and try to figure out if you're making some business logic happen, you don't wanna have to always be calling SNS or DynamoDB in order to do that. But I do find that once you move past those and you go to the integration test, you do want to test it. But is this something where when you run the integration test, do you have to get test every single function? Or do we only just have to really make sure that we're able to connect to the services?

Slobodan: Yeah, that's a great question. And that leads us to another big topic. And this is architecture because you can try to test everything against the real. So, for example, in our code in Vacation Tracker, I mentioned these 80 functions, and I think at least half of these 80 functions are posting some messages to SNS. If I try to test every of these 40 functions against a real SNS that will take ages because for SNS, you need to set up like something that will listen to these messages and things like that, so my test shoot would run for, like, hours or something like that. So instead of doing that, what we do is using ports and adapters or hexagonal architecture, to write our code in a way that we can just basically split everything into different ports and adapters and then we can basically test our business logic against different adapter. So before we do that, we should probably explain ports and adapters and hexagonal architecture because some of the people that are listening to the show will not be, will not know about that. So basically, it's a simple architecture that looks similar to some other architectures, for example, clean architecture and a few others. But the idea behind it is simple. It allows your application to be run and tested in isolation of its eventual environment. So basically your function can be run as a local function, and it can talk to some different, and each of your functions. So your business logic exposes some ports. Let's say that I have a function that receives some messages from Slack and then posts some other message to SNS and saves some data to a database. Instead of connecting to real Slack to get the message to listen to these messages directly from Slack and posting directly to SNS and saving directly to DynamoDB, I want to have three different ports in my code. The first port will be event port that will be able to listen to some different events, and I will have a different adapter for different things, I will have an adapter for Slack that will work in production. But for local testing I will have some kind of local adapter that they contribute from my command line or something else. Then for posting a message on SNS, instead of posting a message directly to SNS, I will have a port for that that will help some interface such as message.post or something like that. And I will have adapter for SNS for production and different adapters for local testing. Maybe just sending a JavaScript event or something like that and, again, for the database, something will assume that data will be stored to a database with some method of database adapter and then database adapter can be DynamoDB adapter that will save the data to DynamoDB, or MongoDB adapter, maybe, or just in-memory database adapter that will save some data in memory and allow us to test everything.

Jeremy: All right, so that was just a ton of information. So, let's break it down a little bit. So we start with let's start with the idea of the ports, right? So when we're building those applications or we’re building some bit of business logic, obviously it needs to interact with something else, it needs to connect to the database like you said, it might need to send a message. Might need to receive a message. The problem is, is that if we build that functionality directly in there and say, you know, maybe we include the AWS SDK with DynamoDB service in there, and the function is supposed to query the database, grab some data, and then maybe send it off to SNS. If it does all of that in that one function, you’re married in a sense, to those services. And you can't, they're not easily testable. That's basically what you're saying, right? So the idea of these ports are to say that we're going to include sort of like, almost like a generic library, maybe like dependency injection. Is that another way to sort of think of it?

Slobodan: Yeah, something like that. But it's called ports and adapters for a good reason. Whenever you're traveling, for example, if you travel to Europe, you want to charge your laptop or your phone. And here in Europe, we have different power plugs for the walls, so you're not able just to plug in your computer in a wall in your hotel room and charge it. Instead, you will need a different cable to do that. But you don't want to buy new cable every time you travel to different countries, because you will end up with a lot of different cables, and it's really expensive and a non-scaleable way to travel. Instead, you have a smaller adapter that will just convert your power plug. Do something that is compatible with a power socket in your hotel room, and that will be small enough that you can carry in your pocket or wherever you want. And it's really cheap. Instead of just buying the cable that will cost you more than $100 this will cost you like $2 or something like that. It's the same with your code. Basically, your business logic doesn't really care. What's your database? You want to save some user? Or let's say that inside our Vacation Tracker we want to save vacation requests. Why would my code even know about my database? It will want to save that vacation request or leave request somewhere, and that needs to be saved in some database. But it doesn't really matter for my code if that database is DynamoDB or I don't know what, maybe MongoDB or something else, as long as that saved and can be read from the database, that's it.

Jeremy: Right. And that's this idea. Like you said, this message.post or the interface. So you would build, your adapter essentially would be a sort of like in the case of the database, like a data access layer, or how we would maybe think of it that way where, you know, DynamoDB might have ‘get item’ or ‘put item’ or something like that. But you would genericize that and expose an interface that says item.get or db.get, db.save or something, and then your individual functions can plug into that and then that way, later on, you could actually change that, and your business logic or your code wouldn't have to change. You would just have to make sure that you had a compatible interface that you could swap out for, I should say compatible adapter, with the same interface that you could swap out later on down the road.

Slobodan: Exactly. So in our code, everything looks much more cleaner now because we have some lambda.js file or lambda.ts file. If you're using TypeScript, which, it's responsibility is basically to require those different adapters or repositories. How we call them, for example, DynamoDB repository to create instances of these repositories and pass them to our main.js or main.ts function, which is basically our business logic. And inside that function, main function that is business logic, we don't require any other repositories and things like that. We get all of them through, like arguments, and we just use them. We assume that whenever we have notification, we assume that that notification library that we got will have notification.send method that will just send some texts as a message to some topic or whatever. Basically, that helps us to like, move away our business logic from integrations and everything. And on the other side, it also helps us to write tests and everything because, as I said, we have a lot of functions that they're doing similar things, like sending messages to SNS or something like that. So instead of testing each function against that, we can now have integration and local unit test for that SNS repository that is posting to real SNS and has that, again, SNS topic and see that that works. And then in our business logic, instead of SNS repository, we can send different notification repository, which can be like a local notification just JavaScript event or something like that. And we can just test against that because that's much faster. And it's still integration between your business logic and SAM adapter.

Jeremy: Yeah. And then, actually, I think that's a really good point about the unit testing, too, because especially if you're doing mocking and you're trying to do some of those things you're always overriding or your stubbing API calls and things like that, and that just gets a little bit complex and kind of messy. And it's easier to, or I think would be easier, using this style of architecture to just swap it out so that when your function calls that you've already got that sort of set up to send back the right messages based off of, you know, the different inputs that it sent. So that seems like a really smart thing to do and then the thing about the SNS topic, or like you're doing this cloud integration piece, that's another thing where if you're, you don't want to test this probably in production, right? You want to test this in some sort of staging environment. But if you're even trying to do integration tests for individual functions, or isolated business logic, you could just change your adapter so that that's using a cloud resource. But it's just using some test resource so that it doesn't, you know, doesn't mess up anything else in your production environment.

Slobodan: Exactly. So I think the best example for that was, a few weeks ago, on a Twitch series that Heitor Lessa did for like, Built on Serverless or something like that. He was showcasing, like with different people, how to build real world applications on serverless and I think you helped him with unit tests.

Jeremy: I was on there, yes.

Slobodan: Then after that, I was working with him on integration tests, and we actually did this exactly the way that you mentioned. So we didn't want to spin up a new environment every time we want to run integration tests, because we want people to be able to run their integration tests when they're doing development on one function or something like that. So whenever you're doing your development, you run your unit tests, you check your integration test, and after that, you push your code to somewhere where something will do end-to-end tests before it goes to production. But with integration tests, let's say we want to test some DynamoDB table. Basically what we want to test in integration tests is that our code will work fine with DynamoDB table, because with unit tests, you can just assume that everything you send to DynamoDB table will be saved. But maybe you're using attribute that is called status or something like that. Some reserved word, and then it will fail with the real DynamoDB table. So what we did was create a DynamoDB table just for tests on the beginning of the test suite, then running tests against that new DynamoDB table. And then when all tests are finished, then we just destroyed that table. So we don't need to deploy all the environment and things, so that we just want inside the code to create some temporary table test against that and kill it in the end. And that's it for integration test.

Jeremy: That's great. All right, so maybe this would be helpful if we could give some examples of, like why would this be beneficial? Like, how has it benefited you? And I know you have a story about MongodDB in there that maybe you could share.

Slobodan: Yeah. So, the first thing we already mentioned thing with SNS and the issue with testing SNS, posting to SNS, is not with posting the message. That's really fast. And I would be able to post as many messages as I want, but the issue is actually validating if that message was posted. So if we want to validate that then I need to have a Lambda function. Or maybe as the SQS that’s subscribed to that SNS topic so I can see if something was posted to that SNS topic. So that takes a lot of time to set up. And it's a complex thing to set up. So I don't want to test that in each of my functions. But the other thing that you just mentioned is with MongoDB and DynamoDB. When we started our product, the early version used MongoDB and some Express application for part of the app and the rest of it was serverless. And we slowly migrated from, first, we removed Express. And then we tried to remove MongDB, and we still have some parts of the application that are talking to some MongoDB. But most of it is now on DynamoDB. But, it took us a lot of time to migrate everything because of, like, writing everything and testing databases itself. But when we had different repositories for DynamoDB and MongoDB with the same interface- so whenever we wanted to say save leave request, we had db.saveLeaveRequest that got the same params and returned the same value in the end. We were able just to swap these two things. We went to our Lambda.js files. And instead of requiring MongodDB, we just required DynamoDB repository, and our business logic just worked because we don't change the business logic itself. We just change the repositories in the end. So we were basically able to swap the databases for a lot of things without, like, changing our business logic.

Jeremy: Awesome. And so just question, too, in terms of how specific you would get with some of these commands. So, I like to take with DynamoDB, for example, I like to basically take every access pattern which might be ‘getUsers’ or ‘getUsersByXyz’ . I like to write a specific function for each one of those. Is that something that you would do as well? Would you get very specific?

Slobodan: Yes, we do the same. The only issue that we found with that approach is that after some time you have a lot of different small functions that are talking to your database. So basically, we were using classes, some kind of class that is talking to, let's say, DynamoDB. But we have a base class. And then on top of that, we have a few different classes that they're extending that and providing interface to do, for example, things related to leaves, things related to users and things like that. Because otherwise we were in the situation that we got so many functions, and it was really hard for you to, like, see if the thing that you want to do already exists or not inside your code base. So, yeah, that's tricky.

Jeremy: Yeah, I have run into that as well. So, all right, so what about, you have an example, too, with Vacation Tracker. So you use Slack right now, but that's something with this architecture, to be really flexible if you wanted to migrate to something besides Slack or you wanted to add more services too, right?

Slobodan: Exactly. We started with Slack because that was a way for us to validate the MVP. And now we have, like, more than 200 teams paying for our product, so we kind of validated that, the idea, we know that people are ready to pay for this kind of product. And in the past few weeks we were getting more and more questions related to Microsoft Teams because it seems that Microsoft Teams is, like, getting more and more popular now. So, one of the things that this architecture will allow us to do, and we're actually considering doing that in the next few, like, weeks is extending our product not to work just with Slack, but to work with Slack and Microsoft Teams. Because with this approach, we have interfaces such as like postMessageToSlack, like post, not even post messages, like, but post, for example, request, to that communication channel, and our communication channel can be Slack. Or we can basically add another communication channel that is different in the end. So we need different repository and different adapter. But the port is the same for our business logic that just, like, posting the request somewhere into someone's chat, and that chat can be Microsoft Teams or, who knows, maybe in future, some different things. We even have web applications, so, it's like it doesn't matter for my business logic, where that request is going. But, adapters are there to post that message to different channels and things like that.

Jeremy: Yeah, because all that logic is hidden. All that logic is hidden inside those adapters. And like you said, your business logic doesn't care. You know, and you can split that adapter, you could probably even put an adapter on an adapter, in the sense where you might have, you know, write to one adapter that then decides, maybe based on who the client is, whether it's going to go to Slack or whether it's going to….

Slobodan: We are exactly doing that, actually, because each adapter is basically using hexagonal architecture in its own. So, whenever I want to test different functions inside the adapter, I'm still applying the same principles as I do for my business logic. So, if you want to test your posting messages to a database and you just want to test it if, I don't know to do a unit test of that, you don't want to check that against a real DynamoDB. Then again, you can pass different, like small mini adapters inside and check if you'll get the response with the right format and things like that. It's especially important for, for example, Slack because most of the time you just need to send some kind of JSON and on the other side, Slack will just show something. And it's really hard for us to do integration tests for that. We can do some end-to-end tests, but not integration tests. So it's really important for us to be able to just read that JSON and validate against different things there, documentation or something.

Jeremy: Yeah, and I think that the other thing that this does too, and I'm hoping that we are moving past the, I'll call it a fallacy of vendor lock in, right? But I do think it's kind of important because the other piece of this is data lock-in too, right? Sort of like, once you choose a database technology and you write all these applications that interface with that, you know. Let's say you start with MySQL, and then all of sudden you realize, ‘Hey, MySQL doesn't scale when you get to, sort of, Cloud scale.’

Slobodan: It's even with, even with your programming language. Imagine that you started your application with, for example, Java or something like that, and then you realize that you're cold starts are too big or something like that. So you want to move to Node.js or go where cold starts basically don't exist anymore. So you're locked in with your language. Not just with, or with your framework, or with many other things. Our Express.js application was some kind of lock in, so lock in is kind of real. But Cloud lock in is basically not that real because I never heard about, like, company that moved from one cloud vendor to another unless they had some really specific use case and a lot of long history of bad decisions and things like that. And yeah.

Jeremy: Or maybe, maybe like a start up credit or something. A reason why people move.

Slobodan: Exactly. I don't think start ups, so even if you have a lot of credits or things like that, that will not help you if you need to rewrite your application because development time is much more expensive than, like, your infrastructure.

Jeremy: Right. So maybe, that might be a great way to end this. Sort of, like, what’s your advice? I mean, like, start early, you know? Or think about testing right from the beginning? And what's your advice?

Slobodan: Definitely. You should start thinking about testing early, because if you start thinking about testing at the moment when you have some issues and things like that, it will take you a lot of time to optimize your application and be able to test everything. Most of the time, people are trying to find complex ways to test their code without, like, changing their code. So you want to keep your code the same as it was and just find a really complex way to test it with different mocks and stubbing libraries and things like that. But my approach is completely different. I always want to change my code a bit to allow myself to write tests easier, because that will help me in future, to just move to, let’s say migrate, to other things, not really to other Cloud provider, but to migrate to, for example, from one database to another. Or even more important, my application is growing, so sometimes I just want to switch to another version of another service that we're building or things like that. So basically, it's okay to change your code a bit. To help yourself and your team to write easier tests and to be sure that you can check that your code is working in a way that it should.

Jeremy: Awesome. All right, well, thank you so much Slobodan, for joining me and sharing all this complex knowledge about hexagonal architectures because it’s definitely something that I think people should learn and definitely approach. Why don't you tell the listeners how they can find out about you and you have a ton of things going on, so, you know, feel free. Tell us all about how we could get in touch with you and these other things you're working on.

Slobodan: Yeah, so you mentioned complex knowledge. It's like complex things about simple things, because hexagonal architecture is quite simple in the basic of it, it's really simple. And then we add a layer of complexity around everything, of course. So, thanks for having me. And yeah, of course, people can find me on Twitter. My Twitter handle is @slobodan_. I'm Tweeting a lot about, like, serverless and things like that. And we're also organizing a Serverless Days Belgrade conference in September. So, if anyone of your listeners want to come to Belgrade, Serbia in September, weather will be really nice and we have a nice lineup of people. You can check that on serverlessbelgrade.com. And beside that, you can check Vacation Tracker, our awesome start up application, vacationtracker.io or my other company, Cloud Horizon at cloudhorizon.com. And yeah, I should probably mention Claudia.js, which is claudiajs.com, which is still one of the best frameworks for beginners because it's not a real framework. It's just the deployment library. For complex applications, you should probably use something more complex such as AWS SAM or now CDK, or, of course. Serverless Framework and things like that. And I also mentioned the book that I wrote with my friend Aleksandar Simović. The book is called Serverless Applications with Node.js. And we actually have 5 free e-book copies that we can give away to your listeners. And finally, you can read more about hexagonal architecture and many other things on serverless.pub website, where I write some things with Alexander and Gojko Adzic. There are a lot of great articles written by the two of them on that website, so feel free to visit it.

Jeremy: Awesome. So we appreciate those codes, by the way. So I think what we'll do is, I'll share this in a Tweet, but if you share the podcast or maybe go and leave a review on iTunes or something like that, we'll enter people in and give away some of these codes that they can read your book. All right, awesome. So thank you so much for being here. I will get all of this stuff into the show notes too, so that, you know, there's a lot to take in here.

Slobodan: Thanks for having me in this awesome podcast.

View Details

About Gunnar Grosch

Gunnar is Cloud Evangelist at Opsio based in Sweden. He has almost 20 years of experience in the IT industry, having worked as a front and backend developer, operations engineer within cloud infrastructure, technical trainer as well as several different management roles.

Outside of his professional work he is also deeply involved in the community by organizing AWS User Groups and Serverless Meetups in Sweden. Gunnar is also organizer of ServerlessDays Stockholm and AWS Community Day Nordics.

Gunnar's favorite subjects are serverless and chaos engineering. He regularly and passionately speaks at events on these and other topics.

  • Twitter: @GunnarGrosch
  • Webpage: https://grosch.se
  • LinkedIn: linkedin.com/in/gunnargrosch
  • YouTube: youtube.com/channel/UCTltOOcP1UZTpEQKQjjO-5Q

Links from the Chat

  • Gremlin: gremlin.com
  • Chaos Toolkit: chaostoolkit.org
  • Adrian Hornsby Lambda Layer: github.com/adhorn/aws-lambda-chaos-injection
  • Thundra: thundra.io
  • Yan Cui's Article: hackernoon.com/how-can-we-apply-the-principles-of-chaos-engineering-to-aws-lambda-80f87e3237e2

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Gunnar Grosch. Hi, Gunnar. Thanks for joining me.

Gunnar: Hi, Jeremy. Thank you very much for having me.

Jeremy: So you are a Cloud Evangelist and co founder at Opsio. So why don't you tell the listeners a bit about yourself and maybe what Opsio does?

Gunnar: Yeah, sure. Well, I have quite a long history within IT. I've been working almost 20 years in IT, ranging everything from development through operations, management and so on them. Um, about a year and a half ago, we started a new company called Opsio. And we are a cloud consulting firm. Helping customers to use cloud services in any way possible and also operations.

Jeremy: Great. All right, so I saw you speak at ServerlessDays Milan, and you did this awesome presentation on Chaos Engineering and serverless. So that's what I want to talk to you about today. So maybe we can start with just kind of a quick overview of what exactly chaos engineering is.

Gunnar: Yes, definitely. So chaos engineering is quite a new field within IT. Well, the background is that we know that sooner or later, almost all complex systems will fail. So it's not a question about if it's rather a question about when. So we need to build more resilient systems and to do that, we need to have experience in failure. So chaos engineering is about creating experiments that are designed to reveal the weakness in a system. So what we actually do is that we inject failure intentionally in a controlled fashion to be able to gain confidence so that we get confidence in that our systems can deal with these failures. So chaos engineering is not about breaking things. I think that's really important. We do break things, but that's not the goal. The goal is to build a more resilient system.

Jeremy: Right. Yeah, and then so that's one of the things that I think maybe there’s this misunderstanding of too is that you're doing very controlled experiments, as you said, and this is something where it's not just about maybe fixing problems, either in the system, it's also sort of planning on for when something goes down. It's not just finding bugs or weakness, it's also like, what happens if DynamoDB for some reason goes down or some backend database like, how do you plan for resiliency around that, right?

Gunnar: Yeah, that's correct, because resiliency isn't only about having systems that don't fail at all. We know that failure happens, so we need to have a way of maintaining an acceptable level of operations or service. So when things fail that the service is good enough for the end users or the customers. So we do the experiments to be able to find out how both the system behaves and also how the organization, their operations teams, for example, how they behave when failures occur.

Jeremy: Well and about the monitoring systems too, right? I mean, we put monitoring systems in place, and then something breaks and we don't get an alarm. Right? So, this is one of the ways that you could test for that as well.

Gunnar: Yeah, that's a quite common use case for chaos engineering, actually, to be able to do experiments to test your monitoring, make sure that you get the alerts that you need. No one wants to be the guy that has to wake up early in the morning when something breaks. But you have to rely on that function to be there so that PagerDuty or whatever you use actually calls the guy.

Jeremy: And you said this is a relatively new field. You know, when you say new, it's like a couple of years old. So how did this get started?

Gunnar: Well, it actually started back in 2010 at Netflix, so I guess it's around nine years now. And they started a tool or created a tool that was called Chaos Monkey. And the tool was created in response from them moving from traditional physical infrastructure into AWS. So they needed a way to make sure that their large distributed system could adapt to failure so that they can have a failure. So they use Chaos Monkey to more or less turn off or shut down EC2 instances and see how the system behaved. So, that was in 2010, then I guess the first chaos engineer was hired at Netflix in 2014. So about five years ago, and in 2017, the team at Netflix published a book that's on Chaos Engineering. I think it's out on O'Reily Media, which is like the book on Chaos Engineering today that's used by most people who use chaos engineering.

Jeremy: So we can get into some more of the details about running the experiments, so I want to get into all of that, but I'm kind of curious, because this is something where, especially with teams now, and maybe as we get into Serverless too, you've got developers that are closer to the stack, there may be less OPs people or more of this mix. So is this like a technical field? Is it the devs that do it, the OPs people, like who's sort of responsible for doing this chaos engineering stuff?

Gunnar: Yeah, that's a good question. I would say that it's a question that's being debated in the field as well. Where does chaos engineering belong? And I think it's different in different organizations. Some teams have specific or some organizations have specific teams that are only working with chaos engineering like Netflix, like people at Amazon as well. But other organizations, they use chaos engineering and it’s spread out in the organization. So it might be operations that works with chaos engineering. But it also might be a DevOps team or just developers as well. But to do the experiments, you need to involve more or less the entire organization. So you use people from from many teams.

Jeremy: Right. Cause you're gonna run, in some cases, you run this earlier on where you’re in testing or dev, you might run some of these experiments, but then you're gonna end up most likely, if you really want to test the resiliency of your system, you're gonna run this in production somewhere. And that means if something breaks, you know, your tech support team or customer service or whatever, they might start getting flooded with calls. So it's probably good to kind of notify everybody, “Hey, we're running an experiment”, right?

Gunnar: Exactly. It might involve people from from all teams within the organization. I guess if a major customer is affected, someone at sales might get a call as well.

Jeremy: Yeah, right, you don't want to be that support rep or you don’t want to be the account manager, probably, unless you figure that out. All right, so let's get into actually talking about the experiments themselves. So let's maybe take a step back and ask the question, why would we run experiments in the first place?

Gunnar: Yeah, well, since the purpose is to find out if the system is resilient to failure, we look at if our customers or if our system has a problem, are our customers getting the experience they should? Is the system behavior, behaving good enough for our customers to get the experience? Or another thing might be that we have downtime, we have issues that are costing us money. So, and as you mentioned, is our monitoring working as it should? So we have quite a lot of things that might intends us to actually do the chaos experiments. And to do it well, it builds confidence. When we do the experiments, we build confidence, and we know how everything within the system and the organization behaves in the face of failure.

Jeremy: Right and we probably already kind of talked about this a little bit, but this idea of an organization being able to handle an outage, right? Because if all of a sudden something goes down and you're like, well we just expected everything to be running. And then all of a sudden something goes down. And either there's cascading failures or all kinds of things like that. If you run these experiments and say, okay, it's like a fire drill, what do we do if X fails? You know, how do we recover from that or what kind of resiliency do we need to build in. So that's a big part of it as well, right?

Gunnar: Yeah, that's correct. A lot of organizations today run what's called Game Days, where you do exactly these types of fire drills, where you inject some sort of failure into either the system or the organization, and like a game, you actually do it and see how would you behave. How do you perform within the organization. And this continues not only until you have resolved the actual issue, you need to make sure that you know how do you report everything. How do you follow up on the failure and how do you solve the issues that you found within your system or your organization?

Jeremy: Yeah. And you mentioned this too, I think in your talk where you know, it's not just about what happens when a system goes down, like what happens when the system slows down? Like does that affect how many customers, like how many orders you get per hour, for example, I think you mentioned Amazon that when their latency goes up by 100 milliseconds. They lose X amount of dollars per hour or something like that.

Gunnar: Yeah, that's correct. I know Adrian Hornsby, evangelists at AWS has a slide in one of his presentations where he shows some numbers on this. And I believe that the example of Amazon is that they have 1% drop in sales for 100 milliseconds, extra load time. And for a company Amazon’s size, that's quite a lot.

Jeremy: That's that's a lot of money.

Gunnar: Yeah, another example I know was that Google has a number of that says that 500 milliseconds of extra load time cost 20% fewer searches on google.com.

Jeremy: Which means 20% fewer ads load, which means 20% less money from the ads. Yes, so I think that's just to me is one of the more fascinating things of this as well is just this idea of injecting latency, and we can talk about that. But you know, all these different things you can do to affect the customer experience and see what effect it literally has on your bottom line. That's just a really cool, it's a really cool thing that I think maybe most companies aren't thinking about right now.

Gunnar: No, exactly. And that's why just business metrics are so important as well. We tend to focus on technical metrics in our field. But when it comes to chaos engineering, the business metrics is probably more important. And what CPU load there is, or how much memory we're using isn't all that important. Our system should be able to handle that, but how does the business metrics get affected when we have issues?

Jeremy: Right? And that's why you want to run these experiments sort of his close to production or in production, because that's where you're going to see actual effects like that, where you actually see when customers are impacted what happens.

Gunnar: Exactly.

Jeremy: All right, so if we're gonna run some of these experiments, you had a bunch of steps laid out, right? So why don’t we start at step one, what's the first thing we do?

Gunnar: Well, the first step is that we need to define what we call the steady state and steady state is the normal behavior of our system over time. So that we, the metrics that we have, we need to know what is the steady state of those metrics. And that might be if it's a business metrics. It might be, how many searches do we have per hour per day or per year? And how many purchases are there on the Amazon.com? But of course, we have the technical metrics, or system metrics as well. And how many, clicks, or how many function invocations are there per hour, for example. Or what is the CPU load? So we find these metrics and we define those so that we can use them when we run our experiment, to be able to benchmark what happens when we do the actual experiment.

Jeremy: All right, so then the next step, so once we know our steady state. And essentially, as you said, and this again, this is based off of a number of different factors, too. So it's also like, how much load, like, you know, what's our steady state at 8 am on a Monday versus what's our steady state at 2 pm on a Wednesday, things like that. So you should have those metrics kind of laid out. And as you said, those business metrics are extremely important as well. So once we have the steady state, then we move on to step two.

Gunnar: Yeah, and the second step is then that we form what we call a hypothesis. So we decide upon how will the system or a hypothesis around, how will the system handle a specific event. By using what we call what ifs, we try to find a way to form this hypothesis. And what ifs might be, what if dynamodb isn't responding? What if the load balancer breaks? What if latency increases by X amount? And by using these what ifs were then able to form our hypothesis. And the hypothesis might be that “if latency increases by 100 milliseconds, our front end will still behave correctly for the customers or end users.” Or if Dynamodb isn't responding, our system will have a graceful degradation, and we will still be able to have service to the end users. So that's the hypothesis that we're then going to try to prove or disprove.

Jeremy: And so is there, because again, I know it's a relatively new field, and maybe there's more about this in the book that Netflix did. But are there like a common list of hypotheses, or is it something that, is just sort of going to be unique to eat to each system.

Gunnar: They are often quite unique, but many of the chaos experiments that we use, they originate in the eight fallacies of distributed computing. You know, that the network is reliable, that latency is zero, that bandwidth is infinite, and so on, so on. So if we base experiments on these, well, we will get a baseline of experiments that we can inject into most organizations or most systems. But then outside of that, well, it's probably based upon what type of system, what services you're running.

Jeremy: And so should we come up with a list of experiments, when we’re coming up with hypotheses, should we come up with something that's going to test everything, or should we try to be, are there more specific things we should focus on?

Gunnar: Well, we usually start by the critical services, so we start by Tier 0 services perhaps at the most critical functionality. And so we find those systems and perform the experiments there, and then we can move, expand based on that. So the systems that will probably affect the most are the most important ones to get started with.

Jeremy: So you would want to start with what happens if the database goes down? What happens if this service goes down or something like that as opposed to let's start injecting latency, to see how that affects customer behavior.

Gunnar: Yeah, most likely, if you run an e commerce site that I suppose the final steps of the purchase process is quite important. You might want to start there.

Jeremy: And I would think too, if you came up with a hypothesis and you said, “Alright, if the DynamoDB table goes down, everything's gonna blow up and the site’s not gonna work anymore”, that before you would run that experiment, you'd probably, if you identify that there's a problem, you probably would fix that before you run the experiment, right?

Gunnar: Yeah, that's correct. And it's quite common that just while you're trying to form your hypothesis or create your experiment, you think of things that you haven't thought about before, and depending on the people in the room, you might start talking with people that you normally don't do and you find things that probably won't work. So then it's better to fix or better, you should fix those issues first and then start over, form a new hypothesis.

Jeremy: Right, and then once you kind of work your way through that, then I mean, there's gonna certainly be things in your system that you're not gonna know that they're gonna have an effect until you actually run the experiment. But like you said, I just think it's one of the things you don't want to say, “Okay, I have a feeling that if we if we take DynamoDB down that our entire system will crash, let's try it,'' because you're pretty sure it's going to happen. So, alright, so perfect. So now we have our hypothesis and now we need to plan our experiments, right?

Gunnar: That's correct. So we have the hypothesis, and based on that, we then create the experiment and we do that just like, we create a plan for it in detail, exactly how to run this experiment. Who does what? And when we do it? But key here is that you should always start with the smallest possible experiment so that you contain the blast radius. And blast radius containment is really important to be able to, first, of course, test the experiment, but to not create a bigger impact then what's needed to be able to see if that part works or not.

Jeremy: So when you say when you say contain the blast radius, so what would be an example of that, like just only kind of adjusting or breaking a single function? Or how would you define that?

Gunnar: The blast radius might be different sorts, but just doing it on the single function is one way of doing it. And you're trying to make sure that you don't affect the entire system that way. But it might also be that you're doing it on the small set of users. So instead of injecting the failure or exposing all users to the failure, you can do it on a small test group, for example.

Jeremy: Yeah, and that actually makes a ton of sense, too, because if you're like, rather than saying what happens when dynamodb goes down, you could say what happens when this function can’t access DynamoDB? And then that's a, like you said, much smaller blast radius. All right, so you got the experiment to find. You planned it to contain this blast radius. What is the, how do you test though, that the experiment worked or didn't work?

Gunnar: Well, yeah, it's important that you have metrics that can measure the effect of the actual experiment. So if you're testing a specific part of the system, you need to know how did the system behave when you injected that type of failure? So once again, it's part of monitoring and having observability into what you're actually doing in the system.

Jeremy: And when you run these experiments, I can imagine, especially with serverless, it gets a little more difficult, right? Because you can't actually shut down DynamoDB, right? So you have to find different ways to kind of get around that.

Gunnar: Yeah, that's true. When we're creating experiments for serverless we, well, we have a higher level of abstraction, so the failure modes aren't the one that we're used to in chaos engineering. As I said, it started out with Chaos Monkey that shut down EC2 instances. Well, we don't have any EC2 instances anymore. There isn't an off button for DynamoDB. So we need to find new ways of injecting this failure.

Jeremy: So now we've got, you know this is a sort of a complex step. But now you've got all these things in place. This is probably where you want to notify the organization. Like, ‘Hey, we're gonna break something now’ or ‘We might break something now.’

Gunnar: That’s really important because, just performing experiments without notifying anyone might lead to consequences that you don't really want. And having success stories around chaos engineering is really important. And success might be that something breaks. It doesn't have to be that the system is resilient. It's a success story, even if you find something that breaks. But the organization needs to be ready to handle whatever happens when you inject a failure. And an important part as well is, when you start your experiment or when you're designing your experiment, make sure to have a way of stopping it. You need to have that big red button to stop it.

Jeremy: Yeah, that makes a lot of sense, because I can imagine that would be ‘Hey, we might break some stuff. Oh, by the way, we broke it. But it's gonna take an hour before we can fix it.’ That's probably not a good thing.

Gunnar: Your only chaos experiment.

Jeremy: It might be the only one. Well, but you're right, because you mentioned this idea of building confidence with an organization. And I can certainly see that if I'm, you know, I'm the shipping department or something like that, and you know, we rely on the system being up all the time, and you're gonna tell me ‘Oh, by the way, we're going to inject some failure into the system or we're gonna try some experiments. It might take down the shipping system.’ And then all of your workers that are in the factory or in the warehouse might stand around for an hour because no new orders are coming in because we broke something. You know, that's the kind of thing where, you know, they might be not super excited about something like that happening. But if you could do it on a smaller scale and you can say, ‘Well, you know, for the billing department, it's okay if we didn't bill people for an hour because we were testing some things. But what we found was, if this happens, we can fix this.’ And that basically means, or we built in the resiliency for it or some sort of backup plan for it, and now we can reduce that type of outage. If that happens, we reduce the downtime or we reduce the failures or whatever. You know, that reduced down to 15 minutes, as opposed to maybe a three hour fix or something like that. And so if you're the shipping department, you say ‘Okay, see, this is a benefit. I take that hour, let them test it. And then I know that if there is some critical failure in the future, all of my people aren't standing around in the warehouse for a day because we were able to figure out what the problem was before it happened.

Gunnar: Yeah, that's absolutely correct.

Jeremy: Alight, cool. So now, speaking of, sort of the measurement side of things. So, step four is we measure it and we learned something, right?

Gunnar: So now we've run our experiment. We've performed the injection of failure that we wanted to do. So now we need to use the metrics that we have in place to to prove or disprove the hypothesis, and depending on what type of hypothesis we have, this might be a technical or system metric. Or it might be some other business metrics. So that we see, did the frontend break when we shut down DynamoDB? For example. So we need to see that is the system resilient to the injected failure? Remember what I said early on? That resiliency isn't that nothing ever breaks. It's about assisting, being able to give good enough service to to the end.

Jeremy: Right, and that's actually a really good point too, because that's that's another thing, too, is with distributed systems. Things break all the time. Messages don't get delivered or you can't access something. And Netflix does this very cool, graceful degradation, right? Where they just don't show recommended movies, or whatever their thing is, if that section isn't working right. So this is that, “don't show an error message if you don't have to.” If you can get away with kind of cutting out a piece of it, you know, then have a graceful way off of letting your system fail.

Gunnar: Yeah, exactly. And just having that graceful degradation, I think the Netflix example is really good there. Having UIs that don't block users. You don't get an error message. It just keeps going. But you don't see the part that's not working.

Jeremy: Yeah, and then with serverless, too. It's one of those things where if for some reason your database was down and you couldn't process payments or something right away, you can certainly buffer them in an SQS queue or in Kinesis or something like that. So you could still be taking orders even if you can't necessarily bill, or you can't charge somebody's credit card. But what's better? Being able to keep taking orders with maybe a few people who enter incorrect credit card information that you have to notify them later and say, “Hey, this credit card didn't work” and maybe get them to give you a new one. Or to basically say to everybody, “Sorry, you can't order anything right now."

Gunnar: Yeah, that's that's part of what you need to do when you run your experiment, you need to learn from the outcome and find ways off doing it in a better way, usually. And that might be a question you have to ask. Should be shut down the entire part? Or should we leave good enough service for some users?

Jeremy: Right. And that's probably where communicating with the rest of the company and sharing all of that success, or I guess, non-success, depending on what it is, with the rest of the company and saying, “Okay, if this broken your system, what would you want to do?” Like, “how would you want to handle a failure in your system?” I think that's a good discussion to have with your entire organization.

Gunnar: Yeah, you involve the product owners or how your organization looks and let it be a business decision in the end, perhaps.

Jeremy: So now we've got these small experiments running, right, cause we're still containing this blast radius and trying to be smart about not killing our entire system. The next step, though, is to turn it up to 11 and see what happens, right?

Gunnar: That's correct. So, if we've done the experiment on a small set of users and the next step if it was a success, we didn't have any failures were then able to expand on the system and scale it up and perhaps have bigger set of users or more countries or more functions. However you contain the experiment early on, now you're able to scale it up, and with that increased scope, you'd usually see you some new effects. Injecting latency into one single function might not reveal any issues, but if you inject into multiple functions at once, you might get an effect that you didn't see you with a small scope test.

Jeremy: Right. Yeah, and then because again you have those cascading failures and things like that as well. So one function goes down or one function can't connect. That's one thing. But if a whole series of them can't or multiple services that maybe share some common infrastructure, something if that goes down, so that's really interesting. And this is also to where I think the sort of the economic tests that injecting latency what happens if we slow down, like on 5% of your users, you may see something there, but if you want to get a larger sample size, that would certainly be where you would want to run the larger test around that.

Gunnar: Yeah, exactly. The business metrics might not, you might not see the effect in the business metrics until you scale the experiment up a bit.

Jeremy: Alright. Awesome. So why don't we now kind of shift this a little bit more towards serverless and just kind of talk about what the challenges are? I mean, we mentioned that you can't turn off DynamoDB, but maybe we can just talk about what are the common things in serverless we might test for like, what are those common weaknesses that you typically see in a serverless application.

Gunnar: Yes, definitely. Well, things we usually see, and this isn’t something that we see only when doing chaos engineering. It's something that we might see in architectures every now and then. It might be that we're missing error handling. So when we have functions receiving errors from downstream services, for example, we aren't handling them in a correct fashion. These are things that we might find when doing chaos engineering experiments in serverless. Timeout values, that's quite common that we don't have the proper timeouts on our services so that perhaps it's not always an issue. But when we get latency, when we inject latency, we might see that intermediate services might have failures when we have the incorrect timeout values. So services at the edge might be affected because there are timeout issues further down. And fallback, missing fallback, rather. That's something that we see every now and then, that a downstream service of some sort in DynamoDB, for example, isn't available. So then we don't have proper fallback for that because we're relying on DynamoDB to be there all the time because that's how cloud works, right?

Jeremy: Well, but that's because every engineer pretty much designs for the happy path, and you just assume everything will work. Like you said, those eight fallacies of distributed computing. So that makes that makes a ton of sense.

Gunnar: And then we have the bigger one with the missing regional failover. Well, quite often, with design systems that aren't distributed over more than one region. So we have them in one single region, and regions they rarely fail, but it might happen. And it doesn't have to be that an entire region goes out as well. It might be that network issues for closer to the user might prevent them from reaching a specific region as well. So you don't always have enough regional spread if you only use CloudFront, for example. You might need to spread out your Lambda functions as well over multiple regions.

Jeremy: You know, that's something to where this is, I see this all the time. It's just people build, serverless application or parts of their application serverless and they just assume us-east-1 will always be there. You know, I mean, there are hurricanes, there are floods. There are earthquakes. There are all kinds of things that could affect, and not just one of the availability zones, cause, obviously, availability zones within each region are actually spread out quite a bit. But you could have something that took out an entire region, and then kind of what happens, what happens there? But also, the other piece of it, too, is that if that goes down, how are users affected, even if you do have another region, like what if we have to now route all US traffic to the EU? Right? Or how does the EU behave, or what's the latency like when they're accessing services in us-east-1. Because, obviously you have that speed of light problem, right? It can only get back and forth so fast. So, yeah, that's interesting, too. And also, just if one region goes down, how do you replicate the data? I mean, there's just a lot in that if you're building a really large distributed application. So anything else, any other sort of common things that we see.

Gunnar: One common thing that I see when we're performing these types of experiments is that, like we said before, we don't have graceful degradation, so that the systems they show ever messages to the end users or, parts just don't work but are still there. So we don't have UIs that are non-blocking. And that's a perfect use case for a chaos engineering to be able to find those on and, well, then fix them.

Jeremy: Actually, I think that's a good point. One of the timeout things on the UI side. If we just set our API gateways to timeout after 30 seconds, then the problem is that, which is the maximum that they in timeout at, 29 and a half or whatever it is, that when that times out, if our front end is waiting for 30 seconds for that API to respond, that's not gonna be a great experience for the users on the UI side. So shortening those timeouts so that it fails faster would be helpful, because then you can respond. If you expect that to respond in two seconds or less, which you should. I mean, you should be a couple hundred milliseconds, but if you expect it to respond that quickly and it doesn't, then you should probably degrade at that point, right?

Gunnar: Yes. Correct. I know serverless hero, Yan Cui, he wrote an article about exactly this and how to use chaos engineering to be able to find these or fine tune the timeout values. So I think that's an excellent use case.

Jeremy: Awesome. All right, so let's talk about serverless chaos engineering. So how do we actually do these experiments? Or what are these experiments, is probably a better question?

Gunnar: Yeah, well, one thing that we have talked about already, it's latency injection, and I think that's probably the easiest one to get started with. Yan Cui, once again, he used latency injection to be able to properly configure timeout values for functions and downstream services. So, he created an article around this and with examples of how to do it. And Adrian Hornsby then from AWS, has created a Lambda Layer for this as well. So you're able to add a Lambda layer to your functions and then inject latency and see how the function and the system overall behaves then. So latency injection is a quite easy one to get started with. And then we have status code injection or error code injection. So instead of having the proper 200 status code as a response from your API gateway, you might get a 404 or 500 errors. Some like that. So you can inject those and see how does my system behave when there is an error message and because that's not something that you might normally get. But by injecting them you're able to test the system behavior that way.

Jeremy: And you could do like concurrency, manipulate the concurrency. Probably?

Gunnar: Yes, definitely. And to be able to test that properly, then you need to be on a quite large scale, usually, to see exactly how it works. But you're also able to configure DynamoDB for the read and writes. Since we don't have the control of the underlying infrastructure, we need to make stuff up to be able to do these experiments in another fashion than perhaps traditionally with chaos engineering. And that's the fun part as well. Since most of the services are API driven, it's quite easy to do. Change us back, back and forth. The stop button is a configuration change in many cases. So you can just use the CLI to do a change back and forth when you're doing your experiment.

Jeremy: Right, and then what about configuration errors? Like if we changed IAM policies, for example, there are other things you could run around that?

Gunnar: Yeah, by doing the configuration changes or permission changes, you might replicate that some service isn't available or isn't responding in the correct fashion. But you might also create a configuration error in a way that might happen every now and then. Because with the way that serverless works, we have so many ways to configure everything. You have configuration on each and every function, so that is something that happens every now and then. It might be on deployment that some configuration change happens or something is missing. And so we're able to do that in a controlled fashion by using chaos experiments.

Jeremy: So could we run a region failure test with serverless? Like how would we do that?

Gunnar: Yeah, that's a bit harder to do it in a controlled fashion. But, we can do it by configuration changes, by doing permission changes. So were able to lock, more or less, lock ourselves out of specific region. But it's hard to get the exact same effect as if the actual region is down. Because you're probably getting errors, that are different from the way it would behave if the entire region is down.

Jeremy: I wonder if you could do that with Route 53? You know, because you can do like some of the health checks and things like that. I wonder if you could inject a failure there and then have it route traffic that way?

Gunnar: Yup. Well, if you have those health checks or you have that type of routing in place in Route 53, you might do it that way. That's right, to more or less inject failure into the health checks.

Jeremy: Alright, so maybe then let's talk about how we actually then run the test. We talked about some of them, but there’s tooling, or there’s some tooling you mentioned. Adrian Hornsby has his lambda layer, but are there other tools that we could use if we wanted to start doing this?

Gunnar: Yes, there are. Looking at chaos engineering as a whole, there is one of the bigger options today is a tool called Gremlin, and they have a SaaS offering for chaos experiments. And one part of that system is an application fault injection. And that way you can inject failure into Lambda functions using Gremlin. And so that's one option. Another option is the open source Chaos Toolkit, which has drivers for the main cloud providers. And you're then able to inject failure into serverless services as well. And then one which I think is a fun one, is Thundra, the observability tool. They have added on option to inject failure into the serverless applications that you're observing through there too, so then you're getting both the observability off the chaos, and you're doing the chaos through the same tool.

Jeremy: And so, with the open source one. Is that something that you can build your own experiments? Or are there plugins or something?

Gunnar: Exactly. It's built around, and I think they called extension, or plugins and that you're able to build yourself. There are a bunch of them out there, but you can easily build your own as well and extend all that. So I would say that the most common way of getting started with it in the serverless space is to more or less build your own. Seems you're able to do many of the things through CLI or the API. You can easily do simple scripts that can perform the chaos and have an easy way of stopping them as well.

Jeremy: But it sounds like if you're building your own stuff and even if you're using any of these other ones, maybe not so much with like Thundra, for example. But you do have to kind of build these tests into your code, right? Like you have to kind of write your code to say “alright, I can fail this by flipping a switch somewhere,” but you actually would have to modify your underlying code. It's not just something that maybe you could just wrap, or is it something you could just wrap?

Gunnar: If we look at the Lambda Layer that Adrian has built that is actually a wrapper that you're wrapping around your functions. So you still need to have it in your code. But if you have a disabled, it should be fine to be there all the time and then you just enabling it when you're doing your experiments.

Jeremy: Because you don't want to change your code If you're running an experiment versus not running experiments. That kind of defeats the purpose.

Gunnar: Exactly, yeah, because you want, when you're running your experiments, you wanted to be asked when in production. And if you have to do deployments every time, as you said, it defeats the purpose.

Jeremy: Alright, so let's say you're a company and you're super interested in doing this and obviously cause it sounds ridiculously fun to go in and maybe break things, but not break things, but so what would you, how would you flip these on and off? Would you do that with environment variables, or how would you start and stop these tests?

Gunnar: I would say that depends on what tooling you're using. If you're using, like gremlin or Chaos Toolkit, they have their on and off switch in the systems, and Thundra as well. But if you're using, for example, the Lambda Layer from Adrian, then you're doing it through parameter store. So you're having a variable there or a parameter where you're able to set it enabled or disabled in a quite easy way than just by updating using in the AWS CLI. And so that way, or if you're building your own, you might do it through to environment variables on the actual function as well.

Jeremy: Awesome. Okay. All right. So I guess, last question here is chaos engineering, I mean, obviously, to me, it sounds like a no brainer. Do you have any advice for people thinking about doing chaos engineering?

Gunnar: Well, what I would say it's that, of course, they should do it. I think it is beneficial for organizations to do it. Even if you don't have a large global scale distributed system, you can still do chaos engineering. But an easy way to get started is by looking at the references that are out there. Chaos engineering is a hot topic today, and there are tons of talks around chaos engineering all the time at different conferences. So look at them on YouTube and read up on it. And just make sure to know that you start small and you do your small experiments first, and then you're able to scale up.

Jeremy: Awesome. All right, well, thank you, Gunner, so much for being here. If the listeners want to find out more about you, how did they do that?

Gunnar: Well, I guess the easiest way is to look me up on Twitter @GunnarGrosch, which is hard to spell, but I think you'll find me.

Jeremy: All right. I will put that in the show notes and also those other tools that you mentioned and stuff, I'll get those in there as well. All right. Thanks again, man.

Gunnar: Thanks for having me.

View Details

About Ran Ribenzaft:

Ran is a passionate developer, with vast experience in network, infrastructure, and cyber-security. He's constantly chasing new technologies, with a current focus on Serverless. He is an open source contributor and currently the co-founder and CTO at Epsagon, a tool for monitoring serverless applications.

  • Twitter: @ranrib
  • Epsagon website: Epsagon.com
  • Epsagon blog: Epsagon.com/blog

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Ran Ribenzaft. Hi, Ran. Thanks for joining me.

Ran: Thank you very much for having me, Jeremy.

Jeremy: So you are the CTO at Epsagon. So why don't you tell the listeners a little bit about yourself and also what Epsagon is doing?

Ran: So I'm Ran, the co-founder and the CTO over at Epsagon, based in Israel, Tel Aviv, a fun place to be, very warm. In my previous roles, I've been doing mostly cybersecurity stuff. So mostly getting into kernels and things that I can't tell you, you probably heard some of them in the news. But let's keep it discreet. And in my recent role, I'm the CTO here at Epsagon, a start up where I'm one of the co-founders, and mainly focuses on monitoring and troubleshooting modern applications. But in general, the idea is to have a single platform where you get both the monitoring capabilities, which is "is my application working properly," "is it meeting the SLAs performance," and so on. And the troubleshooting part, where something bad happens, you need to scroll through the logs and correlate between them and do all this distributed tracing. That's Epsagon in a nutshell.

Jeremy: Alright, great. So I wanted to have you on because I want to talk about observability in modern applications, and by modern applications, we mean cloud native, or serverless, or distributed, or whatever sort of buzzword we might want to use to describe it. But as our applications grow beyond the traditional monoliths, being able to observe what is happening in your applications is a huge part of what we need to do when we are building these modern, interconnected systems. So maybe you could just give me a minute or so on what's the difference between your traditional monitoring and observability systems, and where we are now and what's changed ?

Ran: Definitely. So it starts with the change in our infrastructure in our way we code. So if we used to have this monolithic application running on our on-prem servers, so the things that you wanted to monitor are like: what is the network throughput, and what is the CPU usage, and hard disks and so on. And, you know, just making sure the application, the process itself is alive there. But shifting to more modern application, which I think in my mind the modern application is something that you don't mess with the infrastructure around it. You get most of the services out of the box working for you in a matter of configurations that you just can, you know, tick some boxes that I want this feature and so on. And you just built your own business logic through that where it can run — I don't care whether it's in your server, something like a function of the service or other thing. So this is modern application, and in this kind of modern application, there's a big difference in what you want to monitor. Honestly, we're doing monitoring to make sure our business works. Our application is our business. I want to make sure it works, so things like how much CPU is being consumed or network throughput, or all these kinds of metrics that just show me charts about infrastructure are getting irrelevant over time. Like, for example, if I used to have a chart of how much CPU usage my database is consuming, so now we don't really care if I'm going to a managed database - a fully managed one, not like a semi-managed - I don't really care about the CPU or anything else. I just want to make sure it works, and my application can speak to it at the right timing, at the right performance, and it gets the right results. So that's the first thing. The second thing is mostly about the nature of these kind of applications, which we broke them from being a big monolith, a big single monolith, to multiples of microservices, you can call it microservices, service, nanoservices, but the fact that there was one giant thing that broke into 10 or hundreds of resources, suddenly presents a different problem. A problem where you need to understand what is the interconnectivity between these resources, that you need to keep track of messages that [are] going from one service to another, and once something bad happened, you want to see the root cause analysis. This is like a repetitive thing that you can hear over and over. This root cause analysis, so the ability to jump from the error - the error can be like a performance issue or like exception in the code - all the way to the beginning. The beginning can be the user that clicks on a button on your business website that caused this chain of events. So these are the kinds of things that you want to see where, in traditional APMs, in traditional monitoring solutions, you don't have it. And in the future, once you'll find it more and more like that.

Jeremy: Yes, so you mentioned, you said logs. You said metrics. You said interconnectivity. You talked about a couple of different things, and I think it's probably important for listeners who maybe aren't 100% familiar with what observability is. There's this thing called the three pillars of observability, which are logs, metrics and traces. So maybe we can talk about that and you can sort of tell us why each one of those things is important.

Ran: Yeah, I'll start first with the metrics. So metrics are like the key component that you can ask questions about. Let's say how [many] events that I got per day, how [many] purchases, how [many] events of [this] kind that promote my business [are there] per day or per timeframe that you want to see. These kind of metrics are the base unit that you want to monitor. Now, when something bad happens in this metric, sometimes you need to see something bigger than just a number that will tell you, "Hey, we found out that the amount of transactions that you're seeing per day is lower than 100." So you want to see a trace or, in my opinion, a trace is more like a story. What happened — like tell me the exact event where this metric was below that 100 or the KPI that I've measured. I want to see what you talked with, which resources were being involved, how long each kind of these operations took, and why or how is it different from being a good trace. Now, when you wanna dive even deeper, so you need to get to your logs. Logs are like the ultimate developer utility to troubleshoot problems. I mean, regardless, what you'll see in metrics or in traces, logs are the core thing that developers put in their code in order to troubleshoot and debug their applications and often, you want to see correlated to one another. So, for example, I want to ask these questions: "How many purchases did I have on my website in the specific day?" Now we'll see there is a spike, or the opposite of a spike, some down in their registration. It's probably going to be because of a problem. So I want to see all the traces that correlate to these specific events, and I want to see all the logs that correlates to the traces that have found to this metric. So all three are connected to each other and all this observability, which is a nice buzzword, it's a just a translation of being able to monitor our business in production. That's for me, the things logs, metrics, and traces are just different way to look on my observability.

Jeremy: So maybe we can talk about why these things are a little bit different in monitoring a distributed system versus monitoring a traditional application. You had mentioned breaking things into microservices or nanoservices, which I'm not a huge fan of that word but it's okay, um, but breaking things down into smaller parts and they're disconnected or they're you know they're distributed, right? So what sort of the, maybe what are the options that you have with these modern applications to track that kind of stuff?

Ran: So let's offer some AWS alternative to each one of them. Probably the first thing that each serverless developer or every serverless developer thinks of logs is CloudWatch logs, which is great because it comes out of the box and you're getting the logs shipped. Every print that you'll do, every stdout and stderr will come out to the logs, which is perfect. Honestly, it's great up to a certain scale, but once you're hitting millions of requests, it's really hard to navigate through and try to find what you're looking for. The log that you're looking for. So searching might be a bit of a pain in the logs, but honestly, it works out of the box, so there's no reason not to start with it. The second thing we talked about is metrics. So metrics, there's obviously CloudWatch metrics that can build on top of logs, but also can build on top of custom metrics, which also is great, you get it out of the box. For example, for Lambda or for any resource that you'll use in AWS, you'll get CloudWatch metrics already defined. So, for example, for Lambda, you'll be able to see the amount of invocations, the amount of errors, duration statistics and so on. But honestly, they are distributed across hundreds of metrics. And sometimes you want a single dashboard that will just show you all of these metrics and will tell you when something bad happens. You don't want to configure, if I cross this threshold, that if I'm getting an error in here and if I'm getting something there, it's an endless amount of alerts that you'll need to configure and you want something out of the box that will work for you. And the last one regarding traces, or distributed traces in modern application, we've got X-Ray, which is great for tracing AWS resources. It tracks down almost any request that you'll make using the AWS SDK. However, it doesn't track anything external to AWS, and it doesn't do distributed tracing yet. Hopefully they'll get there soon because that's X-Ray distributed tracing, but at the moment it's still limited. So these are the options that AWS provides. There are tons of other resources that you might, that you can find outside of AWS.

Jeremy: So with X-Ray, though, in terms of being able to get alerted on slow-running processes or resources that are taking a while, is that is that possible with X-Ray?

Ran: So it doesn't come out of the box. Like AWS does build the best infrastructure that you can build applications on top of it. So X-Ray will collect these traces for you, and then you can build your own application that says "Scan all my traces. Scan all the operations against my database." And when I find an event in my trace that crossed my threshold, send [me] a Slack alert, or on whichever platform I'm feeling most comfortable. So it doesn't come out of the box, but it gives you the infrastructure to build great things on top of it.

Jeremy: And you also mentioned that it's great for tracing AWS resources, but what about third-party calls? Or if you're trying to interact with a third-party resource?

Ran: Exactly. So such as every application, it's almost completely hybrid. I, for my years in doing cloud, I haven't found any application that is only serverless or only containers or only on-prem. You find mixtures of all kinds of applications and you can find yourself using, like Redis, which is not part of the AWS. And instead of DynamoDB, MongoDB. And instead of Kinesis, Kafka, that you own because you needed to configure something specifically for you. X-Ray wouldn't be able to trace these kind of things and definitely not distributed kinds of these things. For example, a message that goes through Kafka, it wouldn't be able to trace you from one service to another.

Jeremy: All right, so you mentioned third party apps, and obviously, or third-party products, and obviously Epsagon is one of those, but there are others besides Epsagon. But in this context, how do these third-party monitoring or observability tools, how do they extend what the cloud provider does?

Ran: The main difference between SaaS services and infrastructure solutions is that SaaS service comes out of the box prepared for you with all the integrations and all the needed configurations already plugged and played just for you, so you can run quietly and make it work. Just like you're using a managed Kinesis, instead of building your own Kafka. So you want to a managed monitoring solution, just so you want me to build your own Elasticsearch and build things on top of X-Ray and build another dashboard outside of CloudWatch metrics, it comes much better outside of the box for you and sometimes it brings some more value, more application-wise value like, for example, Epsagon provides cost monitoring or monitoring things that are not necessarily provided by the cloud provider. For example, if you own other services, which are not running on AWS, Epsagon will monitor them as well.

Jeremy: So let me ask you this question, because again you think about CloudWatch logs and you think about, you know, metrics and X-Ray and you're right. There's a little bit of setup involved there. The searching on CloudWatch logs can get kind of slow, but obviously, you know, some people just transport their logs to like Logz.io, or whatever it is, or put it in an ElasticSearch cluster or something like that. But if you're just shipping the logs or trying to aggregate some metrics, you're not getting the whole picture, right, because you're not seeing, like you said, where all these things sort of tie together. So maybe you can tell me or help me answer this question. Why is this such a hard problem with distributed systems?

Ran: That's actually great point, Jeremy. I think that the main issue here is that just shipping out logs, unstructured logs, wouldn't tell you a lot of information about what you're looking for. Because logs are things that developers wrote for themselves in their codes, so once there will be a problem, they'll be able to investigate it. But it's not a thing that you can ask questions on top of it. I can't ask how many transactions, how many purchases did I have on my website today, because it's not a metric. It's a log line. So building things that will do instrumentation and distributed tracing, and will give you out of the box this ability to do custom alerts on frames that you would like and this cost monitoring that I've mentioned doesn't come out of the box and building it, it's really hard. It takes a lot of time, especially when you're doing things in scale, so you need to manage that as well. So you want a service that will be managed for you to do all of these things.

Jeremy: All right, so let's move on to sort of the next, I think that's a good segue. So let's move on to this idea of actually enabling your application to do some of this tracing and logging and things like that. So you mentioned unstructured logs. So obviously, if I'm in Node and I just use console.log I can write texts to the log, or I can create my own structured JSON object and send that in. So I could do that. But that sort of requires me to manage that myself. Again, you mentioned the CloudWatch metric sort of captures the data after the fact and you can kind of parse through the logs and X-Ray, there is some instrumentation that needs to be done there in order to make sure that it's tracing calls to MySQL or other services. So maybe we could just kind of go into this whole idea of instrumenting services and code in general and maybe we can start by, you know, sort of explaining what exactly we mean by instrumentation.

Ran: Perfect. That's a technical question that I like to take. Instrumentation is the way or a technique which allows a developer to, let's call it hijack, or add something to every request that he wants to instrument. For example, if I'm making a calls using Axios to a REST API for my own code to an external or third-party API. I want to be able to capture each and every request and response that is coming in and out from that resource, from that Axios request. Why would I like to do that? Because I want to capture vital information that I'll be able to ask questions about later on. For example, if my Axios is calling Stripe to make a purchase or to send an invoice to my customer, I wouldn't know how long it takes, because I don't want my customer to wait on this purchase page or wait for his invoice to get into his email. I want to make sure of how long it takes so I can measure that, put that as a metric in CloudWatch metrics or in any other service. And then I'll be able to ask, "Well, was there any operation against Stripe that took more than 100 milliseconds?" If so, it's bad, and this is only accomplished using instrumentation. I mean, the other way around is just to wrap my own codes every time that I'm calling Stripe or every time that I'm calling any other service. But with the amount of annotations that you'll have to add to your code, it's almost unlimited, so you won't get out with it without a proper instrumentation in your code.

Jeremy: Yeah, and I can imagine too if you're wrapping every request or you have to do something custom for every request, that's obviously an easy thing for for developers to forget or potentially get wrong. So how do you do it then? So if you're using Python or you're using Node or, you know, one of the other languages that maybe Lambda supports, what exactly do you do? How do you instrument these HTTP requests and SNS requests and things like that?

Ran: Yeah. So you mentioned console.log so I'll give the example in Node. In Node, there's a fantastic library called Shimmer. Shimmer allows you to, since it's a dynamic language as it's not compiled to anything, it can just alter a function in the memory. So, for example, I'm altering Axios.get() to my own function. I'll make sure to get all the details that you've sent to Axios.get(). I'll extract information from it, tell which kind of information do I want it to get for me. I'll send this request to the real Axios.get, and then I'll capture the response, get everything needed from the response and I'll put the response back to you. So it's almost as transparent to the actual operation, but in the meantime, I have collected information both from the request and from the response. This can be done, for example, for AWS SDK library, for any HTTP request library like Axios, Got, Fetch, HTTP and so on, or any thing. Even I can instrument myself into Console.log. So every time that you run Console.log, I'll capture the log and I'll stream it to where it was originally originated.

Jeremy: But how do you do that? Is that something you have to do manually for every call to Axios?

Ran: Yes. So I'm doing it once. I'm doing generic for every call that there will be to Axios and then I'm collecting from Axios, for example, if it's an HTTP request, the URL, the params, the headers, the status code from the response, the headers of the response and every metadata or fingerprint that I would like to collect as part of it that I will be able to ask questions or filter or seeing the trace afterwards.

Jeremy: And so the the information that you collect, what do we do with that?

Ran: Many kinds of things. One of them is to put as a metric. As I mentioned, a metric can be how long it took or how many error codes, a type 501 I get or larger than 500 in the HTTP response code. It could be something from for traceability. So, for example, I want to capture the headers of the request because I know that there is a user ID there, and sometimes I would like to ask questions about: tell me or show me all the traces that belong to that user ID that I've sent to Stripe. So it's good for tracing. And it's also good for logging. I mean, everything that I capture will be used for myself then to explore the logs themselves. Like show me all the headers. Show me the body that I've sent to Stripe because I know there has been some error, and now with the body, I can actually see what kind of payload that was sent to Stripe and what was the response and understand what went wrong in this specific request.

Jeremy: So you take you take all that information and you write that into CloudWatch logs essentially, right? You're not making ⁠— I'm just thinking, obviously CloudWatch logs are asynchronous, so it doesn't slow down your Lambda function at all. Whereas if you were making synchronous calls and you had to write back somewhere, that could slow down the execution time. So you're just doing the asynchronous stuff.

Ran: Yeah.

Jeremy: Okay, that makes a ton of sense. So what about information? So you mentioned capturing like a Stripe API call. What if that contains a credit card number? Like we don't want to write that to a log, right?

Ran: Exactly. So when you're taking care of the instrumentation, you also need to take care of data scrubbing or sensitive information omitting. For example, well, it's defined by the user. For example, because for dev environments, I do want to capture that because I want to be able to troubleshoot faster my dev or staging environment. But for production, for example, I don't want to capture any sensitive information: any passwords, any emails, anything that is, you know, PII or PHI, like information about the health or the identity of my user. So instrumentation needs to be aware of the data its collecting or allowing the user to omit every [piece of] sensitive data, so I won't capture any headers. I won't capture any payloads. I just want to capture the metadata. I want to know that I had an operation to Stripe, it took this amount of time, that the response code was this, like the status code of the response, and so on. So it's more about meta data. So it depends really on the scenario, but instrumentation needs to be aware that it can collect sensitive data and to give it the ability to omit that data.

Jeremy: And I would think that if you were building, I mean, for some of these things too where you're maybe building an interaction into Stripe, you would want to sort of build your own module in between that. So when I wanted to charge a credit card or I want to send an invoice or I want to do some of these things I would write a module that kind of handled that for me like a data layer, right? That I would then wrap that so that that my developers, when including the Stripe component in their system, or in their code or their scripts or their Lambda functions, they wouldn't need to do this instrumentation again.

Ran: Exactly. So what we're doing in Epsagon is instrumenting each and every library in an essence that will give the developer the ability to omit any sensitive data so it won't be collected and won't be sent outside of the runtime.

Jeremy: Alright, so that's really cool stuff. But what about auto-instrumentation? So Lambda layers are very cool feature that I think some people know about that allow you to include or run code before every Lambda function. And I'm pretty sure that's how Epsagon does some of the auto-instrumentation, but so what does that do? What does that mean to auto-instrument something like that.

Ran: So the layers, it's a pretty cool technique that applies currently only to Lambda. I wish we could do this in containers or in EC2, for example, that every spawned EC2 will get some of this structure ⁠— some of this data. What we're doing behind the scenes is adding our layer, so that includes Epsagon, already prepared for every invocation that the Lambda is running and we're hooking ourselves into the runtime and changing the handler to us. So it means that the request that invokes the Lambda will first come to Epsagon and we will ship it back to the original code. So the other instrumentation brings us the ability to let developers or ops guys just to mark a function in the Epsagon page and say, "I want to instrument it," instead of adding even the minimal amount of code, like two or three lines. Just say, "I wanna monitor that. I want to see traces out of this function because I had these metrics and it's not enough. I want to see traces, distributed traces and instrumentation comes as well."

Jeremy: And so if you're trying to instrument Axios or SNS or DynamoDB or any of the other AWS SDK components so you just include your layer and that wraps all of it. Now, does that automatically get instrumented? Or do I have to do something in my code to say, capture all the SNS calls, capture Axios, capture DynamoDB?

Ran: We offer this auto-instrumentation comes totally automated, so you won't need to configure. I want to instrument PG or I want to instrument Axios or SNS request. Everything will be instrumented for you. You can manually specify I don't want to instrument this and that. But you know, for auto-instrumentation, it's just about frictionless onboarding, having the best experience at no time. So that's what we're aiming to do.

Jeremy: So that's really cool. But if you're capturing all of this tracing data, that's a huge amount of data that you're writing to CloudWatch logs. One, that sounds expensive to capture all that information, but maybe more importantly, what do you do with all the data? Like what happens next?

Ran: That's exactly where distributed tracing comes. So the first part of traces is to get all the information via the instrumentation. And then comes the part of correlating all this data. Obviously [then] comes the third part of presenting all this data, which is a problem for itself. But the second part is to trace or distributed analysis for all these events, all these traces altogether, so we'll be able to correlate a message that's going through one service to another or through a message queue through any other third-party. You want to correlate between all these kinds of event and that's exactly distributed tracing

Jeremy: Perfect. So alright, so you mentioned distributed tracing. We talked about distributed tracing. I think we've covered a lot of these topics, but maybe we can just kind of go down this path of why it's necessary, right? So we kind of know how it works. We talked about a little bit of correlating the events and capturing all this log data, structured log data, putting it all together. I think that makes a ton of sense and I think most people kind of get that idea. But what are we gonna be able to know? Like you know why is it necessary that to trace all these things when you're building these traditional ⁠— I say traditional ⁠— but when you're building serverless applications?

Ran: Up until recently, like recently, like two or three years, I would say that distributed tracing is not a mandatory thing that each R&D team needs to have as part of its arsenal of tools. Today, I think it's almost like a crucial or vital thing that you need to have. The main reason is that we already know that applications are becoming more and more distributed. So, for example, once a user is buying something at your store, you want to make sure that it gets the email to him with the receipt and the invoice and so on, as soon as possible because otherwise he's hanging there, waiting for confirmation or waiting for something to get to him. And in a monolithic way, it's been pretty easy because you had something specific, a single thing that will take care of everything. But now we've got, like between 3-300 services that might take care of this operation: one that will get the API request from the user, from the Web server, the other one that will parse the user request. The third might be something regarding billing that will charge through Stripe or through another service. The fourth one could be something that is mailing users, and all of them are connected to each other, with some messages that are running from one to another. It could be like a star, or it can be like 1-to-1, all the way up until it gets to the email service. And without distributed tracing, you wouldn't be able to ask yourself this question: how long does it take for a user once you buy something until the moment he gets his confirmation. Because if it takes, let's say, for example, a ridiculous number. Let's say one minute. It's not good. I don't want my user to wait one minute in my website for confirmation. I want it to be, let's say, sub-second or let's say sub-five seconds. Other than that, it doesn't meet my SLA. And only with distributed tracing can I really measure end-to-end traces and not just a single trace every time.

Jeremy: Well, yeah, and I think that part of the reason why this idea of distributed tracing is necessary is just because of the way that we're building serverless applications now. And even with microservices, it was a little bit ⁠— things were still a little bit more contained, right? So what happened within a microservice? There were a lot of things working together, but it was still sort of a mini-monolith. I always call it that and I get criticized for it, but I'm going to say it anyways. It's sort of a mini-monolith, right? It does a lot of different things. It has a lot of subroutines that that interact with different parts of system. Whereas when you start breaking things up into serverless, now you've got all these small little functions that do one thing well, and you are using event-driven in this event-driven approach where, like you said, somebody places an order on the website. You aren't going to then call this subroutine, then call this subroutine, then call this subroutine. You're going most likely parallelize that, right? You're going to send it out in SNS. You're gonna have it queued with SQS queues. You're going to use something like the event fork pattern. You're going to have messages flying all over the place. Maybe you're going to use step functions to process them, you know, and now this is the new callback pattern with step functions where maybe something has to happen and confirm before it can move on. So you just have a lot of things happening that are all disconnected, and either they're orchestrated or choreographed. But either way, knowing how all that stuff flows through the system and more importantly, whether all of those things succeeded, it's a hugely important thing in order for you to run your business.

Ran: Yeah, it's as you mentioned, everything becomes event-driven and on top of event-driven, it's asynchronous type of event-driven. So I throw a message to an SNS. I don't care about the response. And I know that someone will take care to charge the user for me. And I send an SQS to something that will trigger another Lambda that will send email through a SES. I don't care. I sent the message. Someone take care to send the receipt to this user at this email. So it's becoming more of a problem to do distributed tracing where everything is asynchronous.

Jeremy: Yeah, and if you subscribe to that idea of using asynchronous transactions or asynchronous messaging, splitting things up and then you know, dividing up your teams, splitting up your team, so one team is in charge of sort of managing that Stripe API and all the billing requests and things that happened there; and another team that, you know, that does the inventory; another team that does the the ordering components and things like that. Being able to see that whole picture, especially when you're not familiar with some of these components in the system, I mean as things start to scale, it becomes very, very confusing if you don't have distributed tracing in place.

Ran: You'll hear blames flying out from one team to another. Everyone says it's okay for me. Maybe it's the other team's responsibility.

Jeremy: The "Works for me" trademark. You know, that's the one I love the most. Alright, so I think there's just a couple more things I want to talk to you about. Maybe things like OpenTracing and OpenCensus. So maybe you can give us a little bit of background about where those fit into this idea of distributed tracing.

Ran: Yeah so OpenTracing, I think it was the initial draft for how to do distributed tracing, more about specification on how to collect and what is the protocol between services to be able to build these distributed traces. And then came OpenCensus, which was like a mix between instrumentation and distributed tracing. It's bounded up together to be OpenTelemetry, so it's no more OpenTracing, neither OpenCensus. It's called OpenTelemetry, and the main thing that it brings is how to track a message when it spans over multiple processes or multiple services even if they're asynchronous. The main thing that it does is to let you know that you need to inject an ID, once you're leaving the process and to extract that ID, once you're coming from another process away from sending a message through a Kafka, I'll inject to that message, "Hey, I'm trace #123." So when the second service will get it, it will extract this information from the Kafka and we'll see, "Oh, I'm part of trace 123." I know that everything is clear for me. I'll continue with that trace along the path. So even if it's asynchronous, this inject-and-extract mechanism that will work along the way.

Jeremy: And these, OpenTracing and OpenCensus, or OpenTelemetry, now they don't actually do anything with the data though, right? It's just more of the standard. It's just sort of how they're supposed to interact with one another?

Ran: Right, so there are some implementations on top of them for some more automated, some more data collection, data shipping to somewhere, but out of the box, neither of them - neither of both or this single new one - will do anything. It's more about: This is the standard. This is the way you should capture traces and transfer them between one service to another. And actually I want to say it's not that easy to build on top of that something because you've got a standard. Now you need to work out your own way how to do instrumentation, and you need to capture this, all events, all the information that you need. These annotations might be endless. I mean, capturing every event in your system. You need to build it in your own. Then injecting and extracting ID through HTTP requests, through message queues, through pub subs, through API gateways, through anything. It's almost an endless list that you need to take care for. Then just handling all this data. I mean, if you're a company that's handling billions or tens of billions of events per month, you need somewhere to store all these events. And as I mentioned before, you need to build something that will present and you'll be able to ask all these kinds of questions on top of this data, which is a problem for itself.

Jeremy: Awesome. So last question here because I love AWS. I'm an AWS Serverless Hero, big fan of the things that they do. But CloudWatch and X-Ray and CloudWatch metrics: they are not turnkey, right? I mean, there's a lot of things that I still have to do when I'm trying to do that. And I have used Epsagon and I really like the the added functionality that it gives you with the ability to look at the logs and just putting everything together, seeing all your applications together. It's really good. And this is not an advertisement for Epsagon. But I do appreciate you coming on, and maybe you can tell everybody just what it is about Epsagon, and, you know, maybe third-party libraries in general or third-party services in general. How does it go beyond what CloudWatch and X-ray does?

Ran: Yeah, I like the question because it's the more broader term, regardless of Epsagon. Having a managed solution will give you a peace of mind that you know that you handed over the monitoring problem ⁠— or let's say a different problem, ⁠but specifically about what we're talking, the monitoring and troubleshooting and distributed tracing problem ⁠— to some other third-party. It will take care of giving all the information for you, for your teams, to be able to make the right decisions. This applies to any third-party that is doing right service. Now you can build everything on top of AWS. We're built on top of AWS as well, so it means that anyone can build Epsagon. However, it's not that easy. It's not that easy to build something that scales to that much of information, that comes out of the box with everything you need, that gives you all the ability to trace and instrument all these kinds of events whatsoever, even if they're in AWS or external to them. And that's the differentiation between having an infrastructure solution to more application-wide solutions that are ready for you. Just plug and play.

Jeremy: Perfect. Alright, well, thank you so much, Ran, for being on the show today and for just sharing all of your knowledge with the community. How can people get in touch with you?

Ran: First is on Twitter. My Twitter handle is @ranrib. Feel free to ping me there, I'm with direct messages open, so just, I'm looking forward to hear some of the interesting serverless, microservices and other cloud environment stories. I really love them. Also on Epsagon.com, obviously, which is where I work. And the last one is the Epsagon blog, and my biggest fetish is benchmarking, so I really love to benchmark resources, services and make sure how they work internally. So things like how AWS Lambda is built behind the scenes or which is the best way to send a message from one service to another like SNS, SQS, Kinesis, a direct call or HTTP — all of these kinds of things are things that I'm writing about, so feel free to get into Epsagon.com/blog. Other teammates write some cool things there as well, but mine are the best.

Jeremy: Well, at least you're modest about it. Alright, well we will put all of that in the show notes. Thanks again, Ran.

Ran: Thank you very much, Jeremy. It's been a pleasure

View Details

About Taylor Otwell:

Taylor Otwell is the creator of the Laravel framework, Laravel Forge, Envoyer.io, and Laravel Vapor. Before building Laravel, Taylor was an enterprise .NET and COBOL developer. He now works on Laravel and its ecosystem of tools full-time.

  • Twitter: @taylorotwell
  • Laravel on Twitter: @laravelphp
  • Laravel: laravel.com
  • Laravel Vapor: vapor.laravel.com
  • Laravel Forge: forge.laravel.com

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Taylor Otwell. Hi, Taylor. Thanks for joining me.

Taylor: Thanks for having me.

Jeremy: So you are the creator of the Laravel Framework, which is a very popular framework for PHP. So why don't you tell the listeners a little bit about yourself and what Laravel is and what it does?

Taylor: Sure. So I started programming professionally in 2008 after I graduated college, and I was originally a .NET developer in the enterprise world and started tinkering around with PHP on the side in 2010, and sort of in the fall or winter of 2010, I wrote my own PHP framework, sort of inspired by a variety of things: inspired by Rails, inspired by my experience with ASP.NET MVC, inspired by Sinatra and Flask and all these other frameworks, and sort of put something out there in the PHP world that sort of riffed on all of those ideas and brought them together in sort of a really productive way, I thought. And so I put it out there in 2011 and you can think of it as sort of Ruby on Rails for PHP, mainly. So it has, you know, controllers and routes and a database ORM and queued jobs and all kinds of other stuff to let you build web applications in PHP in a very productive way.

Jeremy: Awesome. So people are probably wondering why you're on a serverless podcast. But recently at Laracon, you just announced Laravel Vapor. So why don't you tell us about that?

Taylor: Yeah, so Vapor is something that I've been working on for about the last nine or 10 months, full-time, 40 hours a week. And it all started really over a year ago. I was just really inspired by sort of the serverless ecosystem, what people like Zeit were doing for JavaScript with their Zeit Now product and I really wanted something like that for PHP and something that could tell sort of the whole story for PHP, because there's a lot of moving parts that Laravel developers expect if they were to go on serverless like, you know, what do I do about my database migrations? What do I do about my queued jobs? And so I wanted to build a product around serverless that sort of made sense for Laravel developers and that they would understand, that would provide a really good experience for them.

Jeremy: So I want to jump into the details of Laravel Vapor. But let's start with some background on Laravel. So you said it was sort of this Ruby on Rails for PHP. So what types of applications do you see people building with Laravel now?

Taylor: Oh, gosh. I've seen everything from help desk, you know, accounting applications. I've seen, you know, all kinds of back-office applications, intranet applications. Of course, I've written Forge, a server management platform on Laravel. I have a zero downtime deployment platform on Laravel. So I've really seen such a variety - hotel room management platforms - almost anything you can think of really, I've seen on Laravel.

Jeremy: So is this something that anything you can build with the Laravel Framework now, you're going to be able to just do serverless-ly with Vapor?

Taylor: That's sort of the hope, you know, that your application will translate well and there's a few differences, you know, when you're operating in serverless, we can get into. But that was sort of the goal, though, is to make it so you can deploy on Vapor, and things sort of work as you would expect, and you could build your application as you're used to in a traditional server environment. You can just deploy it on vapor and sort of have the same experience. That was the end goal I was shooting for.

Jeremy: So does the development workflow change now that you're dealing with different types of resources?

Taylor: Sure, your local development workflow is a little different depending on what you choose to use. You know, most Laravel applications are used to interacting with something like MySQL. We also ship with Vapor, a sort of a docker container, where you can run all of your unit tests against the production PHP build that actually runs on a Lambda side of Vapor. So we try to provide some tooling to make that experience a little better.

Jeremy: Very cool. So you mentioned this interest in sort of the serverless ecosystem, but what were your main reasons for building Laravel Vapor?

Taylor: Because I don't ever want to think about servers ever again. Basically. So it goes back to sort of Laravel Forge where really I built and released Laravel Forge in 2014. Sort of the idea there was, you know, I was building Laravel applications professionally, and I was constantly configuring servers with nginx with PHP, with Redis or whatever I needed. You know, and I was building them on let's say, like, DigitalOcean or Linode or some VPS provider and I had written bash scripts to do all that and automate all that. And so I sort of built a platform around that called Laravel Forge. And you can, you know, link your DigitalOcean account, create a 2GB server, and it installs everything you need, and then you can deploy your application out there. And that's all great. Like that works really well for a lot of applications. But there's still like a lot of headache that comes with that, even though a lot is automated for you. For example, like my operating system goes out of date. I sort of kind of have to worry about SSL certificates renewing. I have to worry about various vulnerabilities. I have to worry about all kinds of stuff I just don't want to think about. And then if I'm load balancing those servers, now I've got, you know, 5-10 servers to worry about. All those same problems just multiplied. And so while Laravel Forge does provide a lot of automation and sort of the traditional server environment can be automated in some fashion. The idea of going totally serverless and just never ever thinking about servers at all, never thinking about, you know, how am I gonna load balance them. All of that sounded really, really appealing to me, after managing, basically, thousands of servers with Forge so that was really appealing. And that's what really drew me to the whole ecosystem.

Jeremy: So was there something that prompted this? I mean, at what point did you look at AWS or Microsoft Azure and say, "Okay, now I can go serverless with this." Because PHP wasn't even supported as a runtime until November of last year when they came out with with custom runtimes for Lambda.

Taylor: Yeah. So there were some key things that happened. I remember the first most important thing was we could get Laravel up and running with a Node shim where when a HTTP request comes in, we use Node to sort of invoke PHP. And people were doing this with other languages too, you know, before custom run times. But always the big problem for me was how are we going to hook into the Laravel queue system in a very nice way because there was not an official integration with SQS and Lambda at the time. But last year, when AWS announced that official integration with SQS, that was sort of one of the last puzzle pieces that really clicked into place to where I was like, hey, I think I could actually build a pretty nice platform for Laravel on Lambda and things would pretty much work how people would expect them to, because I could set up that SQL event source mapping for them — or SQS event source mapping for them and sort of take care of making Laravel's queues work as you would expect. And then the custom runtimes were sort of a cherry on top. We had everything running before that came out with the Node shim approach. But once that came out, we got a huge performance boost by just shipping a custom PHP runtime with a PHP-FPM. For example, I think just like a "Hello World" request on the Node shim in Laravel is something like 30 to 40 milliseconds on the Lambda side, and then once we move to the custom runtime with PHP and FPM and all that, now it's like six or seven milliseconds on the Lambda side for a "Hello World" request, so a huge performance increase being able to ship that. So those two things were really, really key things that shipped last year that made this all possible.

Jeremy: So the compute obviously runs on Lambda because you're using custom runtimes. Are there any limitations, because it only runs for 15 minutes, right? And so if you're doing Laravel queue jobs that might be running longer than that, are there's some workarounds or are those just sort of some of the limitations that are built into the system?

Taylor: I think it's just one of the limitations you have to deal with. So, I mean, I guess you have a couple options. If you can somehow chunk that work up into multiple queue jobs or, you know, if you could somehow find some way to sort of judo the whole problem on your end, that works well, otherwise, you know, I just may not be a solution that works for you. For me, personally, I don't really have any apps that have queue jobs that take longer than 15 minutes, so it hasn't been something I've really messed with. But, you know, it is a sort of a stated limitation in the documentation of Vapor, something you have to either live with or work around.

Jeremy: And I would think that for most web-serving type applications, it wouldn't matter anyways.

Taylor: Typically not.

Jeremy: So the other thing was file uploads, right? So you're using S3 now to store files, so that's a little bit different than uploading files to a Laravel server, for example?

Taylor: Yeah, sure. Yeah, in the old school days, people would probably just have some form on the front and that posted straight to their PHP backend. And then there's the, you know, the global $_files array in PHP that they would interact with. But yeah, on the Vapor side, we really encourage people to just send their files directly to S3 from their client side using JavaScript and I actually built an NPM package that tries to make that a little easier because there's a few steps involved to doing that where we need to generate a pre-signed URL to S3 and then get that back to the client and then send the file using that URL in headers. So I wrote a little NPM package that has a Vapor.Store method where you can pass it a file and it will call a backend route to the Lambda to get the pre-signed URL. Once it gets that back, then it can just send the file directly to S3. And once that's done, then you can kind of ping your backend and say, "Hey, the files uploaded," and we can sort of do whatever we want. From that point, we can manipulate the file or whatever you want to do. So that's another, I would say that and the time limit are sort of the two main differences for sort of the development workflow that developers are gonna have to get used to. I think even if you're a traditional server environment and you're using something like, let's say, Laravel Forge to have a DigitalOcean server, streaming straight to S3 from the frontend is sort of still like a good idea, I think, and not sending big files to your PHP server. So it's already sort of, I would say, a good practice that I would recommend. But with Vapor, it's really more of a requirement that you have to start working that way.

Jeremy: All right, so let's get into the nuts and bolts of Vapor here, because this is a really cool service. And this is a hosted service, right? You host this for your users?

Taylor: Right. This is a hosted service, which means we can control things like team members, permissions, keep a record of deployment history, all of that.

Jeremy: But all of the resources are in the customer's AWS account?

Taylor: Yes, exactly. When you sign up for vapor, the first step after you sign up is to link your own AWS account or even multiple AWS accounts if you want to, so that we can create things on your account.

Jeremy: All right, so let's talk about some of these resources. So let's start with databases. What kind of databases or what can you do with databases in Vapor?

Taylor: So you've got a couple options. You can do just your traditional RDS server with a fixed-size small database instance or a development instance, all the way up to sort of the memory-optimized instances as well. And then you could also do an Aurora Serverless database, which is also MySQL, even though they've recently announced the Postgres serverless support. But right now, we've got those two options for databases. So you can pick one and then just link it to your application in your Vapor.yaml configuration file and deploy, and you're sort of good to go, and Vapor takes care of injecting all the environment variables that Laravel needs to connect to that database. So you don't really have to worry about, you know, how you're going to get your database host, your database to username and password, all that into your Laravel execution environment. Vapor injects all of that for you.

Jeremy: So if you're using databases with Lambda, then your Lambda has to be in a VPC. So what about all the complexity around VPCs and NAT gateways and that kind of stuff?

Taylor: So as soon as you create a project on Vapor, we create, or we ensure, that a VPC in that region exists. And then if we see you deploying a database, like a serverless database, we're going to ensure that that VPC has a NAT gateway attached to it. We're going to ensure that all the subnets and security groups and stuff look okay, and we're gonna make sure that the database has the right subnet group as well when we create it. We try to automate all that. I would say that's one of the more complex pieces of getting an application up and running on Lambda and doing all this. So we try to make that as smooth as possible and keep it from getting out of hand. But it does intelligently take care of that for you. So you don't usually have to think about it when you're using Vapor.

Jeremy: And you can control the database, right? So you have this UI that shows metrics and things like that?

Taylor: Yeah, all of the stuff, or a lot of the stuff that RDS would let you do, we sort of provide a UI on top of it. So you can restore the database at a certain point in time, you can scale the database, if it's a fixed-size instance, and you can get kind of cool metrics like your max connections, which is pretty important in the serverless world, your average connections, how much CPU you're using, how much free disk space you're using If you're using a server that has a fixed disk size. And you can sort of monitor that, as well as configure alarms for that straight from the Vapor UI, which ties into CloudWatch alarms. So you can set an alarm if my max connections is more than 100 for five minutes, then I want you to email me, or ping me on Slack or whatever.

Jeremy: Cool. So what about queues? Because SQS on Amazon is really one of the most scalable services that they've got. So how do queues work?

Taylor: Yeah, I love SQS. And I've used SQS in production even before this. And so what happens is when you deploy a given project in a given environment. So say I'm deploying my Laracon project for the Laracon website in the production environment, we ensure that a queue exists and the name of the queue is sort of conventional where it's like the project name-environment, or whatever, and we ensure that exists. And then we set up the event source mapping between SQS and Lambda so that when a queue job comes into SQS, it invokes our Lambda. And then we intercept that on the Lambda side and we say, "Oh, this is ah queue job," and we send it to the Laravel queue worker for you and so on and so forth. So it feels really transparent. There's really no configuration for queues in Vapor at all. You just deploy and dispatch your jobs, and it works just like normal at literally zero configuration in your YAML configuration file.

Jeremy: And what about caches? Because that's a huge part, obviously.

Taylor: Yeah. So for caches we built a UI obviously on top of Elasticache to create Redis Clusters. I didn't add memcached right now, but we may visit that later. So you can create an Elasticache Redis Cluster, and you can scale it up to however many nodes you want and pick your, obviously, the size of your nodes. And it's sort of the same story as with databases. You attach that to your project environment, and then we inject all of the necessary environment variables so that Laravel can actually connect to that Redis Cluster. And, you know, so things like the hosts, we set your cache driver in Laravel to Redis because it can be other things. And everything sort of again just works, and similarly to databases, you can get some kind of cool metrics, I think, where you can see your cache hits, your cache misses, your hit rate percentage, and then also the CPU utilization across all of your nodes individually. So a pretty cool little UI on top of that that tries to make it as easy as possible, and the same set up as databases where we sort of get the VPC set up correctly and all of that.

Jeremy: And what about for local development? Because that's always sort of a tough thing in serverless right now. And so you can't directly connect to a database or cache in a VPC. You have to use like an SSH tunnel or a VPN. So you take care of all that for us, right?

Taylor: Yeah. So the approach I took there was, for both cache and databases, is I let you create what we call in Vapor a "jumpbox," but I think other people call that outside of Vapor as well. But basically, it's a small T2 nano that we put in your VPC, when you want it, and it just takes a minute to provision. But what we can do is since that's inside your VPC, we can do interesting things like, I built a Vapor cache tunnel CLI command where, when you run Vapor cache tunnel and then give it the name of the cache you want to tunnel into, it opens an SSH tunnel through that jumpbox and then opens port 6378 as sort of a port into your Elasticache cluster. So that means that locally, like here on my iMac, I can open up my Medis GUI for Redis and then connect to port 6378 localhost, and I'm connected to my Elasticache clusters through that SSH tunnel. So that makes it really easy to sort of, especially during development or in like my staging environment, I want to see what's happening in the Redis Cluster. I can see what keys, are there, blah blah blah, and same way with database. I can use that jumpbox as an SSH in my like TablePlus, if you have a database management GUI on your local machine or whatever you want, you can connect over SSH through that box to your database so that you can connect to like your Aurora Serverless database in a nice UI so really actually pretty handy. And then if you want to, when you're done inspecting it or whatever, you could just delete that jumpbox in Vapor and get rid of it. And so they're so fast to provision. Sometimes I just make them when I need them, and then get rid of them later.

Jeremy: Yeah, and I think a T2 instance costs, like, nine bucks a month or something like that.

Taylor: Yeah. Very cheap.

Jeremy: Yeah, and and so that that tunneling technique. So I actually use that pretty much all the time for most of the workflows I have because if you need to connect to Elasticsearch or a database or anything like that, it's just so much easier than setting up a box with the VPN and having to manage all that stuff.

Taylor: Right. And I think the things that I think something that led me to some of those features is that was really nice. Is the whole time I'm building Vapor? I'm sort of deploying Vapor out on to Lambda, so I sort of had this dog fooding my own product on Lambda that's helping me discover sort of those kind of pain points and help me sort of flesh out the product really.

Jeremy: So we talked about S3, and how you're sort of managing all of the file storage using that. But you also do a CDN as well, right?

Taylor: Yes. So when you deploy a project, this is sort of the nice thing about managing Laravel and managing Vapor at the same time as I can make all these nice assumptions about how things work. So when you deploy your project, we extract all the assets out of your public directory, which is where Laravel projects keep like things like their style sheets, their javascript and all that. We upload that to an S3 bucket, which has CloudFront in front of it. So we configure a CloudFront distribution for you that points to that S3 bucket. And then once you're on the Lambda side we inject an environment variable called asset_url and Laravel knows to look for that when generating asset URLs, so that when you generate, for example, you're link to your stylesheet or your link to your script tag to your JavaScript, it automatically has that CloudFront URL in front of the file name. So that makes it really nice to automatically get all those assets on CloudFront. Because that would be kind of a chore to sort of do manually. And we tried to make that as smooth as possible.

Jeremy: And you create all the buckets and do all that stuff?

Taylor: Yep. On deploy we make sure all of that exists.

Jeremy: Okay, so what about metrics? You mentioned that you could get some database metrics and things like that, but what about, like, overall metrics or alerting? What does Vapor do for you with that stuff?

Taylor: Yeah, so on the web and queue side we do metrics like total invocations. And of course, you can set, like, over the last 30 minutes, the last 24 hours, the last seven days, whatever different time periods. So we let you look at http invocations, queue invocations, because those are two separate Lambdas when you deploy your project we actually have a separate Lambda for web stuff and a separate lambda for queue and CLI stuff, mainly because that lets you manage the concurrency limits and memory limits separately for those two environments because I think it typically makes sense sometimes for those to be different configuration values. So you can monitor that, and you can also monitor the average duration of both the web and the queue/cli side. And then you can set up alarms on that stuff. Like if my average duration has spiked over some value for a given number of minutes, I want you to email me or whatever. So pretty useful metrics to monitor. And then we also have kind of a slim UI on top of CloudWatch Logs just for logs in general. Where if I visit the logs tab in Vapor, I can see the latest logs for the past hour for both my web and then a separate tab for the CLI and queue side. And that's kind of nice, because if you go out to CloudWatch, you know you're digging through multiple different log streams and stuff, and it's pretty nasty. So we try to interleave all that.

Jeremy: So you had mentioned earlier one of the itches or one of your own itches that you were trying to scratch was things like certificate renewal and some of that higher level stuff. So, DNS and certificate management that's all built in and managed for you. Correct?

Taylor: Yeah, I sort of bake that into one screen where you can actually purchase domains straight from the Vapor screen. Or if you already own the domain, you could just add it to Vapor. And it picks up on all the records that are already in Route 53. But then you can request the certificate straight from that screen, and what it does is actually request a wildcard certificate for the domain using DNS validation. Or you could do email validation, but we really strongly recommend DNS validation within Vapor. And that sets up the proper CNAME Records for the DNS validation to work and prove you own the domain and all that. And then once that is issued, of course, we start using it when you deploy. We actually require every application that's deployed to have a valid certificate. There's no way to deploy a non-SSL application on Vapor. So we let you do all that and take care of the renewals or really Amazon's taking care of the renewals for you on the certificate side, because we're using the certificate manager right there on Amazon.

Jeremy: So the other thing that you typically do is you'd have, like, a DEV stage, a STAGING stage and a PRODUCTION stage. And that's sort of a typical serverless way that these things are done. But you actually don't need to worry about domains because you have his vanity URLs.

Taylor: Yeah, that's one of my favorite features, too. So I use Cloudflare and actually just purchased a bunch of domains like vapor-farm-1, vapor-farm-2. So I own a lot of these domains. And so what that means is, when you deploy, like, for example, like you said, your staging environment, we assign each environment its own vanity URL. It's kind of like a Heroku-style URL, it's like, gorgeous-mountain-scape-124.vapor.build or whatever. And I add that DNS record that CNAME record to my Cloudflare account, actually, because I own all those vanity domains and point that to your serverless application and and then we add our own Vapor vanity certificate into we import it into the certificate manager so that it all works. So that's actually really handy, because one of them pain points if you're deploying PHP to Lambda right now, it's those... by default, those API gateways have, like the slash stage suffix on them, which can just kind of wreak havoc with various PHP frameworks. They don't know what to make of that extra segment in the path, and so having that clean vanity URL is actually a really nice way when you're getting started to access the application.

Jeremy: All right, so develop locally, and then how do you actually get all of that code onto AWS?

Taylor: Yeah. So when you run the "vapor deploy" command on your command line, we build the whole project with build steps that you can specify, like installing your Composer dependencies. You're running some in NPM stuff, and then once that's done, we actually zip up the whole application and send it to S3, and then we ping the vapor backend that says "hey, this deployment ready to go." Here's where the code artifact lives on S3 and then from there the Vapor backend updates the function configuration, updates the function code to point to the new S3 code we have. We use function aliases and Lambda. So at the very last point we switch, the production or the staging or the testing alias to the new version of the Lambda.

Jeremy: And what about if you're doing CI/CD?

Taylor: So if you're doing that, we actually you can ship the Vapor CLI, it's just a single compiled binary with your code. And so, like, if I'm doing let's say, CodeShip, for example, within my CodeShip build steps, I run my tests, blah blah blah. Then my deploy step, I could just use that Vapor CLI and I get my credentials in there through an environment variable so on my CI server, whatever, I would configure my Vapor API token as an environment variable. And then I could run "vapor deploy" straight from the CI service, and it would use that environment variable to authenticate with vapor and deploy my project

Jeremy: And so you have a CLI tool as well as the full on web interface. Correct?

Taylor: Yes. Yeah, and everything you can do on the web interface you can almost everything else so you could do on the CLI interface. You can't do things like change your password or update your billing plan, but you can create databases you can create caches, you can do all that stuff.

Jeremy: And speaking about the billing plans. So this is just SaaS, like monthly type thing?

Taylor: Yeah, just a monthly SaaS. So right now, I've got a price that I think a launch price will probably like $29 a month and a full price will be like $39. Which is the same price as the PRO level of Laravel Forge. And of course, that's unlimited teams, deployments, projects, whatever.

Jeremy: All right, so that would be one part of the building. But then all of the resources in the customer's AWS account, they would be responsible for the costs of those as well. And then, how much control do customers have? Can they go into their AWS account and actually manipulate some of these resources and tweak them if they want to?

Taylor: Yeah, they could tweak things on their own. We try to like, even if you change things in Route 53 we import those records every so often. Of course you wouldn't want to go too far off the beaten path so that Vapor gets like, confused about the state of things. But if they ever wanted to walk away from Vapor, that is kind of one of the nice thing is, all that stuff is in their AWS account, so they could just kind of walk away and build their own deployment process around Lambda. And I've always liked that approach with Forge. It's simpler for us because we don't have to worry about all that billing on our side. And I think it's just nicer for the end user because we don't have to mark up AWS prices for anything. You still own everything in your own accounts, so I think it's just sort of a nice, clean separation.

Jeremy: Yeah, and there's other things you could add as well. Like if you wanted to use SageMaker or Amazon Comprehend, you could use the AWS SDK and you could integrate with those things. And then that would all be managed in the same account. So, that's very cool. Okay, great. So, let's talk about where you see the future of Laravel going. So, is it serverless?

Taylor: I think that is definitely a big part of the future. And I think the serverless philosophy and Laravel philosophy are very similar. From the very beginning. So when I launched Laravel, the idea was that you could just focus on your code and Laravel handled all of this sort of nasty stuff, like authentication and session management, all of that for you. And I think a lot of the route philosophy of serverless is very similar, where the goal is where you could just focus on providing value and writing the logic that makes sense for your business. And in that way, I think, Laravel's philosophy and the serverless philosophy are very aligned, and they're sort of a nice fit together. And so I hope the future of Laravel is tied in with serverless. And I'm trying to sort of be ahead of the trends here with this, you know, and try to be the first kind of major php platform for serverless out there that sort of tells the whole story from databases to queues to mail to assets and all that. Um, so, yeah, I'm excited about it, mainly because I just believe their philosophies are so similar.

Jeremy: Yeah, and one of the great things about serverless, obviously is the massive scalability of it. And Laravel forges is a scalable product as well. But you're still doing a lot of that manually, right?

Taylor: Yeah, sure. It could be scalable if you, you know, build a load balancer and 10 servers or whatever. But now I'm managing 10 servers that would have to really manage. And so yeah, sure, can you build that kind of scalable platform on a traditional server environment? Yeah, but it just is a lot more headache, I think. And I would rather just do "vapor deploy" and be done with it.

Jeremy: Yeah, no, that totally makes sense. So let me ask you about serverless in general. So what are your thoughts about serverless being sort of the next evolution or the future of Cloud computing? Because obviously there are a lot of people using serverless now, but I still feel like you say to somebody "serverless" and they're like, "Ah, what's serverless?" But this idea of moving up the stack and focusing on your code and getting rid of all of that undifferentiated heavy lifting, I mean, what does the next five years of serverless look like?

Taylor: Yeah, I think the next five years will be huge for serverless, I really do. I think it is the future, because what's the alternative, really? Like more complexity, more configuration files, more weird container orchestration stuff? I don't really think that's the future, you know, that people are gonna naturally gravitate towards. I think people want simpler things. And I think at the end of the day, serverless is simpler. It's going to only get more simple as the tooling gets better, as the platforms get better. And to me, it's the real endgame, you know, of the whole server thing, just deploy your code and you focus on your code and let the provider focus on the infrastructure.

Jeremy: Yeah, totally agree. So, listen, you're obviously doing your part here, and anyone in the PHP community, I think this is just a huge step forward for them, and a big vote of confidence for serverless and the serverless model, and obviously, what you can do with it. I totally appreciate what you're doing, so thank you so much for that and, obviously, for coming on. So if people want to find out more about you, or Laravel or more about Laravel Vapor. How do they do that?

Taylor: Yep. So you can follow Laravel on Twitter @laravelphp. You can follow me personally on Twitter @taylorotwell, or you can email me taylor@laravel.com.

Jeremy: And if you want to sign up for Vapor, where do you go?

Taylor: vapor.laravel.com. That sounds a good thing to add.

Jeremy: You probably want to mention that. All right, Taylor, thank you so much. It was great.

Taylor: All right. Thanks for having me.

View Details

About Erik Peterson:

Erik Peterson is the CEO and founder of CloudZero. Previous to founding CloudZero, Erik was Director of Technology Strategy for Veracode and has nearly 20 years of software industry experience, including senior leadership and technology roles at HP, SPI Dynamics, GuardedNet and Sanctum. Erik has also held IT & InfoSec roles at Moody’s Investors Service, SunTrust Bank, U.S. Embassy Vienna, Austria and the United Nations International Atomic Energy Agency where he provided technical assistance to UN weapons inspectors.

  • Twitter: @silvexis
  • Email: erik@cloudzero.com
  • CloudZero: cloudzero.com
  • CloudZero Twitter: @CloudzeroInc

Transcript:

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Erik Peterson. Hey, Erik. Thanks for joining me.

Erik: Hey. Great to be here, Jeremy.

Jeremy: So you are the CEO at CloudZero in the great city of Boston. So why don't you tell the listeners a little bit about yourself and what CloudZero is up to?

Erik: Sure. So gosh, so I'm a recovering AppSec person, actually, by trade. I think I spent 20 years in the application ⁠— in the security industry trying to move the needle on one thing, which was to get developers to care about security. I didn't necessarily start there, but I I certainly thought a lot about application security through the years and where I think the applications security industry ended up is a good place, focused on the people who create the software that we care about. But about 10 years ago, maybe 11 now, in 2008, I got bit by the cloud bug and I started experimenting with AWS and taking that where I could take it. And I had the good fortune of bringing Veracode, the company I worked at before CloudZero, over into AWS and had a lot of fun doing that and learned a lot along the way. So, recovering AppSec person. Now true cloud connoisseur I hope.

Jeremy: And what's CloudZero all about?

Erik: So CloudZero. It's pretty simple. It gets back to my roots. I want developers to care about cost, right? And so CloudZero, we're the first cloud optimization platform that is specifically built to tie engineering decisions directly at cloud cost. You look at a lot of cloud optimization solutions today, they're focused on the finance team or parts of the organization that are outside of the people who are actually making the decisions writing the code. And so we want to empower DevOps team to make smarter engineering and infrastructure decisions. And we do that by giving them a platform that could allow engineers to understand in real-time the cost ramifications of their actions. So really powerful solution that, ultimately, we're going to help the business manage costs, move faster and drive innovation for it. And we love developers. We're focused on that world.

Jeremy: Do you have any big features coming out that you want to share with the audience?

Erik: Yeah. So we are building a whole set of capabilities for engineering teams to get right into the details of what matters most to them, which is how much are the things that they're actually building costing them, and take out all of the noise. You know, today if you go look at Amazon's Cost Explorer; you look at another product. You see all this data related to cost. All I care about is what is the thing that I'm working on right now? What does it cost me? And how are my decisions affecting that? So we have a number of new dashboards that are coming up for that and a few other little surprises around the corner around anomaly detection coming out this summer.

Jeremy: Awesome. So I wanted to have you on to talk about an extremely exciting topic that I actually, surprisingly, am a little bit passionate about because I do see a tremendous amount of value in this. But I want to talk about cloud computing costs. And obviously, you have quite a bit of experience in this, but where I want to look at this is we now have this sort of very, very granular billing that goes well beyond what maybe a SaaS company might provide. And obviously, you have SaaS bills and that's a metric that you could use. But now that you have cost associated with every sort of cloud engineering action that you take, how do we need to think about this differently? Maybe let's start there.

Erik: So, you know, I think every cloud engineer should view cost as something that they ⁠— their expertise in understanding the bill needs to be something that they feel proud enough to put on the resume. You think about what was the very first Amazon service. It was, a lot of times, it's really easy to say, "Oh, it's SQS or S3." No, it was actually Amazon Billing, right? Because...

Jeremy: That's good point.

Erik: Amazon wasn't going to do anything if they couldn't bill you for it. And over time, they've figured out how to, like you say, get deeper into the metered billing. We have a millisecond billing. EC2 used to be billed by the hour, and now could be billed even tighter than that. You go look at the reports. Everything is kind of normalized to the hour, and it's a little bit more complicated to figure out. But the key kind of thing here is, whether we know it or not, as software architects, engineers, DevOps engineers, when we moved from on-prem in the cloud, we had a whole lot of constraints that just disappeared overnight. And you know, this decision about how much things cost, what used to be made for us, somebody went and bought a bunch of servers. They put it in the basement, and that was all well and good. And then we just tried to maximize our usage of that resource. Now, somebody gave us an Amazon account, and our instincts, as engineers, are a little bit off, because our instincts are how can I get the fastest path to value for my customers, innovate quicker build, you know, new capabilities. And your intuition is to expand to use all available resources in front of you in order to achieve that goal, right? It's certainly what your boss is telling you or the CEO is telling you. And so we go, "oh, I have this infinite scale. Let's go nuts." The problem, the flip side of that, of course, is if I have infinite scale, I also need to have infinite wallet, right? If I don't have infinite wallet, then actually, the reality is I don't actually have infinite scale, and so as engineers, I think we need to move past caring just about performance and uptime, and we need to add a third item to our kind of list of operational metrics, and that's cost.

Jeremy: Yeah, and and I don't know how you knew that I have a bunch of servers in my basement. They're all turned off now, but I literally have a bunch of old servers in my basement.

Erik: All have some dirty, dirty little secrets.

Jeremy: Exactly. You wouldn't believe how many hard drives I have because I didn't want to throw them away when I closed down my data center. But so you mentioned⁠ - and I think this is important - you mentioned this idea of the sort of the purchasing decision, right? And in the past, it has always been okay, we need 100 servers. We need this many copies of Windows server or we need this Oracle license or whatever we need, and those things were fairly easy to plan for. And again, they were these purchasing decisions by sort of the, I guess, the C-levels or the purchasing department or something. And you would follow along with these...

Erik: The powers that be.

Jeremy: The powers that be. Yeah, and you'd have these budgets, but I think that just becomes a little bit more ⁠— it's more difficult to plan for. And we can we get into this more in a few minutes. But I think what's really interesting and what has changed, at least from what I've seen, is now that we have this very detailed and granular billing, we can use, like you said, we could use that cost actually as a KPI for our business to understand how much we're spending for every action that we're taking. And so you could actually see, you know, for a customer that uses X number of Lambda invocations, and this many SNS messages, and this many step function executions, and this much data storage, you could actually calculate very, very closely how much each customer costs you from that cloud infrastructure standpoint. So I just think that's a really, really, really interesting thing that you can do now.

Erik: If you're a SaaS vendor, you know your value delivery chains is built on top of cloud. That's your cost of goods. That's your gross margin. You need to understand that if you're going to deliver a profitable product to the market and and you want that conversation to be part of your entire organization because, I mean, the reality is is that the buying decision is being made by your engineering team now, right? They choose: am I going to use this type of instance or that type of instance? Am I going to implement this kind of code or that kind of code? They make a buying decision every moment of every day. Essentially, every time, every line of code that they write, they're making a buying decision, and and so you have to think about that. And then it gets even more complicated, though, because there are so many intertwined, and particularly in the serverless world, which is so, I think, honestly I'm sure our listeners here will appreciate our point of view, is that we think, we believe serverless is the future of all computing. But, you know, it's even more powerful because you create these very interesting applications that are composed of lots of different services. It's not just Lambda compute. It's I have Lambda connected to SNS passing to SQS, DynamoDB, Kinesis ⁠— all these things flowing together. And I'm not just going to the cheat sheet on Amazon and saying, "well, how much does it cost for one hour of compute?" to try to estimate my costs. No, I now have to think through that whole story, and I think it's kind of a shame that, actually, for most organizations, they consider the state of the art there to be well, let's just try it and see what happens. And a lot of times they try it and test and they go, "Oh, looks like it was gonna cost a couple bucks. Great. Let's ship it." And once it gets into production, it's a much different story, and they just don't ⁠— organizations really struggle with this. It's unfortunate. (10:53)

Jeremy: So speaking of organizations and struggling. So this is sort of like a cultural change, right? I mean, if we think of trying to get our developers to think about costs now. Because in the past it was, "I wrote some code. Here you go." And I think there were a lot of developers who did have, that they were cost-conscious about how much they were spending, you know, depending on how many services they were using and things like that. But I think it's a little bit different now. And you have a term for this, right? You call this FinDevOps? Is that sort of what you mean by FinDevOps? This idea of this cultural change.

Erik: Yes. So I've always viewed DevOps as the culture that comes along with a cloud-driven lifestyle and that it embodies a lot of things. And I've spoken a little bit about the relationship between the cloud and DevOps, and and also SRE as being a practice that you can apply to that culture. The thing that was missing from the, what I thought was really missing, from cloud culture was an appreciation for the spend, appreciation for how much things cost, because there's a real tight relationship between a well-architected system and a cost-effective one. I've looked out over hundreds or maybe thousands of different systems now, and I'll tell you in every time, first place I'll look is the bill, and it'll give me a better insight into the architecture and what's built, sometimes better than any other data source. And so the culture of caring about the cost of things, the financial aspects of it, was missing from engineering, from the engineering discipline, I think, in a big way. And I wanted a kind of draw attention to that. And that's the whole point of FinDevOps. It's also about understanding. Over time, I want an organization to understand, what's the true kind of flow of capital through my system. If every transaction costs me $12 for one customer transaction, but I'm only charging my customers $3 per transaction, that capital flow is is not going to work out well for me in volume over time, right? I need to find the balance. But yeah, I think FinDevO ps has done a good job of kind of drawing attention to this. And I've spoken to a lot of engineers who appreciate this, but their impact or they're kind of introduction to cost has been somebody from the CFO's office coming down to their organization and then spending an hour yelling at everybody because the bill is too damn high. Meanwhile, that person leaves the room, and then the CEO walks in and says, "Why are you guys not delivering more value to the customer more?" And it's complete imbalance between the two organizations. Everybody needs to be able to have one kind of common terminology for understanding as to why we built it the way we built it, how it's delivering value and how much it costs; it needs to be part of that conversation. (14:00)

Jeremy: Yeah, and I think that you have this issue, especially with purchasing departments, that, or whoever the accountants are, seeing these fluctuating bills in the cloud and not understanding. They just say, "Oh, well, it cost X amount of dollars last month. So this month, it should cost the same, right?" But maybe we had more users. Maybe we added users or whatever. Maybe we did something. We added a new feature, and suddenly new features might add thousands of dollars worth of costs. Right? So do you see that sort of butting of heads between sort of the developers or the engineering teams, and then, you know, as you as you call them, the powers that be sometimes?

Erik: There is that tension there. I mean, sometimes it's a healthy tension, but there is that tension there, and it kind of, I mean it goes like this: imagine you needed to explain, let's say, a very complicated system that you constructed, and now you're trying to explain it in French to the Germans, right? You're speaking a different language. And that's the hard part, right? You know, you can ask the question, "Well, why did we spend $20,000 this month on EC2 more than we spent last month," for example. And well, it's because the product team had a new initiative. We had to do a migration. We had to do this. We had to move data from over here. We had security requirements, so we needed to encrypt the data. So we're calling the KMS API a lot. And then that resulted in a whole bunch of new storage and processing. And you're talking, talking, talking, and then you look up and there's just a glazed-over look on the finance guy's eyes and they're going, "Yeah, no, no why did we spend $20,000 more this month? And how much are we going to spend next month?" And they go, "What? I can't talk to you. Get out here." Right? And ultimately you want to tie it back to well, look, this product initiative cost this much money, and we forecast it to be X. And we have an idea, before we actually go down that path, how much it's going to cost and cost has been a part of it. Because there, I mean, for a long time in engineering has been a notion of non-functional requirements, right? What kind of performance requirements do you have? What kind of uptime requirements do you have? And the hard question that I think organizations need to ask themselves is, "Well, what kind of cost or budget requirements do you have?" And at what point are you going compromise the the budget for the user's experience or vice versa? You know, you are you gonna go, "You know what. User experience matters at all costs. Even if it's $1,000,000 in extra spend this month, our users must be absolutely happy." Okay. Make that decision consciously. Today, I don't think anybody's consciously making that decision.

Jeremy: Yeah, and I think that's a really, really good point because you're right. I mean, at some point, we have to trade off certain things. I mean, if we had unlimited scale, then like you said, you would need an unlimited wallet to do that. So I think planning around that is a good point.

Erik: Well, I'll throw in one thing, you know, like this decision. It's not like these decisions are new. But the difference is that they've always been made for us.

Jeremy: That's a good point.

Erik: We had the CFO and the CEO, or maybe the CIO decide, "oh, we need 50 servers of this class and we worked with the teams." And then that's what we put in the basement, right? So now your performance envelope and everything has been made. And when you run out of capacity, everyone kind of at the time, you know, I'd have a good joke about it. Like, "oh, server's down because we're having, everything is so successful. We have 1,000,000 users. Our product is wonderful." But, you know, today people go, "Why is it down? You don't have infinite scale." Well, right. We're so successful, we put the company out of business.

Jeremy: So I think another thing about, sort of where I see this cultural change, is the ability, and you alluded to this, about developers thinking about the actions that they take and how that affects overall costs. And you outlined a developer going back and explaining to somebody, "Okay, well, we had this new initiative. We did this. We had to access KMS more times," or whatever. But I mean, this is something where, you know, how much time should developers be spending on thinking about cost optimization? Because obviously, in a small environment, a small tweak here, a small tweak there, might save you $50 a month, right? But when you get to scale and you have an enterprise serverless application, you might be spending thousands, tens of thousands, $100,000 a month, you know, processing things. Maybe it takes an extra two seconds to do this particular job because you're calling it this way or you're not failing fast enough or whatever. So how much time do you think developers should be spending on cost optimizations? And what kind of experiments could they run to maybe affect the overall bill?

Erik: Yeah, you know, I mean, this is one where there's a lot of different conflicting ideas on this because you ask any product organization, particularly software organization, today and you'll ask them, "What's more important to you: spending all this time at cost, or innovating faster?" And everyone will say innovating faster...

Jeremy: Until the bill comes.

Erik: Until the bill comes, right? And then suddenly it's like, "Whoa, wait a second." And then people tend to think "Well, all right. Well, we'll get around to fixing it when we have a problem." And that was the same situation we got ourselves into with security - application security. It was like, let's get our developers to not care about security right now because it's just going to slow them down. It's a complicated topic. They don't really understand it. Can we just get another team to manage this for us? And then when it's a real problem, they'll come back to us? And that was what the industry tried to do, and guess what was happening? Everybody, even could continue to today, gets hacked, right, left and center. And they realized, "No, no this has to be part of the process of front. It will actually cost us less money." I just heard ⁠— I forget. Moody's just downgraded either ⁠⁠— I don't want to get the name wrong. One of the companies out there got downgraded in the ratings because of their security posture. And I think we need to take a more proactive view to this. Now what made that possible to take a more proactive view for security in the security industry, and it wasn't that the developers, suddenly, we're spending more time necessarily on it. It was that the tooling and the processes, processes in the capabilities of the systems that we're using, all improved to make it possible. I'm a big fan of decision loops, OODA loops, and I think about how can I get cost into that decision loop process so that an engineer could make a quick decision without, an informed or educated buying decision when you're doing things. And that means getting the data to them as quickly as possible. And right now, most engineers get the data at the end of the month in the form of a bill or an angry email or some report that they check even the following day or the following week. And that's kind of ridiculous, right? We live in a real-time world. And if my, you know, I asked this question at a conference recently. I asked everybody. I said, "How long would it be until you knew that your site was down?" Somebody yells out, "It'll be a second. I know it instantaneously." And I'm like, of course, right. "How long would it be until you knew that one of the key transactions, your credit card processing was down on your website?" "[Someone said,] "We'd know in seconds." I'm like, great. "How long would it be until you knew that an engineer on your team wrote a line of buggy code, and it costs $100,000?" Just dead silence.

Jeremy: It would be a while. And unless somebody was checking those bills on a regular basis, you wouldn't see that.

Erik: Well, you wouldn't even see it, even if you, in the moment, I could write a line of code that will cost my company $100,000 in a heartbeat. I can do that as an engineer. That's the power that we have. And so it was dead silence. I think there was a gasp in the back, and somebody finally yelled out, "A month." And I'm like, exactly.

Jeremy: Most likely. Yeah.

Erik: Yeah. You know, there's probably no more critical kind of thing here in terms of doing this. And here's the thing. So first, I think the tooling and the technology has to improve. So that cost can become an operational metric that fits in with the developer lifecycle. And that's the mission that CloudZero's on, obviously, and why we're doing what we're doing. We want to enable that to become part of the engineering team thought process.

Jeremy: So let's get into this discussion about costs being a first-class operational metric. So what do we mean by first-class operational metrics, first of all? In case people don't know.

Erik: Yeah, so what it means is you know, when we're doing design, when we're building, when we do a deploy or we do some tests, we run integration tests ⁠— any type of testing. It is a KPI that we care about when we judge whether or not our application is ready for the world or not. And right now, we care about performance. We care about time availability and things like that. We're not spending enough time thinking about cost as a first class-operational metric. It's important that we look at that and we ask yourselves, "Is that correct or is that wrong?" And we have an idea of what we expect before we release the application to the world. And it becomes part of the KPIs that we track. You know, you walk into an operations center for any major Internet property today and you'll see an operations dashboard that's telling you all kinds of key transactions. And I'm not a huge fan of dashboards. Dashboards are where, I don't know, a lot of things could have died, but at the end of the day, these are the things that people are caring about, and nowhere in any of these dashboards will you see how much money is our current cloud infrastructure costing us? People just simply aren't thinking about that until the end of the day. As companies trying to build innovative products, we're also trying to be, trying to build profitable products. And when I was thinking earlier about you know, when we think about innovation, if I can save you hundreds of thousands of dollars, maybe millions of dollars, because I have taken a little bit more time to think about how I've constructed my application, those are dollars that I can invest back into my engineering process. You know, you speak to any engineering manager about what really helps them move innovation faster, and they'll say, "Head count. More engineers on the team." Even though, sometimes you can't two pizza teams and all that. At the end of the day, stuff is built by people. And if I can add more people to my team or invest more in the technology they're using, then I could move faster. And right now, we take this I think, additional innovation budget that we have, and we almost lazily ship it off to the cloud providers because we think there's no better way to do it.

Jeremy: Yeah, and I think, you know, again, the other thing about first-class metrics, and I guess we're maybe getting a little ahead of ourselves, but that ability to measure a KPI and then just determine whether it's good or bad, and whether that has some sort of impact ⁠— I think the granularity of billing is sort of the perfect KPI to tell you something is right or something is wrong.

Erik: Yeah. I mean, the vision for CloudZero, when I first started working on CloudZero, I was haunted by the fact that the systems we were building were getting way more complicated than what any one person could understand. And my point of view was that we should think of the cloud providers, you know, the cloud as a computer, and the cloud providers as an operating system, and there is nothing that understands really what's going on in that operating system today. We were all too focused on the agents that were telling us what was going on inside of EC2 and you know, in Windows and Linux. And I think that's just the microkernel. We should really not care at all about that. What's going on in this big, complicated system? And so I went looking for every data source of that operating system could provide with the goal of pulling it into CloudZero to build this deep understanding. And the day that I started looking at the billing data was really impactful, because there is no other data source across all of your cloud infrastructure that tells you more about literally everything that's going on. Because if it's happening, Amazon want stability for it. Right? So it is all right there. The challenge with this data source is that it's got great accuracy, but it has horrible latency, and there's no way to correlate this data source with all the actions and activities that folks are taking. When I spin up a new machine, it takes a while before I know the actual cost of it. Or I build out a new system where I have, I set up ⁠— you know that example I gave about KMS earlier. We were working with one customer, and they did a migration from one system to another, and they estimated out what the cost was going to be. And we generated an alert and came back and said, "Wait a sec. You have this cost spike here. Completely unpredicted." And they go, "What the heck is going on?" Well, your team, it looks like here you wrote some code that is calling the KMS API millions of times and they're like, "Oh, that's part of the migration. Jeez, we had no idea." And we're like "Well, what if you just zip up that data into one blob instead of writing it all in these individual components and call the KMS API significantly less?" And they're like, "Well, that would be easy." Instant, instant cost savings, right?

Jeremy: Yeah.

Erik: But the hard part about all that was getting that information to the developers at a time where they actually could even consciously think about the code. Because if I'd come back to that team a month later or a year later, who knows what, and said, "Oh, hey, you know, we found this thing." They'd be like, "What?" I don't even remember what code I wrote two days ago, much less a month ago. Useless advice, right? So that's kind of the last mile for cost optimization as well is getting this data to engineers when they're making the decisions in the moment.

Jeremy: How quickly do you get that data from the billing? You know, maybe not just specifically with CloudZero, but do you have access to that billing data through through an API? Can you get that pretty quickly?

Erik: So quickly, and the answer today really for just about everybody is no, because the fastest that, let's just pick on AWS here for a second, is going to send that data out is maybe every eight to 12 hours. They're gonna drop a big blob of information into an S3 bucket, and you're going to be able to poke around at it, and then maybe they'll decide that they forgot to apply some credits, and so a day later, they'll apply those credits, and then maybe a week later, there'll be some additional modifications that need to be made because they have a specialized arrangement in terms of their negotiated cost and things like that. And then some one time cost will flow into there and all kinds of noise kind of mixes into the thing. And so even if you look at it as quickly as the information's coming out, these eight to 12 hours, it doesn't necessarily tell you the complete story. And so the real magic here is taking that information and combining it with all of the operational activity that's going on in the environment and being able to model out and extrapolate that and use some of these fancy fancy terms like machine learning, and I don't want to be too buzzword-y. I'll fit blockchain in here at some point.

Jeremy: Just in a serverless, machine learning, AI, blockchain system. Yeah, sure.

Erik: The reality is there's data there and all the data sources that your cloud provider gives you, each one individually doesn't tell you the full story. It's what, the answer is, when you combine them all together, you get a much more accurate picture of what's happening. And then you start getting in the realm of where you can talk about costs in a much more relevant timeline than every eight to 12 hours.

Jeremy: Yeah, so then if I was to have a new deployment and maybe I introduced that bug, that $100,000 bug per month, or whatever. So how quickly would I know; how quickly could I find out about that?

Erik: Yeah. I mean, our objective is that you'll find out about it the moment it hits the wire. And that you'll be getting a Slack message or something, saying, "Hey, the change you've just made is going to have serious ramifications." The place that we want to be is just like we know we had an adverse impact on the performance of my application, because we do performance testing after you make a change ⁠— you know how much additional CPU load or response time or or things like that ⁠— what are those things gonna hit, land on? We want to also have that information about how much did the cost profile change in my account? And it may be in test that we only see, like, a couple pennies. But if we're running a production, we'll go, "Yeah, but that couple pennies extrapolated out, we get about 5,000,000,000 transactions or something like that. That's going to result in a 20% decrease in our margin." That becomes a real logical decision. And then when the product team, equipped with that information, goes back to the business and says, "Okay. We've completed this new functionality that you asked us for, but it's going to reduce the profitability of the company by 20%. Should we proceed?" And they might say "Yes." They'll go, "You know what? It's still good enough for us, but we're going to prioritize fixing that in the next release because we want to drive the profitability of the business up." Now they're making really educated decisions that are going be very powerful to building a profitable, nimble business. And they won't be sitting here just kind of grasping at sort of why are we doing cost optimization? Because it feels good? No, because we're trying to build a profitable, profitable business.

Jeremy: So that's a really good segue I think into sort of the last thing I wanted to talk to you about. So sort of around this idea of not just total cost of ownership, right, because we know that ⁠— or we know, I should say we know ⁠— but we assume right and most of the anecdotal data tells us that moving apps to serverless can reduce your total cost of ownership because you're not paying for the operations people. You're not managing those servers anymore. Obviously, the price goes up a little bit, you know, depending on, which service is you're using. But I think it's a really interesting approach to look at what the predicted cost would be and how you could actually fit those into your product roadmap. Like how you would build out your product roadmap thinking about cost as one of the factors, right? So it's no longer just, "Oh, we can add a new feature because we just, we already have 50 or 100 or 1000 servers running. And we could just stick it on those. And maybe the CPU will go up a little bit, but it won't cost us any more," as opposed to saying, "Oh, well, we're adding a face detection feature or sentiment analysis feature, and all of a sudden, now we're hitting up against Amazon recognition or doing sentiment analysis with one of their ML AI services." So that's really interesting to me. Maybe how do you kind of go about using these KPIs to plan for new products?

Erik: Yeah. I mean, the vast majority of things running in the cloud today were systems that were lifted and shifted. And they function and they may cost a little bit much. Corey Quinn had a thing about, he said legacy apps is just an unpopular term for applications that currently generate profit or revenue and that they're everywhere. If we think about the kind of serverless revolution right now, almost all of the serverless activity that I'm seeing is in new application development. Although, there are some notable examples I've seen mainframe applications being moved to serverless. I've seen things that aren't necessarily totally greenfield. I mean, you've probably seen a ton of this, some of this as well.

Jeremyb: Sure.

Erik: But it's still an enormous amount of kind of greenfield development in the serverless space. CloudZero, fortunately, had the decision to make when we started building: are we going to do serverless or are we going to go with containers? We said, "You know what. Let's go with serverless." It was a really good decision. I think we might, because of our own product and the serverless decisions that we made, we might be the only start-up that has a cloud bill that is constantly decreasing, even as we add more customers to the platform. Because every engineering decision, we know the cost of it internally. And it's been really powerful in changing our culture, what we were talking about earlier. But, you know, thinking out into the future, the opportunity to get significant return on engineering investment in improving these legacy applications is going to come from really being able to prove out the value of that re-architecture or rewrite. And today, nobody really has the tools to do it effectively. And so most people are just happy enough to leave well enough alone. But if you have the ability to get it, to analyze that entire operating system, everything that's going on in there, and then come back with a pretty accurate understanding of how that system's working and and identify the parts that could be replaced by a surveillance system, not necessarily all of it, just components of it, so I call this a serverless rightsizing (37:48). I think that the activity that's been built around EC2 rightsizing is, a lot of time, kind of a wasted effort. But think about this notion of serverless rightsizing. Take a legacy system, find the most expensive components of it and then rightsize it onto serverless with a cost-justification or cost-benefit analysis that is done for you in an automated way. That gives you real justification for why you do that. And with serverless, we've seen it. You might see 100X return on that investment. The challenges, of course, is nobody really knows where to start, and they don't have that data in-hand to justify that investment. And so they, you know, they don't have the opportunity to do that. But when the business has that data, they can make that data-driven decision about why they might replace this component of system. That's where I think we're going to see the real serverless revolution take off because that cost-benefit analysis is going to really drive the rationale behind why people are going serverless.

Jeremy: So alright, I'm gonna ask you one more question. This is for any developer who is out there listening who has ever had to talk to the powers that be about getting a budget. So traditional operations, I guess, would kind of fall under the CPAEX where you buy a certain amount of servers, or, you know, it's not really that OPEX thing. It's not that idea that you're being metered, right? Have you seen this argument from purchasing departments? And how do you overcome that? How do you convince the purchasing department that, yes, this metered-building that we don't necessarily know how much it's going to cost is a better way to do it than for us to just have a budget that caps us that something?

Erik: I mean, this is, in a lot of ways, this is actually I don't know if you knew to tie this back or not, if I've told you this story, but this kind of ties back to my origin story for getting Veracode onto the cloud. We had a really interesting project for a client, and we weren't in AWS yet. And we wanted to build this out in AWS because we needed the scale. We were going to need thousands of servers to do it. But I needed to convince the CFO that it was a good idea. So I went to our CFO and I said, "Hey, can we build this project for this client. We think it's going to be really amazing." He said okay, because I needed the company credit card and so you could imagine me having this conversation. He's like, "Alright, so I've only heard nightmare stories about this cloud thing. Here's the company credit card. Your budget is $3000." And about a week later, $2997. We figured out how to do it. About 1500 spot instances with like a heavily optimized, kind of homegrown, autoscaling solution. And it was a wonderful challenge, and we had a lot of fun doing it, and we built a really cost effective product, and that was kind of my origin story in terms of thinking about cost as part of engineering, because I think constraints as engineers is a really powerful thing. It helps kind of bound the problem that we're trying to solve, and so I would suggest to any engineer or any anyone out there listening to this, when they think about it, is try to challenge yourself with a budget, and try to manage towards that budget. Become an educated consumer. When you go to the restaurant, I mean, we all do this naturally. When you go to a restaurant or order things off the menu, we see price tags there. Today, when we order off the menu for whichever cloud provider were using, we don't actually see the price tags there. It's actually kind of, I think it's strange that when you've spin up infrastructure using the console of any cloud provider, it doesn't like pop back a message thta just says "Oh, and what you just did is going cost, you know, x dollars an hour, right?" Why is that missing? Well, it's probably because it's not in their best interest.

Jeremy: Maybe they don't want to tell you.

Erik: They don't want you to know that because yeah, you're going to go to the restaurant and you're going to go like, "Oh, somebody else is paying the bill. I'll have the surf and turf. The wagyu beef. That sounds delicious. Everything has been prepared for me." But so I would really strongly suggest to folks, try to manage themselves to a budget, even if no one else is holding them to a budget. Because even if they think no one's holding them to a budget, somebody, somewhere, is holding somebody to a budget. To think about that and and actually try to be irrational, because most people would say a $3000 budget building anything in the cloud is nuts, but it can be done. And you can really do some very powerful things. And what you will discover when you hold yourself to that budget is this tight relationship between a well-architected system and a cost-effective one. And it will make you a better engineer through that process.

Jeremy: Awesome. All right, well, let's leave it there. So listen, I want to thank you, Eric, for joining me, and sharing all of your knowledge. Where can people find out more about you and CloudZero?

Erik: So they can certainly find us on CloudZero.com, and and read all about us. We've got a great blog there. I hope everybody spends some time providing some feedback. And the usual sources on Twitter and whatnot, which I'm sure you'll put links in. But, you know, true story, my Twitter handle, which is impossible to say, was my original DND character from when I was in high school, so that's the hidden fact that your listeners will, that I've now shared with the world.

Jeremy: Silvexis.

Erik: Silvexis. S-i-l-v-e-x-i-s.

Jeremy: And then @CloudZeroInc. And then also, if people want to email you. Can they do that?

Erik: Yeah. You know they can. They can reach out to me. It's really simple. It's erik@cloudzero.com.

Jeremy: Perfect. Alright, well, I will make sure we get all of that in the show notes and thanks again, Eric. Appreciate it.

Erik: Wonderful. Thanks, Jeremy. Had a great time

View Details

About Mike Deck

Mike Deck is focused on building a community of partners around the AWS Serverless Platform. He spent the first half of his career as a full-stack software developer building enterprise applications and coaching teams on agile delivery methodologies. For the last four years he’s been working as a solutions architect for AWS on the partner team helping both ISVs and consulting partners accelerate their pace of innovation using the cloud.

  • Twitter: @mikedeck
  • EventBridge Product Page: https://aws.amazon.com/eventbridge/
  • EventBridge Documentation: https://docs.aws.amazon.com/eventbridge/index.html
  • EventBridge Partner Information: https://aws.amazon.com/eventbridge/partners/

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly, and you're listening to Serverless Chats. This week, I'm chatting with Mike Deck from AWS. Hi, Mike. Thanks for joining me.

Mike: Hey, Jeremy, thanks a lot for having me.

Jeremy: So you're a Solutions Architect at AWS, and I'm pretty sure most people are probably familiar with what Amazon Web Services does. But why don't you tell the listeners a little bit about yourself and maybe what a Solution Architect does at AWS?

Mike: Yeah, sure thing. So I'm actually on our Partner team, so I work as a Partner Solutions Architect, which means that I work with both our ISV and consulting partners to help them with kind of any technical questions they have, work with our ISV partners on their product roadmaps and how they're integrating with our services, helping them with architectural questions, and things like that. Been at AWS for about four years at this point. Been on the Partner team the whole time, and most recently, I've been kind of specializing in the serverless space. So working with a lot of our great ISV and consulting partners that are doing things around Lambda and API Gateway, as well as with the new service that we just recently launched.

Jeremy: Cool. All right, so speaking of services recently launched at AWS Summit New York, AWS launched this new product called EventBridge, which is sort of this, and you can correct me if I'm wrong, but sort of this cool extension to CloudWatch Events and since you're on the Partner Team or you're the Partner Team Solutions Architect for Partner Integrations with EventBridge, you obviously know probably more about this than anybody else. So why don't you tell us a little bit about EventBridge and sort of what it does?

Mike: Yeah, absolutely. So I think you know, it's definitely accurate to compare it to CloudWatch Events. So really kind of the genesis of this service was that, you know, we saw customers building more and more with this kind of event-driven model, and CloudWatch Events is really a fantastic tool for for doing these kinds of things. I think Forrest Brazeal had a blog post a little while ago about using CloudWatch Events to do awesome event-driven things. We see a lot of customers that are interested in this space. The event for pipelines projects came out, got a lot of traction, and so we realized that, you know, there's really this kind of need to build additional services that make it easier to build these kind of things. So we took the existing CloudWatch Events infrastructure and APIs and kind of extended them to add some additional features around integrating with SaaS providers to create more native event sources that you can use within your AWS applications, and then, yeah, extended those APIs to make it easier for customers to do things like creating custom event buses and patching to SaaS event sources, etc.

Jeremy: So what are some of those use cases then that customers might build with this?

Mike: Yeah, so, you know, I think that one of the obvious ones is: Hey, I want to trigger a Lambda function every time someone creates a new ticket in my CRM, for instance. Right? I want to go and kick off some sort of workflow. Or maybe I'm going to start a step functions or do something like that. Um, we also see a lot of customers interested in doing kind of audit-and-analytics type workloads. So I just want to ingest kind of the full event stream of all of the things that have changed out, you know, maybe in my identity management tool that I'm using. So every time a failed login happens, or every time a new user is registered, I just want to ingest that, throw it into a Kinesis Firehose Data Stream, and put it out into an S3 data lake so I can go and create with Athena or something like that. And then, obviously, you know, ML and and doing kind of AI inference and things like that on all of these various data streams is becoming super popular as well, so this gives you a great opportunity to, yo u know, every time a new email is opened in your kind of customer engagement platform, you can ingest that and add it to your modeler or do some sort of inference on it in order to drive additional kind of business decisions.

Jeremy: Yes. So those are some really cool sort of things that you can do with it. And I remember Forrest's article about sort of using CloudWatch Events and using custom events as basically almost like an SNS topic, in a sense. So maybe let's get into the nuts and bolts of EventBridge, kind of how it works. So with existing CloudWatch Events, you can use a, put events API or put custom events API or something like that, where, and then Forrest explains that in his article where you can send an event and then you can subscribe consumers to it. But why this extension? What's this idea of having separate event buses?

Mike: Sure. So, um, yes. So I guess there's a few different reasons that you might want to have different event buses. So like you mentioned with CloudWatch Events, there's really just sort of the default event bus in each region of your account. We publish all of these sort of native AWS events to that default event bus. And so those just kind of appear in your account automatically without your having to do anything. You know, the ability to create additional custom event buses now lets you kind of segregate these different application domains or different context, I guess, for the various different types of events that you need to ingest across your different surfaces and applications. It also gives you a good way to create separate channels for different SaaS applications that you may have that you're trying to build event-driven architectures on the back of. So, yeah, we can we can talk a little bit more about kind of the specifics of how it all works, but basically, every individual SaaS provider that you're integrated with is going to have its own event bus so that you can write rules against that, and keep all of those different events separated from the different sources.

Jeremy: All right, so I definitely want to get into the SaaS side of things because I think that's probably one of the coolest features of this EventBridge. But let's go back, maybe, and start at a lower level, talk about events, sources and targets. So what events were you, because you mentioned that you can get all of these events that are happening within your AWS account. So anytime a new account's created or new resource is created or something changes, there's these events that are sent to that default event bus. But now with this separate event bus, how would we subscribe sources to that or set targets to trigger Lambda function or other other services?

Mike: Right, right. Yeah. So EventBridge works with this concept of rules. So every event bus has a set of rules associated with it, which allow you to essentially select the events that are interesting. And then from each one of those rules, you can you can create multiple targets that it's going to forward each of the events that match that rule down to. So, for instance, you may decide I want a rule that's going to match all of the sort of ticket created events from ZenDesk. And so you would build a rule that has in it, you know, the detail type equals ticket created and then you'd associate, you know, potentially multiple targets for that. Maybe I want to send all of the new ticket records down to a Kinesis stream so they can go and be processed by some other kind of downstream analytics process, and then I also want a trigger, you know, a step functions workflow. So each of those would be targets and you would associate those two targets with that one rule. And now, any time that ZenDesk publishes a new event to your account that is of type ticket created, that's going to match that rule and then send that downstream to those specific targets. So you can have multiple rules and multiple targets per rule in order to kind of build these sophisticated event-routing models.

Jeremy: So you keep mentioning like SaaS providers, like ZenDesk and stuff, and I totally want to get into that. But before we do that, so this isn't not just for SaaS providers, though, right? I mean, you can create your own custom events. It's sort of, and again this may be in comparison to Forrest's article, but you can sort of use it as SNS almost or like an SNS pub-sub type thing. So maybe explain, you know, what's the difference between this and something like CloudWatch Events or SNS or even Kinesis?

Mike: Yeah, I think, uh, So I think SNS is probably the most similar sort of service, if you want to think about it that way, in terms of SNS gives you the ability to publish messages and then fan those out to multiple subscribers. SNS has the concept of a subscription policy that allows you to kind of filter messages per subscription. Again, you can get similar kind of features to the way that rules work within EventBridge. You know, the big difference on the SNS side versus EventBridge is for custom events anyway, the downstream targets that you have accessible within SNS is more limited than what is available in EventBridge. EventBridge has 17 different AWS services that you can integrate with natively so you don't have to sort of pass through a Lambda function necessarily, if you just want to go and yeah, drop something on a Kinesis Sream or FireHose or kick off a step function, etc. So I think that's one of the big key differences there is the sort of richness of the kind of targets as well as the source piece that we'll talk about here a little bit. I think where SNS really shines is when you're in these super high throughput or really massive fan-out. So if you've got thousands or millions of subscriptions that you want to have for a single topic, SNS is definitely the way to go. Similarly, if you're really trying to push, you know, millions of TPS or something like that through a particular topic, SNS is a better option for when you've got those really kind of massive, high throughput workloads. So those are kind of the key call-outs. And then, yeah, talking about Kinesis a little bit. So Kinesis gives you more of a streaming model, so everyone that's going to consume that stream is going to see every single message on that stream. You know, you're somewhat limited in total number of people that can consume a single stream. And each individual consumer would be sort of responsible for kind of filtering out any messages that they weren't potentially interested in.

Jeremy: And so with SNS, so I think that's that's really good point you make about the services that can be triggered off of EventBridge because with SNS again, you can't trigger - you can't start a step function. You have to write a Lambda function that then calls that step function. So there's just extra sort of processing in there. And then with Kinesis, that's sort of interesting, because this is probably one of, was my biggest question when I first found out about this product, was this idea of sort of event buffering, right? So we typically, especially we have downstream services that might not scale as much as our Lambda functions would, we would put a queue or Kinesis stream or some way to buffer those events. There's no event buffering yet, right in EventBridge? You would still use something like SQS if you did need to buffer events.

Mike: Exactly. That's a great point. So certainly using those types of services, SQS, Kinesis, SNS, like all of the standard sort of messaging services that you typically think of often times, I think, get used in conjunction with, you know, previously CloudWatch Events, and now EventBridge. So certainly having a rule whose target is an SQS queue so that you can buffer all of those events and then have you know, each individual consumer have their own queue that they can work off of makes a ton of sense. So using SQS as that kind of simple queue that gives you gives you that message durability, and all of the delivery and that dead letter queue semantics. But then using a EventBridge to really manage kind of the message routing and filtering pieces, which is where really it kind of excels. So super common pattern there, definitely.

Jeremy: All right, super important question here. How is this priced?

Mike: Yeah, so pricing is exactly the same as the existing CloudWatch Events custom event pricing. So dollar per 1,000,000 events that get published into your event buses basically. So if you're publishing custom events yourself then it's exactly the same as if you're publishing to the default CloudWatch Events bus. Yeah, a dollar per 1,000,000 there. And then similarly, if you've got this hooked up to SaaS event sources, as long as you have an event bus hooked up to that event source, you're charged for sort of each event that the SaaS partner publishes to your account.

Jeremy: Okay, cool. All right. So let's talk about SaaS partners, because I think this is one of those things that I definitely don't want to get lost in this announcement. And there's so many products that AWS comes out with on a regular basis that sometimes it's hard to keep up. But I think the coolest innovation with EventBridge is almost this idea of completely getting rid of this concept of webhooks, right? So now, with the ability for partners to put events into an event bus directly, and we can talk about how the authorization of the stuff works, but this is essentially what it's doing, right? You're getting rid of webhooks?

Mike: Exactly. Yeah. So instead of, you know, going out to pick your SaaS partner of choice, you know, I go out to Datadog or whatever, and I go and configure a webhook. And when I do that, I have to give him a URL and maybe set up some sort of like secret token, etc. Now, instead of that, I'm basically going out to that same partner and just saying, "hey, I want you to send it to my AWS account. Here's my account ID. And here's the region that I want you to to send these events to." And then we kind of take care of everything else there for you. So now you don't have to go and stand up an end point. You don't have to stand up a Tomcat server in your VPC, or ideally, obviously, I guess we're talking on a serverless podcast, so we'd be using API gateway and Lambda, right?

Jeremy: We would, we would. Yes.

Mike: But still, that's not a trivial thing to do necessarily. I mean, I shouldn't say trivial. It's easy to do, but they're still, you know, that's still another thing for you to manage, and another thing for you to kind of another hoop to jump through, as opposed to, you know...

Jeremy: And there's cost involved, right? The typical way that I would set up a webhook is most likely API gateway. Actually, if I was doing a high volume webhook, I would probably use ALB at this point just because it would be a lot cheaper. But I would do that, probably hit a Lambda function and then right into an SNS - sorry, SQS queue - in order to go into a database or something. Or maybe I do a direct service integration with API gateway. But with this, you don't have to set up any of that infrastructure, you're just paying for — is it just reads off the stream or writes to the stream as well?

Mike: It just writes to the event bus, so it doesn't matter how wide you fan out on the back of the event bus. You're only paying for each message that's sort of getting published to it. So yeah, I think that's a great point. And honestly, I mean, I think just the sort of the reduction in developer friction too is actually a huge part of it. So I mean, how quick can you get something stood up where, hey, I just wanna, you know, trigger a Lambda function every time a new object comes into S3. We kind of want that same experience no matter what that event source is, whether it's something from inside of AWS or outside of AWS, we want you to be able to just super quickly say, "hey, yeah, there's this event source. Let me attach my Lambda function. Let me attach my step function state machine" or whatever the case may be, and you're off and running.

Jeremy: So how does the customer go and actually build an integration with, you know, partner XYZ?

Mike: Right. So yeah, for partners that have support for EventBridge, basically, it's a pretty simple process where you go out to the partner's portal or, you know, developer portal or console or whatever they've got, and provide them with your AWS account ID. As part of doing that, the partner is then going to go and create what's called a "partner event source" inside of your account with that account ID. The nice thing about this is they don't have to, you don't have to give them a cross-account IAM role. You don't have to kind of mess with any of that type of permissioning. They'll essentially create that. Then when you go back to the EventBridge console, you'll see a list of all of these new events sources that are available. So, you know, the SaaS partner that you went to will pop up in that list. You can check the box and say, "I want to associate this with an event bus" and then the event bus is there. It's ready to go. You can start adding rules and attaching targets, and start building from there.

Jeremy: Alright, so what about infrastructure as code? Can we set up these rules with CloudFormation at this point?

Mike: Yeah, so CloudWatch Events has CloudFormation support today. You can create rules and targets there on EventBridge for custom event buses. That will be coming soon. It's not available right this minute, but we're definitely planning on adding that.

Jeremy: Alright, cool. Alright, so now, and actually, let me ask this question first: retried behavior, right? So you have 17 different sources that this thing can trigger, or 17 different downstream targets that this can trigger, things like the two retries for Lambda. That's still applies?

Mike: Right, so, yeah, if you think about it, really what EventBridge is going to do is it's going to, for every target that you've got configured on a particular rule, it's gonna make sure that the event gets delivered to that target successfully. Now, in the case of Lambda, what successful delivery means from event bridges perspective is that it was able to asynchronously invoke your function. So when it gets a success back from the Lambda service saying that, "hey, yeah, the invoke call you made with successful." You know, 200, good job. EventBridge considers that event delivered, and it's done. So then, at that point, you're really kind of relying on the standard Lambda retry policy within that kind of async event - I'm sorry - async invoke flow. And so if you've got a dead letter queue configured on your Lambda function, etc., it's all going to work exactly the same as you expect.

Jeremy: And what are the retry policies on SQS queues or Kinesis or step functions, things like that?

Mike: Yep. Basically we'll retry for 24 hours to make sure that the event gets delivered to whatever target you've got configured. You know, again, depending on the service, that ultimately just means that we were able to successfully kind of hand it over there. So you know, when you think about like an SQS queue, that's ultimately just, hey, we're gonna go and make sure that the SQS service now has successfully accepted the message that we send to it. And then obviously the retries and all of the downstream processing is going to be up to you and how you're pulling SQS, etc.

Jeremy: Sure. So we can assume, though with the distributed nature of this, that it's an at-least-once-delivery type model, as opposed to exactly once.

Mike: Correct. Yeah. So this is an at-least-once-delivery.

Jeremy: So make sure you designed for item potency and things like that.

Mike: Right. Yeah, exactly.

Jeremy: And then just I guess one question on that, that I'm thinking of now is so after the 24 hours, if for some reason a service doesn't accept the event, is there some sort of concept of a dead letter queue for events yet or?

Mike: No, no not within EventBridge. So right now, yeah, that was basically, we're going to try for 24 hours, and if we can't do it then consider that to be a failure and that, yeah, there's not right now a good way to kind of react to that. I think in the vast majority of cases, you know, 24 hours is plenty of time for...

Jeremy: Yeah, I would think so.

Mike: ... the services to recover because again, the downstream services that you're integrating with, you know, you're not using just like a standard HTTP endpoint, like the way that SNS would. So it's kind of native AWS services that you're relying on there.

Jeremy: Perfect. All right, so let's talk about partners for a second, because again, this is something that's really interesting. I think, you know, every SaaS provider out there that does webhooks now has some way to configure webhooks. So this is obviously something that they're gonna need to build themselves in order to integrate with AWS, which again, I think the demand will be there. So it'll likely be a smart move on their side to do this. But maybe we can talk a bit about how a partner would go about building one of these integrations.

Mike: Sure, Absolutely. So, yeah, like I mentioned kind of before, you know, the standard customer flow is they're going start at the partners site. So if you go to the EventBridge console, you'll see a list of all the partners that are integrated, and so you can kind of get a quick link back over to the partner side in order to go and start the flow. But really, things don't start until you're over on the partner side so if your partner and you want to build this integration, basically, you just need to have kind of a form inside of your developer console that that allows a customer to specify: hey, this is my AWS account, and this is the region that I want you to send events to. And then when they sort of submit that to you, you make one API call, that's create partner event source, where you basically pass in a name of the event source and the account ID that your customer provided you. So you're essentially creating this new handle that you have access to publish events to, and that is kind of that bridge between your account as the partner and the customer's account without having to do this kind of IAM cross-account role assumption dance, etc., right? So once you've created that event source, the customer then can go back into the AWS console. They can associate that event source with an event bus that they want. But as soon as you've created that that event source as the partner, once that source is there, you can start sending events to it immediately, using an API call putPartnerEvents. So it's basically the same thing as putting a custom event to the default bus. You just get to specify this kind of partner event source instead of a specific event bust in your own account. (22:31)

Jeremy: And a new event bus is created for each partner. And actually a partner could create multiple event buses. Correct?

Mike: Correct. That's a good point. So, a partner, so any individual partner could create multiple event sources that are associated with your accounts. If you wanted to have separate events sources, you know, kind of a common ah example of this would be, you know let's say it's an HR system that's publishing, you know, multiple kinds of events. Maybe some events are about really kind of sensitive data, so salary information about individual employees, and then you've got other events that are about time off requests or something like that are not nearly a sensitive, those could be directed to different events sources, and then each event - each partner event source that gets created in your account will have its own dedicated event bus that you create.

Jeremy: Okay.

Mike: And so then, by doing that now, you can create, yeah, specific kind of security policies on each one of those event buses. So I can have, you know, my sort of sensitive events stream that's locked down. And I can have the other one that maybe is more open and enable for other people to go and write rules against more freely.

Jeremy: Alright, so that's a really cool feature. So I do like that sort of almost it makes it super easy for the partner to integrate where they just have to make that one API call, and then you handle all of that sort of authentication and authorization on your side that can be done back with, you know, done by the customer requesting that stuff. Alright, so that's really cool. So what's the process though of becoming a partner? Because it sounds like you have to sign up for this?

Mike: Right, So right now today, it's a bit of a kind of custom request process. So we absolutely want to onboard more partners and there is going to be a process for kind of going through that, and that's documented our website. I'll make sure that you've got a link for where people can go if they want to sign up to become a partner. But overall, it's kind of a matter of saying, "hey, I want to do this." You know, this is my domain name that we'll use to kind of identify your integration. You'll go out and build things. We'll do a quick validation with you and then kind of finalize onboarding at that point. Obviously, we kind of want to continue to move this forward in automating that process, making it more self-service. But yeah, as of today, if you're interested, definitely, yeah. Go check out the link that we'll share. And you can get started that way.

Jeremy: Alright, well, I'm going to make a public service announcement. I'm going to put every SaaS company on notice. I'll give you, like, three months. If you're not integrated with this, I'm going to stop using you and use somebody else who does. Because I just see this saving a ton of money, making the developer experience so much easier. Yeah, I just I love this. So anything else on the partner side? You know, I'm sure there'll be a bunch of documentation or there probably is a bunch of documentation about how to do all this stuff. Seems pretty easy.

Mike: Yeah. I mean, I guess the one other thing to note about that, it's just, you know, when you think about webhooks, I guess, and kind of all that goes into that, everything from, you know, doing the security, doing retries, kind of managing that yourself, tends to actually get pretty complicated if you want to do it right. I think it's really easy to throw a webhook out there, but, you know, if you're just using a simple auth-token in a header somewhere, that's a pre-shared key, you know that there's certain vulnerabilities there. Also I'm really doing the appropriate, you know, retry semantics and everything. It's way easier to sort of offload that onto our plate. You know, that's one of those classic, undifferentiated heavy-lifting things that we love to solve at scale.

Jeremy: Yeah, absolutely.

Mike: Yeah, anyway would love to work with official partners that want to.

Jeremy: So serverless webhooks. I think that, because honestly, I mean, even if you built serverless webhooks in the past, with all these other serverless features just consolidating that all into one simple call from the from the API provider or from the SaaS provider. And just the other thing I love about that too is there are so many instances where you lose events with webhooks and things like that because whether your infrastructure, something goes on with the infrastructure, or whatever, this is just one of those things where the partner, I'm assuming, will be able to build in their own retries and things like that to make sure that they get delivered. I think you just, you know, again, not that webhooks aren't scalable. But there's just, to me, this seems like just an awesome innovation. So totally, you know, congrats to you guys for coming up with this or for implementing it, at least because it is very, very cool. Alright, so maybe we can just switch to, let's talk less about EventBridge in general, and maybe just kind of talk about event-driven architectures. So maybe for people who aren't familiar with this, because we might have been talking over people's heads, with events passing around and triggering Lambdas and doing all this kind of stuff, maybe you could explain in a couple of sentences what exactly is event-driven architecture and kind of how it fits into serverless too might be good.

Mike: Yeah, absolutely. I mean, I think that it's probably easiest to understand it when contrasted against kind of a command-driven architecture, which I think is what we're mostly sort of used to. So this idea that I've got some set of APIs that I go out and call and I kind of issue commands there, right? So I maybe have an order service. I'm calling create order or I've got downstream from that. There's some invoicing service now, and so the order service goes out and calls that and says, "Create the invoice, please." So that's kind of the standard command-oriented model that you typically see with API-driven architectures. An event-driven architecture is kind of, instead of creating specific, directed commands, you're simply publishing these events that talk about facts that have happened, you know these are signals that state has changed within the application. So the order service may publish an event that says, "hey, an order was created." And now it's up to the other downstream services to, they can observe that event and then do the piece of the process that they're responsible for at that point. So it's kind of a subtle difference, but it's really powerful once you really start kind of taking this further down the road in terms of the ability to decouple your services from one another, right? So when you've got a lot of services that need to interact with a number of other ones, you end up kind of with a lot of knowledge about all of those downstream services getting consolidated into each one of your other kind of microservices, and that leads to more coupling; it makes it more brittle. There's more friction as you're trying to change those things, so that's a huge kind of benefit that you get from moving to this event-driven kind of architecture. And then in terms of kind of the relationship to serverless, obviously with services like AWS Lambda, you know, that is a fundamentally event-driven service. It's about being able to run code in response to events. So when you move to more of this model of hey, I'm just going to kind of publish information about what happened, then it's super easy to now add on additional kind of custom business logic with Lambda functions that can subscribe to those various different events and kind of provide you with this ability to build serverless applications really easily.

Jeremy: Yeah, I like how you, you know, and I've heard like Danilo has said this in the past, which is great, you know, calling them facts, like these events that come through are these —they don't change, right? That's just something happened and you know about it, right? So rather than updating a record in a database that gives you the current state of something, this is that sort of that ledger, right, that immutable ledger that, if you think about it, that has all of that information attached to it. And the other thing, too, that you know, I think this is what trips a lot of developers up, and it's funny, too, because it's hard once you're in it, and once you understand it, it seems very logical. But I think if you take a step back and you look at it through someone who's not familiar with this, the idea of asynchronous invocations is something that I think just, kind of again, it trips people up because we're very used to this request-response model. But the asynchronous nature of a event-driven applications is you just kind of put something out there into the ether and something else picks it up and does something with that. So maybe, a question for you, I guess, is what are some of the patterns or best practices for building event-driven applications? And I don't know if we can cover that in a few minutes, but it's maybe the top line ones. And maybe, actually, how it would apply to EventBridge?

Mike: Sure, yeah. I mean, I think one thing that we already talked about a little bit is this idea of having, you know, kind of an event store, if you will. So this would be something like SQS that that provides this sort of durability of the events that are getting consumed by by an individual consumer. So, like you said, you know, there's this potential for me to create this event. I throw it out in the ether, and then what happens if one of these downstream services, that's really important that it gets this event, happens to be down when that event is produced. So using things like SQS to kind of provide that durability of those events so they can be picked up when that service comes back. I think it is a huge, you know, that's a really important practice to make sure you're thinking about sort of what is the durability needs of the events that I'm producing on. And then a nice side effect of that is now you get this sort of, you know, much better sort of availability characteristics of your system overall, because now, even if I've got an upstream service that's relying on sort of downstream behavior, even if that downstream service happens to have, you know, an outage or is not responding very quickly, maybe it's just in a reduced capacity state, I can continue sort of doing my job. It can continue responding to its clients and customers and continue to operate normally, and then whenever those downstream services come back, they can handle it. The sort of converse piece of that is, of course you need to think a little bit about sort of this eventually consistent data model that you're going end up with. So you're going to have, you know, each one of your services probably has its own data stores, because events are now propagating asynchronously through your system, each of those data stores may have kind of a different version of the world at any given time. And so just kind of being aware of that, understanding it, like you mentioned kind of having this ledger of events and using that sort of event-sourcing pattern to keep track of the state of this system, ends up being really powerful in those scenarios, because now I have this ability to kind of, you know, manage this state of the world as I understand it in a point in time, and then also the ability easily to sort of roll back or understand what the state of the world was at some kind of past point time. Anyway, again, yeah, that's kind of the nutshell description of a couple of interesting practices. Obviously, yeah, we could talk all day about that.

Jeremy: Yeah. I think we maybe opened a huge can of worms. I think the only other thing I would say too just in the context of EventBridge. So I have always, I wrote a post called Serverless Microservice Best Practices for AWS and one of those microservices I define is SNS as being its own sort of microservice in and of itself. Because you could use that to bridge, you know, between multiple services. So if you wanted to use it to choreograph events between multiple services as opposed to using step functions to do orchestrations or something like that, it's a very good way to do that. And with some cross-account capabilities and things, it makes it very interesting. I think EventBridge actually solves that problem and does it better, right? Because now you don't have to go in and specifically create some SNS topic that maybe lives outside of all these other services, which is sort of a weird thing where I often find myself publishing a single SNS topic as its own microservice, as its own thing, and using that with multiple microservices to communicate with one another. So I like how this is just kind of there, and you don't have to worry about it, and like, "oh, did I create this?" or "didn't I create that?" And I'm assuming, you know, you can create as many of these custom event buses as you want. So if you did want to follow that serverless sort of stages mentality, where you have your dev stage, you have your production stage or staging stage, and maybe if those were in separate accounts, actually, those event buses would be in separate accounts anyways, but I think you could do something similar to that if you wanted to.

Mike: Yeah, absolutely. So just to be clear, there is a limit of I think 100 event buses for account by default. So, yeah, it's not necessarily completely unlimited, but you can definitely create a number of these custom event buses, and use that just as you were describing.

Jeremy: Well, now with Control Tower, we could just create many accounts as we want to.

Mike: Yeah, I mean, you kind of a joke, but I think that's actually definitely the sort of the best practice that's going to continue, as account creation becomes easier and easier, definitely having that kind of segregation is huge.

Jeremy: Totally agree. Alright, well, listen like you said, we could talk all day, But why don't we wrap this up so again, thank you so much for joining me, and obviously, just telling us about EventBridge because I think this is a really, really cool innovation. So why don't you tell the listeners how they can maybe find out more about you, and actually, probably more importantly, no offense to you, but how they can find out about EventBridge?

Mike: Yeah, of course. So, I'm on Twitter @mikedeck. Not super super active on there, but I'm more than happy to respond to anyone whose got questions about this topic or others. And then certainly going to the kind of standard EventBridge product page, so aws.amazon.com/eventbridge would be the best kind of place to start off. And then you could certainly just go to the consoles. Well, if you just want to go jump in and start building, that's kind of how I always get started with new AWS services myself, so I would encourage you guys to all go and check that out. And, yeah, I love to hear your feedback, hear about what you're building.

Jeremy: Awesome. Alright, well, we'll get all that into the show notes. Thanks again, Mike.

Mike: Thanks a lot, Jeremy.

View Details

About Chase Douglas

Chase Douglas is the co-founder and CTO of Stackery.io, the leading serverless acceleration software solution. His experience spans the gamut of technical and managerial, specifically focused on how teams of developers build products collaboratively. In prior roles he has been a VP of engineering at a web application security company, technical architect of the New Relic Browser product, and an architect of the multitouch implementation for the Linux desktop.

  • Twitter: @txase
  • Email: chase@stackery.io
  • Stackery: stackery.io
  • Stackery Blog: stackery.io/blog
  • Stackery Changelog: https://docs.stackery.io/en/changelog/
  • Live Stream Series: https://app.livestorm.co/stackery/
  • Portland Serverless Meetup: https://www.meetup.com/Portland-Serverless-Architecture-Meetup/

Transcript

Jeremy: Hi, everybody. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Chase Douglas. Hi, Chase. Thanks for joining me.

Chase: Hey, glad to be here.

Jeremy: So you are the CTO at Stackery, which is in Portland. Why don't you tell the listeners a little about yourself and what Stackery does?

Chase: Yeah. So I'm the CTO and co-founder of Stackery and I've spent my career figuring out how to manage complex systems. And as serverless as an architectural pattern started to take off I was really interested in finding how do we help people adopt that architectural pattern more easily? Stackery is a product, a tool set that makes it easier for anyone from individual developers on up to teams and organizations, but especially at larger sizes, manage to design, manage environments, deploy and at the other side, help monitor their serverless applications.

Jeremy: I wanted to have you on the podcast today because I want to talk about serverless development workflows. And I think when people start moving into the serverless paradigm, we gotta sort of change the way that we think about developing applications and that you know everything from whether you're developing locally or developing remotely, or whether you're trying to do something like offline emulation or trying to do the remote testing things like that. Just what are some of your thoughts on sort of this idea of these serverless development workflows? How does it sort of change things?

Chase: Yeah, serverless itself is a different way of building, developing and including testing applications. And one of the things that we have to step back and recognize is that at the end of the day, we're still developing software, we're still testing software, but we need to find the right ways to be efficient at how we do those. It's slightly different in a serverless world, and so we once we find the right patterns. And once we start to use those as an individual or in the team, things actually speed up once again. So there is an interesting play here. Uh, but it's all about just finding the right mix and match of how to do the things we're familiar with when it comes to the development and testing.

Jeremy: Yes, so that makes a ton of sense. So what I think I'd like to do is sort of dive down into a number of these different topics, you know, break it down a little bit and get into I mean, because again, you're an expert. Stackery obviously, is all about building out these workflows or helping developers build these workflows. So I want to get into these these details here and let's start with sort of just really maybe 30,000 foot view. How has cloud sort of changed the way that we develop software?

Chase: Yeah. So the way that we've always developed software up until very recently was it would, in the end, be running on servers, whether it's in a data center or in the cloud. But these servers were monolithic, compute resource. That meant that typical architectures might be a LAMP style stack. You've got a Linux server, and you've got a MySQL database off to the side somewhere, maybe on the same machine, maybe on a different machine. But mostly as a developer, you're focused on that one server, and that means that you can run that same application on your laptop. So were we become very comfortable. We built up tooling around the idea of being able to run an entire application on our laptop, on our desktop in the past, that faithfully replicated what happens when that gets shipped into production in a data center or in the cloud. With serverless, everything is kind of a little works differently. You don't have a monolithic architecture with a single server somewhere or a cluster of servers, all running the same application code. You start to break everything down into architectural components. So you have an API proxy layer. You have a compute layer that oftentimes is made up of Lambda, though it can include other things like AWS Fargate, which is a docker-based, serverless, in the sense that you don't manage the underlying servers approach. So you've got some compute resource, if you need to do queuing instead of spinning up your own cluster of Kafka machines, you might take something off the shelf, whether it's SQS from AWS or their own Kafka service or Kinesis streams. There's a whole host of services that are available to be used off the shelf. And so your style of building applications is around how to piece those pieces together rather than figuring out how to put those and merge those all into a single monolithic application.

Jeremy: So how then do developers need to think differently? I mean, again, I'm super familiar with the LAMP stack. That was probably where I started. Well, I started with Perl and static text files, but we won't talk about that. But as we got a little bit more advanced and we started using things like the LAMP stack, obviously, it was very easy for us to just either test it locally or to even building in the cloud. It was, or not the cloud on our hosting provider, but we could just easily upload a new file and things would magically work for us. But as you mentioned, things get more distributed. Right? Once we go into this cloud environment, you've got multiple services working together. You don't own those services necessarily. If something breaks with those service is you kind of have to deal with that. So maybe what are some of the limitations that a developer might have to deal with when they're starting to move, you know, their production workloads to the cloud?

Chase: Yeah, for all the benefits you get from serverless, with its auto scaling and its capabilities of scaling down to zero, which reduces developer cost, you do have some things that you have to manage that are a little different than before. One of the key things is, if I've got, like, a compute resource like a Lambda function in the cloud that has a set of permissions that it's granted and it has some mechanism for locating, the external service is like SQS queue or an SNS topic or an S3 bucket. So it has these two things that it needs to be able to function the permissions and locations. So the challenge that people often hit very early on in serverless development is if I'm writing software on my laptop and I want to test it without having to go through a full deployment cycle, which may take a few minutes to ah to deploy the latest code change. Even if it's, ah one character change up to the cloud service provider. How can I actually test with proper permissions and proper service discovery location mechanisms from my laptop? What mechanisms are there to do that? That's something that we are always evolving. But ah, especially here at Stackery, we have some some interesting ideas of how to make that easier.

Jeremy: Well, so what about maybe some of these these ways that we try to replicate the the cloud environment locally? So we can run Docker containers, maybe that simulate this. We've also have, you know, we can do mocking. We can do stubbing. We can do some of these other things. Why are those a good idea or not a good idea?

Chase: Yeah, there's two different approaches. There's the fakes where you run a fake version of DynamoDB. A fake version of S3 and you run those oftentimes by running docker containers on your own laptop that you facilitate how your compute resource, your function can locate it. But the challenge there with running fakes is that docker is a little bit of a beast to run. Ah, and not everyone has 16 gigabytes of memory on their laptop. While you can spin up a fake S3 service and you can spin up a fake dynamoDB table, it can be quite challenging to spin up fakes for all the different services that a modern application consumes. Many applications that we see are comprised of a dozen different services in so faking all those is challenging. On the mock side, that's where instead of running a full fake service for testing purposes, you've got the ability to say, oh, pretend that there's an S3 API that responds with these data when it when certain requests are made to it. The challenge there is, uh, you still have a fidelity challenge of what happens when the service is updated and has new features that you need or slightly changes, or the mechanics are slightly different at different scales. Uh, these are things that can only reliably be tested in the cloud. Ah, and not on a laptops. You get kind of ah ah, partial fidelity out of it.

Jeremy: Right. And I mean, the other thing you have too is when you start to test some of these complex workflows, right? If you're doing you an SNS to SQS and then consuming that with the Lambda or you're implementing something like the Event Fork pipelines that AWS recently released or trying to do, maybe some sort of fan out process and things like that. You're not going to see the sort of how those actually work locally until you deploy those to the cloud.

Chase: Exactly. Yeah, again, it goes back to that idea that your applications now are a graph of connected resources it's not a monolithic service. And so how do you connect all these things together if you have a function that writes to a DynamoDB table, but then there's a stream coming off that DynamoDB table of all of the events that have occurred, how do you model that with mocks and fakes? I'm sure it's possible, but it is not very straightforward. And so then you do it for one service, DynamoDB over here, now you have to do it for all the other services as they have all their different ways of spawning events and interacting with other resources themselves.

Jeremy: And I can imagine that could get pretty tough when you start working with teams, right? And you have maybe one thing is mock, maybe one thing is, you know, like you said, is using a fake to process the data or whatever that the testing there gets pretty difficult.

Chase: Exactly a lot of times when we're talking about fakes and mocks anyways, it's to fake or mock away one part of the system as you develop and test a different part. But what happens when you're done developing and testing that one part and you need to focus on a third part? You may need to tear down those mocks of those fakes. You may need to completely buildup, different, mocking and faking. And so you sometimes see sort of an explosive growth of test framework mechanisms to be able to test everything in lieu of actually testing things in cloud-like environment.

Jeremy: So why don't we talk a little bit about the development tools that are available to serverless developers. And I think that where you have a lot of experience here is obviously with working with customers for Stackery, you get to go into these companies, see how they're doing it now, and give us sort of a really, I think interesting insights into what other companies are doing. So could you maybe give us a typical development workflow? Or what you see is the typical development workflow when you come into a company?

Chase: Yeah, that's ah, it's a great question. There's a lot of people out there as well who are fairly new to the whole serverless ecosystem. And we have learned, as we've engaged with many of them that Stackery, what that looks like and and what they feel. There's a lot where people understand based on what they find online, that serverless has a lot of benefits, and they want to achieve those benefits. Whether it's scalability, managing costs or making cost more predictable. Whatever it is that that drew them to the possibilities of serverless, then they needed to take the next step of "how do I realize that?" What tooling is out there to help me build my application? People oftentimes do a Google search, they come across the Serverless Framework, which is a great entry point into the serverless ecosystem. But then one of the places that starts to stretch too far is as soon as people get beyond building a single function or a function that is in response to one set of events sources and they need to expand into a graph of sources where you might have an API that has a route backed by a function, which needs to access a DynamoDB table which has a stream that's gonna create another function, that's all technically possible to set up inside of the Serverless Framework, but it starts to become very challenging to do that without diving into the raw CloudFormation syntax, which is the infrastructure as code that the Serverless framework compiles down to. So people then start to look around and many of them come to us. Starting to ask questions around, how does this grow beyond a view of just functions that respond to events? Beyond that, how do I manage environments so that I can deploy my own version of our application? And each of my team members can deploy their own versions into different environments? And into production and staging and test environments, and how to manage credentials. And the really interesting and fun thing is that there's there are great solutions out there for all of this. There's AWS systems manager parameter store for parameters. There's AWS Secrets Manager for managing credentials, CloudFormation and SAM, the serverless application model, which is kind of an extension on top that AWS publishes. These all provide great mechanisms, and people just need ways of, of piecing them together in the same way that they're piecing together the underlying serverless Lego blocks to build their applications.

Jeremy: Yeah, and so the other thing, too, is a sort of a CI/CD pipelines that you see a lot of companies deploying. I mean, there's always articles out there. Here's how you do it with serverless. Here's how you do safe deployments, things like that. Is that a challenge, it seems pretty straightforward. I've set a few of these up myself. Is it relatively straightforward? Or is this something that you see a lot of customers having trouble with as well?

Chase: Well, lots of people want to know, how do I set up a CI/CD pipeline for serverless? And in fact, that's one of the top questions we get, but it actually isn't quite just, how do I set up a CI/CD pipeline? Certainly, that's part of it. That's part of any proper application environment. But pipeline is really the key here. They're asking more, they're asking not only once I got code written how do I integrate and deploy that out effectively? They tend to be asking for what is the entire workflow process from how I set up a project, how I manage it inside of a git repository, how I manage that parameterization, the credentials, the processes of workflows that individual developers go through to build out that application. So there's a stand in that people will say, "how do I build a CI/CD pipeline?" because they don't have a term for the everything before that. That's really about the full development workflow and life cycle.

Jeremy: I think you bring up a really good point, cause that is one of those things where I feel like the entire workflow changes. Right? So it used to be like we talked about at the beginning, you make a change to the the code, you uploaded to your hosting provider or even if you did have a CI/CD process, you check that in, you go through the git workflow sort of approval process, and then you automatically kick off builds. I think a lot of that is the same, but certainly when it comes to doing some of the testing, I don't know if it's as easy anymore, as just saying, "Oh, we're going to create a test environment or dev environment or staging environment and then a production environment." I think you have a lot of individual developers that wanna work with their own environment, like maybe have their own test environment and you run into the problem there and maybe not so much a problem, but just something that you have to sort of understand is, you know, how do you create all these separate test environments for your individual developers and maybe even beyond that, what if you are using external service is like a, let's say you're using Aurora Serverless, for example. Do you create a separate instance for every single developer? Do you have a set of shared resources that multiple developers can use? What's the best practice for that?

Chase: Yeah, that again is one of the leading questions that we hear from people. And the challenge there is that AWS has built up all of these services under this idea that you use them within this single account. And over the past few years, they slowly realized that there needs to be some mechanism for creating something like an environment. And so there's ways that we help our customers manage environments if they need to deploy into the same AWS accounts through proper name spacing of resources, but one of the most effective mechanisms for achieving environment isolation is actually setting up separate AWS accounts. So what we suggest to our customers is that they create separate AWS accounts for each of their environments: production, staging, testing, but also for each of their developers. So Bob, Sue, Joe and all the developers on a team, it's actually lightweight enough, and it doesn't cost anything to set up AWS accounts all underneath what AWS calls an organization umbrella and then within tools like Stackery, you've got the ability to tie each of those accounts to different environments so that each developer can have their own environment. Now, the reason that is oftentimes glossed over why this is something new in Cloud Services, certainly was never really done in data centers you didn't have separate environments in a data center for each of your developers, is the fact that you can now actually spin up individual serverless environments and applications at full scale for each of your developers. And it doesn't cost anything or cost very little to run at a steady state. That's one of the key benefits of a serverless architecture.

Jeremy: Yes, I love that idea of isolating individual developers, especially when you start developing something against, DynamoDB tables or even if you're using RDBMS, and you need to have something like Aurora Serverless, which you can scale down to zero and shut itself off if you need to. But yeah, that's definitely great advice. And that's certainly something that I recommend as well. So before we move on to the next subject, let's just go over quickly some of the IDEs that are available to serverless developers. Obviously, Stackery has a web interface right now, but what are some of the other ones? VSCode is a good example. What else could developers use if they were looking for an integrated development environment?

Chase: Yeah, this kind of gets at the heart of one of the key loops that we talk about within Stackery, about how developers build applications. There's sort of this outer loop that is around infrastructure maintenance, management, building. So as you need to create new functions, you need to create new resource is like DynamoDB tables or change the routes of an API. That's this outer loop. There's a process in which you have to build and define that. And as you mentioned, one of the tools that, one of the ways we help with that at Stackery, is we have a visual editor that slurps in your AWS CloudFormation, or SAM, or even the Serverless Framework projects, and it lays it out visually, allowing you to add new resources by dragging and dropping, and wiring them up. But then compiling all that back down into the native raw CloudFormation or Serverless Framework, SAM, what have you. That workflow helps with this outer loop challenge of managing the infrastructure. And then there's the inner loop challenge of how do I edit the code that is in my compute resources and quickly iterate through testing and reediting that code. And so one of the IDEs that people oftentimes use, VS Code is extremely prevalent as it's got a great system for NodeJS and Python, which are the two highest usage languages in serverless. So Visual Studio Code's a great one. Jetbrains and their platform is a great one. PyCharm for python as well. In fact, this has led us to ask ourselves, what could we do to make both of these loops, the outer loop where you're modifying infrastructure, and the inner loop, faster? And so, just this past couple of weeks ago, we launched integrations through a Visual Studio Code Plugin that makes it side by side, you're able to have your template for your infrastructure open and you're able to visually model all the interactions between resources where you can drag and drop. And as soon as you drag in a resource, you see it update in the template and vice versa. And that helps with the outer loop. And then on the inner loop, we also have launched this local invocation mechanism where we were talking earlier about how challenging it is to run your compute code on your own laptop and yet still have the fidelity of interacting with cloud resources without doing mocks and fakes.

[24:50]

Jeremy: That's sort of the last thing that I want to talk about. I want to get into this idea of how do we develop locally but use these cloud resources, right? And so you've used the term, I've heard the term before "cloud-side" or "cloud local." You know, just the idea of bringing the cloud down to your laptop, as you've said in the past. So let's talk about that, right? So before we get into what Stackery does, what are the challenges right now for somebody that's trying to access remote resources when they're developing, maybe a Lambda function locally?

Chase: Yeah, you start with some code, and that code for a Lambda function has this handler that gets invoked. One of the things that I did early on when I was starting to play with this to try and speed up this this iteration workflow is "well, I could write a little wrapper script that invokes that handler code function with some test data just to get it running locally without having to deploy it all out." And, there came along some tools that kind of helped facilitate this mechanism. AWS Sam, their tooling has SAM local invoke where it will take your function code, and it will actually spin it into a docker container and run it as though it's in a proper lambda environment. Ah, the Serverless Framework has a similar thing. But even there you have a challenge where the permissions that your function has is based on the permissions that you have locally on your laptop. Now, a lot of developers, they have permissions on their laptop, but they have sort of administrator permissions. They can, if they wanted to interact with any resources inside of their AWS account. Whereas the function that you're building it's tied to a very specific set of permissions where you don't normally give it full administrator access. So you have to sort of a lot of times you get your code working. And then as a second step, you have to figure out is the code still working when I deploy it to the cloud and I've got the permission set the right way. And then lastly, you've got the challenge of that service discovery piece where if I'm running on my laptop, how does my function know which DynamoDB table it should be interacting with, which SQS queue it should be sending messages to. So you've got to solve these problems through some mechanism, and a lot of people come up with their own little test scripts on the side that help here and there. But there's a real challenge there around, ah, having a workflow that a whole team within an organization can uniformly use and provides them with that sense that they're bringing the cloud to their laptop locally.

Jeremy: Right, and it gets even more complex when you start thinking about different stages, right? So you're always publishing to DEV or to TEST or to PROD or whatever you've named them. And so if you are trying to access your DEV tables or your SNS topic for DEV and so forth, and then you also have the problem to where, even if I do publish to the correct SNS topic in the DEV environment, if I've got other Lambda functions or SQS queues, subscribed to those, they're going to get those, those requests are gonna go through, so even if I'm testing locally. So you just have a whole bunch of things where there's probably not a perfect solution to it. And then you and I have talked about this, I wrote a little plug-in for the Serverless framework. All it does is just resolves the references in there. But you've gone a lot further with that by actually taking care of at least the permission side of it and that service discovery, as you said, right?

Chase: Yeah. So what we did was we piggybacked on top of AWS SAM local and how they invoke things inside of that neat little docker environment. And we went beyond to say, what if we know where you've deployed an application into an AWS account and we can go and ask well, for this function deployed under this name, why don't we go grab all of the environment variables which are often used for sort of the service discovery mechanism you'll put. If your function needs to talk to a DynamoDB table, you'll put the name of that table in an environment variable for the function, using some magic that CloudFormation has. And so, uh, we go out, we grab the environment variables for the that Lambda function, which helps with the service discovery. We also go out and we do something called assuming the AWS IAM role, which is what holds the permissions that the function is granted. We assume it, which means that we get credentials that give us the same permissions. And then we feed that into the local invocation as well. So now you are running your code inside a proper Lambda context, in terms of the operating system image. You have the same environment variable values that your function has in as it's registered inside of your AWS account, and you have the exact same permissions. And so your iteration loop here is on the order of if you make a code change, or you make a code change even if it's a one line change or if it's, you know, adding whole files, whatever it might be, we watch for changes on the file system and we rerun that function. And so we're talking about on the order of about five seconds to fully retest a function that you are that you're developing, all kind of in that cloud local environment.

Jeremy: Yeah, and so the other thing you had mentioned to about the plugin that you developed for VS Code. So if anyone has developed sort of a, uh even the most basic of serverless apps, once you start adding resources, you know, it's always funny to me, especially if you're doing SNS topics with subscriptions, you have to create the permissions for the subscriptions it's like seven different things you have to do for every SNS topic with subscriptions that you want, and I have to have that in your resources. So this is going to take sort of the experience that I get on Stackery now and you're gonna bring that right onto my local laptop. I don't have to communicate with the web or do any of that stuff. It's just gonna run locally.

Chase: Yeah, one of the things that as we built Stackery and it's got this this neat visual editing that, really unlocks the power of CloudFormation and all of the AWS services inside of AWS for people, it actually becomes almost like a learning platform for them as they become more comfortable understanding, oh, if I drag this thing on in on this canvas, all I see it doing is adding these few resources in the CloudFormation template. They start to become more comfortable with it, and then they start to learn, and they start to be able to teach that to the rest of the team. It's been interesting seeing that become, ah, kind of learning platform. But leaving that aside, one of the challenges that some people had with the Stackery visual editing was that there's this context switch where, if I have to create a function and then modify code, I would log in to Stackery in a web browser and I would add a function, I drag it into the canvas and wire it up, and that's great. I would commit that to my git provider, but then I'd have to go to my IDE and I'd have to pull down the changes. And now I'm ready to write code, and when I'm ready to test, I have to push that back up to git, and then I have to go back to Stackery and tell it to deploy. So there's a lot of context switching involved, and so one of the really exciting things about what we released a couple weeks ago with our Visual Studio Code integration is the fact that we're bringing all of that into the IDE, you don't have to leave the IDE, when you need to modify your infrastructure. You don't need to leave the IDE when you need to deploy out to your AWS account. You don't need to leave the IDE when you're iterating on your code. And so this is really that next transformative step in helping people build serverless applications effectively.

Jeremy: Yeah, I mean, and for me, just the amount of documentation surrounding CloudFormation is overwhelming and as good as the documentation is, and it has what it needs. Sometimes the examples aren't always super clear. You always have this issue where it says whenever you have nested sort of nested values, it will say, Oh, this requires this SNS topic value. Whatever it is, you have to click into that and then see what that has, and that may have nested things. So I think you really hit the nail on the head when you said, you know that you can use it as sort of a learning experience or a learning tool, because it does, it gives you all of the possible options. You don't have to go and look at all this documentation, it kind of does it for you. And of course, once you do these things a few times, certain things become sort of second nature to you. But definitely, really, really great feature. So, I mean, I love what you guys are doing over at Stackery, you have a great team over there. I think people could can really benefit from from using this just even if it's just use it as a learning platform. But obviously, you know, to become customers of yours, I think would be great as well. But well, it was so I think I think that probably wraps us up. Um, you know, thank you so much for joining me Chase and for obviously continuing to share all of your knowledge with the serverless community and with what you guys are doing at Stackery. So how can people find out more about you and what Stackery is up to?

Chase: Yeah. So our website is stackery.io and we push things out on our blog very frequently, at least once a week, if not much more.

Jeremy: I like the name, by the way. Stacks on stacks.

Chase: Yeah. Thanks. Thanks.

Jeremy: A good name.

Chase: Yeah. One of the most interesting things is we've got a changelog. And we have some people in our team who are exceedingly witty. So I don't know if a changelog is the most exciting thing to most people. But, you know, in the past month, it's been, ah, adding support for visually editing web sockets API, to the cloud local invocation mechanisms to the IDE integration. All these exciting things, changelog has it first. And, you can always reach us on Twitter as well. Personally, my handle is @txase on Twitter, and I would be happy to answer any questions people might have. Lastly, I would say inside of Stackery itself say that you're thinking of doing serverless. Maybe you've tried it, but you have had a hard time figuring it out. You're just not sure you don't have the confidence of how to build a successful serverless application. If you sign up and you start using Stackery, if there's any questions you might have, one of the most prominent features of our application is there's a little chat box down there. And, we've staffed that all the time and we get back to people, not just about like, oh, I can't figure out how to do this in Stackery but like, oh, I need to add a CloudWatch alarm on a random resource metric. I don't have any idea how to do that with CloudFormation, and we help people get off the ground running. Oftentimes it just takes a couple of interactions to understand how AWS accounts work, how CloudFormation works. We'll be there to help you figure that out and then get you off to the races.

Jeremy: And you guys have a, speaking of teaching, you have a live stream every Wednesday, right?

Chase: Every Wednesday. Yep. We had you on a couple of weeks ago. It was great, it was one of the best ones I think when you talk about the architectural patterns.

Jeremy: Well I appreciate that. And then, last thing again, I know you guys do a lot for the community. You host the Portland Serverless Meetups, right?

Chase: Yes. Every month we've got our Serverless meetup, taking in the best serverless practitioners from around the area.

Jeremy: Awesome. All right, well, I will make sure we get all of that stuff in the show notes. Thanks again, Chase.

Chase: Thank you. It was a pleasure.

View Details

About Marcia Villalba

Marcia Villalba is an AWS Serverless Hero and a software engineer from Uruguay, currently living in Helsinki. She currently works as a Full Stack developer at Rovio, and has her own consultancy company, Unicorn.codes. Marcia also hosts her own YouTube channel, FooBar, creating fun and creative videos and tutorials that focus on how to use AWS serverless technologies and managed services. She loves to help others in the serverless community learn about migrating to serverless, and publishes courses, all of which you can find on her blog.

  • Twitter: @mavi888uy
  • Instagram: foobar_codes
  • Blog: marcia.dev
  • YouTube: FooBar Serverless

Transcript

Jeremy: Hi, everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week I'm chatting with the fabulous Marcia Villalba. Hi, Marcia. Thanks for joining me.

Marcia: Hello, Jeremy. Thank you for having me here and I hope my cat is not meowing because she's already meowing.

Jeremy: That's fine. My kids love cats. I'm sure the people listening will love them as well.

Marcia: She always appears in my video. So it's part of FooBar.

Jeremy: Perfect. So you are a full-stack developer and an AWS serverless hero. So why don’t you tell the listeners a little bit about yourself and what you've been up to lately.

Marcia: Yes. So I’ve been doing serverless since 2015 so Lambda was launched in November 2014. I started quite early on more or less when API gateway was announced and was able to connect with Lambda. And since then I've been just working on different types of projects. My first project was to migrate something to serverless and then I've been working on greenfield projects most of the time. Then one of the things I really like is to create content. So I think, besides my show, that's what I do the rest of the time - my shows, create YouTube content, and a lot of courses. I have a small consultancy where I help companies with workshops and training, some things like that. So I spend all day doing this serverless stuff.

Jeremy: I know that feeling. You have a blog too, right?

Marcia: Yeah, as well. My blog is not really a blog per se. It’s more or less where I gather all the content I create in one place. I used to use Medium but I quit. So I think blog is more like I can do whatever I want with it. It's just a place to have my content. But most of the content I create is video. It's the place I feel more comfortable. So YouTube is the place I hang out. I just put all everything there, but well, the blog is good too.

Jeremy: You do an amazing job with all of the videos. So thank you so much for all that stuff. So I wanted to have you on today to talk about AWS AppSync. So I've seen I've seen you do your talks before. You have a great talk on the subject and you've done videos on it and things like that. So maybe let's start by, just in case people don't know, because I think this is one of those – it’s not obscure if you are in this world, but if you're just kind of getting into it, I think AppSync is maybe an obscure sort of thing. And why that's different than API Gateway. Maybe you could tell us what is AppSync exactly?

Marcia: Well, in a few words, AppSync is is a GraphQL service. So it's managed GraphQL service by AWS. It’s a platform where you can kind of – AWS will take care of all the heavy lifting for the GraphQL. And you just basically need to put your schema and create your resolvers that are a bit. But basically the schema and the resolvers are the two proprietary things that you need in order to have a GraphQL application. The rest is very generic between every GraphQL implementation. So AWS create this platform and then you just do the smallest amount of work you can and get a working GraphQL server. And that's really good for, for example, creating mobile applications. It's really fast to create fast back-ends for mobile applications or to interconnect multiple microservices or do all kinds of things. So it's a very interesting service.

Jeremy: So it's basically like a managed Apollo server.

Marcia: Well, yeah, Apollo has their own platform as well, so they have their own AppSync. So, yeah, this kind of concept is – when I talk in my talk about the ways, because GraphQL is a specification. It’s not something that somebody –well, somebody wrote it down and then different people implement in different ways. So there is, like, three ways of doing GraphQL, as I said. Either you write your whole GraphQL specification yourself. You are a hardcore developer. You want to have your GraphQL done in Cobol and then you do it yourself - I don't know why. Then the most common way is that you use some library like Apollo. That's the most popular library that you just put it in a server, and maintain that server and this kind of serverless way of doing GraphQL is using a platform where all this heavy lifting of maintaining the infrastructure, and do we know the small connections that needs in order to hook up your library to your server. That is already done by the platform, and you just focus on your business logic. That is the main mantra of serverless.

Jeremy: Awesome. So let's actually talk about GraphQL specifically because again, it's another one of those things where I think a lot of people are very used to REST APIs that has been the way to do things for quite some time. I think it was Facebook that came out with GraphQL and…

Marcia: Yes, GraphQL was released by Facebook, but now it's owned by everybody.

Jeremy: So maybe just quickly explain what is GraphQL exactly?

Marcia: So first of all, GraphQL and REST are not like enemies. So you do not have to choose one or the other. And I think that's an important starter for the discussion, because people are like, well have my REST endpoints to use GraphQL. No, we are not talking one or the other. They're two different things. So GraphQL is a specification that when you implement, usually it sits between all your microservices and your clients, and it’s, thus, like an entry point for your application. So you can, instead of having multiple different endpoints and point to the different microservices independently, you can have one entry door and that's really convenient. For example, if you're doing, um, I don't know a mobile app and you have multiple different microservices with different people working on them, then you can combine all these requests and responses into something that is like a contract between the client and the server as a big entity. So that helps a lot of the client developers, because that's one of the big problems when client developers are starting to work on a project, they need to [really understand] the whole backend architecture, and sometimes there is really not a lot of need for them to understand that. So GraphQL will provide a contract where all the possible operations are specified. All that is a strongly-typed language. So all the operations request a response with really clear defined types that they have strongly-typed. So you know exactly what you can put in, what you can get out, and then when you do these operations that you'll get a type back or many types back and then you can, in your request, you can ask for exactly the same, the right fields that you want from this type. Because that's another problem with REST in general for mobile developers that they need to over-fetch a lot of information a lot of the time and do the filtering in the client, and when you're working in mobile apps then if you're fetching a lot of information, that's a lot of bandwidth, and you might need to do a lot of requests to the back-end and GraphQL will unify everything in one response request. So it's in a way quite efficient for mobile development.

Jeremy: Yeah, you mentioned over-fetching, and I think that's one of those things where you run into that with rest APIs where you know, again, you can't – and you can design REST APIs to act as like a GraphQL server, in a sense where you limit what comes back. But you also have that problem of under-fetching, right? So you bring back your list of products, but then you want to get the product details and you have to make additional calls.

Marcia: Exactly. So the idea is that you can fetch many resources, many types, in one request. So you don't need to fetch one and then fetch another one, and fetch another one. So that's kind of one of the benefits of GraphQL when you use it. So in general it’s a good combination to have both REST and GraphQL, so, for example, if you have, in one of the projects I work [on], we had a back end with a GraphQL, but at that time AppSync didn't exist, but we were just working with plain Apollo, and it was really good because we have 16 different types of clients. So from which we maintain four, the rest were maintained by different types of people. Some of them were consultants that were paid to do this specific application for a smart TV, for example, and it was never maintained again. So GraphQL really performed because you don't need to really synchronize a lot of documentation back and forth because with this contract, everything is specified there and then there is no versioning in GraphQL. Everything is done in one contract, so you evolve your API with time, and GraphQL has kind of ways to do that. So if you do it right, it's very easy to be backwards compatible and support all versions with no problem. And that's one of my favorite features because when you're working with REST, having version 1.2.3.1

Jeremy: Does anybody actually version and then point though with REST API? It's always, so I’m like, well, we're on version 1, but it's really like version 9 at this point, because you keep changing things.

Marcia: I have seen version 49, so yes.

Jeremy: Some people do. But what I like about GraphQL is that fact where it's like you can add a new field or you can add a bunch of new fields. You can add nested fields and things like that that that you just keep adding to your schema and really the weight or the need for maintaining backwards compatibility – it’s all done via the clients. Clients can just enhance themselves by using these new endpoints, but you don't have to get rid of them because they're not being sent down in every request for new clients.

Marcia: You have a choice to get rid of them. For example, if you are changing something in your backend and you're not just returning that field. So they have these two concepts that is hidden files and deprecated files. Fields, sorry. So you can play with those concepts and start like removing attributes from your types if you need to or you can tell this type is not valid anymore. You still can return it, but we will not like – it’s a way to communicate with your clients because at the end of the day, sometimes you cannot return always everything back because the backend changed.

Jeremy: Eventually. That's what I actually love that metadata deprecated stuff that you can put into GraphQL. That’s another really cool feature. So let's go and talk about resolvers for a second. Because I think this is one of those concepts too, where people that are new to GraphQL, what exactly is a resolver?

Marcia: I said that, at the beginning, that AppSync is a kind of already a packaged GraphQL platform that where you need to put your schema, that is the contract between the client and the server, and then the resolvers. So the resolvers are the unique way that your GraphQL middleware kind of component will connect to the different microservices. So that's very proprietary from every application. It’s basically a set of instructions. So imagine that one of your microservices, we can call a data source - that's the word that AppSync uses - because they might not be microservices per se. They can be things like a SQL table, like a relational database. So imagine that one of these data sources is a relational database and we have type, that is, I don't know, order. We have a table that is order. So we want to link our GraphQL type to this table. So the resolver will need to understand SQL, so we need to open a connection to that data source, and do Select All from the orders table and if you have some condition or if you have some fields that you need, so we'll need to write this SQL query and then it will need to get the results back from the table, parse it in a way that it becomes JSON, that is what GraphQL puts out in the response. So it's doing all these translations between the client and the server. So it is a very, very critical part from GraphQL.

Jeremy: Yeah, and I hear a lot of people when they say GraphQL server, they sort of think that’s sort of the end-all, but it's really an intermediary, where you have a client request that goes to the GraphQL that sort of does this assembly or calls all these resolvers, puts everything together for you, but it's really not the server. That's not where your data’s coming from.

Marcia: GraphQL is not a database. It doesn't have any information itself. It's just handling requests and pushing it to the other side. So it's kind of just translating in the middle. Like if you were speaking English and I will be speaking Spanish and we'll have somebody in the middle that is translating.

Jeremy: Well, you speak beautiful English, so I wouldn't worry about it. But so all right, so let's get into AppSync now, because this is where I think we can start putting all this stuff together. So you've got GraphQL, you've got the structured language that you can use. You've got this concept of resolvers, so AppSync is serverless or in the sense that we don't…

Marcia: It’s a managed service.

Jeremy: Yes, so we don't have to worry about the back end servers. So how does AppSync work? What does that do? What's that flow via AppSync?

Marcia: So when you start with AppSync, you need to set up your schema, so you need to find all the types of your kind of GraphQL application. What is GraphQL returning and all the operations that the users connecting to these AppSync or GraphQL servers can do. So you will decide, okay, they can create orders. They can update orders. I can see all my orders and these are my types, my orders, my products. You find all that in the schema. The next step is to set up what are your data sources, where the data for populating this schema comes from. So you need to say OK, I have an SQL table, I have a Lambda, I have an ElasticSearch, and you connect the types with the different data sources or the operations with the different data sources and when you do that connection is when you build the resolvers, that kind of linking between the schema and the data sources is when you need to write the resolvers. So when you use AppSync a lot of things come out from the box if you're using AWS services. So, for example, while the schema you need to write that you cannot escape, but the data sources it has kind of really native connection with Dynamo, Lambda, Elasticsearch, RDS. I think the next one is HTTP in general also, you can call whatever you want. So those are the five data sources that you can use pretty easily and your resolvers need to be written in a specific language. So it's using this velocity language scripting, or velocity template language, VTL. I'm very bad with acronyms - sorry. And that's kind of, one of the trade-offs when you're using someone else’s platform, that you cannot decide which language you want to write your resolvers in.

Jeremy: I don't want to interrupt you…

Marcia: Sure, you can interrupt as many times as you want.

Jeremy: Okay, so before we move on though to getting into the VTL. So we mentioned these data sources and you said Lambda, you said Dynamo, HTTP, Aurora, and ElasticSearch. So those are five very different types of data sources. And so when you're talking about building a resolver, when it's DynamoDB or it's Aurora, it’s ElasticSearch, AppSync will actually be the resolver for you, or the resolver will be built in.

Marcia: You have templates So, for example, if you're using Dynamo, you can just go click on Dynamo and then it will show you a template. Okay, you want to fetch one? You want to fetch two items from a list of items? Or do you want to insert something on Dynamo? And you just click on the template and it will create the resolver for you. Really easy. The same with Lambda. If you want to invoke a Lambda, it kind of comes out all from templating. So it's very, very easy to get started to writing those resolvers.

Jeremy: But the idea would be is that essentially that VTL, the template that you create that can do that translation between DynamoDB and your schema, that kind of acts as your server, the other end of the GraphQL…

Marcia: That’s your business logic, kind of.

Jeremy: It does that business logic. And then if you do want to do very specific business logic, though, then you can do that. Um, you could do that with a Lambda function or call another service, another HTTP, or something like that.

Marcia: VTL is very powerful. So that's something. Sometimes it happens, people underestimate what kind of things you can do. You can do really a lot of things only by playing with VTL over Dynamo, it’s super crazy the things you can do. And then if you want to do some complicated operations, you can always call a Lambda and the whatever from Lambda. Then also AppSync has this feature called Pipelines, where you can connect, for example, a Lambda, different data sources together. So imagine that you're fetching a profile metadata, and you have to do like one of the fields in that metadata in Dynamo. We have the profile information metadata and one of the fields is a link to a photo in S3. So then you can hook it with the Pipeline and fetch that image from S3 and create this kind of signed URL that you can return back and show in your web app. You can do that pretty easily with AppSync on this Pipeline concept that they have, that you can connect as many resources, data sources as you need, together to return something to the client.

Jeremy: Yeah, that's a pretty cool feature. So what else about AppSync? What are some of these other features? Because it goes beyond just being a managed GraphQL server.

Marcia: Out of the box, one of my favorite things is security and authentication, because that's always something. If you have a mobile app, then you always have some kind of barrier between like all your data, your personal data and the world. So that comes out of the box with AppSync so you can get started super fast with Cognito. It’s super simple to set up, and then you have username, password or Facebook signing or Google signing, or whatever you like from Cognito. And then there’s also more specific ways to authenticate. But I think the most common ways to use Cognito [are] super simple. To get started, you can use an existing using user pool or create a new one. Then you can even filter inside the fields that you don't want to show this information to everybody else but you. For example, if we go back to the profile example, that you can fetch this profile information, and if it's Jeremy's profile, I will be seeing some filter out data and Jeremy will be able to see all the data, like your email address or your phone number. So you can create that with very simple VTL, connecting Cognito and Dynamo, for example, without writing a line of Lambda’s code.

Jeremy: Yeah, and I think people need to understand the importance of that, right, because I've worked with GraphQL before, and then you've got security where it's sort of like, well, GraphQL doesn't do security. It's not part of the specification there. Your resolvers handle that. If you have to write that code into every resolver and then manage it across all these different things, it gets really complicated, so the fact that AWS has Cognito and you can do the federated logins across all those different social networks and things like that, or plug it into Auth0 or even write your own custom one.

Marcia: You can do whatever you want, and there's quite a lot of – they keep on adding different authentications ways, but I think Cognito, if you want to get started, it’s the easiest way and it's really out, integrated with all the mobile kind of platform development that they have AWS build so all the libraries and everything works really well with Cognito.

Jeremy: All right, so why don't we talk about maybe some of the use cases? Because that's the other thing about AppSync where I think it confuses people. It sits or it originally was sitting under the mobile, it was classified as a mobile service, right, because it did some, I think actually, that you could do web sockets originally. There's some real time stuff that you could do there with it. What are some of those use cases that you can use it for – not just with mobile, but with just a web app or, you know, any other service?

Marcia: So I think the easiest use case is building mobile apps. So, as you mentioned, AppSync supports real-time communication out of the box, so that's one of the benefits. If you have an app, you don't really need to sync it, ask the server every 10 minutes, “Have you updated the information?” So that happens in the background with AppSync? So that's a feature a lot of mobile developers are looking because sure, you can have web sockets with API gateway, whether something the client will need to create all the support for that. But with AppSync, it just comes out of the box when using AppSync. So for me, the strongest use is to build mobile apps, but in general, whenever you have multiple microservices and you need to unify them in some kind of way, I think also it kind of works very well. If you have a client applications, not only mobile but desktop, some things like that, that works very well. Also, it's a very good way for managing when you have these client applications that you cannot control, like in our case, that we have all these like different apps that were managed by externals and in a way, it's kind of easy to have this auto documentation in place because, let's be honest, we are very bad at writing documentation, developers, and when you're in a very agile world, then things tend to go faster than the documentation goes. Having something that kind of enforces the documentation just by its nature, it makes life so much easier.

Jeremy: I was told very early on in my career that the code is the documentation.

Marcia: Exactly. And GraphQL is proof of that.

Jeremy: I don't agree with that [that the code is the documentation]. But anyways, that was what I was taught. So some of the other things it does too is like offline sync, right? So if you make changes on a client, it'll store that information and then when you do get connectivity, that'll sync up, and then you obviously have conflict resolution that it needs to deal with and AppSync does that as well.

Marcia: Yes, that's the simple conflict resolution. So if you have very weird things that you might need to create some kind of resolvers there but out for the basic things, it just does it on the back end. So it's important to understand when you use AppSync in the client, you will need to have a library in your client to handle all of these things. That's kind of — it is not something that happens magically. You will need to use a library to connect to AppSync in your client, either web or mobile app. So it will do all the connection and the magic between the server.

Jeremy: All right, so what if we, I don't know if we can do this via audio. But let's say someone's driving in their car, and they just want to get a sense of wiring up a data source in AppSync. Now, obviously they'd go and watch one of your videos. You've got plenty of them on this. They probably can't do that when they are driving or mowing the lawn.

Marcia: What? No, no, it's not a good idea. Don't do that.

Jeremy: So just quickly, just walk us through what that process might be.

Marcia: So first you create a schema, and then you need to link the schema with the data source. So imagine that our data source is a Lambda and we want to invoke that Lambda. So basically, what we will do is just write one line of VTL that says invoke this Lambda and you pass the parameters that are coming in the request. It's super simple. Then you need to write the Lambda that will do all the magic and then you will return from the Lambda some result and then you need to write another piece of VTL that will be the response resolver that will grab that response that Lambda did that we matched in the destination that can be returned like it is to the client that will just grab whatever is in the in the response of Lambda and push it out. So that's [a] super simple kind of example. If you want to connect to Dynamo, what you do is you write a Dynamo query like the ones [that] are very similar like the ones you write in the recent AWS SDK, just with a little different format. So you will put, I get item, you will create this kind of params object where you put, I don't know, the key, the table, and whatsoever in the VTL. Then you need to write again another velocity template for the response that if you are getting just one item, it's just basically whatever Dynamo returns you just push it out because AppSync knows that format and can do all the transformation for you.

Jeremy: So bottom line is learn VTL.

Marcia: Yes, yes, yes, yes, yes.

Jeremy: Awesome. All right. So what about performance of this? Because I've heard stories, people are like, "Oh, yeah, super slow when I'm trying to connect multiple data sources", things like that. What is your advice around, or what are your thoughts on performance?

Marcia: Like everything in life, you need to understand a little bit the things that you're connecting together when you try to go to production. if you're playing in your home, you might not care that much about performance. But in general, my experience of the people that I've been talking to [about the] experience of AppSync, it's very positive. But when people start telling you about performance, it's because they have done some kind of mistakes in their architecture. So one common mistake is people don't want to learn VTL, so they do Lambdas for everything. And Lambdas are quite fast, but it's an added latency in your system. So if you're calling Dynamo and the VTL can do all that, why call a Lambda that might be cold, that you might need to launch and deploy and call the AWS SDK and return and fetch. So that's one thing. Then another thing is that people like to scan in a Dynamo table.

Jeremy: People love scanning DynamoDB tables. I don't know why.

Marcia: Not an efficient thing you can do in Dynamo table. They're not meant for doing that, and AppSync is so easy to get started that it's very easy for people to start scanning Dynamo tables. And that's where performance gets totally kind of rough. If you don't know scanning Dynamo tables goes in kind of one row at the time and Dynamo is a big data source, so you have lots of items there, in general. Also, it's not meant for do[ing] that. If you want to find data use other ways, you can index that data or then put it in ElasticSearch or I don't know, use something else. Not Dynamo tables. Then another thing that people like to do is to sometimes connect to a HTTP. I know it's a data source, but you need to be aware that if you're connecting to an HTTP data source, then you're adding more latency there because that's going outside to the Internet. So everything you do has some costs in general. It's not like when you're inside your own network and everything is ready to fire. So I think those kind of things, it's good just to have it in mind and see how you can improve that. Can you build a cache in front of your HTTP somehow? Or can you do something to speed that up? If it's something you need to do, it's something you need to know. But sometimes it's better to avoid that and put everything in Dynamo.

Jeremy: Yeah, and I've seen HTTP calls just I mean, it probably takes 90 milliseconds just to set up the HTTP call, and then you have the latency that's added there.

Marcia: Exactly, and then you have, like, you need to [do] some weird querying that people don't know. So they call a Lambda. The Lambda then calls an HTTP, and then it returns back, and I have seen all those things, and then, sure, performance is not the best.

Jeremy: So bad architecture or bad queries. I mean, that's the thing. So I think that's one of those, that's that learning curve when you start working with all these services anyways, understanding how they stitch together and what's the most efficient way to do it. All right, so that's great. So let's actually talk about implementing this now in a serverless project, right? So how would we go about doing that? You mentioned sort of client libraries. What do we have to do in order to start using this?

Marcia: So the first thing is to create the backend. So you can either go to the AWS console and click, click, click, click around, and have your GraphQL backend using AppSync. Or, if you want to have a real production application, you need to use some kind of infrastructure as code (IaC). That's my recommendation. Then you can build the whole AppSync service out of configuration files that you can replicate in different environments if you need. So AppSync supports CloudFormation. So if you're using CloudFormation, you can use that. I've been using it with Serverless Framework because they have a really nice plugin, where you can basically build all your AppSync infrastructure, which has writing some files and creating some really simple specification there. So you create your schema file, you create your resolvers file. And then you have this plugin in your infrastructure page that, like in the serverless.yml, if you're familiar with serverless framework, where you just type all these links. Okay, this is my data source. This is my security. These are my resolvers in this case. These are my resolvers in that case. And it's very simple. And then that means that you can have the same GraphQL backend when for dev, and staging, and then production. You can create three different stages for that, or I don't know, whatever you need in your production. So that's the first step to create good production backend. The second step is to create your clients. So that's, you will need to use AWS Amplify. That's the easiest way, in my opinion. AWS Amplify is a library, a client library, that AWS has created, this open source, and it's kind of, I think it's called like a cloud for enabling cloud native applications or something like that. Just a fancy naming.

Jeremy: CDK. Cloud Development Kit is that the...? Is that different?

Marcia: No, no, no. I don't know how they come up with these names, But yeah, the thing is that this library is really cool because it integrates really seamlessly with a lot of the cloud services. So you can also use it for connecting to API Gateway, if you're using API Gateway. So it's not only for AppSync. So it has features like it creates your authentication so you can connect to Cognito. You can get your session there and then you can use it for a GraphQL for AppSync. So when you set up your GraphQL backend, you will get like an URL back and then some other information. With that information, you put in your configuration file of Amplify and it's kind of working. It's super simple to stitch together and the same way you can configure Cognito, you have to know your Cognito user pool and Cognito entity pool. Some other thing is you put it in your configuration file and then you can start calling the different APIs from that library, like auth. I think signing auth login or something like that. And then it's like GraphQL get or it's just very, very straightforward to use.

Jeremy: So if you want to run, like once, you have that all set up and you're logged in and is it Amplify? No, it's not Amplify, but Cognito, for example, same with Auth0, has the login interface for you. You don't even have to write that in your client application. Like that's just done for you.

Marcia: Yeah, you can do that as well. But if you want to have an application that looks like the whole application there, you can also do that with this Amplify library as well.

Jeremy: Um, so once you're all logged in and you're ready to go now you want to run a query. How do we do that?

Marcia: So the general, you write a query. I usually put it in a separate file in my client application because you see here for me to manage those. So I have a file where I have all my queries on I will write, I don't know, give me all the videos, for example, or all the orders with these attributes I want to see from the order. And then I go to where the logic is executing and I shall use the Amplify library, the GraphQL part, and I just do, "okay, get me this," and you will get a JSON by just passing the query.

Jeremy: We didn't talk about mutations. That's another...

Marcia: No, we didn't talk about queries and mutations. That's part of GraphQL actually.

Jeremy: Yeah, but just so obviously queries are when you actually call it, and that's what the query language . But then you can also send updates and make changes to different things. So you're signed in, you can write your queries, you can run your queries, mutations, so forth. That sign in, that like you mentioned, the Cognito user pool. The library automatically sends that authentication token...

Marcia: Yes.

Jeremy: ...up to AppSync.

Marcia: Yes.

Jeremy: And then AppSync uses that and then all your ACLs, your controls, and that's all done in your velocity templates.

Marcia: Yes. Exactly. And for example, when you pass like your login as Jeremy, and you want to see your profile. So, Amplify will hook in your Cognito ID, and then AppSync will see, oh, your Cognito ID is the same as the one that I'm allowed to show you the email, birthday and whatever, and that's all done in the resolvers. So it will return the whole profile to you so you can see it. But if I want to fetch your profile, it will say no, Marcia is not the owner of this record. So she will be only to see the name and the photo, for example. That's it.

Jeremy: So you mentioned, okay, so you're putting your filter, the ability to filter what certain people can see. We're doing that in VTL.

Marcia: Yes.

Jeremy: Maybe we're doing some of that. Would we do it? Would we offload any of that into Lambda if we were using a Lambda resolver or you still would want to do it at the VTL?

Marcia: Well, then, it's the question on who owns the data, and that's something, at least with AppSync, I'm still trying to figure out how to really architect my application, my graph qualifications, because I've been using GraphQL with microservices, and usually I do the filtering in the microservice because the microservice knows the data, knows who can see it, and I don't want to leak that information out. But with AppSync, at least applications and have been building, they are mostly contained into Dynamo tables and Lambdas. So I think when I'm coding this that AppSync is the owner of this data and, then I do the filtering in the resolvers. So I think it's always a question of who owns the data and who is able —where is the level that you want to leak the information out? I don't know if it's clear.

Jeremy: I actually I totally agree with you because I mean, that's always the problem you have with a lot of the security things where I mean, obviously, when you're using API Gateway and you're using Cognito or OAuth and you can send in, you can do the custom authorization, and then you get back a policy document and then you need to use that policy document, that can restrict certain higher levels of things like whether you have access to a particular endpoint. But if you want to take that information and then actually build logic within your application to say, "don't show this if a user isn't in this group" or whatever, you can go even further and do that at that level. But now you've got some of your security happening at this level, then another there, and I guess that maybe even begs the question: Where do you put your business logic? I mean, how much do you want to be putting business logic into VTL or is that something you still would want to keep out of AppSync?

Marcia: I think this is a really interesting question in general for serverless because when we are working with servers, we know that the server is the gatekeeper of data and [these are] the quite clear boundaries. But now we have the serverless world where there are all these kind of functions and different managed services that have access to the data, so how [do you] regulate that? That's something I think is a bigger question than AppSync. It's something I wonder myself every time I design a serverless application, like who owns this? Who has the power to access? Who can see it all and who can't?

Jeremy: Right. That might need to be a conversation for another day. So just a couple more technical questions that I have. So custom domains with API gateway, you could stick a nice little custom domain on it. You can't do that with AppSync.

Marcia: No, not out of the box, as you can do with API Gateway. So you need to go through the CloudFront distribution and do like Route 53 and hook up there the thingy, because in general, AppSync is thought in a way that nobody really will see your URL because it's an application to like, your client is connecting to your server, so you really don't care. So I think that's the thought when when they have not put that feature in. Maybe they put it in the future. I have never needed it, but you can do it if you go through for the distribution in CloudFormation and do all these kinds of things, then you can link the world that AppSync gives you to the one that you get them.

Jeremy: You just set it up as a custom origin or something like that.

Marcia: Something that, I have not done it so, but I have read the documentation is there if you need to do it. You just Google. I've seen custom domain, and it's the first thing that will appear. AWS information on how to do it.

Jeremy: And it's probably not necessary, like you said, with a normal, if it's your own app, maybe if you're exposing it to like a public GraphQL API, maybe you'd want to do that.

Marcia: Yeah, if you have Facebook or GitHub or something like that. Maybe you care about, most of the use cases at least I've been working with, and in general, people have been working in, it's just like you have this mobile app and you want to hook it to AppSync.

Jeremy: And most people aren't Facebook or GitHub, so you don't have to worry about it.

Marcia: Maybe here Facebook or GitHub, I don't know. Will you put it in AppSync? I don't know.

Jeremy: Probably not. Not if you're Microsoft or Google, probably. So last question, because this was something that I know has been asked before. So API Gateway handles stages, which is kind of a cool little feature. You can't do stages directly in AppSync, though, right? You just have to publish, like, a different version of it or different endpoint?

Marcia: Publish a different application that will have its own lifecycle. That's what I was saying when I was talking about the implementation to use the infrastructure as code, so then you can create multiple kind of stages, fake stages in a way that there are different versions, like different publish applications with the same configuration, just with maybe the name changed and some data sources changed and things like that, that's for each of the stages. So you have to fake it, in a way.

Jeremy: Awesome. All right, well, listen, this was fabulous. Thank you so much, Marcia, for being here and honestly, thank you for everything. Your FooBar videos and stuff like that are excellent. Great, great tools. If anybody's trying to learn serverless, definitely go check that out. So besides that, what else? How else can listeners get get ahold of you? I mean, you're pretty much out there, but...

Marcia: I'm very active [on] Twitter. So that's the place I write. So usually I'm there putting, some tweets, and also on Instagram. I use Instagram for conferences. I go there quite often, so I do a lot of live stories and share a lot of information as it comes out from the conference. So those are the two social networks I use besides YouTube. So you should find me there.

Jeremy: And that's @mavi888uy for Twitter and @foobar_codes for Instagram. And then you have your blog is marcia.dev and then if you go to YouTube, just search for FooBar serverless and you can't miss you.

Marcia: Yes, exactly. I picked the worst keyword in the world. FooBar.

Jeremy: But nobody ever uses that. What you talking about? Well, this has been awesome. I will get all that information in the show notes. Thank you so much.

Marcia: Thank you for having me.

View Details

About Nitzan Shapira

Nitzan Shapira is the co-founder and CEO at Epsagon, a distributed tracing product that provides automated monitoring and troubleshooting for modern applications. Nitzan writes for his own blog, as well as the Epsagon blog as a frequent contributor. You can find him speaking and helping out at serverless events across the globe, including Tel Aviv, where he recently organized the city’s June 4th ServerlessDays event. In addition to his contributions to the serverless community, Nitzan has more than 12 years of experience in programming, machine learning, cyber-security, and reverse engineering.

  • Email: nitzan@epsagon.com
  • Twitter: @nitzanshapira
  • Blog: epsagon.com/blog
  • Epsagon: epsagon.com

Transcript

Jeremy: Hi everyone. I'm Jeremy Daly and you're listening to Serverless Chats. This week, I'm chatting with Nitzan Shapira. Hey, Nitzan. Thanks for joining me.

Nitzan: Thanks for having me.

Jeremy: You are the CEO and co-founder of Epsagon, one of those hot serverless startups out of Israel. Why don’t you tell the listeners a little bit more about yourself and what Epsagon is up to.

Nitzan: Yes, definitely. As you mentioned I'm one of the founders and the CEO of Epsagon. I'm based out of Israel and San Francisco, currently kind of in between. I'm an engineer, a computer engineer with a background in cyber security and embedded systems. It's more low level background. In the recent years also, of course, [I've worked with] the cloud all the way to serverless. Epsagon is a company focused on monitoring and troubleshooting for modern applications. So the entire field of cloud applications that are built with microservices, serverless, managed services, where you don't have access to the host, very distributed — how do you understand what's going on in your production? How can you troubleshoot issues as fast as possible? Do it automatically and in a way that is suitable for this kind of modern environment. For example, using agents is something that you cannot do.

Jeremy: I wanted to talk to you about building resilient serverless applications. I think you have the right experience for this with what you do. But now that we're building serverless applications, and we're going beyond traditional applications as well as traditional microservices - if microservices can be considered traditional - you're starting to break things down into multiple functions. You obviously are using a lot of third-party services or managed services from the cloud provider. My question here to get us started is what is the main difference between a traditional application, whether server based or or container-based in microservices, and moving to this serverless environment?

Nitzan: Sure. I think the main difference is that a lot of the things are out of your control now, which is a good thing, because this is what you want when you go serverless. But on the other hand, you lose control over some of the things that are going on in your application. So when things don't go well, it can be very difficult to know where they broke. Then if you want to build something that's resilient, that's going to work in high scale, in very high reliability and without many surprises, you really have to think about all the different scenarios that can go wrong, which is not just my code had an exception. But maybe I got a timeout; I got an out of memory condition; I got a series of events that didn't go well - synchronous events, perhaps - and it seems that everything worked but actually didn't. How do I know about these problems, even if everything seems okay? The number of problems that can happen is just growing when you go serverless.

Jeremy: I think that makes a ton of sense. Why don't we dive into this and start talking about some of these individual problems or some of these differences, and maybe we can start with troubleshooting? What's different when you're troubleshooting a serverless application versus a more traditional, server-based application?

Nitzan: There are several key differences. The first one is that when you go serverless, you go distributed in a very significant way, more than with containers, for example, because those functions are kind of nanoservices. When you combine them together, we are seeing organizations with over 5,000 functions or more, which is just a very high number of nodes in the graph, if you look at it this way. It's very, very distributed. When something breaks, usually there are many more components involved in the chain of events, so it's going to be much more complicated to track what happened to find the cause of the problem. So distributed would be very important thing.

The other thing is that the new things that can go wrong. All those time outs, all those out of memory conditions, they happen all the time and [it's] very, very difficult to predict them. It's not something people are used to when they work with traditional services. And finally, the possibilities that you have as an engineer or DevOps to understand what's going on in your application is again more limited because you have no access to the host, so you can't install agents and so on. All you get is basically the basic logs and metrics that the cloud providers give you, which makes it even more difficult to know what's going on in the application layer and not just the simple metrics, because they are usually not going to be enough to troubleshoot a complicated problem.

Jeremy: Yeah, and I think with something like Lambda, or any function as a service, these are ephemeral compute. You have mini execution environments, or containers spinning up in the background, but those go away. You can't go back and look at the logs and see what that server did. And really the only logs available to you that are dumped to CloudWatch, for example, those are only there if your application actually sends logs. It's not logged automatically.

Nitzan: That's exactly the challenge, because once bad things happen, usually you didn't think about them before, and then you don't have the information that you're looking for in the log. Then you also don't have anywhere to connect to, to investigate, because, as you mentioned, it's ephemeral. That makes things very difficult because you can't think about everything that can possibly happen and put it in the log. On the other hand, you really have nowhere to go to after the thing happens. So you don't really have anything to do, just by using the logs. This is basically the conclusion.

Jeremy: Also, if you're using a number of remote services or managed services from the provider, where does the debugger go there? How do you see the flow of information? You have a lot of events. You have these highly event-driven applications with information flying all over the place. How do you keep track of that? Where do you see those logs?

Nitzan: Generally you don't see it, and that's the big challenge. This is, of course, why we are building a tool to help to help you. Generally speaking, the events that are going through the system are usually much more meaningful than the logs, from what we saw. If you actually know the events and data that is flowing between different components, that's going to tell a very good story of what happened from the request until the problem that happened. [That] can really help you troubleshoot the things that these events are not going to be in the log unless you specifically wrote it in the long, but usually, this is not the case. Getting those events is something that can really help.

Jeremy: Let's move on to the things that can go wrong in a service environment? Obviously, if you're using compute like Lambda Functions or Microsoft Azure functions, you have your normal code execution errors. You're going to get a "can't connect to a resource" or "can't parse a string," and those will be logged and those will be available to you. But when you start dealing with distributed systems and you're thinking about connecting to SNS topics or SQS queues or other types of managed service, what are the things that can go wrong there, in distributed systems in general, and maybe more specifically in serverless environments.

Nitzan: The things that are more specific are that, in the past, usually you had one big monolithic application, and when something went wrong, it would produce an error and something written to the log, and that would be pretty much the story of what happened. Now, when you are talking event-driven and distributed, in many cases, it's very asynchronous. For example, one function can perform perfectly fine, produce a message to some SNS, and then another function will get this message a little bit later, and then it will fail, but everything seemed fine. So the problem is actually that the message was not in the right contract, for example, between the two services. It's very difficult to to see what went wrong, because if you look at each function, everything seems right. I mean, this one did something right; the other one failed as it should have, but why was the message like this? Because the two teams that wrote these two services didn't actually coordinate together. Suddenly you have another thing that can go wrong, which is actually the agreement between how do we communicate, [and] how do we transfer messages and events between services. These things were not issues in the past because it was all just functions calling other functions, and you're in the same binary process. So this is something very new.

Jeremy: The other thing you have that's different is this idea of the retry behavior. I think most people are familiar with synchronous invocation of a resource where you make a call, it does something and then you get a response back, and maybe that's an error, and then you can deal with it there. But now, as you move to serverless, you start dealing with things like asynchronous or stream based processing, and with asynchronous certainly, your code that calls that resource doesn't know what happens. It just gets a response back that says all right, I got your event, and then something happens down the line. What's the impact of of this retry behavior on serverless applications and how we think about it?

Nitzan: One of the things I'm talking about in conferences, such as the [ServerlessDays] one we had in Boston and you invited me to, is the fact that these retries are something that [are] kind of considered as a good practice, by cloud provider to recover from error. For example, if the function fails, let's try to run it two more times and then see what happens. If it's an SMS message, we're going to run it two more time times as long as the message is new enough. That's something that is not really written in any programming book or software design book, but this is something that the architects of AWS thought would be a good idea. And it is a good idea sometimes, but for the developer, it can be very confusing. When it happens, usually it's very confusing because you just didn't know that this is the same invocation, running one or two more times, and when you have to think about it, it is very difficult to plan. This is where a concept such as idempotency come to action, when how are you supposed to write code that can run multiple times without having a bad effect or bad things happen. Eventually, it comes to the fact that people can't really plan an application that will be retried as many times as wanted with everything going right. It's basically kind of a constraint that you have to live with. You need to try and take it to your advantage when possible, but most of the time, I think many people would prefer to just go to do it the standard way. So don't try and run my code again without telling me, because I'm not sure what's gonna happen.

Jeremy: Yeah, and I think that the thing that's important that you mentioned about idempotency is that obviously if your transactions are getting tried or your events are being replayed multiple times by the cloud provider, your code has to deal with that. For certain transactions, it might not make that big of a deal. But if you are dealing with financial transactions, for example, you don't want those to retry or to submit the same, maybe, charge requests multiple times. If you think about the basics of the retry policies or how those work, the two times for an asynchronous Lambda event makes sense. If you're dealing with an SQS queue, then you have redrive policies that you can put in place so that the message will only be tried a few times, that way messages don't get stuck forever. But if we're thinking about the the redrive here, or maybe just the asynchronous invocation of a Lambda function, what do we do when that Lambda function gets tried three times and then it fails?

Nitzan: First of all, you need to know about it, which most people don't, because again, you have no indication, and the log is not going to tell you. So knowing the fact that something broke and it was retried is going to be very important when it actually happens. I think when you use a service, you need to know the properties and the limitation of that service. If you are writing a Lambda function that's triggered by a Kinesis stream, you need this Lambda to do string processing, because if it's doing something else with the data it's probably not the right thing. You need the Lambda to actually take the data, process it, and send it somewhere, then usually you wouldn't mind if it happens again. So I think it's possible to write microservices or nanoservices in Lambda functions or any other service for that matter that is contained enough so it will be able to handle retries. Then when you combine them together, in theory, it should work. The problem is that people just connect these services to each other without thinking, and they have hundreds of Lambdas, with many Kineses and SNS topics, and everything is running around, and it's not really working, as I suggested, of course, because people develop software fast. But if you had the time, you could actually plan every service to be working the right way.

Jeremy: If you're dealing with these failed events, obviously there's dead letter queues, as part of AWS at least, where you can put a dead letter queue or attach a dead letter queue to a Lambda function. If it's invoked asynchronously and it fails the three times, then that event goes into that queue there. And you can do the same thing with an SQS queue, for example, where if something fails after a certain number of times, your redrive policy will move that into a dead letter queue as well. Then, you have this issue where now you have dead letter queues or multiple queues with events living in them. You have to inspect those events. You have to set up alarms, so you know those events are in there, and then you potentially need a way to replay those events. So what about using something like Step Functions? What are the advantages of using a state machine or using something like AWS Step Functions?

Nitzan: Step Functions has several advantages. First of all, it has the advantage of being asynchronous, so you can actually have several functions - almost calling each other, but not really calling each other - but passing events asynchronously, so you wouldn't wait and pay for the accumulated running time for the functions. That would be one advantage. The second advantage is that it allows you to actually implement different rules and mechanisms in how the application is working that really it's a bit difficult to do without. You can say that if a certain event happened, only then you invoke this function or the other function, so eventually you don't really have to be coupled directly to the data. You can process it in different steps that will allow you to — I think this is a good example of resiliency, because using Step Functions the right way, can really scale very nicely because every step in the step machine will generate an event for the next step. So you don't have to worry about everything at once. You can kind of split your application logic into smaller steps that each one of them is much more likely to succeed. On the design level, anyway, this is how I look at it. You can use something that helps you split your logic in a very accurate way. You just decide exactly what I wanna do in each step.

Jeremy: I love Step Functions because it gives you that ability to do function composition, like you said. When you start thinking about individual functions or single functions that do one thing well, making those all talk to one another and creating the choreography for that is sort of a difficult thing to do. You start to introduce Step Functions into the equation, and now you have a Step Function acting like a traditional monolithic application where it can call subroutines and aggregate that information together and then sort of do something with it. I think Step Functions are certainly a really interesting way to solve that function composition problem. They do get expensive, which is is something to think about, depending on how you're designing your application and what level of control you need. But let's move on to something that you're very, very familiar with which would be monitoring a serverless application. How do we go about doing that? How do we monitor a serverless application?

Nitzan: Yeah, sure. Different ways. You can do it in a simple way, and in a more complex way, depends on the complexity of you application, of course. If you have just few functions - I would recommend using whatever AWS provides because it's already there. You have CloudWatch, so it will provide you with logs and metrics that will allow you to identify pretty quickly if something failed, and then go to the log and find out what happened. That, almost in 100% of the cases that we are seeing, is the first step. Then, the second step would be when you're going to a little more functions, and that I would say 20 or more, suddenly they start to get connected to each other. You start to create some kind of a distributed application. That's where the individual logs and metrics will not really tell the story, because they only provide information about individual components. Many people, at that point, will aggregate the log somewhere so they will just stream the logs into a log aggregation service, such as ELK or anything else. This would allow them to search in the logs and hopefully find problems faster. Then, eventually, you have a distributed application, and in order to really understand what's going on there, especially the more complex stuff, you need some kind of a distributed tracing technology. What is actually distributed tracing is basically to know how different services are sending messages to one another and how it's all connected from end to end. Some companies will implement some techniques of tracing in the logs, so you can have identifiers in the logs as we kind of go through. Then, you can search them in your log aggregation tool, so this would be probably the last step before using a dedicated solution for that. It can work pretty well, and we saw people do it in very high scale, with hundreds and thousands of functions. But at some point, there is also the question of how much time do you want to spend. It's going to take you a lot of time to implement different tools, especially based on logs. The whole point of serverless, of course, is developing fast. Eventually, if you're spending 30% or 50% of your developers time doing that, that's where we recommend considering an automated solution that will do as much of the work for you. Eventually, your hope as a developer, as a development manager, is that your developers will focus on building software that matters to your business and not building software that helps you monitor your business software. This would be when you get to a high scale; usually this is where people look for a solution.

Jeremy: You mentioned some of the tools that AWS has. So they have X-Ray that does some tracing; obviously, CloudWatch Logs does logging. Maybe explain the difference between those two things and why tracing is an important component.

Nitzan: Logging is pretty simple. It basically means usually logs is text. A log is a text file in some way or text data that is written either by the developer intentionally or produced automatically from some system that produces logs and then you get text. So textual data is very common. It's everywhere — but it's still text, so it's not structured. It's not even JSON. JSON, for example, is formatted. It's structured. You can say I want this field. I want this hierarchy. Logs are eventually going to be text-based files. Then using logs, you can do many things, right? Tracing is the way to, again, trace. What is to trace? Let's say you got an HTTP request. A trace will go through the lifetime of the requests, so we can go from one service to another service to the next service. That will be a trace that's going through my system and tells the story of what happened. The data of the trace is involved of what services, what data was transmitted, how it was transmitted, how much time did it take to transmit it from each point, and eventually the order of the events. This will be a trace that really can tell you what happened every point of the way, and you can put it on a timeline to identify bottleneck, to identify where you spend your time. One of the things that we are doing is actually we take the logs and we put them on the trace. For us, the log is just another type of data that, of course it's textual, but it's very useful. It's much more useful if it's in the right context, and on the right timeline. So you have five services. (25:26) You get trace data, you get log data, you get latency. Everything is like a story. So I would say the trace is structured and it's time-based. Log is textual data that somebody will have to structure in order to understand.

Jeremy: So how do frameworks like OpenTracing and OpenCensus help with all of this?

Nitzan: These frameworks are very useful as a standard way to write your tracing data. People said, "Okay, there is this concept called tracing. How can I standardize it so people won't have to invent it every time." OpenTracing, for example, will give you a standard way to create trace, to great spans, to create all those things that eventually tell you the story of what's going on in your system. Then you can implement, in your code, ways to send trace data in the Open Tracing format to some back and that will analyse this data, display it, provide information. These are just ways to standardize the traces and, for example, at Epsagon, we make sure that our traces are OpenTracing compatible. So if someone wants to add their own manual traces, they can easily do it without worrying about the format or, you know, people want to be compatible. Eventually, they always prefer to be compatible.

Jeremy: That makes a ton of sense. Let's talk about X-Ray again for a minute. [With] X-Ray, you go in, you instrument your code, then, as your functions run, it samples it, and you can see calls to databases or calls to other resources, and the latency involved there. But what about calls to an SQS queue and then the function that processes it and then that sends it somewhere else — that flow of data. Can X-Ray show you all that information?

Nitzan: You can do it to some extent. X-Ray will integrate pretty well with the AWS APIs inside the Lambda function, for example, and will tell you what kind of API calls you did. It's mostly for performance measurements, so you can understand how much time the DynamoDB putItem operation took or something of that sort. However, it doesn't try to go into the application layer and the data layer. So if information is passed from one function to another via an SNS message queue and then going into an S3, triggering another function - all this data layer is something that X-Ray doesn't look at because it's meant to measure performance. That's why it would not be able to connect asynchronous events going through multiple functions. Because again, this is not the tool's purpose. The purpose is to, again, measure performance and improve the performance of certain specific Lambda functions that you wanna optimize, for example.

Jeremy: You mentioned automation a few minutes ago, and I think that's a really important concept in terms of instrumenting your functions so that they do the proper tracing and logging. Obviously, in a more traditional application or a monolithic application, you might include your libraries and some of that stuff. But now we're talking about every single function needing to include this instrumentation. And that, in my opinion at least, is sort of a burden for developers to do that — but also pretty easy to forget. You know, say I gotta go back and add this, or maybe even it's a matter of which level of logging you've got switched on. What are some of the options for developers and for companies that want to create these policies to make sure that these are automatically instrumented? Obviously, there's Lambda Layers, which is a is a possibility. But what are some of the other options to auto-instrument functions so that the developers don't have to worry about it?

Nitzan: Yeah, by the way, it's not just worrying. It's not just the fact that you can forget. It's also just going to take you a certain amount of time - always - that you're going to basically waste instead of writing your own business software. Even if you do remember to do it every time, it's still going to take you some time. Some ways that can work [are] in embedded in your standard libraries that you work with. If you have a library that is commonly used to communicate between services, you want to embed that tracing information or extra information there, so it will always be there. This will kind of automate a lot of the work for you. That's just a matter of what type of tool do you use. If you use X-Ray you're still going have to do some kind of manual work. And it's fine, at first. The problem is when you suddenly grow from 100 functions to 1000 functions — that's where you're going to be probably a little bit annoyed or even lost, because it's going to be just a lot of work and doesn't seem like something that really scales. Anything manual doesn't really scale. This is why you use serverless, because you don't want to scale service manually.

Jeremy: And with Epsagon, you have a way to instrument the functions automatically, correct?

Nitzan: Yes, definitely. That's one of the things we do. We actually use Lambda layers, that you mentioned, that you can just do with probably less than a few minutes. You will be up and running with distributed maps of Epsagon, automatically, traced and produced, because we know how to add a layer to your functions through the Epsagon dashboard with one click. This layer goes and instruments the function. It produces all those events and traces from the code while it's running, and our backend can then identify how everything is connected. It's automatic in a way that, in many cases, you don't have to even change the code on your end. So that's very convenient for the developers.

Jeremy: Yeah, definitely, especially if you have 100 Lambda functions already written. You don't want to have to go back into every single one of those and add some new type of instrumentation. But anyway, Nitzan, thank you so much for being here. It is great that you are sharing all your knowledge with the serverless community, and Epsagon's doing a great job. If anybody wants to get in touch with you, how would they go about doing that?

Nitzan: Yeah, sure. First of all, my email is nitzan@epsagon.com. Very simple. I'm on Twitter. It's @NitzanShapira, my full name, and you can also just get in touch with me on LinkedIn. It's pretty easy. I'm very responsive. So, of course, you can check out the Epsagon website. I have a bunch of blog posts that I usually publish there as well.

Jeremy: Awesome. I will get all of that into the show notes. Thanks again, Nitzan.

View Details

About Alex DeBrie:

Alex is an Engineering Manager at Serverless, Inc., a blogger, and a big fan of serverless. He's held a variety of roles at Serverless, from Data Engineer to Head of Growth to Product Manager. He's also an important voice in the serverless community, creating in-depth guides (like DynamoDBGuide.com) to help users build the future with Serverless.

  • Blog: alexdebrie.com
  • Twitter: @alexbdebrie
  • DynamoDB Guide: dynamodbguide.com
  • Serverless, Inc.: serverless.com

Transcript:

Jeremy: Hi, everybody. I'm Jeremy Daly, and you are listening to Serverless Chats. This week, I'm chatting with Alex DeBrie. Hey, Alex. Thanks for joining me.

Alex: Hey Jeremy. Thanks for having me on.

Jeremy: So you are an engineering manager at Serverless Inc. ⁠— that's Serverless with a capital "S," not to get confused. They're out of San Francisco, but you actually work out of Omaha, Nebraska. So why don't you tell the listeners a little bit more about yourself and what Serverless Inc. is up to.

Alex: Yeah, sure. I've been at Serverless, Inc. for two years now. I started originally on the growth team, and now I'm working on the engineering team. But, you know, Serverless, Inc. we're the creators of the serverless framework, which is a tool for developing and managing serverless applications. It really reduces the tedium around setting up API gateway and IAM and all that stuff, and really helps you write your business logic and use AWS Lambda and serverless technologies effectively and quickly. There's a huge community of advocates, plug-ins, and best practices around the Serverless framework. I think we just crossed 30,000 stars on GitHub. So I'm really loving what we're doing here.

Jeremy: That's awesome. Yeah, I think if somebody doesn't know what the Serverless framework is yet, then they haven't been paying attention for the last couple of years. So you also write a blog, and you have a really, really good resource for people who are interested in learning DynamoDB, and people who are using DynamoDB and want to learn how to use it better. That's DynamoDBguide.com. That and your blog ⁠— what's going on with that stuff?

Alex: Let's start with DynamoDB Guide first. This was when I was still on the growth team at Serverless. I was doing a fair bit of content writing, and we were using DynamoDB a lot. I watched the 2017 re:Invent talk from Rick Houlihan, who's this wizard that works on DynamoDB at AWS. He did a talk on some best practices and I just loved it. I think I watched it four times in two or three weeks. This was Christmas break 2017, and I'm like, "I've just got to get some of this stuff out here." I wrote the resource that I wish I had when I started with Dynamo, because I thought I knew it well, and then I saw Rick teach it, and I did not. So DynamoDBGuide.com ⁠— it has a walk-through of all the different API stuff around DynamoDB, secondary indexes, all that stuff, as well as some data modeling examples too.

Jeremy: And your blog is mostly serverless and S3 batch stuff, all kinds of stuff like that, right?

Alex: Yeah, my blog I would say is a lot of, again, sort of like DynamoDB Guide, just the guides I wish I had when I started. I think both with DynamoDB Guide - and then a lot of the content on my blog - it's stuff I was familiar with, and then I want to teach it to people. And then when I teach it, I find I learned a lot of stuff that I actually didn't know. So it helps me, and I hope it helps other people as well.

Jeremy: Great blog, and the DynamoDB Guide is awesome. And yes, Rick Houlihan is a wizard and I don't know how he does some of things he does. But I have watched his 2018 podcast or the 2018 [re:Invent] has a podcast version that I think I listened to maybe 50 times on like .75 speed, so that you can maybe understand it.

So anyways, I wanted to have you on because I want to talk about this idea of serverless purity versus practicality, right? I think that we see a lot of debate on Twitter - and forget about what serverless is and what serverless isn't - but more so, what's the right way? How should we build a serverless app versus how we can practically build a serverless app? I think there's a lot of things around that, whether it's developer experience and that sort of stuff, but what are your thoughts on this sort of debate?

Alex: I think it's pretty fascinating to see. Like you say, if you're on Twitter and you're following a lot of the big time people doing serverless architectures in this space, they have a lot of great tips around best practices, and this is what you should be doing, all that stuff. But I find, as I'm building serverless applications or as I'm talking to customers and users that are building serverless applications, there are times when there's tension between what the best practices are and what their circumstances are. This could be because maybe they're not coming in with a green field application, or maybe they have a data model that doesn't fit DynamoDB or something like that. It's difficult on how you sort of square that with recommending something that you know isn't the best practice or the most pure serverless application, but you also gotta help people ship products, right? I think balancing that tension can be tough at times.

Jeremy: Yeah. I want to dive into a couple of these discussions that we've been seeing on Twitter, and I think, like you said, there are a few champions who sort of lead the effort for each one of these. But let's talk about the API Gateway service integrations. So we know that the typical serverless model would be API gateway to Lambda function, and then access something else. But it's possible to do that without using Lambda, right?

Alex: Correct. Like you're saying, you can do what's called an API Gateway service integration, where maybe you take that incoming HTP event, maybe you validate, authenticate it, maybe twist up the shape a little bit, and then you can put it directly into a different AWS service, like SNS, SQS, Kinesis, something like that, rather than going through a Lambda function first.

Jeremy: If you're just sending the data straight in, and you've got maybe a Lambda authorizer, that's one thing, but what about if you're transforming the data? That's seems like a different beast than writing some Node or some Python.

Alex: Yeah, absolutely. API Gateway allows you to write what are called VTL templates, so it's in Velocity Template Language, which I believe is an Apache project. It's a semi-declarative templating language where you can take some input, like a JSON payload body, the headers, all that stuff, and create a different shape that you want that satisfies the API format of whatever service you're integrating with. It's doable, but I would say not a lot of people have experience with VTL, so it's definitely a learning curve there. It has some quirks and unexpected stuff for people that are new to VTL.

Jeremy: You mentioned the quirks. I'm thinking to myself, here I am writing an application. I've got all my tooling in place. I've got my testing frameworks and I can test all this stuff. Then I say, okay, well, I'm gonna I'm gonna go pure - API Gateway to DynamoDB - and I'm going to write some transformations in VTL. I do that, and then how do I test that?

Alex: That's the tricky question, right? You probably have to deploy your application up to API Gateway, then send in HTTP request, and then check the DynamoDB table to make sure it got there all right. You're probably gonna have more of a a cloud native integration test suite than a local unit test suite that you could go through to validate some of that logic.

Jeremy: But is that a mental burden on developers? I mean, are there trade offs? What's the benefit of doing it versus just saying, look, I can write the transformation in Lambda because I know that. But if I move over to using these VTL templates and things, and I have a hard time testing it, I've got to test it in the cloud ⁠— should I be nervous about making changes or things like that. What are your thoughts on that?

Alex: Great question. To me, I think it does add a fair bit of burden around the development process. I think you really got to balance what your needs are and what your comfort level is with VTL versus running something in Lambda and having full test coverage. To me, my rule of thumb is if you're not doing any sophisticated transformation or if you're not doing any fine-grain authorization, I think it's fine to to use API service integrations. The example that I go to is maybe you have a front end or an IoT device or something like that that's just sending in data and you're only validating the shape of it, or maybe transforming the keys around a little bit, before sending it into SNS and Kinesis. Then, it's going to get processed by a different system. I think that's a totally valid use of API Gateway service integrations, if you want to use it.

On the other end of the spectrum, you mentioned connecting with DynamoDB, and you can write directly to DynamoDB there. I don't love it for more sophisticated use cases where maybe you need to pull off a key from the incoming JSON body. Maybe that's going to be your partition key for the DynamoDB table that you're querying. Maybe you're checking a sub-property on there to make sure this user has access to this thing and rejecting, if not, or if you're reconstructing a complex JSON object after you retrieve the DynamoDB item. Any of that stuff where it gets more complicated, my rule of thumb generally is if there's an if statement or a for loop in your VTL, now it's code. It's not config anymore, and you probably should do it in a Lambda rather than in VTL.

Jeremy: I totally agree with you on that. I love the idea of service integrations, but I think if you start making developers do that mental shift even more so than just moving to serverless, you start to introduce resistance. I don't know if that's the best word. But anyways, so now, we've got maybe a judge's ruling on this. If somebody says, hey, look, yes, I could use the service integration, but I just feel more comfortable using code to do it, what would you say to those people?

Alex: I don't judge those people at all that, because that's mostly where I am on things. I think API Gateway service integrations are interesting and useful in some aspects, but I don't necessarily think it's a best practice, that if you're not doing it, you're doing something wrong. I think it's a choice you can make and and it depends on your circumstances. How often is that code going to be changing? How much control do you need over error handling and messages that you're returning to the client that's calling you, things like that that would determine whether you should use it.

Jeremy: Great. Let's move on to this idea of monolithic Lambda functions versus single purpose Lambda functions. Obviously the AWS best practice is single purpose Lambda function — do something very, very simple. Do one thing well. But I think you and I both know from seeing developers move services to serverless, the most common use case probably is transport an Express app, and have 50 routes in there. What are your thoughts on this?

Alex: I agree generally with the single purpose Lambda function best practice. I think that's that's generally the best way to go, but I think there are two strong exceptions to that. The first one you mentioned is just directly porting over an Express app. It could be porting over an existing Express app, or it could be "I'm very familiar with Express, and and that's what I want to use. That's how I'm productive, but this is the easiest way for me to host Express." This is easier and cheaper even than Heroku or something like that. The first sort of bucket of use cases I would say, where it's OK to have a monolithic Lambda function, are those people that want to run Express, or if you're running Python, maybe you want to run Flask, and handle all your routing within a single function, and you get a pretty great local development experience, but you also get the scaling and easy deployments of the serverless experience as well. I think that's a valid use case for some people.

I think it's also a great on-ramp to serverless because usually what happens is you start with that Express or that Flask app where you're serving Web APIs. Then you decide "I need some background processing, so in addition to this Express app, I'm going to have an SNS or an SQS integration," or "now I need some stream processing, I'm going to use Kinesis." As you do that, you become more familiar with the Lambda model. Now you have all your routes in the the Express app, but you also have these other functions and events that are single purpose, and you start to learn how that works, and then maybe your next project that needs HTTP end points, maybe you reach for a native Lambda single purpose function, rather than going for something like Express or Flask.

Jeremy: You also are going to have the problem too, I think, one of the issues we run into - when you start to build complex serverless apps that that deploy multiple functions with multiple end points and other services that are interacting with them - is you start running into CloudFormation limits as well.

Alex: Yeah, absolutely, and that's the second use case where I recommend people put multiple end points into a single function. CloudFormation only allows you to have 200 resources in a single stack. It may sound like a lot — you're like, how am I ever going to have 200 end points. But the reality is you're not going to get 200 endpoints, because if you ever hook up an endpoint using Lambda on API gateway, it's going to create five resources for each endpoint unit created. It's going to create the function; it's going to create the function version; it's going to create the log group, and it's going to create the API gateway method and path or resource. For every single function you get, you're going to get five resources. Now you're talking max 40-or-so resources or 40-or-so endpoints, but also you got to think about additional things that happen in your service, like DynamoDB tables, SQS queues, SNS topics, other infrastructure as well, so you're probably not going to get even 30 or 35 end points.

Jeremy: How would multiple single purpose functions fit into a concept of a microservice? We often hear the term nanoservice, which I don't really like that term at all, but you have multiple functions working together to form some sort of microservice. What are your thoughts on how you would set that up with serverless?

Alex: I think a lot of the same general microservice principles apply. Find out where your bounded contexts are, and that's where you can split things up. It's probably not one function. It's not the one function nanoservice that you're talking about, but it's probably a whole set of CRUD functionality around a particular resource, and maybe even two or three resources that live together and interact pretty closely with each other, are related to each other. Anything that touches the same databases or shared infrastructure might be in a service. Anything that the objects really relate to each other pretty closely, I would split those into a service. If you can do that and stay under the 200 resource limit and CloudFormation while having a function for every endpoint, I think that's great. If you can't, that's when you start to look into other options of putting it all together into one function.

Jeremy: I really like the Serverless framework, putting everything together within a single microservice, because then you test it altogether. It just makes it a little bit easier to reason about how these services his work together, as opposed to just deploying 30 functions from one serverless YAML file that you have no idea or that they don't necessarily work together or work in concert. That's for me anyways.

All right, let's see if we can make Paul Johnston's ears ring and talk about using relational databases with serverless. That's another one of these topics where you've got relational databases are already RDBMS versus something like DynamoDB. Obviously we've got Aurora Serverless now, which again is very flexible. It grows, but it's still resource limitations, and they've introduced the Data API so a couple of different ways that we can kind of interact with that. If you listen to Rick Houlihan, and you follow the Church of Rick, you believe that DynamoDB is the future, at least for OLTP apps or DSS, and that these are the kind of things where if you can fit your model into it, it seems to make sense. What are your thoughts of relational versus DynamoDB and how it fits into the serverless world?

Alex: As we talked about the beginning, I made the DynamoDB guide. I'm very much a believer and lover of DynamoDB. I think it's just a great tool. Specifically, I think it just works so well with AWS Lambda, right? With AWS Lambda, you have this world of what I call hyper-ephemeral compute, where your compute can scale up to 1,000 invocations in a minute or can scale that back down to zero just as quickly. Something like that doesn't work very well with a relational database where you need to set up a persistent TCP connection and maintain that connection, and in your database, probably has a maximum number of connection limits — connections that can be established at any time. If you scale up to 1,000 instances of your Lambda function, all trying to hit your database, you're going to run into limits, and that's going to be tough to debug. Now you have to set up something like a PG bouncer, like a connection pool, between your Lambda Functions and you're already RDBMS.

The additional problem, too, with these sort of more server-full databases like Postgres or MySQL, or anything like that, is often you want to have those network partitioned where they're not accessible to the public Internet. Now you need to put them in a VPC, which means your Lambda needs to be in a VPC, which means at least for right now, there's a VPC cold start penalty of multiple seconds. It's something that your users are really going to notice if you get that. I think the connection model, and that model of RDBMS does not work well with Lambda, and DynamoDB does work really well. And that's what I love about DynamoDB, HTTP connection, IAM authorization, very, very high scale ability, potential if you need it. But the data modeling aspect is the tricky part in there.

Jeremy: Let's talk about data modeling for a minute, because I can take pretty much any entity relationship model and normalize it, go to 3NF or whatever and I can do that. I can visualize it in my head. I can probably do it without even writing it down, and I could do that. I think that most people who are designing and building databases or building data models, they understand that as well. But when you can't do third normal form anymore, and we're starting to de-normalize that data across all this stuff, there's a learning curve. It's like a giant sudoku puzzle sometimes trying to figure these things out. Does that outweigh the benefit of using DynamoDB? Does it outweigh that trade off on that learning curve?

Alex: That's a great question. I think that's really the crux of this whole issue. I think there are two issues there. One is that short-to-medium term learning curve that you're talking about, and it's a steep learning curve, because you really got to figure that out. It's in such a sensitive area: your data. You don't want to lose data, or having to perform a migration down the road is costly. You really want to do it right the first time, but you often don't have enough experience the first time, so that can be pretty tricky. I think there's that first issue of just learning it. To me, I think it's worth putting in the time to learn it, maybe using it in some smaller areas first, really getting a feel for it, and then figure out how you can use it in more areas down the road.

The second issue with DynamoDB - and this is a more persistent one - is DynamoDB is great, if you know your access patterns and they're not going to change. You know all the ways you're going to read from your database. You know all the ways you're gonna write to your database, and that's going to stay persistent over time. That's going to scale up as as high as you want it to go. The difficulty if you're bringing a new product to market or something like that, your data model is probably shifting as you're adding your features. You're adding new entities. You're changing how you're querying, how you're filtering all that stuff, and DynamoDB is not well suited for changing your access patterns down the road. It's schema-less, but schema-less doesn't mean free-flowing, "do everything you want," because you really got to think about how am I going to access this data? If that changes, you're in big trouble. I think that's one area that's that's still pretty tricky. Serverless is so great for rapid experimentation, really shipping quickly, changing stuff, and focusing on business value. But then, you have a database that's locked in once you've set up your core data model, and how do you both evolve your data model with your app?

Jeremy: I think that for most people, if you start talking about overloaded indexes and adjacency lists and things like that, it gets a little bit confusing. I totally agree with you on the changing access patterns. It's just that I look at this now and I say every time I want to do something, small app - and again we're not talking about a full-blown application with 50 different access patterns - we're talking about maybe a small microservice that does something very specific. There may be 10 access patterns or five access patterns. I look at something like that and I say, do I really want to set up a database for this? It seems like it's possible to store this stuff with all the magic that we've learned from past from past re:Invents. All that magic is possible. We can do these hierarchies. We can do these one-to-many joins, and many-to-many joins, or not really joins, but we can represent these relationships in a DynamoDB.

Is it worth it to take the time to learn these these skills so that when you have that next project maybe you do know the access patterns or that they're going to be relatively simple? Does it make sense to do that? Should this be the de facto? Should DynamoDB be your de facto database, if you're developing a service application on AWS?

Alex: I'm a strong proponent of yes there. I mean, I wrote the DynamoDB guide. I love it, so I think yes. I think it fits so many things about serverless that it's just a nice fit with Lambda. Both the pricing model, the connection model ⁠— everything I think really works well with the service application. You're absolutely right that there is the learning curve. But I think those type of use cases that you mentioned, where you have a microservice and it only has 5 to 10 access patterns, maybe you don't even need a secondary index, or maybe you just need one. That's really the perfect use case for Lambda, where you've got something isolated, just a few types of entities (you can have as many rows those you want). I think that's a great fit for Dynamo and for serverless applications.

Jeremy: I think the other piece of it too is I've had people say "well, yeah, but I can't do counts. I can't do aggregations. I can't do those sort of things." I've always found it very, very easy - whether you're doing scans of the database - to dump that on a regular basis. Or now that you can use DynamoDB streams, you have the ability to replicate that data and either aggregate it yourself as part of that calculation or just dump it into another database. I've done that before, where I've taken Dynamo and dumped it into SQL, or into MySQL, so that you could actually do some magic on the back end. In terms of my app being able to scale and access those things, my users don't need to run complex queries. They just need to get data, put data, maybe see a list of some data that's associated with this, and that, to me is a very, very good use case for Dynamo. With all the other services around that, it seems crazy to me to start with the assumption that you need a relational database. But again, I don't know. I always was on the other side, and now I think I've become a convert, but only because I've seen what's possible.

Alex: True. One thing I want to mention that I just found out today - I think this is so cool - but you can actually copy data straight from DynamoDB into a Redshift cluster. If you're using red shift for analytics, which, at my last company, I did a lot of that, you can go into that Redshift database, and run a copy command. It is fully managed, pulls out all the data from DynamoDB, puts it into a table on Redshift, and now you can query on it, which I think is really cool. You don't need to mess with streams or anything like that. It pulls it all out for you.

Jeremy: That is very, very cool feature. Okay, let's move on to one more topic here. This has to do with optimizing your Lambda functions, or optimizing your serverless apps in general. Obviously you see a lot of people saying, "oh, my Lambda bill was 18 cents this month." I've talked to other people who said, well, if that's the case, then maybe you're not running an enterprise serverless application, if it's only 18 cents. Is there such a thing as premature optimization? Are we trying to make our package sizes smaller? Do we go through all this extra effort? Ben Kehoe, for example, said a number of times that the code that they used for iRobot, it's not worth it for them to optimize that code, because it runs just fine and it doesn't cost a lot of money. Of course, they are enterprise. What are your thoughts on this idea of optimizations?

Alex: That's a great question. That's interesting to hear from Ben Kehoe. He's a great guy to hear from. I think it depends on your use case. It depends on how your application is being used. Ben Kehoe works for iRobot, and it might not be a lot of user-facing stuff, where maybe the robot vacuum is sending up data. Or maybe they're processing data in a background, offline fashion, and it's not like a user-facing thing where a user's going to notice any latency. But on the other side, someone like Brian Leroux, who I really like, he's the creator of the architect framework. He's vigilant on package size, and he was saying the other day that their CI/CD process actually fails the build if their package size is over 5 MB, which I think is really interesting. His point there is that the bigger your package size, the longer your function cold start is going to be, because now it's got to load all that code into memory before it starts executing, and it's going to take it a while. If you're building a user-facing application, now that's something you think about where it's not going to be as snappy for your users. It's something to think about, for sure.

In terms of purity versus practicality there, you need to think about your use cases and what matters to you. If you're not gonna have a user-facing application, I wouldn't worry that much about optimizing it, or if your bill's not that high right now, don't worry about optimizing it. Most importantly - I think this is true of serverless or non-serverless, but I think it's been a focus in the serverless community - focus on building a product that brings value. If speed is something that brings value to your customers because they want a quick, responsive app, then maybe focus on speed. Otherwise, focus on building those features in that core experience that your users are really going to care about. Focus on that first rather than some of the optimization techniques.

Jeremy: I don't think a cold-start latency every once in a while is going to be what puts the nail in the coffin of your application.

Alex: I hope not.

Jeremy: It's likely going to be that you don't get to market fast enough or you don't iterate on it enough. I know me personally; I'm the worst when it comes to front-end. I'm not a typical front-end developer, but I do a couple things here and there. I maintain a React app and I do some of these things. I will spend hours just trying to get something to align right sometimes it seems. It's an incredible waste of time, and the optimization there really isn't worth it. And I liken that to the back-end of building serverless applications.

I think it goes back to what we originally talking about with the service integrations. If you're getting started and you're building an app that you're testing, you're trying to get out there fast, then maybe all of these things we highlighted - build a monolithic Lambda function, if that's easiest way to do it; use Lambda; don't worry about the service integrations, if it's not going to cost you that much more, especially if the app doesn't have a ton of traffic; use a relational database if you need to, if that's the quickest way for you to build an application and model it and figure out what your access patterns are going to be - do that. And you know what? If you're using a bunch of dependencies in order to make the app run, do it. I just want people to jump into serverless and start using it, and then to figure out these things later.

Maybe we just wrap this up on this topic. You hear a lot of people talking about "we should be using service integrations." I don't know if they're saying that you have to use service integrations, but they're saying this is a good way to do it, and let's use DynamoDB because that's a good way to do it. Lately, I've been sort of guilty of that myself. And oh, [they say] "let's make the package sizes smaller; let's use these single purpose functions." What is your advice to the developer who is getting started with serverless. What is your advice to them when they hear all this noise about all these best practices and stuff?

Alex: I'd say jump in. You can you can worry about the best practices and all that stuff, but really, you've got to jump in, start building some stuff, and figure it out. Do a little research on best practices, but don't take it as gospel, because you'll build some stuff. You'll figure out what works for you and what doesn't, and you'll get better over time. Ben Kehoe often promotes that there's a serverless spectrum, or even serverless as a ladder, where you start at just pulling off little pieces of your app and putting them in Lambda or using some managed services. Over time, you start to bite off more and more of that serverless mindset and get into the serverless ecosystem. I think that's true here, too. You don't need to go whole hog the very first time you build a serverless application. Get the core of the benefits, which is the easy deployments, the scaling, the pricing model, the time to market really, and the total cost of ownership. Then, you can get more and more pure as you go.

Jeremy: Well, I don't think we can add much to that. So let's wrap it up. Alex, thank you so much for joining me and sharing all of your serverless knowledge with the community. If people want to find out more about you, how do they do that?

Alex: Sure, Jeremy. Thanks for having me. You can find me at Twitter. I'm @alexbdebrie. My blog, as Jeremy mentioned at the beginning, is https://www.alexdebrie.com/. I've also got my email there, if you want to email me, my Twitter DMs are open. I'm always happy to hear from anyone that has questions in the community.

Jeremy: And you are on GitHub (https://github.com/alexdebrie) and LinkedIn and all of those other social platforms.

Alex: I'm findable.

Jeremy: Great. We'll put all that in the show notes. Thanks again.

Alex: Sounds great. Thanks, Jeremy.