The brutal truth about digital performance engineering and operations.
Andreas (aka Andi) Grabner and Brian Wilson are veterans of the digital performance world. Combined they have seen too many applications not scaling and performing up to expectations. With more rapid deployment models made possible through continuous delivery and a mentality shift sparked by DevOps they feel it’s time to share their stories. In each episode, they and their guests discuss different topics concerning performance, ranging from common performance problems for specific technology platforms to best practices in development, testing, deploying and monitoring software performance and user experience. Be prepared to learn a lot about metrics.
Andi & Brian both work at Dynatrace, where they get to witness more real world customer performance issues than they can TPS report at.
AI coding tools are everywhere, but how do you prove they're actually making engineers more productive?
In this episode of PurePerformance, hosts Brian Wilson and Andi Grabner welcome Michael Reichenbach, Platform Engineer at 1KOMMA5°, to discuss the company's journey toward more than 80% AI-written code. Rather than relying on anecdotes or hype, Michael shares how his team designed a real experiment to measure the impact of AI-assisted development.
We explore the metrics they chose, why traditional DORA metrics such as deployment frequency and change failure rate were not the right indicators, and how they instead focused on "Time to Code" from ticket creation and first commit to merged pull request. Michael also explains how AI lowered the barrier for contribution across the organization, enabling even non-engineering teams to prototype and build solutions faster.
The conversation also dives into the operational side of AI adoption, including AI observability dashboards, budget controls, Slack alerts, usage monitoring, and the surprising decision to intentionally limit AI spending during their proof of concept.
Whether you're evaluating Cursor, GitHub Copilot, or other AI coding tools, this episode offers practical lessons on measuring value, maintaining quality, and scaling AI adoption responsibly.
Links we discussed
Michael's LinkedIn: https://www.linkedin.com/in/michael-reichenbach/
Klaus's LinkedIn: https://www.linkedin.com/in/langenheldt/
Talk at Cloud Native Munich: https://www.youtube.com/watch?v=Ntf0h0vFuMQ
1Komm5 Website: https://1komma5.com/
Kenote from KubeCon: https://youtu.be/P1phxZHJGrA?t=570&is=9DYnbK8VGorMmaXo
Michael's YouTube Playlist: https://youtube.com/playlist?list=PLn-u2xOcMlXlVweZ0aB4pu6VM6Xv6895i&si=MZ6jzF-VwIAY3HFF
"There is no single way to deploy OpenTelemetry at scale—and that’s exactly the challenge."
As organizations adopt OTel across teams and environments, they face tough questions around standardization, configuration, and operating resilient observability pipelines.
To address these challenges, the OpenTelemetry community has introduced Blueprints and Reference Implementations—practical guidance on topics like data standards, consistent agent and collector configuration, pipeline resilience, and intelligent sampling.
In this episode, we’re joined by Dan Gomez Blanco, maintainer of the OpenTelemetry End-User SIG, to explore real-world reference architectures from organizations like Skyscanner, Adobe, and Mastodon.
Tune in to learn how the community is turning OTel complexity into shared best practices—and how you can contribute your own blueprint
I still hear people say, “OpenTelemetry is vendor-neutral, so you can switch any time!”
In this episode, Adriana Villela and Josh Lee (both active OpenTelemetry contributors) help bust that myth.
While OTel standardizes instrumentation and signal transport—and unlocks a rich ecosystem of tools—switching vendors isn’t as simple as it sounds. There’s real cost in retraining engineers, migrating dashboards, SLOs, and alerts, and reworking deep integrations across your delivery pipeline.
We also dive into a key challenge the community is tackling: helping engineers instrument by value, not by default—making it easier to capture the right signals with high quality instead of just collecting everything.
Here the links we discussed:
Adriana's LinkedIn: https://www.linkedin.com/in/adrianavillela/
Josh's LinkedIn: https://www.linkedin.com/in/joshuamlee/
The blog article: https://thenewstack.io/opentelemetry-vendor-neutrality-guide/
CND Austria Talk: https://www.youtube.com/watch?v=1gxLseuaTdM
KCD Prague Talk: https://www.youtube.com/watch?v=pPXG20CXKxQ
OpenTelemetry Project Website: https://opentelemetry.io/
In this episode, we explore how AI is transforming education, from classrooms to corporate training. What changes are needed in schools and universities? How does AI affect both students and educators? And how should companies rethink internal training and hiring to stay competitive?
To answer these questions, we’re joined by Rainer Stropek, CEO of Software Architects and Chairman of Coding Club Linz. With decades of experience teaching at high schools and universities—and helping organizations upskill their engineers—Rainer brings a unique perspective on how software engineering education is evolving.
While many view AI as a threat, Rainer sees it as a “Christmas gift”—opening up endless opportunities to learn, adapt, and innovate.
Tune in to hear why curiosity is more important than ever, how educational institutions can prepare future engineers, and why organizations must step up to ensure everyone has a fair chance to succeed in the age of AI.
Links we discussed
Rainer's LinkedIn Profile: https://www.linkedin.com/in/rainerstropek/
Rainer's Website: https://rainerstropek.me/
CodeClub: https://codeclub.org/en/
Coder DoJo Linz: https://linz.coderdojo.net/
Its rare - but it happens: A guest-free episode of PurePerformance, allowing Andi Grabner and Brian Wilson reconnect to share real-world insights from recent months in the cloud-native and observability space. From KubeCon Amsterdam experiences and the strength of open-source collaboration to emerging challenges like AI-generated contributions, they explore how the industry is evolving beyond the hype.
Your co-hosts of PurePerformance discuss the changing role of observability in the AI-native era—both as a foundation for understanding complex systems and as a tool to monitor AI itself. Brian shares his personal shift from AI skepticism to practical adoption, highlighting how AI can significantly improve productivity when used thoughtfully.
Hope you all enjoy this episode!
In 2011 Heroku defined the 12 factor app to remove emerging bottlenecks as developers tried to scale their output when they moved from building monoliths to microservices. In Platform Engineer we see a repeating pattern called the "8 Factor Platform Producers". AI allows engineering teams to speed up but they face bottlenecks as platform capabilities are not scaling with that demand as they are often depending on a central platform engineering team to be built and maintained.
To learn more about 8 Factor Platform Producers we invited Abby Bangser, Founding Principal Engineer at Syntasso and CNCF Ambassador. She gave an amazing talk at KubeCon in Amsterdam and today walks us through the need of defining both consumers and producers for platforms to eliminate any emerging bottlenecks in Platform Engineering and allow an organization to reap the benefit of speeding up with AI
Links we discussed:
Abby's LinkedIn: https://www.linkedin.com/in/abbybangser/
Abby's Kubecon Keynote: https://www.youtube.com/watch?v=8t0-5cvvMGM&list=PLj6h78yzYM2MXCOWSN9CqqID6OOvF7wxL&index=30
12 Factor Apps: https://12factor.net/
CNCF Whitepaper: https://cloudnativeplatforms.com/whitepapers/platforms/
As the software world is transforming from cloud native to AI-native, observability must transform with it. But how exactly? How do we apply this in an existing enterprise with established processes and practices?
In this PurePerformance episode, Andi Grabner hosts Hilliary Lipsig and Rob Rati to discuss their new book, Observability in the AI‑Native Era. The conversation explores how AIOps, automation, and modern observability must evolve as systems become cloud‑native, data‑heavy, and AI‑driven.
We talk about why old alerting and SLO models no longer scale, how to balance AI with automation and human judgment, and why trust, security, and compliance matter more than ever when machines start making operational decisions. A must‑listen for SREs, platform engineers, and engineering leaders navigating the AI‑native future.
Links we discussed
Book on Amazon: https://www.amazon.com/Observability-AI-Native-Era-Artificial-Intelligence-ebook/dp/B0GHZH1YFL
Hilliary LinkedIn: https://www.linkedin.com/in/hilliary-lipsig-a5935245/
Rob LinkedIn: https://www.linkedin.com/in/roberthrati/
Andi LinkedIn: https://www.linkedin.com/in/grabnerandi/
AI coding agents are fast—but speed alone doesn’t guarantee quality. In this episode, Andi Grabner talks with Lukas Holzer (Straion) about why large context files and “almost right” AI code create new risks for engineering teams. You will learn about the "Lost in the Middle Syndrom" and why many organizations are not getting the promised 10x engineering boost right now!
Andi and Lukas also explore rule adherence, dynamic context generation, enterprise readiness for AI-first development, and how software engineering roles are evolving in the age of AI.
Tune in to learn more ...
Links we discussed
LinkedIn Profile: https://www.linkedin.com/in/lukas-holzer/
Straion Website: https://straion.com/
90 Percent Rule Blog: https://straion.com/blog/90-percent-rule-adherence-straion-coding-agents/
1million tokens Blog: https://straion.com/blog/1m-tokens-wont-save-your-engineering-standards/
In this episode of the PurePerformance Podcast, Andi and Brian sit down with Chris LaBrado—Solutions Architect for AI Enablement, FSO, SRE, and ITSM at HSN/QVC, where he has spent an incredible 27 years shaping operational excellence. Their conversation dives deep into how AI is transforming software creation, enterprise workflows, and even the very role of developers.
Chris shares how the barrier to entry for building tools and automation has dropped overnight thanks to natural‑language-based development: “Everyone can now create automation or tools without having to worry about the syntax.” He explains why AI is rapidly becoming the primary interface into the enterprise—capable of navigating presentations, emails, and complex back‑office systems—and why the future of engineering may shift from human‑oriented coding to AI-driven development models such as MDCD (MarkDown Continuous Development).
The discussion also takes unexpected but fascinating detours into Chris’s background as a former bowling‑industry podcaster, his recent work with generative agents like DynaClaude, his Vibe Coded Root Cause Agent, and a philosophical exploration of AI, creativity, and the concept of singularity.
Amidst all the change, Chris remains optimistic: “AI opens up a lot of new opportunity for everyone willing to adapt. It will result in us creating more things that ultimately help us as humans.” This episode is a thoughtful, energizing look at where software engineering is headed—and why the future might be brighter than we think.
Links we discussed
Chris LaBrado on LinkedIn: https://www.linkedin.com/in/chrislabrado/
Mo Gawdat, former Google Executive on the Singularity "moment of truth": https://x.com/vitrupo/status/2008824930646057380?s=20
CEO of NVIDIA had an interesting excerpt from interview: https://x.com/MinusWells/status/2031974516155695414?s=20
Elon Musk on speed of AI: https://x.com/r0ck3t23/status/2031639621465931903?s=20
AI brain emulation of a fly (e.g. "a sign of the times"): https://x.com/alexwg/status/2030217301929132323?s=20
Elon on fiat currency transforming based on AI manufacturing loop: https://x.com/elonmusk/status/2020202496547844312?s=20
Fiat currency moves to model based on thermodynamics: https://x.com/r0ck3t23/status/2033371028202602547?s=20
In this episode, Andi and Brian welcome back Adam Tornhill—founder of CodeScene and author of Your Code as a Crime Scene—to explore how agentic AI is reshaping software engineering. Adam shares his personal journey from 40 years of hands-on coding to orchestrating AI-generated code, and what this shift really means for development teams.
Together, they dive into new research on the hidden risks of AI-assisted coding, why low-quality or legacy code slows AI down, and how to measure the “AI-readiness” of a codebase. Adam breaks down practical strategies from his latest work on Agentic AI Coding, including guardrails, refactoring patterns, enforced processes, and why test coverage has become a surprising cornerstone for safe, fast AI iteration.
Whether you're experimenting with AI coding tools or planning enterprise-scale adoption, this episode delivers actionable guidance rooted in data, engineering discipline, and real-world experience.
Links
https://codescene.com/blog/agentic-ai-coding-best-practice-patterns-for-speed-with-quality
https://codescene.com/blog/strengthening-the-inner-developer-loop-turn-ai-into-a-reliable-engineering-partner
AI is transforming software engineering—faster than many teams can adapt. In this episode, Andi talks with Wolfgang Heider and Benedict Evert about what it really means to build “AI‑native” software, where prototypes turn into production apps in minutes.
We explore why good engineering fundamentals still matter, how multi‑agent workflows mirror traditional roles, and why testing, governance, and clarity of intent become more important—not less.
We also discuss the future of junior engineers, the risk of everyone reinventing the same solution, and why value—not code generation—is becoming the real differentiator.
Links we discussed
https://www.linkedin.com/posts/wolfgangheider_productmanagement-softwareengineering-ai-activity-7425746505883607042-D1OZ
https://www.linkedin.com/pulse/machines-making-wolfgang-heider-5mvsf
https://www.linkedin.com/pulse/i-built-app-between-final-stranger-things-episodes-wolfgang-heider-5penf/
https://futurelab.studio/ora/
https://futurelab.studio/htmlctl/
Why do we still struggle with resilience in 2026? Is it the growing complexity of systems, the pressure to ship fast, or a lack of education around resilient design? In this episode we welcome Adrian Hornsby from Resilium Labs to explore these questions and learn about chaos, complexity, and the importance of continuous learning!
Adrian has learned his chaos engineering skills while working at AWS for many years. He shares insights from his upcoming book and his experience helping organizations embrace resilience as a continuous learning practice. We discuss:
* Why traditional chaos engineering assumptions break down when AI starts writing your code.
* The rise of AI-powered SRE agents—are they a blessing or a missed learning opportunity?
* Organizational challenges and the importance of tracking near misses.
Links we discussed
Adrians LinkedIn: https://www.linkedin.com/in/adhorn/
Resilium Labs: https://www.resiliumlabs.com/
Upcoming Book: https://leanpub.com/whywestillsuckatresilience
Contributing to Open Source is easier than ever - especially because contributions are needed for documentation, demos, tutorials and code. But how to get started? Where to look for "first good issues"? Is everyone welcome? What are the prerequisites?
Tune in and hear from Diana Todea, Developer Experience Engineer at Victoria Metrics, on how within a year she made it from Zero to Developer and receiving the Contributor Award for OpenTelemetry 2025 at KubeCon Atlanta. Diana shares her journey, how she started, how she found the right topic and how she keeps herself motivated. Diana is also the Co-lead of the Neurodiversity CNCF Working Group and gives us insights into the Merge Forward community.
And don't forget: Call for Papers for Cloud Native Days Romania and Austria are open and both Diana and Andi would be glad to see your proposals!
So - what are you waiting for?
Links we discussed:
Diana's LinkedIn: https://www.linkedin.com/in/diana-todea-b2a79968/
From Zero to Developer Talk: https://www.youtube.com/watch?v=nPrxpEE5GpY
Contributor Award: https://siliconangle.com/2025/11/13/accessibility-meets-open-source-collaboration-kubeconna/
Her latest CNCF Blog Post: https://www.cncf.io/blog/2025/12/04/my-first-kubecon-cloudnativecon-a-journey-through-community-inclusivity-and-neurodiversity/
Start contributing to Open Source: https://contribute.cncf.io/contributors/getting-started/
Diana's Conference Talks: https://github.com/didiViking/Conferences_Talks
Diana on Medium: https://medium.com/@dianatodea/
Articles on OpenTelemetry for beginners:
https://medium.com/@dianatodea/the-unofficial-guide-to-contributing-to-opentelemetry-where-to-look-and-who-to-talk-to-9de04ae75fe0
CNCF Merge-Forward: https://community.cncf.io/merge-forward
CNCF Neurodiversity initiative: https://community.cncf.io/neurodiversity
Cloud Native Days Romania: https://cloudnativedays.ro/
Cloud Native Days Austria: https://cloudnativedays.at/
From Systems Engineer in Aeronautics via many clouds to becoming an SRE in Observability! That's the path from our guest, Alexandra Franz who is a Lead Product Engineer in SRE at Dynatrace. Tune in and learn how their team plans ahead for expected high traffic around Black Friday, Cyber Monday or the Super Bowl. We discuss how regional traffic patterns and differences in available hardware get factored in for capacity management and cost control. We also learn why global cloud outages are stressful - but - how those incidents can also be the reward for a good SRE.
Make sure to connect with Alexandra on LinkedIn: https://www.linkedin.com/in/alexandrafranz/
If you are still treating your AI Coding Agent like a chat bot and not like a development team then this is one more reason to tune into this episode.
In his blog post series 31 Days of Vibe Coding, Jeff Blankenburg walks us through all the lessons learned when bringing an idea to life just with vibe coding. His idea was building a website for collectors of baseball cards. With now more than 950k cards from almost 10k players, he has proven that vibe coding, when done right, can truly boost the output of software engineers.
Tune in and learn about how to effectively use Git Issues as the backlog for your AI, the importance of going through different phases in your conversation with the AI and why it is important to ask the AI the question: "Do you have any questions for me?"
Links we discussed
LinkedIn Profile: https://www.linkedin.com/in/jeffblankenburg/
31 Days of Vibe Coding: https://31daysofvibecoding.com/
Collect Your Cards: https://collectyourcards.com/
Claude: https://claude.ai/
How many people have you met that implemented distributed tracing in the early 2000s? Make it one more after you have tuned into our latest podcast with William Louth.
William, who can't seem to escape the observability space even though he keeps trying, has a track record in the space. He is an innovator and tool builder and is currently reimagining intelligent systems by shifting the focus from data collection to meaning-making. In our conversation we learn about situational awareness and how systems should use symbols to show their current state by also taking into account everything they are aware of happening in their ecosystem.
This podcast episode has been long overdue and opens a fascinating new world beyond metrics, logs and traces!
Links discussed
Williams LinkedIn: https://www.linkedin.com/in/william-david-louth/
Humainary Research: https://humainary.io/research/
Humainary GitHub: https://github.com/humainary-io
Serventis Signs: https://raw.githubusercontent.com/humainary-io/substrates-api-java/refs/heads/main/ext/serventis/SIGNS.md
It started with the prompt: "Create an Uber Clone"! Several iterations and some months later Abhi presents his lessons learned when vibing a Ride Share Platform for RoboTaxis at Cloud Native Days Austria!
"Commit to one tool and go deep. Don't get distracted by all the options you have. Treat your agent like a human! Get better in expressing what you really want!", those are the many lessons learned in Abhi's journey applying the potential of the latest AI agents that are available for software engineers.
Tune into our latest episode and understand what Abhi means when he says: Context is important! Give it Macro Context and do Micro Incremental Improvements!
Links we discussed
Abhi's LinkedIn: https://www.linkedin.com/in/abhimanyuselvan/
Cloud Native Austria Talk: https://www.youtube.com/watch?v=VjMPHWjawxM&list=PLtLBTEzR4SqU9GwgWiaDt10-yOVIN0nzM&index=9
Cursor AI: https://cursor.com/
OpenSpec: https://openspec.dev/
Chaos Engineering is the practice to introduced controlled failures into a system with the goal to improve the overall resiliency! What started with "lets see what happens when we unplug that server" to "lets simulate network latency issues" or "lets kill critical pods and see if the system recovers gracefully" is now seeing new experiments being conducted that are identified by a new companion: AI
In this episode we have invited Bartek Pisulak, Dir of Cloud Quality Engineering at Pegasystems, who has been educating quality engineers on AI-Augmented Chaos Testing in Practice. Tune in and learn about the how AI can improve efficiency in the 5 critical phases of a chaos experiment: Steady State, Hypothesis, Run Experiment, Verify, Improve!
To learn more about the foundational principles make sure to watch some of the conference talks from Bartek listed below:
Links discussed
Bartek's LinkedIn: https://www.linkedin.com/in/bart%C5%82omiej-pisulak-82b94036/
Talk at Cloud Native Days Austria: https://www.youtube.com/watch?v=xUVCKNpMEz8&list=PLtLBTEzR4SqU9GwgWiaDt10-yOVIN0nzM&index=10
Talk at Porto Tech Hub: https://www.youtube.com/watch?v=-ZuEaA2PoTo
Kraken: https://github.com/krkn-chaos/krkn
ChasoEater: https://github.com/ntt-dkiku/chaos-eater
There is only one successful way to adopt new technology, and that is transformational! Sounds like a high-level consulting pitch but our industry has a track record to validate this statement. Just look at the recent web or cloud-native transformations!
Pini Reznik has been helping organizations along the current AI-Native transformational journey. And what a timing: He just published his book on From Cloud Native to AI-Native where he provides a pragmatic approach to leveraging AI from Pioneering to Gradually Scaling!
Tune in and hear from Pini why he thinks that AI projects are not failing because of bad AI, but because they approaching the problem the old and wrong way!
And, stay until the end to hear how it was to write a book about AI using AI!
Links we discussed
Pini's LinkedIn: https://www.linkedin.com/in/pinireznik/
Link to Book: https://re-cinq.com/book
Our previous episode: https://www.spreaker.com/episode/ai-native-the-next-revolution-after-cloud-native-with-pini-reznik--67692567
Prompt Engineering Conference Talk: https://www.youtube.com/watch?v=W7z5XMnvYt8
Don't get stuck using AI to build faster horses. Instead, find the opportunities and rethink your software delivery processes! That, and only that, will help you increase Developer Experience and Efficiency!
This episode is all about how to measure and improve DevEx in the age of Artificial Intelligence. And with Laura Tacho, CTO at DX, we think we found a perfect guest!
Laura has been working in the dev tooling space for the past 15 years. In her current role at DX she is working on the evolution of DORA and SPACE into DX Core 4 and the DXI Measurement Framework.
In our episode we learn about those frameworks but also how tech leaders need to rethink where and how to apply AI to improve overall efficiency, quality and effectiveness!
The key takeaways from this conversation are
* DevEx is all about the identifying and reducing friction in the end-2-end development process
* Tech Leaders need to become better in articulating technical change requirements to business
* As of today only 22% of code in git is really AI generated. Don't get fooled into believing AI is already better
* Back to Basics makes companies successful with AI. That is: proper CI/CD, testing, documentation, observability!
Here the links we discussed
Laura's LinkedIn: https://www.linkedin.com/in/lauratacho/
DX: https://getdx.com/
Cloud Native Days Austria Talk: https://www.youtube.com/watch?v=kZ1F0-XS1l4
Engineering Leadership Community: https://www.engineeringleaders.io/
The AWS US-East problems on Oct 27th was a good reminder how depending we are on globally shared services. Built-in Resiliency is not guaranteed if systems have a hard dependency on a single region of a single vendor. Many of us have experienced systems being impacted that we use on a daily basis - some critical - some not so critical as Andi will tell you when he found out that is beloved Leberkas Pepi App didnt work!
Besides this outage we discuss lessons learned from Cloud Native Days Austria, Observability and Platform Engineering Meetups in Gdansk and Tallinn as well as giving an outline to the upcoming Cloud and AI-Native US Tour from Henrik Rexed and Andi GrabnerAll the links we discussed are here
* Leberkas Pepi: https://www.leberkaspepi.at/
* Cloud Native Austria: https://www.linkedin.com/company/cndaustria/
* Observability Meetup: https://www.meetup.com/observability-tech-community-meetup-group/
* US Tour from Henrik and Andi: https://events.dynatrace.com/noram-all-de-engineering-efficiency-tour-2025-28225/
While Artificial Intelligence seems to have just popped up when OpenAI brought ChatGPT to the consumer market it has its roots in the mids of the 20th century. But what is it that all of a sudden made it into every conversation we seem to have?
Thomas Natschlaeger, Principal Data Scientist at Dynatrace, who has been working in the AI and Machine Learning space for the past 30 years gives us a brief historical overview and describes the critical evolutionary steps and compelling events in that technology that made it to what it is today.
Tune in and hear about how AIs are trained, how they are optimized and most importantly: how their outputs can be tested and validated!
In our conversation we discuss current trends towards small language models that will help model digital twins of our existing roles and how AIs are used to Validate other AIs like we humans do when a senior engineer does pair programming with a junior and with that provides essential feedback on current accuracy and input to improve the outcome of future tasks.
Links we discussed
LinkedIn Profile from Thomas: https://www.linkedin.com/in/thomas-natschlaeger/
Ask Me Anything Session on Davis CoPilot: https://www.linkedin.com/posts/grabnerandi_llm-copilot-activity-7373837743971393536-QgxV?utm_source=share&utm_medium=member_desktop&rcm=ACoAAABLhVQBbh8Jkn_K8din5tsQlMCpXRNzlKU
Voxxed Conference Talk: https://amsterdam.voxxeddays.com/talk/?id=39801
Attention is all you need paper: https://en.wikipedia.org/wiki/Attention_Is_All_You_Need
On September 8 the world saw the npm supply chain attack. Fortunately the community reacted in record time to avert a disaster.
In todays episode we have Constanze Roedig, Key Researcher at SBA Research, who introduces us to the new buddy of SBoM (Software Bill of Materials): SBoB (Software Bill of Behaviors) and her thoughts on how that new approach to fingerprinting software can help cyber security teams.
What's a BoB? It's a detailed runtime behavior profile of software. It expands on the static validation option through SBOMs as it allows security teams to validate the correct execution behavior of deployed software at deploy time or continuously in production. Thanks to eBPF, a malicious behavior such as opening non expected ports or accessing non expected files can therefore be detected.
Listen to Constanze who shares the work she and Vadim Bauer, Owner of 8gear, have done on this topic. You will learn about how software vendors can create their own SBOBs, ship them with their container images and how security teams can get alerted or enforce any detected malicious behavior. Make sure to check out their GitHub repo, star it if you like it and try their hands-on tutorial!
Links:
Constanze LinkedIn: https://www.linkedin.com/in/croedig/
Vadim LinkedIn: https://www.linkedin.com/in/vadim-bauer/O
BobCtl GitHub Repo: https://github.com/k8sstormcenter/bobctl
Cloud Native Summit Munich Talk: https://www.youtube.com/watch?v=XETuwndd_mw&index=11&pp=iAQB
npm supply chain attack: https://www.infosecurity-magazine.com/news/npm-supply-chain-attack-averted/
Defining AI-Native in 2025 is like trying to define Cloud Native back in 2014! We are in the early stages of understanding what AI really means to us. The ecosystem is just evolving, and many organizations are still struggling with re-architecting their digital systems to cloud native patterns!
To learn more about the current transformational wave—the AI-Native Wave—we have invited Pini Reznik, CEO and Co-Founder of re:cinq. We will discuss what we can learn from previous "waves of innovation," why the business must care, and why the primary AI use case should not be just cost-cutting!
Make sure to get a copy of his book or catch his talk from Cloud Native Munich. All links we discussed here:
Pini's LinkedIn: https://www.linkedin.com/in/pinireznik/
The Next Transformation Mini Book: https://re-cinq.com/mini-book
Cloud Native Munich Talk: https://www.youtube.com/watch?v=CHb3TLEV8ZU
Most AI projects still fail, are too costly, or don't provide the value they hoped to gain. The root cause is nothing new: it's non-optimized models or code that runs the logic behind your AI Apps. The solution is also not new: tuning the system based on insights from Observability!
To learn more about the state of AI Observability, we invited back Nir Gazit, CEO and Co-Founder of traceloop, the company behind OpenLLMetry, the open source observability standard that is seeing exponential adoption growth!
Tune in and learn how OpenLLMetry became such a successful open source project, which problems it solves, and what we can learn from other AI project implementations that successfully launched their AI Apps and Agents
Links we discussed
Nir's LinkedIn: https://www.linkedin.com/in/nirga/
OpenLLMetry: https://github.com/traceloop/openllmetry
Traceloop Hub LLM Gateway: https://www.traceloop.com/docs/hub
Did you know that the average salary for a Platform Engineer is 42.5% more than a DevOps engineer? But why is that?
We sat down with Artem Lajko, CNCF Kubestronaut and Ambassador as well as Author of the book Implementing GitOps with Kubernetes. We dive into the role of a platform engineer, the common pitfalls in implementing IDPs and why Backstage and AI won't solve all your problems. And we touch upon a topic hot off the press around Terraform: Its not dead!
Links we discussed
Artem's LinkedIn: https://www.linkedin.com/in/lajko/
Talk slides from Cloud Land: https://lajko10-my.sharepoint.com/personal/artem_lajko_dev/_layouts/15/onedrive.aspx?id=%2Fpersonal%2Fartem%5Flajko%5Fdev%2FDocuments%2FAttachments%2Fcloud%20land%2D2025%5F%2Epdf&parent=%2Fpersonal%2Fartem%5Flajko%5Fdev%2FDocuments%2FAttachments&ga=1
State of Platform Engineering Report: https://platformengineering.org/reports/state-of-platform-engineering-vol-3
Upjet GitHub Project: https://github.com/crossplane/upjet
"Privacy engineering is the art of translating privacy laws and policies into code, figuring out how to make legal requirements such as ‘an individual must be able to request deletion of all their personal data’ a technical reality.", was the elegant explanation from Cat Easdon when asked about what she is doing in her day job.
If you want to learn more then tune in to this episode. Cat, Privacy Engineer at Dynatrace, shares her learnings about things such as: When the right time is to form your own privacy engineering team, why privacy means different things for different people and regulators and what privacy considerations we specifically have in the observability industry so that our users trust our services!
Links:
Cat's LinkedIn Profile: https://www.linkedin.com/in/easdon/
Publications from Cat: https://www.dynatrace.com/engineering/persons/catherine-easdon/
Blog on Managing Sensitive Data at Scale: https://www.dynatrace.com/news/blog/manage-sensitive-data-and-privacy-requirements-at-scale/
Semgrep for lightweight code scanning: https://github.com/semgrep/semgrep
The IAPP: https://iapp.org/
'Meeting your users' expectations' is formally described by the theory of contextual integrity: https://www.open.edu/openlearncreate/mod/page/view.php?id=214540
Facebook's $5 billion fine from the FTC: http://ftc.gov/news-events/news/press-releases/2019/07/ftc-imposes-5-billion-penalty-sweeping-new-privacy-restrictions-facebook
Fact-check: "The $5 billion penalty against Facebook is the largest ever imposed on any company for violating consumers’ privacy and almost 20 times greater than the largest privacy or data security penalty ever imposed worldwide. It is one of the largest penalties ever assessed by the U.S. government for any violation." I think that's still true; the largest fine under the GDPR was €1.2 billion (again for Facebook/Meta)
More than 50% of platform engineering leads don't know how to measure the impact of their platform! Many platform projects fall into common anti-pattern traps that make the platform look great on Day 1 but fail to scale and excite on Day 2!
Daniel Bryant - who's profile tagline is "Helping you build better platforms" - is sharing his thoughts on how to measure the value of your platform, how to avoid common anti-patterns and why he believes that the future of platform engineering is in Platform Democracy!
And of course, we wrap everything up with a discussion around the impact of Agentic AI towards platform engineering. So - tune in!
Here the links we discussed
Daniel's LinkedIn Profile: https://www.linkedin.com/in/danielbryantuk/
Platform Engineering Book for Technical Product Leaders: https://www.amazon.de/Platform-Engineering-Technical-Product-Leaders/dp/1098153642/ref=asc_df_1098153642
Platform Engineering Day Talk: https://www.syntasso.io/post/syntasso-at-platengday-london-presentation-recap
Kratix Website: https://www.kratix.io/
Ai-Driven Platform Engineering Blog: https://www.syntasso.io/post/what-we-learned-building-a-prototype-ai-driven-dev-interface-for-kratix
Platform Democracy: https://www.syntasso.io/post/platform-democracy-rethinking-who-builds-and-consumes-your-internal-platform
Platform Anti Patterns: https://www.syntasso.io/post/platform-building-antipatterns-slow-low-and-just-for-show
Slide Deck on Platform Engineering for Devs and Architects: https://speakerdeck.com/danielbryantuk/platform-engineering-for-software-developers-and-architects-redux
"How do you measure the impact you have with your platform engineering initiative?" is a question you should be able to answer. To show improvement you must first need to know what the status quo is. And this is where frameworks such as DX Core 4 come in. Never heard about it? Then tune into this episode where we have Dušan Katona, Sr Director of Platform Engineering at Ataccama, who is a big fan of the DX Core Four Metrics and who has just applied it in his current role to optimize developer experience.
Dušan explains the details behind those 4 Core metrics: Speed, Effectiveness, Quality and Impact. He also shares how improving those metrics by a single point results in the equivalent of 10 hours saved per developer per year.
And here the relevant links we discussed today
Dusan's LinkedIn Profile: https://www.linkedin.com/in/dusankatona/
DX Core 4 Blog: https://getdx.com/research/measuring-developer-productivity-with-the-dx-core-4/
Marian's JIRA Analytics Open Source Project: https://github.com/marian-kamenistak/jira-lead-cycle-time-duration-extractor
"15 years ago it was enough to be smart - going forward its not a differentiator - being smart will just make you average!". But what is it? What makes great leaders worth following and how do they achieve tripling their value while others keep waiting for their 5% raise?
4 years ago Marian Kamenistak launched the Engineering Leadership Community out of Prague, Czech Republic. Feeding from his experience in the Silicon Valley this community has grown to 1500 members with the mission to create "Leaders worth following". Tune in and hear from Marian on how to think and talk about value impact vs being held up with trying to achieve technical perfection. Why its important to build a network around you, the difference between mentorship and management as well as how to proof the value to your leadership that you bring to the organization!
Links we discussed today
Marian's LinkedIn: https://www.linkedin.com/in/mariankamenistak/
Engineering Leadership Conference: https://www.elc-conference.io/
Engineering Leadership Community: https://www.engineeringleaders.io/
The Leadership Pipeline Book: https://www.amazon.com/Leadership-Pipeline-Build-Powered-Company/dp/0470894563
Scientific research is the foundation of many innovative solutions in any field. Did you know that Dynatrace runs its own Research Lab within the Campus of the Johannes Kepler University (JKU) in Linz, Austria - just 2 kilometers away from our global engineering headquarter?
What started in 2020 has grown to 20 full time researchers and many more students that do research on topics such as GenAI, Agentic AI, Log Analytics, Procesesing of Large Data Sets, Sampling Strategies, Cloud Native Security or Memory and Storage Optimizations.
Tune in and hear from Otmar and Martin how they are researching on the N+2 generation of Observability and AI, how they are contributing to open source projects such as OpenTelemetry, and what their predictions are when AI is finally taking control of us humans!
To learn more about their work check out these links:
Martin's LinkedIn: https://www.linkedin.com/in/mflechl/
Otmar's LinkedIn: https://www.linkedin.com/in/otmar-ertl/
Dynatrace Research Lab: https://careers.dynatrace.com/locations/linz/#__researchLab
As a leader that wants to optimize an organization you are bound to fail if you isolate social (culture and people) and technical (tools and process) changes. When we ask Lesley Cordero, Staff Engineer at The New York Times how to solve this dilemma she answers: "Platform Engineering, it can drive organizational sustainability by practicing sociotechnical principles that provide a community driven support system for application developers using our standardized shared platform architecture"
Tune in to our latest episode and learn more about the importance of leadership to continuously keep up and balance the tension between "Developers" and "Operations", between "End User Experience" and "Developer Experience" and ultimately between "Culture and People and "Tools and Processes"
Links we discussed
Lesley's LinkedIn: https://www.linkedin.com/in/lesleycordero/
GOTO Conference Talk => https://www.youtube.com/watch?v=Jx-XrUONJ-o
QCon 2025 Talk Details: https://qconlondon.com/presentation/apr2025/platform-engineering-practice-sociotechnical-excellence
DevOpsCon 2024 Talk Details: https://devopscon.io/business-company-culture/platform-engineering-devops/
Do you plan for incidents? Do you have a time / cost budget for it in your sprint or quarterly planning? Do you have engineers that are "interruptible"?
We discussed those and more questions with Lisa Karlin Curtis, Founding Engineer at incident.io who teaches us why we need to think differently about dealing with incidents!
In our discussion we learn why modern incident management embraces more incidents that are publicly shared within an organization to foster learning. We learn about how to train more people to become incident responders, how to triage and categorize incidents, how to better plan for them and how to best report on them
We also touch on AI - and how AI-generated code will eventually result in more Incidents which we should use as an opportunity to learn and improve our engineering process
P.S: This was our 10th-anniversary podcast episode!!
Here the links we discussed in the podcast:
Lisa's LinkedIn: https://www.linkedin.com/in/lisa-karlin-curtis-a4563920/
Her talk at ELC Prague: https://docs.google.com/presentation/d/18536WBHBcPEppEeXXP7o5UQOX2XfWoGmfds2CHegHq4/edit?slide=id.g3434e0cba65_0_0#slide=id.g3434e0cba65_0_0
Incident Playbook: https://incident.io/guide
MCPs (Model Context Protocol) is an open source standard for connecting AI assistants to the the systems where data lives. But you probably already knew that if you have followed the recent hype around this topic after Anthropic made their announcement end of 2024.
To learn more about that MCPs are not that magic, but enable "magic" new use cases to speed up efficiency of engineers we have invited Dana Harrison, Staff Site Reliability Engineer at Telus. Dana goes into the use cases he and his team have been testing out over the past months to increase developer efficiency.
In our conversation we also talk about the difference between local and remote MCPs, the importance of keeping resiliance in mind as MCPs are connecting to many different API backends and how we can and should observe the interactions with MCPs.
Links we discussed
Antrohopic Blog: https://www.anthropic.com/news/model-context-protocol
Dana's LinkedIn: https://www.linkedin.com/in/danaharrisonsre/overlay/about-this-profile/
So you think Distributed Tracing is the new thing? Well - its not! But its never been as exciting as today!
In this episode we combine 50 years of Distributed Tracing experience across our guests and hosts. We invited Christoph Neumueller and Thomas Rothschaedl who have seen the early days of agent-based instrumentation, how global standards like the W3C Trace Context allowed tracing to connect large enterprise systems and how OpenTelemetry is commoditizing data collection across all tech stacks.
Tune in and learn about the difference between spans and traces, why collecting the data is only part of the story, how to combat the challenge when dealing with too much data and how traces relate and connect to logs, metrics and events.
Links we discussed
YouTube with Christoph: LINK WILL FOLLOW ONCE VIDEO IS POSTED
Christoph's LinkedIn: https://www.linkedin.com/in/christophneumueller/
Thomas's LinkedIn: https://www.linkedin.com/in/rothschaedl/
In the ever-changing IT world, creating content that stays relevant for long is hard. One of the objectives of "Platform Engineering for Architects: Crafting Modern Platforms as a Product" was to stay timeless by providing practical examples of use cases not necessarily tied to current technology trends.
The book focuses on the importance of building a platform with a purpose, making the impact measurable, and ensuring the platform continuously evolves by continuously including the end users (the engineering teams) in the evolution of the platform.
Tune in to this episode and hear from Max Körbächer (Founder of Liquid Reply), Hilliary Lipsig (Senior Principal SRE at RedHat), and Andi Grabner (Co-Host of PurePerformance) on what made them write a book on Platform Engineering and get some personal insights into what gets the authors excited about their respective topics.
If you have a chance, meet Max, Hilliary, and Andi at KubeCon in London. They will present at Platform Engineering Day and do a book signing at KubeCrawl!
Links we discussed:
Book on Amazon: https://www.amazon.com/Platform-Engineering-Architects-Crafting-platforms-ebook/dp/B0DH5DJFTH
Platform Engineering Day Session: https://colocatedeventseu2025.sched.com/event/1u5mX/platform-engineering-for-architects-crafting-platforms-as-a-product-max-korbacher-liquid-reply-hilliary-lipsig-red-hat
Hilliary Lipsig: https://www.linkedin.com/in/hilliary-lipsig-a5935245/
Max Körbächer: https://www.linkedin.com/in/maxkoerbaecher/
Andi Grabner: https://www.linkedin.com/in/grabnerandi/
One PetaByte is the equivalent of 11000 4k movies. And CERN's Large Hadron Collider (LHC) generates this every single second. Only a fraction of this data (~1 GB/s) is stored and analyzed using a multicluster batch job dispatcher with Kueue running on Kubernetes.
In this episode we have Ricardo Rocha, Platform Engineering Lead at CERN and CNCF Advocate, explaining why after 20 years at CERN he is still excited about the work he and his colleagues at CERN are doing. To kick things off we learn about the impact that the CNCF has on the scientific community, how to best balance an implementation of that scale between "easy of use" vs "optimized for throughput". Tune in and learn about custom hardware being built 20 years ago and how the advent of the latest chip generation has impacted the evolution of data scientists around the globe
Links we discussed
Ricardo's LinkedIn: https://www.linkedin.com/in/ricardo-rocha-739aa718/
KubeCon SLC Keynote: https://www.youtube.com/watch?v=xMmskWIlktA&list=PLj6h78yzYM2Pw4mRw4S-1p_xLARMqPkA7&index=5
Kueue CNCF Project: https://kubernetes.io/blog/2022/10/04/introducing-kueue/
The word "Compliance" reminds many about mandatory training or audits. Two things not everyone gets excited about!
Tune in and meet Michiel de Lepper who has spent most of his career in Security and Compliance. He gives us a different perspective on the importance of compliance, why it exists, how it intertwines with security and threat detection, what it has to do with security posture management and why he thinks its one of the most exciting things in IT!
Links we discussed:
Michiel's LinkedIn: https://www.linkedin.com/in/madelepper/
Blog posts on security and compliance:
https://www.dynatrace.com/news/blog/dynatrace-for-executives-security-compliance/
https://www.dynatrace.com/news/blog/manage-compliance-and-resilience-at-scale-with-dynatrace/
https://www.dynatrace.com/news/blog/dynatrace-kspm-transforming-kubernetes-security-and-compliance/
Feature Flagging - some may call them "glorified if-statements" - has been a development practice for decades. But have we reached a stage where organizations are doing "Feature Flag-Driven Development?". After all it took years to establish a test-driven development culture despite having great tools and frameworks available!
To learn more we invited Ben Rometsch, Co-Founder of Flagsmith, to chat about the history, state and future of Feature Flagging. He is giving us an update on where the market is heading, how the CNCF project OpenFeature and its community is driving best practices, what the role of AI might be and what he thinks might be next!
Couple of links we discussed during the episode:
Ben on LinkedIn: https://www.linkedin.com/in/benrometsch/
YouTube Video on Observability & Feature Flagging: https://www.youtube.com/watch?v=VZakh1_oEL8
OpenFeature: https://openfeature.dev/
To predict the future, it's important to know the past. And that is true for Bernd Greifeneder, Founder and CTO of Dynatrace, who has been driving innovation in the observability and security since he founded Dynatrace 20 years ago!
Bernd agreed to sit down, look behind the covers and answer the open questions that people posted on his LinkedIn in response to his recent observability prediction blog.
Tune in and learn about Bernd's though on the evaluation from reactive to preventive operations, who is behind the convergence of observability & security, why observability can help those that have serious intentions for sustainability and how observability becomes mandatory and indispensable for AI-driven services.
We mentioned a lot of links in todays session. Here they are:
Our podcast from 9 years ago: https://www.spreaker.com/episode/015-leading-the-apm-market-from-enterprise-into-cloud-native--9607734
Bernds LinkedIn Post: https://www.linkedin.com/feed/update/urn:li:activity:7275101213237354497/
Predictions Blog: https://www.dynatrace.com/news/blog/observability-predictions-for-2025/
K8s Predictive Scaling Lab: https://github.com/Dynatrace/obslab-predictive-kubernetes-scaling
Security Video: https://www.youtube.com/watch?v=ICUwRy4JFTk
Carbon Impact App: https://www.youtube.com/watch?v=8Px0BB1U1yk
AI & LLM Observability Video: https://www.youtube.com/watch?v=eW2KuWFeZyY
eBay, Yahoo, Netflix and then 10+ years at Uber. In this episode we sit down with Vishnu Acharya, Head of Network Infrastructure EMEA and Platform Engineering at Uber. Vishnu shares how Uber has scaled over the years to about 4000 engineers and how his team makes sure that infrastructure and platform engineering scales with the growing company and the growing demand on their digital services.
Tune in and learn about how Vishnu thinks about SLOs across all layers of the stack, how they manage to get better insights with their cloud providers and why its important to have an end-to-end understanding of the most critical end user journeys.
Links we discussed:
Conference talk at Observability & SRE Summit: https://www.iqpc.com/events-observability-sre-summit/speakers/vishnu-acharya
Vishnu's LinkedIn Page: https://www.linkedin.com/in/vishnuacharya/
Uber Engineering Blog: https://www.uber.com/blog/engineering/
For the past 10 years Anton has been working at Booking.com - one of the leading digital travel companies based out of Amsterdam. The journey that started as System Administrator has led Anton to be an Engineering Manager for Site Reliability where over the past 3 years he led the rollout and adoption of OpenTelemetry as the standard for getting observability into new cloud native deployments.
Tune in and learn how Anton saw R&D grow from 300 to 2000, why they replaced their home-grown Perl-based Observability Framework with OpenTelemetry, how they tackle adoption challenges and how they extend and contribute back to the open source community
Links we discussed:
Anton's LinkedIn Profile: https://www.linkedin.com/in/antontimofieiev/
Observability & SRE Summit: https://www.iqpc.com/events-observability-sre-summit/speakers/anton-timofieiev
OpenTelemetry: https://opentelemetry.io/
Most services are moving to SaaS - whether it’s email, collaboration, customer relations, or finance. But not everyone can go to SaaS - or at least that’s the initial reaction when navigating certain industries’ rules and regulations.
Milan Steskal - who worked in healthcare for many years - is now helping organizations ask the right questions and find the best solutions as they evaluate their options to move their observability data to SaaS. Tune in and learn about the questions to ask vendors and your internal security, privacy, and compliance teams. Milan also walks us through the capabilities SaaS vendors such as Dynatrace have put in place to protect data sent to the cloud so that it stays safe and only accessible to those needing access.
Links discussed today:
Milans LinkedIn Page: https://www.linkedin.com/in/milansteskal/
Dynatrace Trust Center: https://www.dynatrace.com/company/trust-center/
Blogs on Trust: https://www.dynatrace.com/news/tag/trust-center/
Andreas Taranetz is a software engineer and lecturer at the University of Vienna. He creates a lot of educational content around Web Performance Optimization. For the past seven years, he has also operated Wahlkabine, Austria's top website, for matching one's political views with the parties that are up for election.
This episode was an amazing flashback - reminding us about the time when Steve Souders - the "godfather" of Web Performance Optimization - educated web developers about optimizing CSS, JavaScript, and server-side roundtrips.
Tune in and learn why Web Performance is still such an important topic, how it relates to sustainability, why you should cache on every layer, and what the Static Site Paradox really is!
Links we discussed in the episode:
Andreas on LinkedIn: https://www.linkedin.com/in/andreas-taranetz/
Personal Website: https://andreas.taranetz.com/
We Are Developers Talk: https://www.youtube.com/live/KRemC82gsBk
Wahlkabine: https://wahlkabine.at/
Steve Souders: https://stevesouders.com/
Authentication (validating who you claim to be) and Authorization (enforcing what you are allowed to do) are critical in modern software development. While authentication seems to be a solved problem, modern software development faces many challenges with secure, fast, and resilient authorization mechanisms.
To learn more about those challenges, we invited Alex Olivier, Co-Founder and CPO at Cerbos, an Open Source Scalable Authorization Solution. Alex shared insights on attribute-based vs. role-based access Control, the difference between stateful and stateless authorization implementations, why Broken Access Control is in the OWASP Top 10 Security Vulnerabilities, and how to observe the authorization solution for performance, security, and auditing purposes.
Links we discussed during the episode:
Alex's LinkedIn: https://www.linkedin.com/in/alexolivier/
Cerbos on GitHub: https://github.com/cerbos/cerbos
OWASP Broken Access Control: https://owasp.org/www-community/Broken_Access_Control
Open Source is the Best Thing that happened to IT"! Powerful words from Marcio Lena who has been using and contributing back to open source for the past 20+ years. Besides being a vivid advocate for open source, Marcio also knows the concerns of large enterprises when picking open source projects.
Tune in and follow our discussion about how to identify a healthy open-source project, how to balance between vendor and community lock-in, the power of open standards such as OpenTelemetry, open source business models as well as that contributing to open source is not limited to code but includes documentation, education and advocacy as well!
Links we discussed:
Marcio's LinkedIn Page: https://www.linkedin.com/in/marcio-lena/
CNCF DevStats: https://devstats.cncf.io/
Linux Foundation Events: https://events.linuxfoundation.org/
CNCF Ambassadors: https://www.cncf.io/people/ambassadors/
DORA - the EU's Digital Operational Resiliency Act - will take effect in January of 2025 and is currently top of mind for IT Leaders across all financial service institutions that operate in the European Union. But what is DORA really? Why is this important? How can institutions meet the DORA requirements? What is the role of observability, automation and AI in all of this?
To answer all those and more questions we invited Kay Young, Sr Principal Product Manager at Dynatrace, who has been working with organizations around the globe that have been tasked to implement regulations such as DORA, GDPR, FedRAMP or others.
In our conversation we also touch base on the third-party risk management as well as resiliency testing and incident reporting.
Resources we discussed:
Kay's LinkedIn Profile: https://www.linkedin.com/in/karlien-young-4a156730/
What is DORA blog: https://www.dynatrace.com/news/blog/what-is-dora/
Taming DORA compliance: https://www.dynatrace.com/news/blog/taming-dora-compliance-with-ai-observability-and-security/
Blog on Dynatrace's DORA compliance journey: https://www.dynatrace.com/news/blog/the-dynatrace-journey-toward-dora-compliance/
Beyond DORA compliance: https://www.dynatrace.com/news/blog/dora-how-dynatrace-helps-the-financial-sector-stay-resilient/
NAIS (pronounced like NICE) is a team central application platform that provides DevOps teams with the tools they need build, test, deploy, run and observe applications.
In this episode Hans Kristian Flaatten, Platform Engineer at NAV, walks us through the WHYs, HOWs and challenges of building modern platforms on Kubernetes. Tune in and hear WHY they defined their own abstraction layer for applications, HOW developers benefit from that platform and WHY they developed their developer portal instead of going with other popular available choices.
Links we discussed:
Hans Kristian's LinkedIn: https://www.linkedin.com/in/hansflaatten/
NAIS Documentation: https://docs.nais.io/
"We will overwhelm developers if we give them the same specialized observability, security or deployment tools that are used by their platform engineering, operations, SREs or security teams!" - says Viktor Farcic, Developer Advocate at UpBound and host of The DevOps Toolkit YouTube channel.
Tune in and hear us discuss about making observability easier accessible for developers, what Viktor doesn't like about Kubernetes and how Crossplane- the cloud native control plane framework - can be the gateway to real product-oriented platform engineering!
Here the links we discussed during this episode:
* Viktor on LinkedIn: https://www.linkedin.com/in/viktorfarcic/
* DevOps Toolkit: https://www.youtube.com/@DevOpsToolkit
* Crossplane: https://www.crossplane.io/
Hans Kristian is a Platform Engineer for NAV's Kubernetes Platform Nais hosting Norway's wellfare services. With 10 years on Kubernetes, 2000 apps and 1000 developers across more than 100 teams there was a need to make OpenTelemetry adoption as easy as possible.Tune in as we hear from Hans Kristian who is also a CNCF Ambassador and hosts Cloud Native Day Bergen why OpenTelemetry is chosen by the public sector, why it took much longer to adopt, which challenges they had to scale the observability backend and how they are tackling the "noisy data problem"
Links we discussed in the episode
* Follow Hans Kristian on LinkedIn: https://www.linkedin.com/in/hansflaatten/
* From 0 to 100 OTel Blog: https://nais.io/blog/posts/otel-from-0-to-100/?foo=bar
* Cloud Native Day Bergen: https://2024.cloudnativebergen.dev/
* Public Money, Public Code. How we open source everything we do! (https://m.youtube.com/watch?v=4v05Huy2mlw&pp=ygUkT3BlbiBzb3VyY2Ugb3BlbiBnb3Zlcm5tZW50IGZsYWF0dGVu)
* State of Platform Engineering in Norway (https://m.youtube.com/watch?v=3WFZhETlS9s&pp=ygUYc3RhdGUgb2YgcGxhdGZvcm0gbm9yd2F5)
Has one of the decision makers in your organization decided that you have to go "all in on technology X" because they saw a great presentation at a conference or got a great sales pitch from a vendor? If that is the case then this episode is for you and you should forward it to those decision makers.
Sebastian Vietz, Director of Reliability Engineering and Host of the Reliability Enablers Podcast, shares his thoughts on considerations when picking a technology like Serverless. We discuss the importance of knowing limits, best fit architectural patterns and things that should influence your technology decisions!
Being aware of coldstarts, a 20000 concurrent request limit or 512mb being an ideal size for Lambda are just some of the things we can all learn from Sebastian.
Additional links we discussed:
Sebastians LinkedIn: https://www.linkedin.com/in/sebastianvietz/
Reliability Podcast: https://podnews.net/podcast/ibe8k
More things on serverless: https://serverlessland.com/
When your code runs on more than 6 million systems - many of them business critical - then this is really exciting news for Marco and Wolfgang, Dynatrace OneAgent Java Team members. Their code powers auto-instrumentation and collection of all observability signals of Java based applications running on every possible stack: container in k8s, serverless, VM, on your workstation or even the mainframe.
Tune is as we sat down with Marco and Wolfgang to learn what it means to continuously innovate on agent-based instrumentation with 160+ other engineers across the globe that also focus on OneAgent. They share insights on how they develop their observability code, how they continuously test across all supported environments, what the processes at Dynatrace look like to avoid situations like the recent CrowdStrike outage and how they integrate and collaborate with other communities and tools such as OpenTelemetry!
Things we discussed during the episode
Dynatrace OneAgent: https://www.dynatrace.com/platform/oneagent/
Dynatrace for Java: https://www.dynatrace.com/technologies/java-monitoring/
OpenTelemetry and Dynatrace: https://docs.dynatrace.com/docs/extend-dynatrace/opentelemetry
Jobs at Dynatrace: https://careers.dynatrace.com/
When thousands of systems show a blue screen - which ones do you fix first to quickly bring up your most critical systems? For that you need to know which systems are impacted, which mission critical applications run on it, and which depending systems are also impacted by something like the recent CrowdStrike incident!
We have invited Josh Wood, Principal Solutions Engineer at Dynatrace, who was one of the first responders helping organizations to leverage observability data to identify which systems to fix first to bring critical apps such as ATMs, Self-Service Terminals, POS (Point of Sales), ... back up again quickly.
In this special episode Josh is walking us through the technical details of the CrowdStrike BSOD (Blue Screen of Death), what caused it, how to leverage observability to get a priorities list of systems to fix first and what organizations can do to prevent software impacting issues in the future.
Here the links we discussed in the episode:
Josh on LinkedIn: https://www.linkedin.com/in/joshuadwood/
Josh's blog on CrowdStrike BSOD: https://www.dynatrace.com/news/blog/crowdstrike-bsod-quickly-find-machines-impacted-by-the-crowdstrike-issue/
CrowdStrike Incident Takeaway Blog: https://www.dynatrace.com/news/blog/crowdstrike-incident-revisiting-vendor-quality-control/
WebAssembly runs in every browser, provides secure and fast code execution from any language, runs across multiple platforms and has a very small binary footprint. It's adopted by several of the big web-based SaaS solutions we use on a daily basis.
But where did WebAssembly come from? What problems does it try to solve? Has it reached critical adoption? And how about observing code that gets executed in browsers, servers or embedded devices?
To answer all those questions we invited Matt Butcher, CEO at Fermyon, who explains the history, current implementation status, limitations and opportunities that WebAssembly provides.
Further links we disucssed
LinkedIn Profile: https://www.linkedin.com/in/mattbutcher/
Fermyon Dev Website: https://developer.fermyon.com/
The New Stack Blog with Matt: https://thenewstack.io/webassembly-and-kubernetes-go-better-together-matt-butcher/
"Because I don't want software to go down every single day in my next gig!" is what drives the motivation of Ash Patel, Reliability Advocate and Podcast host of SREpath, to talk about and educate IT professionals on the importance of building and operating reliable systems.
For 15 years Ash used to be Director of Operations at a private health service organization. He has experienced that patients couldn't get the treatment they expected due to unreliable software he was responsible for.
In our conversation Ash talks about how he had to close the knowledge gap on technology but also solve the problem by having engineers understand the pain and the requirements of their end users. One way to educate more engineers is through his podcast called SREpath where Observability has become a hot topic recently. Tune in, hear about the memorable stories from his guests from CapitalOne, IKEA and SquaredUp, and lets move towards a world where software is reliable by default.
Links as discussed today:
Ash on LinkedIn: https://www.linkedin.com/in/ash-patel-srepath/
SREpath Podcast: https://www.srepath.com/podcast/
Clearing Delusions in Observability https://read.srepath.com/p/30-clearing-delusions-in-observability-2af
Boosting your observability data's usability https://read.srepath.com/p/35-boosting-your-observability-datas-3f4
How to Enable Observability for Success https://read.srepath.com/p/40-how-to-enable-observability-for
"Meet your users where they are!" - For Platform Engineering Teams that means understanding the current way your engineers work, understand their pain, and provide a solution that doesnt force them to change their behavior but provides a 10x efficiency improvement. Thats not easy to achieve but is what we discussed with Abby Bangser in our latest episode
Abby is a Team Topologies Advocate, has spent years at Thoughtworks helping organizations transform through Delivery Platforms and is now a Lead at the CNCF Platform Working Group. Tune in and hear our discussions on Why Platform Engineering is nothing new, how to avoid Platform Engineering Teams to become your next bottleneck and silo, why Platforms need to have more than one interface and why the purpose of Platform Engineering should be to bring good Developer Experience to all engineers
Here all the links we discussed during this episode
Platform Engineering Maturity Model: https://tag-app-delivery.cncf.io/whitepapers/platform-eng-maturity-model/
CNCF Platform Working Group: https://tag-app-delivery.cncf.io/wgs/platforms/
KubeCon 2024 Talk: https://colocatedeventseu2024.sched.com/event/1YFdf/sometimes-lipstick-is-exactly-what-a-pig-needs-abby-bangser-syntasso-whitney-lee-vmware
GitHub Issue for Questionnaire: https://github.com/cncf/tag-app-delivery/issues/635
Kratix: https://www.kratix.io/
Abbys LinkedIn: https://www.linkedin.com/in/abbybangser/
Abbys Events: https://www.paintedwavelimited.com/events
Requesting more CPU for your database used to take 6 months of planning 20 years ago. Now it takes the execution of a Terraform script. What has stayed the same all those years is Almudena Vivanco's passion for performance engineering to keep systems optimized. Ensuring that systems are available, scalable and resilient even during spike events such as the upcoming Euro Cup or any holiday specials.
Tune in and hear from Almudena, who is currently working for SCRM Lidl, on how moving to the cloud gave new justification to performance engineering. She explains the importance of connecting business with service level objectives and gives insights on how Lidl makes sure to sell 50000 pieces of pork without breaking the cloud bank
Here the additional links we discussed
Slides from Barcelona Meetup: https://docs.google.com/presentation/d/1h83V4gUyqAmIWeAAtKb4BcRvuJV-XirLk-9Xq077nbw
Video from TestCon: https://www.youtube.com/watch?v=rIP_G-YBy04
LinkedIn: https://www.linkedin.com/in/almudenavivanco/
Making observability available to everyone! This noble goal needs superhero powers in an IT world where there is so much chatter and confusion about what observability is, how to sell the value add besides a glorified troubleshooting tool and how OpenTelemetry will disrupt the landscape.
In our latest episode we have Rainer Schuppe, Observability Veteran (more than 20+ years in the space), who has worked for the majority of the observability vendors. He is sharing his observability expertise through workshops in his home town of Mallorca. Teaching organizations from basic to strategic observability implementations.
Tune in and learn about the typical adoption and maturity path of observability within enterprises: from fixing a problem at hand, to justifying the cost to keep it until enabling companies to become information driven digital organizations! Also check out his OpenTelemetry journey in his blog post series
Here are the links we discussed today:
Observability Heroes Website: https://observability-heroes.com/
Observability Heroes Community: https://observability.mn.co/
Cloud Native Mallorca Meetup: https://www.meetup.com/cloud-native-mallorca/
OpenTelemetry: https://opentelemetry.io/
Rainer on LinkedIn: https://www.linkedin.com/in/rainerschuppe/
eBPF is a kernel technology enabling high-performance, low overhead tools for networking, security and observability. In simpler terms: eBPF makes the kernel programmable!
Tune in to this episode whether you have never heard about eBPF, using eBPF based tools such as bcc, Cillium, Falco, Tetragon, Inspector Gadget ... or whether you are developing your own eBPF programs!
Liz Rice, Chief Open Source Officer at Isovalent, kicks this episode off with a brief introduction of eBPF, explains how it works, which use cases it has enabled and why eBPF can truly give you super powers!
In our conversation we dive deeper into the performance aspects of eBPF: how and why tools like Cillium outperforms classical network load balancers, how performance engineers can use it and how the Kernel internally handles eBPF extecutions.
We discussed a lot of follow up material - here are all the relevant links:
Liz's slide deck on "Unleashing the kernel with eBPF": https://speakerdeck.com/lizrice/unleashing-the-kernel-with-ebpf
eBPF Documentary on YouTube: https://www.youtube.com/watch?v=Wb_vD3XZYOA
Learning eBPF GitHub repo accompanying her book: https://github.com/lizrice/learning-ebpf
eBPF website: https://epbf.io
Liz on LinkedIn: https://www.linkedin.com/in/lizrice/
Use Things you Understand! Learn the fundamentals to understand the layers of abstraction! And remember that we don't live in a world with unlimited resources!
These are advice from our recent conversation with Ernst Ambichl, Chief Product Architect at Dynatrace, who has started his performance career in the late 80s building the first load testing tools for databases which later became one of the most successful performance engineering tools in the market.
Tune in and learn about how Ernst has evolved from being a performance engineer to become an advocate for "Designing and Architecting for Performance". Ernst explains how important good upfront analysis of performance requirements and characteristics of the underlying infrastructure is, how to define baselines and constantly evaluate your changes against your goals.
On a personal note: I want to say THANK YOU Ernst for being one of my personal mentors over the past 20+ years. You inspired me with your passion about performance and building resilient systems
SREs (Site Reliability Engineers) have varying roles across different organizations: From Codifying your Infrastructure, handling high priority incidents, automating resiliency, ensuring proper observability, defining SLOs or getting rid of alert fatigue. What an SRE team must not be is a SWAT team - or - as Dana Harrison, Staff SRE at Telus puts it: "You don't want to be the fire brigade along the DevOps Infinity Loop"
In his years of experience as an SRE Dana also used to run 1 week boot camps for developers to educate them on making apps observable, proper logging, resiliency architecture patterns, defining good SLIs & SLOs. He talked about the 3 things that are the foundation of a good SRE: understand the app, understand the current state and make sure you know when your systems are down before your customers tell you so!
If you are interested in seeing Dana and his colleagues from Telus talk about their observability and SRE journey then check out the On-Demand session from Dynatrace Perform 2024: https://www.dynatrace.com/perform/on-demand/perform-2024/?session=simplifying-observability-automations-and-insights-with-dynatrace#sessions
Whether its GitOps, DevOps, Platform Engineering, Observability as a Service or other terms. We all have our definitions, but rarely do we have a consensus on what those terms really mean! To get some clarity we invited Roberth Strand, CNCF Ambassador and Azure MVP, who has been passionately advocating for GitOps as it was initially defined and explained by Alexis Richardson, Weaveworks in his blog What is GitOps Really!
Tune in and learn about Desired State Management, Continuous Pull vs Pushing from Pipelines, how Progressive Delivery or Auto-Scaling fits into declaring everything in Git, what OpenGItOps is and why this podcast will help you get your GitOps certification (coming soon)
As we had a lot to talk we also touched on Platform Engineering and various other topics
Here are all the links we discussed:
Alexis GitOps Blog Post: https://medium.com/weaveworks/what-is-gitops-really-e77329f23416
OpenGitOps: https://opengitops.dev/
Flux Image Reflector: https://fluxcd.io/flux/components/image/
CNCF White Paper on Platform Engineering: https://tag-app-delivery.cncf.io/whitepapers/platforms/
Platform Engineering Maturity Model: https://tag-app-delivery.cncf.io/whitepapers/platform-eng-maturity-model/
Platform Engineering Working Group as part of TAG App Delivery: https://tag-app-delivery.cncf.io/wgs/platforms/
Can you explain GitOps in simple terms? How does it fit into Continuous Integration (CI), Continuous Delivery and Continuous Deployment? And what are considerations when rolling out GitOps in an enterprise?
To get answers to those questions we sat down with Christian Hernandez, Head of Community at Akuity, who has a fabulous analogy to explain GitOps that I am sure many of us will "borrow" from him. Christian also explains the ecosystem he works in such as ArgoCD, Kargoas well as OpenGitOpswhich aims to provide open-source standard and best practices to implementing GitOps.
We closed the session with some advice around Application Dependency Management, External Secrets Operator and choosing the right Git Repo Structure.
Here are some of the links we discussed:
OpenGitOps: https://opengitops.dev/
ArgoCD: https://argoproj.github.io/cd/
Kargo: https://github.com/akuity/kargo
ArgoCon: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/argocon/
GitOpsCon: https://events.linuxfoundation.org/gitopscon-north-america/
While the mainframe is powering the world's most critical system the words "modern", "open source" or "generative AI" typically don't come to mind. So lets change this!
To do that simply tune in to our latest episode where we have Jessielaine (Jelly) Punongbayan, Sr. Technical Support Engineer at Dynatrace, telling us why she is excited about the modern Mainframe and how it brought her from the Philippines via Singapore and Czech Republic to Austria.
We learn about all the open-source projects and communities she is involved in such as Open Mainframe or Zowethat make it easy to connect the Mainframe with the modern tooling of today's development environments. Jelly shares her stories about the role of good observability, how it connects the distributed and the mainframe world and how it enables development teams to build more efficient systems. And what about AI? Well - you have to tune in and listen to the end!
Here the links discussed in the episode
Writing a COBOL program using VSCode: https://medium.com/modern-mainframe/beginners-guide-cobol-made-easy-introduction-ecf2f611ac76
Using CircleCI to perform automation in Mainframe: https://medium.com/modern-mainframe/beginners-guide-cobol-made-easy-leveraging-open-source-tools-eb4f8dcd7a98
Using OpenTelemetry to capture Mainframe Insights: https://medium.com/@jessielaine.punongbayan/re-imagining-mainframe-insights-through-open-source-tooling-79dd4c937114
Dynatrace support for Mainframe: https://www.dynatrace.com/technologies/mainframe-monitoring/
201 is the HTTP status code for Resource Created. It is also the number of PurePerformance Episodes (including this one) we have published over the past years. None better to invite than the person who initially inspired us to launch PurePerformance: Mark Tomlinson, Performacologist and Director of Observability at FreedomPay
Tune in and listen to our thoughts on current state of automation, a recap on IFTTT, whether we believe that AIs such as CoPilot will not only make us more efficient in creating code and scripts but also lead to new ways of automation. We also give a heads-up (or rather a recap) of what Mark will be presenting on at Perform 2024.
To learn more about and from Mark follow him on the various social media channels:
LinkedIn: https://www.linkedin.com/in/mtomlins/
Performacology: https://performacology.com/
Marcelo Amaral is a Researcher for Cloud System Optimization and Sustainability. With his background in performance engineering where he optimized microservice workloads in containerized environments making the leap towards analyzing and optimizing energy consumption was easy.
Tune in to this episode and learn about how Kepler, the CNCF project Marcelo is working on, which provides metrics for workload energy consumption based on power models it was trained on by the community. Marcelo goes into details about how Kepler works and also provides practical advice for any developer to keep energy consumption in mind when making architectural and coding decisions.
To learn more about Kepler and the episode today check out:
LinkedIn from Marcelo: https://www.linkedin.com/in/mcamaral/
CNCF Blogpost on Kepler: https://www.cncf.io/blog/2023/10/11/exploring-keplers-potentials-unveiling-cloud-application-power-consumption/
Kepler GitHub Repo: https://github.com/sustainable-computing-io/kepler
Its only been a year since ChatGPT was introduced. Since then we see LLMs (Large Language Models) and Generative AIs being integrated into every days life software applications. Developers have the hard choice to pick the right model for their use case to produce the quality of output their end users demand.
Tune in to this session where we have Nir Gazit, CEO and Co-founder of Traceloop, educating us about how to observe and quantify the quality of LLMs. Besides performance and costs engineers need to look into quality attributes such as accuracy, readability or grammatical correctness.
Nir introduces us to OpenLLMetry - a set of Open Source extensions built on top of OpenTelemetry providing automated observability into the usage of LLMs for developers to better understand how to optimize the usage of LLMs. His advice to every developer is to start measuring the quality of your LLMs on Day 1 and continuously evaluate as you change your model, the prompt and the way you interact with your LLM stack!
If you have more questions about LLM Observability check out the following links:
OpenLLMetry GitHub Page: https://github.com/traceloop/openllmetry
Traceloop Website: https://www.traceloop.com/
OpenLLMetry Documentation: https://traceloop.com/docs/openllmetry
After analyzing Distributed Traces over more than 15 years Brian and I thought that everyone in software engineering and operations must be satisfied with all that observability data we have available. But. Maybe Brian and I were wrong because we didn’t fully understand all the use cases - especially those for developers that must fix code in production or need to quickly understand what code from somebody else is really doing without having the luxury to add another log line and redeploy on the fly. To learn more about the observability requirements of developers we invited Liran Haimovitch, CTO at Rookout and now part of Dynatrace, who has spent the last 7 years solving the challenging problems that developers face day and night. Tune in and learn about what non-breaking breakpoints are, how it is possible to "debug in production" without impacting running code and how we can make developers lives easier even though we push so many things "to the left"
I was invited to speak at BankTechShow in Budapest, Hungary where the nations IT leaders in the banking sector presented and discussed the future of banking - both in the cloud as well as what it means for the physical bank branches.
I got a chance to sit down with Adam Gajdi, IT Solutions CoE Lead at K&H, who walked me through the process of their recent new mobile banking app launch. Adam highlighted the importance of observability for both business owners as well as developers. Furthermore, Adam enlightened me with the fact that Hungarian banks are mandated to conduct chaos tests to proof that their systems are resilient in case of data center outages. I was obviously also curious about how AI, LLMs and other technologies are adopted in their sector. Tune in to learn more
Besides attending KubeCon 2023 NA Andreas (Andi) Grabner, co-host of PurePerformance but guest today, has also travelled parts of the US to chat with the broader observablity community on topics such as Platform Engineering, Observability, DevOps, Automation & Security.
Tune in and get a quick recap of all the topics Andi has picked up on his recent trip
Zero-Trust Architectures. Data-Flow Inventory. User Experience First! Those are key initiatives in the public sector to ensure that digital services delivered to citizens around the globe are not only working with a flawless user experience but are also safe from any bad actors trying to disrupt agencies on local, stage and federal sectors.
In this episode we invited Willie Hicks, Federal CTO at Dynatrace, to learn more about the state of observability and security with government agencies Willie has been working with over the past decade. In our conversation we explore the differences between commercial and government as it comes to ROI or how they see competition as a driving motivator.
To learn more about the public sector tune into the Tech Transformers podcast that Willie is co-hosting with his colleague Carolyn Ford.
4% of worldwide CO2 emissions come from IT and like in all other industries we have big potential to not only reduce the carbon footprint but also lower costs.
Tune in to our episode where we have Mario-Leander Reimer, CTO at QAware GmbH, talk about his top 3 suggestions for Sustainable IT: Making the right architectural choices, Right-sizing your environments and shutting down environments not needed!
Mario is also heavily involved in the CNCF and gives us an overview of projects to look into such as Kepler, kube-green, Karpenter or Carbon Aware Multi-Cluster Schedulers.
Here are the links we discussed:
Blue turns Green presentation: https://speakerdeck.com/lreimer/blue-turns-green-approaches-and-technologies-for-sustainable-k8s-clusters-number-kcdmunich?slide=5
Kepler Project: https://kepler.gl/
kube-green: https://kube-green.dev/
CNCF TAG Environmental Sustainability: https://github.com/cncf/tag-env-sustainability
Sustainability Week: https://tag-env-sustainability.cncf.io/cloud-native-sustainability-week/
Martin Spier was one of six engineers to take care of all of Netflix Operations about 10 years ago. Back then performance and observability tools weren't as sophisticated and didn't scale to the needs of Netflix as some do today. FlameScope was one of the Open Source projects that evolved out of that period, visualizing Flame Graphs on a time-scaled heatmap to identify specific performance patterns that caused issues in their complex systems back then.
Tune in to this episode and hear more performance and observability stories from Martin, about his early days in Brazil, his time at Expedia and Netflix and about his current role as VP of Engineering at PicPay - one of the hottest fin techs in Brazil.
More links we discussed:
Performance Summit talk about FlameCommander: https://www.youtube.com/watch?v=L58GrWcrD00
CMG Impact talk on Real User Monitoring at Netflix: https://www.cmg.org/2019/04/impact-2019-real-user-performance-monitoring-at-netflix-scale/
Learn more about Vector: https://netflixtechblog.com/extending-vector-with-ebpf-to-inspect-host-and-container-performance-5da3af4c584b
Martin's GitHub: https://github.com/spiermar
Connect with him on LinkedIn: https://www.linkedin.com/in/martinspier/
Africa is not only the second largest continent in the world - its also top when it comes to adoption of cloud native technologies. I was fortunate to spend a week in South Africa and had the chance to spend a lot of time with Kelvin Klein, Dynatrace Product Manager at Mediro ICT. After two observability events in Johannesburg and Cape Town and several meetings with local tech leaders I got to sit down with Kelvin and learn more about the status of Observablity, Cloud Native and Security in South Africa.
I was fortunate to travel to South Africa and meet many tech leaders in Johannesburg and Cape Town to talk about Observability, Security, Automation, Platform Engineering, DevOps and FinOps. One of those leaders is Amit Chiba, Multi Product Specialist at Nedbank. I sat down with Amit to discuss his personal journey and his projects at Nedbank, one of the leading financial institutions in South Africa. Tune in and hear from Amit how self-service platform engineering helps them to scale observability, how they tackle cloud costs and why he thinks that the future of IT Ops is more Sleep!
Do you measure build times? On your shared CI as well as local builds on the developers workstations? Do you measure how much time devs spend in debugging code or trying to understand why tests or builds are all of a sudden failing? Are you treating your pre-production with the same respect as your production environments?
Tune in and hear from Trisha Gee, Developer Champion at Gradle, who has helped development teams to reduce wait times, become more productive with their tools (gotta love that IDE of yours) and also understand the impact of their choices to other teams (when log lines wake up people at night). Trisha explains in detail what there is to know about DPE (Developer Productivity Engineering), how it fits into Platform Engineering, why adding more hardware is not always the best solution and why Flaky Tests are a passionate topic for Trisha.
Here the links to Trishas social media, her books and everything else we discussed during the podcast
* LinkedIn: https://www.linkedin.com/in/trishagee/
* Trishas Website: https://trishagee.com/
* Trisha's Talk on DPE: https://trishagee.com/presentations/developer-productivity-engineering-whats-in-it-for-me/
* Trisha's Books: https://trishagee.com/2023/07/31/summer-reading-2023/
* Dave Farley on Continuous Delivery: https://www.youtube.com/channel/UCCfqyGl3nq_V0bo64CjZh8g
Only a few can claim they have successfully created a Pure-Serverless architecture and only those really understand the challenges of observing real event driven architectures.
Apostolis Apostolidis (also known as Toli) is one of those people and its why we invited him back to discuss all the lessons learned from his time as Head of Engineering Practices at cinch. Tune in and learn about the evoluation of Serverless observability and the challenges when observing API Gateways, Queues and Step Functions. Listen to Toli's advice on picking one observability vendor, doing your own custom instrumentation and making yourself familiar with the observability data from your managed service provider.
Also go back to our previous episode to hear more from his Engineering Practices for Success and remember that the time to ask about coldstarts is over 🙂
Additional links we discussed today:
Previous Podcast with Toli: https://www.spreaker.com/user/pureperformance/unlocking-the-power-of-observability-eng
OpenTelemetry: https://opentelemetry.io/
AWS Step Functions: https://aws.amazon.com/step-functions/
Dynatrace Business Flow: https://www.youtube.com/watch?v=W0bSzvQrUzA
Codifying Golden Paths that ideally don't need you to build a K8s Operator! This is what Practical Platform Engineering should look like!
In our latest episode we learn from Maurico (Salaboy) Salatino who has been contributing to open source for the past 12 years.
Tune in and learn from his journey of designing and built platforms. He shares his opinion on the Platform Engineering skillsets, how to design for self-service, how to pick the right tools out of the 160+ CNCF project options and shares some of his favorite tools (including Crossplane, VCluster, Argo, OpenFeature, Keptn ...) that should be part of a modern cloud native platform.
Links discussed in this podcast:
* Salaboy on Twitter: https://twitter.com/salaboy
* Salaboy on LinkedIn: https://www.linkedin.com/in/salaboy/
* Upcoming Book: https://www.salaboy.com/book/
* Cloud-Native Snapshots: https://www.salaboy.com/cloud-native-snapshots/
* Diagrid: https://www.diagrid.io/
Reducing the cognitive load by simplifying computing for every developer in an organization! One of the many definitions of Platform Engineering. But what is Platform Engineering for real? Just a new hype? What problem does it really solve? How does it link with DevOps and SRE? Are there any standards or reference architectures available?
To get a new perspective on Platform Engineering we invited Saim Safdar, CNCF Ambassador and member of the CNCF TAG App Delivery Platform Working Group. Tune in and learn about the Platform Maturity Model, how to get involved to shape the field of Platform Engineering, what other people that Saim has interviewed are good to follow and much more ..
Here the links we discussed:
CNCF Platforms White Paper: https://tag-app-delivery.cncf.io/whitepapers/platforms
Maturity Model Working Document: https://docs.google.com/document/d/1bP8-LQ-d41eIdQB3IC2YsncDhawpFLggql2JxwtE0XI/edit
Platform Working Group: https://tag-app-delivery.cncf.io/about/wg-platforms/
Cloud Native Podcast with Alexis Richardson: https://www.youtube.com/watch?v=p6D-NYkVp9E
Patterns and Anti-Patterns: https://octopus.com/devops/platform-engineering/patterns-anti-patterns/
Saim on LinkedIn: https://www.linkedin.com/in/saim-safder/
Are you frustrated with your team's ability to troubleshoot issues in production despite their proficiency in pushing out new builds? The root of this problem may lie in the absence of Observability Driven Development.
In our latest episode we are joined by Apostolis Apostolidis (also known as Toli) who - as Head of Engineering Practices at cinch - has spent his past years enabling teams to adopt the easiest path to value. He is passionate about DevOps and has a strong opinion on how to educate engineers on "Consciously Instrumenting Code for good Observability".
Tune in learn more about good engineering practices, building internal communities of practice, the benefits of traces over metrics and logs and why we need to start adding observability to our CVs and LinkedIn profiles.
Here are all relevant links we discussed in this episode
Tolis Website: https://www.toli.io/
Tolis LinkedIn Profile: https://www.linkedin.com/in/apostolosapostolidis/
Toli on Twitter: https://twitter.com/apostolis09/
WTFisSRE Talk on DevOps Meets Service Delivery: https://www.youtube.com/watch?v=nLrx0BCMl0Y
GOTO talk on EDA in Practice: https://www.youtube.com/watch?v=wM-dTroS0FA
Do you know why customers spend more money at a pub when ordering at a table vs ordering directly from at the bar tender? Do you want to know how to get SaaS vendors to send you their observability & telemetry data? Do you want to know the career path of how an Infrastructure Analyst turned Digitial Readiness Manager?
Tune in to this PurePerformance episode where we sat down with Mark Forrester from Mitchell & Butlers answering all these questions and also drawing the parallels to Observability.
Because observability has come a long way just as Mark: From traditional infrastructure (CPU, Memory, Network) to APM (Service Response Time & Failure Rates), to Real User Behaviour and now End-2-End Business Processes Analytics. Unlocking the potential of Digitial Business Observability lets Mark optimize the end-2-end customer journey to make sure their customers always feel like they are taken care of when trying to order online food delivery, a meal or a drink at a restaurant. As you learn, digital business observability goes beyond your own digital premise and needs to tap into the data of your 3rd party suppliers and SaaS vendors.
To see more from Mark also check out his interview at Dynatrace Perform: https://www.youtube.com/watch?v=rGpduOrPxpU
As far as we know - besides Kubernetes there is only Prometheus that belongs to the prestigious group of open-source projects that have their own documentary. Now why is that?
Prometheus has emerged as the go-to solution for capturing metrics in modern software stacks, earning its status as the de facto standard. With its widespread adoption and a constantly expanding ecosystem of companion tools, Prometheus has become a pivotal component in the software development landscape.
Join us as we sit down with Björn Rabenstein, an accomplished engineer at Grafana, who has dedicated nearly a decade to actively contributing to the Prometheus project. Björn takes us on a journey through the project's early days, unravels the reasons behind its meteoric rise, and provides us with insightful technical details, including his personal affinity for Histograms.
Here are the links we discussed during the podcast for you to follow up:
* Prometheus Documentary: https://www.youtube.com/watch?v=rT4fJNbfe14
* First Prometheus talk at SRECon 2015: https://www.youtube.com/watch?v=bFiEq3yYpI8
* The Zen of Prometheus: https://the-zen-of-prometheus.netlify.app/
* Talk from Observability Day KubeCon 2023: https://www.youtube.com/watch?v=TgINvIK9SYc
* Secret History of Prometheus Histograms: https://archive.fosdem.org/2020/schedule/event/histograms/
* Prometheus Histograms: https://promcon.io/2019-munich/talks/prometheus-histograms-past-present-and-future/
* Native Histograms: https://promcon.io/2022-munich/talks/native-histograms-in-prometheus/
* PromQL for Histograms: https://promcon.io/2022-munich/talks/promql-for-native-histograms/
APIs are powering and empowering software innovation as they enable new use cases on top of existing services. Observability into API usage to answer questions like: how APIs are called, what APIs do, where APIs fail, where APIs are slow, where APIs are misused … has to be on top of mind for architects that decide to build or use APIs.In this episode we welcome Sonja Chevre, Group Product Manager at Tyk, who recently gave a captivating talk at KubeCon about using OpenTelemetry to get insights into popular API frameworks such as GraphQL. We are discussing common challenges for SREs such as that APIs often hide the status of a call behind an HTTP 200 or that debugging individual calls is really hard as details of the call are not exposed by default to telemetry data. We also cover topics such as API-led growth, API as a product as well as open standards such as OpenTelemetry and OpenAPI. Here the list of discussed links during the show:* KubeCon Talk: https://kccnceu2023.sched.com/event/1HyVc/what-could-go-wrong-with-a-graphql-query-and-can-opentelemetry-help-sonja-chevre-ahmet-soormally-tyk-technologies * API Management vs Gateway discussion: https://www.linkedin.com/posts/sonjachevre_apimanagement-apigateway-apisecurity-activity-7061596404854521857-N_cO/ * API as a Product: https://tyk.io/blog/unlocking-the-potential-of-api-as-a-product/ * OpenAPI: https://www.openapis.org/ * OpenTelemetry: https://opentelemetry.io/
Security comes with a price tag, such as additional wait time when going through checks at the airport or when inspection network packages at your firewall.To learn about current approaches to cyber defense and cyber deception we invited back Stefan Achleitner, Lead Researcher Cloud Native Security at Dynatrace. Tune in and learn why it is important to keep changing and using different passwords, why you should monitor all your servers, what zero day vulnerabilities are, the role of eBPF in security and why we have to minimize false positives alarms like the Hawaii Missile Alert! Some of the links we discussed during the podcast can be found here:* Our previous episode: https://www.spreaker.com/user/pureperformance/don-t-look-away-from-the-next-cyber-secu * Hawaii false missile alert: https://en.wikipedia.org/wiki/2018_Hawaii_false_missile_alert * eBPF on isitobservable: https://isitobservable.io/search?q=eBPF * Check if you've been compromised: https://haveibeenpwned.com/ * Stefan’s SolarWinds Article (German): https://intelligente-welt.de/so-funktionierte-der-angriff-auf-solarwinds/ * Stefan's Problem wiht Passwords Article (German): https://intelligente-welt.de/passwort-manager-und-andere-loesungen-fuers-passwort-chaos/
36 million generated OpenTelemetry spans per hour for GraphQL based queries – that’s just one of the stats we discussed with Justin Scherer, Sr Developer and Consultant, who is leading OTel adoption and Shift-Left observability efforts at NWM. For Justin, OpenTelemetry helps commoditize data gathering in modern cloud native environments so that the backend observability platform of choice can focus on answering higher level business impacting questions.If you are about to roll out OpenTelemetry in your organization then take the advice from Justin such as: Bringing Business Leaders early into the discussion! Engage with the OpenTelemetry community! Understand what your Observability Platform already gives you and focus on the gaps! To learn more about OpenTelemetry check out some of the links we discussed during the podcast:
* OpenTelemetry Website: https://opentelemetry.io/
* IsItObservable: https://isitobservable.io/open-telemetry
* Podcast: https://www.spreaker.com/user/pureperformance/adopting-open-observability-across-your-
* LinkedIn Profile: https://www.linkedin.com/in/justin-scherer-198126160/
Organizations that experience Monitoring Data Obesity – having too many arbitrary logs or metrics without context – are suffering twice: high cost for storage and not getting the answers they need!OpenTelemetry, the cloud native standard for observability, solves those challenges and therefore sees rapid adoption from both startups and established enterprises.In this episode we have Daniel Gomez Blanco (@dan_gomezblanco), Principal Software Engineer at Skyscanner and author of the recently published book Practical OpenTelemetry.Tune in and learn about the latest status of OpenTelemetry, lessons learned from adopting OpenTelemetry in a large organization, considerations between metrics and traces, the difference between statistical and tail based sampling and much more Here the links we discussed during the episode:* Chat I had with M. Hausenblas on his podcast the other day: https://inuse.o11y.engineering/episode/meet-daniel-skyscanner * Link to QCon talk (although I believe the video won't be made available till later in the year) https://qconlondon.com/presentation/mar2023/effective-and-efficient-observability-opentelemetry * Recent InfoQ interview covering the talk: https://www.infoq.com/news/2023/03/effective-observability-otel/ * Video on a talk I did with Ted Young a couple years ago during our tracing migration to OpenTelemetry: https://youtu.be/HExcLWA2b8M * Talk at o11yfest 2021 on our tracing migration to OTel: https://vi.to/hubs/o11yfest/videos/3143 * Mastodon: https://mas.to/@dan_gomezblanco
DevOps didn’t die when the world started raving about SRE. And while some proclaim that platform engineering finally kills DevOps it is more an evolutionary process to bring DevOps practices to a new audience that is building and running apps on a new technology stack.
But what about ChatGPT? Can it be the best DevOps engineer you ever had? Will it be able to build and optimize our delivery pipelines? Will it tell us which products to build and how? Which architecture to choose and how to best design it for operations?
Tune in and hear from Stephen Thair, DevOps Thought Leader and Founder of DevOpsGroup, on what he has seen over the past decade working in the DevOps space and why he thinks that while ChatGPT will be disrupting many jobs it is a great opportunity to boost creativity and efficiency for many DevOps and non DevOps folks
Also don’t miss to read Stephen’s 2023 predictions we mentioned in our discussion
The famous tagline from Werner Vogel in 2006 is still used in many presentations promoting DevOps and the autonomy of development teams. But how long does and did this really scale?
Based on our guest Luca Galante, Head of Product at Humantic, organizations that reach 50-100 engineers start experiencing the first bottlenecks. After initial workarounds sometimes leading to Shadow Ops it’s the time where organizations look into building Internal Development Platforms (IDP). This is where Platform Engineering is born by providing “Golden Paths around DevOps & SRE” as a self-service to engineering teams.
Tune in an learn more about the emerging practice of platform engineering, why it already attracted more than 11000 global community members, has an annual dedicated conference and why global analysts are putting Platform Engineering in the Top Trends of 2023! We referenced a lot of material in our discussion. Here all the promised links:
What is Platform Engineering: https://platformengineering.org/blog/what-is-platform-engineering
Platform Engineering Community: https://platformengineering.org
PlatformCon: https://platformcon.com/
Platform Weekly: https://platformweekly.com/
Follow Luca on Twitter: https://twitter.com/luca_cloud
Connect with Luca on LinkedIn: https://www.linkedin.com/in/luca-galante/
While Spring4Shell, Ransomware and attacks on critical infrastructure were the most severe attacks in 2022 the evolving trends in 2023 are around the rising power of AIs, complexity and therefore misconfiguration of cloud native stacks as well as social engineering challenges as part of the post-pandemic shift back towards the office.Tune in and learn from Stefan Achleitner, Lead Researcher Cloud Native Security at Dynatrace, about getting better in securing software supply chain, understanding the impact of attacks and vulnerabilities and why nobody should look away when it comes to detecting and preventing cyber security threats
How do you prepare yourself for the next incident? Not at all? Are you running game days where you simulate incidents? Or are you following the steps of good musicians who are constantly practicing with their band members to always be best prepared for the next big gig!Tune in and hear from Matt Davis, Specialist in Learning from Incidents, how he runs weekly continuous practice and learning sessions with DevOps, SREs, Developers, Marketers or Technical Writers and what the outcomes are.Matt is a regular presenter at conferences. You can meet him at SRECon Americas 2023 where he talks about “Human Observability of Incident Response” Here the other links we discussed during the podcast:* Practice of Practice * Rivers of Opposites * Varieties of Work * Follow Matt on Twitter * Connect on LinkedIn
Did you know that almost 60 years after IBM presented the mainframe 92 of the worlds top 100 banks run mainframes handling 90% of all credit card transactions? We didn’t either until we recorded this episode with Christian Schram, Solutions Engineer at Dynatrace, who has spent the last 20+ years helping organizations optimizing their mainframe environments. Tune in and learn about the mainframe, how the cloud native project OpenTelemetry has made it to the mainframe and what the most common performance patterns are on the mainframe.As discussed check out the following links in case you want to learn more:* A Brief History of the Mainframe World (Blog) * Modernizing the Mainframe (YouTube) * Eliminating inefficiencies on IBM Z (Blog) * End-2-End IBM Z transactional visibility (Blog)
Do you know that 53% of security related issues on Kubernetes are caused by misconfiguration? Me neither!To raise the awareness of how to protect your Kubernetes cluster and workloads from being hijacked we invited Nico Meisenzahl, Microsoft MVP and GitLab Hero, to walk us through a set of best practices that everyone in cloud native should know to contribute to a more secure cloud native environment. In our conversation we cover a lot of what Nico has shown in his recent talks at different container, cloud native and security related conferences.Make sure you check out the slides, github tutorials and recordings from Nico through those links:* Nico’s Website: https://meisenzahl.org/ * Hijack a Kubernetes Cluster YouTube: https://www.youtube.com/watch?v=9wc34MozKok * Hijack a Kubernetes Cluster Slides: https://www.slideshare.net/nmeisenzahl/containerconf-2022-hijack-kubernetes * Hijack a Kubernetes Cluster GitHub Tutorial: https://github.com/nmeisenzahl/hijack-kubernetes * Connect with him on LinkedIn: https://www.linkedin.com/in/nicomeisenzahl/ * Follow him on Twitter: https://twitter.com/nmeisenzahl
If you want to hear more from Nico listen until the end and pick from one of the suggested topics
Incidents happen! And when asking Laura Nolan who was an SRE at Google and Slack, healthy organizations should take proper time to analyze and learn from them. This will improve future incident response as well as overall system resiliency.Tune in to this episode and hear Laura’s tips & tricks what makes a good SRE organization. It starts with doing good write ups of incidents, doing your research on incident reports of software and services that you are looking into using. We also spent a good amount of time discussing root cause analysis where she highlighted an incident that happened at her time at Google and what she learned about outdated alerting.Thanks Laura for a great discussion and lots of insights.
Here are the additional links we discussed during the podcast
* Laura on LinkedIn: https://www.linkedin.com/in/laura-nolan-bb7429/
* Laura on Twitter:https://twitter.com/lauralifts
* Incident Template talk @ SRECon: https://www.usenix.org/conference/srecon22emea/presentation/nolan-break
* What SRE could be talk @ SRECon: https://www.usenix.org/conference/srecon22emea/presentation/nolan-sre
* Howie Post-Incident Guide: https://www.jeli.io/howie/welcome
* My philosophy on Alerting article: https://docs.google.com/document/d/199PqyG3UsyXlwieHaqbGiWVa8eMWi8zzAn0YfcApr8Q/edit
What a year 2022 was! We had 25! episodes with amazing guests from all over the world covering topics from Kubernetes, OpenTelemetry, DevOps, SRE, Cloud Migrations, DNS, Value Streams all the way to Persona Driven Engineering and drawing parallels with Digital Marketing. If you are new to our podcast check out the playlist and listen to some of those we mentioned during our episode!Now its time to say Thank You listeners for the continued support. After 5+ years of podcasting we still see rising numbers of downloads which is the best motivation for us to keep going. Stay tuned as we are going to cover industry relevant topics going into 2023 – or is it year 53? (only those will know that listen to the full episode)
“If I wouldn’t measure it I wouldn’t know it!” or “Build, Measure, Learn! ”These quotes could be from any engineer building new digital services, observing them in production and based on that learn how to improve their software.They are however from Bernhard Dominguez, Digital Consultant at FACTOR, who we invited to the show. Bernhard highlights a lot of parallels between his work planning and executing digital marketing strategies and the world we live in: designing, operating and optimizing complex software systems.Tune in and learn about how important it is to understand your real target groups (=end users), how to define clear goals (=SLOs), how to change from campaign to funnel activities (=User Journeys) and why it is so important to get an outsider’s opinion before implementing your next big project! (=We have always done it this way) If you want to follow up with Bernhard and his work check out the following links we discussed during the podcast:* Bernhard on LinkedIn * FACTOR * Podcast (German): Newsletter Marketing * Podcast (German): Build - Measure - Learn
What’s the difference between People Operations and Human Resources? Why should we stop saying “resources” when we refer to people? And how do you create a unique employee experience for an international company? Anyone coming into an organization really is asking three things: •What is the culture of the organization?•How will they grow and develop within the company? and •How will they be rewarded and recognized for their great work? And these three questions are reflected in three pillars at the base of a company’s employee experience: 1.Culture2.Growth & Opportunity3.Reward & RecognitionSue Quackenbush from Dynatrace gives us an insight into the life of a Chief People Officer, answering those questions and more in today’s episode of Behind The Code.
You have a CISO (Chief Security Information Officer) but no CRO (Chief Reliability Officer)? You blame people if systems crash? You scale your people in the rate of scaling your infrastructure? If you answer any of those questions with YES then you should tune into this podcast as you probably struggle adopting Site Reliability Engineering (SRE) in your organization.
James Brookbank, Cloud Solutions Architect, has dealt with resiliency topics in a large enterprise prior to joining Google. In our conversation he shares advice he gives Enterprises to convert the excitement about SRE into actual implementation. James gave some good guidance on what good and not so good projects are to start with. He gives practical examples on what it means to change your company culture and why there doesn’t have to be an SRE for every service.
In our call we discussed the SRE in Enterprise talk at DevOpsDays Boston and SRECon EMEA as well as their recent book. Here are all the relevant links:
James Brookbank on Linkedin:
https://www.linkedin.com/in/jamesbrookbank/
SRECon EMEA Slides:
https://www.usenix.org/system/files/srecon22_slides_mcghee.pdf
DevOpsDays Boston 2022 Session Recording: https://www.youtube.com/watch?v=__e7b25QOHc
Enterprise Roadmap to SRE Book:
https://sre.google/resources/practices-and-processes/enterprise-roadmap-to-sre/
Dynatrace recently announced Grail – promising boundless observability, security and business analytics in context.
You may think: that’s a lot of nice words that other solutions claim as well. So why should you care about Grail? What is the real problem it solves and how does it solve it?
Tune in and hear from Andreas Lehofer, Chief Product Officer at Dynatrace as he boils it down to two critical issues:
Thanks Andreas for the discussion, the insights on the hidden costs of current approaches, the technical explanation on our architecture as well as giving us some glimpse on what’s coming next.
Show Links:
Dynatrace Grail Announcement:
https://www.dynatrace.com/platform/grail/
Andreas Lehofer on Linkedin:
https://www.linkedin.com/in/andreaslehofer/
Agile has become part of the everyday vocabulary in software development nowadays. But there is still a lot of mystery hidden behind the word. So let’s learn more about it!
We invited Agile experts Andreas Mitter and Julia von Spreckelsen from BearingPoint to join us on the podcast and discuss with us what being Agile actually means, how they became Agile consultants, and what's in store for Agile in the future.
“I was not that interested in coding but more in understanding the impact of software on human beings” says Diana Najda, SRE & Monitoring Lead, when we asked her how she ended up leading the efforts around Site Reliability Engineering.
Tune in to our conversation and learn how Diana is bridging the gap between Dev, Ops and Business by ensuring that the right people get the right telemetry data from their observability platform. She gives us insights into her definition of DevOps and SRE, how she helps teams setting up SLOs (Service Level Objectives) and how she proves the ROI (Return On Investment) into the SRE practices!
Last piece of advice Diana gives everyone interested: “SRE might be buzzword it loses the buzz the more you hear it – BUT - its really cool because SREs make the life of Dev and Ops easier every day”
If you want to connect with Diana reach her on LinkedIn: https://www.linkedin.com/in/diannajda/
“I was not that interested in coding but more in understanding the impact of software on human beings” says Diana Najda, SRE & Monitoring Lead, when we asked her how she ended up leading the efforts around Site Reliability Engineering.
Tune in to our conversation and learn how Diana is bridging the gap between Dev, Ops and Business by ensuring that the right people get the right telemetry data from their observability platform. She gives us insights into her definition of DevOps and SRE, how she helps teams setting up SLOs (Service Level Objectives) and how she proves the ROI (Return On Investment) into the SRE practices!
Last piece of advice Diana gives everyone interested: “SRE might be buzzword it loses the buzz the more you hear it – BUT - its really cool because SREs make the life of Dev and Ops easier every day”
If you want to connect with Diana reach her on LinkedIn: https://www.linkedin.com/in/diannajda/
Serverless and other emerging technologies hide the complexity of the underlying runtimes from developers. This is great for productivity but can make it really hard when troubleshooting behavior that needs deeper insight into those runtimes, platforms or frameworks.
In this episode we hear from Kam Lasater, Founder of Cyclic Software. Kam has run into several walls while he was implementing solutions from scratch using Serverless technologies as well as other popular cloud services. He recently presented a handful of those scenarios at DevOpsDays Boston 2022.
Tune in and learn from Kam as he walks us through two of those challenges he covered during his DevOpsDays talk. If you want to learn more make sure to watch the full talk on YouTube: https://www.youtube.com/watch?v=xB9vsSl93mE
If you want to learn more from or about Kam check out the following links:
YouTube video from DevOpsDays Boston: https://www.youtube.com/watch?v=xB9vsSl93mE Cyclic Website: https://www.cyclic.sh/ Cyclic Blog: https://www.cyclic.sh/blog/ Twitter: https://twitter.com/seekayel Personal Website: https://kamlasater.com/ LinkedIn: https://www.linkedin.com/in/kamlasater/
Over the years we learned how to optimize the performance of our JVMs, our CLRs or our databases instances by tweaking settings around heap sizes, garbage collection behavior or connection and thread pools.
As we move our workloads to k8s we need to adapt our optimization efforts as they are new nobs to turn. We need to factor in how resource and request limits on pods impact your application runtimes that run on your clusters. Out of memory problems are all of a sudden no longer just depending on the java heap size alone!
To learn more about k8s optimization best practices we have invited Stefano Doni, CTO of Akamas. Stefano walks us through key learnings as the team at Akamas has helped organizations optimize the performance, resiliency and cost of their k8s workloads. You will learn about proper memory settings, CPU throttling and how to start saving costs as you move more workloads to k8s.
To learn more about Akamas go here: https://www.akamas.io/
If you happen to be at KubeCon 2022 in Detroit make sure to visit their booth
Show Links:
Stefano on Linkedin: https://www.linkedin.com/in/stefanodoni/
A Guide to Autonomous Performance Optimization with Dynatrace and Akamas: https://www.youtube.com/watch?v=i7MuEjeOvX0
In economic turbulent times leaders get asked questions like: “What’s the return on investment of your DevOps or Cloud Transformation? Did we really get better and more efficient? Or did we just blow a lot of money out the window?”
Connecting business results with your technical initiatives is what would answer those questions. To learn how this works we invited Adam Dahlgren, SVP Product at Allstacks. From Adam we learn about Value Stream Management, how to align with your top level OKRs and how to improve your DORA and SPACE metrics. Because as Adam says in the beginning: “Inspection is coming especially during turbulent economic times and they will question your investment in transformation projects!”
If you want to follow up with Adam check out the following links we discussed:
LinkedIn: https://www.linkedin.com/in/adam-dahlgren/ What are DORA Metrics: https://www.allstacks.com/blog/dora-metrics/?hsLang=en What is the SPACE Framework: https://queue.acm.org/detail.cfm?id=3454124 Allstack: https://www.allstacks.com/ DevOps World sessions from Allstack: https://events.devopsworld.com/widget/cloudbees/devopsworld22/conferenceSessionDetails?tab.day=20220929&search=dora
In economic turbulent times leaders get asked questions like: “What’s the return on investment of your DevOps or Cloud Transformation? Did we really get better and more efficient? Or did we just blow a lot of money out the window?”
Connecting business results with your technical initiatives is what would answer those questions. To learn how this works we invited Adam Dahlgren, SVP Product at Allstacks. From Adam we learn about Value Stream Management, how to align with your top level OKRs and how to improve your DORA and SPACE metrics. Because as Adam says in the beginning: “Inspection is coming especially during turbulent economic times and they will question your investment in transformation projects!”
If you want to follow up with Adam check out the following links we discussed:
LinkedIn: https://www.linkedin.com/in/adam-dahlgren/
What are DORA Metrics: https://www.allstacks.com/blog/dora-metrics/?hsLang=en
What is the SPACE Framework: https://queue.acm.org/detail.cfm?id=3454124
Allstack: https://www.allstacks.com/
DevOps World sessions from Allstack: https://events.devopsworld.com/widget/cloudbees/devopsworld22/conferenceSessionDetails?tab.day=20220929&search=dora
We all want to leverage technology to solve problems. New and shiny toys are appealing to look which sometimes means we loose the insights on the base technologies that powers most of our connected lives, such as DNS or TLS.
In this podcast we invited Philipp Krenn (@xeraa), Dev Advocate Team Lead at Elastic, and learn about DNS, TLS and other bad config changes. We learn about Log4Shell, how the Java Security Manager was a big help in fighting Log4Shell, why its been deprecated and also get his thoughts into CDD (Conference Driven Development)
And if you ever visit Vienna – chances are you meet Philipp dancing Waltz with tourists 😊
Show Links: To learn more from Philipp start with
His personal website: https://xeraa.net/ Twitter: https://twitter.com/xeraa LinkedIn: https://www.linkedin.com/in/philippkrenn His conference schedule (past & future): https://xeraa.net/events/
We all want to leverage technology to solve problems. New and shiny toys are appealing to look which sometimes means we loose the insights on the base technologies that powers most of our connected lives, such as DNS or TLS.
In this podcast we invited Philipp Krenn (@xeraa), Dev Advocate Team Lead at Elastic, and learn about DNS, TLS and other bad config changes. We learn about Log4Shell, how the Java Security Manager was a big help in fighting Log4Shell, why its been deprecated and also get his thoughts into CDD (Conference Driven Development)
And if you ever visit Vienna – chances are you meet Philipp dancing Waltz with tourists 😊
Show Links:
To learn more from Philipp start with
His personal website: https://xeraa.net/
Twitter: https://twitter.com/xeraa
LinkedIn: https://www.linkedin.com/in/philippkrenn
His conference schedule (past & future): https://xeraa.net/events/
What can we learn from the hospitality industry to create a great workplace? Should companies treat their employees the same way as customers? And why should we stop talking about “new work”?
Senior Director of R&D Lab Operations at Dynatrace. Veronika Leibetseder. shares how her experience working in the luxury hotel business inspired her to join the tech workforce and create the future workplace. Her motto? "Your employees are your customers." As much as you want to deliver a great customer experience, you should aim to do the same for your employees.
How do you a design a feature if you don’t know for whom it is for? How do you define SLOs (Service Level Objectives) if you don’t know what your users expect from you? How do you design performance tests and workloads if you don’t know which user behavior to simulate?
In this episode we have Barbara Ogris, Sr Product Experience Designer at Dynatrace, who walks us through the concept of target personas that she helped establish within Dynatrace. It changes product and observability discussions from “as a user I want …” towards “as Archie I have this need …”. Listen in and learn about design thinking, using empathy maps to define your target persona and how this can be applied to many aspects in software engineering.
Barbara on Linkedin
https://www.linkedin.com/in/barbara-ogris-6a0b6011b/
Dynatrace Blog: Terminology matters: how to enhance user experience by aligning names with expectations
https://www.dynatrace.com/news/blog/terminology-matters-how-to-enhance-user-experience-by-aligning-names-with-expectations/
Atlassian persona template: https://www.atlassian.com/software/confluence/templates/persona
MIRO persona template: https://miro.com/aq/ps/templates/personas/?utm_source=google&utm_medium=cpc&utm_c[…]aIQobChMI6J7Juqql-QIVCOJ3Ch3jtwWwEAAYASAAEgLxT_D_BwE&loc=9062705
Adobe XD: how to define a persona: https://xd.adobe.com/ideas/process/user-research/putting-personas-to-work-in-ux-design/
How do you a design a feature if you don’t know for whom it is for? How do you define SLOs (Service Level Objectives) if you don’t know what your users expect from you? How do you design performance tests and workloads if you don’t know which user behavior to simulate?
In this episode we have Barbara Ogris, Sr Product Experience Designer at Dynatrace, who walks us through the concept of target personas that she helped establish within Dynatrace. It changes product and observability discussions from “as a user I want …” towards “as Archie I have this need …”. Listen in and learn about design thinking, using empathy maps to define your target persona and how this can be applied to many aspects in software engineering.
Barbara on Linkedin https://www.linkedin.com/in/barbara-ogris-6a0b6011b/
Dynatrace Blog: Terminology matters: how to enhance user experience by aligning names with expectations https://www.dynatrace.com/news/blog/terminology-matters-how-to-enhance-user-experience-by-aligning-names-with-expectations/
Atlassian persona template: https://www.atlassian.com/software/confluence/templates/persona
MIRO persona template: https://miro.com/aq/ps/templates/personas/?utm_source=google&utm_medium=cpc&utm_c[…]aIQobChMI6J7Juqql-QIVCOJ3Ch3jtwWwEAAYASAAEgLxT_D_BwE&loc=9062705
Adobe XD: how to define a persona: https://xd.adobe.com/ideas/process/user-research/putting-personas-to-work-in-ux-design/
SRE vs DevOps, SRE or DevOps or is it SRE & DevOps? No better person to ask than somebody that has been an SRE for much longer than our industry is talking about Site Reliability Engineering.
Michael Wildpaner, Sr Engineering Director Cloud Security at Google, started as an SRE for Google Maps back in 2006. Fast forward to 2022 Michael has a lot of hands-on experience about the SRE role, the different levels of SRE that one organization can apply and how it connects with DevOps.
Tune in and hear his personal stories from more than 15 years at Google. While not everyone is Google – there for sure is a lot we can take out of this conversation.
Here some of my personal take aways
Core idea of SRE: take engineers that understand distributed systems and “annoy” / guide developers to build better resilient systems from the start
Design for automation: this already starts with naming your infrastructure (aka – don’t use lord of the rings names)
SREs help so that you DO NOT DESIGN yourself into a corner
Observability is the foundation of good SRE as it enables incident management, insights all the way up to user insights
Tip: Ensure new hires understand that you have a blameless culture
As follow up material check out those links
LinkedIn: https://www.linkedin.com/in/michael-wildpaner
Talk at DevOps Fusion 2022: https://devops-fusion.com/en/speaker/michael-wildpaner/
SRE vs DevOps, SRE or DevOps or is it SRE & DevOps? No better person to ask than somebody that has been an SRE for much longer than our industry is talking about Site Reliability Engineering.
Michael Wildpaner, Sr Engineering Director Cloud Security at Google, started as an SRE for Google Maps back in 2006. Fast forward to 2022 Michael has a lot of hands-on experience about the SRE role, the different levels of SRE that one organization can apply and how it connects with DevOps.
Tune in and hear his personal stories from more than 15 years at Google. While not everyone is Google – there for sure is a lot we can take out of this conversation.
Here some of my personal take aways
Core idea of SRE: take engineers that understand distributed systems and “annoy” / guide developers to build better resilient systems from the start Design for automation: this already starts with naming your infrastructure (aka – don’t use lord of the rings names) SREs help so that you DO NOT DESIGN yourself into a corner Observability is the foundation of good SRE as it enables incident management, insights all the way up to user insights Tip: Ensure new hires understand that you have a blameless culture
As follow up material check out those links
LinkedIn: https://www.linkedin.com/in/michael-wildpaner Talk at DevOps Fusion 2022: https://devops-fusion.com/en/speaker/michael-wildpaner/
For some out there SLOs (Service Level Objectives) are the silver bullet to building and operating reliable software. But nothing is as shiny on the inside as it looks on the outside.
In this episode we invited Stephen Townshend, former Performance Engineer now converted to Site (Slight) Reliability. Stephen (@the_kiwi_sre) has experienced the tough side of establishing SLOs within an organization. It’s a constant battle between focusing on reliability and new features and a lack of change in culture.
Listen in and learn about the 9 pre-requisites for SLOs that Stephen has identified such as: having a certain level of observability, define clear business objectives, define ownership and give autonomy or establishing a blameless culture
Stephen on Linked in
https://www.linkedin.com/in/stephentownshend/
Stephen on Twitter
https://twitter.com/the_kiwi_sre
Here the additional resources we brought up during our talk:
Slight Reliability YouTube: https://www.youtube.com/c/SlightReliability
Slight Reliability Podcast: https://www.buzzsprout.com/1698445
Our LinkedIn discussion: https://www.linkedin.com/posts/scottmooreconsulting_7-steps-to-identify-and-implement-effective-activity-6938919857459462144--RI7
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Security is everyone’s business. And as everyone seems to be moving to Cloud Native it's important to understand what the security landscape in k8s, containerized apps, serverless, … looks like.
To learn more about this we invited Anais Urlichs (@urlichsanais), Developer Advocate at Aqua Security and CNCF Ambassador of the year 2021. Over the past years Anais has educated thousands of people on cloud native, devops and security on her YouTube Channel.
Tune in and learn more about the different approaches to security in cloud native, which open source projects are out there and how her advise on embedding security in your day2day work.
Some additional links we discussed can be found here:
Anais on Linkedin: https://www.linkedin.com/in/urlichsanais/ Anais on Twitter: https://twitter.com/urlichsanais Trivy: https://github.com/aquasecurity/trivy Weekly DevOps Newsletter: https://anaisurl.com/ WTFisSRE Talk: https://www.youtube.com/watch?v=0zL61AiOaK0 Anais’s YouTube channel: https://www.youtube.com/c/AnaisUrlichs Aqua Open Source YouTube Channel: https://www.youtube.com/channel/UCb4mfRT5UWpjoUQRcIE2qOQ
Today’s episode is about leadership and how to prepare yourself for moments when you might have to leave for a longer period of time. It could be due to family issues, maternity or paternity leave, going back to school, etc. In many cases, this can be a career killer. You need somebody to take over the responsibility, but you want to keep your leadership role and still have it for when you come back. At the same time, you want to give somebody else the opportunity to grow.
Anita, VP of Delivery, and Thomas, Director of ACE, both work at Dynatrace and have spent the past couple of years sharing leadership in their team. In this podcast episode, we will learn from their experience of how they made it happen and what they’ve learned along the way.
While this episode started out with a recap of April Edwards (@TheAprilEdwards) keynote called “Putting the Ops into DevOps” we quickly got April talk about what measures Microsoft has set to embrace the cultural change needed for their DevOps transformation: Every service has a public health dashboard, putting the customer in the center, make products open source, eat your own dog food, align your objectives with the team, …
Besides this great conversation that finally gave some great input on what cultural change really looks like we learned from her background in Ops, moving to Dev, getting into the cloud and now inspiring Ops teams to have it easier in their job using automation. Tune in, learn and get inspired.
We also talked about the late Abel Wang and how Microsoft UK is supporting Girls Who Code.
Show Links:
April on Linkedin
https://www.linkedin.com/in/azureapril/
April on Twitter
https://twitter.com/TheAprilEdwards
Putting the Ops into DevOps keynote
https://globalazure.at/sessions/#323994
Supporting Girls Who Code in memory of Abel Want
https://www.justgiving.com/fundraising/msbuild2022/?WT.mc_id=modinfra-67727-apedward
With companies needing constant growth to survive and societal pressures pushing people to deliver more and more, it’s becoming more and more common to hear that people are living under extremely high amounts of stress that are leading to burnout.
Julia Simon, community member of the Cloud Native Computing Foundation and is leading the burnout support group, shares her vulnerable story on how it felt to be burned out, how she got there, and how she successfully overcame it. Now, she's using her experience to help others.
Watch Julia's talk at KubeCon 2021: https://www.youtube.com/watch?v=lpiXbfOTNYw Join the burnout support group: https://cloud-native.slack.com/archives/C02JR0MB4V8
Feature Flagging has gained a lot of momentum which we can observe by counting the number of feature flagging solutions. To ensure a good developer experience when implementing feature flags the CNCF OpenFeature project was launched during KubeCon 2022 in Valencia. It is aiming to provide a feature flag standard similar to what OpenTelemetry did for Observability.
Tune in to this podcast where we have two of the founding members Mike Beamer and Todd Baert explain why it was the right time to initiate the project, which problems it solves and what use cases feature flagging brings to organizations.
If you want to learn more about the project check out the following resources discussed during the podcast
WebSite: https://openfeature.dev/
GitHub: https://github.com/open-feature
Community: https://github.com/open-feature/community
ITPro Today Launch Coverage: https://www.itprotoday.com/testing-and-quality-assurance/open-source-openfeature-project-takes-flight-advance-feature-flags
What do you do when your company wants to open a new office location in a new city? What does it feel like to lead a lab with no employees? Where do you start to look for people to join your team?
These and more questions will be answered by Christian Werding and Florian Dorfbauer, Lab Leads of the Dynatrace offices in Vienna and Graz (Austria). With more than a decade of experience leading people under their belt, they will share what they learned from their experience and how they managed to achieve even more growth than originally predicted.
Introducing: Behind The Code The podcast that looks at the human side of tech.
Giulia and Alois are co-hosting this podcast to share stories of the behind the scenes in tech companies. Because there's a lot going on to make sure companies can produce great software.
Listen to this intro episode to find out what you can expect from Behind The Code.
How do you plan for unplanned work such as fixing systems when they unexpectedly break in production? Just like firefighters – the best approach to practice those situations so that you are better prepared when they happen.
In this episode we have Mandi Walls, DevOps Advocate at PagerDuty, explain why she loves Game Days where she is “practicing for the weird things that might happen”. Prior to her current role she worked for Chef and AOL – picking up a lot of the things she is now advocating for. In our conversation Mandi (@lnxchk) gives us insights into how to best prepare and run game days, shared her thoughts on what good chaos scenarios (unreliable backend, slow dns …) are and which health metrics (team health, # incidents out of hours, …) to look at in your current incident response to figure out what a good game day scenario actually is.
Mandi on Linkedin: https://www.linkedin.com/in/mandiwalls/
In our talk we mentioned a couple of resources – here they are:
Mandi’s talk at DevOpsDays Raleigh: https://devopsdays.org/events/2022-raleigh/program/mandi-walls
Ops Guides: https://www.pagerduty.com/ops-guides/
“The most significant body of my SRE work is architectural reviews, disaster and failover planning and help with SLIs and SLOs of applications that would like to become SRE supported.”
This statement comes from Hilliary Lipsig, Principal SRE at Red Hat, as her introduction to what the role of an SRE should be. Hilliary and her teams are helping organizations getting their applications cloud native ready so that the operational aspect of keeping a system up & running and within Error Budgets can be handled by an SRE Team.
Listen in to this episode and learn about the key advices she has for every organization that wants to build and operate resilient systems. And understand why every suggestion she makes has to be and will always be evidence-based!
In the talk we mentioned a couple of tools and practices. Here are the links:
Hilliary on Linkedin: https://www.linkedin.com/in/hilliary-lipsig-a5935245/
KubeLinter: https://docs.kubelinter.io
Listen to talk Helm and Back again: an SRE Guide to choosing from DevConf.cz: https://www.youtube.com/watch?v=HQuK6txYS3g
The world is slowly moving back to having on-site meetings and conferences – such as DevOpsDays in Raleigh, NC where Andi presented on “Oh Keptn, my Keptn”.
Besides presenting Andi also visited several organizations on his road trip through North Carolina and Texas. Listen in and learn what the adoption challenges of DevOps & SRE are, how to define good SLOs (Service Level Objectives) and how to explain the difference between containers and microservices.
Also check out the following links Brian and Andi discussed:
State of SRE Report: https://www.dynatrace.com/info/sre-report/
DevOpsDays Raleigh: https://devopsdays.org/events/2022-raleigh/program/andreas-grabner
SLOConf: https://www.sloconf.com/
WTFisSRE: https://www.cloud-native-sre.wtf/
Keptn: https://www.keptn.sh
OpenTelemetry, for some the biggest romance story in open source, as it took off with the merger of OpenCensus and OpenTracing. But what is OpenTelemetry from the perspective of a contributor? Listen to this episode and here it from Daniel Dyla, Co-Maintainer OTel JS and W3C Distributed Tracing WG, and Armin Ruech who is on the Technical Committee focusing on cross language specifications. They give us insights into what it takes to contribute and drive an open source projects and give us an update on OpenTelemetry, the current status, what they are working on right now as well as the near future improvements they are excited about.
Show Links:
The OpenTelemetry Project
https://opentelemetry.io/
Daniel Dyla
https://engineering.dynatrace.com/persons/daniel-dyla/
Armin Ruech
https://engineering.dynatrace.com/persons/armin-ruech/
List of instrumented libraries
https://opentelemetry.io/registry/
Contribute to OTel
https://opentelemetry.io/docs/contribution-guidelines/
OpenTelemetry Tutorials on IsItObservable
https://isitobservable.io/open-telemetry
When moving to the cloud - have you thought of the performance difference between App Gateway and Application Load Balancers? The disk speed and disk cache limitations impacting Cassandra and or Elasticsearch Performance? Challenges with pre-built containers or resource limits on pods impacting Java Garbage Collection behavior?
These are all performance considerations Klaus Kierer, Senior Software Engineer in the Cluster Performance Engineering Team at Dynatrace, has learned over the past months as he helped performance optimize the Dynatrace Platform as it was expanded from running on AWS Compute to run on Kubernetes hosted in Azure (AKS) or Google Cloud (GKE).
Listen in and learn why Performance Engineering is more important than ever as you are moving your workloads to the “hyper-hybrid-cloud”.
Show Links:
Klaus on Linkedin:
https://www.linkedin.com/in/klaus-kierer-67b83a81/
Blog - When to use Azure Load Balancer or Application Gateway:
https://blog.siliconvalve.com/2017/04/04/when-to-use-azure-load-balancer-or-application-gateway/
K8ssandra performance benchmarks on cloud managed Kubernetes
https://k8ssandra.io/blog/articles/k8ssandra-performance-benchmarks-on-cloud-managed-kubernetes/
Lift and Shift seems to be “the easiest” cloud migration scenario but can quickly go wrong as we hear from Brian Chandler, Principal Sales Engineer at Dynatrace, in this episode.
Tune in and learn how latency can be the big killer of performance as you partially move services to the cloud. Brian (@Channer531) also reminds us about why you have to know about the N+1 query problem and the impact in cross cloud scenarios. Last but not least – Brian gives us insights into why Uber might be one of those companies who can change the SRE & SLO culture within partnering organizations.
Show Links:
Brian Chandler on Linkedin
https://www.linkedin.com/in/brian-chandler-8366663b/
Brian Chandler on Twitter
https://twitter.com/Channer531
Steve Tack has been leading Dynatrace Product Management for the past 10 years. He was one of the few Dynatracer’s delivering the key product announcements from Perform 2022 live from Vegas this year.
In this session we recap the key product announcements, which breakouts to watch and which keynotes you don’t want to miss. To make it easier to follow up follow these links:
On Demand sessions from Dynatrace Perform 2022
Product Announcements: Observability in Multi-Cloud Serverless, Software Intelligence as Code, DevSecOps Automation Alliances, Real-Time Security Attack Blocking
Show Links:
Steve Tack on Linkedin
https://www.linkedin.com/in/stevetack/
Dynatrace Perform On-demand videos
https://perform.dynatrace.com/2022-global
Observability for Multicloud Serverless Architectures
https://www.dynatrace.com/news/press-release/dynatrace-delivers-the-industrys-most-complete-observability-for-multicloud-serverless-architectures/
Software Intelligence as Code
https://www.dynatrace.com/news/press-release/dynatrace-delivers-software-intelligence-as-code/
DevSecOps Automation Alliance Partner Program
https://www.dynatrace.com/news/press-release/dynatrace-launches-devsecops-automation-alliance-partner-program/
Real-Time Attack Detection and Blocking
https://www.dynatrace.com/news/press-release/automatic-detection-and-blocking-of-attacks/
Do you regularly go to the gym or are you just wearing your sweat pants and sneakers at home and think that will do it? Or how about agile practices? Do you think by religiously attending your daily standup your colleagues think your performance testing all of a sudden happens within each sprint?
Leandro Melendez (aka Senor Performo), a DevRel Advocate for k6 load testing, tells us what he has seen in organizations he is helping to transform their performance engineering practices. The true benefit of becoming Agile, DevOps or whatever the next buzz word is, is to enable engineers with automated performance feedback on their changes. Listen in and also learn about Leandro’s ideas on PDD (Performance Driven Development).
Show Links
Leandro on Linkedin
https://www.linkedin.com/in/leandromelendez/
Señor Performo Site
https://www.srperf.com/
K6 Load Testing
https://k6.io/
“Whether open source or commercial – just focusing on logs, traces and metrics is limiting our conversation and missing the point what observability really is!”, says Dotan Horovits, Tech Evangelist at Logz, in his opening statement in this podcast. Listen an and learn more about why observability is not about collecting data. Observability is rather a data analytics problem as it needs to give humans answers to DevOps, SRE and Business questions.
To learn more beyond what was discussed in this podcast listen in to OpenObservability Talks, stay up to date on OpenTelemetry or follow Dotan at @horovits
Show Links
Dotan Horovits on Linkedin
https://www.linkedin.com/in/horovits/
Open Observability Talks
https://openobservability.io/
Open Telemetry Project
https://opentelemetry.io/
Dotan Horovits on Twitter
https://twitter.com/horovits
Who said that automation and data-driven decisions is only for DevOps or SREs? Scalability challenges or quality constraints are just as important to digital marketers like Nina Tollefson, Art Director at Dynatrace.
The pandemic caused many events to transform to a fully virtual. That was also true for Dynatrace’s flagship annual global user conference Perform in February 2021. To successfully transform and deliver a flawless experience for Perform attendees, Nina had to embrace automation in her field of work. Just as DevOps engineers Nina updated her tool chain. Figma was a new digital marketing platform of choice to automate, improve collaboration, become more efficient and scale her digital marketing work to support Perform. Today she can cash in on the automation investment she made last year as Perform 2022 is about to start. Thanks to data-driven digital marketing automation it looks like another amazing experience for our global Dynatrace user base. Thanks, Nina, for sharing your story and allowing us to draw many parallels to the DevOps & SRE world.
Show Links
Nina on Linkedin
https://www.linkedin.com/in/nina-tollefson/
Dynatrace Perform
https://www.dynatrace.com/perform-2022/
Log4Shell was an unwelcome early Christmas present for many IT teams around the globe. Asad Ali, Senior Director Dynatrace Sales Engineering, was involved starting December 9th – helping organizations around the world to react to the new vulnerability threat.
In our discussion we learn how the vulnerability works technically, how runtime AppSec vulnerability detection eliminates false/positives and how teams around the globe are preparing to protect their software supply chain for future vulnerabilities.
To learn more about Log4Shell check out Dynatrace’s Log4Shell Resource Center including educational blogs as well as Asad’s webinar on Detecting and Remediating Log4Shell with Dynatrace webinar.
Show Links
Asad Ali on Linkedin:
https://www.linkedin.com/in/alikingdom/
Dynatrace's Log4Shell Resource Center
https://www.dynatrace.com/resource-center/log4j-vulnerability
Detecting and Remediating Log4Shell with Dynatrace
https://info.dynatrace.com/global-all-wc-dynatrace-for-log4shell-18339-od-fulfillment.html
Encore Presentation - we'll be back in early 2022, until then, here is one of our favorite recent episodes:
To k8s or not – that should be the first question to answer before considering k8s. Granted – in many cases k8s is going to be the right choice but don’t just default to k8s because its hip or cool.
The rise of smart phones clearly created a new demand for “instant gratification” when it comes to interacting with online services through web sites or apps. To ensure services are available at any point in time without any interruption or delay it requires performance engineers to automate performance and scalability engineering into the development processes.
In this episode we invited Mike Kobush, Performance Engineer at NAIC, to hear how he found his way in quality engineering and later on got hooked on performance. He walks us through current performance challenges as his organization is moving to the cloud. Mike also discusses why he is embracing automation as it makes him more desirable as an employee. He shares his goals and motivations such as: Learning something new every day!
Tune in and get inspired by Mike. And remember: it always pays off to spend the extra time with somebody – even if it’s a Sunday afternoon in sunny San Diego!
Show Links:
Mike on Linkedin
https://www.linkedin.com/in/michael-kobush-31-performance/
What are good business level objectives (BLOs) besides conversion rates? Who is responsible for defining them? Who needs to report and who is held accountable?
We invited John Kelly, Sales Engineer at Dynatrace, to answer those and even more questions. John – aka Tech Shady - has been helping our customers over the past years to implement business level reporting for their critical applications. It was exciting to hear that there is much more than your classical availability or conversion rate business metrics. The one we think is really exciting is Engagement Rate. So – tune in and learn for yourself
Show links:
John Kelly on Linkedin
https://www.linkedin.com/in/john-kelly-b22b992/
John Kelly on Twitter
https://twitter.com/JohnKelly17
Java developers love using Spring. But running high performing and scaling Java apps in production takes a little bit more than just compiling your code.
In this episode we have Asir Vedamuthu Selvasingh who has been working with Java for 26 years. In the past 25 years Asir (@asirselvasingh) helped Microsoft provide services to their developer community that make it easier to deploy, run and operate Java based applications at scale – nowadays on the Azure Spring Cloud offering.
Listen in and learn more about observability when deploying apps on the Azure Cloud, which performance and scaling aspects you have to consider and get a look behind the scenes on how your packaged java app magically becomes available across the globe.
Show Links:
Asir on Linkedin
https://www.linkedin.com/in/asir-architect-javaonazure/
Asir on Twitter
https://twitter.com/asirselvasingh
Monitoring Spring Boot Apps with Dynatrace
https://docs.microsoft.com/en-us/azure/spring-cloud/how-to-dynatrace-one-agent-monitor
Observability on Azure Spring Cloud
https://www.dynatrace.com/news/blog/dynatrace-extends-observability-to-azure-spring-cloud/
If there is one thing you take away from this episode then the answer to “Why we should refrain from Reply All on company wide emails”. Jokes aside – as security and performance are not always funny!
In this special anniversary episode we have Mark Tomlinson, System Performance Specialist, talking about the considerations and trade-offs between performance and security. We learn about performance vulnerabilities and why it is important to factor in the additional overhead each layer of security adds to your application stack. It's always a pleasure having Mark on the show – whether it was in the past, present or will be in the future.
If you want to learn more from Mark on the topic of performance make sure to check out PerfBytes that has inspired us to launch PurePerformance.
Show Links:
Mark Tomlinson on Linkedin
https://www.linkedin.com/in/mtomlins/
PerfBytes
https://www.perfbytes.com/
Optimizing or debugging database calls has to become as easy as optimizing your application code based on logs, metrics or traces your observability platform provides to developers. It has to be doable by the development and DevOps teams who are becoming more end-2-end responsible which includes new database services that are running in some managed cloud service.
In this episode we hear from Nimesh Bhagat, Product Manager at Google, how modern database observability supports development and DevOps teams to better understand, optimize and operate their end-2-end service flow. A great project Nimesh has been working on is sqlcommenter which uses OpenTelemetry to continue distributed traces started in the application into the internals of the database engine.
If you want to learn more check out the sqlcommenter documentation or the Google Podcast on Cloud SQL Insights.
Show Links
Nimesh on Linkedin
https://www.linkedin.com/in/nimesh-bhagat-b062354/
SQLCommenter
https://cloud.google.com/blog/products/databases/sqlcommenter-merges-with-opentelemetry
SQLCommenter Documentation
https://google.github.io/sqlcommenter/
Google Podcast on Cloud SQL Insights
https://www.gcppodcast.com/post/episode-247-cloud-sql-insights-with-nimesh-bhagat/
If you need to learn how Prometheus, OpenTelemetry, Loki, FluentD, FluentBit .. help you with your observability requirements in the cloud native and non-cloud native space but you don’t have hours or days to dig into the details yourself then you have a new place to go to get educated within 20-30 minutes: Is it Observable is a new educational YouTube channel by Henrik Rexed, Cloud Native Advocate at Dynatrace.
In this episode we have Henrik explain the motivation of creating Is it Observable, how to best use the videos and the tutorials on GitHub to educate yourself and gave a glance on upcoming episodes that will also include guests from various tool and platform vendors.
Make sure to subscribe to his channel, check out the tutorials on GitHub and give him feedback on twitter (@hrexed) or LinkedIn
Show Links
Is It Observable YouTube Channel:
https://www.youtube.com/channel/UCRiin4u8YZGlVQRZp7qOOaw
Henrik Rexed on Linkedin
https://www.linkedin.com/in/hrexed/
Is It Observable Github Repo
https://github.com/isItObservable
Henrik Rexed on Twitter
https://twitter.com/Hrexed
Tuning the JVM GC to reduce garbage collection time will speed up application performance. If you agree with that statement then I encourage you to listen to this episode where I have Stefano Doni, CTO at Akamas, walk us through 4 Java Tuning Facts & Myths. He is going into details why even in 2021 with great improvements in the JVM it is still important to optimize the JVM specific to the environment, workload and application behavior. If you want some visuals try to catch his presentation from this years Performance Summit called “How AI optimization will debunk 4 long standing Java tuning myths”
To follow up with Stefano check out their resources such as their tutorials on explore.akamas.io, check out their blog posts or videos on their new website akamas.io or follow them on twitter @akamaslabs
Links from the Show:
Stefano on Linkedin
https://www.linkedin.com/in/stefanodoni/
YouTube: How AI optimization will debunk 4 long-standing Java tuning myths
https://www.youtube.com/watch?v=9VvaxATyYsA
Akamas Tutorials
https://explore.akamas.io/
Akamas website
https://www.akamas.io/
Akamas on Twitter
https://twitter.com/AkamasLabs
Understanding the secret behind the turbo button on his first 486 PC motivated our guest to study computer science. That decision started a journey making him constantly learn new technology ranging from coding languages, operational tasks as well as a focusing on improving developer experiences and boosting developer productivity
Listen in and hear from Michael Friedrich (@dnsmichi), The Ops in Dev Evangelist at GitLab, on why it is important to enable developers to design and develop code that makes it easier for DevOps and SREs to operate and automate. “The biggest challenge is code that breaks production but where there is no clear evidence for DevOps & SREs about the root cause”
Make sure to join Michael’s #EveryCanContribute and follow his advocacy such as DockerCon 2021 on From Infrastructure as Code to Cloud Native Deployments in 5 Minutes
Links from show:
Linkedin
https://www.linkedin.com/in/dnsmichi/
Twitter
https://twitter.com/dnsmichi
Everyone Can Contribute
https://everyonecancontribute.com/
From Infrastructure as Code to Cloud Native Deployments in 5 Min
https://docker.events.cube365.net/dockercon-live/2021/content/Videos/emEjNyA4WmBSv8BW2
“Because 9 out of 10 load testing projects fail due to ignorance and outdated thinking about load testing!”. That was the answer Leandro Melendez, aka Señor Performo, gave us when asking him why the world needs yet another book about load testing.
In too many projects Leandro has to remind and educate decision makers and practitioner’s about load testing best practices, how to ask the right questions and how to approach a project from start to finish. His book “The hitchhiking guide to load testing projects” is a fun and edu-taining read for people that are new to the trade as well as seasoned performance engineers.
For more content from Senor Performo check out his PerfBytes Espanol Podcast, his Spanish YouTube channel and all the performance engineering presentations he has been given over the years.
https://www.srperf.com/podcast/
https://www.srperf.com/el-youtube-channel/
https://www.srperf.com/presenter/
Pre-order the book here
https://www.amazon.com/dp/B09C4ZT1LB
Building products that people want to use and activating users to try out new capabilities has to be the ultimate goal of every product manager. User and usage data is the enabler to make the right decisions. But data doesn’t come for free – and making the right decisions is something that data alone doesn’t guarantee
Listen in and learn from Manav Chugh, product enthusiast, medium blogger and organizer of ProductTank Linz, what inspired him to choose the path from Zero to Data-Driven Product Manager. In our conversation we cover how to capture what data, the importance of data privacy, what we can learn from companies that do data-driven design well and why he loves organizations such as www.ecosia.org.
Links from show:
Manav on Linkedin
https://www.linkedin.com/in/manavchugh/
Manav's Medium Blog
https://manav77-chugh.medium.com/
ProductTank Linz
https://www.mindtheproduct.com/producttank/linz
Ecosia Search Engine
https://www.ecosia.org/
Like many frontend developers, Sergey Chernyshev was inspired in the late 2000 by Steve Souders to contribute to and grow the web performance community. Not only did he launch the Meed4SPEEDs as part of the New York Web Performance Meetup. He also worked for meetup.com helping them to improve web performance and user experience. Over the past years Sergey contributed to many projects such as WebPageTest.org, UX Capture Library, Jamstack and more.
Tune in to this episode, learn what has and what hasn’t changed in the quest for better user experience. Get a quick start on Core Web Vitals and current challenges. Most important: get inspired to contribute back to this community.
Llinks we discussed during the podcast are here:
Sergy on LinkedIn https://www.linkedin.com/in/sergeychernyshev/
Steve Souders Page: https://stevesouders.com/
NY Web Performance Meetup: https://www.meetup.com/Web-Performance-NY/
WebPageTest: https://www.webpagetest.org/
Github UX Capture: https://github.com/ux-capture/ux-capture
Jamstack: https://jamstack.org/
WebVitals: https://web.dev/vitals/
CrUX: https://developers.google.com/web/tools/chrome-user-experience-report
Edge Compute:
CloudFlare Workers: https://workers.cloudflare.com/
Fastly: https://www.fastly.com/products/edge-compute/serverless
PWA Stats: https://www.pwastats.com/
Next.JS approach to SSR: https://nextjs.org/docs/basic-features/pages
In his SLOConf talk Production load testing as a guardrail for SLOs and in his blog Production Load Testing, Hassy Veldstra, founder of artillery.io makes the case for load testing in production. It helped him in various organizations to establish SLOs (Service Level Objectives) and change the way engineers think about performance. He got inspired by Building Evolutionary Architectures which introduces the concept of performance as a fitness function.
Tune in into our conversation, hear our arguments pro and contra load testing in the various environments and learn why in the end we agreed on the fact that SLOs – while nothing really new – are a great chance to re-define performance engineering.
Linkedin
https://www.linkedin.com/in/hveldstra/
SLOconf: Production load testing as a guardrail for SLOs - by Hassy Veldstra
https://www.youtube.com/watch?v=Y20K1mJB6tk
Blog: Load testing. In production.
http://veldstra.org/production-load-testing/
Artillery Website
https://artillery.io/
Book: Building Evolutionary Architectures
https://www.thoughtworks.com/books/building-evolutionary-architectures
How do you convince an organization that just went through a 2 year DevOps transformation to continue the journey by applying SRE practices? What is SRE anyway? What are good SLOs? And how do you get development teams to take responsibility for their code in production?
Bart Enkelaar, Lead Site Reliability Engineer at bol.com, not only got their organization to apply SRE practices, define good SLOs and got dev teams to rotate on-call duties. He also followed the advice of Margaret, Chief Platform Officer, to bring his personal passion to the job. This led to inspiring and educating the community about SRE and SLO through music. To see what I mean check out Barts The Game of SLOs – a three part reliability musical from SLOConf or his funny tech conversations at Friendly Tech Chats.
Linkedin - https://www.linkedin.com/in/bart-enkelaar-02242710/
Margaret, Chief Platform Officer - https://www.youtube.com/watch?v=hy1gUEhbnBM
Game of SLOs: A 3 part reliability musical - https://www.youtube.com/watch?v=Y53Pho93i-k
Friendly Tech Chats - https://www.youtube.com/channel/UChHHWkO537q6Yp2dXtJpOzQ/featured
While some think about the late Austrian musician, Dan POP and the CNCF community thinks about modern security when it comes to Falco.
Listen in and hear directly from Dan (@danpopnyc) who, besides doing many things in the CNCF community, also hosts POPCAST where he started connecting technology leaders during the last year. In the podcast you learn a lot about security, the power of eBPF and how Falco aims to contribute to runtime security like k8s contributed to distributed computing.
Here the additional links we brought up during the conversation:
Dan on Linkedin
https://www.linkedin.com/in/danpapandrea/
Dan on Twitter
https://twitter.com/danpopnyc
Popcast Podcast
https://github.com/danpopSD/popcast
Cyber Defenders Career Guide by Alyssa Miller
https://www.manning.com/books/cyber-defenders-career-guide
Cloud Native TV
https://www.twitch.tv/cloudnativefdn
CNCF Tag Security
https://github.com/cncf/tag-security
Falco Tools, Frameworks & Articles
https://github.com/developer-guy/awesome-falco
Falco Blog
https://www.cncf.io/blog/2020/12/14/join-pop-falco-org/
Falco Der Kommissar
https://www.youtube.com/watch?v=8-bgiiTxhzM
Wonder what you learn when building k8s from scratch for a large enterprise? Wonder what you learn when automating delivery by connecting your different DevOps tools together?
Nana Janashia runs one of the most successful technical YouTube channels called TechWorld with Nana where she covers topics ranging from containers, docker, k8s, cloud native and DevOps. She basically takes her lessons learned and explains technologies and concept in a very easy way especially for folks that want to get started with.
In this episode we focus a lot on DevOps, what the right trades of a DevOps engineer are and how to get started. Thanks Nana for your time and all the additional resources we talked about during the episode that are listed below:
DevOps Bootcamp: https://www.techworld-with-nana.com/devops-bootcamp
DevOps Roadmaps for Humans with Bret Fisher: https://www.youtube.com/watch?v=UXf2c76KAyA
DevOps Roadmap: https://roadmap.sh/devops
Linkedin Profile: https://www.linkedin.com/in/nana-janashia/
TechWorld with Nana YouTube channel: https://www.youtube.com/channel/UCdngmbVKX1Tgre699-XLlUA
Have you ever thought about reorganizing data allocation based on production telemetry data? Have you ever thought about shifting compiler budgets to parts of your code that is heavily executed based on profiling information captured from your real end users? Whether the answer is yes or no you will be fascinated by Taras Tsugrii, Software Engineer at Facebook, who is sharing his experience on optimizing everything from compilers, to databases, distributed systems or delivery pipelines.
If you want more after listening to this episode check out his recent talk at Neotys PAC titled “Old pattern powering modern tech”, subscribe to his substack newsletter, his hashnode blog, or the conference recordings of Performance Summit and Scaling Continuous Delivery.
https://www.linkedin.com/in/taras-tsugrii-8117a313/
https://www.youtube.com/watch?v=itOCQvk_LAs
https://softwarebits.substack.com/
https://softwarebits.hashnode.dev/
https://www.youtube.com/channel/UCt50fEvgrEuN9fvya8ujVzA
https://www.youtube.com/channel/UCWf9HxiBudKLzCFtgAAz8XQ
Googles Census, OpenCencus, OpenTelemetry and AWS Distro for OpenTelemetry. Our guest Jaana Dogan, Principal Engineer at AWS, has been working in observability over many years and definitely had a positive impact on the where OpenTelemetry is today. In this episode Jaana (@rakyll) explains which problems the industry, and especially cloud vendors, try to solve with their investment in open source standards such as OpenTelemetry. She gives an update where OpenTelemetry is, the next upcoming milestones such as metrics and logs and what a bright future with OpenTelemetry being widely adopted could bring.
https://twitter.com/rakyll
If you are interested in learning more – here are the links we discussed during the podcast:
https://github.com/open-telemetry
https://github.com/open-telemetry/opentelemetry-specification
https://github.com/open-telemetry/opentelemetry-proto
https://github.com/open-telemetry/opentelemetry-collector
https://github.com/open-telemetry/community
https://o11yfest.org/
Performance Engineering is not about running a performance test twice a year. That is just a poor attempt trying to validate your non functional requirements.
Roman Ferstl, Managing Directory at Triscon, has discovered his love for performance engineering while optimizing code for software used in a space program. He then founded Triscon who is now helping to establish and scale performance engineering at large enterprises. In this episode we get his insights on how he approaches a new project, which bottlenecks to address first and how to motivate more people within an organization to invest in performance engineering.
If you want to learn more don’t miss to check out Roman’s presentation from Perform 2021 titled “Turbocharging your Performance Engineering teams to scale efficiently”
https://www.linkedin.com/in/roman-ferstl/
https://www.triscon-it.com/en/
https://perform.dynatrace.com/2021-americas/breakouts-single-day-3-turbocharging-your-performance-engineering-teams
To k8s or not – that should be the first question to answer before considering k8s. Granted – in many cases k8s is going to be the right choice but don’t just default to k8s because its hip or cool.
In this episode we have Christian Heckelmann (@wurstsalat), DevOps Engineer at ERT, talking about his journey with k8s which started with installing k8s 1.9 on bare metal. He gives a lot of great advice based on his presentation “How not to start with k8s” such as Understand Networking, Don’t use :latest, Set Resource Limits, Train The People, Provide Templates and more.
To get started with Kubernetes we encourage you to look at the YouTube Tutorials posted on TechWorld with Nana.
https://twitter.com/wurstsalat
https://docs.google.com/presentation/d/1EL9OYe-1eOPXh6U8SMHnQxs8pcmr01d-uwoWoFnzUaY/edit#slide=id.g5420f4ebeb_0_5
https://www.youtube.com/channel/UCdngmbVKX1Tgre699-XLlUA
You heard about Continuous Integration, Continuous Delivery and Continuous Deployment. Liquid Software aims to provide the next step towards Trusted Continuous Updates in the DevOps World.
In this episode Baruch Sadogursky, DevOps Advocate from JFrog, explains how as engineers we need to add “Updateability” to our non-functional requirements and how product managers and marketing have to forget about traditional releases but think about incremental delivery of value. Baruch (@jbaruch) also promised to send everyone a hard copy of his book “Liquid Software” if you send him a direct message – so – make sure you do that and also check out the details on our discussion of uniquely identifying artifacts through Build-Info.
https://www.linkedin.com/in/jbaruch/
https://twitter.com/jbaruch
https://drive.google.com/file/d/1PUb67FxM-eTtdyLNGPc-fGTcCJii-keE/view
https://github.com/jfrog/build-info
Software security is about securing websites against malicious attacks or using firewalls to prevent hackers entering your enterprise network. While this is part of software security there is much more that needs to be done – especially as more organizations are developing critical software it is important to protect the whole software delivery lifecycle from any malicious attacks along the supply chain.
In this episode we have Michael Plank, Technical Product Manager at Dynatrace, talk about his latest blog post titled How Dynatrace protects its software development and delivery life cycle against supply chain attacks. We learn about attack vectors from development workstation until production deployment. He covers the strategies ranging from static to dynamic code analysis, vulnerability detection or code signatures. Tune in and learn that building secure software is more than ensuring your users have hard to crack passwords!
https://www.dynatrace.com/news/blog/how-dynatrace-protects-its-software-development-and-delivery-life-cycle-against-supply-chain-attacks/
If you are not a gamer you may have never heard about Cyberpunk 2077. If you are – you may know about the challenges during their latest release.
Dave Farley (@davefarley77), Co-Author of best seller Continuous Delivery, has been an engineering large and complex systems for decades. His work helped elevate our industry around Continuous Delivery and DevOps. In this episode he shares his learnings from failed projects like Cyberpunk as well as his own latest experiences around that picking the latest technology might be fashionable but is not always the smartest choice.
To learn more about Dave check out Continuous Delivery website that also links to his YouTube Channel hosting some of the episodes he was referencing in the podcast.
https://twitter.com/davefarley77
https://www.amazon.com/Continuous-Delivery-Deployment-Automation-Addison-Wesley/dp/0321601912#ace-g9859629705
https://www.continuous-delivery.co.uk/
https://www.youtube.com/channel/UCCfqyGl3nq_V0bo64CjZh8g
Nobody has foreseen the global pandemic that put a lot of chaos in all our lives recently. Let’s just hope we learn from 2020 to better prepare on what might be next.
The same preparation and learning also goes for Chaos in our distributed systems that power our digital lives. And to learn from those stories and better prepare for common resiliency issues we brought back Ana Medina (@ana_m_medina), Chaos Engineer at Gremlin. As a follow up to our previous podcast with Ana, she is now sharing several stories from her chaos engineering engagements across different industries such as finance, eCommerce or travel. Definitely worth listening in as Chaos Engineering was also put into the Top 5 Technologies to look into 2021 by CNCF.
https://twitter.com/Ana_M_Medina
https://www.spreaker.com/user/pureperformance/why-you-should-look-into-chaos-engineeri
https://twitter.com/CloudNativeFdn/status/1329863326428499971
When moving to microservice architectures its time to re-think continuous delivery. Just as many software services rely on a core data analytics engine to make better automated decisions we need to apply the same for continuous delivery. We can assess the risk of every microservice deployment based on data from production and the desired change of configuration. We can assess the potential blast radius and mitigate it through modern delivery options such as blue/green, canaries or feature flags.
Tracy Ragan, Creator & CEO of DeployHub, CDF board member and DevOps Institute Ambassador shares her thoughts on why we need to move to smarter data-driven delivery pipelines. Tracy (@TracyRagan) gives us insights into why not every microservice is created equal and what approaches we can take to better control updates that contain multiple microservice updates.
Also make sure to check out their latest project Ortelius and take Tracy up on a virtual coffee chat as discussed in our podcast!
https://www.linkedin.com/in/tracy-ragan-oms/
https://twitter.com/TracyRagan
https://github.com/ortelius
https://go.oncehub.com/15-30MinuteVirtualCoffeeWithTracy
K8s enables organizations to more easily deploy their containerized solutions as it takes away a lot of the operational tasks which are built-into k8s. This in theory means that you can run your software anywhere and provide it as SaaS offering or deploy it behind corporate firewalls for those customers that demand an on-premise installation.
In this episode we have Marc Campbell, Founder and CTO of Replicated, where they help the k8s community to deliver and manage apps on k8s anywhere. For anyone looking into running their apps on k8s you will learn the challenges of Day 1 (delivery, install) and Day 2 (operation, monitoring, troubleshooting) operations. Marc shares common performance and scalability challenges and how to prepare for them during development.
In this episode we have Marc Campbell, Founder and CTO of Replicated, where they help the k8s community to deliver and manage apps on k8s anywhere. For anyone looking into running their apps on k8s you will learn the challenges of Day 1 (delivery, install) and Day 2 (operation, monitoring, troubleshooting) operations. Marc shares common performance and scalability challenges and how to prepare for them during development.
https://www.linkedin.com/in/campbe79/
https://www.replicated.com/
https://www.heavybit.com/library/podcasts/the-kubelist-podcast/ep-7-keptn-with-andreas-grabner-of-dynatrace/
https://troubleshoot.sh/
https://kots.io/
Stefan Frandl, Development Director, has a single digit employee number at Dynatrace and therefore seen a lot of agile transformation over the past 15 years – growing from a startup in Linz, Austria to now 800+ engineers across globally distributed labs. A visit to several “unicorns” such as Google, Facebook and Slack triggered the latest agile transformation.
In this episode Stefan walks us through the implementation of the changes we discussed with Andrea Holl in her episode on “Scaling Agile at Dynatrace”. He shares the challenges around growing responsibilities of team leads, work left half-finished, overhead on hand-over and cross team collaboration. He then introduces us to the current structure and processes at Dynatrace such as Team Captains, Product Owners and Agile Advocates as well as Dev Directors and Lead Product Engineers. While Dynatrace has seen many benefits already, the journey is still ongoing as Dynatrace is continuously rethinking and improving the way we work and provide value to our customers!
https://www.linkedin.com/in/stefan-frandl-aa86723/
https://www.spreaker.com/user/pureperformance/scaling-agile-at-dynatrace-with-andrea-h
SAFE, LESS or the Spotify Model? Which scaled agile method to apply for your transformation? Or are you unique enough like 44% of organizations based on a European research that are defining their own scaled agile approach to transform successfully?
In this episode we sit down with Andrea Holl, Agile Coach at Dynatrace, and let her walk us through the different scaled agile frameworks. She discusses the pros and cons and why many organizations – including Dynatrace – are coming up with their own approaches. For Dynatrace it was about taking the best from the proven frameworks but adapting them to allow us continue or core cultural values such as full autonomy to teams and flexibility of tools and processes.
If you are on the brink of a transformation make sure to listen to Andrea and how she and her teams have approached that transformational project!
https://www.linkedin.com/in/andrea-elisabeth-holl-b2255a112/
https://www.scaledagileframework.com/
https://less.works/
https://blog.crisp.se/wp-content/uploads/2012/11/SpotifyScaling.pdf
Daylight savings can bring chaos to systems such as rogue processes consuming CPU or memory and therefore impact your critical systems. The question is: how do you systems react to this chaos? How can you test for this? And how can you make your systems more resilient against this chaos?
In this episode we talk with Ana Margarita Medina, Chaos Engineer at Gremlin. In her previous job, Ana (@Ana_M_Medina) was a Site Reliability Engineer at Uber where she helped coping with the “chaos” on New Years Eve or Halloween. Ana gives us great insights into the discipline of Chaos Engineering, that its really about running controlled experiment and that everyone can get started that has an interest in contributing to more resilient systems.
Here the additional links we promised during the recording: Drift into failure, Chaos Engineering Community, Chaos Engineering and System Resilience in Practice.
https://www.linkedin.com/in/anammedina/
https://twitter.com/Ana_M_Medina
https://eng.uber.com/nye/
https://www.amazon.com/Drift-into-Failure-Sidney-Dekker/dp/1409422216
https://www.gremlin.com/community/
https://www.amazon.com/Chaos-Engineering-System-Resiliency-Practice/dp/1492043869
We are sitting down with Sebastian Scheele (@sscheele), CEO and co-founder of Kubermatic, to discuss the challenges organizations have as they are moving their workloads to k8s and realize that managing, scaling and operating k8s is not getting easier the more k8s clusters you allow your application teams to spin up or down. We learn more about the Kubermatic Kubernetes Platform, the Open Source Project, which centrally manages the global automation of thousands of Kubernetes clusters across multi-cloud, on-prem and edge with unparalleled density and resilience.
Thanks Sebastian for answering all the questions we threw at you – questions we have received from many organizations that are moving to k8s but get surprised about the complexity as it comes to properly operating and managing k8s.
Sebastian Scheele Twitter
https://twitter.com/sscheele
Kubermatic Kubernetes Platform
https://github.com/kubermatic/kubermatic
Keptn is now a CNCF sandbox project bringing a new event-driven approach to continuous delivery and operations. While many are just hearing about Keptn the first time, it is interesting to learn more about how it started, which challenges the team ran into, what they learned about K8s, and running an open-source project. We therefore invited Johannes Braeuer (@braeuer_j) and Andreas Grimmer (@grimmer_andreas) – both Keptn project maintainers and contributors – who have been working on the Keptn project since its inception.
Especially for groups that want to start open-source projects or are on the brink of deciding pro or con Kubernetes should listen until the end as Johannes and Andreas tell us what they would do differently now if they would start today based on the learnings from the past 18 months.
If you want to join the Keptn community, make sure to star our GitHub project, join the Slack channel, and join our regular community meetings!
Keptn
https://keptn.sh/
Johannes Bräuer on Twitter
https://twitter.com/braeuer_j
Andreas Grimmer on Twitter
https://twitter.com/grimmer_andreas
Keptn Github
https://github.com/keptn/keptn
Keptn Slack
https://keptn.slack.com/
Keptn Community
https://github.com/keptn/community
Getting visibility into .NET code whether it runs on a developer machine, on a windows server on-premise or as a serverless function in the cloud is the day2day job of Georg Schausberger (@BombadilThomas) and Bernhard Ruebl, part of the Dynatrace .NET Agent Team.
In this podcast we hear firsthand about the challenges in bringing observability, monitoring and distributed tracing to the .NET ecosystem. They give us insights about their continued effort to reduce startup and runtime overhead, the innovation that comes out of Microsoft as they are moving towards open standards and the noble automated approach to always validated things don’t break monitored code with the constant update of libraries and frameworks.
We also got both to talk about their developer experience when working with commercial tools such as Dynatrace and its PurePath technology as well as open source tools when analyzing and debugging their own code or helping users figure out what’s wrong with their code.
In the talk both mentioned other tools which we wanted to provide the links for:
Benchmark.NET
https://benchmarkdotnet.org/articles/overview.html
Ben.Demystifier.
https://www.nuget.org/packages/Ben.Demystifier/
IIS Module tracing
https://forums.ivanti.com/s/article/How-To-Enable-IIS-Failed-Request-Tracing
Georg Schausberger
https://twitter.com/BombadilThomas
https://www.linkedin.com/in/georg-schausberger-6898b6141/
Bernhard Rübl
https://www.linkedin.com/in/bernhard-r%C3%BCbl-084881104/
Successful Cloud Migrations, large scale Kubernetes & OpenShift deployments, making billions of data points actionable and enterprise-wide Citrix & SAP monitoring. These are some of the projects Kayan Hales, Technical Manager at Dynatrace, and her colleagues at Dynatrace ONE help enterprise customers around the world to implement every day.
We sat down with Kayan as we wanted to learn what really matters to many large organizations as they embark on automating monitoring into their hybrid multi-cloud environments. While we constantly talk about cloud native and microservices it was interesting to hear what the global team of Dynatrace experts is doing on a day-2-day basis. Kayan gives us insights how important it is to think about meta data, tagging strategies and automation before large scale rollouts and that one of the first question you need to ask is: who needs what type of data at which time through which channels.
https://www.linkedin.com/in/kayanhales/
https://www.dynatrace.com/services-support/dynatrace-one/
Why do some organizations still see performance testing as a waste of time? Why are we not demanding the same level of performance criteria for SaaS-based solutions as we do for in-house hosted services? Why are many organizations just validating performance to be “within specification” vs “holistically optimized”?
In this episode we have invited James Pulley (@perfpulley), Performance Veteran and PerfBytes News of the Damned host, to discuss who organizations can level up from performance testing to true performance engineering. He also shares his approaches to analyzing performance issues and gives everyone advice on what to do to start a performance practice in your organization.
https://www.linkedin.com/in/jameslpulley3/
https://www.perfbytes.com/p/news-of-damned.html
Imagine a future where we deploy every code change directly into production because feature flags eliminated the need for staging. Feature flags allow us to deploy any code change, but only launch the feature to a specific set of users that we want to expose to new capabilities. Monitoring the usage and the impact enables continuous experimentation: optimizing what is not perfect yet and throw away features (technical debt) that nobody really cares about. So – what are feature flags?
We got to chat with Heidi Waterhouse (@wiredferret), Developer Advocate at LaunchDarkly (https://launchdarkly.com/), who gives as a great introduction on Feature Flags, how organizations actually define a feature and why it is paramount to differentiate between Deploy and Launch. We learn how to test feature flags, what options we have to enable features for a certain group of users and how important it is to always include monitoring. IF you want to learn more about feature flags check out http://featureflags.io/. If you want to learn more about Heidi’s passion check out https://heidiwaterhouse.com/.
Adrian Hornsby (@adhorn) has dedicated his last years helping enterprises around the world to build resilient systems. He wrote a great blog series titled “Patterns for Resilient Architectures” and has given numerous talks about this such as Resiliency and Availability Design Patterns for the Cloud at DevOne in Linz earlier this year.
Listen in and learn more about why resiliency starts with humans, why we need to version everything we do, why default timeouts have to be flagged, how to deal with retries and backoffs and why every distributed architect has to start designing systems that provide different service levels depending on the overall system health state.
Links:
Adrian on Twitter: https://twitter.com/adhorn
Medium Blog Post: https://medium.com/@adhorn/patterns-for-resilient-architecture-part-1-d3b60cd8d2b6
Adrian's DevOne talk: https://www.youtube.com/watch?v=mLg13UmEXlw
DevOne Intro video: https://www.youtube.com/watch?v=MXXTyTc3SPU
Whether you are still researching on whether you need a Service Mesh or simple use a load balancer or if you are already deploying multi hybrid-cloud architectures and Service Meshes help you secure the location aware routed traffic. In both cases: listen to this episode!
We invited Sebastian Weigand (@ThatDevopsGuy) back to our podcast who wrote papers such as Building a Planet-Scale Architecture the Easy Way. In our episode Sebastian walks us through why Service Meshes have gained so much in popularity, what the main use cases are, how you should decide on whether or not use Service Meshes and which challenges you might run into as you expand into using more features.
https://twitter.com/thatdevopsguy
https://files.devnetwork.cloud/DeveloperWeekNewYork/presentations/2019/scalability/Sebastian_Weigand.pdf
Steve McGhee (@stevemcghee) is an expert in post mortems and SRE. He has learned the craft at Google, applied it at MindBody and is now sharing his experiences while back at Google to the larger SRE community. Listen to this episode and learn more about how post mortem analysis can be the starting point of your SRE transformation. How it can help reliability engineering to build and engineer systems that fail gracefully instead of causing full crashes or outages.
Steve also went into monitor what matters and only defining alerts on leading indicators with an expiration date – a fascinating concept to avoid a flood of custom alerting in production!
If you want to learn more from Steve or SRE check out these additional resources he mentioned in the podcast: The SRE I aspire to be (SRECon19) and his 2 blog part series on blameless.com.
https://twitter.com/stevemcghee
https://www.youtube.com/watch?v=K7kD_JfRUY0
https://www.blameless.com/blog/improve-postmortem-with-sre-steve-mcghee
Keep hearing the terms SLIs, SLOs, SLAs, Error Budgets and finally want to understand what they are, who should be responsible for and how they fit into SRE (Site Reliability Engineering)?
Then listen to our conversation with Sebastian Weigand who has been helping organizations modernizing not only their application stacks but also helping them embrace DevOps & SRE. Learn about who is responsible to define SLIs, what the difference between SLOs and SLAs are and what the difference between DevOps & SRE is in his opinion!
Sebastian, who calls himself “That Devops Guy” (@ThatDevopsGuy), also suggests to check out the latest free report on SLO Adoption and Usage of SRE as well as SRE Books from Google to get started with that practice.
https://www.linkedin.com/in/thatdevopsguy/
https://twitter.com/ThatDevopsGuy
https://landing.google.com/sre/resources/practicesandprocesses/slo-adoption-and-usage-in-sre/
https://landing.google.com/sre/books/
Cassidy (@cassidoo) has been building but also educating developers on how to build apps on React, JavaScript, JAMStack and many other technologies over the past years. We got her on our podcast where she gave us insights into React Hooks, how WPO (Web Performance Optimization) plays out in the React world, why it is important to think about state from the start and that its important to always have your end user in mind before even writing your first line of JavaScript.
In the podcast she references additional resources which here are the links for: The performance benefits of Variable Fonts, Mandy Michael (@Mandy_Kerr), Isabela Moreira (@isabelacmor) and A/B Testing with React (YouTube).
https://twitter.com/cassidoo
https://reactjs.org/
https://jamstack.org/
https://uxdesign.cc/the-performance-benefits-of-variable-fonts-79af8c4ff56c
https://twitter.com/Mandy_Kerr
https://twitter.com/isabelacmor
https://www.youtube.com/watch?v=xpfR0rRfcNk
How do you prepare for a 2Mio concurrent user load that lasts for 7 seconds? What does the load infrastructure look like? How do you optimize your scripts? How do you deal with DNS or CDNs?
In this episode we hear from Joerek van Gaalen who has done these types of tests. He shares his experiences and approaches to running these “special event extreme load tests”. If you want to learn more make sure to check out his presentation and read his blog post from Neotys PAC 2020.
https://www.linkedin.com/in/joerekvangaalen/
https://www.neotys.com/performance-advisory-council/joerek_van_gaalen
Have you ever burned 30k because you forgot to turn off your test VMs over the weekend? Have you ever accidentally deleted “the production table” because you thought you were connected to your dev database? We often only hear the good stories and not those that teach us about what we should not do in order to avoid disaster!
Join this episode where Justin Donohoo, Founder and CTO of Observian, tells us horror stories from his professional life that taught him great lessons on what not to do when moving to the cloud, re-architecture because of exponential growth or let the intern do things he/she shouldn’t do.
https://www.linkedin.com/in/jdonohoo/
Goranka Bjedov has seen the different sides of Open Source while she was working for organizations such as Google, Facebook or AT&T Labs. Before she takes the stage at www.devone.at later this year she gives us her take on Scott McNealy’s quote “Open Source is free like a puppy is free”. Tune in and hear her thoughts on how to pick the right tools, languages or frameworks, how to grow a an open source project and what things you should definitely avoid.
https://www.linkedin.com/in/goranka-bjedov-5969a6/
https://devone.at/
Starting your new job as Infrastructure Engineer in a large bank with your to-be boss and his key architects just leaving feels like Chaos! Maybe that’s why Tammy Butow has made a career in Chaos and Site Reliability Engineering. In this episode, Tammy shares her experiences of bring reliability into highly complex systems at NAB, Digital Ocean, DropBox or now Gremlin through chaos engineering. You learn about the importance to know and baseline your metrics, to define your SLIs and SLOs and to continuously run your fire drills to ensure your system is as reliable as it has to be.
If you want to learn more check out Tammy’s presentations on speakerdeck and make sure to join the chaosengineering slack channel.
https://www.linkedin.com/in/tammybutow/
https://speakerdeck.com/tammybutow
https://slofile.com/slack/chaosengineering
3 years ago, Adam Auerbach explained how he helped Capital One to automate performance into the DevOps Delivery Pipeline. In 2020, where IT Security is a hot trending topic, Adam works for EPAM and is back advocating for the same shift-left he as advocated for when it comes to functional or performance testing. But now – its about baking Security into your practices & culture. And he has a cool word for it: DevTestSecOps!
Listen in and learn which types of security checks can be fully automated in the different stages of the delivery pipeline. Also learn how to prioritize your vulnerabilities as you most likely end up with a lot of noise in the beginning. Adam also highlightes the following open source tools that will help in that transformation: getcarrier.io and reportportal.io.
https://www.linkedin.com/in/adamauerbach/
https://getcarrier.io/#about
https://reportportal.io/
DesignOps, just as DevOps or NoOps, is targeted towards increasing the efficiency and collaboration between designers and engineers in order to deliver better, intuitive and consistent user experiences. It requires changes in processes, people and tooling and is heavily driven by enabling engineers to become more autonomous when developing and delivering new value for their organization.
Join this podcast and learn from Ursula Wieshofer (@Ursula_W), UX Design Team Lead, as well as from Fabian Friedl (@fabian_friedl), DesignOps Team Lead, on how we live and breath DesignOps at Dynatrace. You will learn about the recently released OpenSource Design System Barista, how it enables our engineers to deliver consistent user experiences across all sorts of software projects and how we manage feedback through the Barista GitHub project for future innovation.
Also make sure to check out the recent blog posts UX Guilds as well as how we deal with the constantly changing requirements of our design teams.
https://twitter.com/Ursula_W
https://twitter.com/fabian_friedl
https://barista.dynatrace.com/
https://github.com/dynatrace-oss/barista
https://www.dynatrace.com/news/blog/a-running-start-towards-enablement-how-a-ux-guild-will-broaden-our-horizons/
https://medium.com/@holgua/owning-the-unknown-d796e2705ac
We're back to our regular scheduled show!
Kelsey Hightower (@kelseyhightower) has worn many hats as it says on his bio but we also learned from him that he probably doesn’t have that many hats at home as he has been living a minimalistic life over the past couple of years. A philosophy as we learn in this podcast that also goes well when it comes to building your next platform on Kubernetes.
In this podcast we learn about the do’s and don’ts, how you should plan and test for k8s upgrades, which tradeoffs you have to take as it comes to performance, how to think about developer productivity on k8s and why it is important to read up on security as it relates to the software we build, deploy and run on our k8s clusters.
Thank you Kelsey for supporting our community with your time and expertise. Hope to have you back in the future!
https://twitter.com/kelseyhightower
We wrap up the last day of the Dynatrace Perform 2020 conference with an announcement of winners and learnings for the week
Vamos terminando el evento teniendo una platica con nuestro amigo Cesar Quintana quien tambien dio una vuelta por nuestro stand a platicar un poco de su experiencia en esta gran cxonferencia.
Andi Grabner, our man-on-the-street, gets the scoop on:
-NoOps: Reaching zero-incident prod through auto-remediation-as-code with Juergen Etzlstorfer
-Beyond thresholds – Find anomalies and reduce false positives with Thomas Natschlaeger
-Hybrid observability, from enterprise cloud to mainframe and everything in between with Alex Huetter
Andi Grabner, our man-on-the-street, gets the scoop on:
-Observability and beyond - with Thomas Rothschädl
-Automated, AI-powered answers for Kubernetes with Matt Reider
-RUM Roadmap with Alexander Sommer
Travis Depuy of xMatters speaks to Leandro & Brian about how to leverage tools like xMatters to enable proactive alerting and remediation throughout your code life-cycle
Andi Grabner, our man-on-the-street, gets the scoop on:
-Going web-scale with cross-environment features, globally distributed high availability and more - with Guido Deinhammer
-The role of OpenTelemetry in Dynatrace with Daniel Khan and Sonja Chevre
-Build resiliency into your continuous delivery pipeline with Michael Villiger
Tenemos tambien a Gabriel Prioli platicandonos un poco de la conferencua y su experiencia ayudando con la herramienta.
Nos encontramos con el amigo Sergio que nos platica de las novedades que tinene Dynatrace y ia eperiencia que se vive en la confrencia.
Andi Grabner, our man-on-the-street, gets the scoop on:
-Leverage AIOps with Dynatrace with Wolfgang Beer
-Load & performance engineering as a self-service with Rob Jahn
-The right way to deploy canary, blue/green and feature flags with Safia Habib
Tenemos la oportunidad de platicar con Andres Suarez venido desde Colombia participando en la conferencia, quien nos cuenta de su experiencia en este evento asi como de sus aventuras pasadas.
Andi Grabner, our man-on-the-street, gets the scoop on:
-How to improve every user’s mobile experience - with Dominik Punz
-Advanced observability in cloud native microservices and service meshes with Alois Mayr and Sonja Chevre
-Monitoring-as-a-self-service with Kristof Renders
We catch up with Mark on his latest adventures at Barbri, what he has planned next and how he gets business answers from Dynatrace using Digital Business Analytics
https://www.dynatrace.com/perform-vegas/
Here we are once again at the Cosmopolitan Hotel and Casino in Las Vegas, NV where…ONCE AGAIN…we are starting our 3 day marathon from the Dynatrace PERFORM 2020 conference. This is one of the biggest high-tech conference focused on performance disicplines from testing, engineering, architecture, monitoring and system scalability using Dynatrace’s innovative solutions. We’ll be chatting with Dynatracers, discussing news topics, giving away shoes and mini drones and sharing conversations with so many of the partners and attendees at this wonderous conference.
Do you have a clear definition of what Reliability means for your organization? Abigail Wilson, Reliability Architect at CFA Institute, sees this as a key requirement before you start transforming your organization to embrace site reliability, DevOps or Cloud Native.
In the podcast we hear how Abigail went on her journey where she has proven that you don’t need a background in IT in order to become an advocate and change agent for reliability engineering. In her role has bridged the gap between business and IT, she has helped bring stable environments to developers and testers and with those and many other steps has increased overall productivity, quality and stability of their business critical applications.
https://www.linkedin.com/in/abigailswilson/
https://theabigailwilson.com/
What does the Dynatrace ACE (Autonomous Cloud Enable) Team work on these days? How do cloud platform owners implement Monitoring as a Self-Service? How to elevate from traditional performance engineering to Performance as a Self-Service? How to bring the Unbreakable Delivery Pipeline to life? What problems can be auto-remediated and how? And what’s the role of Keptn when it comes to boosting the path to Autonomous Cloud?
In this episode we have Andi, track leader of Perform 2020’s Release Better Software Faster, talk us through what he is expecting to learn and take away from each breakout session. For more details check out his blog post or sign yourself up for a ticket to Perform 2020.
https://www.dynatrace.com/news/blog/getting-ready-a-taste-of-whats-to-come-at-perform-2020s-release-better-software-faster-track/
https://www.dynatrace.com/perform-vegas/
If you lift & shift to the cloud or move things back from the cloud to on-premise you most likely didn’t understand cloud and how it can help you transform your business and organization. A bold statement but very much true so as we learn in our conversation with Mike Kavis, Chief Cloud Architect at Deloitte.
Mike (@madgreek65) has been in technology for 35+ years and was an early adopter of cloud technology as early as 2007 when AWS only had about 6 APIs. He has launched several startups and is now helping organizations to rethink what cloud means for them!
Listen in to this podcast and learn why organizations fail if they don’t understand that cloud is about agility, it’s a platform for innovators and that traditional IT teams have to transform into internal cloud providers that provide services to their development teams for them to deliver better business value faster.
Mike also runs his own podcast – you may want to listen in to the episode he recorded with Andi on Speeding up Digital Transformation with Cloud Native.
https://www.linkedin.com/in/mikekavis/
https://twitter.com/madgreek65
https://www2.deloitte.com/us/en/pages/consulting/articles/digital-transformation-requires-a-cloud-native-mindset-cloud-computing-cloud-value-devops-noops-aiops-software-development-business-value.html
Wait! What? This is our 100th Episode of PurePerformance? For this special anniversary we invited Mark Tomlinson, Performacologist & “The Performance Sherpa”, who also inspired us through his PerfBytes Podcast to run our own PurePerformance Podcast.
While we start with talking about performance in podcasting we move over to learning more about how Mark is establishing a Continuous Performance process at his current employer. We learn about new ways to do performance engineering in a continuous way, how to integrate it with your monitoring and why it is not always important to run the big load tests but rather focus on short feedback cycles.
We want to give Mark credit for what he has done for the performance community and use this to say THANK YOU!! Hope to have you back for many more episodes to come and definitely for episode 200!
https://www.linkedin.com/in/perfsherpa/
https://www.perfbytes.com/
In 2013 the Phoenix Project by Gene Kim, Kevin Bahr and George Spafford sparked the next phase of DevOps transformations. 6 years later Gene Kim (@RealGeneKim) is back with The Unicorn Project, A Novel about Developers, Digital Disruption, and Thriving in the Age of Data.
Developer Productivity is a key focus point of the story in the book and is what Gene has learned from different companies in the last years about. Good engineering companies put their best resources in developer productivity as it benefits every developer and allows them to use their best energy to provide business value instead of solving puzzles. Gene lets us in on his day at Etsy as well as the story from Nokia and the reason they moved away from Symbian – both stories that touch on developer productivity!
If you want to learn more and read about the Five Ideals then download the excerpts from The Unicorn Project.
https://itrevolution.com/the-unicorn-project/
https://twitter.com/RealGeneKim
ChatOps is not new! But many organizations have not understood nor leverage its full potential. The use cases spread from “What’s on todays cafeteria menu?” to “Deploy my latest Git commit as canary and scale based my SLOs!”.
Listen to this podcast and learn from Nestor Zapata and Zohaib Hassan – both working at Citrix – on how they have started their ChatOps journey, how the built trust in the technology and how it helped them transform their organization towards more autonomy thanks to the self-service model enabled through the chat bots they developed. We discuss many of their self-service use cases such as Performance as a Self-Service or even Self-Healing which they implemented through Chat Bots integrated with Slack, Dynatrace, ServiceNow, Jira and other tools.
If you want to see their chat ops in action watch our Performance Clinic on Automate Deployment and Site Reliability with Bots, ChatOps and Dynatrace
https://www.youtube.com/watch?v=xE_LMQ9u7l4
16 years and still growing! Not every open source project has the track history of Spring (www.spring.io), a framework for building modern applications for the java runtime.
Juergen Hoeller (@springjuergen), creator of the Spring framework, gives us insights into how he and his team have grown Spring to where it is now. We learn how they have built a developer community, how they deal with feedback, why its important to interact with your users on a regular basis and where the road is heading. Juergen also shares insights on topics such as scalability, performance, concurrency of the framework as well as how the rise of new java runtimes and distributions still keeps him excited about the future of Spring.
If you want to learn more visit the Spring Framework’s GitHub project and make sure to read up on the latest blogs on https://spring.io/blog/
https://spring.io/
https://twitter.com/springjuergen
https://github.com/spring-projects/spring-framework/
Are you analyzing the dependency between change frequency, technical complexity, and growth and length of code change hotspots? You should as it helps you with tackling technical debt and risk assessment the right way!
In this podcast Adam Tornhill (@AdamTornhill) explains how he is applying data science and forensic approaches on data we all have in our organization such as: GIT commit history, ticket stats, static & dynamic code analysis, monitoring data … He is giving us insights in detecting code hotspots and how we can use leverage this data in areas such as risk assessment, the social side of code changes as well as explaining to business why working on technical debt is going to improve time to market.
Also make sure to check out CodeScene: A powerful visualization tool that uses Predictive Analytics to find social patterns and hidden risks in your code
Adam Tornhill on Twitter
https://twitter.com/AdamTornhill
Adam Tornhill's blog
https://www.adamtornhill.com/
CodeScene
https://www.empear.com
Did you know that Distributed Tracing has been around for much longer than the recent buzz? Do you know the history and future of OpenCensus, OpenTracing, OpenTelemetry and TraceContext? Listen to this podcast where we chat with Sonja Chevre, Technical Product Manager at Dynatrace, and Daniel Khan, Technical Evangelist at Dynatrace, about the past, current and future state of distributed tracing as a standard.
OpenTelemetry
https://opentelemetry.io/
OpenCensus
https://opencensus.io/
OpenTracing
https://opentracing.io/
TraceContext
https://www.dynatrace.com/news/blog/distributed-tracing-with-w3c-trace-context-for-improved-end-to-end-visibility-eap/
In 2018 Adrian Cockcroft was quoted with: “Chaos Engineering is an experiment to ensure that the impact of failures is mitigated”! In 2019 we sit down with one of his colleagues, Adrian Hornsby (@adhorn), who has been working in the field of building resilient systems over the past years and who is now helping companies to embed chaos engineering into their development culture. Make sure to read Adrian’s chaos engineering blog and then listen in and learn about the 5 phases of chaos engineering: Steady State, Hypothesis, Run Experiment, Verify, Improve. Also learn why chaos engineering is not limited to infrastructure or software but can also be applied to humans.
Adrian on Twitter:
https://twitter.com/adhorn
Adrian's Blog:
https://medium.com/@adhorn/chaos-engineering-ab0cc9fbd12a
Adrian Hornsby (@adhorn) has dedicated his last years helping enterprises around the world to build resilient systems. He wrote a great blog series titled “Patterns for Resilient Architectures” and has given numerous talks about this such as Resiliency and Availability Design Patterns for the Cloud at DevOne in Linz earlier this year.
Listen in and learn more about why resiliency starts with humans, why we need to version everything we do, why default timeouts have to be flagged, how to deal with retries and backoffs and why every distributed architect has to start designing systems that provide different service levels depending on the overall system health state.
Links:
Adrian on Twitter: https://twitter.com/adhorn
Medium Blog Post: https://medium.com/@adhorn/patterns-for-resilient-architecture-part-1-d3b60cd8d2b6
Adrian's DevOne talk: https://www.youtube.com/watch?v=mLg13UmEXlw
DevOne Intro video: https://www.youtube.com/watch?v=MXXTyTc3SPU
Susanne Kaiser (@suksr) has transformed her company from monolith on-premise into a SaaS solution running on a microservice architecture: Successfully! Nowadays she consults companies that need to find their “core domain”, break up and re-fit their architectures and organizational structure in order to truly get the benefit of microservices.
In this podcast you learn which questions you need to ask before starting a microservice project, how to find your true “core domain”, how to restructure not only your code but also organization and you get exposed to the concept of Wardley Maps which help you decide what to build vs what to outsource in order to deliver value to your end users the most efficient way.
Links:
Susannne on Twitter - https://twitter.com/suksr
DevOne Conference Page - https://devexperience.ro/speakers/susanne-kaiser/
Wardley Maps Microservices presentation - https://www.slideshare.net/SusanneKaiser3/preparing-for-a-future-microservices-journey-with-wardley-maps
To service mash or not? That’s a good question! Not every architecture and project needs a service mesh but for running distributed microservices architectures service mashes provide a lot of essential features such as service discovery, traffic routing, security, observability ..
We invited Matt Turner (@mt165), CTO at Native Wave, to tell us all we need to know about service mashes. We get a deep dive into Istio, one of the most popular current service mashes, the architecture and how the individual components such as Envoy, Pilot, Mixer and Citadel work together. We also chat about the tradeoff between performance, latency, throughput and service mash capabilities. If you want to learn more make sure to check out Matt’s online content such as blogs and recorded conference presentations on https://mt165.co.uk/.
Native Wave
https://nativewave.io/
Istio vs. Linkerd CPU Overhead Benchmarks by Michael Kipper
Initial Observations:
https://medium.com/@michael_87395/benchmarking-istio-linkerd-cpu-c36287e32781
Second Analysis:
https://medium.com/@michael_87395/benchmarking-istio-linkerd-cpu-at-scale-5f2cfc97c7fa
Keptn (@keptnProject) is an open source control plane for Kubernetes enabling continuous delivery and automated operations. In this session we chat with Dirk Wallerstorfer (@wall_dirk) who is leading the keptn development team. We learn from Dirk why they choose knative as serverless framework to let keptn connect to other DevOps tools in the toolchain, how the event driven architecture works, which use cases are supported and where the road is heading.
If you are interested also check out our Getting Started with keptn YouTube Tutorial, join the keptn slack channel, keep an eye at the keptn community and give feedback after trying out keptn yourself by following the following installation instructions: https://keptn.sh/docs/
Links:
keptn on Twitter - https://twitter.com/keptnProject
Dirk on Twitter - https://twitter.com/wall_dirk
knative - https://cloud.google.com/knative/
Keptn Video - https://www.youtube.com/watch?v=0vXURzikTac
Keptn Slack - https://keptn.slack.com/join/shared_invite/enQtNTUxMTQ1MzgzMzUxLTcxMzE0OWU1YzU5YjY3NjFhYTJlZTNjOTZjY2EwYzQyYWRkZThhY2I3ZDMzN2MzOThkZjIzOTdhOGViMDNiMzI
Keptn Community - https://github.com/keptn/community
Keptn Docs - https://keptn.sh/docs/
Can you explain Cloud Native? What are the key OpenSource frameworks you need to know? How about all these OpenSource Licensing models? Why do they exist? Which one to use? What are the monetization models and why to watch closely how Big IT & Cloud companies are impacting this space?
Carmen Andoh (@carmatrocity), Program Manager at Google and former Infrastructure Engineer at Travis CI, helps us understand how to navigate the Cloud Native & OpenSource world and gives answer to all the questions above. The IT world is changing but its up to us to shape the future by inventing it. If you want to learn more after listening check out the CNCF Trailmap and follow up with Carmen on social media to get access to her material around that topic!
Trailmap
https://github.com/cncf/trailmap
Imagine a future where we deploy every code change directly into production because feature flags eliminated the need for staging. Feature flags allow us to deploy any code change, but only launch the feature to a specific set of users that we want to expose to new capabilities. Monitoring the usage and the impact enables continuous experimentation: optimizing what is not perfect yet and throw away features (technical debt) that nobody really cares about. So – what are feature flags?
We got to chat with Heidi Waterhouse (@wiredferret), Developer Advocate at LaunchDarkly (https://launchdarkly.com/), who gives as a great introduction on Feature Flags, how organizations actually define a feature and why it is paramount to differentiate between Deploy and Launch. We learn how to test feature flags, what options we have to enable features for a certain group of users and how important it is to always include monitoring. IF you want to learn more about feature flags check out http://featureflags.io/. If you want to learn more about Heidi’s passion check out https://heidiwaterhouse.com/.
Self-Healing, Auto-Remediation: Magic words for most IT Leaders! When starting those kinds of projects teams realize their lack of maturity or even understanding of their current IT landscape to even think about Self-Healing. In other scenarios Self-Healing is misunderstood as a band-aid for “keeping the lights on” in order to buy more time for outstanding product improvements vs investing in the core architecture.
In this podcast we invited Jon Hathaway, CEO of HATech, and Jarvis Mishler, Solutions Architect Team Lead at HATech (@hatechllc), to learn about how they help organizations assess and improve the maturity of their IT Systems & processes, which auto-remediation actions they typically implement and why real self-healing is not just about keeping the lights on!
https://www.linkedin.com/in/jonhathaway/
https://hatech.io/
https://www.linkedin.com/in/jarvis-mishler/
https://twitter.com/hatechllc
Did you know that the JVM has 700+ configuration settings? Did you know that MongoDB performance can be improved by 50% just by tuning the right database and OS nobs? Every thought that slower I/O can actually speed up database transaction times?
In this episode we invited Stefano Doni, CTO at Amakas.io, who gives us a new perspective on how to approach performance optimization for complex environments. Instead of manually tweaking nobs on all sorts of runtimes or services they developed a Goal-driven AI-engine that automatically identifies the optimal settings for any application as it is under load. Make sure to check out their website and white papers where they go into details about how their algorithms work, which metrics they optimize and how you can apply their technology into a continuous delivery process
https://www.linkedin.com/in/stefanodoni/
https://www.akamas.io/
Have you ever used USE? Have you ever wondered what differentiates a performance tester from a performance engineer? Want to know how to automate performance engineering into DevOps Pipelines?
Twan Koot, Performance Engineer at Sogeti, is answering all these questions. We met him at the last Neotys PAC Event where he gave an in-depth look on metrics and enlightened us all with USE (a method from Netflix’s Brendan Gregg). In our conversation we explain what USE really is, how to apply it and how a good performance engineer needs to understand more than just response time!
Links:
Twan on linkedin - https://www.linkedin.com/in/twan-koot-a813a8b7/
Twan's deck from Neotys PAC - https://www.neotys.com/performance-advisory-council/twan-koot
Twans Video at Neotys PAC - https://www.youtube.com/watch?v=hV8wpkDUtys
http://www.brendangregg.com/
Brendan Gregg's home page - http://www.brendangregg.com/
eBPF - https://prototype-kernel.readthedocs.io/en/latest/bpf/
BCC - https://iovisor.github.io/bcc/
How many different continuous delivery pipelines do you have in your organization? Do you have dedicated teams that keep them up-to-date and constantly extend them with new tool integrations? Have you already built in capabilities for shadow, dark, blue/green or canary deployments? Is auto-mitigation and self-healing already on your internal pipeline roadmap? Sounds like a lot of manual work?
Keptn (@keptnProject)– an open source enterprise-grade framework for shipping and running cloud-native applications – is going to eliminate the manual efforts in building, maintaining and extending pipelines. Alois Reitbauer, Head of the Dynatrace Innovation Lab, gives us the background on how keptn evolved, which cloud native best practices are implemented as core capabilities, how to contribute to this project and gives us a glimpse into where the journey is going. Visit the about page and join the community and make sure to deploy keptn on your own Kubernetes clusters by simply following the step-by-step guides.
https://keptn.sh/
https://twitter.com/keptnProject
https://keptn.sh/about/
https://keptn.sh/docs/
Nicki (@nicki_23) was bored in finance, started to learn .NET development on the side and eventually won 250k at a hackathon she used for her startup. Now she is a “Digital” Technical Evangelist at AWS and spreads her passion about Serverless through twitch and shares her code examples on github
Tune in if you want to learn more about which things you should know about Serverless and Lambda. We chat about IAM permissions, timeouts, API Gateway and how a CI/CD Pipeline for Lambdas should look like
https://twitter.com/nicki_23?lang=en
https://www.twitch.tv/aws
https://github.com/kneekey23
Is Cloud Native just a synonym for Kubernetes? How to make sense of the sea of tools & frameworks that pop up daily? What can we learn from others that made the transformation and most of all: Where do we start?
We got answer to all these and many more questions from Priyanka Sharma (@pritianka) – Dir. of Alliances at GitLab and Governing Board Member at CNCF (Cloud Native Computing Foundation). In her work, Priyanka has seen everything from small startups to large enterprises leveraging Cloud Native technology, tools and mindset to build, deploy & run better software faster. She advises to start incrementally and whatever you do in your transformation make sure to always focus on: Visibility (which leads to transparency), Easy of Collaboration (which increases productivity & creativity) and Setting Guardrails (this ensures you stay compliant & avoids common pitfalls).
We ended the conversation around the idea of needing “Cloud Native Aware Developers” which can follow best practices or standards such as those promoted by CNCF or OpenSource projects such as keptn.sh
https://twitter.com/pritianka
https://www.cncf.io/
https://keptn.sh/
The .NET Runtime – whether .NET Framework or .NET Core – provides many ways to optimize memory management. But they don’t come in the form of configuration switches as we know if from Java. While there are a handful of settings, the .NET Runtime favors a different approach: asking developers to write memory aware software that follows a couple of core memory aware principles and best practices.
In this podcast we get to talk with Konrad Kokosa (@konradkokosa) – author of Pro .NET Memory Management. In his book he gives developers and operators great tips on how to optimize your .net applications and environments such as #1: start with proper monitoring; #2: reduce memory allocations; #3: well – for this and more you should check out Konrad’s book.
Listen in to a great discussion with somebody that has been working very close with the .NET Engineering Teams over the past years and brings the internal secrets of .NET Memory Management to everyone out there that wants to write Memory Aware .NET Software!
https://prodotnetmemory.com/
https://twitter.com/konradkokosa
Creating and maintaining test scenarios not only takes a lot of time, but means we are creating artificial test scenarios based on what we think users are going to do versus replicating real users behavior. In this episode we invited Thomas Rotté, one of our friends from https://probit.cloud, who solved these problems for their work at KBC Bank. Their solution is an AI that learns behavior of real user traffic, creates a probability model for most common user journeys and uses that model to create automation test scripts on the fly for automated, real user simulating test bots. We also learn how GDPR and other challenges influenced their solution and how they are now working with other tool vendors and enterprises to bring this technology to the market.
How to you scale a startup, a mid size company or an enterprise software organization? Can we learn from the Spartans or the Romans? And how can we explain DevOps to a Dummy?
In this fun filled episode with Emily Freeman (@editingemily), Cloud Developer Advocate at Microsoft, we get answers to all these questions and get inspired to join Emily’s appearance at the upcoming devone.at conference in Linz, Austria where she dives deeper into how to successfully scale development organizations from startup to enterprise. Later in 2019 make sure to watch out for the written version of our discussion on DevOps for Dummies – Emily is using her writing skills to bring it to paper!
https://emilyfreeman.io/
Kurt Aigner gave a session about managing hybrid system complexity, from the cloud to the mainframe and everything in between. He shares a few notes and tips in this discussion.
In this episode, Andi has a coffee and a chat with Dynatrace's Chief Product Officer Andreas Lehofer where they dig a little deeper into AI Ops Enhancements.
Gary Carr shares his experience with Dynatrace IaaS cloud support and a deep dive into how Dynatrace open AI provides intelligence into the capabilities and technologies of Azure, GCP and AWS.
Jimmy Stewart of Kroger, along with Michael Timmers and Kamala Dasika from Pivotal Cloud Foundry, discuss Kroger’s migration to PCF and how they tackle monitoring with the Dynatrace Bosh Agent
Long-time Dynatracer Nestor Zapata chats with us about Citrix’s fundamental shift from reactive to proactive and predictive operations; moving from data sets and charts to AI-powered answers. His session detailed advantages of a “Gen 3” monitoring approach and how to get there.
Ryan Murphy shares his experiences with leveraging Dynatrace to help deliver value to customers and partners, especially focusing on Cloud provisioning and multi-environment configuration management.
JP Morgenthal, CTO of application services at DXC Technologies talks to us about all the things devops you were afraid to ask.
We take a deep look at the new Session Replay news with Senior Director of Product Management Simon Scheurer
Steven Marrocco shares his experience with automated monitoring and management using Dynatrace for virtualized environments leveraging Blue Prism for problem-solving, collaboration and knowledge insight algorithms. No actual robots were harmed in this recording.
E.G. Nadhan, Chief Technology Strategist at Red Hat, talks to us about leveraging open source innovation to maximize performance.
https://twitter.com/NadhanEG
Conversation with Ben Rushlo, VP of Services at Dynatrace, talks to us about dashboarding in Dynatrace
Wolfgang Beer talks about how to make the most of our latest innovations and enable you to automate and manage your operations at web-scale. He shares insights from his session on management zones, API integrations, deployment automation, and best practices using our open AI.
In this episode Dynatrace's Michael Villiger shares some tips about his work with Humana to chose PCF as a key platform to power their digital transformation.He reviews the culture change and new methodology that helped to exposed gaps in existing monitoring practices and tools. And how Humana's strategy for the future using public cloud deployment, continuous monitoring with DevOps, and monitoring tool consolidation.
Sonja Chevre reviews her session on Dynatrace enables Performance Engineering as a Self-Service. She chats about how you can integrate performance testing tools with Dynatrace and how to embed performance diagnostics into the development work-flow for faster automated feedback.
Carmen Puccio, Principal Solutions Architect at AWS, talks to us about the Dynatrace Managed AWS quick-start program we created with him, modeling your monolith as a microservices platform, some of the Dynatrace announcements from Perform as well as Carmen’s snow adventure with guest of the show Mandus Momberg
James and Mark will just chat about everything today, and some NOTD stories and Super Bowl prep! Brian gets ready for his birthday dinner.
Alois Mayr shares a few new ideas about Dynatrace and Kubernetes, in addition to looking into what Dynatrace plans to release this year.
Dynatrace's Adam Dawson gives a few tips on gaining visibility across your entire environment, including infrastructure-only nodes with a single, all-in-one solution. His session showed how to extend Dynatrace to monitor your cloud infrastructure health, in addition to your application performance.
We chat with Henrik about how awesome Francis is, the new integrations anc capabilities for Neotys and Dynatrace, and the upcoming Neotys Performance Advisory Council in Chamonix, France
Peter Hack joins us to chat about how Dynatrace helps you to get better visibility into your PaaS environment; from deployment to automation. He shares a few tips from his deep dive session into OpenShift featuring Red Hat.
We catch up with Chris Morgan of Red Hat here at Dynatrace Perform 2019 and he shares several cool things about latest versions of OpenShift and automated operations and integrations with Dynatrace.
Guido Deinhammer talks about how to make the most of our latest innovations and enable you to automate and manage your operations at web-scale. He shares insights from his session on management zones, API integrations, deployment automation, and best practices using our open AI.
Dominik Punz shares his insights on his session about Dynatrace support for mobile platforms and how to intelligently monitor your mobile apps and what you can discover from troubleshooting, to enhancing user experience.
Daniel Khan talks about how Autodesk transformed customer’s businesses using AWS Lambda by simplifying processes, demonstrating real-world examples about how Dynatrace provides deep insights into Lambda functions and efficiency.
In this episode Jason Westerhouse shares details on how KeyBank is integrating Dynatrace across their DevOps toolchain and driving automated quality gates, continuous feedback and faster incident response time by embracing GitFlow – using Jenkins, containers and release automation to automate CD.
In this episode Trevor and Kristof give some practical guidance on how to make the vision of autonomous cloud management a reality, including the changes needed, the steps required to get there and benefits you will reap once you have arrived.
Florian Ortner reviews some of the top Dynatrace integrations and support for all the major cloud IaaS and PaaS platforms. He shares a few ideas about AWS, Azure, Google Cloud Platform, Pivotal Cloud Foundry, OpenShift and Kubernetes, in addition to looking into what Dynatrace plans to release this year.
Live from Dynatrace PERFORM 2019 in Las Vegas, it's Pure Perfbytes Performance Welcome Reception
If you haven’t heard about the 6-R Migration Patterns, then you probably haven’t heard about the 7-R. The 7th stands for R(e-Fit).
In this podcast we chat with Mandus Momberg (@MandusMomberg), Principal Solution Architect at AWS. Mandus is sharing what he has learned from small to large scale application migration & modernization projects. We learn about Modernization Factories, that it is key to have decision maker buy-in and that the most common migration scenario is R(e-Fit).
Our biggest takeaway are the 3 key measures of success after a migration: Availability, Elasticity & Agility! Now listen in …
https://www.linkedin.com/in/mandusm/
https://aws.amazon.com/migration-acceleration-program/
Migrating from your data center to the cloud is no easy task. In this episode, Patrick Kemble (@PatrickKemble), CTO at SambaSafety, shares their journey from lift & shift to AWS to re-architecting their applications using Cloud Foundry. Along the way, of course, they discovered many important aspects of monitoring.
https://twitter.com/patrickkemble?lang=en
https://www.sambasafety.com/
This episode is a recap of Andi’s presentation at AWS re:Invent where he talked common use cases Operation Teams have been auto-remediate over the years and how now Site Reliability Engineering (SRE) Teams take it to the next level. The key point of Andi’s message is to not only auto-remediate these and newer cloud native use cases in production. It is about shifting-left and preventing them upstream in the delivery pipeline. If you want to learn more check out Andi’s blog or watch the recorded session from re:Invent on YouTube.
Also make sure to listen until the end to learn about how you can mail your Christmas wishes to either Santa Claus or the Christkind!
Blog:
https://www.dynatrace.com/news/blog/shift-left-sre-building-self-healing-into-your-cloud-delivery-pipeline/
Video:
https://www.youtube.com/watch?v=PsI4pc0NtoI
Encore Presentation:
Goranka Bjedov ( https://www.linkedin.com/in/goranka-bjedov-5969a6/ ) has an eye over the performance of thousands of servers spread across the data centers of Facebook. Her infrastructure supports applications such as Facebook Social Network, WhatsApp, Instagram and Messenger. We wanted to learn from her how to manage performance in such scale, how Facebook engineers bring new ideas to the market and what role performance and monitoring plays.
Azure DevOps, formerly known as VSTS, is more than just a set of tools. But what is it exactly? How does it help enterprises to deploy better code faster? Does it only work for Azure or other platforms & clouds as well? How can it be extended or integrated into existing processes and tools?
Abel Wang, Sr Cloud Developer Advocate at Microsoft, is giving us a tour through Azure DevOps and how he has seen it implemented and integrated into existing enterprise DevOps tool landscapes. We briefly discussed the Unbreakable Delivery Pipeline for Azure DevOps that Abel helped implement and is now available on the Visual Studio Marketplace. So – give it a try and see for yourself what Abel and team has built!
Last but not least we also touched upon Azure DevOps for databases and how you to implement regression testing, continuous deployment and canary releases for database updates. Very intriguing topic that we are sure to cover in future sessions in more detail.
https://abelsquidhead.com/index.php/2018/08/03/the-dynatrace-unbreakable-pipeline-in-vsts-and-azure-bam/
https://marketplace.visualstudio.com/items?itemName=Safiahabib.DynatraceUnbreakablePipeline
https://twitter.com/AbelSquidHead?lang=en
Happy Guy Fawkes Night!
In the first episode with Ben Rushlo, Vice President of Dynatrace Services, we learned about things like not getting fooled by Bot traffic, which metrics to monitor and how RUM can replace your traditional site analytics.
In this episode we dive deeper into RUM use cases around user behavior analytics, bridging the silos between Dev, Ops & Business and elaborate on why blindly optimizing individual page load times is most likely wasted time as you won’t impact what really matters: End-to-End User Experience!
In our discussion we also talked about UX vs UI as well as importance of Accessibility. Here two links we want you to look at: Holger Weissboeck on Let’s put U in UX and Stephanie Mcilroy’s presentation at DevOne.
Listen to Episode 70:
https://www.spreaker.com/user/pureperformance/070-exploring-real-user-monitoring-with-
Let's Put the U in UX:
https://www.youtube.com/watch?v=qi19hls9LfY
Stephanie Mcilroy’s presentation at DevOne:
https://devone.us/speakers/#stephaniemcilroy
Data vs. Info article Brian mentioned:
https://medium.com/@copyconstruct/monitoring-in-the-time-of-cloud-native-c87c7a5bfa3e
Did you know that Azure Service Fabric is used by most of Microsoft’s global high scale services such as Bing, Dynamics or Xbox)? It’s a battle tested distributed systems platform that enables developers to deploy, manage and scale their microservices. In this session we have Sravan Rengarajan, Program Manager at Microsoft Azure, giving us an overview of the key use cases, how Service Fabric started and in which direction it is heading. We also learn how you get your own free local version of Service Fabric and why Service Fabric gets us towards real Serverless computing. Additional information can be found on the Service Fabric GitHub codebase – yeah – its all out there on GitHub!
https://www.linkedin.com/in/sravan-rengarajan/ - Sravan on Linkedin
http://aka.ms/servicefabricdocs - learn more about Service Fabric
http://aka.ms/servicefabricmesh - learn more about Mesh
http://aka.ms/tryservicefabric - free clusters to party on!
https://github.com/microsoft/service-fabric - GitHub codebase
Minecraft – the hugely popular sandbox video game – might not be your traditional software to monitor with an APM (Application Performance Management) tool. But Mike Villiger did it anyway in order to learn some advanced concepts in application monitoring such as custom entry points, thread diagnostics, method hotspots or simply to figure out why his mod’ed Minecraft sometimes couldn’t keep up with processing all the changes and skipped cycles. By using Dynatrace – first AppMon now Dynatrace SaaS – he learned more about the internals of Minecraft, how the single threaded architecture calls each mod and why a single mod must not take longer than 50ms to process. Mike gives us insights into which problems he found within Minecraft but more importantly what he takes away for his daily job as a performance advocate and evangelist. If you have any interesting side projects where you use APM tools let us know – there is always something to learn from every project!
https://minecraft.net/
https://twitter.com/mikevilliger
Brett Hofer is giving us his inside story on how he was called for the rescue to break a monolithic healthcare system that, after 3 years of development, was on the verge of having a major business impact on the largest healthcare vendor in the US. We learn about his strategic decisions such as quieting the system, establish traceability and most importantly: setting up a separate team that broke the monolithic while keeping it in sync with the main branch development. Brett Hofer is now Global Practice Lead at Dynatrace where he and his team help Dynatrace customers successfully walking through their digital transformations.
https://www.linkedin.com/in/brett-hofer-2432572/
Ben Rushlo, Vice President of Dynatrace Services, specializes in the Digital Experience. In this episode, Ben talks to us about Real User Monitoring. What happens when good bots go bad? Can Real User Monitoring (RUM) replace your traditional site analytics? If you have RUM, is there any reason to also use synthetics? What performance metrics are the best when it comes to monitoring the end user? How does RUM help you understand the performance of business? Tune in to episode 70 of PurePerformance for answers to these questions.
https://www.linkedin.com/in/benrushlo/
Serverless has been a hot topic for quite a while, but we are still in the early stages when it comes to best practices and tooling. Justin Donohoo, Co-Founder of observian.com, gives us the pros and cons of 4 architectural patterns that he calls: “Microservice / nano pattern”, “Service Pattern”, “Monolithic Pattern” and the “GraphQL Patterns”. Besides these patterns we also learn about common cost traps and how to “architecture around them”. For more information on serverless Justin also shared his recent Serverless Meetup presentation. And stay tuned – there will be more from Justin around secrets, containers and anything else there is to know about cloud native applications.
In our previous episode with Chris Burrell, Head of Technology at Landbay, we learned how they got rid of end-to-end testing in order to speed up continuous delivery. Today we discuss how they still make sure that no code changes in their microservice architecture breaks end-to-end use cases by leveraging Contract-based Testing using Swagger and several tools in the Swagger ecosystem, e.g: diff, code generation … - also make sure to check out Chris’ presentation at yCon called “CDC is dead – long live swagger”.
How can you get your build times down to minutes? Exactly: eliminate the largest time consumer! At Landbay, where Chris Burrell heads up technology, this was end-to-end testing.
Landbay deploys their 40 different microservices into ECS on a continuous basis. The fastest deployment from code to production is 6 minutes. This includes a lot of testing – but simply not the traditional end-to-end testing any longer. Chris gives us insights in contract testing, mocked frontend testing, how they do Blue/Green deployments and what their strategy is when it comes to rollback or rollforward.
For more information watch Chris’s talk on 6 minutes from Code Commit to Live at µCon and his lightening talk on CDC Testing is Dead – Long Live Swagger.
https://twitter.com/ChrisBurrell7
https://skillsmatter.com/skillscasts/10714-6-minutes-from-code-commit-to-live-you-won-t-believe-how-we-did-it
https://skillsmatter.com/skillscasts/11147-lightning-talk-cdc-testing-is-dead-long-live-swagger#video
Have you heard about Load Shedding? If not then dive into this discussion with Acacio Cruz, Engineering Director at Google ( https://twitter.com/acaciocruz ). He walks us through what Google learnt from one of the early outages at Gmail and how he and his team are now applying concepts such as load shedding to avoid disruption of their services despite spikes of load or unpredictable requests. We also discuss SRE (Site Reliability Engineering), how it started and transformed at Google and how we should think about automation, configuration of automation, and automation of automation. For more details – including visuals – we encourage you to watch Acacio’s breakout session from devone.at on YouTube (Load Shedding at Google).
https://devone.at/speakers/#acaciocruz
https://www.youtube.com/watch?v=XNEIkivvaV4
Security is on everyone’s mind. One way to strengthen security of your software and increase the awareness of your engineers is running a Security Hackathon – or a “Bug Bounty Program”. We invited Pascal Schulz ( https://www.linkedin.com/in/pascalschulz/ ), Security Engineer at Dynatrace, to the show to give us more background on HACK.DT – a security hackathon he and his team ran earlier this year within the Dynatrace Engineering Labs. For additional details check out his blog Running a successful internal bug bounty program and ping him on twitter (@PascalSec) in case you have further questions.
https://www.dynatrace.com/news/blog/running-a-successful-internal-bug-bounty-program/
We got to chat with Danilo Poccia (@danilop), Global Serverless Evangelist at Amazon Web Services, on how to best leverage serverless and its new principles to speed up bringing new features to the market. We learn about Event Driven Architectures, Continuous Deployment into Production leveraging Canary and Linear Deployments as well as how to automate testing when pushing your serverless code through CI/CD. Also – did you know that you can run all your Lambda tests locally in your machine? Check out AWS SAM ( https://docs.aws.amazon.com/lambda/latest/dg/serverless_app.html ) and SAM Local ( https://docs.aws.amazon.com/lambda/latest/dg/sam-cli-requirements.html ) for more information.
Make sure to check out Danilo’s Serverless by Design website ( https://sbd.danilop.net/ ) where it you can visually create your end-to-end serverless architecture and get a CloudFormation template to stand up this environment in your AWS account.
Donovan Brown, Principal DevOps Manager at Microsoft, is back for a second episode on CI/CD & DevOps. We started our discussion around “The role of Monitoring in Continuous Delivery & DevOps” but soon transferred over to our recent most favorite topic “The Unbreakable Delivery Pipeline”. Listen in and learn more about how monitoring, monitoring as code and automated quality gates can give developers faster and more reliable feedback on the code changes they want to push into production.
Also make sure to follow up on Donovan’s road show when he shows Java developers how to build an end-to-end delivery pipeline in 4 minutes. And lets all make sure to remind him about the promise he made during the podcast: Building a Dynatrace Integration into TFS and adopt the “Monitoring as Code” principle
15 Minutes of your day! That’s all it takes to make the first step towards applying DevOps best practices. To hear more about this and other suggestions on how to jump start your DevOps transformation tune into this episode where we chat with Donovan Brown ( http://donovanbrown.com/ ), Principal DevOps Manager at Microsoft.
Did you know that over the last 7 years the VSTS team has increased deployment velocity from once every 3 years to once every 3 weeks? Coordinating 50 different feature team that all commit to master daily? If you always thought that transformation like this are only possible in smaller development organizations then be proven wrong by Donovan, a member of the League of Extraordinary Cloud DevOps Activists. If you want instant advice simply tweet using #LoECDA and summon “the league”!
Serverless comes with its own set of best practices, quirks and benefits when it comes to monitoring and performance engineering.
In this episode we have Michael Garski, Director of Platform Engineering at Fender Musical Instruments ( https://www.linkedin.com/in/mgarski/ ), giving us a technical deep dive into lessons learned and best practices they learned when re-platforming their architecture to AWS Lambda. We get to learn about optimizing Cold Starts, Re-Using HTTP Connections, Leveraging API Gateway Caching and finding the sweet spot for CPU & Memory settings to optimize price/performance of AWS Lambda executions.
For more details check out Michael’s slides on Innovating Through React Native Mobile Apps. ( https://www.slideshare.net/AmazonWebServices/innovating-through-react-native-mobile-apps-with-fender-musical-instrumentspdf )
Josh Long ( https://twitter.com/starbuxman ), Developer Advocate at Pivotal, Java Champion and author of 5 books, gives us a great tour through the latest that is happening in the Spring Universe. If you are new to Spring check out http://start.spring.io/ and create your first project within minutes. When it comes to Reactive make sure to check out https://projectreactor.io/ and dive into https://micrometer.io/ to learn more about how to extract metrics from Spring applications. As Josh is constantly traveling the world chances are high you can meet him at a local conference.
Visual Replay gives you full film-like replay of your end users, including clicks, mouse moves swipes and scrolls. It helps you optimize user experience by addressing problems where end users struggle, e.g: not finding that button, an overlay dialog hiding critical elements or a 3rd party browser plugin that messes with your page. It also supports compliance use cases such as allowing you to proof what information you really showed to the end user when they conducted online business with you.
To learn more about this use cases and the technical implementation details of visual replay technology we invited Simon Scheurer, Chief Software Architect at Dynatrace (https://www.linkedin.com/in/simonscheurer/), to this podcast. He educates us on the latest of this disruptive technology!
And besides that we also learn about how awesome Simon’s hometown Barcelona, Spain is.
You wouldn’t build your own Jenkins – would you? Neither would you build your own CRM, Office or Email service. So why are the “cool” DevOps kids still building their own continuous delivery scripts, log analytics and monitoring and showing it off on GitHub or conferences?
In this episode we invited Steve Burton (@BurtonSays), CD geek at harness.io, and discussed the current state of Continuous Delivery and the role of Observability (that’s Monitoring++). We learn about use cases that commercial vendors in these spaces provide out-of-the-box, the APIs they offer to integrate these tools into a larger eco-system and why we believe it’s time to stop building your own tools but investing in building better software for your users. We also learn about Blue/Green deployments, Canary Releases, Continuous Verification and Rollback vs Roll-forward.
If you still believe Blockchain is just about Bitcoin or that Blockchain is a super safe, high performing platform that simply runs then listen in to this podcast.
With David Jones ( https://twitter.com/davidlewisjones ) –AIOps Evangelist – we learn about the different use cases of Blockchain technology, the two top frameworks Ethereum ( https://www.ethereum.org/ ) and Hyperledger ( https://www.hyperledger.org/ ), and also discuss how to monitor both usage and operation of Blockchain to ensure performance for end user applications. As a great read check out his recent blog post on AI-based Monitoring to ensure Blockchain Performance: https://www.dynatrace.com/blog/using-dynatrace-ai-based-monitoring-ensure-blockchain-performance/
New to Kubernetes? Already a pro? In both cases, tune in to this episode, as we have something for both sides of the aisle.
Kubernetes seems to have won the container orchestration game. Major cloud and PaaS vendors are supporting Kubernetes, and attendance at KubeCon in Dec 2017 skyrocketed. Today we chat with Brian Gracely ( https://twitter.com/bgracely ), Director of Strategy at Red Hat. Brian also co-hosts @PodCTL ( https://twitter.com/PodCTL ) – a podcast dedicated to containers, OpenShift, Kubernetes, and Cloud Native. In our chat we learn where and what Kubernetes is right now, where its heading (e.g: providing better onboard experience with developers, more APIs …), why we have to pay attention to Service Mesh ( http://philcalcado.com/2017/08/03/pattern_service_mesh.html ), and why it is important to have a good cross technology monitoring strategy that supports both your brown field legacy services as well as the green field cloud native. We also enlighten you about what the BWI (Brian Wilson Indicator) is!
James Turnbull ( https://jamesturnbull.net/ ) is an author of 10 books on topics like Docker, Packer, Terraform, Monitoring, … and is currently writing a book on Monitoring with Prometheus https://prometheusbook.com/ . We got to chat about what modern monitoring approaches look like, how to pull in developers to start building monitoring into their systems and how to bridge the gap between monitoring for operations vs monitoring for business. Having a monitoring expert like James that knows many tools in the space was great to validate what we at Dynatrace have been doing to solve modern monitoring problems. We learned a lot about key monitoring capabilities such as capturing data vs capturing information, providing just nice dashboards vs providing answers to known and unknown questions and making monitoring easy accessible so that monitoring can benefit both business, operations and developers.
We hope you enjoy the conversation and learn as much as we did. A blog we have been referencing several times during the talk was this one from Cindy Sridharan on Monitoring in the time of Cloud Native: https://medium.com/@copyconstruct/monitoring-in-the-time-of-cloud-native-c87c7a5bfa3e
Steve Pace, Senior Vice President of Global Sales at Dynatrace, discusses the positioning and offerings of Dynatrace in the AWS Marketplace
We are learning more and more about these deceptively simple-sounding improvements and new features coming for Dynatrace, let's break it down a little more.
We stepped aside for just a few minutes to learn more about RedHat's OpenShift products with Chris Morgan. We chat about their experience in building integration between Dynatrace and OpenShift, excitement about the conference announcements and a shared distrust of mustard-based barbecue sauce in South Carolina.
Live from the conference on day 1, the early annoucements, reflections, excitement, caffeine and live stream from the conference: http://perform.dynatrace.com
Markus Heimbach, Team Lead of the Infrastructure and Service Team at Dynatrace, explains the continuous delivery process of www.dynatrace.com really works behind the scenes. 2 years ago the web site team used a traditional CMS (Content Management System) which was slow, error prone, and didn’t deliver the expected end user experience for visitors of our website. 2 years later Markus and his team built a fully automated “Content Delivery Pipeline”. The team decided to leverage Git, static generated web content, immutable infrastructure, and Dynatrace OneAgent monitoring. Production deployments happen twice a day but staging and development deployments – using the same deployment pipelines – happen much more frequently. The result is a very flexible delivery pipeline, fully version controlled content, a very secure and fast website and everything monitored with Dynatrace. Thanks Markus for letting us look behind the scene of www.dynatrace.com
Feature Toggles or Feature Flags are not new – but they are a hot topic as they allow safer and fearless continuous delivery. Finn Lorbeer ( https://twitter.com/finnlorbeer ) gives us technical insight into how he has been implementing feature toggles in projects he was involved over the last years. We learn why he loves https://github.com/heartysoft/togglez, how to test feature toggles, monitor the impact of features being toggled and how to make sure you don’t end up in a toggle mess.
Have heard about “Shifting Left”? Well – get prepared to hear that Shift-Left is not the only solution to building a high quality products. Finn Lorbeer ( http://www.lor.beer/ ) is a Product Quality Specialist working for Thoughtworks. In a recent presentation given at Quest4Quality ( http://questforquality.eu/speakers/finn-lorbeer/ ) in Dublin he explained how being a quality engineer is no longer about being seen as a quality gate (and sometimes bottleneck) in the deliver cycle. Finn is sharing his experience from a recent project with a large German automobile company where he helped transform the development teams to shift-left on quality but not only for the software they develop and test, but also for the type of features they implemented and how they see quality as a whole in their end-to-end delivery cycle. Interesting lessons learned on how to speed up delivery, increase quality and make everyone part of the game!
Erik Landsness, Director Network Operations Center & SRE at Beachbody, talks us through his last 1.5 years in his role where he has been transforming the role and culture of the traditional NOC team from human-based Dashboard analytics to a Automated Self-Healing Zero-Dashboard Culture. While they haven’t yet reached that end state they have made big strides. Erik shares with us how to gradually transform into a modern operations team that automates things that humans shouldn’t do – such as staring at dashboards on walls 😊
Erik is also presenting at Dynatrace PERFORM 2018. Make sure to check out his session to learn first hand!
https://www.dynatrace.com/perform/speakers/
Are you still deploying machines manually? Do you have to login to machines to apply changes? Do you spend hours or even days to detect infrastructure issues messing with your test execution or even production? We have the answer for your pain: Listen to this podcast!
Markus Heimbach leads the Infrastructure and Service team at Dynatrace and explains how they got rid of Snowflakes (not in the political sense), tackled the Configuration Drift issue, and how his team became a Service Organization powering the innovation at Dynatrace R&D. Get a glimpse of his talk track from his presentation at #devone.at - https://speakerdeck.com/markusheimbach/infrastructure-as-code
As another teaser: you will hear about Test Automation of Infrastructure Code, leveraging Docker and Kubernetes (k8s) and how to use and leverage Immutable Infrastructure!
We typically hear about agile transformation being driven from development and eventually pushing it towards operations. But it doesn’t have to be that way as we hear from Nestor and Abeer who helped transform their operations team from Waterfall (Traditional Ops), to Partial Scrum (Intro to Agile) and then Kanban (more defined structures for Ops using Agile principles). Listen in and learn what the differences are between Agile in Dev and Agile in Ops, which metrics they use to measure the success and how they are now pushing towards a DevOps transformation from Ops towards Dev.
If you want to chat live with Nestor and Abeer then take the chance and meet them at PERFORM 2018 ( http://perform.dynatrace.com ) where they give us more insights into their transformation
Most of us remember the DDOS attack last year executed through thousands of Security Camera IoT devices. This raised security questions around IoT but also helped the public to understand that IoT (Internet of Things) is a real thing.
In this session, we learn from Harald Zeitlhofer ( https://twitter.com/HZeitlhofer ) why he rather likes to call this hot trend IoE (Internet of Everything), what the key use cases of IoE are and how proper monitoring of these devices might have been the key to detect the attack before it actually happened.
To learn more about this exciting next big thing we suggest to start with Harald’s latest blog posts on his most favorite topic.
If you believe OpenStack and OpenShift are pretty much the same thing. you better listen to this episode with Martin Etmajer ( https://twitter.com/metmajer ). He explains what OpenShift is, how it differentiates from Cloud Foundry and other PaaS platforms, and which major contribution it can have to successful DevOps transformations.
To put it in his words: OpenShift provides great user experience for developers to push their code changes automatically, packaged as containers, into different environments without having to worry about where and how these containers run or how they scale up & down. You should also check out his presentations from Red Hat Summit on Monitoring and Logging in OpenShift ( https://www.slideshare.net/martinetmajer ) as well as more material on http://www.dynatrace.com/openshift.
Project Jigsaw, G1 as default garbage collector, ahead-of-time compilation, Stack Walking API and many more changes that you should be aware of when upgrading to Java 9. Philipp Lengauer, whom we met at devone.at, gives us all the answers and technical deep dive into all these JVM changes. Especially for performance engineers an episode worth while listening to.
If you want to learn more check out Philipps presentation at devone:
https://youtu.be/Nsg_rhlf4_U?list=PLfi6VUNSzNYmUyeZ2BTM_WmZjgi0FOqRl
If you thought EC2 was the first service offered by Amazon Web Services and if you think 53 in “Route 53” is just a random number then you should listen to this 101 on AWS Podcast. This time we got to chat with Wayne Segar ( https://www.linkedin.com/in/wayne-segar-6222ba57/ ) who has been helping companies to move to new cloud technologies and services such as AWS. Wayne gave us a great overview of the key services in Compute, Database, Storage, Management, Development as well as how Monitoring works with AWS.
If you want to make your first steps with AWS, such as deploying your first EC2 Instance or Application on Elastic Beanstalk, then feel free to follow our 101 AWS Monitoring Tutorial:
https://github.com/Dynatrace/AWSMonitoringTutorials
Why would I move to .NET Core? If I move, can I just recompile my .NET code with the new .NET Core and run it on Linux? Or is there more to it? What is .NET Core at all and what does it provide as compared to ASP.NET Core? Can I still monitor my .NET Applications the same way as in the past or is there a new approach for tracing and monitoring? And is it true that all of this is now available on GitHub as Open Source project?
Get answers to all these questions by listening to this episode where we got to talk with Christoph Neumueller @discostu105 ( https://twitter.com/discostu105 ) and Gergely Kalapos @gregkalapos ( https://twitter.com/gregkalapos ). Christoph and Gergely are two lead engineers for the Dynatrace .NET Agent technology. They are also code contributors to the .NET Open Source and other open source projects such as SuperDump ( https://github.com/Dynatrace/superdump ). Also check out their blogs ( https://www.dynatrace.com/blog/tag/net/ ) to learn more on .NET Core, ASP.NET Core and other .NET relevant performance topics.
And, if you're a Dynatrace customer, make sure to up-vote the RFE to allow hot sensor placement for .net core. ( https://answers.dynatrace.com/spaces/151/product-feedback-and-enhancement-requests/idea/183988/hot-sensor-placement-for-net.html )
Additional Links:
TechEmpower Benchmark ( https://www.techempower.com/benchmarks/ )
Blog about performance improvements as a result of community input ( https://blogs.msdn.microsoft.com/dotnet/2017/06/07/performance-improvements-in-net-core/ )
Download and try out .Net Core ( http://dot.net )
Dynatrace Free Trial ( https://www.dynatrace.com/trial/?vehicle_name=www.spreaker.com )
Visually Complete and Speed Index have been introduced to better measure real end user performance experience. Klaus Enzenhofer @kenzenhofer ( https://twitter.com/kenzenhofer ) gives us a detailed description of these metrics, how they are getting calculated, and which problem they solve. What we also learn in this 101 is why now we finally have these metrics available not just for synthetic monitoring but also for real user monitoring. This can be attributed to the advances in browser technologies as well as to some smart engineering. In our discussion we also cover other recent advances and use cases in Web Performance Optimization – such as the usage of performance markers.
If you want to learn more check out the blogs from Google on Speed Index ( https://sites.google.com/a/webpagetest.org/docs/using-webpagetest/metrics/speed-index ) as well as the blog from Klaus on how Speed Index and Visually Complete made it into RUM offerings ( https://www.dynatrace.com/blog/visually-complete-speed-index-for-real-user-monitoring-rum/ ).
Spoiler Alert: Serverless doesn’t mean that we got rid of servers. We just don’t have to think about them anymore as we can focus on coding functions that get executed when triggered through certain events. Daniel Khan (@dkhan) tells us more about use cases of Serverless or as he likes to call it “Function as a Service” (FaaS). We also chat a lot about monitoring and the challenges of actually monitoring and debugging serverless code. It is still a young technology but constantly evolving.
It sounds like 3 buzzwords, But there is more than that. We were intrigued by the Digital Mastery & Joy ( https://info.dynatrace.com/apm_wc_panera_na_registration.html ) webinar Klaus Enzenhofer @kenzenhofer ( https://twitter.com/kenzenhofer ) did with Panera Bread. In his introductory statement, Klaus cited a recent study from IDG on Digital Customer Experience. The biggest challenges are data silos, poor data quality, redundant data, and missing coordination between departments that manage the individual digital touchpoint channels (Mobile, IoT, Web, Physical, …). In our discussion we find lots of parallels between the problem that DevOps tries to solve and which challenges digital transforming businesses face: Silos! Disconnected Silos! But instead of Silos between Dev & Ops its Silos between your Business Teams that are all strictly focusing on their slice of bread (to reference some great stories from Prashant Karre, Director of Performance Engineering at Panera)
Listen in and join our conversation. Make sure to check out the webinar recording Digital Mastery & Joy ( https://info.dynatrace.com/apm_wc_panera_na_registration.html )
What is Cloud Foundry? And why does Alois Mayr (@mayralois) say that Cloud Foundry is the most opinionated PaaS Platform in the world? Listen to this 101 show to get a good overview of what Cloud Foundry (CF) is, how it started and what offerings are available right now. Also learn what the main use cases are for CF Cloud Operators as well as for Developers that use the platform to push their applications and services.
If you want to learn more or see Alois in action make sure to watch his Full Stack Monitoring on Cloud Foundry PurePerformance Clinic. ( https://www.youtube.com/watch?v=vsQHtdeizXQ&index=65&list=PLqt2rd0eew1bmDn54E2_M2uvbhm_WxY_6&t=1237s )
If you wonder what the top 3 ways are to pronounce Azure, then check out this episode. Also, if you want to learn more about what Azure really provides, why it used to be ahead of the curve, and why Microsoft had to re-invent it to provide services that software companies really needed, you won't want to miss this episode. Martin Gutenbrunner (@MartinGoodwell) gives us a good overview of the key Azure services and use cases that make Azure an interesting platform for many enterprises. It might also be surprising that it is not just a Microsoft-lock in Technology Stack. Besides .NET, there are many technologies that companies can use to run their applications on Azure IaaS, but more so on Azure Service Fabric. Listen in and learn the core fundamentals of Azure, why it might be interesting for you and what role monitoring plays.
If you think Node.js is just a technology used by small start ups then you better listen to this 101 episode. Daniel Khan (@dkhan) – a member of the Node.js community and working group – answers a lot of questions on why large enterprises such as Walmart, Paypal or Intuit use Node.js to innovate. Daniel also explains the internals of Node.js, its event driven processing model, its non-blocking asynchronous nature, and how that enables a list of interesting use cases. We also discuss how to monitor and optimize applications running on Node.js and why that might be different for a developer as compared to an Ops team that runs Node.js in combination with other enterprise software.
What is OpenStack? Oh – it's not the same as OpenShift? So what is OpenStack? If these questions are on your mind and you want to learn more about why OpenStack is used by many large organizations to build their own private cloud offering than listen to this 101 talk with Dirk Wallerstorfer (@wall_dirk). We learn about the different OpenStack core controller services (Cinder, Horizon, Keystone, Neutron, Nova …) as well as the core cloud services (Compute, Storage, Network, …) it provides to its users. Dirk also explains why and who is moving to OpenStack and what the challenges and use cases are when it comes to monitoring OpenStack environments – both for an OpenStack Operator as well as for the Application Owners that run their apps on OpenStack.
Todd DeCapua has been a performance evangelist for many years. In his recent work and publications, which includes Effective Performance Engineering ( http://www.effectiveperformanceengineering.com/ ) as well as several publications on outlets such as TechBeacon ( https://techbeacon.com/contributors/todd-decapua ), he introduces DevOps best practices to improve the 5 S-Dimensions: Speed, Stability, Scalability, Security and Savings.
In our discussion with Todd we focused a lot on Security as it has been a more prominent topic in our industry recently. How to bake Security into the delivery pipeline and why it is such an important aspect. Automation seems to be the key which also includes automating functional checks, performance checks and – as we said: Security!
Related Links:
Follow Todd on Twitter ( https://twitter.com/AppPerfEng )
Follow Todd on LinkedIn ( http://www.linkedin.com/in/todddecapua )
Blog: How to build performance into your user stories ( https://techbeacon.com/how-build-performance-your-user-stories )
For our one year anniversary episode, we go “back to basics”, or, better said, “back problem patterns”.
We picked three patterns that have come up frequently in recent “Share Your PurePath” sessions from our global user base and try to give some advice on how to identify, analyze and mitigate them:
· Bad Multi-threading: Multi-threading is not a bad thing – but if done wrong it doesn’t allow your application to scale. We discuss key server metrics and how to correctly read multi-threaded asynchronous PurePaths. Also see the following blog: https://www.dynatrace.com/blog/how-to-analyze-problems-in-multi-threaded-applications/
· When Micro Service become Nano Services. This was inspired by a blog from Steven Ledoux ( https://www.dynatrace.com/blog/micro-services-when-micro-becomes-nano/ ). It's important to keep a constant eye on your micro-service architecture to avoid too tightly coupled or too fine grained architectures
· Garbage Collection Impact: GC is important but bad memory management and heavy GC can potentially impact your critical transactions. We discuss different approaches on how to correctly measure the impact of garbage collection suspension. If you want to learn more check out the Java Memory Management secton of our online performance book: https://www.dynatrace.com/resources/ebooks/javabook/impact-of-garbage-collection-on-performance/
In this second episode with Goranka Bjedov from Facebook, we learn details about how Facebook monitors their infrastructure, services, applications and end users. Why they built certain tooling, and how & who analyzes that data. We then shifted gears to development where we learned how the onboarding process of developers works and that Goranka herself made her first production deployment within the first week of employment. Join us and learn a lot about the culture that drives Facebook Engineering
Goranka Bjedov ( https://www.linkedin.com/in/goranka-bjedov-5969a6/ ) has an eye over the performance of thousands of servers spread across the data centers of Facebook. Her infrastructure supports applications such as Facebook Social Network, WhatsApp, Instagram and Messenger. We wanted to learn from her how to manage performance in such scale, how Facebook engineers bring new ideas to the market and what role performance and monitoring plays.
In the second episode with Rick Boyd (check out his GitHub repo - https://github.com/DJRickyB ) we talk about how performance engineering evolved over time – especially in an agile and DevOps setting. It’s about how to evolve your traditional performance testing towards injecting performance engineering into your organizational DNA, providing performance engineering as a service. Making it easy accessible to developers whenever they need performance feedback. Rick gives us insights on how he is currently transforming performance engineering at IBM Watson. We also gave a couple of shout outs to Mark Tomlinson and his take on Performance in a DevOps world!
We got Rick Boyd ( https://www.linkedin.com/in/richardjboyd/ ) – Application Performance Engineer at IBM Watson – and elaborated what Continuous Performance Testing is all about. We all concluded its about Faster Feedback in the development cycle back to the developers – integrated into your delivery pipeline. As compared to delivering performance feedback only at the end of a release cycle. We discussed different approaches on how to “shift left performance” with the benefit of continuous performance feedback!
Brett Hofer (@brett_solarch) has been engaged in numerous DevOps Transformation projects mainly for very large enterprises. We got to talk with him on this episode to learn more about how he assesses the status quo when he walks into an organization, what the top blocking items for a successful transformation are and what the best approaches are to implement the recommended changes. Spoiler alert: we talked a lot about IT Ops Automation, building cross functional teams and understanding and defining responsibilities and roles. If you want to learn more about what Brett is doing check out his blogs about DevOps on https://www.dynatrace.com/blog/author/brett-hofer/.
Thomas McGonagle just had his 10 years DevOps anniversary at it was 10 years ago when he got first exposed to Infrastructure as Code through Puppet. He is currently working with F5, helping Big IP Network Teams around the world automate the Network as part of their DevOps transformation.
We met Tom at a recent DevOps meetup in Boston which sparked this conversation on what “Metrics Driven Continuous Delivery” could mean for Network Operations Engineers. What are the metrics to look at? How to engage with the application teams to provision better and automated network resources? How to bake this into the Continues Delivery Cycle?
Besides NetOps Thomas is also passionate about CI/CD. He runs the largest Jenkins User Group in the World out of Boston, MA. If you happen to be around check out their next meetups and DoJo’s: https://www.meetup.com/Boston-Jenkins-Area-Meetup/
Mike Horwitz (https://www.linkedin.com/in/mike-horwitz-a40a139 ) has been working with Mainframe since the mid 80s. In this podcast he explains basic terminology and the challenges that come with the interaction to the distributed and cloud native world. Monitoring end-to-end is a critical capability especially when it comes to cost savings and including the mainframe components in a CI/CD/DevOps environment.
If you want to learn more about common mainframe performance and monitoring challenges check out our YouTube Performance Clinic: https://www.youtube.com/watch?v=8eodOw3gnMA&list=PLqt2rd0eew1bmDn54E2_M2uvbhm_WxY_6&index=55
INFO COMMENTS
We take time out chat with Jason Suss and Dynatrace RUM on RUM, we search for Richard Bentley at the nightclub, Vikram survived the performance puzzlers, Rick Boyd from IBM Watson and Stefan Baumgartner tells us all about his work at Dynatrace and podcasting at http://workingdraft.de
Brian Wilson takes time from his feverish disco dancing to have several interviews with attendees at the Tuesday evening party at Dynatrace 2017 in Las Vegas.
Good morning Las Vegas! We chat today about how to apply logic to Dynatrace Davis, we hear an interesting performance story from Lianggui and an industry update from Vice President of Consulting at CGI, Walter Kuketz.
Wrapping up the day with interviews from Henrik Rexed from Neotys and our new friends Rajesh Jain and Maggie Ambrose from Pivotal Labs.
We continue coverage of Dynatrace Perform 2017 with interviews from John Delfeld of Ixia, our favorite performance geek Andreas Grabner from Dynatrace and strategic partner Ryan Faulk of Faulk Consulting.
Woo hoo!! We're kicking off the Dynatrace conference in Las Vegas at the Cosmopolitan Hotel!
Eric Wright (@discoposse) is a “veteran” and expert when it comes to virtualization and cloud technologies. He introduces us into the field of container and container orchestrations, the vendors in the space, the pros and cons and the key capabilities he things have to be considered when evaluating the next generation virtualization platform for your enterprise.
If you want to learn more check out his podcast - http://gcondemand.podbean.com/ - as well as his publications on https://turbonomic.com/author/eric-wright/
How often have you deployed an application that was supposed to be load tested well but then crashed in production? One of the reasons might be that you never took the time to really analyze real life load patterns and distributions. Brian Chandler (@Channer531) (https://www.linkedin.com/in/brian-chandler-8366663b ) – Performance Engineer at Raymond James – has worked with their Operations Team to not only start loving application specific performance data captured in production. They starting breaking down the DevOps Walls from Right to Left by sharing this data with Testers to create more realistic load tests but also started education developers to learn from real life production issues.
We hope you enjoy this one as we learn a lot of cool techniques, metrics and dashboards that Brian uses at Raymond James. If you want to see it live check out our webinar where he presented their approach as well: https://info.dynatrace.com/apm_wc_getting_started_with_devops_na_registration.html
You can view the screenshots we refer to at:
https://assets.dynatrace.com/en/images/general/Chandler_01.jpg
https://assets.dynatrace.com/en/images/general/Chandler_02.jpg
HAPPY NEW YEAR
Daniel Freij (@DanielFreij) – Senior Performance Engineer and Community Manager at Apica – has been doing hundreds of load tests in his career. 5-10 years ago performance engineers used the “well known” load testing tools such as Load Runner. But things have changed as we have seen both a Shift-Left and a Shift-Right of performance engineering away from the classical performance and load testing teams. Tools became easier, automatable and cloud ready. In this session we discuss these changes that happened in the recent years, what it means for today’s engineering teams and also what might happen in 5-10 years from now. We also want to do a shout out to a performance clinic Daniel and Andi are doing on January 25th 2017 where they walk you through a modern cloud based pipeline using AWS CodePipeline, Jenkins, Apica and Dynatrace. Registration link can be found here: http://bit.ly/onlineperfclinic
Related Link:
ZebraTester Community: https://community.zebratester.com/
Mark Tomlinson, still a veteran and performance god, is enlightening us on his concept of Continuous Acceleration of Performance. Continuous Delivery is all about getting faster feedback from code changes as code gets deployed faster in smaller increments to the end user. One aspect that is often left out is feedback on performance metrics and behavior. In the “old days” performance feedback was given very late – either in the load testing phase at the end of the project lifecycle or even as late as when it hits production. That could be too late and it makes it hard to fix the root cause.
Listen to our conversation on how to accelerate performance related feedback loops without getting overwhelmed with too much data!
Mark Tomlinson, “a veteran” in Performance Engineering, discusses how DevOps is a big opportunity for performance engineering – but also a threat for many that have been in the business for a long time. The big question is: are “traditional performance engineers” using their Load Runners or SilkPerformers at the end of the project lifecycle ready to change? Ready to learn new tools? Ready to think about automating performance engineering into the delivery pipeline and doing that in collaboration with the rest of the engineering team? Ready to “Check your Ego at the door”?
Listen to our conversation where we also discuss how these roles have changed in organizations we recently interacted with.
In Part II with Finn Lorbeer (@finnlorbeer) from Thoughtworks we discuss some of the new approaches when implementing new software features. How can we build the right thing the right way for our end users?
Feature development should start with UX wireframes to get feedback from end users before writing a single line of code. Feature teams then need to define and implement feedback loops to understand how features operate and are used in production. We also discuss the power of A/B testing and canary releases as it allows teams to “experiment” on new ideas and thanks to close feedback loops will quickly learn on how end users are accepting it.
*Related Links:****
Process Automation and Continuous Delivery at OTTO.de
https://dev.otto.de/2015/11/24/process-automation-and-continuous-delivery-at-otto-de/
Are we only Test Manager?
http://www.lor.beer/are-we-only-test-manager/
Sind wir wirklich nur Testmanagerinnen?
https://dev.otto.de/2016/06/08/sind-wir-wirklich-nur-testmanagerinnen/
Finn Lorbeer (@finnlorbeer) is a quality enthusiast working for Thoughtworks Germany. I met Finn earlier this year at the German Testing Days where he presented the transformation story at Otto.de. He helped transform one of their 14 “line of business” teams by changing the way QA was seen by the organization. Instead of a WALL between Dev and Ops the teams started to work as a real DevOps team. Further architectural and organizational changes ultimately allowed them to increase deployment speed from 2-3 per week to up to 200 per week for the best performing teams.
*Related Links:****
Process Automation and Continuous Delivery at OTTO.de
https://dev.otto.de/2015/11/24/process-automation-and-continuous-delivery-at-otto-de/
Are we only Test Manager?
http://www.lor.beer/are-we-only-test-manager/
Sind wir wirklich nur Testmanagerinnen?
https://dev.otto.de/2016/06/08/sind-wir-wirklich-nur-testmanagerinnen/
Das Leben ist hasselhoff
http://giphy.com/search/david-hasselhoff
Gene Kim has been promoting a lot of the great DevOps Transformation stories from Unicorns (Innovators) but more so from "The Horses" (Early Adopters). The next DOES (DevOps Enterprise Summit) is just on its way helping him with his mission to increase DevOps adoption across the IT world.
In our 3 podcast sessions we discussed the success factors of DevOps adoption, the reasons that lead to resistance as well as how to best measure success and enforce feedback loops.
Thanks Gene for allowing us to be part of transforming our IT world.
Related Link:
Get a free digital 160 page DevOps Handbook Excerpt
http://itrevolution.com/handbook-excerpt?utm_source=PurePerformance&utm_medium=organic&utm_campaign=handbookexcerpt&utm_content=podcast
Gene Kim has been promoting a lot of the great DevOps Transformation stories from Unicorns (Innovators) but more so from "The Horses" (Early Adopters). The next DOES (DevOps Enterprise Summit) is just on its way helping him with his mission to increase DevOps adoption across the IT world.
In our 3 podcast sessions we discussed the success factors of DevOps adoption, the reasons that lead to resistance as well as how to best measure success and enforce feedback loops.
Thanks Gene for allowing us to be part of transforming our IT world.
Related Link:
Get a free digital 160 page DevOps Handbook Excerpt
http://itrevolution.com/handbook-excerpt?utm_source=PurePerformance&utm_medium=organic&utm_campaign=handbookexcerpt&utm_content=podcast
Gene Kim has been promoting a lot of the great DevOps Transformation stories from Unicorns (Innovators) but more so from "The Horses" (Early Adopters). The next DOES (DevOps Enterprise Summit) is just on its way helping him with his mission to increase DevOps adoption across the IT world.
In our 3 podcast sessions we discussed the success factors of DevOps adoption, the reasons that lead to resistance as well as how to best measure success and enforce feedback loops.
Thanks Gene for allowing us to be part of transforming our IT world.
Related Link:
Get a free digital 160 page DevOps Handbook Excerpt
http://itrevolution.com/handbook-excerpt?utm_source=PurePerformance&utm_medium=organic&utm_campaign=handbookexcerpt&utm_content=podcast
Guest Star: Anita Engleder - DevOps Manager at Dynatrace
In this second part of our podcast Anita gives us more insights into how new features actually get developed, how they measure their success and how to ensure that the pipeline keeps up with the ever increasing number of builds pushed through it.
We will learn more about the day-to-day life at Dynatrace engineering but especially about the “Lifecycle of a Feature, its feedback loop and what the stakeholders are doing to make it a success”
Related Links:
Dynatrace UFO
https://github.com/Dynatrace/ufo
Guest Star: Anita Engleder - DevOps Manager at Dynatrace
As a follow up to our podcast with Bernd Greifender, CTO and Found of Dynatrace, who talked about his 2012 mission statement to the engineering team: “We go from 6 months 2 weeks release cycles” we now have Anita Engleder, DevOps Lead at Dynatrace on the mic.
Anita has been part of that transformation team and in the first episode talks about what happened from 2012 until 2016 where the engineering team is now deploying a feature release every other week, makes 170 production deployment changes per day and can push a code change into production within an hour if necessary. She will give us insights in the processes, the tools but more importantly about the change that happened with the organization, the people and the culture. She will also tell us what she and her “DevOps” team actually contribute to the rest of the organization. Are they just another new silo? Or are they an enabler for engineering to push code faster through their pipeline?
We got to talk with Bernd Greifeneder, Founder and CTO of Dynatrace, who recently gave a talk on “From 0 to NoOps in 80 Days” explaining the “Digital Transformation Story of Dynatrace – the product as well as the company”
The transformation started in 2012 when Dynatrace used to deploy 2 major releases of its Dynatrace AppMon & UEM product to the market. The incubation of the startup Ruxit within Dynatrace allowed engineering, marketing and sales to come up with new ways and ideas that allow continuous innovating. In 2016 the incubated team was brought back to Dynatrace to accelerate the “Go To Market” of all the innovations. A new version of its Dynatrace SaaS and Managed offering is now released every 2 weeks with 170 production updates per day. Many aspects were also applied to all other product lines and engineering teams which boosted the output and raised quality of these enterprise products.
Are there new Web Performance Rules since Steve Souders started the WPO movement about 10 years ago? Do we still optimize on round trips or does HTTP/2 change the game? How do we deal with “mobile only” users we find in emerging geographies. How does Google itself optimize its search pages and what can we learn from it. In this session we really got to cover a lot of the presentation Pat Meenan (@patmeenan) did at Velocity this year.
Related Links:
Scaling frontend performance - Velocity 2016
**** https://www.youtube.com/watch?v=LdebARb8UJk
WEBPAGETEST
* https://www.webpagetest.org
Google AMP
* https://www.ampproject.org/
*** https://github.com/ampproject/amphtml
Pat Meenan (@patmeenan) is a veteran when it comes to Web Performance Optimization. Besides being the creator of WebPageTest.org he has also done a lot of work recently on the Google Chrome team to make the browser better and faster.
During his recent Velocity presentation on “Using Machine Learning to determine drivers for bounce and conversion” he presented some very controversial findings about what really impacts end user happiness. That it was not rendering time but rather DOM Load Time that correlates with conversion and bounce rates. In this session we dig a bit deeper into which metrics you can capture from your website and presented them to your business side as an argument for investing in faster websites. Find out which metric you really need to optimize in order to “move the needle”
Related Links:
Using machine learning to determine drivers of bounce and conversion - Velocity 2016
**** https://www.youtube.com/watch?v=TOsqP16jnDs
WEBPAGETEST
* https://www.webpagetest.org/
WPO-Foundation Github repository for machine learning
* https://github.com/WPO-Foundation/beacon-ml
Adam Auerbach (@Bugman31) has helped Capital One transform their development and testing practices into the Digital Delivery Age. Practicing ATDD and DevOps allows them to deploy high quality software continuously. One of their challenges has been the rather slow performance testing stage in their pipeline. Breaking up performance test into smaller units, using Docker to allow development to run concurrency and scalability tests early on, and automating these tests into their pipeline are some of the actions they have taken to level-up their performance engineering practices. Listen to this podcast to learn about how Capital One pushes code through the pipeline, what they have already achieved in their transformation and where the road is heading.
Related Links:
Hygea Delivery Pipeline Dashboard
https://github.com/capitalone/Hygieia
Capital One Labs
http://www.capitalonelabs.com/#welcome
* Capital One DevExchange
https://developer.capitalone.com/
Do you speak SQL? Do you know what an Execution Plan is? Are you aware that large amounts of unique queries will impact Database Server CPU and also efficiency of the Execution Plan and Data Cache? These are all learnings from this episode where Sonja Chevre (@SonjaChevre) and Harald Zeitlhofer (@HZeitlhofer) – both database experts at Dynatrace – pointed out database performance hotspots and optimizations that you many of us probably never heard about.
Watch the Online Performance Clinic -Database Diagnostics Use Cases with Dynatrace
https://www.youtube.com/watch?v=pEXfqzE-WQM
Are you still exporting load testing reports into Excel compare different runs manually? Matt Eisengruber – Guardian at Dynatrace – walks us through the life-changing transformation story of one of his former clients who used to spend an entire business day analyzing LoadRunner results.
Through automation, they managed to get her the results when she walks into the office in the morning – giving her more time to do “real” business analyst work instead of doing manual number crunching. Matt shares some insights into what exactly it is they did to automate Dynatrace Load Test comparison, how they created the reports and which metrics they ended up looking at.
Alois Reitbauer (@AloisReitbauer) guest hosts - Mike Jones ( http://bit.ly/mjlnk ) takes us on a journey how the team moved a monolithic application that was built by a remote team to a micro service architecture. Learn how the manage a couple of million lines of code with only 5 people while improving performance and availability. Mike also shares lessons learned on their journey and shares strategies on how to make the transition to micro services while having to keep the lights on for day-to-day business.
Scott Stocker (@sestocker), Solution Architect at Perficient, tells us the background of a recent load testing engagement on an ASP.NET App running on SiteCore. Turns out that even these apps on the popular Microsoft platform suffer from the same architectural and implementation patterns as we see everywhere else. Bypassing the caching layer through FastQuery resulted in excessive SQL, which caused the system to not scale, but crumble. Scott tells us how they identified this issue and what his approach as an architect is to proactively identify most common performance and scalability problems.
The initial idea of the Cloud has long become commodity – which is IaaS. Containers are the current hype but still require you to take care of correctly configuring your container that will run your code.
Mike Villiger (@mikevilliger) – a veteran and active member of the cloud community – explains why it is really PaaS that should be on top of your list. And why monitoring performance, architecture and resource consumption is more important than ever in order for your PaaS Adventure not to fail.
Related article:
http://www.it20.info/2016/03/the-incestuous-relations-among-containers-orchestration-tools/
In Part II, Richard Dominguez, Developer in Operations at PrepSportswear, is explaining the significance of understanding and dealing with bot and spider traffic on their eCommerce site. He explains why they route search bot traffic to dedicated servers, how to better serve good bots and how to block the bad ones. Most importantly: we learn about a lot of metrics he is providing for the DevOps but also the marketing teams to run a better online experience!
Have you ever wondered how to argue with a marketeer about not releasing a new feature or running this campaign? Or in the contrary: how can you show a marketeer that performance engineering and monitoring is as critical to the success of a campaign as the marketing campaign itself?
Richard Dominguez, Developer in Operations at PrepSportswear, is enlightening us about how his DevOps team is cooperating with marketing to have a better shared understanding between business and technical goals!
Microsoft is doing a good job in shielding the complexity of what is going on in the CLR from us. Until now Microsoft is taking care to optimize the Garbage Collector and tries to come up with good defaults when it comes to thread and connection pool sizes. The problem though is that even the best optimizations from Microsoft are not good enough if your application suffers from poor architectural decisions or simply bad coding.
Listen to this podcast to learn about the top problems you may suffer in your .NET Application. We have many examples and we discuss how you can do a quick sanity check on your own code to detect bad database access patterns, memory leaks, thread contentions or simply bad code that results in high CPU, synchronization or even crashes!
The Java Runtime has become so fast that it shouldn’t be the first one to blame when looking at performance problems. We agree: the runtime is great, JIT and garbage collection are amazing. But bad code on a fast runtime is still bad code. And it is not only your code but the 80-90% of code that you do not control such as Hibernate, Spring, App-Server specific implementations or the Java Core Libraries.
Listen to this podcast to learn about the top Java Performance Problems we have seen in the last months. Learn how to detect bad database access patterns, memory leaks, thread contentions and – well – simply bad code resulting in high CPU utilization, synchronization issues or even crashes!
How can you performance test an application when you get a new build with every code check-in? Is performance testing as we know it still relevant in a DevOps world or do we just monitor performance in production and fix things as we see problems come up?
Continuous Delivery talks about breaking an application into smaller components that can be tested in isolation but also deployed independently. Performance Testing is more relevant than ever in a world where we deploy more frequently – however – the approach of executing these tests has to change. Instead of executing hourly long performance test on the whole application we also need to break down these tests into smaller units. These tests need to be executed automatically with every build – providing fast feedback on whether a code change is potentially jeopardizing performance and scalability
Listen to this podcast to get some new insights and ideas on how to integrate your performance tests into your Continuous Delivery Process. We discuss tips&tricks we have seen from engineering teams that made the transition to a more “agile/devopsy” way to execute tests
Have you ever wondered what other people mean when they talk about a performance or load or stress test? What about a penetration test? There are many definitions floating around and things sometimes get confused.
Listen to this podcast and let us clarify for you by giving you our opinion on the different types of tests that are necessary when testing an application. In the end you can make up your own mind what best term to use.
Special Guest: Mark Tomlinson (@mtomlins) of PerfBytes
If you are running load tests it is not enough to just look at response time and throughput. As a performance engineer you also have to look at key components that impact performance: CPU, Memory, Network and Disk Utilization should be obvious. Connection Pools (Database, Web Service), Thread Pools and Message Queues have to be part of monitoring as well. On top of that you want to understand how your individual components that you test (frontend server, backend services, database, middleware, …) communicate with each other. You need to identify any communication bottlenecks because of too chatty components (how many calls between tiers) and to heavy weight conversations (bandwidth requirements).
Listen to this podcast and learn which metrics you should look at while running your load test. As performance engineer you should not only report that the app is slow under a certain load but also give recommendations on which components are to blame.