Interviews with data mesh practitioners, deep dives/how-tos, anti-patterns, panels, chats (not debates) with skeptics, "mesh musings", and so much more. Host Scott Hirleman (founder of the Data Mesh Learning Community) shares his learnings - and those of the broader data community - from over a year of deep diving into data mesh.
Each episode contains a BLUF - bottom line, up front - so you can quickly absorb a few key takeaways and also decide if an episode will be useful to you - nothing worse than listening for 20+ minutes before figuring out if a podcast episode is going to be interesting and/or incremental ;) Hoping to provide quality transcripts in the future - if you want to help, please reach out!
Data Mesh Radio is also looking for guests to share their experience with data mesh! Even if that experience is 'I am confused, let's chat about' some specific topic. Yes, that could be you! You can check out our guest and feedback FAQ, including how to submit your name to be a guest and how to submit feedback - including anonymously if you want - here: https://docs.google.com/document/d/1dDdb1mEhmcYqx3xYAvPuM1FZMuGiCszyY9x8X250KuQ/edit?usp=sharing
Data Mesh Radio is committed to diversity and inclusion. This includes in our guests and guest hosts. If you are part of a minoritized group, please see this as an open invitation to being a guest, so please hit the link above.
If you are looking for additional useful information on data mesh, we recommend the community resources from Data Mesh Learning. All are vendor independent. https://datameshlearning.com/community/ You should also follow Zhamak Dehghani (founder of the data mesh concept); she posts a lot of great things on LinkedIn and has a wonderful data mesh book through O'Reilly. Plus, she's just a nice person: https://www.linkedin.com/in/zhamak-dehghani/detail/recent-activity/shares/
Data Mesh Radio is provided as a free community resource by DataStax. If you need a database that is easy to scale - read: serverless - but also easy to develop for - many APIs including gRPC, REST, JSON, GraphQL, etc. all of which are OSS under the Stargate project - check out DataStax's AstraDB service :) Built on Apache Cassandra, AstraDB is very performant and oh yeah, is also multi-region/multi-cloud so you can focus on scaling your company, not your database. There's a free forever tier for poking around/home projects and you can also use code DAAP500 for a $500 free credit (apply under payment options): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Mirela's LinkedIn: https://www.linkedin.com/in/mirelanavodaru/
In this episode, Scott interviewed Mirela Navodaru, Enterprise and Solution Architect for Data, Analytics, and AI at Swisscom.
Some key takeaways/thoughts from Mirela's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Alyona's LinkedIn: https://www.linkedin.com/in/alyonagalyeva/
In this episode, Scott interviewed Alyona Galyeva, Principal Data Engineer at Thoughtworks. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Alyona's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Arne's LinkedIn: https://www.linkedin.com/in/arnelaponin/
Chris' LinkedIn: https://www.linkedin.com/in/ctford/
Foundations of Data Mesh O'Reilly Course: https://www.oreilly.com/videos/foundations-of-data/0636920971191/
Data Mesh Accelerate workshop article: https://martinfowler.com/articles/data-mesh-accelerate-workshop.html
In this episode, Scott interviewed Arne Lapõnin, Data Engineer and Chris Ford, Technology Director, both at Thoughtworks.
From here forward in this write-up, I am combining Chris and Arne's points of view rather than trying to specifically call out who said which part.
Some key takeaways/thoughts from Arne and Chris' point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Saba's LinkedIn: https://www.linkedin.com/in/sabaishaq/
Decide Data website: ttps://www.decidedata.com/
In this episode, Scott interviewed Saba Ishaq, CEO and Founder of her own data as a service consultancy, Decide Data, which also provides 3rd party DAaaS (Data Analytics as a Service) solutions.
Some key takeaways/thoughts from Saba's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Craziness of the overseas move (including a faulty office chair... long story) are to blame. Back to the normally scheduled one episode a week next week!
Episode list and links to all available episode transcripts here.
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Basten's LinkedIn: https://www.linkedin.com/in/basten-carmio-2585576/
In this episode, Scott interviewed Basten Carmio, Customer Delivery Architect of Data and Analytics at AWS Professional Services. To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Basten's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Olga's LinkedIn: https://www.linkedin.com/in/olga-maydanchik-23b3508/
Walter Shewhart - Father of Statistical Quality Control: https://en.wikipedia.org/wiki/Walter_A._Shewhart
William Edwards Deming - Father of Quality Improvement/Control: https://en.wikipedia.org/wiki/W._Edwards_Deming
Larry English - Information Quality Pioneer: https://www.cdomagazine.tech/opinion-analysis/article_da6de4b6-7127-11eb-970e-6bb1aee7a52f.html
Tom Redman - 'The Data Doc': https://www.linkedin.com/in/tomredman/
In this episode, Scott interviewed Olga Maydanchik, an Information Management Practitioner, Educator, and Evangelist.
Some key takeaways/thoughts from Olga's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Michael's LinkedIn: https://www.linkedin.com/in/mjtoland/
Marta's LinkedIn: https://www.linkedin.com/in/diazmarta/
Sadie's LinkedIn: https://www.linkedin.com/in/sadie-martin-06404125/
Sean's LinkedIn: https://www.linkedin.com/in/seangustafson/
The Magic of Platforms by Gregor Hohpe: https://platformengineering.org/talks-library/the-magic-of-platforms
Start with why -- how great leaders inspire action | Simon Sinek: https://www.youtube.com/watch?v=u4ZoJKF_VuA
In this episode, guest host Michael Toland Senior Product Manager at Pathfinder Product Labs/Testdouble and host of the upcoming Data Product Management in Action Podcast facilitated a discussion with Sadie Martin, Product Manager at Fivetran (guest of episode #64), Sean Gustafson, Director of Engineering - Data Platform at Delivery Hero (guest of episode #274), and Marta Diaz, Product Manager Data Platform at Adevinta Spain. As per usual, all guests were only reflecting their own views.
The topic for this panel was how to treat your data platform as a product. While many people in the data space are talking about data products, not nearly as many are treating the platform used for creating and managing those data products as a product itself. This is about moving beyond the IT services model for your data work. Platforms have life-cycles and need product management principles too! Also, in data mesh, it is crucial to understand that 'platform' can be plural, it doesn't have to be one monolithic platform, users don't care.
Scott note: As per usual, I share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Carol's LinkedIn: https://www.linkedin.com/in/carol-assis/
Eduardo's LinkedIn: https://www.linkedin.com/in/eduardosan/
Continuous Integration book: https://www.amazon.com/Continuous-Integration-Improving-Software-Reducing/dp/0321336380
Measure What Matters book: https://www.amazon.com/Measure-What-Matters-Google-Foundation/dp/0525536221
Inspired by Marty Cagan: https://www.amazon.com/INSPIRED-Create-Tech-Products-Customers/dp/1119387507
Empowered by Marty Cagan: https://www.amazon.com/EMPOWERED-Ordinary-Extraordinary-Products-Silicon/dp/111969129X
In this episode, Scott interviewed Carol Assis, Data Analyst/Data Product Manager and Eduardo Santos, Professor and Consultant, both at Thoughtworks. To be clear, they were only representing their own views on the episode.
From here forward in this write-up, I will be generally combining both Carol and Eduardo's views into one rather than trying to specifically call out who said which part.
Some key takeaways/thoughts from Eduardo and Carol's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Jessika's LinkedIn: https://www.linkedin.com/in/jmilhomem/
In this episode, Scott interviewed Jessika Milhomem, Analytics Engineering Manager and Global Fraud Data Squad Leader at Nubank. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Jessika's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Marisa's LinkedIn: https://www.linkedin.com/in/marisafish/
Karolina's LinkedIn: https://www.linkedin.com/in/karolinastosio/
Tina's LinkedIn: https://www.linkedin.com/in/christina-albrecht-69a6833a/
Kinda's LinkedIn: https://www.linkedin.com/in/kindamaarry/
In this episode, guest host Marisa Fish (guest of episode #115), Senior Technical Architect at Salesforce facilitated a discussion with Kinda El Maarry, PhD, Director of Data Governance and Business Intelligence at Prima (guest of episode #246), Tina Albrecht, Senior Director Transformation at Exxeta (guest of episode #228), and Karolina Stosio, Senior Project Manager of AI at Munich Re. As per usual, all guests were only reflecting their own views.
The topic for this panel was understanding and leveraging the data value chain. This is a complicated but crucial topic as so many companies struggle to understand the collection + storage, processing, and then specifically usage of data to drive value. There is way too much focus on the processing as if upstream of processing isn't a crucial aspect and as if value just happens by creating high-quality data.
A note from Marisa: Our panel is comprised of a group of data professionals who study business, architecture, artificial intelligence, and data because we want to know how (direct) data adds value to the development of goods and services within a business; and how (indirect) data enables that development. Most importantly, we want to help stakeholders better understand why data is critical to their organization's business administration strategy and is a keystone in their value chain.
Also, we lost Karolina for a bit there towards the end due to a spotty internet connection.
Scott note: As per usual, I share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Darren's LinkedIn: https://www.linkedin.com/in/darrenjwoodagileheadofproduct/
Darren's Big Data LDN Presentation: https://youtu.be/vUjoJrl_MEs?si=WzB0sBStVIAyqDJs
In this episode, Scott interviewed Darren Wood, Head of Data Product Strategy at UK media and broadcast company ITV. To be clear, he was only representing his own views on the episode.
Scott note: I use "coalition of the willing" to refer to those willing to participate early in your data mesh implementation. I wasn't aware of the historical context here, especially when it came to being used in war, e.g. the Iraq war of the early 2000s. I apologize for using a phrase like this.
Some key takeaways/thoughts from Darren's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Wendy's LinkedIn: https://www.linkedin.com/in/wendy-turner-williams-8b66039/
Culstrata website: https://www.culstrata-ai.com/
TheAssociation.AI website: https://www.theassociation.ai/
In this episode, Scott interviewed Wendy Turner-Williams, Managing Partner at both TheAssociation.AI and Culstrata and the former CDO of Tableau.
TheAssociation.AI is "a global nonprofit business organization …focused on bridging the disciplines of AI, data, ethics, privacy, robotics, and security." It is focusing on things like networking and knowledge sharing to drive towards better outcomes including ethical AI.
Some key takeaways/thoughts from Wendy's point of view:
Wendy started out with her perspective that in some respects, data has become "a four letter word". There's so much data and everyone is trying to use it but everyone also feels inundated. Instead of being data-driven, we are data-flooded or data-dragged. And there is a major lack of tying the data work to the actual strategy and execution. Where do we need data to support our decisions? We need a strategy to get that data in place.
Relatedly, Wendy sees the major breakdown between strategy and goals when it comes to data. There may be a strategy to grow a product but how much growth is feasible, a good target? Why is that growth feasible? What does the data say about growing that product, e.g. the market dynamics and your positioning in the market? So when goals are set, it is a 'finger in the air' type guess as to how much it could grow or worse, simply how much leaders want a product to grow. And then what data do we need in place to enable the team managing that product to actually be able to grow it that much? How do we enable them to make smart tactical decisions?
Basically, it's a lot of things looping back on each other. We need data to set good strategic decisions. But we need a strategy to set up our ability to capture and analyze that data. We need data to make better tactical decisions. But most companies lack the ability to make good tactical decisions to get the necessary data in place. It's top-down driven but far too often the ones who understand what needs to be done for and with data are too far down in the organization and there's a communication gap. Thus, there is a significant lack of being data-driven. We have to admit the problem first. How to fix all of that is another fun process 😅
For Wendy, far too many organizations have data as an afterthought. And that leads to subpar understanding of what's actually happening with the business and lacking the information to fine-tune their strategic decisions. There isn't a strong strategic connection flowing from the business strategy to the data work. Execs aren't spending the time to really follow-through on exactly what the tactics should be to get the right data in place. Scott note: we literally have a panel on doing that, tying the data work to the business strategy and vice versa 😎 episode #251
In many parts of many organizations, e.g. Marketing, Wendy sees there being very competent people who just don't really understand how to do data well. They need a great partner. Should the marketing leader be focused on what data sources they need and why? Or should we be able to translate their needs into the work? But first, we need to actually be able to partner and they need to understand their needs. Data people can supercharge their efforts but the business partners need to lean in. Scott note: in data mesh, part of the role is enabling them to get better. We need people to up their fluency but doing data mesh or not, everyone starts somewhere and we need to help them level up.
On the flip side, Wendy also sees how often data people are stuck focusing on the data work instead of the business aspects of what that data work is tied to. Without the business context, all you are doing is pushing 1s and 0s. What do business partners need and why?! There needs to be the ability and the courage to just hammer out the understanding differences or the problems will persist.
Wendy also gave some specific examples of too many cooks in the kitchen relative to certain measurements. Instead of there being one official perspective or measurement for something like usage of a cloud product, in a previous role there were many measurements across engineering, finance, marketing, sales, etc. And every single one was different because they all used slightly different methodologies and even sources. So when they tried to look at success of the product, everything told a different story. And when they tried to have a simple bill the customers could understand, it was just not possible. While single source of truth is a complicated and overloaded term, one official source of truth for a question is something you should be able to rally around.
A big problem in many organizations is people are only focused on their own job and lose sight of the bigger picture and especially how they play into that bigger picture of the organization's success according to Wendy. Even if your role isn't directly improving the customer experience, your work can have a positive impact on that if you drive towards that goal. Sometimes, politics around data also gets in the way of collaboration across teams and lines of business.
Wendy talked about another persistent problem in data: the service model. If your data teams are only focused on supporting other teams, you can lose sight of your big picture impact as well as the impact of bad data. She believes data teams need to spend more time creating their own data around their impact and also quantifying the costs of data issues. What are the actual impacts to the organization? And do execs outside the data team understand data well enough to understand those impacts? If they don't understand data, can they even trust it enough to rely on it?
Circling back to the bigger picture, Wendy believes that teams can drive significant process improvements if they just understand the impact of their work - especially through data - upstream and downstream. What do they actually need from others? Who is consuming their data and why? What impact will changes have? How are communications set up to prevent issues and create strong understanding and trust? And then of course, try to automate as much as possible to lower the burden on everyone involved in the data flowing around :) As part of that, please just hire good product managers 😅
Wendy said, "You will never be as successful as you can be as a data organization if you're not able to influence your IT partners, your product teams, your business teams." Data is a team sport, data is about making the organization better. You need others to play with you or it won't work.
When thinking about actual business transformation around data, Wendy said, "There is no transformation without automation." Historically, doing data work has required a lot of effort. The business side just wants to leverage the data, help them automate as much as possible. Otherwise many - most? - business partners won't want to engage with the data and leverage data to improve their work - it's too much effort. Also, removing the friction from data work helps people identify the friction in general business processes. So automating the data work allows them to more easily identify and then address "business choke-points".
For Wendy, too many aspects of data work are treated as wholly separate disciplines instead of treating it as all part of one whole. Security, privacy, compliance/regulatory, performance, etc. We have to shift it left but also stop trying to treat them as discrete challenges to overcome instead of interoperating aspects of a working, scalable solution. Think data by design 😎 That's why she created TheAssociation.AI, "a global nonprofit business organization …focused on bridging the disciplines of AI, data, ethics, privacy, robotics, and security." She said, "there is no security, there is no privacy, there is no ethics, there is no AI without data," but we also need organizations to actually implement their policies into their data and data work. There isn't going to be ethical AI without someone leading that charge and TheAssociation.AI is looking to push that effort forward.
In wrapping up, Wendy circled back to the start. What is the point of doing data work. She said, "What's the point of being focused on the data if you don't understand the business that the data is supposed to be used for?" Being a data leader, especially the CDAO, is VERY tough because you often don't own much of the infrastructure if at all and have to do your work essentially via influence. But if you build the right relationships and understanding of the business, you can still have a major impact and drive significant value for your organization.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Amritha's LinkedIn: https://www.linkedin.com/in/amritha-arun-babu-a2273729/
In this episode, Scott interviewed Amritha Arun Babu Mysore, Manager of Technical Product Management in ML at Amazon. To be clear, she was only representing only own views on the episode.
In this episode, we use the phrase 'data product management' to mean 'product management around data' rather than specific to product management for data products. It can apply to data products but also something like an ML model or pipeline which will be called 'data elements' in this write-up.
Some key takeaways/thoughts from Amritha's point of view:
Amritha started the conversation on some key differences between software product management and product management around data - whether specific to 'data products' or not. One similarity is the focus on solving a particular user problem but in data, you might have to build something to address multiple users' problems. A much bigger difference is that in data, you often don't own the entire process as you might be reliant on others to source your data. In software, you are generally building the data sourcing because you own the interaction creating the data. How the data is stored and collected throughout the upstream process impacts what you can do.
The user problem, the business value, and the user journey are some key guides to doing data product management well for Amritha. Keep coming back to those as you build out your solution. Focus on understanding what the user really needs and work backwards to the sources. And then of course focus on making sure you are actually addressing user needs when you deploy the solution. There are many reasons a data element may not be performing up to expectations so be prepared to deep dive; is there a problem with what you've built, what's feeding your data element - maybe sources have changed or there is a quality issue -, or is it just not performing to expectations because the hypothesis was wrong?
Amritha dug a bit more into some challenges specific to product management in machine learning and AI. While data scientists want clean data, when possible they want to even be part of the process of selecting the cleaning methodologies - even that can impact the data enough to change outcomes. So really start from the process of bringing them in as a stakeholder as soon as you can and don't throw data over the wall at them. And if you already have something developed, share your methodologies and help them figure out if it's the right fit for them or if something new needs to be developed. Again, we want reuse but we also want solutions that address their problems. Always a hard set of needles to thread.
"As a product manager, it's just part of the job that you have to work backwards from a customer pain point." Amritha questions if you are even building a product if you don't have a customer. What is the business value of the work? For a product manager of product without a customer, are you focused on your own thoughts and biases rather than the needs of consumers? "So the point here is that at any given point, you have to be cognizant of who are you building this for, why, and what that is the primary customer. And the secondary is: who else if I build this, what are the impacts it will have on my secondary customers, or other downstream or interacting applications?"
Amritha talked about one crucial rule in product management: prioritize. There are many use cases you _could_ solve but are they actually worth the effort? Think about what will impact your organization the most. Don't try to solve every use case and don't try to make products that can serve every potential customer - focus on delivering value. Scott note: this can be a slippery slope in data mesh. You want to take on use cases you actually can tackle when you are learning. Don't only go for the biggest value but also tackle problems where the juice is worth the squeeze, where the outcome is worth the effort.
In product management, Amritha believes it's absolutely crucial to understand the art and the science. The science is more about is this product specifically meeting the needs it was designed for. Basically, measuring the level of success and determining if that's good enough or especially is it _still_ good enough. But even that last bit can be a bit of art. The real art is all about communication and building relationships. If you build the world's objectively best product but no one trusts it or understands it enough to use it, it's not a valuable product. You must build strong relationships and have the tough conversations with stakeholders, earning their trust, to align on what needs to get built and why as well as when a product isn't meeting expectations. Establish regular lines of communication so it's not that the only time you talk to your customers, it's bad news or big changes. Continue to extract information from them to drive to business value.
When it comes back to the science, that's when Amritha believes you should dig into the why something isn't meeting expectations from the technical perspective :) And have some patience around that. Sometimes it's a blip on the radar, not anything more.
When figuring out what products/data elements you might want to build in a specific area, Amritha recommends digging into the potential workflows and user journeys. Start to really think about what you think could exist and why. But, instead of trying to ideate only yourself, go and talk to people and listen for their pain and points of friction. They may not even realize they have pain but you can find the challenges that people will want to address. Again, work backwards from the user journeys to discover what products you should build 😅
Amritha talked about how to make maximize the chance that what you're building will be used/valuable. A lot of it is simply digging in deep with potential customers in the ideation phase to make sure this will actually drive value. There are ways to do that but a lot of it is simply spending the time to really understand the likely impact of what you're building. As Alla Hale said in episode #122, "What would having this unlock for you?" Also, ask, "what if we don't do this, what is the impact of not doing this?" And make sure to get validation as you're building. It might be the value hypothesis was wrong or that you're building something that is the wrong or suboptimal way to address the challenge/opportunity. You can save yourself a lot of headaches and rework. It's all about that collaboration to drive to value.
In wrapping up, Amritha talked about how changes, especially in data, are inevitable. Make sure to communicate with consumers so they have realistic expectations. Sometimes those are proactive changes but often, you don't have that much control over changes, especially coming from upstream in data. Look to build in a way that can adapt and leverage a "loosely dependent architecture".
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Nailya's LinkedIn: https://www.linkedin.com/in/nailya-sabirzyanova-5b724310b/
In this episode, Scott interviewed Nailya Sabirzyanova, Digitalization Manager at DHL and a PhD Candidate around data architecture and data driven transformation. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Nailya's point of view:
Nailya started the conversation on the need for application, business, and data architectures to all be aligned. She gave some history about how application and business architecture were brought into alignment when we started with microservices and digital transformation. But now we have to add data into the mix which makes things even more difficult. We need both the application and data architectures to be designed to specifically support the business goals.
When it comes to actually transforming your data architecture, Nailya believes the transformation should be led from the business side where possible. At the very least, the business side should be involved. They understand the business needs best so they can help direct the transformation to serve those needs. Great data work that doesn't support the business needs is often just a well-designed money pit 😅 You obviously need the strong data expertise but a transformation led exclusively by the data team is far less likely to align to business goals and priorities.
In a successful past digital and data transformation, Nailya used two simple questions to the key business stakeholders: What should this transformation enable? And how should we enable it? It gave the business stakeholders a chance to fully lay out their pain points as well as some ideas how to address them. That way, the team had a very broad perspective and could come back to each of the stakeholders with solutions that worked for them somewhat tailored to their needs and thoughts. You need repeatable patterns/approaches but you also need people to feel seen and heard in order to drive buy-in where you address their specific pains and ideas.
When asked about adapting data mesh to an organization's specific challenges, Nailya pointed to how every culture is so different and you need to take into account how people internally exchange information and work together as you design how you want to go forward. Overly decentralizing - so not doing any federation - or decentralizing too quickly won't work well. You have to find your balance between centralization and decentralization throughout your journey.
One interesting buy-in point Nailya mentioned was cost control over data work. Because teams have traditionally been charged for data resources and work by central data teams, they were not as involved in managing costs. Data mesh empowers business teams with tools to control cost-effectiveness of their data, and thus they can identify easier which data is valuable for their business and requires investment, and which data or data processing operations are redundant, and thus, a source of savings. They can see it as a chance to do things better and align better on what work is worth doing - the central data team might have done work that the domain doesn't see as valuable when really considering it more deeply. At first, it might be only for their internal-facing to the domain data work before we can get them bought in that they are now responsible for also sharing their data with other domains.
Nailya talked about her experience with a large-scale data mesh implementation. They focused on first enabling teams to own their own data. So again, giving them the chance to gain transparency to their most valuable data as well as define, align, and prioritize their data initiatives. Then, they started to work to incentivize and better enable them to share their data with the rest of the organization identifying new use cases and data customers. This may delay the biggest benefits of data mesh - high-quality, reusable data across domain boundaries - but it does mean that teams aren't struggling to own their data at the same time as learning to share it with others; this also helps with the incentivization challenge as they can take advantage of their data first for themselves before being asked to focus on sharing it with other domains.
Data governance in data mesh will - unsurprisingly - be hard for every organization in Nailya's view. If there are domains that already know how to handle their data well, work to enable them to better share their information but also don't try to push them towards central ways of working. If they can safely secure and share their data in a way the rest of the organization can consume it, don't get in their way. But you should also look to create frameworks and standards for those that aren't as mature to help guide them along. Scott note: this is an adjustment for those that are already somewhat decentralized with their data work. Again, adjust for your circumstances!
Nailya also recommends central committees to ensure teams are meeting some degree of conceptual consistency and also technical/architectural consistency. That way, you can really find your scalability patterns and best practices. Scott note: if you decentralize/federate all at once, this might be the best bet. But many - most? - are going domain by domain so this function is already embedded in an enabling team.
In her experience, Nailya believes that aligning your data and digital transformation is very important to be able to succeed at data mesh. Again, you need to align the application and data architectures with the business architecture. Really take stock - is your data transformation part of your overall digital transformation, are they at the same level and should be partnered, etc.?
Nailya again circled back to engaging with your business partners/stakeholders to have them help design your transformation efforts. They will lean in from feeling seen and heard but also, they know the most critical pain points best and can help direct where you should focus first.
The conversation finished up around getting and maintaining a budget around your data mesh implementation. For Nailya, this is typically far more political than many might expect. But if you have the proper top-down support and management attention/buy-in, you should at least be able to get going. You have to show value along the way but that should be part of the prerequisite to start: what value are you trying to capture and how will you measure it?
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Jen's LinkedIn: https://www.linkedin.com/in/jentedrow/
Martina's LinkedIn: https://www.linkedin.com/in/martina-ivanicova/
Xavier's LinkedIn: https://www.linkedin.com/in/xgumara/
Xavier's blog post on data as a product versus data products: https://towardsdatascience.com/data-as-a-product-vs-data-products-what-are-the-differences-b43ddbb0f123
Results of Jen's survey 'The State of Data as a Product in the Real World' (NOT info-gated 😎👍): https://pathfinderproduct.com/wp-content/uploads/2023/12/2023-State-of-DaaP-Real-World-Study.pdf?mtm_campaign=daap-study&mtm_source=pp-blog&mtm_content=pdf-daap-study
In this episode, guest host Jen Tedrow, Jen Tedrow, Director, Product Management at Pathfinder Product, a Test Double Operation (guest of episode #98) facilitated a discussion with Martina Ivaničová, Data Engineering Manager and Tech Ambassador at Kiwi.com (guest of episode #112), and Xavier Gumara Rigol, Data Engineering Manager at Oda (guest of episode #40). As per usual, all guests were only reflecting their own views.
The topic for this panel was data as a product generally and especially how can we actually apply it to data in the real world. This is Scott's #1 most important aspect to get when it comes to doing data - especially data mesh - well. It's the holistic practice of applying product management approaches to data. It ends up shaping all the other data mesh principles and is a much broader topic than data mesh is in his view. But it can also be quite simple in concept when you really boil it down, it just takes patience and focus.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Tom's LinkedIn: https://www.linkedin.com/in/tomdw/
Data Mesh Belgium: https://www.meetup.com/data-mesh-belgium/
Video by Tom: 'Platform Building for Data Mesh - Show me how it is done!': https://www.youtube.com/watch?v=wG2g67RHYyo
ACA Group Data Mesh Landing Page: https://acagroup.be/en/services/data-mesh/
In this episode, Scott interviewed Tom De Wolf, Senior Architect and Innovation Lead at ACA Group and Host of the Data Mesh Belgium Meetup.
Some key takeaways/thoughts from Tom's point of view:
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
May's LinkedIn: https://www.linkedin.com/in/may-xu-sydney/
In this episode, Scott interviewed May Xu, Head of Technology, APAC Digital Engineering at Thoughtworks. To be clear, she was only representing her own views on the episode.
We will use the terms GenAI and LLMs to mean Generative AI and Large-Language Models in this write-up rather than use the entire phrase each time :)
Some key takeaways/thoughts from May's point of view:
May started with three general approaches organizations are taking to generative AI (GenAI): 1) building their own LLMs from scratch, 2) fine tune specific, pre-trained existing LLMs, or 3) leverage pre-trained LLMs as is. Many organizations may want to do the first but it is prohibitively expensive to train your own LLMs from scratch just for the compute and you also need (very expensive) people with very specific expertise to do so. Tuning pre-trained models will likely become the standard approach for many organizations. However, being able to leverage LLMs on internal data in general requires "existing good quality data and solid data architecture."
When considering training a model from scratch, May also pointed to time as an issue. Typically, it takes at least three months to properly train an LLM from scratch. Parallel training is helpful but you need to fine-tune results and retrain so you can't just throw compute at it and make the process that much faster. So again, you need high quality data - and you need a LOT of it - plus a fair amount of time plus a ton of money. Once you are in production, it also takes a lot of money and effort to keep them running and tuned properly 😅 Luckily, according to some surveys Thoughtworks did, most organizations recognize training LLMs from scratch isn't the right call for them just yet.
May is seeing a trend of people moving away from the 'bigger is better' mentality. More people are starting to explore more targeted and specialized models that have fewer parameters. And often, for specific tasks, they perform better than the first L in LLMs. So we may see a trend towards more and more targeted LLMs/models. Scott note: Madhav Srinath really leaned into this in his episode, #264.
Humanity in general has benefited greatly from machines through predictability and reliability according to May. Essentially, if they are made well, you essentially know what you should/will get from machines. But GenAI is designed specifically to act like humans and humans are not predictable and often not that reliable. So people have to get used to interacting with machines that may give wrong answers and are designed - in a way - to do so 😅 We can't expect predictability and reliability from GenAI.
Relatedly, when thinking about where is GenAI the right choice versus like traditional machine learning/AI, May believes you really have to dig into the tradeoffs. If you really understand the problem set and what you are trying to accomplish, traditional ML/AI is probably the better approach for you. You need to really understand where the strengths of GenAI will play and feed it the data/information it needs to succeed, otherwise you'll be asking an uniformed and unpredictable entity to solve your most pressing business problems. That's probably not going to go well…
May talked about going back to the basics of problem solving when it comes to Generative AI: what problem are you trying to solve instead of what way are you trying to solve a problem and then finding your way back to the problem. It can sound obvious but really, many are in such a rush to leverage these tools, it's crucial to stop and consider. Start with the problem before the solution 😅
GenAI may also surface a number of internal business challenges that aren't spoken about or people have essentially given up on tackling previously according to May. We have a new tool in the toolbox so people want to see if it will be useful to tackle something they haven't been able to address well previously. Lean into GenAI as a conversational lubricant. GenAI may not be the right tool for every one of these challenges but it means there is more internal conversation and sharing :)
From what May is seeing, many to most organizations are still in the early experimenting and PoC phase with Generative AI. They are trying to figure out what opportunities GenAI brings and also what risks. Despite the hype, people are taking their time but they aren't as focused on initial return on investment, more to validate if they can actually leverage GenAI to create value. Also, there is strong trend towards domain-specific LLMs rather than general purpose ones, e.g. financial sector or media specific models.
May finished on the idea that data mesh and other data management paradigms are crucial to doing something like GenAI right. There is still a strong need for quality data that is accessible, interoperable, privacy-aware, secured, etc. to be able to leverage GenAI well.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
IRM UK Conference, March 11-14: https://irmuk.co.uk/dgmdm-2024-2-2/ use code DM10 for a 10% off discount!
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Ole's LinkedIn: https://www.linkedin.com/in/ole-olesen-bagneux-2b73449a/
Piethein's LinkedIn: https://www.linkedin.com/in/pietheinstrengholt/
Samia's LinkedIn: https://www.linkedin.com/in/samia-rahman-b7b65216/
Liz's LinkedIn: https://www.linkedin.com/in/lizhendersondata/
Ole's book The Enterprise Data Catalog: https://www.oreilly.com/library/view/the-enterprise-data/9781492098706/
Piethein's book Data Management at Scale (2nd Edition): https://www.oreilly.com/library/view/data-management-at/9781098138851/
Liz's blog: https://lizhendersondata.wordpress.com/
In this episode, guest host Ole Olesen-Bagneux, Chief Evangelist at Zeenea (guest of episode #82) facilitated a discussion with Piethein Strengholt, CDO at Microsoft Netherlands (guest of episode #20), Liz Henderson AKA The Data Queen, a board advisor, non-executive director, and mentor in digital and data at Capgemini (guest of episode #106), and Samia Rahman, Director of Enterprise Data Strategy, Architecture, and Governance at SeaGen/Pfizer (guest of episode #67). As per usual, all guests were only reflecting their own views.
The topic for this panel was modernizing master data management (MDM) and applying that to data mesh. It's a very challenging topic to cover because even people's general definition of MDM can be pretty different and there is a question between simply mastering data versus trying to globally compared to locally manage master data. It's a very tricky topic in data mesh. I sometimes use the term 'mastered data' because I think it is far more applicable in that situation than 'master data' - audio goes through a mastering process, so must the master data to actually reach a certain quality level. 'Master data' is more the core linking data. But even that is still just one person's definition.
Scott note: As per usual, I share my takeaways rather than trying to reflect the nuance of the panelists' views individually. Also, there was a bit of a misstep around intros if that gets a bit lost in there.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
IRM UK Conference, March 11-14: https://irmuk.co.uk/dgmdm-2024-2-2/ use code DM10 for a 10% off discount!
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
IRM UK Conference, March 11-14: https://irmuk.co.uk/dgmdm-2024-2-2/ use code DM10 for a 10% off discount!
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Sue's LinkedIn: https://www.linkedin.com/in/suegeuens/
In this episode, Scott interviewed Sue Geuens, Director of Data Governance and Product Data at Elsevier. To be clear, she was only representing her own views on the episode.
We use the phrase MDM to mean master data management throughout the episode.
Some key takeaways/thoughts from Sue's point of view:
Sue started the conversation as other data governance experts have - the word governance strikes at least discomfort if not fear into the hearts of many of our colleagues. We need to expect that discomfort and be active in dispelling the myths around data governance as it really is about achieving better outcomes for all. But that means more carrots than sticks, which can be a tall task when it comes to things like regulatory compliance. Basically, it's not easy 😅
Another aspect Sue pointed to is that many - most? - data people really like to talk data. So, instead of talking to outcomes, they talk about the data work, and data work for the sake of data work has kind of been one of the big historical challenges of data - instead we need to focus on the value that comes from the data work. If your business partners are already uncomfortable simply by the phrase data governance, not leaning into their value from the data work and target outcomes is likely to lose them even further. Start the conversation with what they might want from you, not what you might want from them.
Sue specifically said she starts partnering with people by focusing on those target outcomes and how might she be helpful to them. Especially, what are their expectations of her? By trying to walk in their shoes, she can come to better conclusions and find working solutions. It's about getting them to lean in. Scott note: and then she can trap them! In a virtuous bi-directional value trap of course…
Relatedly, prioritization in data governance is key in Sue's view. What are the problems that really matter? While the "who shouts loudest" test may not point to the most valuable problems, it often points to the problems people value most and thus you can find willing partners. Trying to enforce others to care about their data is a hard road but if people are ready for your help, you can make a huge difference and they are willing to lean in. Those are also likely to be your biggest advocates once you help them, gaining your governance efforts more momentum by leveraging champions.
There are many reasons why Sue believes people are skeptical of master data management (MDM). Historically, there were two big reasons MDM projects failed. The first is not really focusing on integrating MDM into the data so not having the governance, quality, and metadata embedded into the data and processes. The second is the drive towards perfection. Instead of focusing on what was good enough, there was this focus on the 'golden record'. That led to inflexibility, poor scaling, high costs, etc. Good data work isn't about being perfect, it's about being good enough.
Sue circled back to her focus on working with people. Good governance isn't about perfect data, it's about getting people to care about the quality of the data. That means working to get them to understand what is good enough and why should they care. It's not all just empathy - there needs to be some oversight and making it part of their job - but with humans in the loop, your data quality will be much better if you get people to care about who else uses their data and why.
When it comes to actually getting people to understand data governance work - whether MDM or anything else - Sue recommends personalizing your communication. While that may not scale perfectly, again, find your key stakeholders and partners. Stories about data work in a vacuum just don't resonate - Scott note: is there a physics/sound joke in there as there is no air for sound to resonate in a vacuum…? 😅 - Getting people to understand that the work has a purpose and it really is useful to specifically them is crucial. Don't talk to the 1s and 0s of data!
When it comes to specifically data ownership, Sue has seen just how scary that ownership word can be. It's not an easy task but we need to find ways to instill people with the excitement around ownership without the fear. Again, easier said than done but it's about getting things to the right place not about doing something right now. It will take time but it's better to do it right.
If you don’t take care with implementing data mesh well, Sue believes it will be a far bigger mess than if you didn't try data mesh at all. (Scott note: strong agree) You need to focus again back to what are you trying to accomplish and what data needs to be put in place to do that. MDM in data mesh should be about "ensuring that you get the right data for the right purpose at the right time for the right person."
In wrapping up, Sue emphasized the need for personalizing your communication around getting people to do data work with care and prioritize it. You need to be able to speak to them in their language and get them excited about the impact of the work.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key points:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Frederik's LinkedIn: https://www.linkedin.com/in/frederikgnielsen/
In this episode, Scott interviewed Frederik Nielsen, Engineering Manager at Pandora (the jewelry one, not the music one 😅).
Some key takeaways/thoughts from Frederik's point of view:
Frederik started with a bit about how their initial data mesh journey started - and it wasn't great 😅 It was led by management consultants and was focused on real-time data with a very tangible use case. However, two things came from it: 1) a better understanding of what data mesh should actually be used for and 2) buy-in around a very specific use case at the highest levels. So while there was a misinterpretation of data mesh and the use case wasn't the best fit, there was still excitement about the term - and somewhat the actual meaning broadly internally. Making it tangible got people to see the potential benefits.
Cost transparency has been a major driver for data mesh internally according to Frederik. Because the costs in a large monolithic stack are very opaque, decomposing the architecture has led to a far better understanding of the cost of individual pieces of work. Because inflation concerns were a big factor for retail in 2023, there was a bigger focus on cost reductions. Being able to give teams the freedom to take different approaches but making them responsible for the costs has led to better cost efficiency - teams can choose more costly methods but those decisions are more exposed. Also, because you have much finer-grained control, there are far more levers to pull when it comes to cost savings, e.g. shutting off test and dev environments at night or scaling up and down dynamically.
Frederik talked about a common pattern when moving to data mesh: some teams are more data mature than others. There will be plenty of teams that need help when it comes to data mesh, especially building good data products. They are considering creating a sort of golden path or easy button approach for those teams that aren't as mature as well to make things relatively pre-configured instead of having to make many complex decisions.
When driving buy-in at the wider level for data mesh, Frederik talked about pitching data mesh as an entire organization transformation versus pitching use case by use case. He believes it's probably better to focus on the use cases but it can be hard to focus on the complete picture of everything you need when you are also focused on specific use cases. It's always a balance between what is needed only for the use case and what is good for the overall company approach to data.
For Frederik, there are two big company strategic priorities: personalization and omni-channel experience (experience across in-store and online). So much of what they have been focusing on is finding use cases that tie into at least one of the priorities because then there will be executive support. Constantly tying the data work back to what people care about shows an understanding of the business instead of doing data work for the sake of data work. However, these are very big challenges across many domains and teams. So making sure to do things in a scalable way and finding the right balance between data products with still high interoperability is crucial.
When discussing bottlenecks, Frederik talked about how the measure for the centralized data team becoming a bottleneck was when the time between a data request and the actually delivery was expanding. The backlog was ballooning even though the data team was quite productive. Many people will feel the pain of the increasing time to delivery, leverage that while still showing a productive team. If you are executing well but aren't succeeding, you need a new strategy.
Frederik talked about the fact your data technology and architecture decisions will incentivize certain behaviors. A monolithic platform incentivizes monolithic ownership and handing off work, responsibilities, etc. When they introduced Kafka, it enabled them to push ownership upstream to data producers because the new technology allowed data producers to more easily own their data. It's of course difficult to incentivize your desired behaviors but always think about what you want to happen and try to make that the easy/happy path.
When it comes to ownership of data, Frederik thinks maturity really matters. When you want to go down the path of data mesh, trying to get every domain to really be advanced with data is just not that realistic. Some teams just don't see data as their focus so if they won't leverage much data for analytical or ML/AI use cases, they are less likely to want to own their data. And less capable quite frankly.
Circling back to tangible use cases, Frederik talked about one really key use case that saw a really big uptake that they couldn't really accomplish before going the data mesh route. Being able to tie something to actual impact, that really helped people get more interested. Whether that is a business capability or directly impacting a business metric. Similarly, when trying to find new use cases, the team did a lot of user journey mapping. The data for that user journey lives in many systems so you need lots of teams participating to make the data available but it can have a big impact on business. Many companies probably can't do something that complex in their existing architecture. You can use the inability to do amazing things in your existing architecture as a potential selling point.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Quick Summary Points
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Mandeep's LinkedIn: https://www.linkedin.com/in/kaurmandeep80/
In this episode, Scott interviewed Mandeep Kaur, Enterprise Information Architect at Nordea Asset Management. To be clear, she was only representing her own views on the episode.
Nordea has been on their data mesh journey for a while and Mandeep has been trying to figure out best practices for the hundreds - thousands - of micro decisions in a journey. So how do we get comfortable with making so many calls?
Some key takeaways/thoughts from Mandeep's point of view:
Mandeep started by discussing one of the key challenges in talking about data mesh: there are so many areas to cover that we often discuss things abstractly. Those abstractions are based on a significant amount of research, discussions, and related work that create a mental model. When we communicate those abstractions, it's hard to communicate the mental model as well. The listeners just aren't as deep into it so much of it goes over their (our) heads. So, we need to get far more specific with anecdotes and examples. We also can't forget the value of 1:1 conversations to drive to deeper understanding. That might not always be the most scalable but it is the best way to prevent misunderstandings. Basically, communication around data mesh is hard! Go talk to people. Scott note: Data Mesh Understanding exists for this reason…
When looking at how to get specific internally with data mesh communication, Mandeep is always on the lookout for her ambassadors or champions. Within a domain, they have strong domain knowledge to connect what you are trying to achieve with data mesh to what the domain is specifically focused on. And they can obviously communicate well in the language of the domain. Connecting the changes data mesh brings to real world problems helps people understand the what and the why.
There is a lot of risk of analysis paralysis in any data mesh implementation according to Mandeep. There are hundreds of 'micro decisions' but if you focus on the core aspects of what you're trying to do, that should guide you to the ones that matter the most. A bit of don't sweat the small stuff. Always come back to the value proposition because you can change things as you learn more. That's not to say be sloppy or careless, there are important aspects like using the right architecture, having strong ownership/accountability, product thinking, etc. But data mesh is as much about learning to get it right as getting it right. And always return to your trade-offs. What aren't you willing to trade-off and why? Once you answer that, more and more solutions become tenable and you can weigh the pros and cons.
Mandeep started to dig into the crucial first question to a data mesh implementation: what is the value you hope to get out of it? And there are different answers for each organization. Those answers will start to inform where you should focus and when in your mesh journey. That will help you set your plan because "a target without a plan is just a dream". And when you form your data mesh plan, think about what you have to adapt to your organization and why. This is not a copy/paste approach! You almost certainly will have competency gaps so how do you plan to fill those gaps and make progress while doing that? Or do you have to fill those gaps before starting the journey because they are journey blockers? Really consider the journey, not only the target outcome. Relatedly, set some milestones for your journey to help you measure your progress and celebrate the progress you've made. They might not be the best success measures once you're further along in your journey but that's okay, you can adjust. That's product thinking.
Mandeep wanted to stress three quick points: "1) don't overthink it. 2) bring value out as soon as possible. 3) evolution before completion."
Many business users still see technology as a threat rather than an enabler in Mandeep's view. They aren't even thinking about data yet, they are still on just tech 😅So there is a lot of work in communication to get them to see data as a major innovation enabler, something to drive their part of the business to new heights.
Another interesting aspect Mandeep talked about was that self-serve might actually be seen as threatening to data consumers. Previously, they controlled - to some degree - their own ability to get access to data but now, it's on the producing team and consumers only get what producers are willing to share. The consumers created the business value by doing the analysis and transformation and now that is pushed much more onto the data producers. Will consumers feel their power and importance is diminished? If the value of data work is attributed to the producers, will data fluent consumers still lean in to leveraging data as much as they did previously?
Mandeep returned to product thinking and her view that a product is only a product if it's providing value. You start from the value you are trying to deliver and work backwards. Build your KPIs around actually delivering value instead of simply creating data products with the hope they create value.
When thinking about data as a product, Mandeep encourages everyone to have conversations about it in their organization and what it will actually mean and look like in their specific organization. Because it's easy to assume everyone is on the same page when they really aren't. And that confusion will bite you in the end with more friction than clearing it up early.
Mandeep believes that in a transformation journey, it almost always starts as somewhat disjointed - a disruptive phase before the planning phase. Part of going on a journey is preparing for that journey. Once people are aligned, that is when you can really start all heading forward. You need pioneers or leaders, those front runners to show people it's safe. But it will still take some time before that alignment. Don't get concerned when that happens even if it feels like everyone should align after the first presentation 😅
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
JGP's LinkedIn: https://www.linkedin.com/in/jgperrin/
Amy's LinkedIn: https://www.linkedin.com/in/amy-raygada/
Andrew's LinkedIn: https://www.linkedin.com/in/andrewrhysjones/
Andrew's website: https://andrew-jones.com/daily/
Andrew's book: https://data-contracts.com/
Data contract standard project Bitol: https://lfaidata.foundation/projects/bitol/
JGP's blog: https://jgp.ai/
In this episode, guest host Jean-Georges Perrin, Data Innovation Consultant at ProfitOptics (guest of episode #130 and panelist in episode #227), facilitated a discussion with Amy Raygada, Senior Data Product Manager at Swiss Marketplace Group (guest of episode #165), and Andrew Jones, Principal Engineer and Author of the book on Data Contracts (guest of episode #29). As per usual, all guests were only reflecting their own views.
The topic for this panel was all about data contracts and how do we go about getting them in place. Much of it was about the general concept but some of it was specifically about how do we think about data contracts applying to data mesh. This was the first topic I really did a deep dive into in early 2022 and it has evolved but is definitely still evolving.
Scott note: As per usual, I share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Alexandra's LinkedIn: https://www.linkedin.com/in/dralexdiem/
In this episode, Scott interviewed Alexandra Diem, PhD, Head of Cloud Analytics and MLOps at Norwegian insurance company Gjensidige.
Gjensidige's approach closely aligns with data mesh but they are starting with a focus on consumer-aligned data products as they have a well-functioning data warehouse and are not looking to replace what isn't broken.
Some key takeaways/thoughts from Alexandra's point of view:
Alexandra started with a little about her background coming from academia into the commercial world and how that shaped her views of things. She was the first data scientist at a company so before she could really do data science, she essentially had to work as a software engineer; that meant learning many good software engineering practices. When she moved to data, she thought 'we should use these practices here too' and then she also came across Zhamak's posts on Martin Fowler's site and it all started to click.
Specifically at Gjensidige, Alexandra was brought in to lead of team of software engineers acting as an enablement team plus a platform team. Their role was to focus on bringing these good software practices - e.g. DevOps, automation, testing, etc. - to the data/business analysts to help them build data products. It has evolved to be more sophisticated but the team is still about enabling people to build data products.
Gjensidige already had embedded analyst teams in many of their domains so when Alexandra and team started to roll out the data mesh implementation, there were already a number of data-capable folks who understood the actual business aspects of the domains. That meant there wasn't the typical pushback on the domain actually owning their data products, it was more about enabling them to do so and build a maintainable and scalable product. This process of pair programming between her team of software engineers and the domain data experts means her team becomes more and more data fluent while the domain learns how to write good software code. They specifically leverage a model of two of her team and two of the analysts in a data product creation team to provide enough information exchange but not too much overhead. That intimate understanding of what has been created also helps her team to help find reuse in other domains - they more deeply understand what has been built and can direct teams towards it quickly. That speeds time to market as well for the new team. Lots of wins all around!
When looking at the central enablement team's strategy, Alexandra strongly believes in a minimum viable data product approach. Her team only has a handful of people and they have 25 analyst teams to work with. The team has to focus on getting each analyst team to capable via the first data product - again with only two analysts on the team - and then letting those two analysts propagate the knowledge to the rest of their own teams. Otherwise, the central team would be too overloaded. So again, the focus is on teaching the analyst teams how to build good data products and then moving on. Otherwise the central team just isn't scalable or you have so many people in the central team that it becomes far harder to find patterns and share information. The domains have to deliver value themselves so teaching them to do so and then moving on is a sustainable strategy.
When communicating with the rest of the organization, Alexandra rarely uses the term data mesh. She points to data product and self-service platform as things that resonate with people and help communicate what she's actually focused on doing: generating value. Most people don't care that the way you are generating value is data mesh. It's simply a mechanism. 'Lean' into that. Scott note: lean is a bad pun here because she mentioned how helpful The Lean Startup is to focusing on value generation.
One very interesting note Alexandra talked about was training the domains in reusing data. Historically, it's been very difficult to reuse data because you didn't have the information about how it was created and didn't really have a reliable source. Getting the spreadsheet from a colleague each month isn't that reliable 😅 so, you will likely need to train your domains on reuse, especially finding sources to reuse and how to see if something fits their purpose. That can be the producing team too, teaching them how to share what they've built to other parties that might want their data.
Alexandra noted that most of the data products they are building use an existing clean and well understood data source: the cloud data warehouse. They are leveraging a hub and spoke pattern from that warehouse for their products. People already know and trust the warehouse so it made sense to them to start there. Essentially, everything ends up as consumer-aligned data products in a sense. Relatedly, for Alexandra and team, they don't see a need to adhere to every aspect of data mesh, especially at the start of their journey. She said, "[I] don't really see the point of having to destroy value before I should be able to generate new value. I can very happily just generate value on top of the value that I already have." They had some things that were working well already and breaking it all down to fit the paradigm didn't make sense to them. However, she is aware of the additional challenges this can bring and made the conscious trade-off.
Scott note: this is an obvious data mesh anti-pattern because the upstream isn't directly from source systems - the teams building the data products don't control the source or their source-aligned data products. But if you don’t have an existing bottleneck from your cloud data warehouse, why fix something that isn't broken? This may become a bigger challenge later - Zhamak has written why not owning source data creates challenges - but if they are willing to take the tradeoff and understand those tradeoffs, is it a bad approach? I don't think so _in their case_ because the data warehouse is functioning well/isn't a bottleneck.
Alexandra talked about how to really embrace a culture around 'minimum viable x' in data. In data science, at least there is a good understanding of hypothesis testing but even then, it's often hard to embrace the necessary 'fast fail' model touted by things like The Lean Startup. Trying to understand how to hypothesis test value is also difficult and people have historically seen the challenges in iterating on anything data related. So there is a learning curve but also generally a necessary cultural change to embrace hypothesis testing and fast fail around data.
On advice to her past 'data mesh self', Alexandra gave a reasonably common response, circling back to an earlier point: stop talking about data mesh, at least early in the process. Data mesh is a set of guiding principles, not the answer. Talk to people about changes to their ways of working and target outcomes. Why are we taking on change? People hear data mesh and expect it to be some technology or technological approach. You can use the name when people ask for what you're calling the approach but selling it as doing data mesh doesn't help your business partners get it. It becomes a much more tiresome approach to specifically focus on data mesh instead of the ways things change and what matters to them.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key Points:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Ryan's LinkedIn: https://www.linkedin.com/in/ryancollingwood/
In this episode, Scott interviewed Ryan Collingwood, Head of Data and Analytics at OrotonGroup. To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Ryan's point of view:
Ryan started off with some framing of how he looks at tech approaches in general but especially how he started looking at data contracts. Most paradigms are presented as if every organization is very tech-y, like a tech startup. With data contracts, much of the content "…there was this assumption that you had multiple teams of people that had a fairly high degree of technical sophistication, … or maybe even data was their primary focus." So when a less tech-y company wants to leverage the paradigm, there is always some adjustments necessary 😅 and when it comes to those types of companies, it’s so much more about the people than the way most paradigms are presented. It makes some sense because every org's ways of working and culture are different but it still can feel very removed from reality for less tech-heavy companies.
When focusing specifically on data contracts, Ryan's company is far more batch than streaming. So trying to even leverage the best advice (Scott note: I highly recommend Andrew Jones for that), he had to adjust some aspects to a world where things were a bit more messy and with teams that aren't as data mature. When approaching how to tweak data contracts to still work, he asked the rhetorical but crucial question: "What are the trade-offs that I can make, while still being true to the value and the benefits that I want to get out of this?"
Ryan moved into what he sees as the minimum viable value aspects of data contracts. You need two parties, you need an agreement of some kind that is recorded, and you need access to data that conforms to the agreement*. As to the parts of the agreement, Ryan focused on two factors at the start: semantics and data quality. If people can't understand the data can they use it? If they don't understand the quality, can they really trust it enough to rely on it? So they worked to create a data dictionary and also provide people a better understanding of the different angles on data quality.
Often, when comparing with what was presented for a tech-heavy company to what is possible at a more regular organization can be disheartening according to Ryan. The idea that the end picture at your organization should look like the one presented is pervasive. So it's not only hard to adapt the approach but then you wonder if you even captured the value 😅 Can you even call it 'data contracts' or whatever you are working on?! Imposter syndrome is very common here. Scott note: you could definitely call what Ryan and team are doing data contracts :)
Ryan also talked about how in data contracts, you must build for change. Change is the only constant after all. So creating systems that don't handle change well is a great way to manufacture more headaches down the road. Much like in software testing, you can more easily tell when something no longer works and needs to be changed. And when the data team is the actual data producer - if the data team are the ones transforming the data, that's often the case or at least is the only group of people consumers talk to with a centralized data team - they are much more sure that what they are doing is correct.
Another key learning Ryan had along the journey was that when displaying data quality, make the metrics more easy to understand to the layperson. Historically, data quality has been measured with complex statistics. Most people can't easily read the charts from that to understand what's going on. Make the data quality metrics understandable so people can see progress but also get a sense of how well they can rely on data. It is a sad truth that you can deliver value but if you can't get others to see that value, it isn't valued. Showing that value gets people to lean in.
Ryan dug a bit deeper into creating systems that act with empathy. If you approach data contracts as consumers only get what the producer shares, that doesn't end up serving the end needs that well. But if you are treating the contracts as the culmination of multiple conversations, the producer can start to really understand the impact of bad data. How much work do data consumers have to do to actually use the data? This is where empathy and product thinking come in.
"…data, as we know, it is merely a side effect of activity, of stuff happening." Ryan believes we need to move past the 1s and 0s thinking in data and focus on what it reflects and how that impacts the people in the organization. Conversations can be hard but they give you the context necessary to maximize the impact of your deep systems work. Talking with people can help both parties bridge the gap between understanding what is happening in the real world versus the data 😅
Internally in Ryan's org, they wanted to review their general processes. Part of that was the uncomfortable truth that change, especially to processes, impacts the data. So that review created a great opportunity to start to implement data contracts. It wasn't about telling people they were doing data contracts, it was about getting people bought in to what value could be delivered if they did data quality and trust better. It just happened to be via data contracts.
When actually starting out, Ryan looked for one ally that was willing to take on some of the complexity of dealing with data contracts and saw the potential benefits. Instead of trying to convert the whole organization, it was contained and let Ryan learn how to implement data contracts well in his specific organization. That initial success gave him the confidence to move further and the success story to entice additional partners/allies.
Ryan discussed the push and pull of data quality and value. While it might be valuable to have a long history of data, is the cleanup worth it? Really have conversations and make hard choices that align to return on investment instead of merely do consumers want it. Similarly, people need to confront the idea of data being right or wrong. They need to consider what is the cost of some data being wrong, especially slightly off. If that's for a regulator, potentially high. But if it's your weekly marketing leads report and it's off by 0.2%, how big of a deal is that? And how much trust is lost if it's wrong? Can we get people to understand data is never 100% clean/right? Getting people to act on signals will likely be somewhat challenging but it's a better way to navigate than trying to wait for exact measurement in many - most? - cases.
Ryan wrapped up back on dealing with yourself and others with empathy. You might not get it right at first but if there's trust, you can iterate towards better together. That goes for your data, your processes, and your relationships.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Kate's LinkedIn: https://www.linkedin.com/in/katecarruthers/
Kate's 'Data Revolution' Podcast: https://datarevolution.tech/
In this episode, Scott interviewed Kate Carruthers, Head Of Business Intelligence at the UNSW AI Institute and Chief Data & Insights Officer at UNSW (University of New South Wales). To be clear, she was only representing her own views on the episode.
UNSW is not currently implementing data mesh but are preparing to be able to do so. This is a great lesson in building up the capabilities to move forward towards your goals but not rush.
Some key takeaways/thoughts from Kate's point of view:
Kate started out with a bit about the catalyst for her current data journey towards data mesh. About 10 years ago, she saw that universities and especially UNSW were going to "undergo a very big digital transformation and that data would underpin it as an organization. [So] we would need to be on top of our data if we were going to be able to ride that wave." She also gave some color on what running a data office at a university entails. At UNSW, it's split into three general areas of administration, learning + teaching, and pure research.
There are some major challenges when it comes to providing data capabilities - especially self-service - to the academic research arm of a university according to Kate. They all have their own ways of working and want - demand? - freedom to work the way they want. Yet, the data team's job is also to "keep them safe." That safety has many facets as well. And the research capabilities of a university can mean some truly world-changing interdisciplinary collaboration. But that only happens if the teams can actually, you know, collaborate 😅
When it comes to the non-research area, Kate believes data mesh is an even better fit. "At the end of the day, data mesh is about controlling the bits that you need to control, and giving people the freedom to do what they need to do, safely."
As many guests have noted, Kate believes when it comes to your organization's data journey, "technology is kind of the least of your problems." It’s about people and often even getting them to recognize the problem with their ways of working and how better data maturity will help alleviate their problems. It's not just the data itself but their understanding and relationship to data.
Kate and team built a quick cloud warehouse PoC that showed people the ability to onboard new data sources in weeks instead of taking up to six months. Showing them instead of simply telling them really won people over. People could connect moving to a cloud data warehouse to business benefits. They also anchored it all to business needs. Yes, rebuilding their architecture to move to the cloud was going to be work but it meant speed to new data use cases and easier management.
When Kate was working to tie her team's work to the overall business strategy, she remained focused on the human relationships and people aspects of doing business. She really recommends building relationships with "customers" of your data work because then they feel comfortable to come to you with more types of problems and challenges. And sometimes that kind of culture/approach isn't for everyone and that's okay. If people aren't willing to treat customers as people, they aren't right for her team.
When asked about her frequent use of the word "safe", Kate talked about keeping people from misusing data or even misusing the trust people who provided that data - e.g. the students at UNSW - gave the organization. Anytime someone wants to share sensitive information like PII, there is a data controller that needs to review the justification. Keeping that human in the loop means there is a real understanding and consideration of 'is this okay?' On the flip side, the team has been proactive in sharing information that someone should have access to, e.g. a professor being able to know who is in their class and being able to contact them.
Kate mentioned that when they implemented the data controller review, the data producers were much happier. Previously, they had no real say in how their data was used but now, they are listened to. It also strengthened relationships because consumers had to actually collaborate with producers to get access to their data. It's creating interesting conversations and people can get more creative around data to achieve their goals with more data safety. And her investment in hiring a bunch of business analysts has created some great value leverage points.
Going back to keeping people and data safe, Kate talked about their struggles with data puddles - where people are copying data into lots of areas instead of accessing the data where it is. And they aren't securing that data well, which leads to more challenges and potential issues. But it's still a process to give people all the access they need and make that copying data less attractive. As like many areas, it's a work in progress 😅
Kate sees the attractiveness of moving fast but believes people need to focus more on sustained incremental change, that they overestimate the value of the former and underestimate the value of the latter. It's similar to transformation versus a change that will revert. Fast changes are far less likely to stick or even work. And people feel less of the suddenness and fight against it far less if at all when it is gradual incremental progress.
Another point Kate emphasized was that people need a mental map for change. If they don't understand what is changing and why, they will inherently fight back, even if the change is good for them. It's simply human nature to not want change. So take away the fear of change to make it easier for people. Basically create change with and through people instead of pushing change on them whether they want it or not 😅
The conversation wrapped up around GenAI, especially because Kate is involved in the UNSW AI Institute. She is seeing the open source large-language models (LLMs) improving at a rapid pace, sometimes multiple times a day. And there is a lot of promise even if things are early days. At UNSW, they are figuring out good ways to leverage GenAI in education instead of trying to fight against it like some math teachers did against calculators. It's here to stay so they have to adapt and adopt.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Benny's LinkedIn: https://www.linkedin.com/in/bennybenford/
Iulia's LinkedIn: https://www.linkedin.com/in/iuliavarvara/
Nailya's LinkedIn: https://www.linkedin.com/in/nailya-sabirzyanova-5b724310b/
Stefan's LinkedIn: https://www.linkedin.com/in/stefan-zima-650229b7/
In this episode, guest host Benny Benford, Founder and CEO at Datent - a data transformation focused consultancy/community - and guest of episode #244 facilitated a discussion with Iulia Varvara, Advisory Consultant in Digital and Organizational Transformation at Thoughtworks (guest of episode #268), Nailya Sabirzyanova, Digitalization Manager at DHL (guest of a soon-to-be-released episode), and Stefan Zima, Data Transformation Lead at Raiffeisen Bank International AG (guest of episode #270). As per usual, all guests were only reflecting their own views.
The topic for this panel was transformation when it comes to data and data mesh in general but especially understanding how organizational transformation must play a large part in a data mesh implementation to be successful. And that transformation is not simply making changes, it is making _lasting_ changes. Organizational transformation is a crucial aspect of doing data mesh even if it's not spoken about all that often.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Sean's LinkedIn: https://www.linkedin.com/in/seangustafson/
In this episode, Scott interviewed Sean Gustafson, Director of the Data Platform at Delivery Hero.
Delivery Hero has been on the data mesh journey for longer than most organizations, at least over 3 years.
Some key takeaways/thoughts from Sean's point of view:
Sean started with a bit about how he sees his role as leading the data platform team. It is very challenging but still important in his view to try to shape culture even through the data platform. There are so many places in data mesh where there is friction, how do you make things easier as everyone transitions to product thinking and decentralized ownership? Just because you have mandates from the top, people need new ways to accomplish new goals. Make your platform reflect the type of data culture you want. Instill in people the understanding that they can and should participate in your data culture/work. Easier said than done of course.
Relatedly, Sean believes the data platform should show people the right way to do things, give them that easy path where possible. But still give them the freedom to do some aspects … not so right 😅
Treat your data platform as a product is something Sean strongly believes in. And to do that, you need someone acting as a product manager. It's not rocket science, we know how product management works and it's not very different when it comes to building a data platform. But you need someone specifically focusing on user needs. And part of that role is also to advocate new features and using the platform. Just because you built it, that doesn't mean people will use it.
When asked about iterating to good, Sean talked about how in product management, good practice is about making constant and small improvements but also balancing the bigger picture/big bets. It's not always about the big new platform but sometimes, it's okay to shake things up - make small bets when small bets are good enough but make big bets when necessary. But you have to do that by balancing the short-term and long-term picture. Fail fast and iterative improvements are crucial to good product thinking in software and we need to apply that to data. But again, big changes are okay if you properly build to them instead of trying to flip a switch. He specifically mentioned that it will be hard to iterate to a platform that does decentralized ownership well from one that was highly centralized. Not impossible but at least consider building that out more from scratch.
Sean talked about Generative AI and how it's starting to change lots of people's views internally about data. While previously, many software teams were at best reluctant/hesitant to model their data, there is a big interest from the software engineers to directly interact with the large language models (LLMs). Tools like dbt previously brought many new people to the data party, making it easy to model data - at least structurally - so hopefully GenAI will mean more people learning to model their data. There are inherent challenges but the more the merrier when it comes to people working to produce good data. We just have to make sure they learn how to do it well 😅many who are new to data modeling do it… not so well…
When it comes to product management, you need to measure how well you're doing. For Sean, that of course extends to the data platform. While KPIs can be somewhat hard around your data platform, that doesn't mean you get to slack off and not measure things. At Delivery Hero, right now they are using surveys to measure a number of things around their data platform rather than trying to measure things automatically without context. It also creates a lot of conversations in the data platform team about what are you trying to do and why, which prevents a lot of waste. It's not perfect but it's getting better. Scott Note: this is why I am writing a book on success factors then one on success metrics in data mesh 😅 this is HARD
Sean talked a bit about APIs and how much data products _should_ be treated like APIs. Not just versioning but tracking usage and having users register to use them. There's a lot to learn from how APIs evolved so we don't have to make the same mistakes in data. Scott note: Zhamak comments on this VERY frequently that API approaches are crucial to data mesh
When talking to software engineering people, Sean has found using data terminology, especially data mesh terminology, doesn't really resonate with them. We probably need to come up with new terms - or potentially use the terms Zhamak took from software and just make them about data too instead of inventing new terms. But be prepared for it all to fall back to that most software people will see the back-end systems as more important than the data. If you get them over that hump, it's far easier to get them bought in on data mesh. You may be able to win them over by showing them how the data is used internally.
Incident management in data is still pretty nascent in Sean's view. While on the software engineering side, there are very well established processes, often in data it has been more slapdash at best. No escalation, no prioritization, no formal process, no post mortem + shared learning, etc. The traditional measure around data issues - how much money did we lose - often isn't applicable to data. So we have to rethink what matters and why because our prioritization is often skewed.
Sean wrapped back to the start about how important culture is. Not just getting your organization to be data driven but setting up more and more people for success in your organization through their work with data.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key Points:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Contact email: Swimwith[at]gulpdata.com
Lauren's LinkedIn: https://www.linkedin.com/in/laurencascio/
Chris' LinkedIn: https://www.linkedin.com/in/censey/
In this episode, Scott interviewed Lauren Cascio, Chief Fish Wrangler, and Chris Ensey, CTO at Gulp Data.
From here forward in this write-up, L&C will refer to the combination of Lauren and Chris rather than trying to specifically call out who said which part.
Some key takeaways/thoughts from L&C's point of view:
L&C started with discussing how many organizations view their internal data landscape/estate and how it's not a complete picture. There tends to be a perspective that an organization's data is only useful for their internal use cases and often that each set of data is only useful for one type of use case. And L&C just haven't seen that be true - internally, most orgs have data that could be useful to existing use cases . How that typically manifests is data silos where data that should be shared isn't because people aren't aware it exists. Or the other side is that data producers have no real idea of how their data is being used downstream by other parts of the company. Externally, most companies' data is often very useful to other organizations in entirely different sectors.
When asked about why lines of business have such a hard time understanding what data other LOBs have, L&C talked a bit about the technical challenges but much more about the organizational. In many - most? - organizations, lines of business have treated their internal data as overly precious, making sure it was structured specifically so they could use it. Trying to get them to structure it so others can use it is an emotional hurdle because it can feel like giving up that control. Thus, it's been hard for other LOBs to even know about the data across the organization, let alone use it. Add to that the challenge of businesses treating the data team and especially their infrastructure as a cost center and that further impedes their data journey. When it's hard to make the tangible business case for updating your data infrastructure, it's easy to fall further behind. If much of this sounds familiar, it's frequent hurdles towards implementing data mesh.
L&C talked about how typical it is where organizations understand they want to visualize their data but without a specific goal in mind. Just visualizing without an expectation of what it will be used for is not product thinking. What information do you need to make your external facing products better? What will cause you to act? Instead, it's about "what does the data tell us" which is not often aligned with taking actual action. That leads to wasted cycles and money; it also often leads to wasted buy-in from upper management - they really only have limited patience, spend it on what matters. Be crisp on what goals you are going after then develop the data and analysis to help you actually go after those goals.
It's far easier to get exec buy-in on selling your data externally than investing further for internal use in L&C's experience. That's because there is a tangible outcome at the end of the road. Look to try to shape your asks for additional funding based on that principle: a tangible ROI makes decisioning easier.
For L&C, there are many use cases that could be unlocked in most organizations if only people knew what data was available. Finding ways to discover and share more about what data you have internally is very helpful. Yes, a data catalog is great but finding better ways to make people aware of the available data will unlock new valuable use cases. An audit for what data to sell externally is one way to spark these conversations but there are many others :)
L&C pointed to two differing types of companies regarding selling their data. The first is low margin businesses. Because they are so reliant on volume, they end up with a considerable amount of data that they could potentially monetize. The other type of company is early stage companies that have yet to reach product market fit, especially B2B. They often think their data will be very valuable but selling data becomes a distraction far too easily. Focus on your core business, not small external monetization streams.
On the somewhat controversial topic of data monetization, how people's information is protected versus leveraged, L&C believe there is a greater good in general to your information being shared. While something like GDPR gives the perception of your data being protected, it's not really all that true - everyone's data is out there already 😅. Meanwhile, there is lots of potential good that can come out of more comprehensive data sharing, e.g. better information to fight diseases from more patient information or lower cost of items in retail stores from data generated in loyalty programs.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Important points:
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Stefan's LinkedIn: https://www.linkedin.com/in/stefan-zima-650229b7/
In this episode, Scott interviewed Stefan Zima, Data Transformation Lead at RBI (Raiffeisen Bank International AG). To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Stefan's point of view:
Stefan started with a bit about his background and current focus. From having worked in data and transformation, the current transformation needs - data and digital in general - for many organizations are accelerating. There is a need to move towards more modern platforms and a self-service orientation of course but the organizational change especially around mindset cannot be overlooked.
As for why the banking industry is embracing data mesh and being data driven, Stefan mentioned that over the last five years or so, there have been many major industry changes. One aspect is the rise of fintechs that are built from day one with data in mind - to stay competitive, more traditional banks need to be able to move extremely quickly and the best - only? - way to do that intelligently is with data. Another aspect is simply the domains are getting more and more capable around leveraging information to gain competitive advantages so they need to feed that with quality data to win in the market.
When the bank started to see more and more friction - e.g. more pushback from the data protection officers to comply with ever advancing and increasing regulation - Stefan and team realized they needed a new approach. Instead of trying to make improvements to the existing processes, there needed to be new ways to get things done instead of relying on existing bottleneck creating interdependencies. That's part of what led them to data mesh - taking points of friction and shifting them left and reducing handovers meant faster time to market with new information and use cases.
Stefan talked about how many people in transformation and data push too far, too fast with their vision. When people are stuck in their day-to-day, getting them to imagine the company that could be in 5 years is at best inspiring but often not and it certainly doesn't help them today. Help them connect that vision of that 'data-driven future' to what you will do for them now. Don't expect people's mindsets to shift overnight.
Relative to transformation, Stefan discussed how important internal communications is. There is of course the messaging to drive understanding but also the communication by the leaders to show support. If you lack visibility of your top-down support, it's much harder to get people to take your efforts - data mesh or otherwise - seriously.
At RBI, Stefan noted that they are through their data mesh PoC phase. They were able to prove out significant value to upper management and get upper management to further buy-in. With that buy-in, there was more communication internally, getting more and more people aligned that data mesh was the way forward. But of course, data mesh is in some ways just a label and approach, it's not the point.
Stefan talked about what can we take from Agile transformation and successful business transformation in general to use for data/data mesh transformation. A big focus in Agile is on communication and transparency. When people feel informed and heard, they are far more likely to buy-in. Humans struggle with uncertainty. When it comes to data mesh, there will be all kinds of new roles and responsibilities so you need strong communication to keep people informed and not feeling lost. Even if that is about trying something and seeing if it works. That honesty and transparency will have far more people leaning in than trying to issue proclamations. And of course, be ready for some politically driven issues because we are changing the way people operate.
The ever present 'what is a data product?' conversation also happened at RBI according to Stefan. And it was important to define it in their own world - every organization has their own needs, understandings, and requirements/restrictions that will mean data products look slightly different. Specifically for RBI, a dashboard isn't a data product but is a business product, meaning there is still a clear ownership model. They also focused a lot on automation and risk assessment/mitigation. In a heavily regulated industry, risk is always a crucial factor. When it comes to working with the Data Protection Officers, it's important to make it safe and worthwhile for them to say yes when saying no is easy and prevents all risk.
Stefan then went into a bit about where RBI is headed around self-service and why that's so crucial to the company's data mesh ambitions. For him, self-service is far more than just giving people access to data. Yes, you need a platform but you need upskilling and data literacy, an understanding of compliance, an understanding of your overall data ecosystem, and an understanding of the tooling and processes. You also need to build a platform that can be leveraged by non-experts. You need to make it easy for the general populace of your organization to actually consume, understand, and produce data.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn.
Transcript for this episode (link) provided by Starburst. You can download their Data Products for Dummies e-book (info-gated) here and their Data Mesh for Dummies e-book (info gated) here.
Vanessa's LinkedIn: https://www.linkedin.com/in/vanessaeriksson/
Sid's LinkedIn: https://www.linkedin.com/in/siddharthin/
Stefan's LinkedIn: https://www.linkedin.com/in/stefan-zima-650229b7/
Duncan's LinkedIn: https://www.linkedin.com/in/duncan-cooper-1113722/
In this episode, guest host Vanessa Eriksson, the first CDO in Sweden and the head of data advisory company Vanessa Eriksson AB facilitated a discussion with Duncan Cooper, Chief Data Officer for Northern Trust Asset Servicing, Sid Shah, Head of Data Monetization and Platform at Airtel (guest of episode #258), and Stefan Zima, Data Transformation Lead at Raiffeisen Bank International AG (guest of episode #270). As per usual, all guests were only reflecting their own views.
The topic for this panel was about the leader's role in a data mesh implementation and what these four panelists have learned in that role. This was the second iteration of a panel we will likely have about every six months or so - the first was episode #215 from April of 2023.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Iulia's LinkedIn: https://www.linkedin.com/in/iuliavarvara/
In this episode, Scott interviewed Iulia Varvara, Advisory Consultant in Digital and Organizational Transformation at Thoughtworks. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Iulia's point of view:
Iulia started with a few basics about general transformation and digital transformation, whether that includes data or not. To really be able to embrace a digital and data-driven future, organizations need to embrace product thinking across the entire organization. They need to align their strategy and operating model to be adaptable and flexible. If the organization has already embraced product thinking, few have really pushed that to data in her experience. But if the organization is new to product thinking entirely, then starting from the data side could create a strong catalyst because data is probably one of the hardest concepts to apply product thinking to - after taking on data, product thinking is far easier to grasp in other areas of the business.
In general, Iulia believes that mindset changes don't come from mandates. Instead, by implementing new ways of working, people's mindsets will start to shift once they see the impacts/benefits of those new ways of working. They see the benefit and change their mindsets. But that will of course take time and concerted effort - the mindset change by decree is faster but doesn't typically stick. Show don't tell.
As other guests have noted, Iulia pointed out to maximize the chance of a data mesh implementation succeeding, you have to take into account the existing ways of working, the organizational and team operating models. Yes, certain aspects will need to change but trying to completely change an organization's operating model is going to be too disruptive. Instead, align a transformation paradigm to how the organization already works so people can evolve and adapt. Don't throw them in the change deep end and don't throw the baby out with the bathwater. Of course, this also means there isn't some blueprint for data mesh that will work for all organizations.
In Iulia's book, the first two pillars of data mesh - domain-based data ownership and data as a product - are the two that have the biggest impact on the organizational operating model. She said, "When you start thinking about your data in terms of products, and put your user in the center of your attention, you try to organize all your efforts around the user needs. Right? You create this connection between the data team and the value." That is a big change to how most organizations work around data and it will take effort to make it happen.
In general, Iulia recommends that for any large operating model change, you really need to clearly communicate multiple things. What are the changes, why are you making them, what is the actual target outcome/goal, what are the measures of success, etc. That way, people can measure how well things are moving forward and more easily prioritize. "Because there would be so many things to be done at the beginning, that team really needs to have a clear understanding what to start with." Transformation will mean tens of changes, understanding where to start and why are crucial.
Specifically to product thinking, user value is your Northstar for Iulia. It will inform your strategy, vision, and business goals. Those business goals will be split into hypotheses of value for how you can reach the goals. This is where you start to allocate teams, to the actionable items from the hypotheses of value. But it all comes back to focusing on user value. Steer your work through feedback loops to focus on that user value and you have a great shot at implementing product thinking/focus well.
Iulia pointed to something many miss when it comes to treating your data as a product. If you don't have a long-lived data product team, it can cause many issues that significantly undercut the value of building data products. One is that you often lack the subject matter expertise in the data product team, so the information encapsulated is not nearly as deep or as relevant to the topic area for the data product. Another is that if the team isn't long-lived, will they really have the time and psychological safety to run experiments and innovate?
Similarly to data as a product, Iulia recommends going small, then sustaining, then scaling when it comes to domain ownership. Basically, start from one to two domains and go broader over time. Trying to reorganize your organization on day one so one or two domains can own their data in that time-frame, that's a TON of effort. Don't try to revolutionize your company to do data mesh, evolve and build the understanding as you go broader. You need to prove out value first too before you go broad.
In wrapping up, Iulia returned to the concept of funding the teams, not the work, and especially long-lived teams. When you fund the teams, they are able to focus on finding value. There isn't an expectation of the teams to be prescient, always knowing what will be valuable. And there isn't a need to simply react to tickets instead of finding what will be of value. The other aspect is that you can understand what should be decommissioned. Far too often, data work continues well past when it is valuable. But with a product mindset, teams can constantly be focused on user value and shut down things that no longer drive value.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key Points:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Akins' LinkedIn: https://www.linkedin.com/in/akinslawal/
Schema for Success: https://www.schemaforsuccess.com/about
In this episode, Scott interviewed Akins Lawal, a Data Strategist. To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Akins' point of view:
In Akins' view, many more people need to slow down a bit to really consider what they are building, why, and for whom. Far too often, data people - and tech people in general - want to build something but we need to focus on good information architecture. We need to consider a lot about the specifics of why we are building something and also who will use it and how in order to maximize positive outcomes.
When starting with good information architecture, Akins recommends asking a lot about what the system will do and why. Map out the user flows even, how will people access information. That way, you can start to back into what kind of systems and approaches will support what you are trying to accomplish - it's not about what you are trying to build, it's about what are you trying to accomplish and building something that can help.
As you start to plan out your information architecture, Akins recommends you start finding your champions - you need people to rally around to get things moving forward. Then you plan out your process - there are many different tried and true processes for sharing context/information but it's important to find one that works well with the use case and your organization. You should be looking for happy mediums between all involved because no one will get everything they want if you're doing it well.
In Akins' view, if you want to become data driven as an organization, you need to focus on hiring learners, not just for current skill sets. Data - how to analyze it, leverage it, share it, etc. - is a lifelong problem/challenge. New tech and approaches always emerge. You want someone focused on staying up-to-date on best practices, not specifically proficient in one tool that could be close to obsolete in a few years.
Leadership buy-in is far and away the most important factor to new initiatives succeeding, according to multiple studies from groups like HBR and McKinsey. So Akins recommends to make sure you are aligning with those leaders and showing them the benefit of improving your data initiatives. A big reason so many data initiatives fail is that lack of buy-in and support. The best laid plans are still far less likely to succeed without support from above.
Akins talked about how maturity models can be a very helpful tool for finding what's already working in your organization relative to data work. You can find the patterns and then assess if you can make those into repeatable patterns for the rest of the organization or not. It's all about creating a situation where things can mature.
In person collaboration for Akins just kind of 'hits different' - you are better able to exchange context if you're in the same room and able to whiteboard. Potentially that's around driving meaning and trust? Virtual tools still have not caught up to the in-person collaboration capabilities. So if you aren't in person, he believes you will likely need to meet more often just to ensure you really understand each other.
Akins then talked about the importance of leveraging data to empower people and weaving that data understanding and empowerment into the fabric of the organization. How do you empower people to make more and better decisions with data? That's how you move towards being data driven.
Something Akins likes for driving data sharing and usage is incentivization: how are you rewarding people for sharing and leveraging data? Leaders should be giving teams and people that are sharing and using data well lots of accolades and potentially other rewards so others want in on the action.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key takeaways:
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Madhav's LinkedIn: https://www.linkedin.com/in/madhavsrinath/
In this episode, Scott interviewed Madhav Srinath, CEO at Nexusleap.
Overall, we are super early in the Generative AI cycle and hype is huge. This discussion is one of early impressions, not fully formed answers. It's far too early for that.
Also, FYI, there were some technical difficulties in this episode where the recording kept shutting down and had to be restarted. So thanks to Madhav for sticking through and hopefully it isn't too noticeable. Generative AI will mostly be shortened to GenAI throughout these notes. LLM stands for large language models which power GenAI.
Some key takeaways/thoughts from Madhav's point of view:
For Madhav, a lot of what Generative AI has become is the concept of data mining with a personable interface. We've been trying to create a way to dig into data that is unstructured and get some insights or information - something that is structured - for a while. The concept isn't new but we've finally found something that might actually be able to do it well and make the outputs easy to consume.
Right now, in Madhav's view most of the emerging GenAI use cases have been pretty shallow, such as to help write an article. It probably can get far deeper but it's quite early days. However, we probably need to put a human in the loop to actually make sure the answers LLMs are giving are correct. That might be the best option in his view, something like a guide plus guardrails driven by a human to make sure these LLMs don't hallucinate. In a way, this isn't all that different from other machine learning work - black boxes tend to have unexpected consequences.
Madhav's view is that it's totally okay to start working with GenAI at your organization as long as you understand GenAI/LLMs have quality issues right now and are really only at the MVP stage in many senses - especially if you are using them internally on your own data. There will probably be good ways to put a wrapper around them to prevent improper data usage/leakage and prevent hallucinations as well.
Starting with domain specific questions/problems is where Madhav thinks people should focus their GenAI work. If you try to feed an LLM a ton of information from many sources, you can't really be sure of the logic it uses to generate answers versus keeping the inputs tighter and asking about more specific business areas. Keeping that tighter focus and having many LLMs across the organization gives you an ability to more tightly focus your models on specific topics. You can then attempt to add additional focus areas to those LLMs once you have the model performing well on one topic.
While LLMs aren't magic, Madhav is seeing an emerging use case where people point the LLMs at maybe two data products and ask it to infer relationships between them. It might discover something people haven't thought of before. You still need a human in the loop or you end up with something like the Pastafarian 'belief' that global temperature rise since the 1700s is caused by the decreasing number of pirates globally - correlation doesn't equal causation. It's not a magic wand but it could help people find more places where data is already interoperable or should be. Then, he's seeing once those relationships are discovered, a differently tuned GenAI model is used to actually infer some information based on those newly discovered relationships. Again, specialized models.
Madhav doesn't believe most organizations should be training their GenAI models from scratch. Instead, go and find the open source models and add your necessary information to the training - basically, why start at zero to train it to one when you can start at 0.7? Leverage the work others are doing so you don't need super expensive LLM training focused engineers. By starting with an existing base model, you can tune it based on your own answers instead of trying to feed it very specific data and supervise the base-level initial training. Leverage your business subject matter experts and get them to share their tribal knowledge with the LLMs as well.
Circling back to the idea of layered LLMs, Madhav talked about how some organizations are having models specifically tuned to answer questions about data but then there is a secondary LLM that is focused on checking the answers the first one gives for sanity/correctness as well as around governance, e.g. security and privacy concerns. Again, that separation of the work. And it's not nearly that expensive when it's all done in a serverless way - if your LLMs are becoming cost prohibitive, you are probably not running them in a cost effective way.
Madhav has a few strong feelings around what organizations should be doing with LLMs. The first is that most should not be trying to train their own LLMs - the time and cost just don't make that much sense when open source models are advancing at frankly remarkable speeds. The second is that GenAI is probably more helpful for data producers than data consumers. It really can make producers far more productive, e.g. letting them generate insights on their own data or helping them to find good interoperability points with their data and other data products.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Emily's LinkedIn: https://www.linkedin.com/in/emily-gorcenski-0a3830200/
Amy's LinkedIn: https://www.linkedin.com/in/amytobey/
Alex's LinkedIn: https://www.linkedin.com/in/alex-hidalgo-6823971b7/
Alex's Book Implementing Service Level Objectives: https://www.alex-hidalgo.com/the-slo-book
In this episode, guest host Emily Gorcenski, Head of Data and AI for Thoughtworks Europe (guest of episode #72) facilitated a discussion with Amy Tobey, Senior Principal Engineer at Equinix and Alex Hidalgo, Principal Reliability Advocate at Nobl9. As per usual, all guests were only reflecting their own views.
The topic for this panel was applying reliability engineering practices to data. This is different than engineering for data reliability which is focused on data quality specifically.
The overall concept is taking what we've learned from reliability engineering across disciplines but mostly in software, especially SRE/site reliability engineering, and bringing those learnings to data to make data - especially data production and serving - more reliable and scalable. Scott note: this is probably one of the most frustrating topics in data for me because it feels like it's basic foundational work yet most organizations aren't tackling this well yet if at all really. The best starting point for an organization is simple awareness and starting to have reliability engineering conversations around data. And you will probably feel like you're behind after listening to this. Everyone is behind on this 😅even most orgs aren't doing SRE well so applying it to data, that's no surprise.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Corrin's LinkedIn: https://www.linkedin.com/in/corrin/
In this episode, Scott interviewed Corrin Shlomo Goldenberg, Senior Product Manager of the Data Platform at BigPanda.
It's important to note that BigPanda is not at the stage yet where data mesh makes sense but this is a story of getting production of data into the heads and hearts of the application development team, which is a crucial aspect to doing data mesh well, whether it's done pre data mesh or as part of the journey.
Some key takeaways/thoughts from Corrin's point of view:
Corrin started with the tale of BigPanda and how she started building out their data, ML, and analytics capabilities. When she came in, they didn't have the infrastructure or really the focus on a scalable platform for storing and analyzing their internal data. They were doing a lot of this for external clients but hadn't moved to doing it internally, which is pretty common in B2B startups. But BigPanda wanted to do a data driven transformation of their business model so they had to change the situation around their internal data.
There is always a balance for when you start collecting data at scale in Corrin's mind. At a B2B startup, you need to ask how early should it be for the company but the same is applicable for an early-stage offering at a larger organization. Most development teams aren't tasked with dealing with creating the necessary data until far later in an offering's lifecycle but it would be nice if you could include it at the start. But it definitely isn't free so there is always a balance and the conversations need happen, hopefully earlier than later.
Corrin's tipping point for when you should really start to press development teams on creating necessary data is when it becomes hard to answer simple 'how many' type questions. It is also an easier conversation than a hypothetical one. If it takes more than a day to get basic information on how your customers are using your product, that's obviously an issue that's only going to grow. It's also a pretty tangible place to start.
When they started to build out the data platform, Corrin said it just made sense to start centralized. If the R&D team wasn't really thinking about data, trying to upskill them enough to take over the work entirely was probably a bridge too far. Plus, if your data requirements aren't complex enough to require decentralization, decentralization is often just an extra layer of complexity. So they moved to a high communication model where people can see what data work is happening even if it's controlled by the central team. They can slowly upskill the development teams to understand data instead of trying to hand over ownership prematurely.
Corrin talked about working with the team to understand the product mindset to data. Start from the why - it's easy to fall into the trap of trying to do everything because it might have value. That's what happened with data lakes that became data swamps. Focus people on the why and you can bring them more and more into working with data.
Similarly, while Corrin and team didn't have a lot of pushback on getting things done, she was very cognizant of prioritization and cost/benefit. Again, focusing on 'the why': what is most important and when? Why are the requirements like this? Can we cut the cost down by storing for less time and/or refreshing less often? When you say 'real time', what do you actually mean? Etc.
Corrin has been seeing good results from having strong ownership conversations. While the central team still owns the data, they are partnering with the domains as the domains still need to own the concepts and the understanding of the information. While this might not work at a large scale, it's perfectly normal and functional at a 300 person company. Scott note: centralization isn't the enemy until it becomes a bottleneck 😎
As with all global companies, BigPanda has some challenges around communication, per Corrin. Time zone differences and of course differences in focus are just two of them. So she recommends spending a lot of time to communicate to stakeholders about what you are building and why. It's easy to assume that because you build out a data product, people will use it but you have to work with people to ensure they actually use what you built.
Corrin pointed to the fact that many companies in the B2B space feel they aren't "data oriented" enough. She gave a few tips for how to become more data oriented but also has empathy for people feeling that - it's pretty common, most B2B companies feels they aren't as data oriented as everyone else. Similar to data mesh, where everyone believes all the other companies are far down their path. It's simply optics - companies project a better image than the reality of their situation with data.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key Points:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Jimmy's LinkedIn: https://www.linkedin.com/in/jimmy-kozlow-02863513/
In this episode, Scott interviewed Jimmy Kozlow, Data Mesh Enablement Lead at Northern Trust. To be clear, he was only representing his own views on the episode.
Also, FYI, there were some technical difficulties in this episode where the recording kept shutting down and had to be restarted. So thanks to Jimmy for sticking through and hopefully it isn't too noticeable that Scott had to ask questions without hearing the full answer to the previous question.
There is a lot of philosophical discussion in this conversation but tied to very deep implementation experience. It is hard to sum up in full without writing a small novel. Basically, this is one to probably listen to over just reading the notes.
Also, Scott came up with a terrible new phrase, asking people to "get out there and be funky."
Some key takeaways/thoughts from Jimmy's point of view:
Jimmy's role is a rather unique one in that he is literally tasked with enabling data mesh to happen and make sure they are focusing on the right things. That means a myriad of different things but a lot of it is ensuring the communication and collaboration happens where it needs to while still focusing on the big picture. It might be like an American football coach that coaches the entire team but is still calling plays and making the minute decisions during the game too. It's a big set of tasks to take on.
At Northern Trust, they started their mesh implementation with the innovators according to Jimmy. That's a common pattern for all tech innovation, not just in data mesh - find the people enthusiastic to try something new so you don't have to spend half your time driving buy-in. There obviously also had to be a need where data mesh could work and would be a differentiator.
For Jimmy, one big complexity factor he is seeing is around data products. He believes that you really start to drive value of a data product higher the more relevant and interoperable data sets you can include in the data product - within reason of course. But as you add that fourth or fifth data set, it gets complex to maintain interoperability and consumability even at the data product level. But that complexity is where a big part of the value really lies in data mesh - that it pushes organizations to take on that complexity but in a scalable way.
There is a bit of a push/pull in data mesh for Jimmy: organic data product growth from new use cases versus the centralized team pushing for the value from more and more interoperable data - basically creating data products that fill the gaps between existing data products to enable additional use cases. New use cases emerge the more data products you have that are well crafted and that link data across domains; but the question becomes do you push for new data products that don't have a specific use case - inorganic growth - in order to create a tipping point where those use cases can quickly emerge? It would mean less work for consumers and faster time to market if the data is already available instead of working with the producing team. But will the data products be made well enough to serve use cases? Are people ready to go and discover pre-made data products instead of ones tailored to their needs? Should we only apply this to new data products - how would a central team spot when additional attributes are needed to take an emerging data product from only serving a use case to being more globally valuable? It's still a question in data mesh as to whether to pursue inorganic data product creation.
At the start of their journey - and even though they are two plus years in, it's still relatively early - Jimmy and team are focusing on existing, known use cases as they build out the existing available data and improve their capabilities. Capabilities not just to deliver new data products but how to deliver incremental data that fits well into their already existing set of available data on the mesh. It's about building out the entire picture instead of focusing too much at the micro level. Scott note: balancing that micro and macro level is hard - extremely hard? - but the earlier you get good at figuring out how to add value at the overall mesh level while serving use cases, the more value you will deliver with each incremental data product.
Jimmy talked about with data mesh, even though we can deliver scalable data products relatively quickly, there can be more initial friction for new use cases. E.g. teams have to go and collect the necessary information to do the governance well instead of trying to add the governance at the end. And some might be frustrated or not bought in that the upfront friction is worth it. So he's trying to show the value far exceeds that initial extra friction but people will of course resist new ways of working. Such is the nature of working with humans 😅
When asked what does success look like for them, Jimmy pointed to showing value. It's important to note that isn't merely delivering value but being able to show that value. As Jerry McGuire said, "SHOW ME THE MONEY!" Being able to show value generation helps to build momentum and adoption. If you are proving value then it's not nearly as difficult to get incremental investment. Teams want to participate and capture value too. Excitement builds. But it's also important to note that what success looks like will change - maybe not wholesale but at least in part - in different phases of your implementation.
For Jimmy, there are two important aspects of your implementation, essentially the setup and the knockdown. You have to set up your implementation for success by building out the platform and capabilities but getting teams to actually adopt is still crucial. Just because you built something amazing, you still have to work with people to understand the shift in mindset and approach to get them to buy-in and adopt. A great platform that no one is using isn't really a great platform…
Trying to keep momentum of the whole mesh implementation until you reach critical mass is very challenging. Jimmy talked about how trying to get teams with quite different capability and speed levels to work together can be hard. It's not as though any organization is built to all move together as one - it's not a car that's built to move as one unit, it's more like a group of cars - so you need to really focus on the coordination, collaboration, and especially communication. ABC - Always Be Communicating 😎
Jimmy believes that the central data team in data mesh should be a key point of leverage. They can jump in to help a domain early in their journey to get something delivered while raising that domain's capabilities. A central team can bring repeatable patterns to find easy paths for new domains. That way, the domains still learn by doing but they don't have to learn by repeatedly failing. But once a domain is capable enough, the central team tries to move out quickly to give the domains autonomy to do what's valuable. They are also building a community of practice for practitioners to share insights with each other, providing even more leverage and discovering more repeatable, high value patterns.
When asked about bringing a domain up to speed and the question of complexity, Jimmy strongly believes that you should start simple. You don't hand the 7yr old who wants to help you cook the knife and have them go wild on day 1. Get them into a groove, get them confidence and understanding how to deal with data, then you can start to think about adding complexity. But dealing with data is complex enough, keep it simple for them to deliver value initially.
When thinking about success and things like a new domain's time to first data product, Jimmy believes you can learn where there is friction but the actual times vary quite a bit. So you might see lengthening times for new domains launch a data product but it's a good thing because you are dealing with domains who really aren't sure what they are doing and have to learn a ton about data in general and their own data - that means you are penetrating the less data savvy parts of the organization. All things equal, you want to go faster but just take things with a grain of salt.
There's also the fun of trying to thread the needle of data modeled for the initial use case - so fit for purpose - yet also modeled so it fits well and is interoperable with the rest of the data in your available mesh of data products. Jimmy said this is especially true of domains just getting up to speed with dealing with their data and data modeling so prepare for them to move slower and need more help.
When asked about where there is friction in the process of bringing on new domains that we _shouldn't_ try to reduce, Jimmy pointed to learning and understanding. People need to take the time to understand how to do data work and all that but also understand the new ways of working and why the organization is going in this direction. It's that old are you trying to get them to do the steps you say or achieve the target outcome you give them. Learning takes time, don't rush it.
Jimmy's role is pretty unique as far as other organizations telling their story. He's focused on helping people focus in the right areas but also on helping them connect the dots. There is such a big picture when you think about the entire information scape of an organization and helping people to connect to each other and see where they could enhance that bigger picture is highly valuable. This drives better value while also reducing miscommunications, duplication of work, and wasted time. Scott note: this is somewhat similar to the 'Data Sherpa' concept I've mentioned repeatedly that just about everyone looks like I'm a madman when I bring up.
The question of what to decentralize versus centralize is a tough one for every organization doing data mesh. Jimmy pointed to the fact that where
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key points:
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Sid's LinkedIn: https://www.linkedin.com/in/siddharthin/
In this episode, Scott interviewed Sid Shah, Head of Product Data and Analytics at Airtel, a large India telecom operator. To be clear, he was only representing his own views on the episode.
Before we jump in, it's important to note that Airtel are doing data mesh on 'hard mode'. Because of regulatory requirements/restrictions, they are all on-prem. That means extra challenges when it comes to securing compute resources.
Some key takeaways/thoughts from Sid's point of view:
Sid started by talking about some of the challenges of data and analytics in a very large organization that has grown through acquisitions - data is siloed across the organization and you can't get sight of the customer journey across lines of business. To address this, a few years ago they started landing data from across the organization into a central data lake. You can probably see where this is going - a central team getting more and more overloaded while there is an ever increasing demand for data. Add to that a data platform where the only teams sophisticated enough to use it being in the central data team and you have major bottlenecks, poor quality governance, the business teams weren't sure what to trust or what data really meant, all combining for a bad analytical experience. So that was the stage for why they started to look at data mesh.
Even before coming to data mesh as a potential solution, Sid and team thought the biggest problem was poor data ownership. The teams who knew the data couldn't own the data even if they wanted to - they didn’t have a platform they could actually use or the data fluency to own their data work. The other issues could really be traced back to that - poor quality, bottlenecks, low trust, etc., it all came back to poor data ownership.
According to Sid, coming across data mesh and all the other organizations that were implementing - including the stories from this podcast and the community - validated that they weren't the only ones experiencing these data and analytics issues. And it gave them not just validation but hope and a set of talking points and proof points to leverage internally for driving understanding and buy-in.
Sid and team understood early on that changing data ownership is not a "trivial exercise". It's not just changing the responsibility, you have to enable domains to actually be able to properly own their data or it's just the same problem of bad data and bottlenecks but just new people to point the finger at. They decided to not try to solve challenges at the small scale but come up with an approach that was going to really target the pain at the enterprise level, so that's why they arrived on data mesh.
Two things that helped drive data mesh forward as a concept at Airtel were a general understanding of data mesh and its aims at the senior leadership level in the organization and again feeling that pain for long enough that people were ready for a fix. A third aspect was creating a way for data mesh to actually happen at Airtel specifically - how could data mesh work in their organization without trying to change everything about the organization and how people/processes worked? Essentially, how can you explain how this would all work to people at their level of involvement - senior leadership doesn't worry much about what tech your platform uses - and that your mesh implementation won't disrupt the entire organization in a bad way while addressing a lot of the causes of data pain.
As a number of guests have pointed to, when speaking to the business, Sid recommends to speak to their business pains. Data quality issues, not being able to nimbly react to the market based on data, reliability issues of things breaking, etc. The business wants to deliver on their objectives, not talk about how you are building the platform. Think about time to experimentation as well - if you can take the time to deploy an experiment in the market from weeks to days or even hours, what can that do for them? The product and engineering orgs love the sausage factory tour so keep it to just them :)
When looking to drive buy-in and understanding of data mesh with the technical teams, the engineering and product teams/leaders, Sid brought in a number of people from other organizations doing data mesh. People got used to the idea of data mesh and also, if they are seeing many internal presentations on data mesh, they understood the momentum was building so they'd have to get on the bandwagon. Scott note: this is part of why Data Mesh Understanding exists - not everyone can easily bring in those outside speakers.
When they were just starting out their journey, a number of domains gave the quite fair pushback of they didn't have the skill sets to do data mesh, to own their own data. They also didn't have the resourcing to take on the additional work, both people and compute - recall, Airtel is all on-prem. So Sid and team had central resources - including people - to loan to those domains who wanted to own their data, who were bought in, but needed that help. Another aspect was to find champions and develop hero stories, where someone not that advanced around data was able to successfully build a dashboard or similar to inspire others and lower the perceived bar to doing the data work themselves.
Sid and team knew that if the domains couldn't successfully own their data, the work would fall back on his team. So they had a bit of an extra incentive to make it work 😅 So they made sure to only take on the amount of work they could handle instead of onboarding every domain they could/that was interested right at the start. Just make sure you communicate effectively/transparently. Scott note: It's absolutely normal and reasonable to push domains out in time if you don't have the capacity to work with them, whether that is resources or platform capabilities around their use cases.
At Airtel, Sid and team didn't succeed on their first attempt at data mesh - in fact they "failed miserably". They weren't fully prepared and didn't have a lot of the issues figured out. Luckily their company didn't throw in the towel but the failure helped them understand what really mattered. It was about making data mesh possible, making it a part of the strategy so teams could feel like it was okay to lean in, and building out the resourcing - platform, compute, and people sides - to make this an actual possibility.
Sid didn't see his early data mesh role as trying to convince every domain, win over everyone upfront. It was about getting out the first few use cases and working with a few domains. It was more about finding collaborators that could help drive to a scalable platform. It was about finding the friction points as domains look to own their data and then addressing those friction points. Really focus on what you need to accomplish and don't try to take on the world :)
It's going to be important - and somewhat challenging - to keep all your stakeholders engaged at the start of your journey according to Sid, especially those not participating early. It's important to keep engineering and product leaders engaged but also the people doing the actual work in the domains and the data team. And of course, you need to keep the business stakeholders engaged as they are the ones who will do a lot of the work but also get the biggest benefit. But it's also crucial to understand each group will require different approaches to keep them engaged. Focus on that incremental progress and find leverage points instead of trying to tell them about everything. For Airtel, the product and engineering teams were excited about faster product delivery and quicker delivery of insights.
For Sid, it was important to communicate to leadership that the data mesh journey would take a few years and to get their buy-in that this was a long-term approach, not a quick quarterly fix. But most teams really only care about what is going to be coming, what value will be delivered, in the next quarter or two. Again, focus on delivering updates about what the audience cares about and understand you will have multiple personas to cater to.
At Airtel, budgeting, especially around data, has been done centrally and not in each domain according to Sid. So, for the first year of their data mesh journey, the data team isn't even showing people exact budgets/costs for their domains around their data work but they intend to build a perspective on overall expected costs. This is a big change in thought for the domains from when the costs were all on a central shared data team. The data team didn't want to switch from essentially an afterthought to billing the domains too quickly but they do want the domains to consider costs.
The return on investment conversation has been somewhat difficult to nail down in many use cases for Sid. What value do you attribute to faster time to market? If something now takes 1 day instead of 4 weeks, there is a very clear business value BUT what is the measurement of that business value? Time savings is nice but ends up being pretty low in the grand scheme at Airtel. There are also new use cases that could not have been done before. But overall, it's hard to point to an exact return figure. Scott note: this is where you can flip this around on the lines of business and ask them to tell you how much value it generates for them to do a use case or to get to market quicker/easier and do smaller scaled, fine-tuned experiments.
Sid and team are still figuring out things like the domain onboarding model and many governance aspects in the platform and out. Should they have an embedded expert in things like data modeling go into domains to help them get up to speed? Now that sharing data is far easier, how do you make sure data is well governed across a number of dimensions like quality, security, privacy, etc.? It's not all figured out and that's okay :)
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Andrew's LinkedIn: https://www.linkedin.com/in/andrewsharp27/
Nicola's LinkedIn: https://www.linkedin.com/in/nicolaaskham/
Kinda's LinkedIn: https://www.linkedin.com/in/kindamaarry/
Jay's LinkedIn: https://www.linkedin.com/in/jaycomoiii/
In this episode, guest host Andrew Sharp, Principal Consultant - Data Governance & Data Protection at The Oakland Group (guest of episode #172) facilitated a discussion with Kinda El Maarry, PhD, Director of Data Governance at Prima (guest of episode #246), Nicola Askham, AKA The Data Governance Coach, an independent data governance consultant (guest of episode #129), and Jay Como, Strategic Advisor at Curate Insights (guest of episode #92). As per usual, all guests were only reflecting their own views.
The topic for this panel was how do we do data governance well including how do we get started around data governance in data mesh. There's a lot to learn about how to improve your governance but there are no blueprints unfortunately. You have to do the work specific to your organization.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Amy's LinkedIn: https://www.linkedin.com/in/amy-tang-edwards/
In this episode, Scott interviewed Amy Edwards, Formerly Director of Analytics and Product at Vista. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Amy's point of view:
Amy started the conversation around measurement in data. To become data driven, you have to actually measure things. It doesn't have to be the perfect measure but you need to track and understand your progress. In fact, it almost certainly won't be a perfect measurement and your measurements will evolve over time too as you want to focus on different aspects of a data product or mesh journey. In her own journey to get back in shape, she focused early on exercise quantity over quality because it was about forming the habit. After she saw success there, she started to focus more on performance and speed of her runs.
It's also important in Amy's mind to understand what are you actually trying to achieve in a certain phase of any implementation. For data products, there is the development phase, once it goes into production, maturity stage 1, etc. What looks like success or is important will change and evolve in each phase and as you progress in your mesh journey. That evolution is not driven only by that your measurement framework gets better, it's that what you want to even focus on changes. So, starting with number of users of a dashboard was a good initial metric = it was early phase, relatively easy to track, etc. So as a data product matures, what you track should change.
For Amy, even in a more data literate environment, it can be a pretty hard adoption/learning curve for consumers with each individual data product. So, you want to try to get more users early in a data product's life to encourage developing new use cases. But people don't just generally start using a new data product so someone needs to do more handholding and work than just creating the data product. Part of that might be the data team or data product owners doing more of the interconnection work between data products so users can understand and be confident - otherwise, they will probably stick to what they already know. Focus on increasing and deepening engagement, that's what drives your culture towards being data driven.
Amy talked about how there is a bit of a push-pull when thinking about those better interconnections between data products. A product owner or manager should be thinking about how can their data product better connect to the greater data landscape inside the company. And then should that be pushing data from their data product into someone else's or pulling in data from another data product to make their data product more valuable. Find those paths of least resistance for consumers to get useful information to drive increased adoption.
For Amy, there is an interesting challenge when it comes to balancing high reusability versus customization for data products. If you do not create the demand for data with the consumers, then even the most reusable data product won't see a lot of use. So they focused more on customizing data products to use cases than is probably general practice in data mesh. But that meant the users had an easier time adopting and leveraging the data products. You can't create too many highly customized solutions in the long run, you just can't support that many data products. But data products that are perfect in theory with very low consumption is probably far worse than a situation where you will need to refactor in the future but you have people eagerly consuming data. It was still manageable to use this approach quite far into their mesh journey.
At Vista, there was some duplication of work in Amy's view, where multiple data products might have some of the same data. But, she viewed that as a conscious choice and that it's better to have the data more easy to consume than to really put the hammer down on not have duplication of data. Again, they really focused on driving usage with an eye on improving things later. Scott note: this is an interesting note and balance to strike. I honestly don't know how I feel. If you're aware and it's communicated and people understand, as long as it's not causing issues, I guess it's okay?
One key metric for Amy along the early drive to becoming data-driven was how often people were self-serving their analytics. Essentially, how often were they going and looking at dashboards or outputs of some kind to make their decisions rather than leveraging a data analyst to answer questions for them or worse, waiting for a data analyst to come tell them the insights. It can be hard to measure specifically but you can get a decent idea of momentum around that kind of metric.
A really interesting insight Amy had was that when they first started trying to work with the broader organization, many of the people were not yet confident with their analytical capabilities and providing them a ton of flexibility around their analytics dashboards overwhelmed them. So, instead, they started by providing more static views to let people learn how to consume data and analytical information before giving them more advanced self-service tooling. And to get the general business users more comfortable, they paired those business folks up with data analysts to walk them through what data existed and how they could query it and leverage it. It let the business users get comfortable in the shallow end, not throwing them in the deep end.
Amy talked about how the role of the data analyst might not be the most glorious when you are in a transition phase. They were using data analysts as change agents, helping to train the business people to get better and better with data. It would be great if the data analysts didn't have to do that but there is that necessary transition period where people are improving their data fluency. After getting the general populace to a better data fluency, the analysts could focus more on higher-value, more in-depth analysis.
When asked about how a data product team should look and function as part of the greater organization, Amy talked about how an end-state, highly data-driven organization might look versus how you will probably start out. If we want to manage data as a product, we have to think of data management and generation at the domain level as just part of the product function/software development. That may still be a separate data product team within the domain but it's not a completely separate unit/function. But when you are starting your data mesh journey, you almost certainly will not be at that capability level from a platform standpoint or general developers being that data literate - there will probably be a need for a separate data product team that is led more from the data side than the product side. But with that setup, beware of silos even within the same domain, try to make sure the data product teams are integrated into the work of the domain.
Amy talked about the number one most important aspect to a successful data product or use case: an engaged consumer stakeholder that is strongly aligned. It's very hard to manufacture demand by creating the most attractive data product in the world. You want a stakeholder who is helping drive the development decisions and is chomping at the bit to get the data. Don't let the data product teams develop in isolation either. Make sure they are iterating with the stakeholder.
When building out your data mesh, Amy recommends focusing a lot on the foundation first. You will be building out some unsexy data products, really the core of what you'll build on later, as the early part of your journey. But that means you can focus more and more on building to value later. Of course you need the buy-in, funding, and momentum to go this route. Crawl, then walk, then run.
Quick tidbit:
When building out a data product team, Amy recommends your first hire being a data engineer. They are the ones who "build the rough plumbing", the pipes in the walls and without that, the water doesn't move anywhere throughout the house, in or out.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Takeaways:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Paul's LinkedIn: https://www.linkedin.com/in/paul-cavacas-32a36158/
In this episode, Scott interviewed Paul Cavacas, Senior Manager of Data and Analytics at Ocean Spray.
Quick note before jumping in: Ocean Spray is just at the beginning of their journey - in their pre-implementation phase - and there hasn't been a lot of resistance yet internally. That might make a few people jealous 😅 but there's a lot of interesting things Paul is doing to ensure that they are ready to decentralize what makes sense to decentralize at the right time. There is a lot to be gained from not rushing in. Also, apologies that Scott's audio is a bit weird, he had yet to build his makeshift sound studio in the Netherlands.
Some key takeaways/thoughts from Paul's point of view:
Paul started off with a bit about why they are headed down the data mesh path. For a large internal project, Paul had to become an expert on so many aspects of the company and that's just not scalable in the long-term or if he's on vacation. So, he's started to decentralize the data capabilities - slowly - as teams understand what data they will likely need to own in the long-run. And he's not playing data ownership 'hot potato', he's making sure they are prepared in the right ways.
At Ocean Spray, Paul shared that until recently everything tech - including data - was very centrally owned. In some areas, the IT team knew possibly more about the business processes than even the business people in those domains. So the company is going through all of their software and applications to decide how that should look in the future. Data mesh plays well into that rethink because central ownership scales until it doesn't and limits flexibility.
As there isn't a rushed timeline, Paul has been able to put together a complete idea of what the data mesh roadmap will look like. But he also understands that it could look completely different as he learns more and starts trying to actually implement different aspects. There are some existing data sets/assets out there that could pretty easily become data products in the right environment so that is where they are targeting first. They are working with the teams to transfer some part of the ownership, especially around documentation of use cases and SLOs. The central team is pairing to take existing data assets, decompose those into their data products and help people get on the path to real ownership.
Paul recommends what Brian McMillan talked about in depth in episode #26: finding people within domains that are at least somewhat tech savvy and want to advance their careers. Work with them to get them more and more up to speed. Ownership is not something that gets transferred in a day - treat it with more respect than that. So that's finding receptive people inside the receptive domains. Yes, it won't always be easy but why make the buy-in complicated at the start if you don't need to?
Right now, Paul is building out some of the technical underpinnings of the platform they plan to build. If there are teams that want to move more quickly, they can start to test things out now. As long as those teams understand things aren't fully automated and they may have to change things about what they build now when the company starts to fully move to data products. One big piece he is anticipating is the need for testing and data contract mechanisms. But exactly how to do that is still a challenge and will be learned along the way. He's anticipating a workable but not perfect solution to start. Build to useful and then improve.
Paul circled back on the idea of finding the right partners over the right use cases/domains. Having engaged and excited partners, who know you can up their own data capabilities and drive value for them too, will make your early journey far easier than going for the most "valuable" data. You are also likely to get better feedback because they are bought in to collaborating with you! To find those partners, potentially look at how teams present their results internally. If they are presenting with lots of advanced figures and almost a flair around data, that is great sign.
How much data ownership/work gets decentralized and when is a key remaining question for Paul. He's aware that he'll have to test what works and iterate as he learns but there are plenty of domains that are too small to justify them learning a ton about how to own data when there just isn't that much data/data work to deal with. There will be a shared ownership model between the central team and the domains. Scott note: this works up to a certain scale and in certain types of organizations. Shared ownership in a very large organization rarely works that well for all that long - too much political infighting and challenges but it's an interesting pattern for smaller orgs that seems to be working well.
Paul's plan for assessing the quality of data products is to create a rubric scoring system - asking people to rate them across multiple dimensions like usability, data quality, SLA compliance, etc. And that the scores or how they are measured may change across time. At the start of a data product's life, when it's still in beta, those scores can be invaluable to iterate towards value but then consider throwing the historical scores out once it hits that v1.0. That's because there is a useful aspect of feedback depending on what you are trying to achieve and bad historical scores could hinder the success of a data product when it is now very high quality and valuable.
For Ocean Spray, their first few data products are going to be source aligned, combining a lot of important sales information. That way, those people who want raw data can still get at it but then they can build out more and more views/data products for the users on top of those. That way, there is still the scalable/productized underlying production of the raw data and then more fit-for-purpose outputs for the different users.
Paul is not letting perfect get in the way of progress. Data contracts have to get to a place where we aren't locked onto schemas as something that can never change. But no one has really come out with a better solution yet, so that's what he's doing to start. It's better than nothing so go with it while you figure out better ways.
Paul finished with a bit of advice around working with a few domains at the start of your journey. That way, you can take the learnings and understand the needs from multiple domains to abstract to a better solution for the organization rather than one overly tied to one domain's needs. Scott note: people seem pretty 50/50 split on working with one domain or 2-3 at the start of your journey. It's an interesting question.
Other unique factors of Ocean Spray:
The corporate structure is a co-op of growers so there isn't some massive pressure to grow at all costs.
Domains have been able to get access to other domains' data relatively easily for a long, long time. It hasn't been cleaned and prepared for them but there is an existing culture of sharing.
They are moving more and more to 3rd party applications rather than custom-built, which means data isn’t necessarily in an easy to consume format by default. (maybe not all that unique?)
Because many domains are quite small, the central team will likely still own most if not all of the data work for those domains.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key takeaways:
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Use code DATAGOV23 for 35% off ebook copies of Designing Data Governance from the Ground Up here: https://pragprog.com/titles/lmmlops/designing-data-governance-from-the-ground-up/
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Lauren's LinkedIn: https://www.linkedin.com/in/laurenmaffeo/
Designing Data Governance from the Ground Up (Lauren's book): https://pragprog.com/titles/lmmlops/designing-data-governance-from-the-ground-up/
In this episode, Scott interviewed Lauren Maffeo, author of the book Designing Data Governance from the Ground Up and adjunct Lecturer at George Washington University. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Lauren's point of view:
Lauren started with a somewhat common refrain for this podcast: the pace of data practice maturation - around governance and data practices as a whole - is just not keeping pace with innovation in other aspects of software. Even the pace of conversation is not maturing as fast. Cybersecurity is maturing very quickly for example but we're just not seeing that in data. So companies are just not ready to really derive a lot of value from things like machine learning (ML) or natural language processing (NLP).
One of the big issues around industry maturity and data governance for Lauren is that there isn't even a large community around the topic. So there isn't a larger cohesive conversation around data governance best practices. Scott note: it's really hard to have a broader conversation too because approaches and practices vary widely and data governance has about 15 varied subtopics that each deserve their own focus rather than being lumped under a huge umbrella of 'other' that governance has become.
Most cybersecurity breaches, Lauren pointed out, are caused by internal employees making a mistake. So how do we think about that relative to data? Is that people creating low quality data and/or is it people not being data fluent enough to actually make good decisions based on data? In cybersecurity, there is a big emphasis on training people to see what an attack looks like - should we take the same approach in data regarding bad quality data? "You really have to embed data literacy into very strategic ways of communicating with the organization and educating them that way. Without that approach, I think very little progress can actually be made."
Lauren talked about how few companies are really going broad with their data literacy programs, training up a large number of their employees. There is a lot of talk about that as part of data governance programs but few are walking the walk. And she believes it's not that hard to get people to a relatively data fluent level - understanding SQL, being able to more easily spot low quality data, etc.
"We'll do the data governance later," is something Lauren has seen and heard in conversations. Governance is seen as something that can be layered on like a coat of paint at the end of a car being manufactured. But because good governance is intrinsic to data quality and matching to the actual business use case, trying to do it later rarely leads to good results.
When asked about selling the return on investment of data governance work, Lauren admitted that it's often quite nebulous but data governance is so key that people know they need it despite not being super clear on the specific value of the work. And you can roll out your data governance tech, policies, and processes at a reasonable pace, creating some definitions and a sandbox to show people how it will work. She is really big on the idea of a sandbox to get people used to new governance practices and tech. It isn't as though everything changes suddenly, it's that you're working towards better data practices that will drive value for the organization. Fail fast is "the essence of innovation in tech" so we need to embrace it far more - but still safely and sanely - in data.
"We also can't afford for leaders of any department to not know what quality data looks like for their teams, because their success, the success of their teams depends on having quality data that their customers trust," Lauren said. So we all need to be in this together and have domains really owning and understanding their data. That can't be on a central data team.
Gamification is one thing Lauren is seeing work for improving data literacy/fluency. It is a great pathway to creating a data-driven culture. Make it fun and give out rewards :)
Lauren wrapped on a simple message. Automate your standards. It is easy to have your tech and standards/processes quickly lose alignment if you aren't making things easy for people via automation.
Learn more about Data Mesh Understanding: https://datameshunderstanding.com/about
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this rerelease of episode #130 with my partner in our weekly data mesh roundtables Jean-Georges Perrin. There are a lot of interesting things to take away from this. A biggie is to have an early thesis about what to drive towards - what will drive value early? Doing data mesh doesn't simply create value. And you need to build momentum. There's a lot here to learn about how to apply good software engineering practices to data with data mesh.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Data Mesh at PayPal blog post: https://medium.com/paypal-tech/the-next-generation-of-data-platforms-is-the-data-mesh-b7df4b825522
JGP's All Things Open talk (free virtual registration): https://2022.allthingsopen.org/sessions/building-a-data-mesh-with-open-source-technologies/
JGP's LinkedIn: https://www.linkedin.com/in/jgperrin/
JGP's Twitter: @jgperrin / https://twitter.com/jgperrin
JGP's YouTube: https://www.youtube.com/c/JeanGeorgesPerrin
JGP's Website: https://jgp.ai/
In this episode, Scott interviewed Jean-Georges Perrin AKA JGP, Intelligence Platform Lead at PayPal. JGP is probably the first guest to lean into using "data quantum" instead of "data product". JGP did want to emphasize that as of now, he was only discussing the implementation for his team the GCSC IA (Global Credit Risk, Seller Risk, Collections Intelligence Automation) within PayPal.
Some key takeaways/thoughts from JGP's point of view:
JGP started the conversation talking about how in his team, he's really leaning into the idea that software engineering and data engineering are not that different. Zhamak has discussed this too. We should focus on sharing practices so we all create better software and infrastructure. For JGP, data engineering work in most organizations has followed a very waterfall approach. However, his team has been mostly working in an Agile manner. Therefore it wasn't a huge switch to their ways of working - like it is at many organizations - once they started doing data mesh. And luckily, there was already an appetite for changing the way they were tackling data challenges.
In the spirit of being agile and capital A Agile as well, PayPal set out on their data mesh journey. They wanted to do an MVP but what was the P? Minimum Viable Data Product/Quantum? Minimum Viable Platform? Both? Minimum Viable Mesh? JGP recommends looking at what you want to deliver as a minimum unit of value. PayPal already had extensive data platform expertise so they were able to focus on delivering data products/quanta (plural of data quantum) but they worked in parallel to build out their initial data quantum and mesh capabilities. As many guests have noted, it's dangerous to only do a minimum viable data product/quanta.
PayPal has been building data platforms for a long time. As mentioned by JGP, they were one of the pioneers of the self-service data platform concept. But data mesh offered a path to faster and easier data discovery, to making it easier to use data in a governed way, and to increased trust in data by the data consumers - their first consumers being data scientists. A big benefit of addressing those needs is those data scientists are able to better tell if the data they access is the right data for their use case.
One thing JGP emphasized that's significantly helping PayPal move forward is standardizing APIs across data quanta. Those are not data access - or analytical - APIs as JGP thinks those will just never work all that well. Instead, as their audience is data scientists only to start, everything anyone needs other than the actual 1s and 0s of the data is accessible via Python APIs. The metadata, the observability/trust data, etc. Then, the data scientists use notebooks to work with the data. But standard APIs means data consumers only have to learn one interface. This is similar in concept to what many are doing with data marketplaces - one standardized way to interact with the information about the data quanta.
PayPal is using the terms data product and data quantum as two separate things. A data product is simply a product powered by data and analytics. Those have been around for quite some time. But PayPal is looking at data quanta like side cars, used specifically to power more and more of their data products going forward.
PayPal have invested heavily in making data contracts work well per JGP and earlier PayPal guest Jay Sen. They've been building APIs to make it far easier to consume data contracts as people learn more about a data quantum. And as mentioned before, they can consume observability metrics via API as well. When asked about how are they setting their actual contractual terms, the data producers initially put out some contractual terms and then may adjust those terms as data consumers request. It's important for data producers to not set their data contract obligations too strictly unless there is a user-based need.
JGP made the good and often unspoken point: the term domain has lost a lot of its meaning. It can mean a very high-level domain like Marketing, Finance, Sales, or HR. Even in software companies, a domain could be Product. But at PayPal, they are being quite strict about what they mean for domain in data mesh: it is a small scale sub-domain - think two pizza team size - and they enforcing a strict 1:1 relationship of one data quantum per domain; and of course, not cross domain source data quanta too. That way, each small domain can focus on creating a great data quantum instead of worrying too much about how big each data quantum should be. The scope should never get that huge at a two pizza team size.
Back to APIs, PayPal is implementing an API-first approach. APIs for the data quantum control plane, observability APIs, and data discovery APIs. It's the preferred way of working for their initial consumers - data scientists. However, as mentioned previously, JGP does not believe analytical APIs - that is APIs designed to do things like filtering and returning many hundreds to thousands or more results - are really feasible. Definitely not now and possibly ever. So APIs are great for getting at the metadata but not the data for analytical use in his view.
JGP wrapped up in sharing how our tooling must evolve so we don't have to think about such a hard wall between analytical and operational. There will always be analytical and operational workloads but our systems can evolve to support both. We aren't there yet though.
Quick tidbit:
If you are just delivering data, the 1s and 0s, you are not delivering the necessary trust.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this rerelease of episode #65 with Abe Gong all about how people are implementing data contracts in the wild. There are so many ways people can just do only defensive data contracts and I think that is such a missed opportunity. Maybe it's where you will have to start but there's a much better way and we talk a bit about why I think that is so distressing that people aren't talking to each other.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here
Abe's Twitter: @AbeGong / https://twitter.com/AbeGong
Abe's LinkedIn: https://www.linkedin.com/in/abe-gong-8a77034/
Great Expectations Community Page: https://greatexpectations.io/community
In this episode, Scott interviewed Abe Gong, the co-creator Great Expectations (an open source data quality / monitoring / observability tool) and co-founder/CEO of Superconductive.
One caveat before jumping in is that Abe is passionate about the topic and has created tooling to help address it. So try to view Abe's discussion of Great Expectations as an approach rather than a commercial for the project/product.
To start the conversation, Abe shared some of his background experience living the pain of unexpected upstream data changes causing data chaos / lots of work to recover from and adapt. Part of where we need to get to using something like data contracts is to remove the need to recover in addition to adapting and move towards controlled/expected adaptation. Abe believes that the best framing for data contracts is to think about them as a set of expectations.
To define expectations here, this would include not just schema but also the content of data, such as value ranges/types/distributions/relationships across tables/etc. So for instance, a column may be a one to five for rankings and then the application team changes it one to 10. The schema may not be broken - it is still passing whole numbers - but the new range is not within expectations so the contract is broken.
At current, Abe sees the best way to not break social expectations is via
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here
Scott Hawkins' LinkedIn: https://www.linkedin.com/in/scott-hawkins-8934393/
In this episode, Scott interviewed Scott Hawkins, Principal Data Architect at ITV.
Scott views data mesh as a mechanism for change. Your company culture and your understanding of it are crucial to establishing data mesh well, driving that buy-in. Ask yourself: what challenges does data mesh actually address and hopefully solve, how will it impact the business not the tech, and what does it change. Your organization might not be ready for data mesh. Or a specific domain might not be ready. And that's okay!
As ITV moved forward, they found a "good enough" solution via a global ID. It's not perfect, there might be some overlap - such as one person might have a different global ID for their online subscription versus their broadcast subscription - but it is far better than what they were doing. And it allows for interoperability/joins across the data. This is a big improvement - don't let perfect be the enemy of good or done.
One thing working for ITV is deploying a "team-in-a-box" to help domains move forward - similar to an internal consulting team. Each situation is different so each box they are given is different. The team-in-a-box concept also means it is somewhat easier to build common best practices internally. Coming to the table with defaults has really helped ITV.
Per Scott, there are 3 good ways to drive buy-in for the domain teams:
Continuing on the driving buy-in, Scott recommends working with the domain managers to generate a viable/valuable carrot for the entire team. Explain to those leaders why it matters, work with the leaders to revamp the KPIs if the KPIs are getting in the way of delivering a good data product. This is why exec-level buy-in is so crucial - it is pretty hard to start modifying team KPIs/OKRs without it! Talking to teams to understand who they (the individual and the group) operate is crucial to developing the right path for them.
Scott also talked about making failure an option. You can try to work with a domain and if it isn't working, it's okay to move on. You don't need to get everyone onboard on day one or sharing their data on day one. If you design incentives well, people will want to participate eventually. Until then, it's okay to walk away from that team.
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this rerelease of episode 48 with Scott Hawkins. I thought this episode was crucial because it is foundational to understanding a lot of good techniques and practices for getting lots of the organizational pieces of data mesh right. You will learn a ton on this one!
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this rerelease of episode 150 with Carlos Saona. eDreams' approach is very unique and interesting because it was essentially all on its own so there are a ton of useful learnings to consider if they are the right fit for your own organizations.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Carlos' LinkedIn: https://www.linkedin.com/in/carlos-saona-vazquez/
In this episode, Scott interviewed Carlos Saona, Chief Architect at eDreams ODIGEO.
As a caveat before jumping in, Carlos believes it's too hard to say their experience or learnings will apply to everyone or that he necessarily recommends anything they have done specifically but he has learned a lot of very interesting things to date. Keep that perspective in mind when reading this summary.
Some key takeaways/thoughts from Carlos' point of view:
To make data producers feel a better sense of ownership, 1) look for ways for producers to better leverage their own data; 2) maximize your number of consumers for their data quanta so there is a quicker time to identify issues with the data product - more eyes means more who can spot issues; and 3) create automation to easily/quickly let domains identify sources of data loss rather than searching: with proper setup, you can make it easy to identify if the data pipeline is the problem. If it's not, then the issue is in the domain.
When Carlos and team were looking at building out how to tackle their growing data challenges a few years ago, they were looking at request for proposals (RFPs) from a number of data consultancies around building out a data lake but just were not convinced it would work. Then they ran across Zhamak's first data mesh article and decided to give it a try themselves. Until more recently, Carlos was not aware of the mass upswing in hype and buzz around data mesh so their implementation is very interesting because it wasn't really influenced by other implementations.
When they were starting out, Carlos said they didn't want to try to create a single, overarching approach. It was very much about finding how to do data mesh incrementally. They started use case by use case and built it out organically, including the design principles and rules - they knew they couldn't start with a single data model for instance. But it was quite challenging iterating towards that standard data model.
When choosing their initial use cases to try for data mesh, Carlos and team had some specific criteria. They rejected anything that needed a very quick turnaround because it wouldn't let them have enough time/space to try things, learn, and iterate. They did plan ahead by creating foreign keys to data products that didn't exist to make interoperability down the road when they would exist easier. And they were very honest with stakeholders about what early participation meant - and what it didn't mean; that way, it was clear what benefits stakeholders could expect.
According to Carlos, while they had executive support and sponsorship for data mesh, that wasn't enough to move forward with confidence at the start. They needed to have a few key stakeholders that were engaged as well and wanted to participate. It was also okay to have some stakeholders not engaged but just informed of what they were trying to do with data mesh. You don't have to win everyone over before starting.
Five things Carlos thinks others embarking on a data mesh journey should really take from their learnings: 1) it's okay to not have everyone really bought in or especially engaged upfront but they will have to participate - make their eventual participation inevitable. 2) Really emphasize that you are learning in your early journey, not that you have it figured out - and factor in learning when doing estimations and promises. 3) Don't try to design your data model from the beginning; you need to learn via iteration - you will start to find your standards to make it easy to design new data products. 4) When treating data as a first class citizen, it's important to understand that will take additional time. Reserve the team's time to create and maintain their data quanta. 5) Let the use cases drive you forward and show you where to go.
Carlos' philosophy is, within reason push as much of the burden onto the consumer as you can. Obviously, we don't want consumers doing the data cleansing work - that's been one of the key issues with the data lake - but the costs of consumption should fall on the data consumers as they are the ones deriving the most benefit. So eDreams makes the consumers own stitching data products together for their queries and makes them pay for the consumption. This minimizes the costs - including maintenance costs - to producers.
One very interesting and somewhat unique - at least as far as Scott has seen - approach is how truly small Carlos and team's data quanta are. Thus far, they have really adhered to the concept that each data quantum should only be about sharing a single type of domain event and really nothing more in it. This again makes for lower complexity and maintenance costs for data producers. They are considering changes with upcoming BI-focused data products so that is to be determined.
Carlos believes - and Scott exceedingly strongly agrees - it is not feasible for your documentation for your data quanta to be fully self-describing. You can't know someone else's context. You need to write good documentation so people can still understand what the data product is and what it's trying to share but if you do not have knowledge of the domain, it would be a considerable amount of effort - essentially impossible to do it right - to fully explain the domain and how it works in the documentation of each data product. Getting to know how other domains exactly work is outside of the scope of the data mesh.
At the start of their journey, the data team was in control of all the use cases, who was consuming, and who was producing, according to Carlos. But, as they've gone wider and there is a self-service model for data consumers, more and more of the use cases are directly between the producers and consumers - or the consumers are consuming without much interaction with producers if they already know the domain. It could become an issue with people trying to understand data from lots of different domains for the sake of understanding but it hasn't been an issue so far.
To date, Carlos hasn't seen many problems around versioning. They thought they would have many more issues with versioning than they have which Carlos believes is from keeping their data products as small as possible and using domain events. When they have had versioning, the retention window for the data has been relatively short so the versioning has been relatively simple to move to the newer version. And because most people are getting their data from source-aligned data products, changes have a smaller blast radius - they won't affect data products that are downstream of a downstream of a downstream data product. Domain events have been enough because their main stakeholder has been machine learning. They are now working on a different kind of data quanta for consumers such as BI, and they plan to include more governed versioning there.
One of the biggest challenges early on according to Carlos was that domains didn't really feel the ownership over the data they shared. So to increase the feeling of ownership, they first looked for ways for producing domains to use their own data - as many other guests have mentioned. Second, they tried to maximize additional consumers of data products by looking for use cases. That led to faster feedback loops if there was a problem - more eyes on the data - so producers discovered issues sooner. And third, the platform team helped identify issues that might be in the system or in the data platform/pipeline process - if there was data loss, there is automation to help identify if it is on the platform side; if it's not on the platform side, then it is an issue with the domain. That one automation has led to a lot less time searching for the cause of data loss rather than fixing data loss.
Carlos and team built in a few different layers of governance. The first is a universal layer for standard metadata in each data product, like when something happened, who is the owner, the version of the schema, the existence of a schema, etc. These are enforced automatically by the data platform and you can't put a data product on the mesh without complying. Producers must also tag any PII or sensitive information like credit cards. Then, a second layer is policies for data contracts between producers and consumers. As many guests have suggested, they have found having default values for SLAs in data contracts provides a great starting point for discussions between data producers and consumers.
"You can have your cake and eat it too," using domain events per Carlos. You don't want direct operational path queries hitting your data quanta as they are designed for analytical queries - they will have a separate latency profile. But at eDreams, the pipeline that writes data quanta to the analytical repository is implemented with streams that can be consumed in real-time by operational consumers (microservices).
Other tidbits:
When launching a new data product, there must be a settling period - consumers must understand that things are subject to change while the producer really figures things out.
You want to avoid duplicating data. But you REALLY want to avoid duplicating business logic.
Data products should have customized SLAs based on use cases. You don't need to optimize for everything. Let the needs drive the SLAs.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this rerelease of episode 133 with Ammara Gafoor. There is a ton to learn from this one and reflect back on.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
ammara.gafoor@thoughtworks.com
Ammara's LinkedIn: https://www.linkedin.com/in/ammara-gafoor/
Data Mesh in practice article series from Ammara and colleagues:
In this episode, Scott interviewed Ammara Gafoor, Principal Business Analyst at Thoughtworks who has been working on a few client projects related to data mesh including one for well over a year.
Before jumping in, it's important to note that much of Ammara's learnings come from an implementation in a 100K+ employee company split into 21 high-level domains. So the definition of domain in this episode revolves around that context of a very large business unit, not a two pizza team size sub domain.
Some key takeaways/thoughts from Ammara's point of view:
Ammara started off the conversation sharing about how she and her team "had it all laid out" for the plan to standardize how they'd bring each domain up to speed on data mesh - from the introduction of new ways of working to being ready to participate in the data mesh implementation in just six weeks. And then reality struck. Each domain is different and much like trying to explain the benefits or implementation of data mesh, a single approach for all audiences just didn't work well so they adapted. Every domain is unique and required its own unique approach to make implementing data mesh in that domain work. There are of course some commonalities but each of the 13-14 domains that are part of the data mesh implementation thus far has had its own unique challenges.
So, Ammara shared some stories about working with different stakeholders. Often, the first stakeholder they encountered was an IT sponsor for the domain itself - either an IT leader in the domain or an IT counterpart for the domain. This persona typically wanted to bring them in and welcomed them with open arms. And while they were often bought in on data mesh, there was a push - from IT and often the business side - to only speak with IT. So Ammara and team had to work to get permission to also include the business people in the conversations about their proposed data transformation. Because without the business support and knowledge, your data mesh implementation is likely to fail. How many episodes have said tie your data strategy to your business strategy? But, the business people often have what they need currently via shadow IT. So why would they want to give that up? It's an emotional response to be asked to give up what you have for the greater good and the long-term.
There is the concept of immediate returns - you build a dashboard and there is immediate potential value - versus the mid- to longer-term returns from things like building your data platform and building out your data governance capabilities. Ammara has seen many times there is not any incentive to wait and focus on the mid- to long-term returns - if your funding this year is based on results this year, focusing on your results 2-3 years out is often doesn't feel like an option. They won't get rewarded for that long-term work. And most domains don't even have the capabilities to do said mid- to long-term high-value work. But to do data mesh right, we need to incentivize patience - and incentivize and provide the capabilities to do things right for the long-haul instead of just the short-term, low stakes wins.
According to Ammara, as part of a successful data mesh implementation, there is the technical stream - the Team Topologies meaning of work stream - but you must also work on the operational stream at the same time. And the product stream too. If you don't look to change domain's KPIs to align their operational work to data mesh "you won't prioritize it - you cannot prioritize it." You need to put a metric into place to measure progress - it doesn't even have to be a great measure! It's a way to start the conversation. There is too much of a hangup in data mesh around trying to get things perfect the first time. Get it done, measure it, iterate on it, and move forward. Don't let perfect be the enemy of done and/or good. Don't fall to bikeshedding.
The cost of change and the cost of failure in data historically have been very high per Ammara. But we have new economic models with cloud that make that no longer true. We now have "the privilege to be able to fail". Failure wasn't an option historically. But that's such a foreign concept to many, it will cause some to push-back. They have lacked the psychological safety to fail. And we have to understand why they are pushing back and work with them to understand that failure in a highly agile environment is incremental learning.
After picking the 2 most obvious use cases in a domain - again, the very large business unit concept of a domain -, Ammara believes it will reveal a 5-6 of the foundational source-aligned or "source oriented" data products of the domain that will be able to power most use cases. So just start building the MVP of those source-aligned data products because they will support other use cases down the road as well.
On Personas, Ammara laid out a few she and team have run into:
The IT sponsor - typically a Data Architect or Data/Analytics Lead; bought in to data mesh, likely after feeling the pain points as Zhamak has laid out. Trying their best to go wide on getting people bought in on data mesh and has some - but not a ton of - social capital to influence. Their social capital is more with the IT/data people and less on the business side of the domain. They are critical to get things moving.
The Business Owner - generally supportive of the data mesh initiative but doesn't have the time - or the incentive - to spend time on the data mesh implementation. You're trying to get their support by the promise of making their lives easier.
The Sideline Watcher - sees data mesh as probably 'yet another data trend'. Not pushing back but not taking a stance. Waiting for the tide to turn one way or another before making their own waves.
The "Yes to Your Face" - will say yes to you and then just go do whatever they were going to do anyway… These are inevitable - try not to take it personally.
The Product Owners - they are building the dashboards or the analytical solutions, desperate for the data. They really WANT to work with you but don't know exactly how - how can they get the resourcing and we're asking them to rethink the way they do their work. Help them figure out how they can partner where possible.
The data lake (or other historical data paradigm) builders - have spent so much time and effort to build a viable data lake/warehouse/etc. Often fight you because you're going against everything they've built. It's not personal against the data mesh team but it is personal if you put all their hard work aside. But they can build data initiatives very well, try to work with them and let them know you're building off the knowledge they've gained if not their direct work.
For Ammara, a lot of the data mesh literature and conversations feel like they say there are new roles and therefore there isn't room for many existing data roles, like the data warehouse or data lake builders/maintainers. But she thinks that's not a great idea - and Scott agrees. They are subject matter experts in how the domain's data flows and systems actually work and can be excellent guides to bringing more people into the data fold as they themselves pick up new skills. Trying to hire your way to a data mesh is not a great idea… No one is redundant, everyone has valuable knowledge for Ammara.
You need to make your IT sponsor successful in order for your data mesh implementation to go broad in that domain so that means learning the - and communicating in the - language of the business according to Ammara. That might mean you have to deal with the horror of PowerPoint Presentations. And as many guests have said, the selling points and implementation details of data mesh don't stick with the broader audience the first time. Repetition, reframing, holding of hands, etc. You won't succeed if you try to just message once. Be prepared to repeat yourself. And then repeat yourself again.
Ammara gave an example of why data mesh can really help improve communication and drive to common language. In manufacturing, there is the concept of "on time, in full delivery" as a very crucial KPI. And the domain had analytics teams constantly asking to build this for the different manufacturing lines while at the same time, the business side said they didn't have the information. How could that be when there were 10+ completed "on time, in full delivery" projects that had been funded? So once Ammara and team removed the data team from the picture, the business folks were able to talk with the regular IT team and they came to a shared, common understanding of what was actually needed and what was missing. It's pretty easy to lose sight of what the actual need and use case is when people are siloed by function.
It is crucial to understand the three streams of work model, per Ammara. The operating stream is "building the cadence for IT and business to communicate" in order to prioritize. This helps identify which data products will be built. The product stream is identifying the actual data products that need to be built, as in what are the scope and boundaries. The technical stream is about building the data product and the platform needs. Each of the three streams should have equal weighting. This is another way to think about your MVP thin slice, you must encapsulate some of each capability, each stream.
As previous guests have noted, many domains build data products that benefit themselves first in Ammara's experience. This obviously makes it easier because there is more buy-in and no cross-domain communication and prioritization friction. But that is just the initial stages of a data mesh implementation - still in phase 1 before going truly broad. More domains are moving to support use cases across domains so phase 2 might be up soon.
Ammara does not believe source oriented data products, ones that are difficult to understand outside the domain, should not be made freely available on the mesh; they should not be made available to business users within the domain or to other domains. And her reasoning is very sound: if the data products are difficult to understand, it's easy to misuse them and they are more likely to change with the source systems so breaking changes/versions are more common. Other domains can consume the information from those source oriented data products in specially designed consumer oriented data products instead of directly from source oriented data products. Data scientists are a bit of another story as they are data literate enough to do some spelunking but even then, data scientist beware.
Ammara is also seeing an interesting pattern relative to source oriented data products. When you really start to map out a lot of obvious use cases for a domain - and remember, the size of a domain in this context is quite large -, it might seem like you need a large number of source oriented data products. But when you zoom out further, it becomes clear that you can actually shrink those into a much smaller number, that 5-6 data products mentioned earlier for that domain.
The way things are evolving at Ammara's current client is 3 layers relative to data products and use cases. For each use case, there are one or more - typically two it sounds like - consumer oriented data products. Then each consumer oriented data product is derived from or powered by typically three to four source data products. So the domains are able to create multiple consumer oriented data products off the same set of 5-6 data products. But it's still early days and will likely evolve further :)
Encourage people to think business need first instead of data first according to Ammara. Think about what business outcome you are trying to achieve and then work backwards to what data you need to address that. If we are just sharing information without intention, it can lead to misuse of data - will people really...
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this rerelease of episode 177. As stated in the original show notes, this is one to revist often as it is a great level-setting on why are we doing what we do in data mesh.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
This is likely to be an episode to revisit. Zhamak explains a simple concept - data should not be copied unless it is owned by a data product - but the why is multi-layered and important. It might be one of the most important yet underestimated aspect of data mesh because when done right, it truly ensures trust in data - for consumer but also producer. There's a lot of nuance in how Zhamak is thinking about this but the actual application is quite easy :)
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
This is a rerelease of a previous episode as we are on a hiatus related to Scott's recent health issues. This is a very important episode for many reasons so please do check it out again :)
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here.
Data Mesh Learning Meetup Presentation: https://www.youtube.com/watch?v=7iazNKG8XQo
Sarita's LinkedIn: https://www.linkedin.com/in/saritabakst/
In this episode, Scott interviewed Sarita Bakst, Managing Director and a leader in Firm-wide Data Management at JP Morgan Chase. Sarita was previously on a Data Mesh Learning meetup and is helping lead the firm's data mesh governance charge.
Per Sarita, historically, data governance has meant controls and gatekeeping to most people - typically getting in way of innovation. So there needs to be a focus on changing that narrative, not just through words but actions to show it's not the case. In data mesh, you need to ensure that domains can make good decisions on governance and seek out subject matter experts when it makes sense.
Sarita covered that one of the key issues in the way governance has been done is the people making decisions - the central governance team - don't have the real understanding of the data. When those decisions are put in the hands of the people who really know the data, but with guardrails and guidance, the fear is lifted about can we actually use this data and how. This opens up lots of new opportunities to leverage your data.
Sarita strongly recommends starting with purpose-built data products. Find a use case and build data products to serve that specific. And data products MUST be about unlocking business value. You don't need to serve up all of a domain's data on day one, in version one of that domain's first data product - make it extendible and reusable so you can find additional consumers and expand over time.
Get out of your own way on data governance in data mesh. You are going to learn and your approach will evolve as you learn. It's okay to not know everything upfront, set yourself up to not get in trouble - put the proper guardrails in place - but you won't know everything. Think of designing your risk controls as toll-gates and make sure they aren't bottlenecks.
Have standards (not standardization) so people don't have to invent things from scratch. Standards for interoperability, naming, etc. - think of them as guiding principles instead of rules. Make sure domain owners know who to contact and when on governance subject matter expertise.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
This is a rerelease of a previous episode due to my health related issues. New episodes will begin again in a few weeks. Please enjoy this very important episode of Data Mesh Radio.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Khanh's LinkedIn: https://www.linkedin.com/in/khanhnchau/
In this episode, Scott interviewed Khanh Chau, Lead Architect for the Data Mesh Initiative at Northern Trust.
Khanh believes you have to be passionate about making data better to do a good job implementing data mesh. And it is DEFINITELY a journey so you need patience and vision. Also, each journey is unique, you can't just copy/paste from another organization. You need to make failure okay - but you should look to make it easy to fail fast, measure, and adjust.
Khanh talked about the need for exec buy-in before heading down the data mesh path. They got that exec buy-in by proving that the total cost of ownership of data was quite high as the consumers had to do a LOT of work to get the data to usable.
When speaking internally, the business people were very excited to participate if it meant they could get quality data. Some of the IT/data engineering folks were harder to convince. It was especially hard to get them to shed layers of not-useful technology.
Some IT teams were easier to convince - they had felt the impact of a few too many middle-of-the-night data downtime incidents. Other teams hadn't felt that pain so there were harder to win over. There was also the incentive of additional possibilities - data mesh meant they could do things they couldn't do before.
Khanh talked about making the platform the easy and right path for 80% of use cases. They focused on making things easy to configure; basically: what transformations do you want to do and then it automatically provisions the pipelines. Their goal was to make it easy to make good progress quickly; their time to initial deploy went from 2-3 months per data service to 2-3 weeks per data product and they hope to drive it down further.
Northern Trust has been moving forward with data mesh for about 7 months as part of their high-level digital transformation initiative. On the data side, they had previously focused on data virtualization and data federation but it was not delivering the results they wanted. It was not as scalable as they wanted - it was taking 2-3 months to launch each new data service. They also did not have great information on who was consuming the data and why.
For their data mesh proof of concept, Khanh and team set a timeline of 9 weeks. They needed to prove value by then or data mesh would be a very tough sell internally. Khanh talked about the need to sell data mesh as a paradigm shift in order to get people out of technology-focused thinking.
Northern Trust decided to take a pragmatic approach e.g. not pushing all aspects of data...
Due to health-related issues, we are on a temporary hiatus for new episodes. Please enjoy this release of episode 25, potentially the most important episode of Data Mesh Radio.
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Due to lingering severe health issues and not trying to rush myself back, I will be doing a 3 week hiatus around new episodes. In that break, I will be rereleasing some old episodes that I think people really need to listen to or possibly reflect on if they haven't listened to them lately. Here's to better health in the near future!
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Burce's LinkedIn: https://www.linkedin.com/in/burcegultekin/
Beth's LinkedIn: https://www.linkedin.com/in/beth-bauer-102449/
Ghada's LinkedIn: https://www.linkedin.com/in/ghada-richani/
Michael's LinkedIn: https://www.linkedin.com/in/mjtoland/
Michael's 'Super Chicken' reference talk by Margaret Heffernan: https://www.ted.com/talks/margaret_heffernan_forget_the_pecking_order_at_work
In this episode, guest host Burce Gültekin, Chief Data and Analytics Officer at FrieslandCampina facilitated a discussion with Ghada Richani, Managing Director, Data & Technology Strategy & Project Management Office at Bank of America (guest of episode #206), Beth Bauer, CEO at her own consulting company PosiROI (guest of episode #218), and Michael Toland, Senior Product Management Consultant & Coach at Pathfinder Product Labs. As per usual, all guests were only reflecting their own views.
The topic for this panel was how do we tie the business strategy all the way down to the data work via the data strategy...
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Christoph's LinkedIn: https://www.linkedin.com/in/christoph-spohr-2a33b619b/
In this episode, Scott interviewed Christoph Spohr, Lead Architect of Big Data Platforms and the Product Owner of Data Mesh at Volkswagen Group. To be clear, Christoph was only representing his own views on the episode.
Quick Note: apologies that Scott's audio is a bit weird, he had yet to build his makeshift sound studio in the Netherlands.
Some key takeaways/thoughts from Christoph's point of view:
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Takeaways:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Ale's LinkedIn: https://www.linkedin.com/in/alejandracabre/
In this episode, Scott interviewed Ale Cabrera, Senior Data Quality Product Manager at Clearbit. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Ale's point of view:
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key takeaways:
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Kinda's LinkedIn: https://www.linkedin.com/in/kindamaarry/
In this episode, Scott interviewed Kinda El Maarry PhD, Director of Data Governance at Prima. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Kinda's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Samia's LinkedIn: https://www.linkedin.com/in/samia-r-b7b65216/
Khanh's LinkedIn: https://www.linkedin.com/in/khanhnchau/
Sheetal's Links:
Sheetal's LinkedIn: https://www.linkedin.com/in/sheetalpratik/
First Paper at ICEB Conference: https://easychair.org/publications/preprint/qZ3m
Interview in Harvard Business Review: https://hbr.org/resources/pdfs/comm/BeyondTechnology_CreatingBusinessValuewithDataMesh.pdf
Medium Blog on Saxo Bank Implementation: https://blog.datahubproject.io/enabling-data-discovery-in-a-data-mesh-the-saxo-journey-451b06969c8f
In this episode, guest host Samia Rahman, Director of Enterprise Data Strategy, Architecture, and Governance at life sciences company Seagen (guest of episode #67) facilitated a discussion with Sheetal Pratik, Director Engineering and leading India Data Integration Platform at Adidas (guest of episode #24), and Khanh Chau, Director of Cloud Data Architecture at Grainger (guest of episode #44). As per usual, all guests were only reflecting their...
You can simply use the link here https://us06web.zoom.us/j/87380091930?pwd=NVhUS3ZDOS9hSFFBbU1sQm9HQUJHZz09 or find a link on LinkedIn events: https://www.linkedin.com/events/weeklydatameshopenroundtable-da7091373106551734272/comments/
Link to past recordings playlist: https://www.youtube.com/playlist?list=PL9tZROTS_Pi7TjjhLR-sCtwQVeth6Jv7e
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Benny's LinkedIn: https://www.linkedin.com/in/bennybenford/
Benny's Substack: https://www.datent.com/p/elevating-data-to-a-profession-why
In this episode, Scott interviewed Benny Benford, the former CDO at Jaguar Land Rover (JLR) and who is currently building out a community around data transformation. To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Benny's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Takeaways:
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Kim's LinkedIn: https://www.linkedin.com/in/vtkthies/
Gemba Walk explanation #1: https://kanbantool.com/kanban-guide/gemba-walk
Gemba Walk explanation #2: https://safetyculture.com/topics/gemba-walk/
PayPal Data Contract Template OSS: https://github.com/paypal/data-contract-template/tree/main/docs
Start with why -- how great leaders inspire action | Simon Sinek | TEDxPugetSound: https://www.youtube.com/watch?v=u4ZoJKF_VuA
In this episode, Scott interviewed Kim Thies, at time of recording a Leader on the Enterprise Data Team at PayPal and now SVP, Client Innovation & Data Solutions at ProfitOptics. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Kim's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Key summary points:
Please Rate and Review us on your podcast app of choice!
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Smriti's LinkedIn: https://www.linkedin.com/in/smritikirubanandan/
Smriti's HLTH Forward Podcast: https://hlthforward.buzzsprout.com/
In this episode, Scott interviewed Smriti Kirubanandan, a Healthcare and Public Health Data Expert at a large consulting firm. To be clear, she was only representing her own views on the episode. Much of the challenges and opportunities discussed in this episode are more on the US side because of the not-so-well-functioning healthcare system there.
Some key takeaways/thoughts from Smriti's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
Get involved with Data Mesh Understanding's free community roundtables and introductions: https://landing.datameshunderstanding.com/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Frannie's LinkedIn: https://www.linkedin.com/in/frannie-farnaz-h-a7a11014/
Jill's LinkedIn: https://www.linkedin.com/in/jillianmaffeo/
Alla's LinkedIn: https://www.linkedin.com/in/allahale/
In this episode, guest host Frannie Helforoush, Technical Product Manager/Data Product Manager at RBC Global Asset Management (guest of episode #230) facilitated a discussion with Alla Hale, Senior Data Product Manager - Digital Capabilities at Ecolab (guest of episode #122), and Jill Maffeo, Senior Data Product Manager at Vista (guest of episode #151). As per usual, all guests were only reflecting their own views.
The topic for this panel was broadly data product management and the role of the data product manager in a data mesh implementation. Data Product Manager is still a very nascent role so there is still a lot of confusion around it :) If I were to sum up the feeling of the conversation very succinctly, it would be: it's early days, have patience.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Tania's LinkedIn: https://www.linkedin.com/in/sofia-tania/
Presentation: "Data Mesh testing: An opinionated view of what good looks like": https://www.youtube.com/watch?v=stNZQESndAA
In this episode, Scott interviewed Sofia Tania (she goes by Tania), Tech Principal at Thoughtworks. To be clear, she was only representing her own views on the episode. Scott asked her to be on especially because of a presentation she did on applying testing - especially important for data contracts - in data mesh.
Scott note: I was apparently getting extremely sick throughout this call so if I ramble a bit, I apologize. Tania's dog also _really_ wanted to be part of the conversation so you might hear us both chuckling a bit about her antics. And Tania has some really great insights so I probably asked her probably the hardest questions of any guest to date. She did a great job answering them though! A lot of the takeaways are about are we actually ready to do a lot of the necessary testing to ensure quality around data, which I don't think has a clear answer yet :)
Some key takeaways/thoughts from Tania's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Takeaways:
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Brenda's LinkedIn: https://www.linkedin.com/in/brenda-contreras-9649a47/
In this episode, Scott interviewed Brenda Contreras, VP of Engineering and Architecture at Self Financial.
Some key takeaways/thoughts from Brenda's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
So here are the summation points of this episode:
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Iryna's LinkedIn: https://www.linkedin.com/in/irinakukleva/
Mobey Forum: https://mobeyforum.org/
In this episode, Scott interviewed Iryna Arzner, Head of Group Customer Growth, Retail Banking at Raiffeisen Bank International (RBI). To be clear, she was only representing her own views on the episode.
Scott note: I mostly use the phrase line of business or LOB instead of domain in this write up but they are mostly interchangeable.
Some key takeaways/thoughts from Iryna's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Khanh's LinkedIn: https://www.linkedin.com/in/khanhnchau/
Balvinder's LinkedIn: https://www.linkedin.com/in/balvinder-khurana/
Yushin's LinkedIn: https://www.linkedin.com/in/yushin-son-30362b1/
Carlos' LinkedIn: https://www.linkedin.com/in/carlos-saona-vazquez/
In this episode, guest host Khanh Chau, Director of Cloud Data Architecture at Grainger (guest of episode #44) facilitated a discussion with Balvinder Khurana, Technical Principal and Global Data Community Lead at Thoughtworks (guest of episode #135), Carlos Saona, Chief Architect at eDreams ODIGEO (guest of episode #150), and Yushin Son, Chief Architect of Data Platform & Data Products Engineering at JPMorgan Chase. As per usual, all guests were only reflecting their own views.
The topic for this panel was an architect's view of data mesh, especially from an architecture lead standpoint. There are many challenges architects face in data mesh, managing the micro level minutiae, down to the data product output and input port decisions but balance that with crucial high-level decisions. Balancing the near-term and long-term vision and roadmap/North Star.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views...
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Aaron's LinkedIn: https://www.linkedin.com/in/aaron-wilkerson-81bb21a/
In this episode, Scott interviewed Aaron Wilkerson, Senior Manager of Data Strategy and Governance at Carhartt. To be clear, he was only representing his own views on the episode. Apologies for the lawn work sounds around the middle of the episode :)
Before we jump in, this episode contains a lot of really good framing on how data leaders can actually partner with business people to drive to what matters for them. How do you extract what matters to the organization and to each specific business partner? And then how do you tie the data work to that? So while this episode is not heavy on data mesh specifics, it's really important to really considering the business partner's point of view and how to work with them to drive value for the organization.
Some key takeaways/thoughts from Aaron's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Takeaways:
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Frannie's LinkedIn: https://www.linkedin.com/in/frannie-farnaz-h-a7a11014/
Post on the Product Trio concept by Teresa Torres: https://www.producttalk.org/2021/05/product-trio/
In this episode, Scott interviewed Frannie Helforoush, Technical Product Manager/Data Product Manager at RBC Global Asset Management. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Frannie's point of view:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
My overall point here is that why do so many folks in data hate Agile? Because in data it's so rarely done well basically. Things are done because 'that is the way they are supposed to be done' instead of 'because this will make our teams happier and more efficient'. And quite honestly, Agile isn't for every organization. The spirit of Agile probably should be for every organization so maybe go read the Agile manifesto but in data, the one size fits all approaches are obviously breaking more and more. So work with your teams and talk about what you want to achieve and collaborate with them to get there. Yes, easier said than done but I believe in you.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Tina's LinkedIn: https://www.linkedin.com/in/christina-albrecht-69a6833a/
In this episode, Scott interviewed Tina Albrecht, Lead Coach for Data-Driven Transformation at Exxeta.
Some key takeaways/thoughts from Tina's point of view:
Tina started with some of the big questions you will need to consistently ask yourself in a data transformation: is this good enough? Is this good enough for now? What is changing and how is it changing? What are the milestones we want to hit and, looking back, that we have hit? And often overlooked: how is/was the team feeling during the transformation? Those questions can help you narrow in on how well your transformation is going - and let's face it, data transformation doesn't stop. A good three layer system to think about when breaking this conversation down is: the surface layer of what happened in the past and where things are now; one layer deeper is where could they go if they had stronger data capabilities; and the bottom layer being okay, but how can we get there?
When actually measuring whether a current solution is "good enough", Tina recommends two measures or questions to consider. The first is: how happy are people with the current process? Happiness is a decent measure of effectiveness. If people are happy, working hard to improve the process probably doesn't make financial sense, the return isn't there. The other aspect is to think about how much effectiveness of the process is lost to constraints and bottlenecks. This analysis will give you a good insight into the process and provide the perspective on if it is currently good enough.
Tina talked about two ways most data mesh implementations seem to be going wrong. The first is how teams are interacting with each other: the team and domain setup. There are often breakdowns in how teams collaborate together so things that should be explicitly owned and bridged between the teams just keep getting dropped. The other is data ownership where domains won't take data ownership and/or don't understand data ownership. We need the platform team to own enabling domains to leverage the platform but the domains keep trying to push work back to the central platform team or not owning data well enough so the central team has to step in to help. If that happens, the platform team can't get the necessary platform work done to do data mesh and they become that centralized bottleneck again.
In data mesh, Tina believes there is still a lack of understanding by domains of what data ownership means, what they are actually supposed to be responsible for. You can help domains better understand data ownership by making sure they have necessary embedded data engineering talent within the domain to actually be capable of owning data as more domain members learn how to own data. And you need strong governance capabilities to help teams understand how to interoperate data between domains easily.
Tina talked about with one client, they are rotating embedded data engineers between domains - and the central platform team - so the domains become more data capable and you have a wider knowledge base about data across the organization as the data engineers share that knowledge with each other and the organization. And just having a simple community or guild for the data engineers wasn't enough, they needed to go with the embedded model instead. That central hub and a central team managing the data engineers' careers has been very important to keeping people happy.
Similarly, Tina talked about how while Team Topologies is a great tool for organizing your teams in data mesh, it's only a tool. If you don't understand your value chain, if you don't really focus on how you create value via data, it won't save your data mesh implementation. Start from value first.
Is data mesh right for your organization? When Tina is assessing that question for clients, a great starting question is simply what changes, what value would come from doing data mesh? If there isn't a clear vision as to what would be better and how that would drive clear value - and a large amount of value too, data mesh is not a light undertaking - then will data mesh really align with the business strategy and drive value? Scott note: I think these questions are REALLY crucial to answer. If you don't know what will change if you do this well and how that ties to business value, you shouldn't do it :)
For Tina, in general in data work, there are two big areas where value is blocked or lost along the value chain. As she mentioned earlier, lack of clear ownership and responsibility is a big one. People understand how value is generated but it's unclear who owns what and so major needs along the value chain aren't met - basically, no one thinks they own crucial aspects and so they don't get done. The other is again simply bottlenecks - where are things blocked or where are dots not connected? Once you identify issues in either aspect, you should have your areas to target for change to drive more value related to those specific processes.
A key aspect of transformation for Tina is deep clarity. There are so many things that are changing, who owns what and what outcomes do they own? What is actually being done and why? There needs to be strong governance leadership that lays out many aspects rather than leaving things to chance. It doesn't have to be heavy-handed governance but ensuring things will work together - and that someone _owns_ making them work together well - is the best way to ensure a successful data transformation, data mesh or otherwise. And you have to stay on top of things, it's not a one-and-done kind of transformation.
Intentionality around communication is also crucial to successful data transformation for Tina. Overcommunication is a virtue. Data isn't about the 1s and 0s, it's about sharing information. So you need to be explicit in setting expectations and creating mechanisms for people to exchange information, especially across domains. Have regularly scheduled workshops to actually get people exchanging crucial context - it's like a good relationship, you need to continue to work on your communication. Many think about that information exchange at the actual 1s and 0s level but we need people to exchange information with each other constantly too. Otherwise there are too many incorrect implicit assumptions and again, balls get dropped and value is needlessly lost.
Tina made a somewhat comical but very true point: when you are doing a large-scale change to how you do data work, if there isn't anyone saying they are confused, it's probably a bad sign. Because large-scale change is difficult and inherently change will be at least a bit confusing for most so if no one is speaking up, they probably have some bad implicit assumptions you need to address but you don't know what they are. At least with confusion, you can drill into where they don't get it. Lean into confusion because it creates the perfect situation to actually exchange context and drive people to the same page.
While it is incredibly difficult to provide an exact value of data and data work - it will be valued differently by different people - Tina still asks people what is the purpose of doing the work. Why do we care about this data? That will tell us the general value of it if not a specific dollar figure.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Paolo's LinkedIn: https://www.linkedin.com/in/paoloplatter/
Paolo's Medium (multiple data mesh articles): https://medium.com/@p-platter
Agile Lab's website: https://www.agilelab.it/
Manisha's LinkedIn: https://www.linkedin.com/in/evermanisha/
'A streamlined developer experience in Data Mesh' blog post by Manisha: https://www.thoughtworks.com/insights/blog/data-strategy/dev-experience-data-mesh-platform
'A streamlined developer experience in Data Mesh (Pt. two)' blog post by Manisha: https://www.thoughtworks.com/insights/blog/data-strategy/dev-experience-data-mesh-product
'Data Mesh Accelerate Workshop' blog post by Thoughtworks: https://martinfowler.com/articles/data-mesh-accelerate-workshop.html
Max's LinkedIn: https://www.linkedin.com/in/max-schultze/
Max's Data Mesh Learning meetup presentation: https://www.youtube.com/watch?v=QwtTdP2wKFo
(he has many more on YouTube! https://www.youtube.com/results?search_query=max+schultze+data+mesh)
Data Mesh in Practice ebook he co-authored (Starburst info gated): https://www.starburst.io/info/data-mesh-in-practice-ebook/
JGP's LinkedIn: https://www.linkedin.com/in/jgperrin/
JGP's 'Data Mesh for All Ages' book: https://jgp.ai/2023/01/20/data-mesh-for-all-ages/
JGP's website (lots of data mesh content): https://jgp.ai/
JGP's Blog Post 'The next generation of Data Platforms is the Data Mesh': https://medium.com/paypal-tech/the-next-generation-of-data-platforms-is-the-data-mesh-b7df4b825522
In this episode, guest host Paolo Platter, CTO and Co-Founder of Agile Lab (guest of episode #3) facilitated a discussion with Manisha Jain, Data Engineer at Thoughtworks (guest of episode #220), Jean George Perrin (AKA JGP), Intelligence Platform Lead at PayPal (guest of episode #130), and Max Schultze, Associate Director of Data Engineering at HelloFresh (guest of episode #21). As per usual, all guests were only reflecting their own views.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Mike's LinkedIn: https://www.linkedin.com/in/2mikealvarez/
In this episode, Scott interviewed Mike Alvarez, Former VP of Digital Services leading the data mesh implementation at a large healthcare distribution company. He's now working on his own startup.
Some key takeaways/thoughts from Mike's point of view:
Mike started off with the general need for large companies to change their approach to analytics at scale. We've been doing a lot of the same things for the last 30 years and they aren't quick enough to respond to changing business needs - 6+ months and $1M+ to get to your first query just doesn't make sense anymore - did it ever? And we can do better now. The business side of companies shouldn't have to wait for data and see the world change well before a solution is delivered. We need to move at the speed of business.
Regarding shadow IT, despite leading a central data/IT organization, Mike doesn't hold it against domains. The lines of business can't deal with the bottlenecks of going through a central team and try to build things themselves. However, it's rarely all that scalable and certainly isn't built with sharing to the rest of the organization in mind. Shadow IT just isn't built with a product mindset so it becomes brittle and dilapidated quickly. So the central team is a bottleneck but the decentralized approaches don't scale. Add in the teams generally not really truly understanding the data they are ingesting or even often producing and it's a recipe for data underperforming expectations. Of course, this is what Zhamak identified and why she created data mesh.
Mike talked about when considering a new approach to data, he didn't want to do data mesh for the sake of it. What was the problem they wanted to solve? Could an existing approach or platform do what was necessary? What were the organization's past failure modes or times when things didn't meet expectations and what were the common through-lines or patterns? And then he took those past unmet expectations and used them for understanding as well as driving buy-in. The definition of insanity is trying the same thing over and over and expecting different results. So if data projects were constantly not meeting expectations, shouldn't we change the way we approach data? And treating data as a product seemed like a great start, which led to selecting data mesh :)
While data mesh can feel like the right call to some immediately, it's not likely to be the universal reaction at any organization. Mike and team spent a number of months driving to how this could work and building up the buy-in and momentum to even start on their data mesh journey. This isn't an overnight approach, you really need to think deeply about how it could work and - back to those potential failure modes - how it could go wrong so you can prevent heading down bad paths as best as possible.
But what really drove Mike's interest in data mesh as a possible solution was how it could enable the teams closest to the customer to react to customer and market needs, especially changes in customer demands/wants/challenges. It is about empowering the teams to move at the necessary pace to stay ahead of the competition instead of waiting for a centralized team to give them access to leverage their own data or the data of teams close to them in the organization.
For Mike, the value of data mesh isn't about the technology shifts, at least not yet. It's about the operating model shift - giving domains the capabilities and empowerment to handle data. We are trusting them to own their data and giving them the ability to do so in a scalable way. We are giving them the ability to react in a much quicker and more meaningful way. All of these can get people leaning in to doing data mesh. But they don't care whether it's data mesh or any other paradigm. That's where data people need to connect the dots for them, how can this work and what benefit does it have for the domain. And what are the actual changes for them?
To get the most out of data mesh, Mike believes you have to have a strong vision of what are you actually trying to achieve. It's not an approach to take on lightly. You need to really think about aligning everyone around that shared vision and build as a community effort. How do you take the principles and new approaches and focus on delivering business value - for their own domain and the broader organization too?
Mike believes a big part of doing data mesh is kind of the social contract around enablement and empowerment. Sure, teams can go off in their own direction but if they give up some of their autonomy to stick to the centrally provided tooling - which makes governance far easier -, you need to give them something in return. In their case, Mike and team gave the gift of automating away a lot of the toil work :D
On advice to his past data mesh self, Mike talked about early in a mesh journey, people believe they are aligned on vision but they probably aren't. Holding all of data mesh as a concept and then contextualizing it to your specific organization is a massive amount of work and cognitive load. Trying to get someone to fully understand that upfront without seeing the progress, you will almost certainly have some misalignment and misunderstandings. Instead of the specifics, focus on the target outcomes, what are you trying to achieve? If people align on the benefits, you are more likely to gain and retain momentum. And it will take a lot of effort to get most people committed to the vision, just be prepared for that.
Mike talked about the three key aspects of a product: viability, feasibility, and desirability. Feasibility it a crucial aspect to consider in data - especially data as a product thinking - because often, something just isn't likely to work for a number of reasons. And when there is the desirability but not feasibility, you really need to communicate why it's not going to happen. With data mesh, there can be a misconception that the switch has been flipped and we can do any data work we can think of - and that we should! But that prioritization process and understanding - and then communicating - what is the current art of the possible is important. Always be communicating about what you are doing when and why.
Domain understanding is crucial to really understanding data as a product for that domain in Mike's view. How do we move from trying to serve data sets as if that is the product to creating the information that will be most useful to consumers about the domain in a productized way? And then iterating towards more and more value as you improve the data product or suite of data products representing the domain. Easier said than done of course.
Mike asked the provocative question of do we still want to seek the fabled "single source of truth." It can be a bit like the dog chasing its tail - when you catch the tail, then what? Are we trying to perfectly clean data or are we trying to drive value from data? Is the juice worth the squeeze or can we drive better value - and especially nimbleness - by taking a slightly different view? Scott note: Zhamak urges people to consider "the most relevant source of truth" because there are multiple perspectives on the same things that can all have value, you have to decide what is best.
Mike warned that some use cases are politically untenable or even toxic. Especially early in your journey, consider will participants actually want to know the information. Yes, in the abstract, we want everyone to be perfectly data driven but humans aren't and won't ever be. Don't ignore that and tackle something that will be more hassle than it's worth.
In wrapping up, Mike had two points. The first is learn to work incrementally. That has been somewhat of the antithesis to how data work has historically been done but it's incredibly important. The second point is to really lean into empowerment and the art of the possible. We don't really know what might happen when we empower thousands of our colleagues to be better able to leverage data. Be excited and open to the journey of finding out what value they create.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Key Takeaways:
More on Postel's Law: https://ardalis.com/postels-law-robustness-principle/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Himateja's LinkedIn: https://www.linkedin.com/in/himatejam/
Himateja's AWS ReInvent presentation (link starts at her part): https://youtu.be/y1p0BGsPxvw?t=1991
In this episode, Scott interviewed Himateja Mandala, Senior Data Engineering Manager and Head of the Data Mesh Data Platform at Disney Streaming. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Himateja's point of view:
Himateja started the conversation with the situation Disney Streaming was in that matched many organizations right now: many data platforms but not one that will really fit with data mesh, even with augmentation. So she and her team decided that because the existing platforms were too hard to change to meet the needs of a data mesh implementation, they'd need to build their data mesh platform from the ground up.
When you have new key personas leveraging the data platform, even if those are data engineers embedded into the domains, Himateja recommends rethinking how data work is done. What do people need automated and by default, like security? How do you create monitoring/observability that helps people easily pinpoint issues as they come up? How do you make data accessible by default at the data product and greater mesh level? Etc. In a decentralized, federated data approach, ways of working and needs will be different so dig into what are the actual pain points instead of solving the same pain points of previous implementations.
Himateja shared that while people may think data products are pretty similar, they end up relatively different based on use case. Audience also really mattered when trying to figure out what capabilities people required early in the journey - execs were often more focused on data privacy and security and data scientists were focused on data quality. It's hard to focus on the business context at the platform level because many people are used to doing that via request. The data products themselves need to own business context.
Data contracts are crucial to maintaining data quality in Himateja's view. While they are certainly helpful to data consumers, they are also very helpful to data producers because - with proper observability - data product owners can quickly identify and address quality issues as they emerge instead of waiting until consumers complain and downstream data is wrong. That proactive alerting and then response helps everyone better trust the data. However, data contracts are still a work in progress because not everything is easy to define in a contract, there are definitely gray areas that are improving but not great yet. Scott note: and that's okay, we can't get everything perfect upfront, we have to iterate towards better :)
Himateja then shared a lot about what the data platform team that she leads set out to do at the start of their data mesh journey. One aspect was to create a center of excellence approach, standardizing how data engineering work is done to create data products across the 15+ teams running on the platform now. They did that by starting to drill into pain points and doing lots of listening to potential users. They needed to take a different approach rather than just yet another data platform.
Preventing the central data platform team from becoming a central data engineering team was a worry for Himateja: how do you prevent being a bottleneck and empower teams to do what they need to do? Especially at the start of a journey? As many guests have pointed to, automation and blueprints have been crucial. Teams pushed back initially at the thought of managing their own infrastructure but they realized it gave them the ability to move at their own pace - no more waiting in a prioritization queue for necessary infra. Another key milestone was developing tools to make it easy for cross domain communication and data sharing. Domains at Disney Streaming actually wanted to share their data with each other and the data mesh platform/implementation made that possible - it was previously very difficult to trust data but now that quality metrics were clearly defined and tracked, data sharing and usage between domains increased significantly.
Prior to doing data mesh, Himateja shared that data engineers in domains had no real visibility into data infrastructure - provisioning timeline or any other aspect. They'd push a ticket and wait for things to happen. But with data mesh, since they own the infrastructure, they can go at their own pace and can understand much more about any delays. So they understand and better control their own timelines, which makes them far happier. And then once they’ve gone through the process of spinning up infrastructure the first time, the next time they can be that much faster. They are better able to move at the speed of their business.
Himateja shared about the process of evaluating new platform capabilities requests. As many past guests have noted, you need to establish a process to abstract away the requirements from individual use cases to find a generalized approach. Otherwise, you end up with yet another overburdened platform that you can't evolve. A specific example at Disney Streaming was enabling their Apache Kafka clusters to better communicate across domains which were leveraging individual cloud accounts. Instead of building a solution to share data for each technology they use in the platform, they built a system to better enable sharing across accounts with proper access control and privacy. And sometimes, the answer is that you can't support a unique requirement via the platform - that's okay and often is the right call.
At Disney Streaming, Himateja and team implemented a very interesting approach to access control via RBAC (role-based access control). There are a few levels of data usage clearance but if you are at access to PII level or access to financial information level in one domain, it's the same clearance level for all domains. This might not work for heavily regulated industries but it's working very well for them. The work to decide who has access to what data is done ahead of time instead of constant requests. They just think very carefully about each use case and who should have access and why but there isn't a need to manually grant access. And there is of course oversight to see how people are using data to make potential changes. Scott note: this is a really interesting approach and I'd love to hear people's feedback.
In wrapping up, Himateja shared how they are strongly limiting their blast radius around sharing data with partners/vendors. They have accounts that are not able to get access to any other accounts where they give those partners/vendors access so there is not a way for them to access data they shouldn't be able to see. It's a simple security pattern but others should consider adopting it in her view.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Important points to consider about being a data driven organization:
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Lynn's LinkedIn: https://www.linkedin.com/in/lnoel/
DAMA New England: https://damanewengland.org/
Correlation/Causation XKCD: https://xkcd.com/552/
In this episode, Scott interviewed Lynn Noel (pronounced Knoll), Data Governance Lead at AIM Consulting. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Lynn's point of view:
Lynn started with her being the lead data governance person for a data mesh implementation at a client. It's allowed her to have governance be part of the value drivers for all aspects of mesh. She said data governance in mesh is about making data "safer, better, and easier to use and share." The value proposition of doing governance well in data mesh starts to emerge when you think that to really get the most out of data, you need teams to feel like they can rely on data, not just use it.
It's crucial in all aspects of data, but especially with data mesh, that governance is an enabler instead of a blocker according to Lynn. As a governance leader, you can't be trying to put up blockers but it's crucial for the team to understand why aspects of governance are helpful and add value. You should be building alongside - especially the building aspect, being part of the work - to keep things moving while also making sure you aren't setting yourself up for trouble later, whether that is quality, regulatory, or otherwise.
Lynn said "not all governance is created equal" - it's important to establish where you can and cannot compromise. Infrastructure security and data classification are both non-negotiables. Both wrap in to external threats and internal access control, including to regulatory requirements. Risk management and compliance just cannot be ignored so make them part of the build process and ensure everyone understands why they matter. It's a core of product thinking around data.
There are multiple ways to 'shift left' in governance in Lynn's view. The first is shifting governance left in the project timeline - it should be part of every aspect of building your data mesh from the start - platform, data products, etc. You don't layer governance on at the end. We also should be shifting some of the responsibilities left to the developers - while giving them the capabilities and knowledge to handle what they can and come to experts for what they can't.
Lynn has had some success with talking to business users about not providing them a platform but an "ecosystem". It lets them envision something on a bit more of a tangible level, that they will build out their own habitat but it plays into a larger ecosystem. So they are comfortable in their own habitat which lets them better deliver value to themselves and others. Of course, your ecosystem can't infringe on other ecosystems so it is a good analogy for how domains all interact.
When looking at what to automate as part of your platform, Lynn recommends strong communication with users. You can surprise and disorient them if you automate the wrong things. E.g. there are certain aspects of discovery analysis that look a lot like standard tests but one is incremental work where you want to dive deeper into the data and one is part of daily work without value add from doing manually. Automate the second but not the first :)
Lynn brought up that your organizational setup before moving to data mesh will have a very large impact on what you should focus on when in your journey. If you are coming from an overly centralized approach, it will mean that you understand the benefits of integration but might have some challenges with letting go of central control. If you are coming from a very highly decentralized world, you might need to focus much more on alignment because domains may have had freedom to do as they please with little thought for users in other domains.
In organizations doing data mesh coming from a very highly centralized approach, Lynn says be prepared for a lot of dark data. Because the officially blessed IT systems were overly rigid, many people were doing work outside of the systems, often in Excel or similar. Getting people to trust you that you'll take good care of their data and they can still maintain access will be hard. There are many fun governance challenges including naming conventions to tackle with data like this.
For Lynn and her team, it's been very important to focus specifically on answering what's in it for me from users. Me in the personal sense and me in the organizational and/or domain sense too. And having people whose role involves owning organizational change management is very helpful as part of a product team. There is literally a person or people responsible for getting alignment and driving buy-in.
Quick Tidbits:
In data, we too often "confuse the pipes for the water" - losing the forest for the trees. We focus too much on the plumbing instead of the resource people want: understandable, usable, trustable data. Focus on the outcomes and business value.
It's easy to scare people by using the 'federated computational governance' term. Work with teams to show them where we can automate the past pains of governance so everyone can work on what drives incremental value not manually applying policies.
Regarding a data warehouse: "I never saw a plain vanilla implementation survive anybody's business data model." Your plan will not survive first contact with the business model as is. Be prepared to iterate on your plans in data mesh. But now, that doesn't break everything like it did for the enterprise data warehouse.
Similarly "no taxonomy survives first contact with the data." You have to build so you can update and iterate, prevent brittleness.
Don't use the phrase feature for your platform, use 'capability' instead. Make sure you are building the capabilities to support multiple use-cases and start to prioritize capabilities based on what use cases they can unlock.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Eric's LinkedIn: https://www.linkedin.com/in/ericbroda/
Eric's Medium: https://medium.com/@ericbroda
Phill's LinkedIn: https://www.linkedin.com/in/redshirts/
Liz's LinkedIn: https://www.linkedin.com/in/elizabeth-negrotti-calloway/
In this episode, guest host Eric Broda, an Executive Consultant in the financial services space (guest of episode #38) facilitated a discussion with Liz Calloway, a Data Governance expert in Financial Services (guest of episode #92), and Phill Radley, Principal Data & AI Strategy Consultant at Thoughtworks. As per usual, all guests were only reflecting their own views.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views. This will be the standard for panels going forward.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Manisha's LinkedIn: https://www.linkedin.com/in/evermanisha/
'A streamlined developer experience in Data Mesh' articles by Manisha:
Part 1 - Platform: https://www.thoughtworks.com/insights/blog/data-strategy/dev-experience-data-mesh-platform
Part 2 - Product: https://www.thoughtworks.com/insights/blog/data-strategy/dev-experience-data-mesh-product
Article on Lean Value Tree: https://rolandbutler.medium.com/what-is-the-lean-value-tree-e90d06328f09
Blog post on mentioned data mesh workshops: https://martinfowler.com/articles/data-mesh-accelerate-workshop.html
In this episode, Scott interviewed Manisha Jain, Data Engineer at Thoughtworks.
Some key takeaways/thoughts from Manisha's point of view:
Manisha started the conversation with her thoughts on how to get going with data mesh, on-boarding any domain but especially your first domain. Work with a small team aligned with that domain to find how data mesh can align with how the organization works and thinks - this will be different for every organization. That alignment is crucial to getting people comfortable and driving buy-in. People have to be comfortable with how it will work and what are their responsibilities. As Manisha said, "…only when they're comfortable with that concept … will [it] make sense to go ahead and explore more."
According to Manisha, when you are bringing teams up to speed, it's really crucial to get on the same page on what you mean and what you expect from them. They often confuse data and data product for example. The differences can be subtle but are important to understand. As Chris Haas also stated in his episode, they are using the Lean Value Tree method to break down target outcomes into explicit assumptions and more manageable aspects of work. What are the bets you want to make and what are the hypotheses you are testing?
Your initial workshop(s) with a domain can also be a lesson in how to deliver value using a data mesh approach and prioritization. Manisha talked about how when working with a domain, you might identify multiple potential use cases. But you need to choose what is a priority to do now and why. This can surface what are the top one or two use cases and also show the domain how to prioritize as use cases continue to emerge in the future. The use case(s) they select to prioritize then directly lead to discovering the data products that needed to support the use case. And then you identify what skills and tooling are needed to actually execute and build and then maintain the necessary data products. Then, you can start to back into what a team working on the necessary data products (and potentially platform) look like. You can use that Lean Value Tree concept to really get specific because far too often in data work, things are left too vague. Scott note: Get specific, get explicit, chase away vagueness - but of course leave LOTS of room for experimentation and iteration as you learn and build.
When asked more about workshop dynamics, Manisha shared how they try to keep them from being too heavy on the domain - get a few people, maybe 2-4, who really understand the domain and can represent the business aspects, not just the data and/or technical aspects. Each workshop has its own goal as an outcome but it's important to first align data mesh to organizational goals, the business strategy. Then you can get into data mesh specifics. They call their workshops 1) accelerate, 2) discovery, and 3) inception.
Manisha shared some crucial dynamics when working with your first domain that do get easier as you bring on additional domains. In the first domain, it's crucial to really narrow in on understanding and definitions including roles and responsibilities. Data product owner is a new role, what does it actually mean? And there's the initial platform work too. But as you bring on your third, fourth, fifth, etc. domain, there is internal learning to share with the new domains. There is more clarity around what a data product is - they can even see already built data products - and roles/responsibilities. But you will need to definitely do a gap analysis to figure out how to best enable each domain as each domain is unique. So there is a balance - look to maximize reuse of platform, processes, organizational changes, etc. but don't look to force new domains to adhere to exactly how previous domains went through the journey.
For Manisha, it's very important for the platform team to think in terms of capabilities. Deliver capabilities, not technology, to the domains. Work with early data product teams closely and focus on what they are trying to do instead of how you want to solve the technical aspects. Focus on specifically what are they trying to achieve? Also, the platform team needs to consider what mesh-level capabilities are necessary when. Don't try to deliver a complete platform at the start - your platform is a product and minimum viable product, make sure you understand what minimum means and don’t go overboard.
The platform team can focus on a few simple things to drive to a good initial outcome/partnership with domains in Manisha's view: 1) how does the work create business value? What do the domains need to do to actually drive value? 2) How will users trust data, what does trust mean and what's needed? 3) How do we make it possible for domains to create and manage a data product that is usable and discoverable? By focusing on the task at hand and then mapping to capabilities to support that task, you can prioritize and deliver something useful and valuable without boiling the ocean. You don't need to try to include every capability at the start, that is a bad anti-pattern. Get close to the use case and find friction. You will also learn to recognize reusable components of the platform but some reusable components might not be evident at the start.
Manisha then went further into finding and identifying reusable components. The things that are most unique to each data product are the data modeling and data transformation in her experience. Almost every other aspect of spec-ing out and building a data product are reusable, merely customized to the data product itself. Finding the necessary SLAs and SLOs by working with consumers, that is a reusable process. How your SLAs are actually measured, the definitions around those SLAs are reusable. The infrastructure and CI/CD is reusable. The overall data product blueprints are reusable. So look to make these reliable as your organization learns how to build data products to make for easy reuse.
On data modeling and interoperability, Manisha shared that it's crucial to let domains evolve how they model their data as they learn. And interoperability, especially to support a use case, is of course important; but you will likely see a need for interoperability standards emerge when it's needed - basically, don't try to build all your standards ahead of time. That might be creating an enterprise data model with a different name :)
When asked specifically about sample data models and automated data modeling tooling, Manisha pointed to them being a double-edged sword. While they can be helpful, most (all?) data products need more custom data modeling to maximize their value. Essentially, the tools can get to a decent initial data model but domains should look to improve them. If Platform teams offer automated modeling tools, they should make sure there is a big caveat to their usage .
Manisha recommends you make sure your initial domain has strong enough data talent - whether existing or embedded - to communicate the basic needs to the platform team. Regular developers are often not going to be data fluent enough at the start to drive to exact data infrastructure needs like a data engineer could. But be careful not to over index towards tech too. Every domain will need people skilled in creating value through data modeling but you probably won't need people as advanced in data infrastructure later - the platform is already built by that point :D
It's important to differentiate what the platform should offer and what the data product developers should handle according to Manisha. The platform, at least the aspects around data product creation, should be focused on making it quicker, easier, and more reliable to create, deploy, maintain, and evolve data products. It sounds easy but it's actually easy to lose focus on that. Look for friction points in the creation and management lifecycle and automate what doesn't add incremental value. E.g. a data product developer shouldn't have to manually add data to the catalog so look to automate it - and yes, not everything should be built upfront :) Scott note: she added some good flavor around data product boundaries but it's very hard to summarize
Within the platform, Manisha believes it's very important to maintain team boundaries because shared resources become a bottleneck and pretty quickly can become very hard to manage. This is why Zhamak has been so clear on the data product as an independently deployable unit of architecture. Manisha gave the example of even the namespace for data products in the data catalog should be reserved for that one team so teams have a dedicated space to put all their data products.
Manisha gave some early mesh journey advice:
1) back to data product specification, you should create something that gives teams a very clear idea of what a data product is and encompasses. Scott note: still waiting for someone to open source their data product creation template…
2) if, as Zhamak says, data products are our unit of value exchange in data mesh, then making it easier to exchange value is crucial. Start to create standardized input and output ports so you can easily ingest and serve data. ETL shouldn't be a concept, it's ingest or serving only.
3) really focus on making it easy to discover and then implement SLOs and SLAs. Being able to understand and trust data is crucial to being willing to rely on it. That trust comes from good communication around SLAs.
Manisha believes learning the language of the business is crucial for data people. You need to extract the actual business value drivers and build to those so you have to be talking the same language - unfortunately for data people, the language that aligns to business value is usually the business language :) Look to ask more business user-focused questions than trying to get technical.
Quick Tidbits:
"… the data product spec should [at a] minimum talk about the data set ports, domain, service level agreements, how do I share my data, what does data sharing look like…" - Make your data product specification easy to understand what someone will create and what a consumer will receive.
Again, focus on a streamlined developer experience that keys in on autonomy. That's the way to a scalable data mesh implementation at least on the platform side.
There is a responsibility on both the platform and the product teams to understand responsibilities and collaborate to drive to that...
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Key Takeaways:
Postel's Law: https://ardalis.com/postels-law-robustness-principle/
Semantic Diffusion article Zhamak mentioned: https://www.martinfowler.com/bliki/SemanticDiffusion.html
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Beth's LinkedIn: https://www.linkedin.com/in/beth-bauer-102449/
Beth's Website: https://posiroi.com/
Harvard Business Review article on 'The 3 Elements of Trust': https://hbr.org/2019/02/the-3-elements-of-trust
In this episode, Scott interviewed Beth Bauer, Founder and CEO, PosiROI. FYI, there are lots of nuggets in this one for people creating a data strategy or trying to tie your data work to value creation.
Beth's ADEPT^2 Framework (covered briefly near the last 10min of the episode): Analytics - Acuity - Data - Decisions - Engagement - Enablement - People - Processes - Technology - Trust
Some key takeaways/thoughts from Beth's point of view, much of which she helped craft:
Beth started the conversation with a bit about her background and then got to swinging a bit at current practices :) She talked about the need to have high-level data strategy vision but also be able to understand things in the weeds. Oftentimes those details in the weeds are important to being able to execute. Just don't get lost in them!
Setting a flexible, evolvable strategy and vision is crucial in Beth's view. You need to have a vision of where you want to go before you can figure out what actions you want to take. And then you need to look at milestone goals and break down what needs to be done when. Look to find incremental value delivery instead of putting all your value eggs in the basket that won't pay off until 2-3 years down the road. No one is willing to wait that long and needs will probably shift along the way. It's definitely okay to make long term bets that will require work to deliver value years down the road but focus on delivering value before that too. Priorities will change and so will your understanding of what will deliver the best value. So set out what needs to be done and when and do get going and be ready to reprioritize as you learn!
Beth talked a little about the cost of collaboration. If that is loosely coupled collaboration but not necessarily coordination on most to every step, great. But there is always a friction cost to closely controlled collaboration. Try to avoid work that requires too many dependencies - get aligned and work together towards a common vision. We really need to have trust that other parties can and will deliver. If you need to be working that closely, do you really trust each other?
Agile can be as much of a hindrance as it can be a help in Beth's view. If it's about taking the bigger picture and constantly breaking it down into achievable pieces, great. But if the focus is on completing work instead of driving to bigger picture value - if the work doesn't have an extremely tangible connection to value - you get lost in the cogs of the machine. You need a strategy that can handle how the cogs work but the cogs aren't the point of the machine - what are they actually driving?
Trust is made up of 3 elements according to an HBR article (linked above): relationships, judgment, and consistency. Relationships are over 50% of that trust. Beth believes data is essentially the judgment and the consistency but to get people to use the data we have, the data we can provide, we need to build relationships. You can't just show them the data, it doesn't even get them halfway to trust! Scott note: one comment I make is the difference between someone simply using the data and someone relying or depending on it. It seems like a small differentiation but it's not. One is using the brick as an accent to the building and one is building with the brick as a key element of structural support.
Beth pointed to how while the business folks don't need to know exactly how "the sausage is made" relative to data, they do need to understand more than just 'here's some data for you.' It's about sharing the necessary context. If you want sausage analogies, what flavor, what are the ingredients, what is the shape as in patty versus link, etc. They don't care about the data processing techniques but they should care about - and can gain value from - how was the data transformed from a business perspective. Scott note: this is where I talk about sharing information versus data. Just the 1s and 0s of data have no value without context as Beth said. So embed the context, focus on sharing the context, otherwise it is just values in a more complicated spreadsheet :D
Creating data sourcing strategies - at the micro and macro level - are important in Beth's view. Don’t overly rely on external data, that's costly. What data do you already have internally that you should leverage? What data could you be generating internally that you aren't? Dive into specifics and create a scalable way for lines of business to figure out good paths with sourcing data - internally and externally - going forward. Make sure to look at things from a cross domain lens too and also think about privacy, regulatory, etc.
For Beth, many organizations have trouble keeping the data work aligned to the business strategy. So there needs to be a specific focus on making the work matter, driving to business value. Yes, at the micro level but also on the whole - what business objectives and business outcomes is the data work supporting? Getting "down in the weeds" can also be very helpful, the details do matter as they make it “fit-for-purpose”.
On business strategy and data, Beth echoed the view many other guests have shared that creating a data strategy not aligned to the business is not a smart practice. But almost as egregious is not using data to help power your business strategy but this is extremely commonplace. Data is synthesized knowledge of what is going on in the world, often how the organization is interacting with the world. Why wouldn't you want to leverage that for shaping your strategy?!
In data, Beth has seen the benefit of the MVP (minimum viable product) methodology and they are wonderful if used correctly. However, they are often not used correctly :) Innovation doesn't have a steady timeline - it's messy. MVP timelines are tough so focus on getting to something viable instead of hitting a deadline - and communicate that to stakeholders. MVPs are about making sure you are on the same page and then iterating to better from there.
Beth talked about the need for continuously doing gap analysis with your data and business capabilities. The world is ever changing and new challenges and needs will constantly come up. Plus current capabilities can atrophy. You might be hindered in projects because you need some special capability - especially think legal/regulatory compliance - and you should know that _before_ doing the work :)
Two key questions Beth uses when people ask for data/data work: 1) what are you going to do with this? And 2) How much value do you think this will generate? You don't need to get super specific but people need to at least have a good idea of what the work will unlock and the value of that. If it won't cause any action, why do the work? If the cost outweighs the value, that should be known so you can work to balance that equation by cutting costs and/or finding more value.
For Beth, to do data right, we need shared responsibility. There is the technical piece of course but the business aspect is just as important. "…we need to realize that nobody's anything without each other." We need to drive to address current gaps and we need short, medium, and long-term strategies that drive necessary work in the short, medium, and long-term. Don't get overly focused on the near or the long-term. But it's not just about doing the data work, especially the technical data work. What value and change does this work actually drive?
"And what I found is that largely, a lot of organizations, the challenge is, with really good data management comes really good transparency into how things work. And that really causes pushback on the power structures, and particularly in the 'how it's always been done' power structures. Because if it now points to a way that things can be done better, you start to get into things. Things are happening behind the scenes that have nothing to do with data, and everything to do with people's perceived value of themselves to the organization - without (them) thinking about how they can evolve to actually move from what they're doing today to doing it better." Scott note: it's crucial to help people see how they can move to doing more valuable things - their time of toil is behind them and we can unleash the value creation :)
Quick Tidbits:
In data, there is often a rush to get things done instead of get to automation. It's often - but not always - the right call to slow down to do things right and set yourself up to do them faster/better/more scalably as you move forward.
Lineage, especially for how data was generated and transformed from external sources, is really crucial to increasing trust in that data.
"Data for the sake of data is useless." Scott note: PREACH!
"The digital transformation, to me, is the biggest misnomer and misguidance that we've ever created." A transformation has an end. This is a journey.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Here are 8 more leverage points to consider when trying to work with your initial domain
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Guy's LinkedIn: https://www.linkedin.com/in/guytaylor/
Deep Work by Cal Newport: https://www.youtube.com/watch?v=xJYlhhT7hyE
In this episode, Scott interviewed Guy Taylor, Director of Data Science and Analytics, as well as the Director of Experimentation at Booking.com. To be clear, Guy was only representing his own views on the episode.
Some key takeaways/thoughts from Guy's point of view:
Guy started off by talking about data literacy and how the analogy of literacy - it's not only the ability to read but also write - carries over to data well. Data literacy or data fluency, it's not just can someone consume data but can they also produce data, can they share information in a way that can be 'read' by others? After all, we aren't trying to share data, we are trying to share information but via data.
When Guy embeds people from his data team into domains it "is with the express purpose of doing education, making sure that we are having the conversations around what things mean, what our expectations of those things are." Instead of embedding people into domains to do most of the work, they are focused on helping other people get to a level they can handle far more of the necessary data work. Which is quite often not the deep data work but bridging the communication gaps and getting on the same page. That is especially important for expectations - mismatched expectations is one of the most prevalent and damaging challenges to data work. So Guy is asking data team members to spend a lot of time making sure the producers know how to manage those conversations and drive to what is actually of value instead of what was initially requested.
In his experience, there is tendency for data people to try to jump to the tooling to solve issues according to Guy. Going back to expectations, if you try to solve without the expectations setting and leveling conversation, you will likely not deliver what consumers expect. You may see it as solved but they sure don't. That's where you get the dreaded "the data is bad" feedback because there aren't clear metrics and expectations. If you align - and as Ghada Richani mentioned in her episode (#206) stay aligned through collaborative prioritization - then there is a much better chance of delivering value and making all parties happy.
A comment Guy made was that there is an over tendency towards action in tech and especially data. People see a problem and they want to jump to trying to fix it instead of getting the necessary information first. And it may be no action is the best answer too. Just because there is pain, that doesn't mean action is necessary immediately.
Right now, Guy sees the industry conversations around data contracts and data sharing agreements as slightly naïve: people seem to be thinking this is about data integration between systems instead of data sharing between two parties. And that producers should declare every aspect of what they are producing instead of consumers being part of the conversation. Consumers need to share what they are trying to achieve, how they will use the data, etc. so producers understand the value and what would disrupt that value creation. There needs to be accountability and responsibility falling on consumers too. The contract portion can serve as the technology interface but that doesn't replace the need for conversation.
Ethics in data is always going to be an interesting but challenging problem in Guy's book. A good place to start is the social contract aspect: how would this be viewed by society? As an organization, start down the ethics path by creating and agreeing to a set of principles. Create good ways for people to seek and receive useful feedback regarding ethics. And honestly, your company ethics will change and it's important to reevaluate your ethical choices, especially as you learn more - your organization will have made mistakes and that's typically not something to lose sleep over, fix it now and know you’re better. Basically look to "do the right thing."
While it can feel good to 'make progress', Guy believes in the Deep Work by Cal Newport type philosophy. People have the greatest impact when they have the time to really think and process. Yet, in today's work world, that is a rarity for anyone. If we are asking teams to really take on data ownership, we have to work to prioritize the time to learn - and that includes processing time. Yes, people learn by doing but not only doing :)
Guy talked about trying to clear the space for teams to learn something new, including the impact to the cognitive load capacity of teams, especially when it comes to data ownership on the domains. When people know there are expectations of them and their work but those expectations aren't explicit or clear, that's unnecessary cognitive load - domains need to have crisp and clear expectations - and if the expectations of them by consumers and the data team aren't super clear yet, communicate that. But try to get to more detail to make the implicit very explicit.
Another common friction point Guy pointed to is the lack of understanding of your impact on others. He believes again that most people want to "do the right thing" but they don't always know there is even a problem to address. How are your actions impacting downstream data consumers? Again, there is a responsibility on those data consumers to generate the conversation! Sharing that context allows people to do the right thing because they are aware.
Currently, most teams in most organizations seem to be overloaded with work and cognitively overloaded in Guy's view. It would be lovely to wave a magic wand and fix that but it's not possible. So, we have to work to pull out the complexity and give our teams the ability to do their best work but high value work has interconnected complexity so you can't take a machete to it and expect good results. Look to prioritize what really matters when you can and break things down into manageable chunks.
In talking again about data ownership, Guy believes that producing teams should not have to own too much of the downstream consumption. They should own the transformation and sharing of data after it's been transformed but the consumer should own the insight or metric - they asked for the data so they need to have some responsibilities and accountability too. Scott note: I don't entirely agree it should be that for every use case but I do like a very crisp line of ownership if it works.
Guy uses the Marie Kondo approach when looking at orphaned systems or processes - evaluate it and is this "sparking joy", is it creating value. If it is, great, let's get it into a good shape and put it into the right hands ownership wise. Not shove something to someone while it is still in data disrepair but get it functioning well and hand it over. But if something isn't creating value, shut it down. For too long in data, there has been a hesitancy to shut things down because at some point they might create value. Don't fall into that trap.
It's important to get to an experimentation - a test, learn, and then iterate - approach in Guy's view; he is the Director of Experimentation after all! A good way to get people to see the value of experimentation is how it provides far better incremental value delivery. Instead of huge projects with big budgets that take years and rarely deliver the value expected, what if instead you took the overall goals and broke it into manageable pieces and delivered value over time as you get closer and closer to the project vision. You'll be more nimble and get a return on investment far quicker.
A few tidbits from the end of the conversation:
Skunkworks can be a great approach to trying things out and seeing if there is value. Don't try to move the skunkworks directly to production but you can do some fun and useful innovation that way.
Good leaders set their teams up to innovate. They prioritize the time to try new things and let their people explore, let them go off the "paved road".
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Kim's LinkedIn: https://www.linkedin.com/in/vtkthies/
Mike's LinkedIn: https://www.linkedin.com/in/2mikealvarez/
Ferd's LinkedIn: https://www.linkedin.com/in/ferdscheepers/
Omar's LinkedIn: https://www.linkedin.com/in/kmaomar/
In this episode, guest host Kim Thies, Director of Intelligence Automation at PayPal facilitated a discussion with Ferd Scheepers, Chief Information Architect at ING, Mike Alvarez, Former VP of Digital Services at a large healthcare distribution company (guest of episode #236), and Omar Khawaja, Head of Business Intelligence at Roche (guest of episode #96). As per usual, all guests were only reflecting their own views.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views individually.
Before we jump in, I think the main takeaway here would be a data mesh implementation leader's journey can be a lonely one. Find peers and exchange information. You can reach out to me (Scott) but there are also many leaders that want to exchange information with each other. The other is the meaning of journey: it's never done; be prepared to continue to push - it can feel Sisyphean but it's important to keep moving forward and expect to continue to drive buy-in.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Marcie's LinkedIn: https://www.linkedin.com/in/marcie-stoetzel/
DGIQ Conference (Marcie and Samia Rahman will be speaking in June; in San Diego, CA, USA): https://dgiq2023west.dataversity.net/registration-welcome.cfm
Square hole video Scott mentioned: https://www.youtube.com/watch?v=cvwN_O5ypTE
Article about the "leaf sheep" sea slug Marcie mentioned: https://www.bbc.com/travel/article/20210324-the-odd-sea-creature-powered-by-the-sun
In this episode, Scott interviewed Marcie Stoetzel, Principle Product Manager of Enterprise Core Data at Seagen. To be clear, she was only representing her own views on the episode.
Before we jump in, Seagen might be a bit of a special case in how much the domains can leverage each other's data around a key record type. There is a lot to learn but you might not be able to find use cases that are as broadly impactful to many domains at once.
Some key takeaways/thoughts from Marcie's point of view:
Marcie started off with a bit about her background as a teacher and as well on the commercial side of the healthcare space which has shaped her view of teaching/learning and also healthcare data needs. Part of Seagen's goal in using Enterprise Core Data instead of Master Data Management (MDM) as a phrase was the moving away from the connotations with slavery (master) and that this data is core to the enterprise, that this is crucial to the organization, not just a data management practice or task. Enterprise Core Data at Seagen is about creating a way to make data that many domains leverage the same to prevent lots of domains doing the same work and make interoperability FAR easier. MDM also is typically managed centrally instead of enabled centrally and managed at the domain level (federated governance) so you have to rethink a lot to do Enterprise Core Data instead in a data mesh setup. Trying to map 1:1 to MDM in a historical data approach won't work well.
At Seagen, Marcie and team decided to tackle one type of core data record first - healthcare professionals - rather than trying to unify every type of record across healthcare - don't boil the ocean or bite off more than you can chew. They are working on creating the platform for domains to manage their core data records which creates more sharing opportunities and even higher quality data - teams can better cross-reference information. While they are still pre-production - it's early days - even bringing this to domains' attention is sparking conversations between domains about potential collaboration and new use cases.
Marcie and team are winning converts by showing them what the platform will be able to do for their own domain but also keeping an eye on that interoperability and leverage provided to other domains. That means that domains get some value from participating even if no other domains participate. If other domains do participate, then everyone gets more value from each other. They started with a simple value proposition - this will make handling your own data easier - and then created a group collaboration incentive - the more domains that participate, the better the information and the less work everyone has to do to get to better outcomes :)
When asked about incentivization complications around domains wanting to focus on their own goals, Marcie mentioned that as the domains are starting to find cross-domain use cases, the organization can realign goals to be about focusing on those common goals where everyone wins. What drives the most value for the business and how do we incent that kind of outcome/behavior? Scott note: this is an interesting nuance that I haven't heard before.
How they found the cross domain use cases was also interesting. Marcie and team met independently with different business domains to extract what each team felt could be a good output of working with a common enterprise core data platform, as well as tying these individual value props back to the larger goal of helping more patients. The data team then literally took all those outputs and showed each business domain that uses HCP (healthcare professional) data the value across Seagen of having an enterprise core data platform. This sparked collaboration ideas between domains for how to drive even more value from the core data platform.
Marcie is seeing that domains have to spend a LOT of time to cleanse and match data. And downstream consumers of their data have to do the work too as they don't know it's already been done upstream. The quality requirements for most use cases are pretty high. So creating a way for domains to much more easily interoperate data will save them a LOT of time and effort. The core data platform will hopefully prevent lots of domains from having to do a lot of that quality checking work and it will also increase quality by having more sources of information to verify data is high quality / correct - basically, the checking of quality becomes far less arduous.
One thing that is working well at Seagen for Marcie and team is doing lots of demos and small proofs of concepts. Similar to doing sprint demos, teams are buying in because seeing is believing. They are showing the business real, realized value that will come from their participation. So teams are leaning in. Similar to what Karolina Henzel mentioned in episode #104, there can be a LOT of value in addressing data quality issues for domains.
Marcie talked about the value of maintaining focus on a thin slice. Instead of trying to get many domains bought in on the data mesh concept, there is a specific use case and they are only working with a few domains at the start. There are clearly defined and scoped benefits. Again, thin slice. And focusing on the end-to-end solution to working with this data for the domains has also helped to get and keep everyone on the same page.
Quick Tidbits:
Look for ways to share knowledge in fun and interesting ways. Upskilling can be a bit intimidating, make it more gamified and less high pressure. Humanize (or in Kye's case, dog-ize) it a bit.
Really embrace an attitude of learning - be vulnerable and transparent. "Explore, discover and mature data together." "Engaging conversations, exploration, curiosity, and a safe space" are crucial.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Key Takeaways:
Semantic Diffusion article Zhamak mentioned: https://www.martinfowler.com/bliki/SemanticDiffusion.html
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Jyotshna's LinkedIn: https://www.linkedin.com/in/jyotshna-karki-81a24038/
In this episode, Scott interviewed Jyotshna Karki, Data Engineer at Novo Nordisk. To be clear, she was only representing her own views on the episode.
Some key takeaways/thoughts from Jyotshna's point of view:
Jyotshna started off the conversation with a bit about her background, especially in data engineering and the need to be and stay curious. There are so many new approaches and technologies that could provide significant benefit to consider. Think in that product mindset and look to evolve your approaches and tech stack to create more value.
Specific to Novo Nordisk's data mesh journey, Jyotshna and team saw the writing on the wall for their data lake setup. While their centralized data lake was doing well and people were happy with it, there were increasing consumer and producer demands and the central data team was still required to help teams create their data products. Having a centralized team in the middle of every use case just wouldn't be efficient. Then, they hit some cloud service limits which caused some major headaches as well. All this led to looking to decentralize via data mesh.
At Novo Nordisk, many domains already had significant data capabilities and there were people building data products anyway according to Jyotshna. What they really needed was a way to empower and enable teams to more easily create and manage those data products in an interoperable way and lower the bar. So the central data team was to focus on the platform but there wasn't a huge need to upskill all the domains. There is still another centralized team of data experts to help domains that aren't as data fluent. Scott note: while this is not super uncommon, most organizations are not this lucky :D
Specifically to the pharma industry, Jyotshna shared some of the pre data mesh compliance/regulatory issues that were better addressed with data mesh. Domains needed to work with regulators but it was hard for them to really see exactly how the data was stored as it was managed by the central team, which is part of compliance. It was all part of the central data lake AWS account and those teams didn't have the ownership or visibility they needed. But with data mesh, the teams now have the visibility to their own data storage and access to audit logs and data governance.
Jyotshna shared that at Novo Nordisk, there was so much demand to participate in their data mesh, the data platform team and any centralized data capabilities - to assist the domains that didn't have high data fluency - worked with multiple teams to start. This helped them to define the requirements for their data mesh platform to support multiple data domains. While this is a data mesh anti-pattern, it went well for them as many of the domains again were quite capable with data engineering and data analysis. There were also many domains that wanted to contribute aspects to the platform so there were good feedback loops between the platform team and many domains. Scott note: Don't go this route unless your domains are already highly data fluent/capable. Working with many domains at the start can create a high-risk scenario instead of thin slicing.
Jyotshna and team are focused on enabling proof of concepts more than trying to automate everything right at the start. She noted they are focusing on understanding the problem deeply and moving fast to get proofs of concept into people's hands and then circling back to automate when there is more need and things are slightly more stable. Basically, they are being agile. It also has led to more modular components and reusability - they can get things out in a prototype phase and then think bigger picture how to deal with similar problems instead of point solutions.
In order to prevent tight coupling and keep modularity, Jyotshna and team started to actually remove things like data pipeline blueprints, reusable components, and bootstrapping accounts from the data mesh platform. While that might feel counterintuitive, they wanted to create a community specifically around things like blueprints so the central team wasn't managing them, community members were managing them. Look to create central sharing mechanisms but with a decentralized ownership and contribution model. Community-led innovation is more scalable than centralized knowledge ownership.
When thinking about platform maturity and if they need to pay down any tech debt, especially around certain features, Jyotshna and team benchmark quality levels and compare those to the actual business needs. Being in a heavily regulated industry, some aspects of compliance just are non-negotiable, you must meet them. But there are places where 'good enough for now' is a completely acceptable and correct answer. Some signals they use are support tickets and direct feedback around different aspects of the platform. They are also starting to build KPIs but it's a work in process.
One interesting aspect of doing data mesh has been less duplication of work per Jyotshna. This is a target goal of data mesh of course but it came about naturally as now that people can reliably find and access data, they don't feel a need to build it themselves.
Jyotshna said "make sure that it is reliable enough for people to depend on this data". Part of your platform and your overall mesh is to make it easy for consumers but also producers to trust the data. If you have a black box process, can producers really trust it? And evolution of your data products plays a part in trust too - a consumer can trust that the way data is presented is still relevant to the business.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
3 use case aspects or patterns to lean away from when trying to work with your initial domain:
And the 4 you should look to lean into:
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Chris' LinkedIn: https://www.linkedin.com/in/christopher-andr%C3%A9-haas/
Article on Lean Value Tree: https://rolandbutler.medium.com/what-is-the-lean-value-tree-e90d06328f09
In this episode, Scott interviewed Chris Haas, Advisory Consultant at Thoughtworks. To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Chris' point of view:
Chris started out with a rather blunt but crucial statement: when thinking about data mesh, you have to identify at least one domain that will actually take ownership of their data. A successful data mesh implementation can't be entirely IT driven. And domain data ownership and coupling that with data as a product are typically very much not the natural order for most large organizations. You will need to invest into domains to make them capable of owning their data - showing that you will invest in their success can help win them over - so you want to make sure it will be money/resources/time well spent.
For Chris, it's very important for there to be at least a domain-wide understanding - and hopefully buy-in too - for what the domain is trying to do around data. That can be about data mesh or simply how they are changing their relationship to data. It won't work well to do data mesh from just upper management buy-in and pushing that down as a mandate. And that requirement can lead to deciding that data mesh isn't right for a domain and that is perfectly okay and reasonable - not every data management challenge is a nail for a data mesh hammer.
Lean value trees are an important tool for Chris and team when speaking with a domain about data mesh. What are they actually trying to achieve - the mission - and work backwards from there. Break down the goals, then figure out the assumptions or bets around those goals. This helps you stay focused on what you are trying to achieve - is it deliver a data product or is it address the use case and create business value?
So the output of a lean value tree helps align the team on a mission or missions which align to the business goals. Then, when thinking about use case prioritization, you need to balance those business goals, the amount of expected value of a use case, and the expected amount of work to deliver that use case*.
At the start of a journey, Chris recommends to find a use case that benefits the producing domain if at all possible. Yes, we want domains to publish data to benefit the entire organization but if a domain is going to be the test subject and invest their money, their people's time, their people's cognitive load, etc., it will be hard to find a domain willing to do that for another domain without decently strong incentivization. And at the start of your journey, those incentivization and community mechanisms are likely hard to come by. If you don't have these challenges and domains are all happy to help each other, consider yourself very lucky.
Based on interactions with a number of clients and prospects, many organizations would rather focus on the technical aspects of data mesh first over the organizational and Chris wishes that weren't the case. While building out the technical aspects is no easy task, if you aren't ready to actually do the day-to-day work, what are you building the tech to support? Scott note: this is extremely common and is also a very common comment from consultants. The tech feels more tangible and it's easy to say yes/no than squishy operating model discussions. But they are crucial to doing data mesh right.
Chris believes an early data mesh alignment on the organizational model doesn't have to be - and shouldn't be - disruptive. You can have a domain start to carve out new ways of working without doing a major reorganization. There should be a team building the platform and a product owner for the platform but 4-5 people building is reasonable and they can live in a central data team, that's a normal pattern, no reorg needed. As for the data product team, you want them to be part of the domain if possible and it should be a mix of highly data fluent people and domain subject matter experts but again, it can be a small team. So a small team carve out but not realigning the entire domain. Scott note: remember domain is an overloaded term. Some domains are sub-domains that are 3-5 people but typically we mean the line of business which can be 10K+ people.
For Chris, when asked about data product teams, he said you should look to have them be long-standing teams. You need an owner per data product but often the development teams will be charged with creating multiple data products rather than only one. Especially as you build the knowledge of how to build good data products, that team can become more efficient. But do not treat your data products like projects, they should be managed/evolved by long-standing teams, otherwise they are more likely to fall to data disrepair.
As a consultant, Chris recommends working with a consultancy :D but the point is that your data product team in the domain should be considered a new team and that new team needs a budget. It might be reallocated budget but hopefully it's new budget. But the consultancy angle is that if the implementation doesn't deliver value, it's easier to cut the consultancy than reassign or lay-off employees. Scott note: I think this is a valid concern but many can't get the budget and many organizations are doing this work entirely internally.
Once you've proven value from and a capability to do data mesh, a typical pattern is to set up a transformation office that will assist other domains in their journey according to Chris. That way you can maintain coordination, find reusable patterns - tech, architecture, people, process, etc. patterns -, prioritize, upskill, etc. That office is helpful in getting teams up to speed but also setting expectations - domains won't magically transform their data capabilities overnight. And look to build an internal community function to encourage cross domain communication and hopefully collaboration.
For Chris, that community aspect is really crucial to identifying cross-domain use cases. The data product owners should be communicating with each other to discuss what they've created in case that sparks new ideas for new use cases. For finding what data products you need to support a use case, start from the mission, then the consumer-aligned data product that would best serve that mission. Once you know what you'd want to have as the end product, you can start to find the necessary source-aligned data products.
Chris wrapped up with a few useful tidbits:
Don't worry about trying to serve future, unknown use cases with data products you are building now. Build for reuse and build for evolution and then evolve when those new use cases emerge.
Source-aligned data products should be - or at least start out - "very specific and very small." Don't try to cram too much in. Scott note: I think this can create some discoverability issues but it is a pattern I am seeing more. See Carlos Saona episode #150 for how that looks in an implementation
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Ammara's LinkedIn: https://www.linkedin.com/in/ammara-gafoor/
Ammara's Articles on data mesh Part 1: https://www.thoughtworks.com/insights/articles/data-mesh-in-practice-getting-off-to-the-right-start
Part 2: https://www.thoughtworks.com/insights/articles/data-mesh-in-practice-organizational-operating-model
Part 3: https://www.thoughtworks.com/insights/articles/data-mesh-in-practice-product-thinking-and-development
Elif's LinkedIn: https://www.linkedin.com/in/elift/
AtScale article on data mesh: Principles of data mesh and how semantic layer brings it to life: https://www.atscale.com/resource/data-mesh-principles-semantic-layer/
General AtScale article: The Semantic Layer’s Critical Roles in Modern Data Architectures: https://www.atscale.com/resource/the-semantic-layers-critical-roles-in-modern-data-architectures/
Ryan's LinkedIn: https://www.linkedin.com/in/ryandolley/
Super Data Bros Blog: https://superdatablog.substack.com/
Super Data Bros YouTube: https://www.youtube.com/@superdatabros
In this episode, guest host Ammara Gafoor, Principal Business Analyst at Thoughtworks (guest of episode #133) facilitated a discussion with Elif Tutuk, Global Head of Product at AtScale, and Ryan Dolley, Independent Data Consultant and one of the SuperDataBros (guest of episode #183). The focus topic area was what role does Business Intelligence (BI) have in data mesh - e.g. where does it sit and who owns it - and how do we enable BI to really drive significant value in a data mesh implementation.
A few other episodes that would be good to get a broader picture here on related topics in addition to Ammara and Ryan's episodes are #199 with Brent Dykes and #192 with João Sousa.
Scott note: I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views. This will be the standard for panels going forward.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Srinivas' LinkedIn: https://www.linkedin.com/in/spaluri/
In this episode, Scott interviewed Srinivas Paluri, CTO at Rentbase. Srinivas was previously part of a data mesh implementation as the Senior Director of Data Engineering at Zillow.
Some key takeaways/thoughts from Srinivas' point of view:
Srinivas started by saying how important architecture is, yes, but how it's too hard to try to get everything right up front. Instead, set yourself up to be able to change your architecture as you learn more. Fail fast is crucial, you need to be able to get moving sooner rather than later. Otherwise, there is too much risk in building something that doesn't fit needs or even more likely, never building anything because you're always waiting for the perfect technologies.
To do engineering - especially data engineering - well, Srinivas believes engineering needs to be part of the decision making process. If not providing input, at least then being aware of strategic shifts and working with key stakeholders to shift systems to better align to new changes. You can't change your business model and not expect it to not require a shift in what data you need and how you work with it! And you could have started collecting data sooner for the business model shift too :)
As data mesh aficionados know, centralized data engineering can create bottlenecks - and almost certainly will at scale. Srinivas recommends you show people the impact of these bottlenecks through specific examples. You can use that to drive better buy-in for decentralized, domain-based data ownership. It can be hard to quantify the exact impact but it's important to try to at least communicate the outcome of those bottlenecks, how did it impact day-to-day business operations and capabilities.
When trying to push data ownership to domains, it's crucial to make sure you give them the support to actually own the data per Srinivas. That's technology and that's capabilities/understanding. If you are saying data ownership has value and you can show the value, then the organization should support that value, it's not free, it's worth investing in.
Srinivas believes you should focus on making gradual changes rather than sudden shifts in a data mesh or other large-scale data implementation. There needs to be commitment to making change to do it right. And as part of that, look to create and foster close collaboration with users. You need teams to be blunt and honest to help you get to where you need to be, to get to a valuable outcome. That feedback will help you improve your processes and platforms.
If Srinivas could give 3 pieces of advice to his former self about data mesh:
Another thing Srinivas learned looking back on his time at Zillow is to focus on communicating to people what does data mesh change for the organization and especially for them. There is a vague sense of data mesh changing the way we work but what is the actual value we expect to drive. Not a specific dollar amount but what's the vision of the organization of once you reach a relatively data-informed, data-driven state? Faster reactions to market changes? Better identification of new opportunities? Significant cost savings? And again, talk to people about what changes for them, driving value and also responsibility/role wise.
If you want a somewhat visceral approach to showing people why the central data engineering team has become a bottleneck, Srinivas recommends asking how long would it take to clear your current backlog with your current team if no new tickets came in. What about how big of a team would you need to actually clear your backlog based on the number of tickets coming in - is anyone's backlog actually decreasing? What could your organization be doing if you didn't have that bottleneck? What more value could you be creating by really enabling teams to understand and leverage the organization's data? Yes, there will be a cost but we have to invest to create value :)
When driving data mesh buy-in, Srinivas again went back to the need for product management in general. Talking with the producers about why this matters; how you are shifting their prioritization, not just adding additional work; how they can actually achieve this and what are meaningful milestones / incremental value deliveries; etc. Prioritization - and communication around prioritization - are likely to lead to your easiest, happiest paths.
Data mesh product management should not only focus at the micro level - the data products and the platform - but also at the mesh level according to Srinivas. To really drive value at scale, you need interoperability and things working in harmony. Product management is not just about managing the product but the entire suite of products. Think about your data products as an overall suite of information to serve many use cases and drive a huge amount of value.
For Srinivas, when looking at setting your priorities for data work, start with what are the overall organization's priorities, what are the priorities of your business partners? If you want to do data work that isn't valuable to them, will it be valued even if it drives the expected value? Look to find data work that are both valuable and valued and that support the overall organization's priorities. Scott note: this is especially key as you are getting going on a data mesh journey.
Everyone on the data engineering team should understand how their work supports the company's priorities and drives value for Srinivas. And the exec team should understand how the data engineering work maps to value creation too. Sometimes, that can be harder with platform work and the like, but it's important for everyone involved to understand how the data engineering work drives value. That helps map to prioritization decisions as well.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Takeaways:
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Ghada's LinkedIn: https://www.linkedin.com/in/ghada-richani/
In this episode, Scott interviewed Ghada Richani, Managing Director, Data, Analytics, and Technology Innovation at Bank of America. To be clear, Ghada was only representing her own views on the episode, not that of the company.
Some key takeaways/thoughts from Ghada's point of view:
Ghada started discussing balancing speed, structure, and control in your data mesh implementation. There are those that want to build everything upfront and boil the ocean but there are also those that want to get to value as soon as possible without taking the product mindset to heart. Work with both sets of people to keep them deeply informed and show them why a balanced approach works better. If stakeholders are very close to the journey, they won't be pushing back on timeline - they can see where prioritizations are changing and the learning is happening as data products - or other aspects of your mesh - are being built. In fact, let them control prioritization where it makes sense so they are the ones causing timelines to stretch and they made the tradeoff decisions.
Keeping stakeholders closely informed also has benefits around control in Ghada's experience. They can understand the tradeoffs relative to governance challenges like regulatory compliance. Exposing the actual evolution of the data product itself to stakeholders helps stakeholders feel comfortable with the process and that compliance/regulatory concerns are addressed.
For Ghada, starting from requirements for a use case doesn't work well, people aren't sure and they get stuck in the details instead of the big picture. Instead, work with them to focus on what they are trying to achieve, what their deliverables are, and then work backwards to figure out what they need to meet their own deliverables. And those deliverables better be tied to value somehow :)
While driving buy-in from data producers, Ghada recommends making them a clear stakeholder in the process. She's found that really deeply informing them of how their data will be used and the value it will drive often gets them excited to participate. Of course, you need to work with them to prioritize the work but showing them the value - or potential value - of a great use case often helps set that prioritization. You also want to make sure to highlight their work, either for them or preferably making the stage for them to present the value delivered from their data, giving them credit and visibility. When you do those things and give people true ownership, not just requirements, many data producers are far more willing to get involved.
In general, when trying to get approval for data work, while Ghada recognizes it can be hard, she has a few good approaches. One is to look at different aspects of what you are trying to improve. Say a process or product line drives significant value for the company. What could you do to tangibly improve the value it delivers? Not as one giant project, break it down into more tangible improvements and seek a budget to tackle one or a few so you can prove out value, getting a budget for additional improvements. Another is to know your audience. This might seem simple but really, you have to learn what drives your counterparts and find a way to communicate the benefits in their own language and address something that matters to them. Make it digestible and hard to resist wanting to tackle the challenge. It's definitely more art than science.
One way Ghada has found to drive buy-in from reluctant data producers is to assign the cost of not doing something to them. Essentially, there is a benefit, a value to doing the proposed work - whether that is increased revenue, decreased cost, decreased risk, increased speed, etc. So, there is a negative of not doing the work the proposed work and you ask the reluctant data producer to officially own the cost of not doing that, own that business risk. Many have become far less reluctant to participate :)
For Ghada, there are two ways in general to measure the value of data work - economic value and impact value. Economic value is slightly easier to conceive if not that easy to measure - if you make improvements to a process or say create a new product line, you measure the incremental revenue it drove or the amount of cost savings. Impact value, the team(s) impacted by the changes have to give the value measurement - what is the value of speeding up a process, improving the data quality, lowering the associated risk, etc. Neither are exact measurements so it's crucial for stakeholders to understand that it's about triangulating and assessing value, not an exact amount of return. And the stakeholders again have to be the ones that assess value. Only they can say what an impact would mean for them, the data team doesn’t have the context to do that. And you need an organizational environment where the forecasts are seen as forecasts, not commits.
It's okay to have failures in data work, according to Ghada. As many past guests have also noted, experimentation is about trying, learning, and iterating. Sometimes the learning is that this won't work or isn't worth the effort. Getting to that learning quickly and iterating to value - or stopping work when that's the right call - is crucial to driving significant value from data work. Your culture must allow for failure or you just won't take on initiatives that are higher risk but higher reward and where the reward justifies the risk. You need to see getting to value and getting something directionally right as a win so you can iterate towards more value.
In Ghada's experience, for high profile, high visibility, high intensity projects/data product builds, it's not unusual to check in 2-3 times every week with all the stakeholders. While it may feel like overkill, you can find miscommunications or friction early and even more importantly you can identify and work to address challenges and risks as they emerge, e.g. if someone is disengaging. Instead of the data team going off and doing a bunch of work to deliver at the end of a months-long project, it's tight feedback loops and iteration and changing priorities through close collaboration. And have a highly visible accountability model - if someone isn't delivering, that should escalate to the executive sponsor to figure out prioritization and an appropriate response.
On the platform side of things, Ghada is very happy with their use of data virtualization for their virtual query layer. As teams have learned how to build and mature their data products, data virtualization has meant they can expose what a mature data product looks like even when the underlying data product is not yet mature. The underlying data creation and curation process is not fully productized or robust in many instances but consumers don't have to care. The views presented to users are controlled by subject matter experts and serve as a type of interface or output port of a sense.
More on data virtualization:
1) Sometimes, that virtualization layer can lead to query performance challenges but usually, that's tied to someone trying to do too large of a query all at once instead of breaking it down appropriately.
2) Data virtualization has made exposing connections between data products much easier. It's just creating another virtualized view. Connections need to be discovered/surfaced manually but beyond that, it's quite easy to do interoperability if the data fits well together.
Discovering and mapping domain boundaries is really crucial in data mesh according to Ghada. And it will get easier as you go along. You really want to consider what are you trying to accomplish with a data product and not have many things loaded into one data product, or it will become overloaded and be hard to evolve/improve well. Data owned by a team that are not the subject matter experts is a likely occurrence but you should look to rectify it quickly. Teams building data products that consume information from upstream data products should not take unnecessary dependencies. At BofA, they create a virtual view that combines the upstream data from the source data product with the data product rather than that downstream data product taking a dependency. They are also having many domains that are represented by one data product but some have more than one data product. The boundaries and the governance are far more important to get right than trying to match a certain number of data products to domains.
Every data product should have a defined purpose in Ghada's view, that's how you can find your data product boundaries. But a data product should also not take on additional purposes, that's scope creep. That doesn't mean it can only serve a single use case, reusability is crucial but when someone tries to find the right source for accomplishing a goal with data, it's best if they have to consider fewer options but still get all the data they want/need. Yes, easier said than done :)
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Stephen's LinkedIn: https://www.linkedin.com/in/galsworthy/
In this episode, Scott interviewed Stephen Galsworthy, former Head of Data at TomTom. Obviously, given he's no longer with the company, he was only representing his own views.
Some key takeaways/thoughts from Stephen's point of view:
Stephen started off with a bit about his background from working for companies consolidating information into a product they sell, such as TomTom taking in tons of data to create maps to sell. However, using data as a key aspect of your products doesn't necessarily translate into collecting and analyzing information about how your products are actually used which would allow you to drive improvements like new features or even new products. Organizations that are focused on obtaining and analyzing information of usage and other market dynamics are likely to be those that win in the market more often going forward. Their user experience knowledge is going to be a much tighter feedback loop.
The AI flywheel is something Stephen mentioned that can create a virtuous cycle. You are creating data around interaction points with your products that feed AI to make the products better. The better they are, the more people interact with them generating more interaction points, meaning you can make your product even better. Essentially, collecting and analyzing data to make your product better -> you make your product better -> more people use the product -> more data to analyze to make your product better. However, if it's not inherently part of your business model, trying to pivot to that information gathering practice can be a tough sell for customers and/or partners. Oftentimes, partners aren't even allowed to hand over data. Think about how you'll effectively collect data as early as possible even if you don't start collecting it then.
A good way to build momentum around the importance of data to an organization is start small according to Stephen. It might sound great to try to convert the entire organization to being data driven at once but Rome wasn't built in a day. Show some successes from working with data, find something that couldn't have been done without leveraging data like a new product, and then share those successes internally and encourage more parts of the organization to try leveraging data.
Stephen believes that data should rarely be THE deciding factor in decisions. The human in the loop is crucial. And it's crucial to make that clear to execs - you don't want them to think data is a magic wand / silver bullet but they also don't want the data making the decision. Data is a touch-point in decisions, it's used to create better feedback loops as you iterate towards good solutions. Leveraging data well often isn't really about making the right call, it's about finding better and better ways.
It's very important to frame the role of data in your organization well - it's what differentiates high performing organizations in many cases based on Stephen's history. Again, it's not a silver bullet but data makes it so you can make smaller bets to quickly get to a better solution via tight test, learn, and then iterate loops. Again, you want to make data a companion to execs instead of an either/or to their intuition and experience. Data can help them arrive on the right decision.
Stephen laid out a data maturity journey for an organization or a team: you need to start with collecting data - it seems obvious but you need mechanisms to actually collect data as noted earlier. Once you start collecting data, you don't need to become the most data driven team in the world overnight; start to process and analyze the data - find bits of information and insights to assist as you decide if there should be a formal data analyst/data scientist or not. From there, drive to faster experimentation and improve. The faster you can get solid information and iterate, the better. That will naturally lead to more data democratization as people see the value of the data and want to get involved. That can easily build from a team level up to larger organization level as well. But it's a journey, it's not a switch to flip.
Being cognizant of the cost of data acquisition has been relatively easy for Stephen since his past employers have been involved in hardware/IoT and there was a distinct cost to pipe the data back into the organization. But it's easy to lose sight of the cost of data collection for many organizations. He recommends looking at what data you want to collect by specifically what you want to figure out and why it might inform a decision. Just collecting data for the sake of data isn't good. Scott note: you really want to have these conversations as early as possible because the earlier you have the data, the more you can shape your decisions. If XYZ metric will be crucial for a decision in 6 months, you don't want to start collecting the data for XYZ in 6 months.
For Stephen, the challenge of data producing teams understanding downstream dependencies on their data is getting better but is still not fixed. But we shouldn't focus on just understanding who is using our data - it's the difference between producing data and producing data as a product. If you are producing data as a product, you should be actually interacting with your consumers so it's inherent that you should know who is consuming - and crucially, you know why they are consuming. But there is of course a cost to producing data as a product, lots of engineering time, especially if you don't have the tooling and capabilities, so that should be incentivized whether you are doing data mesh or not.
The best incentive, the 'best carrot' rather than stick, for getting teams to really work on sharing their data as a product has been the insight flywheel from Stephen's point of view. The team sharing their data, if the consumers then generate many useful insights that are helpful back to that producing team, that's how data at an organization-wide scale hopefully works. The 1+1+1+1 = 10 kind of approach but it's also hard to make sure that will happen and that teams make sure to give information back to the producing team as appropriate. You want to try to foster an org culture of 'if another team wins from my data, that is a win for our team too' but easier said than done :D
In wrapping up, Stephen talked about how can you get a team that isn't data driven to start heading down that pathway. There was the data maturity journey mentioned earlier but this is about developing a passion in someone on that team for data and then letting their enthusiasm for - and results from - data get everyone else on board. You may have to deploy someone into that team but make sure there is someone in the team driving them to be more data driven because of passion instead of simply try to use the stick.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Debra's LinkedIn: https://www.linkedin.com/in/privacyguru/
Debra's Shifting Privacy Left Podcast: https://shiftingprivacyleft.com/
Katharine's LinkedIn: https://www.linkedin.com/in/katharinejarmul/
Katharine's book: https://www.oreilly.com/library/view/practical-data-privacy/9781098129453/
Samia's LinkedIn: https://www.linkedin.com/in/samia-rahman-b7b65216/
Quick acronyms to know: PETs - privacy enhancing technologies; SMEs - subject matter experts
Scott Note Warning: there is some nerding out about how awesome it could be if some advanced privacy approaches and PETs were implemented at a broad scale across the industry to protect individual's privacy. It's pretty early days so warning about getting your hopes up :)
In this episode, guest host Debra Farber, privacy expert and host of the Shifting Privacy Left podcast facilitated a discussion with Katharine Jarmul, the author of the upcoming book Practical Data Privacy and Principal Data Scientist at Thoughtworks (guest of episode #157) and Samia Rahman, Director of Data and AI Strategy and Architecture at life sciences company Seagen (guest of episode #67).
Scott note: given this is a newer area, I wanted to share my takeaways rather than trying to reflect the nuance of the panelists' views. This will be the standard for panels going forward.
Scott's Top Takeaways:
Other Important Takeaways (many touch on similar points from different aspects):
Privacy and data sovereignty are going to be intermingled in interesting ways in data mesh. Querying data where it is instead of piping it all over the world* will help maintain privacy and comply with laws - many countries don't allow data to be exported as is.
see Zhamak's Corner 13 episode #173 that covers some of what querying data where it is means and that's not necessarily about source systems but it does mean not moving it without necessity
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Kiran's LinkedIn: https://www.linkedin.com/in/kiran-prakash/
Kiran's article on the 'Curse of the Data Lake Monster': https://www.thoughtworks.com/insights/blog/curse-data-lake-monster
In this episode, Scott interviewed Kiran Prakash, Principal Engineer at Thoughtworks.
Some key takeaways/thoughts from Kiran's point of view:
Kiran started off by talking about a blog post of his with a colleague from 2019 called "The Curse of the Data Lake Monster" - lots of clients were building big data lakes and it wasn't providing the expected value. There was an expectation that if you ingested and processed as much of your data as you could, it would create great use cases and lots of value. But it didn't happen. Value doesn't just happen without concerted and concentrated effort. So Kiran asked why aren't we applying product thinking to data to figure out and focus on what matters to drive value. What would happen if we focused on outcomes instead of platforms? What if we measured value not how many terabytes were processed and stored?
A key reason Kiran feels the 'Curse' happened was the strong separation between business and IT. That separation meant both were seeking different goals instead of collaborating. IT was focused on building things instead of solving business problems and the business side was focused on doing what they could, not building out a scalable and robust data practice. Conway's Law in action. We saw microservices really help to tackle those same issues on the operational plane so the data side of the house was definitely ripe for some product thinking-led innovation.
Data mesh avoids a lot of the issues of past data approaches by not leading with the technology - the first two principles are not tech focused. For Kiran, many (most?) of the organizations that are getting data mesh right are respecting Conway's Law and shifting their architecture and organizational approaches together. But in thin slices so as to not put too many eggs in one basket and make reasonable progress. And they are getting the exec sponsorship because while you don't want to reorganize your entire company upfront to do data mesh, you do need some top-down pushing to actually drive the necessary org changes when appropriate.
According to Kiran, while many people think data mesh has a high barrier to entry, that shouldn't be the case. There should definitely be a target operating model at the organizational level and organizations need to keep that in mind but it's not as though, again, you reorganize the organization all upfront. Organizations also need to really answer the question of what are they trying to do with data and what would doing data mesh well drive for them - if that's not crisp, they probably aren't ready to do data mesh because their business strategy isn't aligned to or contingent on doing data well.
Once the organization has the target operating model and vision down, Kiran recommends that domains should start to set their own specific goals aligned to the broader organizational vision. They should work on some hypotheses on how to achieve those goals and how they plan to measure their progress towards those goals. Start to build out your thin slice approach to making progress towards your vision and goals. Don't get super far ahead of yourself, look to progress at a meaningful but reasonable pace and tackle as little as is necessary now while still making sure you are aligning with the organizational vision. Keep your eyes on the prize, don't take on too much now. And yes, easier said than done.
Kiran pointed to a quote he read about if you modernize your legacy software stack but don't change your organization, you will need to do the same modernization in about five years. The same goes for data - if you are taking on data mesh from a tech-first approach, you'll just have to do all the same modernization again and you won't get a lot of the potential benefit from data mesh. Decentralizing the architecture will only really change things if you change the organizational aspects too. People, process, technology.
For Kiran, we need to really start to focus more on measuring value outcomes instead of inputs - how many terabytes or operations per second isn't directly tied to value. Teams need to have it made clear what is actually valued and valuable. In many large organizations, there are often less clear links between data work and business value so you have to educate and incentivize teams to do the high-value data work.
When thinking about trying to measure the return on investment in data work, especially data mesh, Kiran recommends starting by breaking it down into more tangible goals and measuring the value of achieving those goals. Then you can start to say how did the data work contribute to achieving those goals. But a data team can't really know the value themselves, whether inside the domain or not. And by breaking things into smaller goals and objectives, you can more quickly iterate towards value with tight feedback loops. Instead of large-scale projects, you build to larger and larger objectives through breaking things down and achieving meaningful micro progress that leads to large macro value.
Kiran talked about thinking of use cases as value hypotheses: you believe it will have value and thus you are making a bet. And it's okay to get things wrong, just limit the scope of the mistake so you can use the missteps as learning so the larger macro bet has a much higher chance of paying off. This is iterating to value, those tight feedback loops. If you don't have a culture where it's okay to be wrong, okay to fail, then data mesh is potentially (Scott note: almost definitely) not right for you.
Minimum viable product is often neither minimum nor viable in Kiran's experience. If you can't put something pretty rough in front of stakeholders, you waste far more time and effort building in wrong directions and are less likely to hit on success. But that's often out of the control of the product team. So it's a catch-22: do you put in a lot of effort to get it well past MVP or do you risk losing face? So we need a culture where we can actually do thin slicing well to really derive the most value out of building software, whether that is apps or data products. That incremental value delivery is really crucial to maintaining nimbleness as you scale. If you can't actually deliver in thin slices, it can significantly increase risk as you are making larger bets. But Kiran's seen that if you spend the time to explain the need for thin slicing and that what they are looking at is the 'sneak peak' and you just want feedback, most people get it and are reasonable. But you need to communicate about it :)
Kiran used a phrase Martin Fowler uses often from Ralph Johnson: "Architecture is about the important stuff. Whatever that is." When thinking about decisions that are hard to reverse, spend a lot more time and care, but the ones where it's easy to reverse, those probably don't make up the core of your architecture. When building your architecture, it's again important to build incrementally instead of trying to get it perfect from the start. Think about necessary capabilities, not technologies. In data mesh, that is about data product production and then monitoring/observing, data product consumption, mesh level interoperability/querying, etc.* Start to map out what you need and then think what level you need from a capability standpoint and when. You don't need to build out capabilities for when you have 20 data products when you have 1-2 data products.
Quick tidbit:
Leverage the 4 key metrics from DORA to measure how well you are doing your software engineering as applied to data - 1) Lead Time to Changes (LTTC); 2) Deployment Frequency (DF); 3) Mean Time To Recovery (MTTR); and 4) Change Failure Rate (CFR) https://dora.dev/
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Vanya's LinkedIn: https://www.linkedin.com/in/vanyaseth1809/
In this episode, Scott interviewed Vanya Seth, Head of Technology for Thoughtworks India and Global 'Data Mesh Guild' Lead for Thoughtworks. To be clear, Vanya was only representing her own views on the episode.
Some key takeaways/thoughts from Vanya's point of view:
Vanya started with a bit about her background and how deeply entrenched she's been in the microservices space - that played into the overall conversation a lot. Both Vanya and Scott agree if we want to do data mesh right, we really should take learnings from microservices and DevOps so we don't have to relearn what they already did the hard way.
For Vanya, data mesh is at a similar inflection point to where microservices was a decade ago - people were extremely skeptical that developers and operations could even work together, much less around combining them in a singular approach with DevOps. It's hard to imagine a post monolith world when all your career and experience are with monoliths. We have to be somewhat kind to those people in understanding that change is hard and scary :)
But, as a counter, for data mesh Vanya believes (and Scott agrees) we must try to prevent creating the same fear of missing out (FOMO) that microservices had. For many, if your organization wasn't doing microservices, it wasn't seen as a cool place to work and that all the best developers were at companies doing microservices. We don't want that in data mesh because it will lead to lots of wasted effort for companies that shouldn't be doing data mesh now or potentially ever.
According to Vanya, there are a few really good indicators an organization might be ready for data mesh. Before we get into the 3 she listed, a few things that might be indicative of indicators (Scott note: I know, I know, silly Scott phrasing) are constant displeasure of the kinds of initiatives they've been doing in the data and AI space - there is a constant pressure to prove the value of data and AI investments but really, an inability to do so. Long and lengthening cycles to return on data work/projects. A biggie is an ever-growing platform that is trying to do too much and hasn't been delivered - trying to boil the ocean.
So the 3 indicators data mesh could be a good fit that Vanya listed were:
1) Investments in data and AI aren't delivering expected value and it's hard to actually point to the value that is being delivered. Users aren't getting "the right data at the right time with the right quality".
2) Large and growing central data teams where trying to scale is done by throwing more people at the problem and it just isn't working. When automation would be better, they add people.
3) Confusion around who owns data when and why. Who owns the handoff between systems? Who owns the documentation and metadata around data? When someone has a question, how hard is it to find who owns the data?
Vanya highly recommends using value stream mapping to understand how you drive value with business processes and especially where are value leakages; this can be data or not, and should be applied to both analytical and operational data processes. You can understand better your business processes and expected outcomes - if something didn't meet expectations was that because expectations were wrong or did something happen along the way to lose value? Value stream mapping gives you an objective and neutral starting point and helps identify problem areas - value leakage - where you can prioritize what to tackle first.
In microservices, Vanya pointed to how challenging service discovery started to become until tooling came along - specifically mentioned Consul - so we really don't have to reinvent everything in data mesh. The tools out there, especially those in the open source space, are really making nice progress - specifically mentioned DataHub - compared to where they were 2 years ago at the infancy of bleeding edge data mesh adoption. Overall, we should 1) look to existing tools to see if we can use them as is; 2) look to extend existing tools where possible to cover incremental needs specific to data mesh; and then 3) look to create new tooling that is required for data specific challenges. Again, don't reinvent the wheel.
For Vanya, one thing many organizations struggle with in data mesh is the self-serve platform - what is the goal? Circling back to an earlier point, it's not about building the most amazing, ocean-boiling platform. It's about stitching tools together to automate the toil away - how can you create a holistic user experience to focus on doing the value-add? The value of the platform to the users is the abstractions away from the tools that make it easy to focus on what needs to be done to drive value from data, not play with the shiny tools. Focus on enabling interacting with the data, not the tools of the platform.
"Choose your blast radius" is a key phrase for Vanya. Think about scope appropriately and don't try to bite off more than you can chew. You don't have to reorganize your entire organization on day one to do data mesh, that is far too much of an upfront cost and makes failure a massive cost. Look at how it was done well in microservices: thin slices, not taking a sledgehammer to the monolith. Gradual evolution is sustainable, a revolution either succeeds or it doesn't - don't take on risk that isn't actually beneficial!
"Nothing succeeds like success itself," was another line from Vanya. It's crucial to get to an early win or two to show off to the rest of the organization proving data mesh delivers value and getting them interested in participating. 'Hey, we did this and it was a big win, who's next?!' It's not just about showing value, it's about showing there was a reasonable encapsulated timeline, not just promises. That incremental value delivery creates momentum and the more momentum you have, the more you can get people on board.
As many past guests have noted, it is a pretty bad (Scott note: fully terrible) idea to build the platform and then bring it to the users when it's done in Vanya's view. There are far too many unexpected friction points and finding those and tackling/automating away the actual friction is where the platform adds value, not bells and whistles. You want to find those as they emerge and work on tight feedback loops - that's product thinking! And if you don't make evolvability a first class concern, you are not building your platform as a product either.
For Vanya, it's pretty easy for tech people to focus on the tech, whether that is in data or not. But the overall organization doesn't care about the tech, they care about what they can do. So it's crucial to find the ubiquitous language and make your implementation and platform about what are people trying to do. The user isn't accessing S3, they are accessing the Inbound Marketing Conversion data product. S3 is simply a mechanism to accessing the data and insights.
When considering your thin slice early in your data mesh journey, it's okay to have a very unbalanced slice in Vanya's view. This has been mentioned before but it's important to reiterate. If you only need a bit of one of the pillars but you do need more capability in another of the pillars, that's absolutely okay. Don't build today for all the problems of 6mo from now. You want to focus on tackling the toil of today.
Quick tidbits:
Vanya's phrase "innovation in queue" is when an organization keeps putting off their innovation agenda for more immediate concerns - everything innovative ends up getting deprioritized in the queue.
Most data mesh journeys are taking six to seven months to really prove out doing data mesh and value. Scott note: this seems to be standard for larger organizations but a complete POC means faster follow-on for additional use cases. It's a balance!
In some organizations, if data is not really valued the CDOs or CAOs are looking to implement data mesh to show the value of data but their organizations often aren't ready and trying to do data mesh just creates more challenges than benefits.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
This episode is part of the greater AI/ML conversation I had with Zhamak but it's super important to emphasize the importance of trust - enough so that I created a separate quick episode on it. Not just trust in the data itself but that there is easy access and there will be going forward. A lot of the things we have done in data historically has been defensive in nature - especially grabbing a copy of the data now because who knows when you'll get access to it again.
What if we can implicitly trust that there has been care and foresight in preparation of the data I find, that there is an owner I can ask if I'm confused or curious, that my access won't suddenly go away or that what's there won't suddenly change without my knowledge? In ML/AI, the data scientists have done things in ways that made sense to their situation and challenges. What happens when we make trust inherent? What incremental value does that drive?
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
This episode is part of the greater AI/ML conversation I had with Zhamak. To start, Zhamak recognizes we aren't where we want to be in terms of capabilities - ways of working or tooling - to make this a reality just yet. But, if we can make it so data scientists can trust and easily consume from data products - that we create data products that don't care what use case type - regular analytics or AI/ML - can we remove a lot of the complexity they face? Do they need feature stores for data they aren't transforming? If they can get continued access and know the quality, why create a separate process that has fragility instead of trust the data product owners upstream?
I wasn’t smart enough in the moment to talk about do we need to have a copy of the training data itself for reproducibility but folks smarter on ML than I am can answer that one, probably in the affirmative. But overall, there is a lot of complexity in the way we do AI/ML because data scientists can't trust the sources of their data and they feel the need to take control because if they don't, their models break. So we need to earn their trust and show them a better way. But again, we aren't there yet, so let's work to make this a reality in the future.
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Brent's website and book: https://www.effectivedatastorytelling.com/
Brent's LinkedIn: https://www.linkedin.com/in/brentdykes/
Brent's Data Analytics Marathon Forbes article: https://www.forbes.com/sites/brentdykes/2022/01/12/data-analytics-marathon-why-your-organization-must-focus-on-the-finish/?sh=2af698743c3b
In this episode, Scott interviewed Brent Dykes, Chief of Data Storytelling at his own firm, AnalyticsHero. Scott asked Brent to be on after João Sousa pointed him to Brent's content.
Some key takeaways/thoughts from Brent's point of view:
Brent started with a bit about his background and why he titled his book "Effective Data Storytelling: How to Drive Change with Data, Narrative, and Visuals." There are a few places many organizations fall down in driving change via their analytics whether that is a failure to generate actual insights, a failure to communicate insights well enough to drive action, or a failure to actually take action on the insights - he's most focused on the communication of insights, a place often overlooked. You can find the best insights in the world but if you can't communicate those insights well enough, no one will understand them and/or understand the potential impact of acting on them. Communicate well enough to drive change!
The analytics marathon is one of Brent's big analogies for explaining where organizations fail along the path to taking action on their insights. There is data collection, which pretty much all organizations do. Then data preparation into data visualization. But this is where many orgs fall off because they are simply reporting on what's happening, the descriptive analytics and not actually driving to diagnostic analytics. Instead of doing deeper analysis, they believe their problems lie in what data is collected so they try to collect more data, thinking it's simply a lack of information instead of a lack of analysis. And then of course, once you do the analysis, you still have to communicate and take action.
For Brent, a few common indicators an organization will likely have a good analytics practice include: 1) an executive sponsor for being or becoming data driven. Possibly the entire leadership team. 2) a general commitment to driving actions from data where possible. It's "how we do things." 3) A test and learn culture in the organization that's supported by data.
If an organization isn't yet data driven, isn't doing analytics that well, Brent recommends getting to wins and slowly moving your executive sponsorship up the ladder. It might start at a Director level and then after you build momentum, people will take notice and you can climb to VP level and then C-Suite level. It's about showing the value of analytics and plugging along so you have proof points when you move the conversation higher in the organization. Rome wasn't built in a day and neither is a good, organization-wide analytics practice.
As data initiatives have become more ambitious, it's often meant ownership has become more murky according to Brent. What was once data that was essentially only for the generating team is now a potential core value asset and driver for the organization. And that opens you up for much more misunderstanding. Focusing on making sure information is understood - not just data is made available - is crucial to making good decisions with your data. There's often what the metric means and what others assume it means.
Brent shared his views that we need both active and passive ways of sharing context around data. Passive is the metadata, the documentation and the like. If we want to scale, passive is crucial. Self-service can't just be a pipe dream. But too often, people in data want to only do passive and ignore the people-to-people conversation. But often that's key to nuanced data or crucial to working with key people making big decisions based on data.
For 1000s of years, humans have been passing information via stories - human brains have evolved to share information via stories. We inherently want to know where the story goes. For Brent, mastering that storytelling with data and about data is the best way to convey the information we generate and discover with our data. If you don't communicate insights to those who can take actions, in a way they can understand, they won't take those actions :)
For Brent, execs rarely want to hear how the sausage was made via data. You want to show them what you've discovered and what they should do with that, not how you came upon that. It can be important to show people the sausage making isn't that hard though, especially trying to enable a team to do self-serve analytics. Really consider which is more appropriate to the situation.
It's pretty easy to like what the data is saying and point to the data as backing you up when you agree with it in Brent's experience. But a truly data-driven culture will focus on updating their thoughts and processes based on what the data says, improving their understanding via the data instead of trying to bend the data to support our hypotheses. It's about getting to a place where you try to remain more neutral until you hear what the data says and shape your vision around that.
On data-driven versus data-informed as semantics of what we're trying to get to, Brent likes the idea of data-driven. For him, it means really leaning into the data. Part of that is understanding and accepting that sometimes the data is wrong or we didn't ask the question in the right way - that's just getting to data maturity. And data-driven companies recognize the value of learning - there is still value from experiments and moves that didn't have as much benefit as expected when digging into the why. You have incremental understanding even if not direct incremental business value. And it sets you up to do better on the next iteration.
When asked about where will business analysts fit in data storytelling in the future, Brent sees them like personal trainers - they won't be doing the work but showing people how and assisting them until they can get to a level the BAs aren't as needed. That's for the analysis and especially the insight communication. Pure self-service analysis is nice in theory but you need a way for people to get help and make sure they aren't hurting themselves :) If more people are far more capable, that means the BAs can focus on the more valuable, large-scale questions.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Ananth's LinkedIn: https://www.linkedin.com/in/ananthdurai/
Schemata: https://schemata.app/
Data Engineering Weekly newsletter: https://www.dataengineeringweekly.com/
In this episode, Scott interviewed Ananth Packkildurai, Author of Data Engineering Weekly and the creator of Schemata.
Scott note: we discuss Schemata quite a bit in this episode but it's an open source offering that I think can fill in some of the major gaps in our tooling and even ways of working collaboratively around data.
Some key takeaways/thoughts from Ananth's point of view:
Ananth started by sharing a bit about his background. Despite writing the Data Engineering Weekly newsletter, he sees his experience as somewhat between a data engineer and a data analyst. That gave him the ability to see the full end-to-end journey of how data was handled at many different organizations. He consistently saw that analytical data outside of the application scope was an afterthought because developers were focused singularly on their application, not how it fit into the greater scheme, especially on the analytics side.
For Ananth, the data marketplace is a useful concept for many organizations when thinking about data contracts. It might be a bit more of a data bazaar than like Amazon in certain ways as there can be a bit of collaborative negotiation - 'oh, you have XYZ to offer, what about ABC, could you do that?' We need standardized ways to discuss/document data to make it far easier to share data, or at least start the conversation off from an informed standpoint when collaborating to get the most useful data created and shared. We need programmatic ways for producers to share what data they have available including their expectations like SLAs and consumers to request data they want with their expectations.
Scott note: It's crucial to understand that data contracts are less about the actual contractual terms and more about the establishment of a relationship that is covered through the contract terms. There are expectations but the contract isn't the entire relationship between the data producer and the data consumer. Essentially, the relationship includes the contract but just having SLAs will not resolve many of the issues people have around data contracts/sharing.
Similar to something Chris Riccomini mentioned in episode #51, Schemata is looking to provide feedback to producers about what broke downstream when they made a change. Or more valuably what will break before a commit is deployed. Data producers haven't had much of this feedback historically - e.g. "if you make this change, it will break your data contract expectations on the schema front because of…". But, Schemata is also designed for producers to see how well what they are offering fits with what other domains are offering when thinking about how well does my domain or potential new data product integrate into the overall organizational data sharing landscape.
On consumer-driven testing in data contracts/agreements, Ananth thinks there are two aspects: structural and behavioral. Structural is what you'd expect and what most people discuss in data contracts - mainly schema validation, is it backward compatible, is it strongly typed, is the required metadata complete, is there a registered owner, are the SLAs defined and complete, etc. The behavioral is similar to what Abe Gong talked about in episode 65 about what are the expectations, does the data behave the way people expect such that it can actually be leveraged for their use case. A key, widespread reason we need consumer-driven testing is producers rarely really understand how data consumers will use their data or are using their data already. Thus, that behavioral testing can inform the producer - along with actual human to human conversations - about how consumers will be/are leveraging data.
One general issue many teams have according to Ananth is the consumer doesn't really understand the cost or complexity of doing something around data creation. E.g. the producer of one domain might not store the user ID so to get every user ID is an expensive database call. So a consumer creating a pull request instead of a demand/request for data means you can start from a deeper conversation about what the data will be used for and why it's structured like it is proposed in the PR. This is also much more in a developer in the domain's workflow of using git. It's all just far less vague even if the initial proposal is infeasible - a producer has far more information about how the data might be used to start to iterate towards a workable solution.
According to Ananth, many people looking at Schemata have seen the need for years but there hasn't been a great way to implement what Scott calls "making the implicit explicit" around data sharing/data contracts. And this isn't a typical problem at a small company but once you get to a certain scale, the need for decentralized data modeling starts to become very evident. But with decentralized data modeling, it's pretty easy to put yourself in a bad spot because there is no collaboration layer so you create data silos / things that just don't interoperate well. Much like thinking federated governance versus decentralized governance in data mesh.
Schemata has a concept of a core domain that then every incremental entity or event you model, it will automatically assess how well the new event or entity is connected to the core domain. The theory is to quickly figure out how well what you are building will connect into the greater whole of the organization through the core domain. It gives you quick feedback on what is in process and you can easily add more fields to better match the core domain if a producer wants. It isn't a blocker, it's giving feedback to someone creating a pull request - data producer or consumer - about how well the resulting data model would fit in the organizational data landscape.
Ananth discussed how data creation is really a human-in-the-loop challenge - autonomous data creation is just not very valuable now and might never be. We need a collaborative platform to create data that is truly valuable and understandable but especially usable. The crucial aspect is to make a tool that integrates into people's workflows instead of yet another screen that further fractures the data management experience. Schemata is trying to be like Snyk - automatically scanning and giving people actionable advice but with little effort on their part. Where are your likely pain points? How could you address them? You can more easily set a goal of remediation/improvement and figure out how well you are doing. What are the top 2-3 things you could focus on to make the data you share that much better/more valuable?
A big thing many overlook in creating data contracts is about defining the value and/or cost of something happening according to Ananth. It's about getting people to the table to discuss something concrete and make sure people are on the same page. Instead of requirements, it's a collaborative discussion. Alla Hale in her episode #122 talked about every conversation, you should have something to show the other party, whether a full prototype or a post-it note with a little drawing. So getting to clear contract/agreement is far easier if you have a system that defines an owner, defines the parameters you need, makes sure the implicit aspects are explicit so both parties can fully agree, etc.
One thing Ananth - and Scott - keep running across are stealth data consumers creating one-sided data contracts. Essentially, the consumer has created their consumer-side testing and is consuming but the data producer has no idea they are consuming their data. Or many don't even really do the testing/contract model to protect themselves at all. The first the producer hears about their consumption is when something breaks for the consumer. With Schemata, at least there is a contract in place and stealth data consumers just have to inherit existing contractual bounds. Scott note: I hate stealth anything in data, let the producer know or they will potentially make breaking changes that could be prevented if they were just aware.
According to Ananth, we can really learn a LOT from the DevOps movement that has become more the platform engineering movement on the microservices side. If we try to push ownership to domains/data producers without the tooling to help them verify they comply with governance and that things are working okay, that's a lot of extra work on the data producer end. It's why we are seeing so damn much pushback from domains about not wanting to own their data - it's just way too much of an ask. Data producers just don't have enough information about what might be an issue when they try to make a change and it causes unnecessary friction. We need to make both the producer and consumer more productive, so that people can develop and deploy without tons of manual intervention.
Far too many teams are using tooling to solve single problems and while that one-off tool helps address a singular issue, it creates an even more disjointed data management workflow in Ananth's view. It's easy to focus too much on the spot challenge instead of the overall challenges in data management, the holistic process. Tooling fragmented with cloud and it made sense as we figured out new approaches and patterns - and VCs were quite free with their money - but we need to think about the whole process as one again now. Zhamak has mentioned this multiple times as well.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
All about why we need more explorers in our data mesh implementations and that it's too early for experts :)
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Andrew's LinkedIn: https://www.linkedin.com/in/andrewpease123/
In this episode, Scott interviewed Andrew Pease, Field CTO of North Europe at Salesforce. To be clear, he was only representing his own views on the episode.
Some key takeaways/thoughts from Andrew's point of view (mostly written by him):
Andrew started off by discussing the general way that organizations evolve. It's pretty natural for most to evolve into silos and the larger the organization, the deeper the divides between the silos and the harder it is to bridge those divides. With Conway's Law, IT systems/approaches also then often develop into silos. There is a lot of required intentionality to prevent evolving into silos or lessen divides that have already formed. And there is no "silver bullet architecture" to overcome the challenges silos create or undo the silos.
One of the big dreams of being data driven is putting timely and actionable data - "what do you want to tell them?" - in the workflows of business people. But, according to Andrew, many organizations attempting to do that look at it as an all-or-nothing kind of goal and that's just not reasonable. You won't get it "perfect" at the start. And that's okay, it doesn't make it not worth doing. As part of that process, it can be very important to reiterate that data is there to help not replace people - AI should mean augmented intelligence, it's there to help the human in the loop be better.
There are two major opposing forces re data quality in Andrew's view. First, you never get a second chance to make a first impression so your data quality has to be up to a certain level before showing to potential consumers. But conversely, the only way to get to actual quality data - essentially what matters, why it matters, and what quality levels are acceptable - is to get data in front of consumers and then iterate towards the required quality. Feedback loops are crucial to actual data quality so you can optimize for what matters. Your data consumers must understand that data quality is a team sport so they need to participate too.
Andrew brought up his concept of "statistics trauma" when discussing improving people's data fluency - essentially, many have a bitter taste from past statistics/math and/or data related work/school. So to get execs more data driven, you need to sensitize them to data but in a careful approach. That falls to the CDO and it can be challenging but is quite rewarding when it works. It's as much about communication as anything else in data.
In data, Andrew believes there needs to be far more bi-directional conversations. Data consumers need to tell data producers what they need and that can include data that doesn't exist yet so the producers need to start capturing it. So the earlier a data consumer can tell a data producer about their needs, the more likely they will get what they want down the line. Data mesh helps there because it's not the central team trying to understand and take requests to the producers. By cutting out the data team in the middle, you have a better chance to get to what data consumers want more quickly. But we can't lose sight of something that many seem to overlook - we can't just inform data producers of what we want them to produce and maintain, we need to properly incentivize and enable them to do so.
In Andrew's view, there is obviously value in collecting feedback on what data is viewed as valuable. But it's going to have bias - essentially it's valued but might not be valuable - so you should develop more concrete ways to measure what data work is useful and valuable. We should track what is being used but also how well what we thought would be valuable actually performed - that way we might better know what additional data might drive incremental value. Your feedback loops should include both quantitative and qualitative measurement where possible.
If you make ultimatums around data usage - you're either with us as a data user or you're against us - you won't get buy-in per Andrew. Mandates just don't get the buy-in some people believe. So you need to work to figure out why someone is not leveraging data. Again, make it less intimidating and make it rewarding. If you threaten people - do this or we'll fire you - you will simply get people adhering to the letter instead of the spirit - we want to make using data useful AND fun. Gamifying learning about data and data hackathons are two great ways to accomplish that.
Around data, if we want a "yin yang synergy" between business and IT, both parties have to meet the other _more_ than halfway in Andrew's experience. Both sides have to be willing to partner to improve. There isn't a silver bullet way to accomplish it but embedded IT in the business and vice versa can certainly help. You could rotate people across different business units. Etc. However, it's very important if you are in a decentralized organization to make sure you share best practices.
Andrew said, "the most complex system that we have in our organizations isn't a computer, it's the people who are operating the computers." There is a major change in the way our brains work between learning something and trying to get a point across. Some people are good at switching between those quickly - e.g. in a meeting - but many aren't and it's important to not leave them behind. So communication is crucial to get right and think about the broad group you are trying to work with. Sometimes data should be brought in to the discussion to make a point but sometimes it should purely be about increasing data fluency.
It's easy to try to focus on hiring for data skills in many roles but really, every organization should at least consider data training as part of a new employee training according to Andrew. Obviously don't forget existing employees but immersing people in data, especially the data of the organization, from the start pays off in the long run.
Quick tidbit:
Data consumers need to understand what is actually possible with data. E.g. lead scoring based on a person's name and email address is not a reasonable request.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
So, continuing the conversation about AI and ML's place in data mesh, we start the episode with Zhamak discussing an unnecessary complication we've created in data - why do data sets/assets only have to serve one user or even user persona? Yes, product thinking is about creating reuse but are we thinking reuse across regular analytics and ML/AI at the same time? We need to make it easy to give access in the language of, that native mode of access of, the data consumer. We shouldn't have to care what it is used for, regular analytics, ML, or anything in between.
There's also this very painful bifurcation between upstream data production and data science where the second data enters the data science realm of influence, it's copied over and you lose sight of it for discoverability, governance, security, quality, etc. They pull it in and then it's essentially impossible to track. That creates all kinds of problems. So why don't we extend data mesh into what they are doing? Do they need to make copies of the data in the feature store? If they have a trusted source of access to the data, do they care?
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Ebru's Twitter: @ebrucucen / https://twitter.com/ebrucucen
Ebru's LinkedIn: https://www.linkedin.com/in/ebrucucen/
In this episode, Scott interviewed Ebru Cucen, Lead Consultant at Open Credo. To be clear, Ebru was only representing her own views on the episode.
Some key takeaways/thoughts from Ebru's point of view:
Ebru started by sharing her background where she was a software engineer and trainer including training people on SQL before moving into data/data science. As a software engineer, it was crucial to at least model and understand data well enough to ingest and store it for the application side. The big challenges for software engineers really came in integrating that data into the monolithic data warehouse and then keeping it well integrated as the application evolved. The monoliths were bottlenecks on the software side and the integration into the data monolith was just becoming too much of a major bottleneck for all. As the DevOps and microservices movements picked up steam, the speed to reliably changing the application significantly increased. That increased speed created more and more challenges in integrating into a monolith - the applications drifted far too quickly to easily work with the data monolith. But monoliths aren't necessarily the wrong choice for all, just at scale and especially scale of complexity, they become a massive bottleneck.
When talking about versioning, Ebru talked about the many copies of data challenge - which one is the right one to use and can I trust it? There are people doing incredibly important work on data where they can't reliably trace it to source and know they are working on the right version. And with no clear ownership of data, nothing ever gets cleaned up so finding a reliable, repeatable source of data is very hard. So people copy the data they do find to their work area lest the source goes away, creating more copies. So we've figured out how to do versioning relatively well on the software/microservices side with APIs but we haven't figured it out for data - whether that is versioning the analytical API or the data itself. It's far too hard to make our data assets maintainable right now, thus the big push to data mesh.
For Ebru, when asked specifically what is the most important aspect of versioning in data - code, schema, API, or the data itself - she chose the data itself. This is a somewhat controversial choice but her reasoning was traceability - what actually happened to the data and when did it change? She expects that the code versioning, we'll have more version control systems and many people already manage their data related code work in git or other systems.
Another point Ebru made was that software development hasn't really had a focus on aligning itself to what version of data it is using. When you do a production deployment, the database is the database, it's tied to the application. But when we start to think about how we actually deploy software going forward, if it is referencing external data as part of that, the version of the data source it's leveraging obviously matters far more and we need far more coordination to ensure the software is referencing what we need it to. There is not enough tooling out there to easily manage this coordination and it's causing far too many issues.
Scott note: this is a really incremental thought here but VERY hard to explain. Historically, most services have been more or less wholly contained in what data they use or they access information from other services via a versioned API on the operational plane. So the coordination is less challenging. We have not really figured out well how to do that for data intensive applications - this is partly why everyone is building data products, whether data mesh or not, but it's still challenging if you don't think about providing a steady access mechanism and a way for a consumer to know what they are accessing hasn't suddenly changed without their knowledge. See the episode on my rant on data contracts and how it's not just schema and constraints.
We just can't escape Conway's Law according to Ebru. While many people have applied it to the operational plane, we really need to think about how Conway's Law applies to data. The way we exchange information can't only be the data itself, we need to get better at how we actually communicate and collaborate internally or gaps in how we communicate will be reflected in the data and our data integrations. Without fixing the way we work together and communicate, the producers and consumers will not collaborate well enough to leverage our data to the fullest extent.
Ebru believes that right now, it's still far too hard for producers to reliably publish clean, trustable, and understandable data. We haven't developed great ways of working and the tools are definitely not there yet. So if we try to push ownership on them too quickly, it will not go well. They have historically published what they want and we need to make it far easier to publish what consumers specifically want or they won't likely want to participate.
Data mesh is a sociotechnical approach but for Ebru, there is a lot of talk about the social and the technical is still lacking. There are so many tools but they don't work together that well natively and most only do a few very specific things - you could need 5+ tools to accomplish just the ingestion part of a use case. There is also a major challenge on the testing side - can you observe what changes would occur before writing the tests?
In general, we need to change our ways of working in data to enable much faster feedback cycles in Ebru's view. She was working on a project where everyone was in close collaboration and you could try things out and get feedback in the same day, meaning there was far less time spent building toward a solution only to find out the data wasn't available or there were other challenges. With better data ownership, we can go from idea to ingestion to testing in a short period of time, significantly improving how productive data science team members are.
Scott note: if you listen to early data mesh presentations from Zhamak, she talks more about data science/machine learning/AI than regular old analytics. This is that data bazaar/data marketplace kind of concept in action.
Ebru believes we need to take more learnings from microservices, especially the concept of Lego pieces. In data, we haven't really built incrementally to really achieve good value - it's often been all or nothing. But cloud means we have a chance to do things differently. That iteration means we can fail faster too - if we have an idea but we can't get the right data or even get enough of the right data, instead of building for weeks, we can change course. It's important to realize you can't ask any question to any data as well - sometimes you have a question that just can't be answered with the data you have or can get and that's okay.
To do data well/better, Ebru believes we need to create psychological safety and an ability to fail safely. That means we will have to train data consumers far better on how we work with data - a 95% confidence interval doesn't mean what most believe. And our understanding of data evolves too so consumers must learn to evolve their understanding. Human interaction is far more crucial than many want to believe in doing data well.
In data, as Zhamak has mentioned this trend towards super fractional roles, Ebru believes there is far too much focus in many organizations on what specifically is "my role" instead of what is the team's role and how can we make sure we accomplish our objectives. This fractional thinking of course creates more friction and challenges and handoffs - handoffs are always a place of lost context. So work to have teams focused on accomplishing team goals instead of individual ones.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
João's LinkedIn: https://www.linkedin.com/in/joaoantoniosousa/
João's Medium: https://joao-antonio-sousa.medium.com/
Brent Dykes' LinkedIn: https://www.linkedin.com/in/brentdykes/
In this episode, Scott interviewed João Sousa, Director of Growth at Kausa.ai. To be clear, he was only representing his own views on the episode.
The "four types" will often be throughout this summary. The four types refers to the types of analytics: descriptive - what is happening; diagnostic - why is it happening; predictive - what might happen in the future; and prescriptive - what actions should we take.
Some key takeaways/thoughts from João's point of view:
João started the conversation discussing the four types of analytics: descriptive - what is happening; diagnostic - why is it happening; predictive - what might happen in the future; and prescriptive - what actions should we take. Most of analytics work over the last 30 years has been the descriptive and both descriptive and diagnostic are typically owned by the analytics team. Data science, ML, and AI have moved the needle for doing predictive and prescriptive analytics the last few years. But diagnostic analytics remains underserved.
That diagnostic analytics gap exists for a number of reasons in João's view. On the people side, diagnostic analytics requires two sets of skills/knowledge: the analytical + technical and the business + domain. Without the domain knowledge, it is far harder to connect the dots around the why - a key concept in data mesh in shifting data ownership left*. Yes, we know sales in this region are falling, but why, what changed? João believes diagnostic analytics requires the most domain knowledge of any of the four types.
On the tools and processes side, João believes diagnostic analytics is far less developed than any of the other four types. Dashboards are great for descriptive analytics - what is happening - including some exploration but they are difficult to use to actually understand the why, drilling down to the root cause. Culture around diagnostic analytics is another large issue for many organizations - there are many varied approaches and lots of differing views on the actual value of doing deep diagnostic analytics.
João has three different diagnostic analytics immaturity stages before getting to a well-functioning approach. The least mature is "stuck in the what," where the business stakeholders are the ones trying to do diagnostic analytics with low data fluency* to drive to the why. They are only reporting the what, the descriptive analytics. The second maturity level is "the usual suspects" - essentially, the team builds lots of slice and dice dashboards and then just monitors things using that, they don't think to keep adjusting their angles and dig in. The third maturity level is "need for speed." The domains have the capabilities - usually via embedded analysts - to analyze their own data but are almost always siding with speed versus comprehensiveness of analysis. The world and business are changing fast but it takes time to do good analytics well to generate an actual insight.
Scott note: this brings up the question of where diagnostic analytics lives in a data mesh implementation. If domains have high data fluency, then presumably they can do their internal analysis but what happens if the information to drive to the why is cross domain? This is why I believe many domains are likely to have their own business analysts but organizations will still have a centralized business analyst team too.
On the question of who does the diagnostic analytics in most organizations - an empowered and highly data literate domain or a centralized analytics team - João said it depends. In a low data maturity organization, it's typically the analytics team - hopefully pairing closely with the business. In a higher data maturity team, it's about upskilling the subject matter experts in data and providing the right tools so they can do the analysis themselves.
João shared two signals you might be "stuck in the what", that you need more diagnostic analytics maturity. The first is in your weekly or monthly review meetings, you are talking about what is happening and there are only some high-level guesses as to why - "oh, that's _probably_ because we changed the website" - and not much more. Nothing is data-driven answers or even hypotheses. The second is reflecting that you haven't taken any real data-driven actions with a large impact recently. If you aren't driving your actions from your data, it's likely you aren't answering the "why" questions.
It's a lot harder to detect if you are in "the usual suspects" phase of maturity per João. The data teams aren't getting lots of additional requests. The business people are generally happy because they have dashboards that show a lot of information sliced and diced in how they typically look at things. But they are only testing existing hypotheses and not really coming up with fresh/new insights. So two signals are that there aren't really any new insights or hypotheses and there aren't many requests from the domain to the data team. The third signal, one that's indirect, is that because there is that lack of incremental data work and requests, the data teams start to become more disconnected from the business.
Teams that are stuck more in the "need for speed", João said while it's a better place to be, it's still frustrating. Teams are always trying to balance thoroughness of analysis versus speed. So some signals you are there is the pressure to cut corners on thoroughness of analysis in the name of speed - that actually happening is another signal - and constant high-priority interruptions for diagnostic analysis, juggling too much and putting aside the long-term work to take care of the fast turnaround requests.
When teams break past the immaturity stages for diagnostic analytics, João pointed to a few things high performing teams do well. The first is to segment questions/requests into tactical, strategic, and operational. Strategic questions are typically more big picture so they change less frequently and thus are typically less urgent than operational or tactical requests. Strong teams also adjust their thoroughness versus speed depending on what the situation calls for. Lastly, they automate as much as possible - there is still a human in the loop but repetitive tasks aren't value-add tasks for someone to do.
João shared Brent Dykes' definition of an insight, which is probably much more strict than many use. First, it must provide a shift in understanding - so not "we found this anomaly", it changes what people know. Second, it must be unexpected - so those teams stuck in "the usual suspects" won't meet this because they are only testing against the expected. And third, it must actually matter, it must be relevant and/or aligned to what stakeholders care about. João added his own criteria of it must be on time and communicated effectively. These are all necessary to actually drive the right action.
João wrapped with a few tips for improving your diagnostic analytics: First, show the value of drilling down in to the why - find a few easy initial use cases to really show the value, not your most difficult questions that will take months to really answer. Second, have the data and business people collaborate more closely so the data people can better understand requests and business people can start to think about new analytical approaches. Third, really get clear around your data role definitions: who does what and why and what _aren't_ they supposed to do. Fourth, start to get very clear on expectations, improve that communication so everyone is on the same page. Fifth, plan ahead and don't get stuck in firefighting mode grasping for straws - it's too easy to approach diagnostic analytics in an unstructured, reactive manner. Finally sixth, look to automate away the repetitive parts as much as possible.
Quick tidbit:
Beware the 'boring' label for diagnostic analytics. Many data people want to focus on the more technically challenging predictive or prescriptive analytics. Show people diagnostic analytics is valued.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sponsored by NextData, Zhamak's company that is helping ease data product creation.
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
Humans by our very nature categorize things - otherwise how can we really differentiate? How can we learn about new ideas and experiences if not finding a way to store them in our mental models. And in data, we've been treating diagnostic and descriptive analytics as an entirely different category to the predictive analytics of AI and ML. The way we partition the world in data is around how data will be used and then prepare the data as such, to be very fit for purpose. What if instead we partition around the data domain and don't really care about who or how things are used - we want to serve all consumers - what changes? Can we create data that is simply usable by many? Does that actually reduce complexity overall by not owning data production designed to specific purposes? Do we really need to treat AI/ML as if their consumption is all that different or special?
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Data Mesh Radio episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Karen Passmore (CEO at Predictive UX) led this discussion with Wannes Rosiers (Product Manager at Raito) and Alice Parker (Data Engineer at DNB). This panel was held in partnership with Data Mesh Learning - you can see a link to the video here: Panel: Data User Experience - An Introduction (Data Mesh Learning and Data Mesh Radio)
Alice's LinkedIn: https://www.linkedin.com/in/aliceparker/
Wannes' LinkedIn: https://www.linkedin.com/in/wannes-rosiers/
Blog post 'The Importance of UI/UX - and why Raito’s first hire was a designer': https://www.raito.io/post/the-importance-of-ui-ux-and-why-raitos-first-hire-was-a-designer
Raito blog: https://www.raito.io/blog
Karen's LinkedIn: https://www.linkedin.com/in/karenpassmore/
Predictive UX: https://www.predictiveux.com/
Some key takeaways from panelist Wannes Rosiers:
Scott note: I am pretty new to thinking and dealing with UX in data so I took an opportunity to write down some of my own takeaways that may or may not agree with any or all the panelists. Hopefully you'll find them useful.
The most key takeaways from Scott's view and learnings:
A lot of additional takeaways from Scott:
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Vikas' LinkedIn: https://www.linkedin.com/in/vksnov9/
Vikas' Twitter: @vikaskumar9 / https://twitter.com/vikaskumar9
Vikas' email: vikaskumar9 [at] gmail
In this episode, Scott interviewed Vikas Kumar, AVP and Head of Data, AI, and ML at CNA Insurance. To be clear, he was only representing his own views in this episode.
Some key takeaways/thoughts from Vikas' point of view:
According to Vikas, 2010 through the early 2020s the focus has been on moving the data to the cloud to better drive value. And now that more and more of our data is in the cloud, we are starting to see much broader adoption of things like ML and AI. The cloud gives us the promised but under-delivered scalability of the "big data" technologies along with the flexibility to move quickly and experiment. Cloud can also mean it's easier to bring non-data people into the mix to drive better collaboration between the data people and the business people/domain. So cloud gives us this massive scale and data availability but we still have to learn to better leverage our data, drive value from it - we are still in pretty early days there as an industry.
A big outcome of the mass movement of data to the cloud is how much time is spent on data management versus getting value from the data according to Vikas. DBAs used to spend 60%+ of their time just managing the data but data people's time is now focused on getting value and probably only 10-20% is spent managing the data specifically. But cloud can be a double-edged sword too - if it's very easy to create new data products or beta data products, you have to be very careful to not create overlap/duplicate work/data products. It all comes down to governance and your operating processes to prevent that.
As an industry, we are getting much better at serving data reliably at scale according to Vikas but we still struggle with the gap between the data is available and the data is able to be used by consumers in the business domains. We are still working on figuring out where to meet in the middle between handing people reports and maybe dashboards - a kind of old school approach - versus upskilling them to very high data fluency so they can build everything themselves.
When asked that question - do the data people have to learn all the business context or vice versa - Vikas gave the very data mesh answer of "it depends." But that makes sense because there shouldn't be a single prescribed method, you have to look at how your organization works and fit with that model. And you probably want to meet somewhere around the middle. Otherwise, you will cause unnecessary friction. So look to your general ways of working, cross train people, get people exchanging context about what they are trying to achieve and instill a culture of feedback and collaboration. That's how you can actually execute well on a data mesh strategy.
Vikas talked about your data strategy north star being about getting value from your data, reliably and at scale. So, you need to be realistic about where you are in that capability journey right now. As a data producer, you need to assess can your data consumers do everything necessary if you give them raw data or should you be curating it for them so they can actually leverage the insights. Work to find the high value return data work early instead of trying to do the most complicated aspects of data. It's okay to start small, no shame there.
A data product should always map to a target business outcome according to Vikas. But that shouldn't be the only factor. The reason for creating a data product should be trying to achieve that outcome so use that as the north start for the data product but we must build in a way where data products can be reused - sometimes with some additional work - for additional use cases. And it's really crucial to have a data product owner that is discovering and focusing on the objective of the data product. How can you provide the business meaningful data that meets their objectives, that should be a key objective of every data product.
When asked how do we balance focusing on the long-term wins instead of the quick - but typically small - wins, Vikas talked about the need to create a holistic view of your data and build a very strong foundation for how you will deal with data in general. That makes it so you can jump on the quick wins when you find them but you also have a steady foundation for making much bigger bets going after long-term big wins. But with a shaky foundational layer for your data, those long-term big wins are much less likely to pay off. And that foundational aspect comes in at the data product level too - build data products that can be easily extensible when it makes sense because they are built to be extensible from the start. Kent Graziano in the recent data modeling panel railed against having to rebuild every time you extend a data product, don't do that :)
For Vikas, there are many value streams for a data product - most people focus on the data set itself but it could be the governance work or the collaboration conversations between producer and consumer. We need to focus less on the data product as the exact output instead of the data product being the vehicle for delivering value but the overall product work itself significantly enhances the value of the data product.
Data governance seems to be the part of data mesh that confuses a fair number of organizations so they ignore at their significant peril according to Vikas. While you might not have to build every aspect of your governance upfront, it's crucial to think about how you will apply governance. And to truly get to the ideal of a self-serve platform, governance needs to be a simple part of the ways of working. Saving that for later is not going to end well for many organizations. And while access control is hard, we need to get far better at understanding who is using what and _why_. How long should someone get access to data? Forever access should be a non-starter. And how do we make it easy to grant that expiring access?
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
I share my strongly held belief that every data product consumer MUST register their use case with the producer and why it is so crucial and important. I sum up with these points:
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Bruno's LinkedIn: https://www.linkedin.com/in/brunoaziza/
Bruno's Medium: https://medium.com/@brunoaziza
Bruno's YouTube (Carcast videos): https://www.youtube.com/@brunoaziza
In this episode, Scott interviewed Bruno Aziza, Head of Data and Analytics at Google Cloud.
Some key takeaways/thoughts from Bruno's point of view:
Bruno started off with something you don't often hear a vendor say: "The number one barrier to your ability to drive value of data is not your technology, it's your people and how you organize your team." So while you can't buy your way to a data mesh, you also can't just flip a switch and be doing data mesh. You need to build your organization's capabilities to a degree they actually can derive value from their data.
It's also not easy for a data leader to necessarily create the necessary change per Bruno's conversations with data leaders. Many don't get a true seat at the executive table. And even if they do, if there aren't enough people "devoted to the data opportunity," it will be a very hard road to drive the data function to where it can add significant value. Bruno also dove into what he's seeing that makes for a high data literacy rate at customers - changing the day to day interaction, the habits of working with data. Making data part of many more people's roles and making it an intentional part of the company practices/habits builds an incredibly deep bench of data talent across the organizations. So his three components to positive change in your data approaches as a company are a strong data leader, the proportion of people committed to data work, and daily practices involving data.
While we know being data driven has an advantage - data driven companies are 162% more likely to surpass their revenue goals per a study - Bruno sees a few reasons why only 27% of companies are actually data driven now. To be data driven, you need to reliably produce data at scale, hence creating data products. And to do that, you need to build out the capabilities to handle data at scale - and not skip the governance :) But the end goal is to provide a reliable way to create value from data. That's really it. The best way to reliably do that is via data products in his view.
Bruno is seeing people go through three phases in getting to a reliable, scalable way to turn data into value. Phase 1 is the data ocean - it's not a lake, that's landlocked. The second stage is data mesh, allowing people to autonomously innovate with data but relying on central resources. And the third stage is a data factory. Scott note: the factory analogy might be rough because 1) feature factories are a very bad software pattern and 2) factories are notoriously about producing the same things at scale. And while we want scalable ways of creating data products, they should be more fit for purpose to use cases (but of course reusable as well) in my view.
The data product manager role is crucial to getting data products right according to Bruno. You need someone to be the CEO for you data product, that is focused on the actual value the data product drives and how reliable is the data product creation/maintenance. What more should be added to the data product? How is it used? To drive that cultural shift, you need a strong leader of the data organization that is empowered to make the right changes.
For Bruno, there are two factors that significantly increase the chance of an organization successfully becoming data driven. The first is an organization-wide mandate that data matters and that people must participate in the change and leverage data. Especially if the CEO is bought in on the data opportunity and the need for more and better data for themselves especially and the organization more broadly. The other is the attitude + aptitude to actually go out and build a scalable capability to build data products. And that's far easier said than done. That can be driven centrally or in a distributed way but you need people to step up and own the data.
The centralized data team model is becoming harder and harder for companies to scale according to Bruno. The team needs to be constantly ahead of the curve and they don't have the ability to learn all the necessary context so they quickly get overwhelmed by requests. This was a key factor in Zhamak creating data mesh as a concept. But the teams that are just fully decentralizing are creating data silos and making it increasingly hard to answer cross domain questions. So the organizations that are doing a federated approach with a strong sense of overall collaboration are winning - there are things that are centralized and things that are decentralized and each organization needs to figure out what works for them but balance is crucial. Find the right approach for the job.
Bruno talked about those companies that are focusing more on empowering domains than on the bigger picture of how domains can also work together. No major surprise but it creates data silos because everyone has different definitions so nothing is easy to integrate/interoperate. This is leading to the rise of the idea of the universal semantic layer.
Quick tidbits:
Data leaders should be involved in the hiring process for business people. That way, you can start to build a relationship early and help select someone who values data and has a decent data fluency. You don't want to be left out of the process.
It's absolutely okay to have domain-only data products that are very specialized to that domain - basically data on the inside in a data product. It's also - per Bruno - okay to have very centralized data products that are pretty core across the organization. But look for places to build reusable data products to get the most leverage from your data work.
To do data products and data product management right, you can't only focus on the data product launch. Maintenance and growth/evolution are crucial aspects of product thinking.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
We missed our window to record so I am interpreting what Zhamak is saying in her Medium post about why did she creator her company and the general state of the tooling market around data mesh. I also added on a the full 50min recording of our second recording which I had broken up into episodes 4-7.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
I will have much more to say on federated governance in data mesh but to reiterate my points:
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Neda's LinkedIn: https://www.linkedin.com/in/neda-abolhassani-ph-d-61354329/
OSDU Ontology: https://github.com/Accenture/OSDU-Ontology
In this episode, Scott interviewed Neda Abolhassani PhD, R&D Manager at Accenture Labs. To be clear, she was only representing her own views in this episode.
There's some very specific language about ontology in this episode but I think it's quite approachable for most people as a good understanding of ontology, the difference with taxonomies, and some specific insight into developing and applying an ontology.
Some key takeaways/thoughts from Neda's point of view:
Neda started off with a definition of ontology: "So literally, an ontology is a formal explicit specification of a shared conceptualization. I know that it has lots of jargon, but I'm going to explain it to you. So it is an abstract model of concepts, properties, relationships, and it is standardized, it is machine readable. And it is not just the instance level data, it doesn't include the instance level data, but it includes the schema and the type level information and how stuff should be connected in your domain."
When asked if it's best to start top down or bottom up when thinking about building an ontology, Neda said either is acceptable but the main advice is to start from the business questions you want to answer. After all, this isn't an exercise for fun, there needs to be a business purpose. And look for open ontologies for your problem statement or industry. There are a number of ontologies that have already been created that you can leverage, extend, and/or use for inspiration. There is no real reason to reinvent the wheel.
When building your ontology, Neda recommends keeping it as generic as possible. That way, you can apply it to multiple domains with no conflict. But it still has to meet your needs obviously. There are ontology editors to make things easier as well but it's important to set your ontology up to evolve as your understanding of your organization evolves and as your organization itself evolves. You can even do version control of your ontologies to make collaboration far easier as multiple parties look to improve the ontology simultaneously.
For Neda, ontologies and knowledge graphs go hand-in-hand. It's okay to have an ontology for the global organization and another one for a specific domain if that's of value. Ontologies are typically about communicating externally from the domain or whatever grouping you are representing. And knowledge graphs are for integrating data from different sources or domains. And for knowledge graphs, you need an ontology and a data model for it.
Ontologies are richer than taxonomies because while both capture the definitions, ontologies also have description logic. That description logic gives you a better ability to define things like unions, intersections, restrictions, and equivalences. So ontologies are broader than just concepts and terms.
Neda then discussed the OSDU or open subsurface data universe - an open source data platform for subsurface data in the oil and gas space. She specifically saw a gap in OSDU where companies loading their own data into the OSDU format was pretty challenging. It required a lot of subject matter expert time to match schemas to the OSDU format. So Neda and team developed a technique using a knowledge graph and AI techniques to try to automatically match and map data in a company's own schema to the OSDU format. And as stated earlier, a knowledge graph needs an ontology :)
As Neda worked to build out the OSDU ontology, she looked at the OSDU canonical data format and reviewed the schemas to understand what embedded choices were made so she could ensure she added that in to the ontology. She looked at the ontologies specific to related spaces or even some that were part of the OSDU area of interest like seismic data. However, the ontologies that existed for oil and gas were mostly outdated and didn't really cover what was really useful and interesting. An aspect of OSDU that made developing the ontology easier was that the schemas were not changing very often so there wasn't a constant remapping and versioning challenge.
So, circling back to data mesh, Neda believes it's important to leverage a knowledge graph to really ensure good interoperability between domains and data products. A data catalogue - or other mechanism for discovering data in data mesh - that only has information about the individual data products and not how they interconnect won't have as much value.* When should you actually start to develop and deploy your knowledge graph is again something that requires more study and feedback. All else equal, the earlier the better, but of course as things are developing and changing rapidly early in your data mesh journey, trying to _also_ update your ontology will be a ton of extra work. Time will tell.
*Scott note: yup. Isn't that just high quality data silos? Even if they interconnect, if people can't easily understand and find the interconnections, you likely lose a LOT of the value of data mesh. Whether knowledge graph and ontologies are the best approach remains to be seen however.
Neda covered an aspect that is really important for all things data mesh: how to measure when things are good enough versus they need updating :) for her, it's about what she said at the start: what are the business questions you are trying to answer? If you are still able to answer those well enough, you probably don't need to change your ontology. But if those questions have changed considerably and your current implementation is not able to answer those questions well, your ontology will need to be updated - maybe some new concepts will be added and some old concepts deleted. You do want to be careful to try to keep things backward compatible as you deploy a new version of your ontology.
Evolving ontologies is a challenging thing but if you designed your ontology well enough at the start, you probably don't need to do it all that often according to Neda. You should design your ontology in a generic enough way so it can handle new use cases without every little new aspect needing a whole new ontology version. However, that doesn't mean your ontology should never evolve. Things change or need clarification and you should be willing to be adaptable. Scott note: this is where Zhamak sees challenges with ontologies: if they are overly centralized and overly rigid, they prevent people from expressing real meaning at the data quantum level because they are trying to fit the definitions of the data quantum into the ontology.
In wrapping up, Neda shared her views on how to really get started on building out a good ontology and knowledge graph. It will require your data people to learn enough about domains from the subject matter experts to develop the ontology. Be prepared for that to be a bit confusing as sometimes learning a lot of domain knowledge can "discombobulate" your data people. And it won't be a super quick exercise. But Neda believes it will pay out in the end and add a lot of value.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
We wrap up another recording of Zhamak's corner talking about how do we actually start to look to build data products in a post pipeline data world. Data tools right now are kind of duct taped to each other and duct taped to the pipeline - how do we rethink starting from the end product - that mesh data product - and hook the tools to that to make interacting with it better. If you build a system that truly focuses on intentionality and responsibility that people can see, it creates trust. Away with the data black box!
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Chris' LinkedIn: https://www.linkedin.com/in/charles-dove-b4715723/
In this episode, Scott interviewed Chris (Charles) Dove, Data Architect at Endava. To be clear, he was only representing his own views.
Some key takeaways/thoughts from Chris' point of view:
Chris started with his view that while tooling is getting better in general, most tools are still very lacking in how their metadata plays into the greater organizational view of data which means we can't do some pretty basic things. Or at least there aren't comprehensive tools that make easy sharing of context easy across teams because of many metadata incompatibilities/challenges. So we need to get to a way to show the semantic context-related metadata as well as the transformation metadata in one place that is also understandable by the business users. A hard order to be sure but fundamental to enabling the vision of companies being actually data-driven.
There is a very common problem in organizations that comes from an implicit taxonomy and homonym problem according to Chris. The classic example is the definition of a customer but it goes far deeper than that. Some bit of data often has a very specific meaning in a source system and/or domain but then a different business unit looks at it with their own interpretation of what it means and misses the nuance, the differences. So do you have an enterprise taxonomy or do you try to document the exact meaning differences or do you not let people have access to data in case they misunderstand? Not as easy of a choice as many would like to prevent these misinterpretations or misaligned data mixing.
An interesting and very crucial nuance Chris mentioned about data sharing: the business people in a domain are often consuming their own data through a different lens. The data is embedded into an application for them so the interface makes much of the nuance, the meaning explicit. But that meaning isn't included in the data by default - the column title in the table doesn't have that meaning. So those business people - the data producers - often struggle to understand why people are confused or don't get the nuance. So it's important to make sure the data producing domain understands the interface others use to consume their data. Scott note: this is benefit I hadn't considered of having domains consume from their own data products. If nuances about the data aren't explicit, if the documentation isn't good enough, will they get confused about their own data? Will that force them to do better in building their data products?
Chris hit on a common problem many are having in data mesh - and data in general: what depth of documentation and explanation is necessary for data to be useful and not misused? People automatically assume some level of knowledge of the domain simply because others are in the organization. 'You work here so you understand my domain' type of attitude. So we need to make sure people can at least understand what they don't know and give them a way to get up to speed on what they need to know about a domain. Self-service can be a recipe for disaster if people can't understand when they are missing the necessary understanding/meaning.
While domain-specific acronyms can have a lot of embedded information in them for people with knowledge of a domain, they are often a major hindrance to those trying to learn about the domain in Chris' experience. Instead of focusing on exactly what your team calls everything, focus on the concepts and why they matter. As Shakespeare said "what's in a name", don't be enamored with sharing context via domain specific language. Referring back to Vlad Khononov's DDD episode, the internal domain language is the ubiquitous language but the published language is what is used to share with the rest of the organization. Focus on that published language - how can things be understood easily by those outside the domain?
For Chris, the point and meaning of data literacy isn't what most think - it's about getting people to understand what data they have and the general meaning/context so it can be communicated with the rest of the organization. It's understanding how data can be used and shared, not the exact technical aspects. It's about getting to a capability to share context and understand other's context around data without getting overly technical. When there is a use case that emerges, not every single person in the company needs to be able to create and maintain a data product. Basically, the concepts matter far more than everyone learning SQL.
In Chris' view, tribal knowledge is a very dangerous place to be. You have amazing and extremely valuable knowledge but it's trapped in people's heads. What happens if they leave? We all know about tribal knowledge but it's especially important in data because again the context and nuance, not just the column name, matters :) So extract that valuable tribal information, get it into a consumable format for the entire organization. It frees up the time of your most knowledgeable people too as they aren't answering questions all the time - extract once but leveraged by many :)
Good documentation, good knowledge sharing isn't about anticipating every challenge and writing out the fix or entirely preventing it according to Chris - that's not feasible. It's okay to get things into a knowledge base rather than the perfect metadata tool at first. You want to improve but if you are waiting for the perfect solution, you won't ever move forward with your data documentation. So get something out, put it in front of others, ask for feedback, and improve. And as stated, documentation doesn't have to answer all questions - it's something to make sure people generally understand a certain set of data and if they have a deeper question, they have a clear question escalation path for who to ask.
It's easy to get lost in data by focusing on 'data as the point' in Chris' experience. Data is merely a vehicle for exchanging information. But there are lots of interesting technical challenges in dealing with data so data people often lose the plot. Without the context around the data, it's useless too, so we have to focus on delivering it as one packaged unit. Scott note: this is what Zhamak keeps referring to as a data product container or a unit of data - that it isn't merely the 1s and 0s but the context, the user experience, the lineage, etc. wrapped in one package so it is usable as is.
For Chris, the biggest issue right now in data, especially for something like data mesh, is the trapped metadata problem. It's something Scott has mentioned repeatedly: most tools that touch your data in some way at best generate metadata that is trapped in that tool or it's extremely difficult to extract and integrate that metadata into other tools. And when people write custom code to do transformations, they often don't generate the metadata at all! So trying to get the full necessary picture of what's happening around our data is extremely time-consuming and difficult.
Chris called out the need for specification around metadata so we can at least bring it all into one place. Only a few vendors have moved to making it possible to even extract most of the metadata they create - he noted Atlan and data.world - but hopefully more vendors are pushed - or dragged - into doing the same. OpenMetadata or other early projects may provide a good way to start developing some standards for how things are described, shared, and/or stored. But again, trapped metadata is a lock-in pain that vendors are seemingly unwilling to let go of unless their hands are forced.
To really move forward with how we all approach data - as an industry and at the organizational level - we need to change the way people think and feel about data according to Chris. But change forced upon people only _might_ change their way of working and usually doesn't. So we have to focus on changing hearts and minds or the behavioral changes won't actually net the necessary care changes to the ways of working and understanding we need in data. Easier said - to change hearts and minds - than done but actually changing how we work requires empathy not mandates.
Chris finished on two points: the first is to really change the way an organization does data, they have to understand how data fits into their overall strategy and how treating it as a product impacts the work. Something may be valuable now but that value might fade. It's okay - and even very healthy - to end of life any data work that is no longer valuable. And the second point is that reuse is really key to generating strong business benefit from data. The cost of getting data to a point you can leverage it is typically high, look to make it reusable and find valuable ways of reuse as much as possible.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Juha Korpela (Chief Product Officer at Ellie Technologies) facilitated this panel on data modeling in data mesh with Veronika Durgin (Head of Data at Saks) and Kent Graziano (The Data Warrior, former Chief Technical Evangelist at Snowflake). This panel was hosted by the Data Mesh Learning Community in partnership with Data Mesh Radio.
Veronika's Links:
Veronika's LinkedIn: https://www.linkedin.com/in/vdurgin/
Data Vault North America User Group: https://www.meetup.com/dvnaug/
Kent's Links:
Kent's LinkedIn: https://www.linkedin.com/in/kentgraziano/
Kent's Website: https://kentgraziano.com/
Kent's Twitter: https://twitter.com/KentGraziano
Data Vault Alliance: https://datavaultalliance.com/
Juha's Links:
Juha's LinkedIn: https://www.linkedin.com/in/jkorpela/
Ellie Technologies' website: https://www.ellie.ai/
This write-up is from Scott Hirleman's point of view:
As someone without a ton of depth in the data modeling concepts, here are some of my key takeaways that should be taken with a grain of salt :) I decided not to write up everyone's opinions but more what are my takeaways:
Data modeling in data mesh will probably be far more similar than dissimilar to data modeling in a more centralized world. The focus on the business concepts is crucial. Far too often we try to start from the technical instead of what are we trying to achieve and it's crucial to not fall into that trap. Getting the technical aspects for interoperability wrong can be a pain but if things work together technically but not at the business level, that's a lot of sound and fury signifying nothing - essentially that's a lot of cost for work and compute that doesn't lead to actual business value.
One thing I'll note is every one of the guests is a Data Vault proponent. I'm not sold that it's the right way in the long run for data mesh - I feel like we need to evolve data modeling concepts for a more distributed, federated organizational approach. But from what they said, Data Vault does sound like a very solid base to start from - start from the business concepts first and focus on what you are actually trying to accomplish. Data modeling for the sake of data modeling is not something anyone should want to do.
As with just about everything else in data mesh, data modeling should be about limiting your blast radius of potential negative impacts as you get to fast initial and incremental feedback. Get to that iteration, take on things that matter but don't make it a big bang. Fail fast and all that :) This is not about taking requirements and going off to your own world at the domain level. It's even more crucial to have overall communication/cohesion as we enable more and more people/domains to own their data.
The most important aspect for developing lasting interoperability via data modeling is the business concepts. It's not that hard to do the technical interoperability once you figure out how things should work together. Starting from the technical feels easier but is a recipe for losing the value of the bigger picture. Typically, technical-based implementations are less extensible because the technology decisions are embedded into the solution instead of an enabling factor. Veronika said something like "focus on the words and meanings and not the data types." She also said that starting with technical integration focuses far too much on what can be done with the existing implementations instead of what do we need to drive value and how can we improve the existing implementations to create more value.
As with most things in data and software engineering, it's okay to build your data model in an opinionated way but maintaining flexibility, especially as you are doing your initial development, is crucial. Major change can be quite costly - especially if you have to change the entire foundation of what you are doing. Kent railed against the need for rework. Build your data model and subsequent data products so providing another view or angle isn't nearly as difficult and requires only the work to do that, rather than changing everything else you've done to also accomplish the new data view - in data mesh, this would often be a different API or a new table sharing the data for a different use/perspective. Again, look to prevent rework and ensure flexibility.
A massive concern with data mesh is data silos, the worry that if you have a bunch of domains doing their work separated and not in communication, nothing will interoperate. So you probably do need some kind of centralized group - whether that is their main role or part of their other responsibilities - helping domains do their work in the context of the greater organization. Note, that is what loose coupling from microservices means. Fully decentralized would be no connections versus things work together but can scale independently - that is decoupling. Having people who are there to help is the key to federated instead of fend-for-yourself so data architects are a crucial aspect of data modeling in data mesh.
While there is no centralized data model in data mesh - they aren't flexible and mean you lose a ton of context from trying to force things to comply with that model - there obviously can be centralized guidance and direction, a standard set of data models, etc. Think about a well-functioning federated government - maybe not the US… There are people doing work in the centralized function but it's about enabling those at the more local level to do things the right way. Juha quoted someone with something like "governance is not about leading people to do things right, it's about setting them up to do things the right way". That centralized team can't know what's right for specific situations - because they lack the localized context - but they can specialize in enabling doing things the right way. Kent claimed there is an enterprise data model and that can quickly go the wrong direction but if you interpret it well - that there are clear relationships across the business that are crucial to model well in your data, that are fundamental business truths you should reflect in your data - it can mean much less learning of deep domain specific context because you understand how domains fit the organization.
A number of people believe data modeling must be all about one view or perspective to rule them all. That is where data mesh fundamentally pushes back. You can have one view you agree on as an organization - such as revenue - but others should still be free to publish something that is similar in meaning but from another view. Much like in Domain Driven Design (listen to Vlad's episode for more, #171), there should be a 'language' (broad definition, think interface and terminology) of the domain to maximize the context of information shared in the domain and a separate 'language' used to communicate to the rest of the organization. That way we can still maximize context for business value locally but communicate globally in order to also maximize global data interoperability, which is a crucial organization-wide business value driver.
Kent mentioned another worry of data mesh that is often closely aligned with the data silo worry: master data management-related nightmares. While we absolutely have to reinvent MDM for data mesh - look for a few panels on that in the near future? - it's pretty clear it's bad to potentially have 10 different definitions that might filter to someone who doesn't understand the nuance and differences. Especially if that exec asks a simple-seeming business question and gets 5 different answers. Data trust is gone. So we have to be clear in tackling that problem and strongly communicating. Maybe not mastering data but mastering ways of answering typical questions?
Overall, I think you will learn a ton just like I (Scott) did :)
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
This is likely to be an episode to revisit. Zhamak explains a simple concept - data should not be copied unless it is owned by a data product - but the why is multi-layered and important. It might be one of the most important yet underestimated aspect of data mesh because when done right, it truly ensures trust in data - for consumer but also producer. There's a lot of nuance in how Zhamak is thinking about this but the actual application is quite easy :)
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter.
In this episode, Scott and Zhamak covered the misconception many have when she says "leave the data where it is" - it's about leaving the ownership with those who should have it, not leaving the data in the source system. If you only do source system, all you have is current state, so your data isn't even immutable or bi-temporal! They also discussed the need to be smarter about processing data - should it be at the source or should it be on query? There isn't a universal approach but we also shouldn't have to move the data around just to process it, bring the processing to the data. Lastly, we need systems to get smarter around efficient processing. We have far too much manual work on making data processing efficient.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Andrew's LinkedIn: https://www.linkedin.com/in/andrewsharp27/
Blog post "Data Mesh – Is this the evolutionary trigger to reinvigorate Data Governance?":
https://www.theoaklandgroup.co.uk/data-mesh-is-this-the-evolutionary-trigger-to-reinvigorate-data-governance/
In this episode, Scott interviewed Andrew Sharp, Data Governance Lead at the consulting company The Oakland Group based in Leeds in the United Kingdom.
Some key takeaways/thoughts from Andrew's point of view:
Andrew started the conversation off with a potentially controversial - but probably often agreed with - statement: of the four pillars of data mesh, federated computational governance is the most challenging and is the least mature in ways of working/patterns. Organizations are starting to learn and make their way forward but it's still a major challenge. Be prepared to explore and find the right path for you and your organization.
According to Andrew, most organizations are already not doing that well with data governance in the traditional sense so trying to figure out how to do it in a federated approach will be tough. And in data mesh, the computational aspect of federated computational governance means things are automatically applied where appropriate. That's very hard to do when you know exactly what needs to be done so it will be doubly hard in data mesh where we are still figuring it out. But changing your data governance approach can be an opportunity, not just a threat to existing status quo. How can we leverage the change we are doing to governance to be better than we ever were before? Far easier said than done but it's not only challenges.
To do data governance right in data mesh, Andrew believes it is more likely to require a major shift to generally how the industry approaches data governance; organizations will need to make big changes - over time - rather than just a few tweaks to better align with data mesh. But, it is very early days and that all remains to be seen, just a prediction. Scott note: I strongly agree with this belief. I think people are looking for ways to not invest effort in aspects of data mesh but I think many have noted the automated/scalable governance work pays significant dividends as your implementation goes wider.
But, Andrew wanted to stress that while we need major shifts, it's almost more like tectonic shifts than seismic shifts which often result the volcanic eruptions and earthquakes. Large but not moving quite as quickly - the big bang change approach to governance is overly risky. Why put all your eggs in one basket rather than try incremental improvements? Data mesh is all about trying, getting feedback, and iterating to improvement and governance shouldn't be any different. Build up the momentum around your changes and work with people to communicate where you are headed and why.
When discussing evolution of data governance and sort of traditional data governance roles and people that have been working in governance for a long time versus new people moving into the space, Andrew believes it is crucial for those doing the traditional type of data governance to grow and adapt their skills, especially technically. Will roles require additional responsibilities? Will domains have embedded data governance-focused people as their main role? Or will most of data governance at the domain level be split to responsibilities handled by roles not exclusively focused on governance? He doesn't expect widespread redundancies but do prepare for some changes. That said, it can be very much of a pendulum action instead of a shift that stays - so potentially look for an overly technical focus for a year or two before it settles into a better equilibrium.
"Turkeys voting for Christmas" is a phrase Andrew used relative to perception of the work many governance teams are doing in data mesh. Essentially, if turkey is a traditional Christmas dinner, are these governance teams that are helping lead the work to federate governance eliminating their own roles? He doesn't believe so and Scott STRONGLY does not believe so. Look at federated government - it isn't fully decentralized, that is just silos. Data silos are bad. So you need central coordination points and planners. Where the balance falls for governance responsibilities remains to be seen.
Historically, the central governance team has been doing all the heavy lifting because they are the ones trained to do so according to Andrew. But if we use a fishing analogy, we can see why central teams are happy to participate - if we have to fish and provide food for everyone in the organization, that's a LOT of fish you need to bring in. Instead, give them rods, teach them to fish. You can still focus on the big value fish - e.g. going and catching a swordfish or a tuna - but by breaking the work load down into manageable chunks, everyone can move faster and focus on creating more value where they have the best context. The less coordination we need across teams, the less unnecessary friction there is.
Other quick tidbits:
Most understand why data ownership is crucial. But many domains are not willing - or not capable - to take real ownership of data immediately. So gradual capability building and ownership handover is probably necessary.
The role of data governance professionals in data mesh is still in flux. Will there be embedded roles in domains or will it merely be skillsets as part of broader roles? Either way, there is likely to be a significant shortage of highly capable data governance people while the need for those people is greater in data mesh than traditional approaches.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Alexa's LinkedIn: https://www.linkedin.com/in/alexandrawestlake/
Alexa's Medium: https://medium.com/@westlakealexa
In this episode, Scott interviewed Alexa Westlake, Senior Data Analyst at Okta. To be clear though, she was only representing her own views.
Some key takeaways/thoughts from Alexa's point of view:
When asking people generally what do they think of data, Alexa has seen many fall into one of two camps: either data is transformational or data is expensive. And in truth, both are probably right. There is massive investment in data but most organizations are still struggling to scale.
"The data is bad" is such a common refrain across the industry but Alexa believes that is like data without context - essentially meaningless. What aspect is bad? What makes it bad versus good? Why is it bad? Do we need better collection processes in place? Do we need more expertise in transforming and analyzing data? And when answering those questions, it's often very difficult to figure out what are the overarching problems that you can tackle instead of building point solutions.
Alexa has seen it can be a bit of a scary proposition to try to get exec sponsorship to generate new data specifically for the purpose of analytics. So most companies start with what data they have today, and don't get overly ambitious, at least until they are proving they can deal with what they have now well. But, at the same time, many fall victim to the sunk cost fallacy of 'we've spent a bunch on this platform already, we have to focus on scaling it out', often throwing good money after bad. You can't just constantly change your platform but ignoring problems simply because the decision was already made is a recipe for disaster.
According to Alexa, a lot of the challenges in data come from the decision makers not really understanding their internal data ecosystems so we need to make it easier for execs to make better decisions. That ML model might have 40+ pipelines in some form or fashion that feed into it, of course it's likely to degrade. And there is often an urgency to solve the problem of today with a solution that addresses that challenge today for that specific use case - fighting the symptoms instead of tackling the cause of issues.
While something like data mesh - or any other large scale data transformation initiative - is a big change, Alexa believes we shouldn't make that a giant leap instead of small steps. You have to get to small wins as you're turning the ship to keep earning the right to steer the ship. And you shouldn't revolve the entire transformation around a negative, a pain point. If you do, you'll end up focusing on the pain too much instead of the goal - it's hard to sustain focus and drive around pain. Look to focus on the motivation behind why you are doing data transformation work but also what are the incremental results.
A few ways to burnout your data team Alexa mentioned: not providing them work they feel is meaningful; putting low priority on data work in general but especially on the interesting insight generation; data not being part of the critical path to business success. Essentially, if people aren't working on interesting, important, and valued work, they will leave.
Alexa believes it's crucial to focus on your communication when working with the business - they aren't well versed in data terminology or often even data concepts. So focus on communicating to them what matters and why instead of using data jargon. Most decision makers make so many decisions across many contexts, work to make it as easy as possible for them.
"Never jump and hope," Alexa said about measuring and maintaining momentum. To do data work right, you can't be an order-taker - taking the requirements, going away to build, and then presenting the 'thing' at the end and it's done. Get closer to the decision makers, understand what their expectations and needs are. Are they shifting? Are they expecting too much? Have the constant flow of information bi-directionally to make sure you are still doing valuable - and valued! - work.
There is often a lot of pain around data for business stakeholders but they typically can't directly identify the exact source of the pain in Alexa's experience. Collaborate with them to figure out the actual pain so you aren't treating symptoms. And work with them to explain that just like in medicine, exercise, diet, etc., there isn't a miracle cure or improvement. Data work takes time. Explain why it will take time and drive to what actually matters to address. The common example is 'what does real-time mean for you and what value does real-time drive?' It's often '2 hours is fine, just sick of 24 hour delay.'
For Alexa, it's very easy to try to build out governance centrally because it maintains a feeling - if not a reality - of control. The reaction of most humans to the unknown is fear. But you can make slow improvements to build trust and momentum - as Laura Madsen also talked about in her episode. If you don't invest the time to empower and enable your employees around data, you won't get good returns from it.
To do any data work, but especially something like data mesh, right you need champions in the domains to help. Both to help you move things forward but also that constant communication loop as the world - and thus requirements - changes. Focus on them being your partner and treating them like a customer of your product.
It's very easy to fall into the trap of putting too much emphasis on the data work, according to Alexa. It is not likely to make or break your organization but it can be a significant competitive advantage. Data can unlock your value potential and sharpen your competitive advantage.
Quick tidbits:
"Without literacy, all your analytics is is expensive."
Alignment is one of the hardest things to do organizationally and is far more crucial in data work than most believe. Be prepared to repeat yourself over and over.
It's easy - and often fun - for data people to focus on the raw data you have instead of the insights you can generate. When talking to business stakeholders, how often do they care about the data itself versus what it means? Focus on communicating in what matters, not the speeds and feeds aspects of data.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter. Oh, and she's hiring.
What tech is already available that could be used for data mesh? There are so many amazing approaches and technologies in data but they've been used for the pipeline approach only. We need to think more like developers - not accepting the grunt work or death by a thousand cuts of data - and take a hard look at what we've done historically in data and what should be replaced.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Salim's LinkedIn: https://www.linkedin.com/in/salim-syed-11981521/
Capital One Software: https://www.capitalone.com/software/
Capital One Slingshot: https://www.capitalone.com/software/solutions/
Capital One blog post on their use of Snowflake for data mesh: https://www.capitalone.com/software/blog/operationalizing-data-mesh/
In this episode, Scott interviewed Salim Syed, VP of Engineering at Capital One Software.
Some key takeaways/thoughts from Salim's point of view:
Capital One is on its own data mesh journey and what they learned and built internally inspired them to build an offering called Slingshot. When asked about where he would tell others to start their own data mesh journey, Salim mentioned that we can't solve our data issues - whether data mesh or anything else - through just technology. At Capital One, they started on the organizational aspects, breaking into discrete lines of business (LoB) and then creating units of data responsibility with a hierarchy. However, the data hierarchy wasn't the same in each, they left it up to LoB to determine - a large domain having 3-4 and a small domain having 1.
According to Salim, a big reason why they have so many people focused on risk is similar to what Sarita Bakst mentioned in episode 52: you want to give people as much access to data as you can while minimizing risk. So set yourself up to give people access when it's valuable / necessary for their job. There are even instances at Capital One where you need to complete training to get access to data. And they have active risk monitoring constantly in place as well to make sure they don't miss anything. Again, doing that gives you more piece of mind to leverage more data.
When you start federating your data ownership, you can't only give people the tooling and the authority, that won't work well in Salim's view. So, in addition to all of that, you need to focus on a usability layer. Think about email - it's pretty easy to work with an email client but what if you had to put all the plumbing and add your headers and all that yourself. Just giving people tools, patterns, policies, etc. and expecting them to be able to handle it all isn't realistic, make it easy to leverage, think about their user experience.
As an example, Salim talked about the process for creating a data product. Instead of interfacing directly with 6+ tools, there is a workflow with automated processes to make it a smooth process. As many past guests have noted, reducing friction to sharing data is a key element of driving value from data mesh. It's not just called a data platform, it's called a self-serve data platform for a reason :)
Salim shared some advice other recent guests have touched on: really focus on the job to be done, who is doing it, and ensuring your infrastructure and governance stay in sync. Focus on the job to be done, whatever that job is, instead of the tools. Again, reduce friction. Focusing on the persona of who is doing the work is also crucial - Audun Fauchald Strand and Gøran Berntsen from NAV in episode 37 talked about building a great data platform no software engineer would want to use. Use product thinking, who is using it and how do they typically do their work? And lastly, ensuring your governance and infrastructure stay in sync - if there is an update to the data product, does that automatically update the data catalog? How do you prevent drift between systems or kludgy manual fixes?
Data discovery done right is about a few things according to Salim. Ensure people can find relevant data easily is baseline but also get them to as much understanding as possible. Even sensitive data, what are the quality metrics, the general data shape, etc.? And make it easy to immediately request access when you find data you want to use which immediately triggers a request to the data owner with a business justification. And the relevant policies for that data product are automatically part of the approval process so the data owner doesn't have to remember the policies themselves. Again, reduce friction to getting the job done for all parties.
Salim talked about a few things they overlooked at the start, one for personas and one that was causing a lot of friction. As mentioned earlier, at Capital One there are a number of risk managers but the early iterations of the platform didn't cater to their experience. Which meant access requests were delayed and risk monitoring was tougher to do. So make sure to consider all your personas that will be using your data mesh - ignore at your business value peril. The other aspect was how much manual effort was involved in patching production data so they addressed that as well.
To be able to federate the actual infrastructure management to the domains, Salim and team knew they couldn't just hand over the tools. Again, the personas in the LoBs wouldn't have the expertise to manage the infrastructure. So they focused on exactly what Salim mentioned throughout: the experience. How could they empower the lines of business to own their data infrastructure without the domains having to manage their data infrastructure? So the team built out a platform with capabilities with experience at the core but with additional aspects like DBA best practices, guardrails, and cost management as part of the platform.
Scott Note: from here forward, there will be some discussion of Capital One's Slingshot offering. This is not an endorsement of the product at all but it is interesting and germane to the conversation (read: Salim was not selling, only telling, which is a-okay on the podcast). Cloud cost is near and dear to my heart and it's important - whether you build or buy - to not overlook cost management.
So, all these challenges of how they addressed their cloud cost management via their platform led them to believe there was a market for this type of solution per Salim. Capital One has a history of creating cloud cost tooling as they were the creators of Cloud Custodian (https://cloudcustodian.io/). Scott Note: unpredictable and/or high costs have been a major concern/pushback to data mesh since early 2021.
One general issue with on-demand / cloud computing is cost inefficiencies. There is always a lot of waste and it's often actually more cost effective to ignore than chase it down unless you know where inefficiencies lie. So Salim and team found it useful to not try to automatically clean up cost inefficiencies - that pretty much never works - but to highlight them relatively quickly and offer potential recommendations and/or help. And they lowered their own Snowflake costs quite a bit in the process.
The bigger benefit according to Salim was the team put proactive questions in place for when teams were provisioning their data infrastructure. Often, it can only take a few minutes to save a large percentage of money if only people know the knobs to turn. But you don't want everyone to have to be an expert - extract the information from them based on their needs and create a recommendation system. This is just yet more on the experience side - don't build cloud cost experts in every LoB, make it so they can make the right decisions as often as possible quickly and easily. You should also look to build in cost forecasting tools as part of your experience so people aren't hit with a surprise bill - the surprise Cloud bill is so common, it's a meme on Twitter.
From what Salim is seeing, most companies - or more correctly, central data/platform teams - are pretty reluctant to federate ownership of the infrastructure to domains. He believes that is because of things the LoBs don't understand like cost controls and best practices but that if you allow the central team to set guardrails and best practices, they will be more willing to give up control. Remains to be seen.
Salim finished with "in our experience, data mesh works with central policy, central tooling, but federated ownership."
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Amy's LinkedIn: https://www.linkedin.com/in/amy-raygada/
In this episode, Scott interviewed Amy Raygada, Senior Director and Analytics Product Manager at Swiss Marketplace Group (SMG).
Some key takeaways/thoughts from Amy's point of view:
Amy started by talking about something many other guests have probably felt but few have said: when signing up to do data mesh, you won't really know for sure what it will be. And that's okay, your journey will take you places you didn't expect and have hurdles and obstacles you can't see or predict. But that's all okay, you can learn and iterate along the way. Expect the unexpected.
When Amy started interviewing at SMG, the team was not as familiar with data mesh. But for the past six months it's been a key part of her focus. She paired up with the head of data engineering and worked to brainstorm before moving forward - that pre-work lasted about 2-3 months. There was also a lot of other change happening in the general technology and data landscape at SMG - moving from on-prem to cloud, moving from monolith to microservices, moving off legacy technology, etc. - but that meant they were able to get everyone together for a 2 day workshop and really look at things from a fresh perspective. At the same point, people are not used to abrupt changes and with data mesh, there will be a LOT of changes - look to implement that over time, don't be in a rush.
SMG decided to start with a single domain - leads - instead of multiple domains. The initial use cases were useful for the leads domain as well, especially by significantly improving data quality. As part of enabling that domain, Amy and team are working closely with them to teach how to handle data, what data ownership actually means.
The leads domain was chosen because they had a significant need for help with data quality and had moved more to cloud and microservices than other domains. It was also a smaller, more manageable problem than say sales, which is in a major transition to Salesforce. They had three non-mesh data products that were impacted from poor leads data quality so there was a lot of downstream issues that they could address, a lot of incremental business value to drive by fixing the data quality issue. The other domains they considered were in big transitions so it would be harder to get where was necessary to drive value in a PoC with a lot more work and risk.
Data contracts have been a key goal and key driver for Amy and team. The goal is to better define what you are trying to do with data, what do you as a data owner need to actually do and deliver and what can the data consumer expect. The driver aspect is that data owners have something more concrete so they are willing because they only have to deliver what they say they will. It gives a limited scope to their data work.
A focus for Amy and team is to be the enablers only, building out the platform and teaching people how to own but not do/own as central data ownership was causing issues in the first place - it's what data mesh specifically moves us away from. The domain teams will certainly need "babysitting" but that is expected and can mean more information flowing to the platform team to make improvements too. Data ownership isn't a one or a zero, it's a process.
Amy believes it's important to really pair with domains to share the logic with them behind data mesh - why are we doing this change - and not make it feel like you are changing their ways of working instead of adding new, value-add capabilities to the domain. This close relationship has allowed them to do a data contract demo where you can show the domain what happens when they make a change that will violate their contract. That way, they understand what happens to downstream consumers - possibly themselves - when they make a breaking change. And that the platform alerts them to a breaking change too so they have a better chance of preventing issues themselves.
Similar to what Chris Riccomini mentioned in episode 51, Amy and team are implementing automated schema validation checking at the pull request level. This prevents breaking changes from going through with consumers being the first to know about an issue. It also kicks off a conversation about should this change be made and if so, how will they do versioning. And Amy knows this can overwhelm some people but the team she is working with understands the pain so they are eager to prevent that pain. They are also looking at data reviews - similar to architecture reviews - to assess if operational system changes will impact the data. Abhi Sivasailam in episode 9 mentioned they are doing a similar process.
Amy believes - and Scott agrees - patience is crucial. Getting the first domain into a really good spot and then enabling them to share their story and their learnings will be crucial when they try to go to additional domains. Not being in a massive hurry means teams have the time and space to learn how to own data instead of piling a huge additional workload on top of an overburdened team.
Educating the general company about what they are doing with data mesh has gone well. Amy created a Miro board using Barr Moses' old joke of data mess to data mesh. So they are working to explain the what and the why to everyone involved but in a simplified way. Talk about what changes - the new responsibilities and what those mean and drive. Talk about how you can bring everyone to the table and especially what benefits data mesh has for them. Really focus on the practical of what are they being asked to do and why.
While many stakeholders in their initial domain were anxious to engage at first, now that there is proven value, according to Amy those hesitant stakeholders are much more willing to pair up. They are providing test cases to the data team so they can quickly validate value and iterate together. They are already seeing the benefit of the work with other stakeholders and it's getting those previously hesitant stakeholders excited to team up. So if you want buy-in, look to provide some value first and then show/prove that value. Yes, easier said than done.
Quick tidbits:
Make sure to team up with any domains that have had success so they can help you sell other domains on working with you. They can help educate and also show that you are actually providing the business value you claim.
It's super easy to get bogged down by metrics. Push back on big metrics requests - why do you actually need this? As Alla Hale mentioned in episode 122: "What would having this unlock for you?" If it doesn't unlock value, why do it?
Having prioritization meetings weekly keeps everyone on the same page, heading in the same direction.
It's easy to get wrapped up in what might happen. Focus more on what's in front of you and what you are trying to do. Don't cross bridges before you come to them :) there are countless bridges in data mesh, focus more on the now.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter. Oh, and she's hiring.
What do we need to develop to provide a better experience for data product developers and consumers? How can we make it easy for consumers to move from discover to trust to lean to use? We still have a long long way to go with analytical APIs, they are basically only for sharing raw data at the moment - how can we better share the embedded information? And how can we give data product developers the ability to do the right thing by default where we optimize for ease of use and also for more generalist developers to use - instead of hyper-specialized tools experts? There have been two very separate universes of operational and analytical planes - how do we change that and leverage software engineering practices in analytics?
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Alice's LinkedIn: https://www.linkedin.com/in/aliceparker/
No Silver Bullet: Essence and Accident in Software Engineering by Fred Brooks: https://www.cgl.ucsf.edu/Outreach/pc204/NoSilverBullet.html
IBM Research paper mentioned: https://dl.acm.org/doi/10.1145/3290605.3300356
Microsoft Research paper mentioned: https://dl.acm.org/doi/10.1145/2884781.2884783
In this episode, Scott interviewed Alice Parker, Data Engineer at DNB.
Some key takeaways/thoughts from Alice's point of view:
Alice started by talking about her recent Master's thesis which was studying human computer interaction, specifically around data mesh at DNB. And it's a huge topic. To start though, she emphasized the difference between user experience (UX) and user interface (UI). Your experience is impacted by much more than just the visual buttons you click - or knobs and levers you move if on a more industrial setup - so user experience is much deeper than many think. Zhamak has mentioned user experience - in multiple contexts - as a crucial factor in getting data mesh right. It's a rarely explored topic in data, especially designing your architecture to provide a better experience for data producers and consumers.
According to Alice, there is even an ISO standard focused on experience. It's about "how to help users achieve their goals effectively, efficiently, and with satisfaction, how you can minimize risks, and how you can enhance the maintenance of tasks to be completed effectively." And breaking that down into each piece can be a pretty deep conversation. How do we ensure satisfaction of data users? What is risk when it comes to data? Risk of interpretation? Misuse? A crucial aspect of UX is understanding the capabilities and limitations of the people using it.
As Alice noted, "systems and technologies evolve incredibly quickly, but unfortunately, humans don't." And we can't think of every person as being the same with the same needs and capabilities. Which unfortunately means, one size will not fit all when thinking about building your platforms for data mesh. We have to design to serve persona needs. And to actually understand persona needs, we need to speak with them.
When listing out the different personas, just on the data consumer side Alice mentioned data analysts, data engineers, data scientists, data stewards, and business owners. On the producer side you have data engineers, data scientists, software engineers, and business owners. Then in kind of the in-between, other personas you have platform engineers, data governance people, etc. So you have to think about what each persona needs and then design for each persona. The personas even have different terminology, different ontologies they use. If only this were that easy :)
What this all leads to is going and actually talking to your potential users to find their needs - which is what Alice did with interviewing data consumers for her Master's thesis. Similar to what Jen Tedrow mentioned in episode 98, you need to go and listen to their pain points and then abstract away the use cases to see what are the bigger needs. And you can do that in a very informal way too. But you can really only get that feedback through conversation and explicit effort. Scott Note: this is EXTREMELY true… the number of times I've asked for feedback and gotten crickets is… yeah…
Alice noted it's important to let people know when their requirements won't be met, or won't be met on their timeline. That open communication will get them to trust you. It's okay to say no to a change to user experience, especially if you have good reason for it that you explicitly communicate to them. You can work with them to maybe find the quick wins that gets them additional value now. Prioritize what can be done now and explain why for your prioritizations.
Documentation is one crucial aspect that has been lacking in data according to Alice. Yes, data mesh calls for data quanta to be well documented but we need to ask who are we actually documenting for? If you have 5 different consumer personas for your data product, is it documented so all of them can actually use it? And then how do we make it as easy as possible for data producers to actually create and update that documentation? Does documentation only have to be written? What are some low friction ways to share the context of what a data quantum is all about?
It's crucial to think about incremental progress - and showing that incremental progress - on your user experience in Alice's experience. Every system everywhere will have frustrated users. Such is life. And you can't solve most challenges in a day. But look for ways to iterate towards a better UX and circle back and show people you are improving it - if that's data product creation cycle time, show them your prioritizations that sped up that cycle time and show them the improvements, show them you listened and reacted. They might still be grumpy but at least they have reasons to be less so.
As other guests have noted, Alice has seen most people are very willing to share about their current challenges. If you go with the right attitude - to listen and empathize - you can learn a ton. People want to feel seen and heard. And it might help more with your prioritization than you'd expect. As Alla Hale said in episode 122: "what would having this unlock for you?"
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Dacil's LinkedIn: https://www.linkedin.com/in/daciluhernandez/
In this episode, Scott interviewed Dacil Hernandez, Director of Data and AI for Northwest Europe at Nagarro.
Some key takeaways/thoughts from Dacil's point of view:
Dacil started out by calling herself a purple person, meaning not red or blue, not only technical or only business focused but a combination. As part of that, she echoed many past guests - the tech is the easy part, creating value is the hard bit. It's easy to lose focus on creating value and playing with the cool toys, building really amazing things. But do they create business value?
It is crucial to have business counterparts as key partners, key stakeholders, in your data mesh journey according to Dacil. It is so easy for IT and the business side to get disconnected, for use cases and needs to change but the data getting shared doesn't change. And with everyone saying business first, we are still focusing on tech far too often. Dacil likened data practices to being a teenager - we keep hearing business first but don't listen, we do what we want :) And we need to really get crisp on what "business first" actually means.
Dacil believes it's crucial to get out of the IT desk request -> data loop we've been stuck in for so long. The business needs to come to the table and we need to bring them to the table. If the business side isn't part of the conversations, you can't get them to understand what data ownership means. They may say they will take on the responsibility but they can't really grasp how to do it well. So we need to partner with them to get what we need and we need to work with them to improve their understanding and their capabilities around data ownership.
"I need your help to help you" is something core to Dacil's work with the business. Too often, IT is left to wait for requests instead of working together to get what people need. It can unlock a number of additional use cases too.
When creating or improving a data strategy, Dacil recommends including incentives that mean people see/feel the need to participate. Leverage a fear of missing out or FOMO. Make sure people understand what will happen and the rules of engagement. Find the incentives to get them to participate by adding value back to them. Of course, easier said than done.
Dacil mentioned an interesting idea a company is implementing: using gamification to find data quality issues. So instead of it being this really bad thing to discover data quality issues, the people who found the issue get rewards. So that also will drive better data consumer literacy and drive up trust - they are checking data for "does this make sense" instead of just consuming data and learning in the process. Where else could friendly competition/gamification work in data? Look to create friendly and positive energy around your data work and show how it contributes to company value too.
For far too long we've had a vicious and costly cycle of bad requirements and requests leading to bad results and data according to Dacil. So, we need to ask far more questions - why do you really need this. As Alla Hale mentioned in her episode, not in a push back way. She said "what would having this unlock for you?" So take what they are looking for and why and repeat it back to make sure you are on as close to the same page as possible. And then keep communicating while you are building. Drive to that small prototype to make sure you are driving towards value together and aligning on expected outcome.
There's a maturity level to differentiating between what people want and what they say they want in Dacil's experience. And it is also often hard to differentiate between what they say they want and what they are trying to achieve. Always dig into what are they trying to achieve or you will create lots of wasted work. Again, have the conversation.
It's very easy to measure the wrong thing in data quality according to Dacil. She brought up an example where phone number was 100% complete for every record but most of them were not real numbers. So someone said they had perfect quality based on "there was something in the field" but it was unusable, wrong information. So when you look to data quality measurements, tie to value. What is actually valuable here? If it's operational data plane, that might be speed. If it's mailing address for sending out holiday cards to all your customers, accuracy is probably better than completeness. If a few clients don't get a card, that's probably better than sending out lots to fake addresses.
In Dacil's experience, business is typically the first team to notice data quality issues. And that means their trust is broken. Trust is hard to build but much harder to rebuild. How can you be data driven if you don't trust your data? People need to understand that issues will come up but put the rules in place - and show the business - to prevent that same issue from happening again. It was a data downtime incident, treat it just like you would with a software incident.
Other tidbits:
Data ownership is often like a hot potato, no one wants to catch it. You can't throw that responsibility to the business and expect them to know how to handle it.
Talk to potential data consumers before creating a data quantum or anything similar in data. You won't know what you can offer that will be valuable to them until you know what they want. So have the conversation, drill into what is of value and why, then collaborate together to drive to that value.
It's easy to lose sight that optimizing to turnaround time doesn't typically optimize time to actual value and you create hard-to-support data assets.
Break your changes into much smaller pieces. Big changes are more prone to failure and are harder. Make incremental, small-scale progress.
Measure if centralization is actually your challenge before you look at implementing data mesh. If it's not, will data mesh be worth it for you?
Really think if you are mature enough to really do federated governance and decentralized data ownership. Centralized governance is a bottleneck but that will likely be far better than chaos if you aren't ready.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Alex's LinkedIn: https://www.linkedin.com/in/alex-bross-33853837/
In this episode, Scott interviewed Alex Bross, VP of Data Engineering at Fifth Third Bank. To be clear, Alex was only representing his own views on the episode.
Some key takeaways/thoughts from Alex's point of view:
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
LinkedIn: https://www.linkedin.com/in/brandonbeidel/
In this episode, Scott interviewed Brandon Beidel, Director of Product at Red Ventures.
Some key takeaways/thoughts from Brandon's point of view:
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter. Oh, and she's hiring.
What can we do now relative to data mesh with what we have? People want to move, not wait for the tools to evolve. We can start to shift in anticipation of tooling getting better. It might not make things a ton better now, but when tools start to emerge, then we can jump ahead quickly. Learn from what happened in the API revolution and don't compromise on interoperability - that will just lead to high quality data silos, which is not a great outcome. And we need to get to a place with data where consumers have a delightful experience going from discover to learn to trust to use with little friction.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code...
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Katharine's LinkedIn: https://www.linkedin.com/in/katharinejarmul/
Practical Data Privacy (Katharine's book in early release): https://www.oreilly.com/library/view/practical-data-privacy/9781098129453/
Katharine's newsletter: https://probablyprivate.com/
'Privacy-first data via data mesh' article by Katharine: https://www.thoughtworks.com/insights/articles/privacy-first-data-via-data-mesh
danah boyd [sic] website: https://www.danah.org/
In this episode, Scott interviewed Katharine Jarmul AKA K-Jams, Principal Data Scientist at Thoughtworks.
Some key takeaways/thoughts from Katharine's point of view:
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter. Oh, and she's hiring.
Who will be the data product developer in data mesh? There has been a misconception in Zhamak's view that the application developers should be the ones focused on building the data products as well - but she thinks they already have a full-time role :) But, we need someone applying software engineering practices and data know-how to building data products. Right now, to do data work, you need way too much tool knowledge instead of data understanding. We have hyper specialized data roles - ML engineer, data engineer, etc. - when we should have data developers that can tackle these challenges better.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code...
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Mozhgan's LinkedIn: https://www.linkedin.com/in/tavakolifard/
In this episode, Scott interviewed Mozhgan Tavakolifard, Data and AI Lead for the Nordics at Accenture. To be clear, she was only representing her own views on the episode.
Before we jump in, most of the conversation was about external data marketplaces rather than internal data marketplaces within an organization. It's also important to note that data marketplace technology and implementations are still in the relatively early stages - it's quickly evolving and maturing.
Some key takeaways/thoughts from Mozhgan's point of view:
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here. You can download their Data Mesh for Dummies e-book (info gated) here.
Mariana's LinkedIn: https://www.linkedin.com/in/mariana-hebborn-phd-118035117/
In this episode, Scott interviewed Mariana Hebborn, Lead of Data Governance for the Healthcare Sector at Merck Group Germany (not Merck, the pharmaceutical company).
Some key takeaways/thoughts from Mariana's point of view:
Data Mesh Radio Patreon - get access to interviews well before they are released
Episode list and links to all available episode transcripts (most interviews from #32 on) here
Provided as a free resource by DataStax AstraDB; George Trujillo's contact info: email (george.trujillo@datastax.com) and LinkedIn
For more great content from Zhamak, check out her book on data mesh, a book she collaborated on, her LinkedIn, and her Twitter. Oh, and she's hiring.
What tech is already available that could be used for data mesh? There are so many amazing approaches and technologies in data but they've been used for the pipeline approach only. We need to think more like developers - not accepting the grunt work or death by a thousand cuts of data - and take a hard look at what we've done historically in data and what should be replaced.
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): AstraDB
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1I0dD4CGQyVxfQHzEF6EbEqN-ossHUnGtjt2zRrL_bPw/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Jill's LinkedIn: https://www.linkedin.com/in/jillianmaffeo/ (https://www.linkedin.com/in/jillianmaffeo/) Developing Interoperable Channel Domain Data (blog post): https://vista.io/blog/developing-interoperable-channel-domain-data (https://vista.io/blog/developing-interoperable-channel-domain-data) In this episode, Scott interviewed Jill Maffeo, Senior Data Product Manager at Vista. Before jumping in, Jill gives a lot of very useful examples of outcomes they've been able to drive that could be abstracted to apply to your own organization's business challenges. Outcomes like better customer segmentation, faster time to launch new offerings, etc. If you are having difficulty with stakeholder buy-in, especially for someone in marketing, this episode could help you frame things in their language.
Some key takeaways/thoughts from Jill's point of view: "When you're thinking about interoperability, it's just playing nice, right?" If you think of interoperability as a key part of your culture, it's easier to implement. Let people know why interoperability is good for them and the whole company. Taxonomies help drive interoperability because there is already an established language even if things don't fit perfectly. New concepts can emerge and your taxonomies should change. But it makes the interoperability discussions have at least a common starting point. Taxonomies are a living thing - make sure they aren't overly rigid and be prepared to continually evolve and improve them. Within your taxonomy structure, if there is a reason for things to be unique for a domain or use case, that is okay. Look for potential ways to also convert that data to best fit your taxonomy but you don't want to force a square peg through a round hole. Taxonomies really start to add a lot of value at scale. They are somewhat costly upfront with likely moderate return on investment early; but, if you do them right, will pay back a lot as you move forward. They make historical analysis - especially with interoperability - far easier because you've done the work ahead of time. Taxonomy, when done well, is about balancing standardization and flexibility. Much like most things in data mesh, it's about finding the right balance for your organization. Most customer journeys are cross domain. The more you make domain data interoperable, the more insight you have into the customer that can drive better business with your organization as a whole instead of only trying to locally optimize value by domain....
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/101DSo19l1KiocNNX7JFgbL6X2illb6lLjuZs5dEXXCI/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Carlos' LinkedIn: https://www.linkedin.com/in/carlos-saona-vazquez/ (https://www.linkedin.com/in/carlos-saona-vazquez/) In this episode, Scott interviewed Carlos Saona, Chief Architect at eDreams ODIGEO. As a caveat before jumping in, Carlos believes it's too hard to say their experience or learnings will apply to everyone or that he necessarily recommends anything they have done specifically but he has learned a lot of very interesting things to date. Keep that perspective in mind when reading this summary. Some key takeaways/thoughts from Carlos' point of view: eDreams' implementation is quite unique in that they were working on it without being in contact with other data mesh implementers for most of the last 3 years - until just recently. So they have learnings from non-typical approaches that are working for them. You should not look to create a single data model upfront. That's part of what has caused such an issue for the data warehouse - it's inflexible and doesn't really end up fitting needs. But you should look to iterate towards that standard model as you learn more and more about your use cases. ?Controversial?: Look to push as much of the burden as is reasonable onto the data consumers. That means the stitching between data products, the compute costs of consuming, etc. They get the benefit so they should be taking on the burden. Things like data quality are still on the shoulders of producers. You should provide default values for your data product SLAs. It makes the discussion between consumers and producers far easier - is the default good enough or not? ?Extremely Controversial?: At eDreams, you cannot publish data in your data product that you are not generating. In derived domains (e.g., customer history), “generate” includes the derived stitching. NOTE: Go about an hour into the interview - not episode - for more specifics. When starting with data mesh, there must be a settling period - consumers must understand that things are subject to change while a new producer really figures things out for the first few weeks to months. You want to avoid duplicating data. But you REALLY want to avoid duplicating business logic. Be careful when selecting your initial data mesh use cases. If the use case requires a very fast time to market, while it has value, you likely won't have the time and space necessary to experiment and learn. You need to find repeatable patterns to scale in data mesh. Hurrying is a way to miss the necessary learning. Look ahead and build...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/14bbqPIo_6FcpBy0IdwXCtGlTZjb9FN9gO9J1q27Y68A/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Adelle's LinkedIn: https://www.linkedin.com/in/adelle-mcdonald-79a9a2139/ (https://www.linkedin.com/in/adelle-mcdonald-79a9a2139/) In this episode, Scott interviewed Adelle McDonald, Customer and Origination Lead at ANZ Plus, a bank in Australia and New Zealand. Some key takeaways/thoughts from Adelle's point of view: To drive buy-in, the 1:1 conversations with domain owners - the business leaders - you will have to tailor your conversations to each person. Listen to their pain and reflect it back to them. Focus on an ability to quickly pivot with low cost. That can mean things aren't as product-worthy to start but it means you can evolve towards value more quickly. Addressing domain owners' pain points gets them looking at you as a partner. They will be much more willing to work with you, especially as you partner to provide actionable insights. ?Controversial?: ANZ Plus is embedding data leads into domains to handle the data quanta for the domain and also build the team what they need from data. As part of that, they are slowly building up the domains' capabilities to handle their own data. This minimizes friction and creates buy-in but is likely not long-term sustainable - ownership will need to be transferred. Very important to tie the data quanta to use cases - driving value for users means focusing on use cases. Developers or software engineers owning data is complicated. Make it so they can start to make small changes and learn in a safe way instead of dumping all ownership on them at once. Ownership and knowledge aren't a switch you flip. Using a git-based, pull request approach, developers can attempt data work without manual stitching so they learn to do the work themselves; but it can still be easily overseen by someone with more data expertise. One way to potentially drive executive buy-in is joint/collaborative KPIs. So it's not just about their domain's results but how well they work and drive results with another domain. ?Controversial?: It's okay to have a data asset with murky long-term ownership at first. If usage picks up, you want to convert it to a proper data quantum but we need to be able to test the waters with things and see if they are actually useful first. Clarity comes with usage. When creating anything data related, use a software development lifecycle (SDLC) approach. Domains may create something exclusively internal to the domain but once you look to share externally, you have rules and standards and best practices. Move from the pipeline...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1ZUvYRq9zr8TRGJUpBJ5TqMM9kpUjKTTUV9EwTy_OiHA/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Gunjan's LinkedIn: https://www.linkedin.com/in/gunjanaggarwal/ (https://www.linkedin.com/in/gunjanaggarwal/) Gunjan's Medium: https://gunjan-aggarwal.medium.com/ (https://gunjan-aggarwal.medium.com/) In this episode, Scott interviewed Gunjan Aggarwal, Head, Digital Data Products and MarTech Strategy at Novartis. To be clear, she was only representing her own views on the episode. Some key takeaways/thoughts from Gunjan's point of view: Set your overall data product strategy - for when you are in stage 2, going wider with data mesh - earlier in your journey than many may think. It's easy to focus only on use cases instead of the bigger picture. Make sure to align early on who owns what - what are the clear boundaries between roles. Otherwise, with the amount of change data mesh drives, there will likely be unnecessary chaos. Get specific. Don't fall to the 'Data Field of Dreams' - "if you build it, they will come." Focus on building to actual problem statements. Involve people early, make them accountable, give them skin in the game and they will care. "The more you ask why, the more clarity you will get." Really dig in deep into the reasoning for creating new data products or ARDs (analytics ready datasets). If we have this data product, what will it unlock for us? It's crucial to avoid the trap of building data products specifically to use cases. You must have the bigger picture in mind and focus on reusability instead of only solving one set of challenges. Can you extend an existing data product? Data people should have domain knowledge where possible. That way, they can push back on requirements that don't make economic sense, that don't maximize the return on investment. 4 part approach to designing data products: 1) find clarity on the problem statement; 2) assess who are the personas that will benefit from it; 3) dig into what you already have available; and 4) focus on serving value to the problem statement in a way the persona can use. ?Controversial?: scalability is more important than time to market when it comes to data products, especially as you develop a broader set of data products. Tech debt around scaling is hard to combat just as you are delivering strong value with the need for scale. And it limits additional use cases leveraging existing data products if they can't scale. Look to provide as many easy paths as possible for new data products. Templates, blueprints, standard schema, a global taxonomy, etc. They don't have to use them but they are great...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/14uVlti-v7tHuZVE3KKOGaeqDPKE018LOVFfdVjUWhmQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Dmitriy's Twitter: @squarecog / https://twitter.com/squarecog (https://twitter.com/squarecog) The Missing README book: https://themissingreadme.com/ (https://themissingreadme.com/) Building Evolutionary Architectures book: https://www.oreilly.com/library/view/building-evolutionary-architectures/9781491986356/ (https://www.oreilly.com/library/view/building-evolutionary-architectures/9781491986356/) In this episode, Scott interviewed Dmitriy Ryaboy, CTO at Zymergen and co-author of the book The Missing README. Some key takeaways/thoughts from Dmitriy's point of view: Organizational design and change management is "like a knife fight" - you are going to get cut but if you do it well, you can choose where you get cut. There is no perfect org, and there will be pain somewhere but you can influence what will hurt, and make it not life-threatening. There is too much separation between data engineering and software engineering. Data engineering is just a type of software engineering with a focus on dealing with data. We have to stop treating them like completely different practices. When communicating internally, always focus on telling people the why before you get to the how. If they don't get why you are doing it, they are far less likely to be motivated to address the issue or opportunity. This applies to getting teams to take ownership of the data they produce, but also to everything else. There is often a rush to use tech over talk. Conversation is a powerful tool and will set you up so your tools can help you address the challenges once people are aligned. Paving over challenges with tech will not go well. Build your data platform such that the central data platform team is unnecessary in conversations between data producers and data consumers. That way, your team won't become a bottleneck. Try to reduce cognitive load on the users - they shouldn't have to deeply understand the platform and its inner workings just to leverage it. "Data debt is debt forever." You can certainly 'pay it down' but data debt typically has a much longer life than even the initial source system that supplied the data. Take it on consciously. Looking to hire or grow full-stack engineers for an ever-growing definition of stack (backend, frontend, security, ops, QA, UX, data...) is probably not a great idea, we can't keep piling new domains on people and expect them to be good at all of them. Instead, look to build full-stack teams, and tools that look and feel sufficiently similar that e.g. "data...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1QIeIbx8CfqMJW78Gf3hBgj4LU1a7vujNrQdMGwyQYPo/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Karin's LinkedIn: https://www.linkedin.com/in/karin-hakansson/ (https://www.linkedin.com/in/karin-hakansson/) In this episode, Scott interviewed Karin Håkansson, Data Governance Lead consulting on a large data mesh implementation. To be clear, Karin was only representing her own views on the episode. If there is one theme that resonated throughout the conversation, it was have more and deeper conversations with your data governance teams. Don't make assumptions and as a data governance team, don't make decrees - explain what you are trying to accomplish and why. And look to help where possible instead of telling people how to do things.
Some key takeaways/thoughts from Karin's point of view: Data governance teams need to get closer to both data producers and consumers; but when doing that, emphasize it isn't about control or oversight. You are there to help. Data producers and consumers need to share with the governance team about their pains and challenges - the governance team in most orgs is there to help. So get specific about where you need help. Emphasis on specific! To get closer to data producers and/or consumers, identify a few obvious data problem areas and just reach out to those involved/impacted. Focus on asking not telling when digging into the pain points. And involve all stakeholders in assessing potential solutions to get them bought in to the eventual solution. A control-based data governance approach doesn't add value to your data. We need new ways of working in data governance focused on adding value, not adding roadblocks. Data mesh creates an environment where we need to look at things differently. And it provides the excuse to look at them differently too. Challenge your base assumptions. If your organization isn't open to finding and implementing new ways of working, it will be challenging - at best - to implement data mesh. Really ask if you are ready for the change before implementing it. Data governance teams need to focus on finding long-term solutions and balancing those with temporary fixes. Be extremely clear about what is a temporary fix and why it's temporary and must be replaced. Always keep people informed as to progress - even if it is "no progress". Humans aren't good with uncertainty. Data governance teams can easily be overwhelmed if you don't develop good ways to prioritize. Find the most important - not necessarily the most obvious - pain points and assess the value of fixing them.
Karin started the conversation...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1cIQo0dVDI2rKqYkzVJvhj-ipjd3t7kC66aWyGDmRBk8/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Scott's LinkedIn: https://www.linkedin.com/in/scottmztaylor/ (https://www.linkedin.com/in/scottmztaylor/) Scott's Website: https://www.metametaconsulting.com/ (https://www.metametaconsulting.com/) Scott's Book: https://technicspub.com/data-storytelling/ (https://technicspub.com/data-storytelling/) (save 20% off with code DATAWHISPERER) Scott's YouTube playlist: https://www.youtube.com/playlist?list=PLashWxBySOAOKUtvn2NTBQeENJTLB8NoS (https://www.youtube.com/playlist?list=PLashWxBySOAOKUtvn2NTBQeENJTLB8NoS) In this episode, Scott interviewed Scott Taylor AKA the Data Whisperer. From here forwards, "Scott T" will be used to represent Scott Taylor to differentiate from your host. It's also important to note Scott T differentiates 'data management' from 'analytics', so things like data governance and infrastructure fall under data management.
Some key takeaways/thoughts from Scott T's point of view: To actually reliably successfully obtain funding for data management initiatives, you need to articulate the value of the work to the business - why does this matter? Storytelling is key. The quickest way to lose your shot at securing funding for data initiatives is to focus on the how instead of the why. Most execs do not care at all how it gets done. They hired you, the data leader, to handle that. To get funding, focus your messaging on tying your data strategy and proposed data initiatives to the business strategy. Why does this matter - how will this drive the business strategy forward? Be prepared for cynicism from the business side - many have heard about how this or that data/technology initiative will be the silver bullet for far too long. Digital transformation is an enormous opportunity to change how your organization deals with data. Use that as a good point of leverage for funding your data management - it's a necessary foundation for a digital enterprise. 'Story' is about constructing a narrative and 'telling' is about effectively communicating. So storytelling is about effectively communicating what is the target outcome and why will this drive value. Data storytelling is an art, not a science. You need practice to get good at it. Don't make your 'practice' be overly high stakes. Talk to a lot of folks on your side to hone your talking points. You don't publish the first draft of a book do you? Find your editors. Focus on empathy and listening to find your audience's aspirations and pain points. Then tie your messaging about your data initiatives to those...
Inverse Conway Maneuver Definition: https://www.thoughtworks.com/radar/techniques/inverse-conway-maneuver (https://www.thoughtworks.com/radar/techniques/inverse-conway-maneuver) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1uyz7WhEasFe-CqkJvBrUR9p9kLTsHi08K74ZCZfVbZk/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Glovo's meetup group: https://www.meetup.com/glovo-tech-talks/ (https://www.meetup.com/glovo-tech-talks/) Javo's LinkedIn: https://www.linkedin.com/in/javiergrandag/ (https://www.linkedin.com/in/javiergrandag/) Javo's Twitter: @JavierGrandaG / https://twitter.com/JavierGrandaG (https://twitter.com/JavierGrandaG) Pablo's LinkedIn: https://www.linkedin.com/in/pabloginerabad/ (https://www.linkedin.com/in/pabloginerabad/) In this episode, Scott interviewed Pablo Giner Abad, Global Director of Data and Javier "Javo" Granda, Senior Data Manager at Glovo. From here forward in this write-up, P&J will refer to Javo and Pablo rather than trying to specifically call out who said which part.
Some key takeaways/thoughts from P&J's point of view: It's okay to not fit the exact or complete picture of data mesh in your early journey. Focus on what matters to your org and implementation and focus on learning over trying to be perfect. Iteration is possible and not too costly with data mesh. That's sort of one of the main points of data mesh. When selecting your first use case, look for high value and low dependencies. The less cross-team coordination work needed to actually get to an initial end data product that has value, the better. And buy-in is much easier if the producers are one of the consumers too :) When starting out, really look at how thin of a slice you can get away with for your MVP. Be prepared to make some hard compromises. Make them with your eyes open. It's tech debt but taken on consciously. Focus on solving your problems of today instead of trying to solve all your future problems. Fixing the challenges of today will set you up to fix the challenges of 6 months from now in 6 months. Focus on reducing cycle times to creating and iterating on data products more than you probably think you should. It's easy to get focused on delivering new data products instead of the capabilities to deliver new data products but that will cost you more and more as your data mesh implementation matures. An important quote to remember re product thinking: "If you aren't embarrassed by the first version of your product, you shipped too late." Your data products don't need to be perfect when launched. Just using your domain mapping from your operational/DDD side as your data domain map is likely to lead to some big challenges. Look to how your data flows to figure out good data domain mapping. Misaligned domain maps between the operational side and data side can also cause issues
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1MHCMG57ZaQsn436xVfqaXO73sOiS05ASvPmnjkqs5zU/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Darshana's LinkedIn: https://www.linkedin.com/in/darshana-thakker/ (https://www.linkedin.com/in/darshana-thakker/) In this episode, Scott interviewed Darshana Thakker, Architecture Director at BCG Platinion. Some key takeaways/thoughts from Darshana's point of view: Data can be the enabler to achieving your business goals but only if you actually tie your data work to your business strategy and goals. Speaking data jargon when talking to the business stakeholders makes it hard to actually communicate. Focus on speaking to business outcomes and business value. When selecting your first use case(s) for data mesh, evaluate domains on four different metrics: business value, capabilities, eagerness, and feasibility. There isn't a golden formula but it's likely some domains/use cases will rise to the top pretty quickly. Look to - as best as you can - quantify the incremental business value from something like data mesh. That can be at the micro level - the single use case - or the macro level. But getting specific instead of "let's be data-driven" will lead to better buy-in and partnering with the business side. You need to prepare to evolve your architecture. You can't have things set in stone. But you also can't just throw things against the wall and see what sticks. You need a balance between overly rigid and overly flexible. If your centralized data function isn't a bottleneck, isn't the cause of your data challenges, data mesh probably isn't right for your organization - or at least it doesn't directly address your challenges. Time between identifying useful data and making it reliably available is a good place to look to measure if the centralized team is your bottleneck. A good buy-in driving question for data mesh - or any data initiative - is "what's the cost of doing nothing?" Will you miss opportunities? Lose market share? How much are you "paying" on your existing tech debt? If we don't move, what is the cost to our future business prospects? When looking at buy-in, you need to drive from the top down - so you can actually make necessary large-scale org-wide changes - and bottom up - so you are working well with the people doing the actual implementations. Data mesh feels a lot like the Agile movement in that people expect it to be a silver bullet instead of a framework for thinking - and re-thinking - about how you approach your work. As also recommended by previous guests, look to a vertical thin slice for your data mesh MVP. Make sure you include all necessary...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1CAIgDnlSNR0u9jlgV88zRuVYd0GMznSdt-F_qTJqOA8/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Squirrel (OSS data platform) GitHub: https://github.com/merantix-momentum/squirrel-core (https://github.com/merantix-momentum/squirrel-core) Alireza's LinkedIn: https://www.linkedin.com/in/alireza-sohofi/ (https://www.linkedin.com/in/alireza-sohofi/) In this episode, Scott interviewed Alireza Sohofi, a Data Scientist focused on building the data platform at Merantix Momentum. Some key takeaways/thoughts from Alireza's point of view - some written directly by Alireza himself: Where possible, look to build your platform in a loosely coupled way. It will make it easier to extend and evolve; and domains can replace pieces, mix and match components, or even extend the functionalities when it makes sense. It's easy to fall into the trap of building a platform that is hard to evolve and support. Be very conscious about what you want to include - and not include - in your platform. Don't try to solve every challenge with a point solution. To effectively share data - and the information it represents - software engineers / domains need to really understand their own data, including data modeling. That can't be easily outsourced. A platform team's job is to build the tooling so those domains only need to deal with the data, not the data engineering. If you want a scalable platform - in many senses of the word scalable -, your platform should be relatively generic. It must also be easy to extend and augment. Focus on providing flexibility and ease of customization. One size definitely won't fit all. Packages and templates are both useful but templates are typically more user friendly and easier to customize - start with templates when possible. If there is a need for customizing or extending a package or template, it's better to first build it within a domain (with the help of the platform team if necessary). The generalized version of the new feature is then contributed to the platform. This leads to a more integrated domain-platform, more robust first release of new features in the platform, knowledge sharing, and avoiding bottlenecks that may arise if only relying on the central platform team. Platform teams need to A) dog food the platform - you will learn far more by using it; B) provide good methods of communication for domains to give feedback and requests; and C) find better ways to exchange context with your domains regularly, e.g. pair work and scheduled informal chats. The platform consists of several tools that should not only work well together, but should also work well...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1QX1bEWoa2S6tZYH-KeYp_FZVAGZq0E4s0M6vQdxvTBg/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Balvinder's LinkedIn: https://www.linkedin.com/in/balvinder-khurana/ (https://www.linkedin.com/in/balvinder-khurana/) In this episode, Scott interviewed Balvinder Khurana, Principal Data Architect at Thoughtworks. Some key takeaways/thoughts from Balvinder's point of view: Data mesh is NOT a silver bullet and not everyone is ready to do data mesh - others have stated that but it's crucial to repeat. A data mesh doesn't happen in a vacuum - you need to assess if you are really ready and does it align first to your business strategy and second to your data strategy. If you decide to move forward on a data mesh implementation, really consider how you will measure progress and success against business goals. To evaluate data mesh appropriately, consider what business value having better data practices would bring to your company and is your company aligned into lines of business or would you need to reorganize your business. Are you prepared to extend your line of business practices to data? A common failure pattern in analytics has been not looking at the Intelligence Cycle - changing your operational systems and processes as a result of insights. Don't just generate insights, insights must generate action! Data mesh must avoid this too. Even if existing centralized data team setups have significant bottlenecks, data consumers typically eventually get their needed data. Those data consumers can see something like data mesh as a risk - will they still be able to eventually get the data they need? Is eventually getting to faster access to new data worth the perceived risk? If you have resistance to data mesh, look at delivering necessary capabilities to your data producers and/or consumers in a small-scale, incremental way and not specifically tying that in their mind to data mesh. Tie those incremental capabilities to business value. Look to constantly communicate the what and the why of improvements to your platform to drive engagement. What is the purpose? Why should they care? It's easy to fall into the trap of trying to iterate everything constantly. There is a concept of "good enough for now". Don't be focused on getting everything perfect, that juice is not worth the squeeze. A good signal to reevaluate your domain boundaries is: a dashboard is still used but the owner no longer uses it. They don't want to own it any more so you need to find a new owner. That probably means changing business boundaries if the original user doesn't use it anymore. To drive buy-in for data
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/16boNM9Jptb5Bdy5qDKVl4YOnk949NbFiCym_u5oxc-o/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). ammara.gafoor@thoughtworks.com Ammara's LinkedIn: https://www.linkedin.com/in/ammara-gafoor/ (https://www.linkedin.com/in/ammara-gafoor/) Data Mesh in practice article series from Ammara and colleagues:
In this episode, Scott interviewed Ammara Gafoor, Principal Business Analyst at Thoughtworks who has been working on a few client projects related to data mesh including one for well over a year. Before jumping in, it's important to note that much of Ammara's learnings come from an implementation in a 100K+ employee company split into 21 high-level domains. So the definition of domain in this episode revolves around that context of a very large business unit, not a two pizza team size sub domain. Some key takeaways/thoughts from Ammara's point of view: There often is a hangup around data work, especially relative to data mesh, where people want to get it all right, all perfect the first time. That's never going to work. Get something decent out there, test, and iterate. Perfect is the enemy of done. No bikeshedding! If you don't look to change domain's KPIs to align their operational work to data mesh "you won't prioritize it - you cannot prioritize it." Make it easy for domains to prioritize data mesh work if you want it to get done. ?Controversial?: source oriented data products should not be made available to business users within the domain or to almost anyone in other domains - at least by default - as they are difficult to understand for anyone other than the highly data literate people in the domain. "Don't make things that you don't need yet." Build data products for use cases you've identified. Think of
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1uQ-o4BewpWADsUWTUXz6vayygnIiqAnp-q0X_PS6We4/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Data Mesh at PayPal blog post: https://medium.com/paypal-tech/the-next-generation-of-data-platforms-is-the-data-mesh-b7df4b825522 (https://medium.com/paypal-tech/the-next-generation-of-data-platforms-is-the-data-mesh-b7df4b825522) JGP's All Things Open talk (free virtual registration): https://2022.allthingsopen.org/sessions/building-a-data-mesh-with-open-source-technologies/ (https://2022.allthingsopen.org/sessions/building-a-data-mesh-with-open-source-technologies/) JGP's LinkedIn: https://www.linkedin.com/in/jgperrin/ (https://www.linkedin.com/in/jgperrin/) JGP's Twitter: @jgperrin / https://twitter.com/jgperrin (https://twitter.com/jgperrin) JGP's YouTube: https://www.youtube.com/c/JeanGeorgesPerrin (https://www.youtube.com/c/JeanGeorgesPerrin) JGP's Website: https://jgp.ai/ (https://jgp.ai/) In this episode, Scott interviewed Jean-Georges Perrin AKA JGP, Intelligence Platform Lead at PayPal. JGP is probably the first guest to lean into using "data quantum" instead of "data product". JGP did want to emphasize that as of now, he was only discussing the implementation for his team the GCSC IA (Global Credit Risk, Seller Risk, Collections Intelligence Automation) within PayPal. Some key takeaways/thoughts from JGP's point of view: Data mesh as it's been laid out by Zhamak obviously leaves a lot of room for innovation. For some, that's great. For others, they want the blueprint. And it's okay to wait for the blueprint. But JGP and team are excited to innovate! PayPal's 3 main initial target outcomes from data mesh: A) faster and easier data discovery, B) easier to use the data in a governed way, and C) increase data consumer trust in data. PayPal's initial data consumers are data scientists so their platform and data quanta are built to serve that audience first. Really consider what you want to prove out in your MVP. Is that minimum viable A) data quantum, B) data platform, C) data mesh, or D) something else? Only doing a data quantum probably sets you up for trouble and a platform only won't be tested until it has data quanta on it. Data contracts are crucial to making trustability actually measurable and agreed upon. Otherwise, it's far too easy to have miscommunication between data producers and consumers, which leads to lack/loss of trust. Producers, don't set your data contract terms too strictly when first launching a data quantum. There's no need to over-engineer - despite how interesting that can sometimes be... For too long, we have tried to keep software...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1I92-lwPsbB_ofuBuoSAe9O2SLQBSE9HRGAgn0ZfoZjc/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Nicola's website: https://www.nicolaaskham.com/ (https://www.nicolaaskham.com/) The Data Governance Podcast: https://www.nicolaaskham.com/podcast (https://www.nicolaaskham.com/podcast) Nicola's LinkedIn: https://www.linkedin.com/in/nicolaaskham/ (https://www.linkedin.com/in/nicolaaskham/) In this episode, Scott interviewed Nicola Askham, a data governance consultant known simply as The Data Governance Coach and host of The Data Governance Podcast. Some key takeaways/thoughts from Nicola's point of view: The key point of data governance: ensure the data we use is the right data for the right people to better address business challenges. Everything you do with governance should circle back to that. In doing data governance right, you need to set yourself up to take in feedback and iterate. You absolutely won't get everything right upfront. It's crucial to set expectations that your data governance approach will evolve as you learn more. If you see data mesh as being about making better data more accessible to your current data consumers, that's a very big opportunity wasted. Aim to significantly expand your pool of data citizens. Not everyone should be a data scientist but data should play a role in much more people's jobs. To get going in data mesh, you need to get your data governance to "good enough" and start moving forward. Think about what you need - is it the very complicated standard to last the next decade or is it about getting people to understand and trust the data they can now access? Probably the second... To drive buy-in for data governance, you should tailor your message to the audience. It's very hard to have universal appeal around a specific selling point of data governance but data governance can - and should - drive value for everyone. Every data governance approach should be tailored to the organization but it should start from a few building blocks: A) policy; B) processes and standards; and C) roles and responsibilities. (More info below) A centralized data governance team making decisions about what to do with specific data will not scale - they just can't have the context/knowledge needed. So federated governance has been the sensible approach for a long time, it's just not necessarily easy to do right. Or at least it's quite easy to do wrong. Central governance teams are crucial - they make it easy for federated teams to do what's necessary to comply with regulations and internal standards but with as little friction as possible. The central...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/19AQsVHHz3Gonkj3L5APC3WFJcjTZiy_bIhJrdWHl2iM/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Winfried's LinkedIn: https://www.linkedin.com/in/winfried-adalbert-etzel/ (https://www.linkedin.com/in/winfried-adalbert-etzel/) MetaDAMA podcast (episodes in both English and Norwegian): https://podcasts.apple.com/us/podcast/metadama-en-helhetlig-podcast-om-data-management-i-norden/id1572639573 (https://podcasts.apple.com/us/podcast/metadama-en-helhetlig-podcast-om-data-management-i-norden/id1572639573) In this episode, Scott interviewed Winfried Etzel, Information Strategy Consultant at Bouvet, Board member of DAMA Norway, and host of the MetaDAMA podcast. Some key takeaways/thoughts from Winfried's point of view: Transformation teams - a team that helps domains to transform and upskill by embedding in the domain for a time - is an exciting pattern for data mesh. Transformation teams can collect information more easily and find patterns as they are closely collaborating with more domains. Think of transformation teams like a personal trainer - they help you get in shape so you learn what exercises to do and how but then you can work on your own. And you can engage them again if you need more help. How can we give guidance and enable change at the same time - the transformation teams shouldn't do the work for the domains but need to work with them. Can we go broad in the organization if we only have a limited number of transformation teams? Do we need to be in that much of a hurry where that's an issue? There's definitely some vendor washing going on around data mesh. It's often difficult to determine what is new and what they've recycled to sound new. In data, we should look to take many learning from software engineering but also other disciplines. History, political science, law, manufacturing, etc. all have something to teach us to better approach sharing our information with each other. "Data is a message to people in the future." What do you want to tell them? Data products are how we communicate with those people in the future. Data maturity models are crucial so you can self-reflect - are we really willing and capable to do what we need? And how can we measure how well we are maturing our capabilities? We need to adopt better change management principles in data if we really want to create data citizens. And you should explain how data is important to their role to drive buy-in. Potentially look at using a big bang approach to organizational changes if it is to something like upskilling or even ways of working. Doing a big bang approach for new responsibilities...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1ls5QawrOffb3VGIfZCmPYHG0v729Ye7Gu7oKRaOYi8c/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Blanca's LinkedIn: https://www.linkedin.com/in/blancamayayo/ (https://www.linkedin.com/in/blancamayayo/) Pablo's LinkedIn: https://www.linkedin.com/in/pablodoval/ (https://www.linkedin.com/in/pablodoval/) In this episode, Scott interviewed Blanca Mayayo - Product Manager, Data Platforms - and Pablo Alvarez Doval - Lead of Data Platforms and Principal Data Architect - at Plain Concepts. From here forward in this write-up, B&P will refer to Blanca and Pablo rather than trying to specifically call out who said which part. Some key takeaways/thoughts from B&P's point of view: It's easy to fall into adding fit-for-purpose capabilities to your data platform but don't. Stay focused on managing your platform as a product - all aspects of it have lifecycles - and you can't try to fit every use case, especially before there is a need. If transformation, especially data transformation, is not tied to the business strategy, that is a major recipe for failure. You likely won't deliver good business outcomes. "Beware the proof of concept" - too many try to do a proof without the actual concept. What are you trying to prove and how will you decide/measure if you proved it? You can have everything necessary for a data initiative to succeed lined up - the sponsors, the will, the budget - and still fail. Nothing is 100%. Three common data initiative failure modes: 1) focusing only on the technology aspect and not does it meet needs and can we maintain and pay for it; 2) only treating it as an urgent tactical needs instead of playing into the broader data strategy; and 3) not considering how to actually do change management. Your platform is a product too - data as a product isn't just about mesh data products - think about capability lifecycle and how you communicate upcoming changes - especially deprecation - and help users migrate to the new capabilities. Acclimatize people to change and evolution. Most people in data aren't good with - or at least used to - evolution and preparing for said evolution because the cost of change for data has been so high. To do data as a product, you need the right balance of curious developers, risk/risk management, and capabilities. The most likely places to find reusability in your platform will be the mechanisms around data product production and maintenance - the lineage, CI/CD, data quality monitoring, security/compliance, etc. Reuse is crucial for a data platform - look to have data transformation and storage reuse of course but also really focus on...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/11HZ1Cf-LbJxGDKONwfJLxYQspsXVECPNkhcOM7KseyA/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). Get in touch with Simon and Sunny: data@esynergy.co.uk A webinar (info gated) Sunny and Simon did on data mesh: https://events.esynergy.co.uk/data-mesh-experimentation-to-industrialisation-on-demand (https://events.esynergy.co.uk/data-mesh-experimentation-to-industrialisation-on-demand) Simon's LinkedIn: https://www.linkedin.com/in/simon-massey-82718a3/ (https://www.linkedin.com/in/simon-massey-82718a3/) Sunny's LinkedIn: https://www.linkedin.com/in/sunnysjaisinghani/ (https://www.linkedin.com/in/sunnysjaisinghani/) In this episode, Scott interviewed Sunny Jaisinghani and Simon Massey who are both Principal Consultants at the consulting company esynergy. They have been involved in multiple data mesh implementations including at a large bank. This episode could also have been titled: Aligning Incentives, Reducing Friction, and Continuous Improvement/Value Delivery but it doesn't roll off the tongue very well.
From here forward in this write-up, S&S will refer to Simon and Sunny rather than trying to specifically call out who said which part as that leads to confusion.
Some key takeaways/thoughts from S&S's points of view: We are all still early in our learnings about how to do data mesh well. There is still a ton left to learn. Which is why people should share what they are learning more broadly. Helping others will help you. Data mesh, whether it's your overall implementation, your platform, your data products, your ways of working, etc. is all about evolution, incremental improvement, iteration, etc. You don't have to get it perfect upfront to drive immense value in the long run. Acclimatize people to iteration and lower the pain of change. Don't go for the big bang approach, find ways to continuously deliver incremental value. That builds the momentum necessary to drive data mesh broad in your organization. "Data mesh is 25% technology and 75% ways of working." Once you start to get into a groove with the organizational ways of working, that's when the value force multiplier in doing data mesh starts to take off. Fast time to market with new data products, quick iteration cycles, low-friction cross-domain collaboration, etc. But it takes time to really figure out how to do data mesh in your organization. Your data mesh will inherently have a huge scope. Try to keep that scope as limited as possible as you are getting moving - especially the near term. It is very easy to try to "feature stuff" your data mesh implementation, especially the platform and...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1LPYc0Be6Fq8latz41FNKLkDNiOfgiFxYjfFdWxxo4dU/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies e-book (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). In this episode, Scott interviewed Alla Hale, a Data Product Manager at Ecolab. To be clear, she was only representing her own views rather than anything on behalf of the company. She is hiring in Barcelona as well, see https://jobs.ecolab.com/job/15770737/data-product-manager-barcelona-es/ (here) Some key takeaways/thoughts from Alla's point of view: The most useful question in your quiver as a data producer is "what value would having this unlock for you?" It's not about pushing back, it's about skipping to collaborative negotiation. How can you work together to unlock business value? It's important to remember this when there is a data request: "they are coming to you because they need your help." Act accordingly - with empathy and patience. You need to really consider your data user experience (DUX) for your data products. How can you quickly get people past figuring out "what the product is" to leveraging the data product to drive value? You want users to enjoy using your data product. User stated requirements often do not match actual user needs. To maximize the return on your data work, look to exchange context to find the needs instead of just taking requirements at face value. And do so with patience and empathy. No prototype, no meeting = having something tangible - even if that is simply a process map on a Post-It note - for people to react to. Otherwise, what will the conversation be about? How do you prevent the meeting from being a waste of time unless there is a specific topic to address? We need to take lots of learnings/practices from tangible/physical goods product management when thinking about data products. We have users who have needs. How can we best serve those needs and drive value through that? It's always about serving the users' needs. Another product management learning - how can we do fast prototyping? Prototypes have an actual cost, even in data/software. What gets us to value quickly? How can we capture value early as we iterate towards product quality? All products - and their features - have a lifecycle. Be prepared to prune features or an entire product if they no longer drive more value than the cost to develop + maintain. Discussing sunsetting and pruning should happen with users even in the development process. Far too often, data consumers are used to data assets being available in perpetuity - even as they degrade. We need data consumers to be part of an active conversation about if they are still using something and how to create more...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), https://pixabay.com/users/itswatr-12344345/ (ItsWatR), https://pixabay.com/users/lexin_music-28841948/ (Lexin_Music), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/16yvSgx6S1tMsdEuUlIwNjclOYkw3RugdIldjaK9GnKw/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodmcenter&utm_term= (here). You can download their Data Mesh for Dummies (info gated) https://starburst.io/info/data-mesh-for-dummies/?utm_campaign=starburst-brand&utm_medium=outbound&utm_source=&utm_type=&utm_content=dmradiodnvid&utm_term= (here). In this episode, Scott interviewed Elena Samuylova, Co-Founder and CEO at the ML model monitoring company - and open source project - Evidently AI. This write-up is quite a bit different from other recent episode write-ups. Scott has added a lot of color on not just what was said but how it could apply to data and analytics work, especially for data mesh. Some key takeaways/thoughts this time specifically from Scott's point of view: A good rule of software that applies to ML and data, especially mesh data products: "If you build it, it will break." Set yourself up to react to that. Maintenance may not be "sexy" but it's probably the most crucial aspect of ML and data in general. It's very easy to create a data asset and move on. But doing the work to maintain is really treating things like a product. ML models are inherently expected to degrade. When they degrade - for a number of reasons - they must be retrained or replaced. Similarly, on the mesh data product side, we need to think about monitoring for degradation to figure out if they are still valuable or how to increase value. Data drift - changes in the information input into your model, e.g. a new prospect base - can cause a model to not perform well, especially against this new segment of prospects. That data drift detection could actually be a very useful insight to pass on as an insight - has something changed with our demographics? If so, what? When? Do we know why? Concept drift - the real world has changed so your model is not performing as expected - is a crucial concept in data and analytics too. Are we still sharing information about the things that matter? In a way that is understandable? Are we encapsulating what's happening in the real world in our mesh data products? Concept drift feels similar to semantic drift in the analytics world. So we can look to potentially take deeper learnings from how people approach and combat concept drift from ML and apply it to data mesh. How can we monitor degradation in mesh data products and prevent that degradation our data and analytics work? Historically, reports drifted further and further from reality with no intervention because the pain of change was so high. Are we fully reliant on the domain to know? Can we use software to help us detect semantic drift? Very early days on that one. ML models are designed to do one thing very well. Unfortunately, we don't have a good framework for reuse at the model level in ML. Maybe at the ML feature level? ML models have expected...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1ls5QawrOffb3VGIfZCmPYHG0v729Ye7Gu7oKRaOYi8c/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Argyris Argyrou, Head of Data, and Konstantinos "Kostas" Siaterlis, Director of Big Data at Orfium. There is a ton of useful information on anti-patterns, what is going well now, advice, etc. in this one.
From here forward in this write-up, A&K will refer to Argyris and Kostas rather than trying to specifically call out who said which part in most cases.
Some key takeaways/thoughts from A&K's points of view: On a data mesh journey: "It's not a sprint, it's a marathon." Pace yourself. It's okay to go at your own pace, don't worry about what other people are doing with data mesh, do what's right for you. Really focusing on the why and showing people results was a far better driver to buy-in and participation than any amount of selling about data mesh as a practice. Calling it data mesh when trying to explain it to people outside the data team didn't go well either... Orfium's "Data Doctor" approach - a low friction and low pressure office hours for a general staff data engineer - has really helped people help with data challenges and in spreading good data practices but without the "Doctor" becoming a bottleneck. The Data Doctor's role is to answer questions and provide guidance but not do the work for people. Then, take what was discussed and the best practice and document it for others to learn from - providing good leverage for scaling best data practices. In a smaller company like Orfium (~250 people), it's hard to justify a lot of full-time heads to implement data mesh. And trying to treat a data mesh implementation like a side-project also creates issues. There isn't a great answer here on exactly what to do except possibly take things slower than most startups are used to. Your data will still be waiting for you a few months later. If you are having difficulty driving broad buy-in, showing people what data mesh can do in action really helped at Orfium. Once they saw the approach delivering value, they wanted to participate. When trying to drive buy-in, specifically talking about data mesh didn't work well with non data folks. It's very easy to get confused around data mesh for data folks - just imagine it for non data folks. Trying to use Zhamak's articles as the optimal early state - where you need to be just to get moving - requires far too much work. Get to a place where you can try, learn, iterate, and repeat on your way to driving value. It's a journey! It's probably not a great idea for your first use case to be your most advanced or complicated - you will build your platform to focus on serving those needs instead of general affordances. Jen Tedrow's episode covers this quite nicely. Really assess how much additional work your data products will be for a data product owner. For Orfium, it was something to add to the existing product managers' plates as it wasn't a huge incremental burden just yet. Consider splitting your mesh data product ownership between business context ownership and...
For more great content from Zhamak, check out her book on https://www.oreilly.com/library/view/data-mesh/9781492092384/ (data mesh), https://www.oreilly.com/library/view/software-architecture-the/9781492086888/ (a book she collaborated on), https://www.linkedin.com/in/zhamak-dehghani/ (her LinkedIn), and https://twitter.com/zhamakd (her Twitter). https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1Zkr674u2niwHefNtVrCnwuwIZZi_amYJtThCvu29NKU/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Riya Singh, Business Insights Manager at Iterable. Some key takeaways/thoughts from Riya's point of view: ~4 years ago, Iterable was in essentially "spreadsheet hell" with lots of manual data work and no standard way of storing or sharing data across domains. While domains had good data capabilities, the integration and coordination between domains was very difficult at best. Most exec questions can't be answered by the data from a single domain so cross domain data integration became a key factor in Iterable continuing to grow. How could they make crucial decisions informed by data if there was so much manual work to try to integrate ad hoc? Could they really trust something done manually each time? Fast time to market for simple, base level capabilities of their data platform was much more valuable than trying to nail every feature upfront. Data consumers understood it wasn't perfect data at the start but it led to much faster exploratory data initiatives which led to valuable insights sooner. You might have a much higher ROI buying tools than trying to really get by on low-cost but not feature-rich tools. If you build a very cost-efficient data platform that no one wants to use, is that actually valuable? How much time will you spend managing the tools or is it worth it to outsource that to a vendor? Combining data across sales, marketing, and product meant Iterable could tailor marketing messages and find better prospects, measure marketing return on investment (ROI), and cost optimize their operations and product among many other new insights. As teams that previously weren't directly interacting start to have more conversations, gaps in your data - whether in data created/collected or data shared - will emerge. Filling those gaps will mean you can answer more high-value questions to drive the business forward. At Iterable, when there is a specific use-case identified for cross-domain data integration, the central data team takes over ownership of what would be considered a consumer-aligned data set in data mesh terms. With only 4-5 domains, Iterable doesn't need to decentralize the data team yet. The cost of decentralizing is far greater than the benefit right now. Iterable found the most value by doing exploratory data analysis then quickly moving to minimum viable consumable form. Then, they work to continue to improve the data set. This approach means a fast time value by grabbing the low hanging fruit while continually driving to better data and incremental value. But to do this, consumers must be very aware of what they are getting when :) A key way to keep stakeholders informed and bought in is by constantly keeping them updated on progress - Jen Tedrow talked about this in her episode. Keep people informed of progress or ongoing investigations so you can stay coordinated and all parties understand decisions along the way. At Iterable, conversations between domains are happening weekly. That...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/17LTNtckPujHjShS_tt7N9_e8omi7uPnkLls25lY6VzA/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Marisa Fish, Director of Information Management at American National Bank. To be clear, Marisa was only representing her own views on the episode. Some key takeaways/thoughts from Marisa's point of view: Understanding your data value supply chain - the way you derive and deliver value from your data - should be the crux of data and analytics work. The data value supply chain breaks down into sharing the data itself, sharing analytical insights about the data, and managing the data. All three are crucial to creating value from your data. Intentionality is crucial - instead of being reactive, stop and ask what are we trying to accomplish and what value will it drive. Then you will focus much more on high value-impact work. Similarly, think about system engineering work as "mission engineering" - what is your mission in doing your work? Does the work you are prioritizing serve the mission? When sharing information, start from: what is the point, what am I trying to drive with this information exchange? Are you trying to share one person's way of thinking or insights or give others the capability to derive their own insights from the new information? Both are very valid and useful but it's easy to talk past each other if you're not on the same page. So much of the way most organizations work with data is about the known knowns - the data consumer knows what data they want and what questions they want to answer with the data. We need to enable people with questions to find the right data to address them and people to also do data spelunking with data they aren't sure what it might tell them. Look to the Library and Information Sciences space for how to approach that. We need data librarians, not data publishers. Data publishers are about putting data on the shelf and serving only the known knowns. Data Librarians are there to help people find the information they need to address more of the unknowns - the value of curiosity in driving incremental valuable insights. There is a major mismatch in most organizations between what insights the business units are producing and the key questions the C-level execs care about. Consider creating a Chief Data Analyst type role to pair with execs to make sure insights are produced to support their initiatives, not just answer their questions as they come up. Think ahead, build ahead. Data teams need to take far more practices from general engineering - not just software engineering - so we learn how to better understand requirements. When requirement gathering, expecting the data consumers to know all of their requirements upfront can lead to data consumers asking for the world and a bad mismatch between asks and needs. Look to new ways to exchange information about requirements including the Japanese Obeya technique. Spend the time to ensure you understand how data consumers will derive value from the information you will share with them. That will give you a better...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1J4ikSSLv_VgY0yZYqz_xqwzB7BC0P4Q8P5PaCNZ18WU/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) Data Governance In Action: What Does Good Governance Look Like in Data Mesh - Interview w/ Shawn Kyzer and Gustavo Drachenberg In this episode, Scott interviewed Shawn Kyzer, Principal Data Engineer, and Gustavo Drachenberg, Delivery Lead at Thoughtworks. Both have worked on multiple data mesh engagements including with Glovo starting 2+ years ago. From here forward in this write-up, S&G will refer to Shawn and Gustavo rather than trying to specifically call out who said which part.
Some key takeaways/thoughts from Shawn and Gustavo's point of view: It's very easy for centralized governance to become a bottleneck. Make sure any central governance team/board that is making decisions has a way to quickly work through backlog through good delegation. Not every decision needs deep scrutiny from top management. To do federated governance right, you need to enable the enforcement - or often more appropriately the application - of policies through the platform wherever possible. Take the burden off the engineers to comply with your governance standards/requirements. Domains should have the freedom to apply policies to their data products in a way that best benefits the data product consumers. So if there are data quality standard policies, the data product should adhere to the standard for measuring completeness as an aspect of data quality but might be optimized for something other than completeness. The cost of getting anything "wrong" in data previously has been quite high because of how rigid things have been - the cost of change was high. But with data mesh, we are finding new ways to lower the cost of change. So it is okay to start with policies that aren't complete and will evolve as you move along. If you have an existing centralized governance board, that will sometimes make moving to federated governance ... challenging at best ... so you will need a top-down mandate to reshape the board. Look to meet the necessary representation across your capabilities (e.g. product, security, platform, engineering, etc.) but not create a political issue if possible. Look to add incremental value through each governance policy. And look to iterate quickly on policy decisions where you can. Create a feedback loop on your policies to iterate and adjust. It's okay to not get your policies perfect the first time, you can adjust them. Really figure out what you are trying to prove out in your initial proof of value/concept. If it's full data mesh capabilities, that can easily take 4-6 months. An interesting incremental insight: Zhamak has warned about organizations trying to scale too fast as an anti-pattern that may result in lots of tech debt or a failure of your implementation. An interesting incremental insight: in all of the data mesh implementations S&G have worked on thus far, the initial data product has not had any PII as that adds significant complications probably beyond what the value add of including PII would be in most cases....
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1-QkTaba0nnTgz95iht0uIhwXtSTV9msKLNPky7upmfs/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Martina Ivaničová, Data Intelligence Engineering Manager at the travel services company Kiwi.com. Some key takeaways/thoughts from Martina's point of view: The most important - and possibly one of the most difficult - aspect of a data mesh implementation is "triggering organizational change". Driving buy-in for something like data mesh is obviously not easy. As you are getting started, look to leverage 1:1 conversations to really share what you are trying to do and why and how this can impact them and the organization. These 1:1 conversations are crucial to developing early momentum. On driving buy-in for data mesh, really think about how to limit incremental cognitive load as much as possible on developers/software engineers. If you can keep cognitive load low, you are much more likely to succeed - succeed in driving buy-in and succeed in delivering value. When sharing internally about data mesh, it's important to focus on what it means to the other person. Using "data mesh" as a phrase can lead to a lot of confusion for people not on the data team. Make it clear what you are trying to accomplish - the what, the why, and the how. Using data-as-a-product as the leading concept resonated and worked well. Kiwi.com started driving buy-in by working with the engineering upper management, then found a few valuable and achievable first use cases to move forward. And they have kept cognitive low on the engineering teams while they learn how to deliver data as a product. If possible, the easiest way to drive buy-in is by finding a use case that is beneficial to the producing domain. If not, then look to spend the 1:1 time to really share why this matters. Kiwi.com is getting software engineers in domains to commit to simply sharing their data, not even really structuring into data products. So the software engineers in most cases are really only focused on maintaining high-quality data sharing mechanisms - read: pipelines. That is a relatively low initial cognitive load/low workload ask. Analytics engineers are creating the data products from the sourced data to satisfy consumer needs. Martina and team want to move to software engineers handling more of data product creation/management over time but it's a process. They plan for analytics engineers to upskill the software engineers by pairing with them closely. It might initially be more important to find a way to evaluate and iterate on what data is shared and how than getting to the most complex or valuable data product. You want to build the muscle around sharing data first before trying to go too big too soon. It's important to know what you are trying to prove out in your initial data mesh related deployment. It's okay to prove out you can produce data products before proving out you can build out the full mesh. A key success metric for a data mesh journey could be how many direct conversations and then actions come from data producers and consumers speaking
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/14w6rvgdPGBhqyIJPVp7y4z5lvzdo9ZsVkl4boA0AemY/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Laura Madsen, CEO at Moxy Analytics and author of the book Disrupting Data Governance. For the purposes of this write-up, when discussing data governance, it refers to the way many large organizations handle data governance at scale - a way that is very rigid and causes bottlenecks. We all know we can't stereotype or group every org together but general trends can be observed. Some key takeaways/thoughts from Laura's point of view: A big issue with today's data governance is that the concept of data stewards - the people who own the data concepts - is from 30 years ago and hasn't changed much despite the demands and scope changing dramatically. The data governance committee/council structure most organizations use is inherently inflexible and ineffectual. Those making the decisions don't really understand what's happening with the data under the covers and those who do understand have little ability to influence the wider committee outside their own domain. And thus, they become a major bottleneck. Data governance committees can be quite useful if they focus on communication and context exchange rather than driving decisions and work forward. To drive change in your data governance practices, you need to disrupt but not destroy. Start to break down the big picture into much smaller, bite-sized chunks that when you improve on them will incrementally drive value - Agile provides a good framework to approach this. You will absolutely have to throw out a LOT of your current data governance practices - over time - as you replace them with better ways of working. You will need to really evaluate each practice and assess if it will drive value or should be replaced. "Marie Kondo" your data governance practices - really look at your processes one by one and ask "does this spark value?" Reference: https://storables.com/storage-ideas/marie-kondo-method/ Current data governance practices does no provide incremental value to most organizations, they are about compliance and risk mitigation. If you can drive value creation, you can more easily drive change. People want to enable value creation or at least are hesitant to stop it in most orgs. Look for small ways to drive incremental value to build momentum. The current data steward and data ownership model essentially rewards innaction more than action. Action has risk and risk mitigation is a large part of the data steward and data owner's role. We need to change that relationship and reward enabling valuable use of data but within compliance. Laura is a fan of the hub-and-spoke model for data governance - and in general. To make hub-and-spoke work, 1) everyone has to really work on strong communication and 2) the central governance team cannot fall into the trap of trying to fix the data themselves, they must empower and enable the teams to fix their own data. Data governance teams must stop writing policies -> compliance and InfoSec should be doing that. Policies
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1abIWwqb5IBD9cqyZwXmF9Nwn-FicdP7iK4YAvmx6i40/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Anitha Jagadeesh, Principal Enterprise Architect at ServiceNow. To be clear, she was only representing her own views on the episode. Some key takeaways/thoughts from Anitha's point of view, some of which she specifically wrote: It is absolutely crucial to tie the data strategy to the business strategy. The business strategy must drive the data strategy which drives your data architecture. Architects need to lead the way in digging into use cases to get the specifics on what data producers are trying to solve for data consumers. Then, those architects can find the common patterns across use cases to tie to your organizational data strategy and also tie to your data architecture guidelines and principles. That way, instead of addressing challenges via point solutions, you can drive organization-wide choices that support many use cases via your data architecture. Architects also need to ask the probing questions to continuously tie work back to the business strategy and value or expected outcome for customers. If you aren't driving the business strategy forward, if you aren't helping the big picture, is the work worth doing? When it comes to data, companies shouldn't be entirely offensive - trying to leverage data for as much value as possible - or defensive - trying to minimize risk as much as possible. So organizations that have been very conservative need to push to be creative/offensive and high-risk organizations will get themselves into trouble if they don't start going defensive too. As we build the data strategy we have to catalog our data assets/products and contracts to access these data assets/products – internal, external, and third party. Next steps we have to enable active metadata to ensure the catalog is always current. Data contracts - especially SLAs and SLOs - are really crucial to driving reliable and scalable data practices forward. How can people trust what they are consuming without having to check it themselves unless there are very specific parameters and documentation of what they're getting? The data space needs to rework the way we approach data contracts. We need to be careful to not head down the same paths/ways of working - just with different names - that we've tried and didn't work. But we also need to focus on what we've learned from different approaches instead of reinventing the wheel where appropriate. Hopefully data mesh can thread that needle. When thinking about how you should split into domains, look at the business strategy. How does your organization tackle business challenges? That should inform how you create domain boundaries. One of the biggest challenges in data at the moment is centralization versus decentralization and/or federated. How far to go towards one or the other side across many many decisions is really crucial to your data strategy. Look for places to centralize support of multiple use cases but not take the decisions out of the hands of people who...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/15QaPbgVvALLawq19qe62-CRxVFbawjo1BGOrvt9rX3g/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Gretchen Moran, the Senior Director, Data Products at the National Geographic Society (NGS; the non-profit arm of National Geographic). Some key takeaways/thoughts from Gretchen's point of view: NGS is a bit unique in that they don't have a widely deployed data architecture so they do not have a lot of habits to unlearn. Starting with a greenfield means likely more training and learning/experimenting will be required but at least no institutional unlearning. To move forward with data mesh, organizations must be able to embrace change - and the pain that it will inevitably bring - and embrace ambiguity. You need to move forward and figure it out together but also be okay with failure as a learning experience as you test what works for your organization. To win the hearts and minds of data producers, show them what high-quality data can mean for the organization and their domain/role. Work closely with them, understand their context, hold their hand to bring them along and align them to the vision of data mesh. It's easier to drive buy-in widely if you find the organizational influencers and win them over. It is the domino effect in practice. Partner closely with the influencers early on to drive your initiative forward. For NGS, they are working with a single initial data producing team for their proof of value. The data mesh world seems to be split a bit between working with one or two to three teams in the initial proof of value stage. "Any technology effort is still a people effort." We have yet to learn how to leverage the knowledge and context of people without data knowledge in general in the data and analytics space. This is what data mesh tries to unlock but we are still figuring out how to do it well. It's very easy to intimidate people with data. We need to make tech and especially data much less intimidating to push broader adoption. The business context of those who aren't yet data literate can be extremely valuable. We need to lower the actual bar to leveraging data but also lower the perceived bar to leveraging data. "Metrics + outcomes = value" - without outcomes attached, metrics have no value. Automation is going to be key to many aspects of data mesh. Upskilling people to leverage data will only really pay off if it doesn't mean a large increase in the amount of work to leverage data. User experience is crucial to getting the most value out of your data. Think about your data user experience (DUX) and bring in designers to help optimize the experience and really focus on data as a product thinking. NGS is still trying to find who should own generating and sharing insights on data combined from multiple domains. Is that a centralized insights team? Does that push us too far back towards centralization? It's still early days but those insights are crucial to driving value from data. We will see where new insights come from in data mesh. Will it be more insights from data consumers as they...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1TtT53TI4j-lwsbCnNHWhtX2rTJudxSyLpm8ylId6fM4/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Liz Henderson AKA The Data Queen, Executive Advisor at Capgemini. To be clear, Liz was only representing her own views on the podcast. Some high-level takeaways/thoughts from Liz's view: To drive buy-in and engagement in a data strategy - especially with people outside the IT/data team - focus on the "why". Why are you doing this initiative or approach? What business goals is it supporting? Also on driving buy-in for a data initiative, start by listening instead of selling/pitching. Focus on the business needs and work backwards to show how data can help address those needs. To be successful with a large change-management data initiative, you need the patience, leadership, courage, and will to push forward. And the budget - don't forget the budget :D You can't have an effective data strategy if it isn't directly tied into the business strategy. Really consider how data can help to support and execute on the business strategy. Data strategy in a vacuum away from the business is a recipe for trouble. It's very easy to get overly focused in data on what you are delivering instead of why you are delivering it and who it is supposed to serve. If you want to be successful, you need to focus on the latter two. And look to deliver continuous incremental value rather than a back-end loaded value delivery. Change management is very easy to get wrong in data. Really consider if you can not only get the ball rolling, but keep it rolling and in the right direction to implement a large change. Loss of momentum can mean loss of funding. If data mesh follows a similar pattern to data literacy, it's likely to be 3-4 years from initial large swell in hype around data mesh - whether that was late 2021 or more now - until we really see a clear picture of how more organizations have implemented. There needs to be a time for trial and error. Data literacy can only get you so far - you need data storytelling and visualization as the business people need to be able to understand what the data is saying to drive decisions. The 3 most common ways organizations go wrong with data are 1) technology-first / expecting to buy your way to a solution to challenges; 2) not asking the "why" questions; and 3) not having a data strategy.
Liz started the conversation talking about how data is an asset to the company - not treating data as an asset, an important differentiation - and started with a common theme for the conversation: "why?". Why is data important to the business? Why are we doing the data work we are doing? What is the purpose?
When speaking with company executives, especially those outside the IT/data team, Liz starts every conversation by digging into what is the business trying to do. What is the business strategy? How does data currently play into that strategy? What role do they want data to play in the business strategy? What role might data play if everything were perfect? You can't have an effective data strategy...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1VenyJNXvT--N19GNgVcXKGDyiFhhmO4_RI4cFh1U1mw/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Karolina Henzel, Data Enablement Tribe Lead at T-Mobile Polska. FYI, businesses and domains are fundamentally similar in this conversation and are used essentially interchangeably. Some high-level takeaways/thoughts/summarizations from Karolina's view: Business transformation and impact is what really matters. Digital transformation is just a mechanism to transform your business into being more digital native and focused. Data transformation is just a part of digital transformation. Transformation should all come down to driving positive business impact. To drive something like data mesh forward, you really need top management support, likely a C-level executive sponsor. Otherwise, it is very easy for work to get deprioritized and pushed out. Don't take on a large-scale data initiative unless there are specific business challenges to address. Don't do data mesh for the sake of "being data driven"; what are the issues and why will addressing them help your business? Explicitly define the problems and the pain points. To drive change, look for "change agents" in the domains. They are people with the will and capabilities to drive large-scale change. They aren't always easy to find but once you start to identify them, patterns will emerge. The big pain points T-Mobile Polska was facing were: 1) poor/inconsistent data quality; 2) data discovery difficulties; and 3) slow time-to-market for new data and insights. T-Mobile Polska was able to move forward with data mesh because business representatives in the domains were bought in that addressing the data pain points would drive incremental business value - there would be a return on the data work investment. Look for quick wins and how to deliver continuous incremental value. If it is all about producing a big bang, you will very likely lose momentum, prioritization, and funding. Continuous value delivery is crucial to keeping people excited about data work. T-Mobile Polska's data quality issues were caused mostly by a lack of accountability/ownership and not adhering to standard definitions across domains and reports. Lack of standard definitions - or at least very clearly differentiated definitions - can cause numbers to not match across reports. And that makes people not trust the data. So drive domains to clearly define terms and, if possible, look to create standard definitions across the organization. Find KPIs that are focused on what actually impacts the business. Dave Colls mentioned fitness functions to break down measuring progress against big challenges into smaller measurements. Karolina and team are making sure the end result is business impact, not technical-only change. You can drive buy-in that data producers should provide high quality data products by showing producers the impact their efforts are having. Can be a chicken and egg issue at first but you can take the results from one domain or business and show them to another to drive buy-in....
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1ls5QawrOffb3VGIfZCmPYHG0v729Ye7Gu7oKRaOYi8c/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Mahmoud Yassin, Lead Data Architect at ABN AMRO, a large bank headquartered in the Netherlands. Some high-level takeaways/thoughts from Mahmoud's view: It's very difficult to do fully decentralized MDM, which led to some duplication of effort - that can mean increased cost and people not using the best data. ABN tackled this through their Data Integration Access Layer - similar to a service bus. They are using that centralized layer - called DIAL - to help teams manage integrations that are both consistently running and on-the-fly. It helps monitor for duplication of work instead of reuse. If Mahmoud could do it again, he'd focus on enabling easy data integration earlier in their journey to encourage more data consumption. Cross domain and cross data product consumption is highly valuable. The industry needs to develop more and better standards to enable easy data integration. Data mesh and similar decentralized data approaches cannot fully decentralize everything. Look for places to centralize offerings in a platform or platform-like approach that can be leveraged by decentralized teams. Most current data technology licensing models aren't well designed for or suited to doing decentralized data - it's easy to pay a lot if you aren't careful - or even if you are careful! A tough but necessary mentality shift is not thinking about being "done" once data is delivered. That's data projects, not data as a product. Try to keep as much work as possible within the domain boundary when doing data work. Of course, cross-domain communication is key but try to limit the actual work dependencies on other domains if possible. A data marketplace enables organizations to more easily create a standardized experience across data products and make data discovery much easier. You don't necessarily have to tie your cost allocation models to the marketplace concept. Sharing what analytical queries/data integration "recipes" people are using has been important for ABN. It drives insights across boundaries and also creates a lower bar to interesting tangential insight creation/development. You should consider not allowing integrations across multiple data products by default. Producers should be able to stop integrations - for compliance purposes or because the integration doesn't actually provide good/valuable/correct insights. Traditional ETL development is about translating the business needs to code. But centralized IT usually can't deeply understand the business context and needs so they deliver substandard solutions. If you consider that business needs evolve, it gets even worse.
Mahmoud started his career in data as an ETL developer so he saw the ever increasing issues with the traditional enterprise data warehouse approach in large organizations. Then he moved on to working with the common way people have approached data lakes - managed by a centralized team - and the issues seemed pretty similar to with data warehouses to him....
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info: email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1P-Xjxgz7GafDgkCnBF67Tj5LuQXcjvQerU5FLxyyxNM/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Stéphanie Bergamo and Simon Maurin of Leboncoin. Stéphanie is a Lead Data Engineer and Simon is a Lead Architect at Leboncoin. From here on, S&S will refer to Stéphanie and Simon. Some key takeaways/thoughts from Stéphanie and Simon's point of view: "Bet on curious people", "just have people talk to each other", and "lower the cognitive costs of using the tooling" - if you can do that, you'll raise your chance of success with your data mesh implementation. Leboncoin requires teams to share information on the enterprise service bus that might not be directly useful to the originating domain on the operational plane. They are using a similar approach with data for data mesh - sharing information that might not be useful directly to the originating domain by default. Leboncoin presses teams to get data requests to other teams early so they can prioritize it. There isn't an expectation of producing new data very quickly after a new request, which is probably a healthy approach to data work/collaboration. Embedding a data engineer into a domain doesn't make everything easy, it's not magic. Software engineers will still need a lot of training and help to really understand data engineering practices. Tooling and frameworks can only go so far. Be prepared for friction. Similarly, getting data engineers to realize that data engineering is just software engineering but for data - and to actually treat it as such - might be even harder. Software engineers generally don't know how to write good tests relative to data. Neither do data engineers. But testing is possibly more important in data than in software. We all need to get better at data testing. Start with building the self-service platform to solve the challenges of the data producers first. You may make it very easy to discover and consume data but if the producers aren't producing any data... If your software engineers are doing data pipelines at all before starting to work with them in a data mesh implementation, you can probably expect they aren't using best practices. It's pretty common for good/best practices to be known by only a few people inside an organization, such as with a specialty-focused guild. Look for ways to cross-pollinate information so more people are at least aware of best practices if not able to fully implement them yet. Trying to force people to share data in a data mesh fashion didn't work for Leboncoin and probably won't in most organizations. Find curious developers and help them accomplish something with data, that will drive buy-in. As part of #10, data products often start as something serving the producing domain and then evolve to serve additional use cases. They start by serving a specific business need and evolve from there. Look to build your tooling to enforce your data governance requirements/needs. Trying to put too much on the plate of software engineers probably won't go well.
Around the time Zhamak's first post...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB); George Trujillo's contact info; email (george.trujillo@datastax.com) and https://www.linkedin.com/in/georgetrujillo/ (LinkedIn) Transcript for this episode (https://docs.google.com/document/d/1kLgkxRsiFr1G5BDbFcZaVMJhb3A_UcGc0N5n6vo7fCk/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Erik Herou, Lead Engineer of the Data Platform at H&M. To be clear, Erik was only representing his own views and perspectives. A few key thoughts/takeaways from Eric's point of view: Data mesh can work well with a product-centric organization strategy as both look to put ownership and product thinking in the hands of the domains. To develop a good data/enablement platform for data mesh, look to work with a number of different types of teams. That way, you can see the persistent/reusable patterns and capabilities to find ways to reduce friction for future data product development/deployment. H&M had an existing cloud data lake that was/is working relatively well for existing use cases. But the team knew it likely wouldn't be able to handle where they wanted to go with many more teams producing data products of much higher quality and potentially sophistication. When implementing data mesh - or any data initiative really - it is easy to fall into the trap of doing things the same way you did before. The "old way" feels safe and it was/is still working relatively well for H&M. So they treated their data mesh implementation as almost a greenfield deploy. Because of the long-term focus on making it low friction and scalable to share data - the consumers will come as you make them more data literate - most of the early data/enablement platform work has been focused on helping data producers. A common pattern in data mesh but your constraints and needs may not match. Erik's team is focused on enabling data producers first specifically so his team doesn't become a bottleneck. It is easy for a platform team doing any part of the individual work to become that bottleneck. Consider how much organizational change you require before starting to create mesh data products. H&M did a large amount of that organizational change, other companies start in their current structure and evolve as they learn more. Both are valid and can work well. Specific to H&M, a strong track record of good return on investment in AI meant there was less pushback than in many organizations when they started driving buy-in for implementing data mesh. In the historical data warehouse world, there was less need for data literacy because most people were pushed reports but also couldn't do much, thus not "getting themselves in trouble". If we move to a more self-serve approach, that means we need much better data literacy - it can be a big risk to allow access without understanding. Otherwise, it could be like turning a six year old loose in a fully stocked kitchen where they intend to "make dinner". Data catalogs could really help push forward general data practices but we still need to have actual conversations too. Being able to ask someone about what data means and similar high context exchanges are crucial. "If you have a complicated business, you have complicated data." If your mesh data products don't maintain loose...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1ls5QawrOffb3VGIfZCmPYHG0v729Ye7Gu7oKRaOYi8c/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Andrew Padilla, who runs a data and software consulting company - Datacequia - and serves as editor of the Data Mesh Learning community newsletter.
This one is a bit more philosophical about sharing information/knowledge so it's one to sit and think over. Things in quotes are direct from Andrew.
Some key takeaways/thoughts that come from Andrew's view of data mesh and the data space in general: To move from sharing the 1s and 0s of data to actually sharing knowledge, we need to harmonize data, metadata, and code - "the digital embodiment of knowledge". That's where Andrew hopes the mesh data products can head. Software development isn't cutting it for sharing knowledge. Will data product development? Do we need to move to knowledge-centered development instead? Remains to be seen. We still don't know how to model well - in data - what is going on in the real world. What are the experiences of the organization? Can we really define an "organizational experience"? Event storming tries but seems to fall short often. We must learn to treat organizations like living entities. Organizational experiences cross multiple domains and the types of experiences will change, will evolve - possibly quite quickly. We again have to get better at modeling those and evolving how we share knowledge about the experiences. Knowledge graphs are the best way we have currently for combining information across domains. We still haven't fully figured out how to leverage our cross domain knowledge though. Historically, we've bent our ways of working to the limitations of the machines. We need to spend more time on bending the machines to better match the way humans store, process, and share knowledge. Data centricity is an interesting concept but might take our current imbalance of data versus operational focus too far towards data. But that might be what is necessary to really get to balance. It remains to be seen. But it's crucial to understand a data-first focus isn't necessarily a knowledge-first or knowledge as a first class citizen approach. It's important to understand that mesh data products are a means to an end in data mesh. Yes, they are crucial to sharing information but they are there to serve a purpose, not that they are the purpose. In data mesh, it can be easy to focus too much on creating data products of immediate utility or that are high value in and of themselves. But it's important to think about how data products together create value - and maybe not immediate value - to really drive forward our understanding of the organization's knowledge and experiences.
Andrew started the conversation with his hope and vision for data products - or the data quantum - in data mesh. Historically, data, metadata, and code are not often grouped together and even less frequently are they in harmony. They belong together as that harmony creates a higher level abstraction to share knowledge, not just the 1s and 0s of data. To get data mesh right, Andrew believes you have to really figure out how to build mesh data products with that harmonization in mind....
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1P3b8399IKJvBJhNxp0Tep_DtG5hu3rtLY1fFMoMHEeU/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Jen Tedrow, a Product Management Consultant at Pathfinder Product Labs who is currently working with a large client on a data mesh implementation. She was only representing her own perspective in this episode. Some key takeaways/thoughts from Jen's point of view from the conversation: A data mesh vendor assessment is likely to be different than almost any other vendor assessment you've done before, especially if you aren't evolving something existing. There is so much more to cover and the overall platform needs to meet your needs, integrate with how you handle data on the application side including integrations, comply with your governance standards, fit within budget, etc. That's a lot of needles to thread. Spend considerably more time doing the discovery process in your data mesh vendor assessment than you would for a normal vendor assessment. There are a lot of potentially hidden needs / wants and it is far better to surface them early. By digging deep into stakeholders' desired outcomes, you can understand what you need to deliver but also, you can get insight into driving buy-in. Address the challenges preventing them from the desired outcomes and they will feel seen and heard. As many guests have said, lead with empathy. Change is painful. But if you are realistic with people and make them feel seen and heard, it will be much less painful. When speaking with potential users, again really spend the time to make them feel seen and heard - reflect back to them what you heard. And have them share what is their ideal state. You may not be able to fully deliver on it but it's important to understand where they want to go. As you are learning new information, share that in a continuous stream with stakeholders so they understand the recommendations you are making along the way and at the "end" of the assessment - it doesn't really end when you finish the assessment, so "end". Be prepared for there to be capability gaps - possibly significant - between what you want now and what is available in the market or that you are able to build in your budget. There are just a number of capabilities that aren't really part of any vendor offering at the moment. Waiting until everything is perfect will mean you are waiting for a while! It's crucial to focus on what do you value most right now and also in the future when making vendor assessments. There are so many nice-to-haves with a data mesh implementation but you'll need to compromise. Jen and team developed a good framework for evaluating offerings as to whether they fit actual needs or just wants. The three most important aspects to meet right now for Jen and team through the self-serve platform were user experience/low barrier to usage, automation, and ability to easily integrate the tools together and with the existing stack. When figuring out what capabilities you need from your platform, create high-level task-based use cases, not systems requirements. This will prevent you from steering too much towards trying to serve any one specific use or getting bogged down in tech compared to...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Community open discussion meetup hosted by Eric Broda: https://www.youtube.com/watch?v=OwtQ37WYK1g (https://www.youtube.com/watch?v=OwtQ37WYK1g) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1JyEGeFHqLz3_WdGviNZeAYMJZQzI8ttS9r8I5ffqqwo/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Omar Khawaja, Head of Business Intelligence at Roche Diagnostics. To be clear, Omar was only representing his own viewpoints and learnings, not necessarily those of Roche.
Some interesting thoughts/takeaways from Omar's point of view and learnings: If you are going to make progress in a data mesh journey, you must be okay with "good enough". Perfect is the enemy of good and done. Measure, learn, and adjust along the way but get moving and keep moving. It's okay to make mistakes - recognize and correct them. Echoing a number of past guests, change management and organizational challenges will take a large portion of a data mesh implementation leader's time and effort - likely far more than most would expect. Focus on empowering people and showing them why this can work for them. And what it means for them. Data mesh cannot be your entire data strategy. If you are implementing data mesh, it must only be part of your data strategy. Start from the why. Why undertake something as transformational and difficult as implementing data mesh? What business value will it deliver? Data-as-a-product thinking is the true heart of a data mesh implementation. It's far more than just creating data products. Data product discovery is crucial, much like discovery in regular product management. Take considerable learnings from product management in other disciplines. Focus on outcomes in day-to-day data work. What are you trying to deliver? What is the value in it? For whom? How will we measure if we are successful? And were we actually successful? We need to get data people to rethink creating point solutions - sometimes called project management thinking - where they deliver a dashboard and the dashboard itself is the focus. This leads to fragility that could be prevented by focusing on the entire data lifecycle to create the dashboard with the dashboard - and many other chances for data reuse - as an output. Roche is being quite flexible around who develops data products - it is all about the capabilities and needs. Often, it is data engineers in the domains, enabled by the central platform team. But it can be data/business analysts or software engineers too. If the data product isn't overly complex or if a business analyst really understands data, why can't they be the data product developer? It would have been the definition of insanity - trying the same thing over and over and expecting different results - for Roche to just move from an on-prem data lake that was having scaling and quality issues to a cloud data lake. Many other aspects needed to change. The organization needed to unlearn and relearn a number of things and data mesh was a great vision for where they could go. Roche saw some duplication of work across data products so they adjusted and made their data product discovery and design phases very public. Making it public can increase collaboration early in a data product's life as well so you might find additional data consumers in the development phase.
Omar started the conversation with a definition of what Business...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Community open discussion meetup hosted by Eric Broda: https://www.youtube.com/watch?v=OwtQ37WYK1g (https://www.youtube.com/watch?v=OwtQ37WYK1g) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1r2GDDj3IQ0L4UO3iGv-5sLPCU0rKErjrbUzN7vMTzTY/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Dave Colls, Director of Data and AI at Thoughtworks Australia. Scott invited Dave on due to a few pieces of content including a webinar on fitness functions with Zhamak in 2021. There aren't any actual bears, as guests or referenced, in the episode :) To start, some key takeaways/thoughts and remaining questions: Fitness functions are a very useful tool to assess questions of progress/success at a granular and easy-to-answer level. Those answers can then be summed up into a greater big picture. You should start with fitness functions early in your data mesh journey so you can also measure your progress along the way. To develop your fitness functions, ask "what does good look like?" Focus your fitness functions on measuring things that you will act on or are important to measuring success. Something like amount of data processed is probably a vanity metric - drive towards value-based measurements instead. Your fitness functions may lose relevance and that is okay. You should be measuring how well you are doing overall, not locking on to measuring the same thing every X time period. What helps you assess your success? Again, measure things you will act on, otherwise it's just a metric. Dave believes the reason to create - or genesis of - a mesh data product should be a specific use case. The data product can evolve to serve multiple consumers but to start, you should not create data products unless you know how it will (likely?) be consumed and have at least one consumer. Team Topologies can be an effective approach to implementing data mesh. Using the TT approach, the enablement team should focus on simultaneously 1) speeding the time to value of the specific stream-aligned teams they are collaborating with and 2) look for reusable patterns and implementation details to add to the platform to make future data product creation and management easier. We still don't have a great approach to evolving our data products to keep our analytical plane in sync with "the changing reality" of the actual domain on the operating plane. On the one hand, we want to maintain a picture of reality. On the other, data product evolution can cause issues for data consumers. So we must balance reflecting a fast-changing reality with data consumer disruption, including downstream and cross data product interoperability. There aren't great patterns for how to do that yet. There is a tradeoff to consider regarding mesh data product size. Dave recommends you resist the pull of historical data ways - and woes - of trying to tackle too much at once. The smaller the data product, the less scope it has, which makes it easier to maintain and the quicker to deploy and feedback cycle. But smaller-scope data products will increase the number of total data products, likely leading to harder data discovery. And do we have data product owners with many data products in their portfolios? Dave recommends using the Agile Triangle, framework to figure out a good data product scope (link at the end).
Dave mentioned he first started discussing fitness functions regarding...
Sign up here for a Data Mesh Therapy session here: https://calendly.com/data-as-a-product/data-mesh-therapy (https://calendly.com/data-as-a-product/data-mesh-therapy) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1hNNILAXWGAZ8xEyjUxe7oiek88R0C_PcFIKo3EBgRfQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Jay Sen, Data Platforms & Domain Expert/Builder and OSS Committer. While Jay currently works at PayPal, he was only representing his own view points. Some key takeaways/thoughts from Jay's view: When you get to a certain scale, any central team should focus on, as Jay said, "Empower people, don't try do their jobs." That's how you build towards scale and maintain flexibility - your centralized team likely won't become a bottleneck if they aren't making decisions on behalf of other teams. To actually empower other teams, dig into the actual business need and work backwards to a solution that can solve that. If there is a solution already in place that isn't working any more, look to find ways to augment that rather than trying to replace or reinvent the wheel. Self-service is a slippery slope - it often solves the immediate problem of time to market but also creates next level challenges. A big issue is that when you remove the friction to data access, you are throwing challenge of finding right data on consumers plate. Data contracts are great when everybody aligns on a single contract and there are enough tools to support the contracts. But they also create a proliferation of data to enforce the contracts required by multiple consumers - thus, they often don't survive the real world. The data catalog space is finally getting some needed attention. But there are still a myriad of issues that need solving. Will those be solved by technology or by leveraging a "data concierge" remains to be seen. It's insanely easy to overspend in the cloud. Everyone is vaguely aware but cost should be part of every important architectural discussion. You can drive business value but it absolutely must also be focused on the cost as return on investment is far more important than simply return.
Jay took a few lessons from working on a central services team in a company of ~200 people. Having a centralized team was doable at first but as the org scaled, it quickly got complicated. As a centralized team, it's very easy to become a bottleneck but Jay learned a lesson that has continued to help in his subsequent roles: "empower people, don't do their jobs." Focus on reducing the friction to others doing their work instead of doing it for them.
Easier said than done so how do you empower people? Per Jay, you must understand the business aspect and what the requestor actually needs. That isn't really going to get communicated well in a ticket most times so you should have a high context information exchange to take what they need and convert it into a workable solution. And often, there is already a solution in place but it's just not handling the job anymore. So you want to consider if you should solve the same issues in a better way. It's much easier to do a greenfield deploy but brownfield is an inevitable facet of enterprise data work.
Per Jay, a few good things to remember: 1) frameworks and technology come and go but the concepts are the things that stick around. Focus on solving issues by leveraging technology and frameworks, not relying
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1-xh94Bry34KhWCgFE-b2GPzWUKGRkqh9fVpKoAoPl_Y/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Jay Como, Head of Finance Data, and Elizabeth (Liz) Calloway, Director of Finance Data Products at Silicon Valley Bank. To be clear, they were only representing their own views and experiences. Some key takeaways/thoughts from the conversation: The governance team should "wear them down with empathy." Take the time to share your context, learn their context, make them feel seen and heard. That will get them to see you as a partner and good governance is truly about partnering, not mandating or being a gate/hurdle to get past. Great governance is the pathway to great data. Great data leads to great decisions which lead to great outcomes. Share that path to great outcomes so people can see a clear answer to "why are we doing this?" Governance isn't just risk mitigation, it can be a significant - if almost always hidden/secret - value driver. To drive governance buy-in from data producers, again, lead with empathy. Let them in on the "why" - why does this matter? What is the business value? How can this benefit them? "Help me help you" is a good approach to talking to internal teams about data governance. You are there to drive value for them, take work off their plates when appropriate. You can further drive buy-in through helping teams get to quick wins. While the long-term is obviously important, incremental value-add is better than a big bang approach. Provide a constant stream of value, including executing on where you are helping, will drive teams to want to work with the central governance function. Drive value and the buy-in naturally comes with it. Good data governance is necessary to avoid fees, fines, and the huge revenue/business impact bad data can have. But don't use those as a boogeyman, don't use fear to sell good data governance. It's easy for a centralized governance team to become a bottleneck. Focus on not solving all the problems for teams but being there to help when they need it. You are the backstop, not the stop sign. Be the support, not the roadblock. To prevent governance in general from being a bottleneck, you must have flexibility and pragmatism. If exceptions to requirements are necessary, then those exceptions are often a valid response to other constraints/pressures. But be very explicit about the reason and type of exceptions and also very explicit about the expectations for if/when those exceptions will be remediated. Good governance is about incremental improvements, not trying to do everything as a big bang. Set expectations, move forward. Fix for today and prevent the same issue in the future.
Jay started the discussion talking about the concept of governance-as-a-service - in other words, providing a service to internal stakeholders instead of a mandate of comply or else. And while governance may not be the most "sexy" aspect of data, it certainly is one of the most crucial. For Jay, without good governance, data is often not nearly as correct and clean as it could be so internal stakeholders aren't making as good of decisions as they could. He laid out a simple framework of great data leads to
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1C8KjRoqwiwoZaSQ9LcuZ4Xew_TCx1A69VnEOIeZkLp0/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Björn Smedman, Engineering Manager at Communication Platform-as-a-Service (CPaaS) company Sinch.
Some interesting thoughts or takeaways: A good indicator for when decentralizing your data team might make sense is the cognitive load of a centralized data team. How many systems - including a measure of how complex - are they managing? How much of their time is spent in meetings, especially trying to understand context/requests? Is there starting to be combative prioritization from multiple domains? It can be very beneficial and scalable to apply data mesh principles to non analytical use cases, especially sharing data for application purposes. It is still often difficult to prioritize creating a data product for machine learning without knowing the business value of the ML model. But the ML team needs the data first before they can figure out the business value of the ML model. You have to make speculative bets. If you see the data platform team start to dig into the semantics of a use case, that's a red flag that people are trying to leverage them as a data team. And while you want a centralized data platform team, you probably don't want them to become a centralized data team.
Since December 2020, Sinch raised nearly $2 billion USD. With this funding, they have made a number of sizeable acquisitions, with the company growing from 500 employees to over 3,000 in about a year. This has led to some interesting challenges in sharing data in a hyper-scaling environment.
Per Björn, data is a very key part of Sinch's plans for growth. Sinch's operational systems are often very transactional, as some product lines can process tens of thousands of monetary transactions a second, so data that might be typically shared on the operational plane in other companies is shared on the data plane lest the operational data stores deal with billions of events, making the data challenges even more complex than for most organizations. Then add in the regulatory requirements of telecom.
Björn helped lead the move to decentralizing the data team. When Björn joined, the central data team organization was 4 teams and 25 people. The data function was previously centralized and that was becoming a bottleneck, even for the legacy business. Now that the company had acquired a number of other sizeable companies, that central data team setup clearly wasn't going to scale. The company reorganized around business units and started to build data and analytics teams inside each BU.
For Björn, who started in December 2021 just as Sinch started acquiring new businesses, the central data team was clearly not going to be able to meet the needs of this new organization that was about 8x larger than a year earlier. There was too much cognitive load on the team, especially trying to understand the product lines of five distinct business units, many of which were entirely new to the company.
Björn gave a few good indicators of what to look for when considering if you should decentralize your data team. A big one is team cognitive load. Cognitive...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1G80qJuECjV6heh0W5aOVIej1G9UUug1urb3nOLrJ_Mg/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Luca Paganelli, Data Architect at the Italian utility Gruppo HERA. To start, some interesting points and/or key takeaways and questions: Introducing new concepts and ways of working around data slowly - not looking to make a hard shift - has worked well. When Gruppo HERA debuted their new data strategy manifesto, none of it was a surprise and it was already relatively in-line with the way many were talking about data and moving forward on their data journeys. HERA's Data, Analytics, and Intelligence Automation (DAIA) team is not forcing domains to comply with HERA's data mesh-inspired guidelines but instead working with them closely to help the domains achieve their data related goals - delivering the "right thing". That gives the DAIA team strong influence to direct the domains' approach to data work without pushback and gives domains better confidence in the guidelines and can mitigate analysis-paralysis risk. This lack of rigidity and strong rules created a better sociotechnical environment to innovate but it can mean nothing really feels standardized because the domains can still choose to go a different direction. The paradigm-shift was initially "steep" for both IT and domain owners. But domain owners realized how much better they could serve themselves and external data consumers if they took over more data ownership. IT was afraid to give up control but started to buy in to the leverage and expertise they can provide by empowering the business domains to do great things with their data. A concern with not having broad standardization is bespoke solutions so it is hard to create broad reuse. There is also a challenge of people not being sure how much they can trust the data products. The DAIA team believes the tradeoff is worth it to drive initial buy-in with domain owners. Defining data products has been a struggle. There is a chicken and egg issue of 1) needing to understand who from the business should be involved in designing a data product but 2) data domains must be discovered to know who are the subject matter experts from the business to involve. For HERA, they are looking for data products to first serve their domain owners. This can be a slippery slope as domains may have valuable information but that isn't useful for them to analyze for their own purposes. So then other domains can't get to that data. But getting domains to freely share their data is a common incentivization problem. Purely technical focused data products will probably not serve demand. We need to focus on sharing information - what is the data saying, what is it about? Information is more than just the 1s and 0s of data.
Gruppo HERA had to develop a sophisticated and reliable way to do their reporting to regulators but had not focused nearly as much on their internal data and analytics. But a progressively larger number of experimentations (spanning BI to AI) emerged where data started being used to drive the company. About 2.5 years ago, they developed a new team around data, analytics, and intelligence automation (DAIA) to start to rectify that...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1CrNt8qo72qGtU1dOdz4PMDVMfqeIZOAP9v5dFzpSKcI/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Joe Reis, CEO/Co-Founder of data consultancy Ternary Data, Co-Host of the Monday Morning Data Chat, and author of the upcoming book Fundamentals of Data Engineering. Some key points or takeaways specifically from Joe's point of view (not necessarily those of the podcast): Find quick, high-value wins. Too often people focus on the big wins and those become overly complicated and end up in failure. Most software engineers don't understand data well enough to be data product developers in data mesh, at least yet. Data mesh is a polarizing topic. And that makes sense as it is pushing boundaries. Many hope it can come to fruition but it is a bit of a utopian view. The future of data engineering is to move past managing pipelines to much higher-value work. Speed to achieving wins with data - with a clear return on investment and trust - is the first thing you should focus on. Get this right and you can have the "luxury" of building great data products.
Joe started by discussing the kind of nebulous area within software engineering and data that data engineering has always played - sit between the source systems and the data output, converting the data in the source systems into something consumable for data users. Previously, that was mostly about making sure reports got pushed through and you hoped people derived insights. Now it's more about pipelines. But the way we store information in source systems, it is not in the format or shape we need for analytical purposes. So there needs to be a go-between.
A big trend in data engineering currently for Joe is the abstraction of tooling. Some of that can be good - makes people more productive - or bad - means it's harder to understand what is actually happening under the covers. But for Joe, it's probably worth it to use the abstractions as they are able to do the heavy lifting and data engineers can focus on the higher value work. We might be coming to the end of the "pipeline monkey" era of data engineering so we can shift more focus to the data output, DataOps, orchestration, security, etc.
For Joe, the biggest value-add the data engineering team can have is getting wins quickly. When asked about speed to returns versus repeatability, Joe said that the speed is more important, especially when you are trying to prove out the value of your data team. Trust is crucial, so you have to be careful to not move too fast, but trying to do big-bang projects is often a recipe for failure in his view.
When asked what could be the signs an organization is ready to implement data mesh, Joe mentioned that if an organization is already seeing "wins" with data across a number of teams/domains, that's a very good sign. But you can't only have a few teams getting those wins as that means the overall organization data maturity is still probably low.
Joe made a good point about how polarizing data mesh can be. When he speaks with some organizations, there are a few leaders who simply reject the idea outright. But many also simply don't see data mesh as ever being possible specifically in...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1J5LuzBWwHZIp1cXPKctyb88iXCQDPAhVDlwiMGJ3wfM/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Jesse Anderson, Managing Director at consulting company Big Data Institute, host of the Data Dream Team podcast, and author of 3 books, most recently Data Teams. To start, a few takeaways from Jesse's perspective on the choosing technology side: You should make sure you have the right team in place to make good technology decisions - the team needs to be in place first Before selecting any technology, it's crucial to understand what you are trying to accomplish. And to understand that the technology will provide help in addressing the challenge but won't solve anything itself Focus on: is this the right tool or solution for us now and in the future? What is the roadmap and vibrancy of the solution? "Technology must earn its keep", meaning you should understand the total cost of ownership and what is your expected return on investment Data tooling cycles are probably going to be 10 years at the most - prepare for obsolescence so you aren't overly reliant on any one technology
And some takeaways from Jesse's point of view on decentralizing data teams: Currently, software engineers aren't ready to be data product developers so you'd need embedded data engineers to handle creating and maintaining data products in data mesh But many data engineers are not willing to be embedded into domains Managing the dotted line versus solid line of reporting between a functional team and the domain is very difficult There are a number of cracks where crucial data can fall into and fail to find a good owner in a decentralized structure, especially aggregate data products
Jesse started the conversation on how important people are to getting things right with data, especially making technology decisions. The chicken and egg question is do you need to have the right people in place first or do you want to make technology decisions that will attract people. In Jesse's view, you need the right people in place first as they will be the ones to make the right decisions on technology selection. The most important question for Jesse when selecting technology is what are you trying to accomplish with technology. If you don't focus on the target outcome, that is not going to work out well. And you should know, in general, what most of your use cases will be for the technology - use that to assess what is the right technology to choose.
Also, for Jesse, "technology must earn its keep". Just because you made a decision on using that technology at one point, it must continue to be of more value than its cost. And you want to strongly factor in your long-term total costs, as best as you can estimate then, when looking at adding a technology. This is important for build versus buy, can you continue to keep something running, is the long-term roadmap a match to your goals and vision, etc.
Jesse also pointed to how different data is to the operational side relative to technology cycles. Considering Hadoop, where Jesse focused in his time at Cloudera, 10 years - or even less - is realistic for how long data technologies might be around. Thinking in those
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1WRNYqsgfM-S4jt4ExOxnW8kuhvcaX1sh2CIGJFzvlHE/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Immanuel Schweizer, the Data Officer for EMD Electronics. Some interesting thoughts and questions from the conversation: Good governance starts at data collection - what are ethical and compliant ways to collect data from the beginning? This points to intentionality around data use stretching into the application - what should you collect that might not be part of the day-to-day application function but that might might lead to generating insights that will be used to generate a better user experience? And what are the ethical concerns? Should we initially create data products to serve specific use cases or should we focus on sharing data first and then shaping what people consume most into data products? EMD is approaching data products from a different angle than most, using the second approach. When looking at data mesh, should you start with the high data maturity teams or work to pull everyone up to at least a decent baseline maturity level? If you work with the most mature teams, will their challenges really be applicable to the not-so-mature domains? Can you find good reuse patterns to scale your mesh implementation? Domain owners are much more willing to share data if they understand use cases for how their data will be used and maintain control to prevent misuse. Reluctance comes from an incomplete picture causing concerns - the more visibility into how data can be and is being used, the more willing domain owners are to share. But understanding your end-to-end data supply chain is tough, especially to start. How do you evaluate when to spend the time with a domain to get them data mesh ready? If you need a high value use case to justify spending time with that domain, are you leaving many domains behind? This ties to #2 and #3. Set your target picture but be ready to adjust your target picture along the way. The world is ever changing, don't lock in to an expected target outcome. Good data governance is about speeding up 1) access to and 2) usage of data. EMD launched a data literacy program where the employees spend the majority of a 10 week timeframe learning about data and how to make use of it. For Immanuel, making things tangible relative to data makes people much less hesitant to explore and use data. You should make using data a "part of the job" so it is tracked and part of the review process. Otherwise, you are missing out on a key incentive to leverage data. How many people in your organization wish they could be leveraging data more often to make decisions? What's holding them back? Is it tooling, knowledge, incentivization, access, etc.? How can we democratize insights? So much of insight generation is one-off, how do we make that scalable, shareable, and repeatable?
Per Immanuel, EMD's data mesh journey is not that typical in that they are still getting their arms around centralizing data in a constructive way. It was previously locked away in the domains. So, they are starting their data mesh - or decentralization - journey by centralizing data in a certain sense. Wannes Rosiers mentioned this at DPG
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and link to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1HyanBGukQ1zXUz25EOwtpClCfB-aoSYh3Aw2VZ9QAWI/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Jean-Michel Coeur, who is the Head of the Data Practice at the consulting company Sourced Group. Jean-Michel has developed a simple three question framework that works well with people asking for data, especially business counterparts. The questions typically lead to collaboration instead of confrontation and gets data consumers to share what they want to accomplish with the data instead of what is their request. It feels more like a friendly chat than an interrogation or "prove to me why this is worth my time". He also recommends following up each question with "the reason I am asking is..." to explain specifically you aren't pushing back, merely information gathering. The three questions: Do you know what this is for? Do you know who is going to use it? Do you know how they are going to use it?
Jean-Michel developed his three question framework after watching people struggle for years to properly request data and/or properly understand the use case of data consumers, often delivering solutions that did not meet business needs, wasting everyone's time. Oftentimes, the technical person wouldn't ask the right questions or they couldn't even get to the end data consumer so they didn't really understand the reasons for the data ask.
For Jean-Michel, the first question - Do you know what this is for? - helps to set the tone. It is not "why do you want this?", which often makes people defensive. He tells the person making the ask that with more context, his team can better understand how to make what they deliver better. And sometimes, the person making the request will realize they aren't really sure what it will be used for and can go back to the end user. A key is to not be a gatekeeper to the data, both in reality and perception.
The second question - Do you know who is going to use it? - starts to drive towards who will consume the data output and how - the use case is pretty important for delivering valuable data after all. For Jean-Michel, asking it in this way can often empower the person making the data request to lead the journey rather than undercutting them to get to the end user. It makes them part of the team. This is far better than the reaction he got previously when asking "who is the end user and what do they want?"
The final question - Do you know how they are going to use it? - means you can own the data user experience all the way through to the end consumption. Oftentimes, people deliver just the data into a warehouse or similar but if that is used in a dashboard, the dashboard is the user experience. This final question also helps data producers to suggest additional value-add features or to push back on some ill-defined requirements. E.g. "real-time" rarely actually means real-time. And Jean-Michel recommends understanding what you are delivering well enough to be able to answer the question of "okay, what does this mean?" After all, data projects/products are meant to deliver value and the value is informing business decisions with data.
And again, a crucial part of the conversation to add after...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and link to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Ole's Book (O'Reilly Early Release): https://www.oreilly.com/library/view/the-enterprise-data/9781492098706/ (https://www.oreilly.com/library/view/the-enterprise-data/9781492098706/) Transcript for this episode (https://docs.google.com/document/d/1wWCM6AQJ0GtGmYJ7XFR-kSUfWv2jzxVp4jUApNDSyMk/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Ole Olesen-Bagneux, an Enterprise Architect who focuses on data at GN and the author of an upcoming book on data catalogs with O'Reilly. To be clear, Ole was only representing himself and not GN. The two main topics, which are somewhat intertwined, were: 1) how can we better understand and handle the concept of a domain when discussing data; and 2) how can we build systems that better enable us to search "for" data, not just search "in" data that we know exists? Some practical advice and general conclusions from Ole: Leverage what the Library and Information Sciences discipline - which is centuries (millennia?) old - has formed around the domain concept. It will help you better dig into the actual business dealings of the domain first before trying to focus too much on the technical/software aspects. The software aspects hinder your initial domain mapping - especially in depth - and business context understanding when you start from a DDD perspective in data. Spend a lot more time on enabling people to understand what data is available. We focus a lot on optimizing for searching "in" data, but we don't spend near enough time setting up our systems to allow people to search "for" data. To do that, work seriously on your metadata tools system and look for ways to harmonize data across those tools.
Ole started the conversation sharing his view that Domain Driven Design (DDD) has some shortcomings when used especially for data domain mapping and in general in data. In his view, DDD is overly tied to software engineering so there is too much of a technical bent to understanding and even mapping out domains. He recommends taking domain analysis and domain theory learnings from the Library and Information Sciences discipline and using that to start your domain mapping and then look to bring in DDD after you get a good initial understanding of your domains. DDD and domain analysis can work together harmoniously, they don't really contradict, but domain analysis focuses on the knowledge first instead of the technical first. While Ole was inspired by Zhamak's book as well as the book by Piethein Strengholt, his believes domain analysis lowers the significant friction and often frustration organizations feel when trying to start doing DDD for data. Domain analysis digs much more into what the domain does and why instead of how the domain communicates via software. He believes that data mesh should focus more on the information sharing and less on the software and that DDD will overcomplicate your domain mapping. For Ole, DDD is overly concerned with modeling domains into software but you need to get to a deeper understanding of your domains and organization first before focusing in on your model. It may be that you truly can't fully communicate your domain's context in a data model either and it's good to know that upfront and take steps to communicate in other...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1AbcwrOmp5rbB1BFtQjpHZXj03ei-ur3ibgcY9G2DvEQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Shane Gibson, CPO/Co-Founder of AgileData.io and Agile Data Coach. A few takeaways from Shane to start: - Agile methodology is about finding patterns that might work, trying them out and deciding to iterate or toss out the pattern. It's going to be hard to directly apply software engineering patterns to data but we should look for inspiration there and then tweak them. - Any time you look at a pattern you might want to adopt or evaluate if a pattern is working for you, ask yourself: will this/does this empower the team to work more effectively? - Applying patterns is a bit of a squishy business. Get comfortable that you won't be able to exactly measure if something is working. But also have an end goal in mind for adopting a pattern - what are you trying to achieve and is this pattern likely to help you achieve that? - Share your patterns to not only help others but to get feedback and maybe ideas to iterate your pattern further. Shane's last 8 years have been about taking Agile practices and patterns and applying them to data as an Agile Data Coach. And those patterns required a lot of tweaks to make them work for data. A big learning from that work is that when applying patterns in Agile in general, and specifically in data, each organization - even each team - needs to test and tweak/iterate on patterns. And that patterns can start valuable, lose value, and then become valuable again. Shane gave the example of daily standups drive collaboration as a forcing function but then lose value when that collaboration becomes a standard team practice. If there is a disruption to the team where collaboration is no longer standard practice, daily standups could get value again. So how do we apply these Agile concepts to data? Currently, Shane sees no real patterns emerging in the data mesh space - it is quite early as patterns often take 5-8 years to develop and data mesh is maybe 12 months in to even moderately broad adoption and is such a wide practice area, there are many practice areas that patterns will need to cover. But, that lack of patterns makes it quite hard for even those who want to be on the leading edge of implementing data mesh instead of the true bleeding edge - having to invent everything yourself is taxing work! So we need companies to really take existing patterns, iterate on them, and then tell the world what worked and what didn't. If people aren't sharing patterns, that's going to make it hard to adopt data mesh for many organizations. Shane believes that it will likely be pretty hard for many organizations - or at least many parts of large organizations - to give application developers in the domain the responsibility of creating data products. If your domains aren't already quite technically capable in building software products, it's going to be very hard for them to really handle data needs. So looking at domains that are using large, out-of-the-box enterprise platforms or SaaS solutions instead of rolling their own software, will they really have the capability to manage data as a product? If their domains don't...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/17Fl1bLCmzqh-mlozccziGnWMslsPOfvJdejhgp3FIlw/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Vincent Koc, Head of Data at the merchant platform company hipages. To start with some big takeaways from Vincent: If you aren't comfortable with an agile mindset and ambiguity, bleeding edge probably isn't for you - and that's okay! You and your organization need to be comfortable with failing, learning, and then iterating To get data mesh - or really any big change initiative in data - right you should focus on change management much more than you probably think It ends up being the secret sauce or crucial lacking factor much more often than the tech Think problem-specific, not technology specific It's easy to over-engineer the problem - technologists want to technology In general, consistency is key to achieving widespread success in data One domain having a major success won't lead to broader org-wide success if you don't look to leverage reusability factors to make consistency across other domains easy - a bunch of great but non-consistent solutions doesn't add up to a valuable whole picture
For Vincent, every organization considering data mesh should ask if it is really the correct approach for them. Data mesh really isn't for a large subset of organizations, whether that is right now or even ever. If your organization doesn't have an appetite for change, it's going to be very tough to move towards data mesh. If you want to implement data mesh, he recommends embracing an agile methodology e.g. fast feedback and trial and error. When thinking about splitting your data monolith into domains, Vincent recommends taking a lot of learnings from what works well in the microservices realm. You shouldn't decompose everything all at once - that just creates chaos. You can split out larger domains one by one and then figure out if you need to split them further when there is more value in doing so. Peel them off instead of a big bang approach. Vincent believes that, in general, ~20% of your teams will consume ~80% of your data team's time and energy. There are a few ways to work with those teams to reduce that but it is also somewhat a fact of reality. Whether that is because those domains are more prominent, noisy, well loved, or for many other reasons. That data work disparity often leads to those areas being more data mature. When discussing disparate data maturity, Vincent talked about the need to drive all domains that will participate in something like data mesh to at least a common base level of maturity. You have to have a relatively mature domain to be at what he referred to as "mesh-level capability". Domains that aren't at that capability will still need to rely more heavily on centralized data teams and capabilities as they improve their data maturity. And it's okay to work with them closely to up their maturity level - just telling them to catch up is probably not going to work, there will need to be a bit of hand-holding. Vincent believes embedding data analysts - whether you call them data analysts, analytics engineers, or something else - into domains is crucial, especially if you are going to attempt to implement...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1OcU2AbWkY-7kLkmjAHW6-aN0IlZa3Hjmv_uSFKqoJ9I/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Paul Andrew, Technical Architect at Avanade and Microsoft Data Platform MVP. Paul started by sharing his views on the chicken and egg problem of how much do you build out your data platform and when to support your data product creation and on-going operations. Is it after you've built a few data products? Entirely before? And how that discussion becomes even more in a brownfield deployment that already has existing requirements, expectations, and templates. For Paul, delivering a single data mesh data product on its own is not all that valuable - if you are going to go to the expense of implementing data mesh, you need to be able to satisfy use cases that cross domains. And the greater value is in cross-domain interoperability, getting to a data product that wasn't possible before. And, you need to deliver the data platform alongside those first 2-3 data products, otherwise you create a very hard to support data asset, not really a data product. When thinking about minimum viable data mesh, Paul views an approach leveraging DevOps and generally CI/CD - or Continuous Integration/Continuous Deliver - as very crucial. You need repeatability/reproducibility to really call something a data product. In a brownfield deployment, Paul sees leveraging existing templates for security and infrastructure as code as the best path forward - supplement what you've already built to make it usable for your new approach. You've already built out your security and compliance model, make it into infrastructure as code to really reduce friction for new data products. For Paul, being disciplined early in your data mesh journey is key. A proof of concept for data mesh is often only focused on the data set or table itself, not actually generating a data product and much less a minimum viable data mesh. It's pretty easy to put yourself in a very bad spot because taking that from proof of concept to actual production is going to be a very hard transition and telling users it will take weeks to months to productionalize is probably not going to go well. Be disciplined to go far enough to test out a minimum viable data mesh. Paul emphasized the need for pragmatism in most aspects when implementing a data mesh. Really think about when to take on tech debt and do so with intention. When shouldn't we take on tech debt? And how do we pay down tech debt and when? There is a balance between getting it done and technical purity. How do we choose what features to sacrifice? What is the time-value to money aspect, or how much importance do we have on getting it done sooner rather than more completely? These are questions you'll need to ask repeatedly. Similar to what previous guests mentioned, Paul is working to encourage and facilitate the data product marketing and discovery process - discussing with data consumers what they want, pie in the sky thinking. Then taking that and speaking with data producers and figuring out pragmatic approaches and what is simple to deliver. Is one aspect going to be very difficult? Go to the consumers and let them know it will delay...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and link to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1ozNbiGSetfyoBHAWofygmGLgFN0PKza1U0RlaoChMgc/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Tim Gasper, VP of Product at data.world and the co-host of the Catalog & Cocktails podcast. They covered two main topics - 1) the skeptic's view of data mesh and 2) Tim's/the data.world team's "ABCs of Data Products" framework. Skeptics have a few main pushbacks on data mesh in Tim's view. Tim listed the top 6 that he sees and then discussed them with Scott.
Overall, Tim and Scott agreed that a lot of the pushbacks are probably...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1N9U9hCubjIs1waFjQxx7pxxjvOLxAM4WaQBugXFt9OQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed João Rosa, Principal Consultant at Xebia. They discussed domain driven design for data, the importance of intentionality in preventing chaos, being effective instead of efficient, and the concept of a hyper object. To start at the end, João talked about the need to embrace complexity when dealing with software - and we need to treat data and analytics as a software process. If we try to abstract away the complexity, we lose the nuance and that nuance is what can make all the difference in terms of the value of your data. Software is not like manufacturing where complexity is very costly. This was a pretty broad-ranging conversation starting with Domain Driven Design - or DDD - for data. João believes we should apply the principles of DDD to everything controlled by software - and when thinking of data as a product, data is definitely controlled by software. One of the big challenges with bringing something like DDD to data is that there aren't tools - and most challenges in the data space have historically been addressed with a tool-first approach. There is a desire to move quickly and just solve challenges but it's not possible to do that with DDD in João's view. A very interesting point of view João has is developing software is a learning process and working software is a consequence of that learning. With the move to cloud and the easy consumption of new tools, creating data is very easy. But João believes that in an enterprise, there needs to be very clear boundaries and contracts between domains to prevent overlap and confusion. The conversations between teams are hard because all of them are context-dependent. Even at the software level, your interface to your data products is a form of communication. João brought up the manufacturing-oriented philosophy of software development and why it causes so many challenges. It is very much about efficiency and lean development. That works well when you are producing physical goods but he doesn't think it does for software. Small incremental changes to software are not costly in a CI/CD world but the creation of software is expensive. So we need to move away from the manufacturing approach. But that would mean management releasing more control, which many are not willing to do. For João, there is also a major value to discovery about what you've already deployed. How are people using it, what is the market / consumer-base telling us? But in general, we spend far too much time focused on new features and not discovering new things about what is already in production. And those small incremental improvements are often the things that generate real value - and if the investment is small to generate good returns, those small changes are a significant point of potential value leverage. João brought up Kent Beck who said "once software arrives to production, it changes itself". Measuring that feedback is crucial. Data mesh, if done well, can really set up organizations to succeed because it can make people effective rather than efficient - we create data products that are easy to use but...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and link to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Scott shares some thoughts on some recent data mesh FUD and market confusion. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1F1bpl2wQkPuqqd1rhsT6IoK_ZkmOY9BhMGJd04UYj4A/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Katie Bauer, a Data Science Manager at Twitter in their Core-Tech group. To be clear she was not on representing Twitter, only her own opinions. The main topic of discussion was how to measure the value and success of your data projects/implementations. Some very useful advice from Katie that can feel a bit obvious when said but is VERY often and easily overlooked: measure for what would make you drive actions. If getting a 10x higher than expected or 90% below expected result isn't going to change your decision, while it may be interesting information, is it really important? If not, don't waste the time to measure it. Especially early on in your data measurement maturity. The point is also to get to an objective evaluation, not overly precise measurements. Set yourself up to improve and iterate. Don't make this hard on yourself. She also gave the pithy statement: what is valuable is not necessarily valued. Katie has a cake analogy that plays into data maturity well. Think about your need and the other person's capability regarding making a cake. Do you need a fancy cake for wedding or is this for a 3 year old's birthday party? One, you probably want to be special. One, if it vaguely resembles something from TV and tastes decent, the consumer will probably be happy. Is the other person capable of making a super fancy layered red velvet cheesecake or is a cake mix in a box probably more up their alley. How mature are the parties on creating measurement data and how mature or advanced do you need the output to be? Katie started the conversation talking about some survivorship bias / other biased ways of measuring. Often, she has seen throughout her career that people having success seek to prove their success via metrics instead of find the metrics that matter the most. That has some pretty obvious flaws so we need to move forward towards better measurement practices. For Katie, measuring the value of data science is pretty meta. Katie recommends starting out with some really easy measurements around engagement and usage. If it's a platform, what are your daily active users, weekly active users, and/or monthly active users - and what is the actual most useful metric? Should people actually be leveraging your project daily? Think about what is your addressable market and what percent of that market you have. And NPS (net promoter score) is a very lagging indicator. When thinking about metrics, there are two things that really stand out to Katie: first, what is your useful granularity? Don't get overly precise if you don't need to. You want an objective evaluation and anything past that can become overkill, which has an inherent cost. And second, what is your useful time-scale? Is this on a micro-scale, where the task should take 5min to complete so a difference of 5min is a big deal? Or is it a much longer time scale? When thinking about what to measure, ask yourself what does your company value. Is it shipping, usage, cleaning up tech debt/deprecation, etc.? Katie threw out a bit of a mind bender: what is valuable is not necessarily...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and link to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1b4hO2-i4gk_mxE57sRAuwZ_R58BWditeR6mdoBKliFQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Emily Gorcenski, Head of Data and AI at Thoughtworks Germany. Emily has put out some great content relative to data mesh. As a data scientist by training, Emily has a data consumer bent in her views on data mesh. She is therefore often focused on how can data mesh help "me" (her) as a data consumer. SLAs and SLOs come right out of the site reliability engineering playbook from Google. Overall, systems reliability engineering practices are crucial - Emily asked why don't we bring the rigor of other engineering disciplines to software engineering? So, what is an SLA and an SLO? Per Emily, an SLA is a contract between two parties - hence why agreement is in the name. This agreement should be written around an SLO with the SLO serving as a specific target. That can be uptime or latency in the microservices realm but with data, SLOs can get a little - or a lot - more tricky. The theory around developing an SLO is for it to directly connect to business value. Emily believes that when we think about SLOs and data, we shouldn't apply SLOs directly to the data but should shift those SLOs to the left and have SLOs in the software engineering practice that apply to data. Emily mentioned another antipattern for SLAs in general, which is not connecting them to SLOs. But when it comes to data, most teams don't even have any SLAs, connected to an SLO or not. As an industry, software engineering has figured out how to offer great SLAs to external parties but many organizations still struggle to offer good SLAs internally. For Emily, software-focused SLAs can even result in worse outcomes for data. If an SLA is about uptime, it might result in pushing bad data into a system so a service can maintain its SLA. When developing SLAs, Emily recommends starting with conversations and negotiations between both parties. If 5 9s of uptime is not valuable to your consumers, why build to ensure 5 9s? Dig into actual user needs and what will actually drive user value. And start to differentiate between infrastructure focused SLAs - like is the data product available - and data SLAs - like is the data updated and does it meet quality thresholds. Emily then started to talk about some of the fun very specific SLAs around data and what does data availability mean. These SLAs can get complicated but they can start to really drive towards what is actually valued by the consumers, what is the actual value of the data so you can then start to negotiate to drive a high return on investment. Again, we can avoid pre-optimizing for facets that consumers don't care about. Per Emily, good SLOs will tell you what you should improve. We should make sure our SLOs are decomposable to again, get quite specific when useful and/or necessary. It is much more difficult to do in data than in general software engineering - we can't think about data in a binary way such as accurate or not it is much more of a continuous spectrum. Emily recommends to look at the error budget concept and think about how we can apply that to data. Emily believes SLOs can help you to avoid building...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and link to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/18fQai0a4y0xTBGsLz2fDay2dJ7IG0I0fnUvX60kyoIM/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Ramdas Narayanan, Vice President Product Manager of Data Analytics and Insights at Bank of America. To be clear, he was not representing the company and was sharing his own views. Ramdas came on to discuss lessons learned from building effective data sharing at scale on the operational plane over the last 5-10 years so we can apply those to our data mesh implementations. A key output of the conversation is a guiding principle for getting data mesh right - your goal is to convert data into effective business outcomes. It doesn't matter how cool or not cool your platform is or anything else - drive business outcomes! It's easy to let that get lost in the tool talk and everything around data mesh. Per Ramdas, when looking at creating a data product, or really any data initiative, you need to align first on business objectives and that will drive funding. In the financial space, that is direct literal funding but even outside, you should have the same mindset. Make sure you get engagement and alignment across business partners, technologists, and subject matter experts. How are you using technology to address or solve the business problem? Ramdas has seen that if you don't focus on creating reusable data, you can create silos - you need cohesive data sets, not bespoke data sets for every challenge as that just doesn't scale. You should also study the data sources you are using - is there additional useful data you could add to your dataset or could you use that data for other purposes - keeping an eye out for additional data to drive business value will really add a lot to your organization. When working with developers, Ramdas recommends helping them understand how the business is going to consume and use the data and then figure out if they should deliver data as something like an API or web service or more of a custom batch delivery. It is important to also work with data consumption teams to be reasonable in their consumption demands - getting them to modernize can be a challenge and that can put an unreasonable burden on producing teams. Ramdas talked about how crucial conversations and culture are to getting data projects/products right. Sometimes the conversations can be tough but often they really aren't and there just needs to be open exchange of context and information, especially aligning on business objectives. Projects that fail typically have poorly defined business objectives or lack alignment. Per Ramdas, it is important to educate the business people on what data exists and even what data doesn't. That clouded vision of what data is available creates a lot of frustration - we need to get better in general at data discoverability so the business folks can know what is available and get access easily. Ramdas has seen repeatedly that good context via rich metadata also leads to better context sharing at the person-to-person level as it generates additional conversations. To emphasize that point a bit more, Ramdas believes that data discovery is the main spark for sharing context. Otherwise, we are at best exchanging data as 1s
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1ER4MVmZfRdG3itAIoLy0ZD-V9f-P4pSFEXfgaJfpcIU/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Justin Cunningham, who worked as a tech lead and data architect on data platforms at Netflix, Yelp, and Atlassian over the last 8.5 years. In that time, Justin was involved in initiatives to push data ownership to developers / domains. To sum up one of Justin's points he touched on repeatedly - he recommends to create a pool of low effort data which will inherently have low quality. Use that for initial research into what might be useful. Focus on maximizing accessibility - you can have governance and use things like join restrictions or give consumers an ability to self-certify that they are using the data responsibly. Once you get the use cases, then you go for the data mesh quality data products. Justin saw a lot of success at Yelp focusing on data availability - getting data to a place it could be found and played with - was a bigger driver for success than focusing initially on data quality. Once people discovered what data was available and how they might use it, the organization was able to work towards getting that data to an acceptable quality level. Another point Justin made was figure out which you want to optimize for in general: getting things right upfront or testing and changing. He believes in optimizing for change. Create an adaptive process and optimize for learning. Keep it simple and focus on value delivery - it will set up more tractable bets. At Yelp, they were trying to ETL a huge amount of data in their data warehouse to build reports for the C-Suite. But they were never really going to get enough data ingested to really meet their goals. It was taking them 2 weeks to create each new set of ETLs and that was just creation, not maintenance - it was looking like they'd need 5x the number of people. What Justin found the most useful at Yelp was to focus on getting as much "usable" data in an automated way. They achieved this initially through the data mesh anti-pattern of copying direct from the underlying operational data stores and building business logic on top of it. But, that data getting into the hands of the data team meant there could be an initial value assessment - once they proved out there could be value in the data, the conversation with developers was much easier to get them to care about providing clean and reliable data. Justin mentioned the same thing Wannes Rosiers mentioned in his episode: there are operational and analytical workloads but there should absolutely not be that separation when it comes to data. Data from operational systems is useful for analytics and vice versa. One thing that really helped developers understand how to share data was thinking of data sets as being similar to public APIs. At Netflix, there were just too many bespoke data sets - made it very hard to manage quality. What they found that worked was a data certification program for data sets, creating tooling to prove a data set was complete and accurate. That and upping the amount of focus on data set reuse significantly helped them to combat the data sprawl. Back to data accessibility and availability versus...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/16JheFqeHHGKj8t1hpUPM9oZlI2UEOzDUkgKzjFH4tXk/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Doron Porat, a Data Infrastructure Leader at the SaaS company Yotpo. Some crucial points Doron made: 1) Be kind to yourself when you make mistakes - it's worse to stagnate so don't be afraid of change and making choices 2) Build versus buy is always tough but don't let your ego get in the way and push you towards building everything 3) If you do buy, build a close relationship with your vendors to help influence the roadmap and have an outlet if you are having issues 4) A data platform team's job is to drive usage as usage means creating value - drive towards that and set your KPIs around platform usage 5) There will likely be many different types of consumers of your data platform - work to improve / optimize the user experience for most folks Doron is a technologist at heart so for each decision she instinctually wants to build instead of buy. And at the start of building out the data platform for Yotpo, that was typically her decision. But as the demands for more and more capabilities from the platform, the increasing ubiquity and quality/scalability of as-a-service offerings, and the growing need to drive usage and developer happiness instead of manage cool tech, she started to consume more and more managed services. When you are building out the platform, vendors, in the long run, can often better serve your needs because they have a whole lot of people focused specifically on making what you use better. You need to make bets on the vendors getting to where you need them to be and sometimes they don't pay off either. To up the chance of making those bets pay off, Doron recommends building relationships with your vendors to influence their roadmaps and get help when necessary. Doron strongly recommends putting together a framework for evaluating build/buy decisions. Some of the factors she considers are how extensible is the offering, is this taking on too many challenges in one solution, cost, the cost of later migration to or from a managed service, open source compatibility, etc. One thing Doron talked about that many teams seem to struggle with is the ego hit of saying someone else managing a service we use will drive more value. That's always tough but needs to be addressed. Doron talked about the strong need to drive your platform forward, not just be responsive. Provide a roadmap, set time aside for innovation, etc. Doron made the point there is a difference between learning to leverage a tool and operate it. You want to make sure you build out the knowledge around leveraging a tool whether you are operating it yourself or not. It's an interesting balance when you build that you don't focus too much on one or the other. Doron talked about the general job of a data platform team is to drive usage because that usage means creating value. Serving critical needs for users is crucial to driving adoption and focus the user experience on the business logic - no matter how cool the tech is, a data platform team's job isn't to expose that to users. And your team's KPIs should reflect usage. However, there are all kinds of users...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1Z1KdsEVjlKtGNAnp6UvK93Zm3TqTH8Z7MJNv7GY7ptk/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Samia Rahman, Director of Data and AI Strategy and Architecture at life sciences company Seagen. Samia is helping to lead Seagen's early data mesh implementation after helping with two implementations at Thoughtworks since the start of 2019. For Samia, interoperability is about taking information from two systems and combining them to get a higher value. A simple definition but a good one. Two potential key takeaways: 1) don't try to plan too much ahead for developing interoperability standards but definitely keep an eye out for places where you could start to develop those standards. And your standards really, really should evolve - you don't have to nail them right out of the gate. 2) your interoperability will also evolve - you don't need to make every data product interoperable with every other data product and you can start with basic interoperability first. The more you can standardize around unique identifiers, the better, but it's okay to not get it right first thing out of the gate. Samia started her career - and even before in school - focusing on software, especially end-to-end development. A repeating pattern for her has been how crucial contract testing is to getting things into a trustable and scalable state. We've had them in hardware and software for a long time and if you don't have easy testing, those systems often get replaced pretty quickly. Those tests are the safety net to allow for fast and reliable evolution. And that evolution is a key theme for this conversation - set yourself up to iterate and evolve as you learn. Work to not paint yourself in a corner Data standards, including specifically for interoperability, are everywhere in the life sciences space - FHIR, FDA has lots, etc. but it's still not great for truly sharing the meaning of the data. FAIR is trying to get there but the interoperability and domain knowledge isn't really standardized yet. Samia strongly recommends not getting ahead of yourself on interoperability and standards. It's perfectly okay to start small - iterate and build on your standards for interoperability, To start have some key identifying "linkers" done. Get things out in front of consumers so they can explore and give feedback and use that to power your iterations. Incrementally building towards a standard is crucial. If you are going to build a standard, reusability should be your first goal. If it is only for a single use case, that isn't a standard, it's just an implementation detail. Samia again recommends contract testing / a schema checker. And definitely leverage existing standards. It's also not a huge deal if you have more than one standard internally. You don't need one standard to rule them all. Per Samia, if you implement versioning, data consumers are usually very willing to work with data producers as they evolve data products. But without versioning, you are just pulling the rug out from underneath them. And right now, there isn't a lot of good info on versioning data out there, nor tooling. The need to evolve data products is why absolute self-service is...
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Scott shares his views on the importance of collaboration via negotiation, not requests, to make your data mesh implementation a success. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1JpbuFe-lZh2y7XiH_TttGCRP-nn5lrq16PF-PtdK5Pc/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Abe Gong, the co-creator Great Expectations (an open source data quality / monitoring / observability tool) and co-founder/CEO of Superconductive. One caveat before jumping in is that Abe is passionate about the topic and has created tooling to help address it. So try to view Abe's discussion of Great Expectations as an approach rather than a commercial for the project/product. To start the conversation, Abe shared some of his background experience living the pain of unexpected upstream data changes causing data chaos / lots of work to recover from and adapt. Part of where we need to get to using something like data contracts is to remove the need to recover in addition to adapting and move towards controlled/expected adaptation. Abe believes that the best framing for data contracts is to think about them as a set of expectations. To define expectations here, this would include not just schema but also the content of data, such as value ranges/types/distributions/relationships across tables/etc. So for instance, a column may be a one to five for rankings and then the application team changes it one to 10. The schema may not be broken - it is still passing whole numbers - but the new range is not within expectations so the contract is broken. At current, Abe sees the best way to not break social expectations is via getting consumers and producers in a meeting to talk about the upcoming changes and prepare, such as with versioning. But, as tooling improves, Abe sees a world where we won't even need a lot of those meetings going forward - either because data pipelines can be "self-healing" and automatically adapt to changes upstream or because metadata and tools for context-sharing will reduce the need for meetings. Abe sees two distinct use cases in general for data contracts or more specifically how people are using Great Expectations to implement data contracts. The first is purely defensively - put some validation on the data you are ingesting to prevent data that doesn't match from blowing up your own work; the second type is when the consuming team shares their expectations with the producers and there is a more formal agreement - or contract - with a shared set of expectations. The first often leads to the second, via an agreement conversation that happens after there was an upstream breaking change. Abe also mentioned there is a third constituent on data contracts in the room: the data. Sometimes the consumers and producers may agree on what they expect, but if that’s different than what’s in the actual data, then it's hard or dangerous to move forward. The data has a veto. There was an interesting discussion on the push versus pull of data contracts - should the producer team create an all-encompassing contract or should we have consumer-driven contracts? Would producer-driven contracts be too restrictive, preventing the serendipity insights data mesh aims to produce? Would consumer-driven contracts mean multiple contracts for each data product that the producer agrees to? Is that sustainable? So, to sum it up, the idea of a set of explicit expectations around a data product that are the result of collaboration between producers and consumers sounds like where we should all head if possible. If the expectation set is only...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1KGYF2GKYhaoRUl_FTYXZkloqaahw-v-Oe15HbZLcn9c/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Sadie Martin, Senior Product Manager, Data Platform at Q4 Inc about applying a product mindset to data in general. This is really crucial to getting data as a product right but also in building out your data platforms and even some processes for data mesh. Scott's summation of some key points: Anyone can apply a product mindset, not just the product manager Giving yourself the time before starting work to investigate and create you measurement framework, including your baselines, is crucial to measuring data work progress and choosing where to focus Approach your data work with intentionality Really understand what you are trying to accomplish and what your immediate customers/consumers are trying to use the data for to accomplish.
Sadie started as a data analyst where the team didn't have a product manager - they were doing a lot of work and weren't sure if things were likely to work or even if what they did had a positive impact after it was done. So she started to take on some of the task of answering those questions and transitioned into being a product manager for data. So, what is a product mindset? For Sadie, the easy definition but with lots of hidden depth, is "it's all about really understanding the problem". For most organizations, really thinking about the problem you are trying to solve is new relative to data. There may be a data request but what product or process is that data contributing to and what is that product or process trying to solve? Sadie believes measuring the problem is really crucial. Once you figure out what you are trying to solve, what is the scope of the problem? How are you going to measure if you are actually solving the problem? Especially is it better than what you were previously doing? She also talked about the importance of customer-centricity - really why are they making a data ask? Should this really be a one-off or a repeatable process? Did they ask for the complete set of what they need? Etc. One crucial insight Sadie has brought from product management to data is to be willing and ready to throw things away. If it ain't working, don't be too precious. That's a very different mindset than we've historically had relative to data. There's also the idea that processes can devolve quickly so ensuring when you start a repeatable data process, understand the effort to keep it going. While it feels counter-intuitive, Sadie laments that for most, it's often quite difficult to get the buy-in that you need data to measure if your data work is actually providing value. It's still worthwhile to do however. You need to take the time to do spikes and investigate ahead of time and slow down enough to set yourself up to measure results. Just continuing to go off assumptions and gut feelings is going to put you in a vulnerable spot to a competitor doing the work. Sadie looks at measuring the success of data work in two ways. The first feels obvious once said but really isn't: start by measuring the baseline. Without that baseline, you can't measure if you're having an impact. And lots of data work proves to be low value or negative value - you tried a hypothesis and it isn't working so stop and move on. How do you get to that answer fast? You measure the incremental change for the effectiveness.
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1ELI-maJxAipwHVuaR_mFHLqSD6AfNnnJqtqyj2gchtQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) This episode is part of the Data Innovation Summit Takeover week of Data Mesh Radio. Data Innovation Summit website: https://datainnovationsummit.com/ (https://datainnovationsummit.com/); use code DATAMESHR20G for 20% off tickets Free Ticket Raffle for Data Innovation Summit (submissions must be by April 25 at 11:59pm PST): https://forms.gle/gRa8hCFQAkmbtmSJ7 (Google Form) In the last of the interviews for the Data Innovation Summit Takeover week, Scott interviewed Henrik Göthberg, the Founder and CEO of consulting company Dairdux, the Co-Founder of the Airplane Alliance, and the Chairman of the Data Innovation Summit. Let's start with some conclusions/advice from Henrik: When working with other departments, in data mesh or not, you need to start from respect, empathy, and understanding for people in different roles. When you think about maturing a domain or process, a big bang approach very rarely works. You need to think about evolution, not revolution. To find a good pathway to maturity, start with the domains already on the leading edge, the innovators; trying to get the laggards to catch up instead of focusing on those who see value in maturity is going to lead to pain and likely not much progress. Start with less complicated and high risk challenges so you can learn and develop the right muscles to do things easier in the future. Focus heavily on reuse - reusable data, yes; but also templates and other "easy path" enabling things. To succeed in data mesh, you need to get to a place where you can have broad reusability. Reusable data, reusable processes, reusable templates, reusable tooling, etc. In a data mesh implementation, start with an initial domain but move on to adding a second domain quickly if possible. Templates will get you to value quickly. It's okay to skip automating or building out a great solution for certain pieces of your data mesh implementation. What will get you in trouble is building half-solutions that end up as major pain points. This is the biggest source of unintended tech debt. If you business people don't understand they own the processes and the data, your data mesh implementation is much more likely to fail.
Background and other color: Henrik covered his journey from 2012 to present in most of the first 30 minutes - from joining a domain to add analytics capabilities to that domain to building out a large data and analytics central team at the same company to joining a new company in 2019 to help them implement a new data strategy which has evolved into implementing data mesh. Henrik joined Vattenfall to build out the data and analytics team inside the sales org. They had a multi-country domain with different maturity levels across each country. They needed to improve the data and analytics capabilities and operations in all three countries so they could have strong data and analytics capabilities at the country and European level. The team had some technical savvy but they were struggling with actually getting the data - the data was locked into the source systems. It was difficult to even do basic customer analysis and data science, not to mention anything fancy. So they needed a lot of help in maturity. In 2015, Henrik became the Business Intelligence Officer at Vattenfall. That meant taking ownership of the...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1VjP8lguv5L3LNdgMV4MBnNw8SMCe_QNV83qr0a8NtRs/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) This episode is part of the Data Innovation Summit Takeover week of Data Mesh Radio. Data Innovation Summit website: https://datainnovationsummit.com/ (https://datainnovationsummit.com/); use code DATAMESHR20G for 20% off tickets Free Ticket Raffle for Data Innovation Summit (submissions must be by April 25 at 11:59pm PST): https://forms.gle/gRa8hCFQAkmbtmSJ7 (Google Form) Scott interviewed Jarkko Moilanen, a Data Economist and Country CDO ambassador for Finland. Jarkko will be presenting on "Data Monetization and Related Data Value Chain Requires Both Data Products and Services" on May 6th in track M6. Jarkko is approaching measuring the value of data and then trying to extract that value from data from many different angles. He is thinking about data products, data as a service, data as a product, etc. Per Jarkko, treating data like a product can apply a LOT of the learnings from the API revolution - this time around we can skip a lot of the sharp edges. APIs are about an interface to value creation - how can we treat data products the same way? We discussed the difference between return and return on investment. A data initiative may have a very high return but if the investment to get that return is too high, it's a bad initiative. How do we get to figuring out what quality level we need to solve our challenges - there is no reason to go for 5 9s quality if that doesn't move the needle. Jarkko coined a new concept on the call - the half life of data value. For a large percent of data, Jarkko believes the value of the data starts to fall considerably over a relatively short period of time. How can we extract the value when it is most valuable? If the half-life is weeks, days, hours, or even less? And how do we set ourselves up to get the most "bang for the buck"? Jarkko is firmly in the camp of intentionality when it comes to data. We can't keep betting on "this data might have value" or collecting data for the sake of collecting it. The data cleansing after the fact is difficult - what was the context at the time? Can you enrich the data further? Etc. - and the cost to do so is typically quite high compared to the value. And you keep incurring costs just to keep data around. If you ascribe to his data half-life theory, the value diminishes quickly so stop keeping around so much data you aren't using! Jarkko's data economy model has three layers that he adapted from the API economy: The bottom layer is private/internal to the organization only use. In Jarkko's view this is typically for organizations that don't have the capabilities to productize their data, that have low data maturity. If they can move toward productizing, it will enable reuse - not just for themselves but potentially third parties. The middle layer is to have closed sharing agreements with other organizations, creating data sharing ecosystems. These ecosystems are typically very limited in the number of other organizations involved. These are often also about creating value for a joint purpose, e.g. two suppliers sharing data specifically to meet a customer need. To do this, you need to productize your data enough to make it generally understandable and relatively easy to use. The top layer is the completely public data marketplace layer. Organizations participating at this layer are...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1nbE4TX34t5B5cJzyLUtNspym0LH7B02whMVqtAvDuy0/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) This episode is part of the Data Innovation Summit Takeover week of Data Mesh Radio. Data Innovation Summit website: https://datainnovationsummit.com/ (https://datainnovationsummit.com/); use code DATAMESHR20G for 20% off tickets Free Ticket Raffle for Data Innovation Summit (submissions must be by April 25 at 11:59pm PST): https://forms.gle/gRa8hCFQAkmbtmSJ7 (Google Form) Scott interviewed Daniel Engberg, Head of AI, Data, and Platforms at Scandinavian Airlines. Daniel will be presenting on "Structuring an Enterprise-Wide Data Organization" on May 6th in track M4. A key point Daniel made right away was that organizational structure should be tailored to accomplishing your goals - so we have to know what those goals are first. What are the capabilities we need to meet those goals? "Traditional" companies are often locked into their structure - silos by competence; so data engineering in one silo, marketing in another, sales in another, and so on. Daniel is interested in figuring out how we can split up the competencies to create cross-functional, cross competency teams but not cause chaos to the organization as a whole. Daniel gave an example of creating a cross functional team early in the pandemic as there were some very big threats to the business - being an airline when no flights are happening is a scary place. The cross-functional team was able to move so much more quickly than the way the company tackles challenges when it is business-as-usual, achieving their goals in a few days instead of what typically would have taken months. This cross-functional work also created new information sharing connections across the entire company which continues to create additional value. What Daniel learned from that experience, he is trying to replicate as best he can to make it the new business-as-usual instead of a one-off. As the head of AI, Data, and Platforms, he is working to infuse members from his team directly into more projects so they can be part of the teams and decisions instead of handling requests after decisions are made. It also gives his team members the ability to rationalize goals so there is a better ability to do maybe 80% of what would be requested with only 20% of the work in a month instead of the whole 100% in 6+ months - where is the value cut-off? Negotiate instead of take requests. For Daniel, product owners must start working to gather the competencies they need on their own cross-functional teams. But that can cause issues when domains start to hire when they lack strong knowledge in that competency. E.g. domains hiring data scientists when they have no idea how to find a good data scientist whose capabilities match their needs and goals. Or do they even want a data scientist instead of a data analyst? Then the career growth aspect gets challenging too - does a product owner need to know how to grow the career of 10 different types of widely varying roles. We talked about the challenges of dotted lines versus solid lines between a functional manager and a competency manager - who do you listen to? Can we have two solid lines for reporting structure? Daniel believes people want managers who understand their day-to-day work. As stated, hiring into domain teams directly is very tough. Competency Leads need to ensure the company has the right...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released This episode is part of the Data Innovation Summit Takeover week of Data Mesh Radio. Data Innovation Summit website: https://datainnovationsummit.com/ (https://datainnovationsummit.com/); use code DATAMESHR20G for 20% off tickets Free Ticket Raffle for Data Innovation Summit (submissions must be by April 25 at 11:59pm PST): https://forms.gle/gRa8hCFQAkmbtmSJ7 (Google Form) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1ls5QawrOffb3VGIfZCmPYHG0v729Ye7Gu7oKRaOYi8c/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) Knowledge Graph Conference website: https://www.knowledgegraph.tech/ (https://www.knowledgegraph.tech/) Free Ticket Raffle for Knowledge Graph Conference (submissions must be by April 18 at 11:59pm PST): https://forms.gle/Gy8KSMNDxbBfib2Z6 (Google Form) In this episode of the Knowledge Graph Conference takeover week, special guest host Ellie Young (https://www.linkedin.com/in/sellieyoung/ (Link)) interviewed Philippe Höij, Founder at DFRNT. At the wrap-up, Philippe mentioned that data architects should be able to communicate in ways other than PowerPoint. We need new and better ways to express ourselves, and the way things are connected. We will always need metadata around our data, we need text to express our ambiguity; but we don't have great ways to express things that are slightly ambiguous - not fully formed but also mostly known. A good tool allows to more easily query your model of the world to iterate and increment on it. That is where knowledge graphs can be the most helpful. Ellie responded, "It's not difficult, it's just complicated." Philippe shared his journey towards knowledge graphs, especially thinking about the AIDITTO project he and a team built out of the "Hack the Crisis Sweden" event in 2020 around COVID-19. He needed a way to prototype, visualize and collaborate on data and the connections between data at scale. A regular data model does not convey enough information about what the data is and how it relates. Ellie then shared some insight into the difficulties around collaborating on data across organizations and people in her climate change work at Common Action. Collaborating across organizations, all with different ways of working, you need a common "language" or way of communicating relative to data but can't easily develop a shared schema. Knowledge graphs provide incremental capabilities for collaboration. Philippe talked about when collaborating across organizations, you still have the needs master data management (MDM) tries - and often fails - to address but there is zero capability to manage the other organizations' data flow into the shared data "pool". Philippe was having issues with open-ended knowledge graphs like SparQL or OWL - they needed composable data structures to be able to be flexible in the case they can't fully decompose a concept, especially as the concept or their understanding of the concept evolves. For Philippe, TerminusDB was a big win because it allowed for composable data structures and much easier querying across the graph. Ellie discussed the origins of TerminusDB being about collaboration across many entities/organizations so it has a much different approach to accepting data that doesn't necessarily conform to a schema or data model. The "git for data" concept in TerminusDB also really was a big win for Philippe as it made experimentation much easier. Ellie shared some of the challenges in her work at Common Action around working with many different entities, many of which are small and not that data literate - or "data native" - and how they need to enable collaboration without rigidity as things are so dynamic. Philippe discussed the need for enabling people to collaborate in a "messy" environment - the world is changing and trying to spend all your time and effort categorizing it into a single schema isn't...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Scott gets grumpy and calls out some misbehaving vendors. And adds some fun new music to the episodes. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/19cdK0WancUlNT6mlaBxuXyiUcz0AlPxuZ2Tcjkc9tVA/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) Knowledge Graph Conference website: https://www.knowledgegraph.tech/ (https://www.knowledgegraph.tech/) Free Ticket Raffle for Knowledge Graph Conference (submissions must be by April 18 at 11:59pm PST): https://forms.gle/Gy8KSMNDxbBfib2Z6 (Google Form) In this episode of the Knowledge Graph Conference takeover week, special guest host Ellie Young (https://www.linkedin.com/in/sellieyoung/ (Link)), founder of Common Action, interviewed Olivier Wulveryck, Senior Consultant and Manager at OCTO Technology. In the first two thirds of the interview, Olivier and Ellie chatted about a lot of concepts specifically around data mesh and then they linked in the concepts around knowledge graph in the last third. For Olivier, a knowledge graph is the map for the data that is available - each data product or node in a data mesh is the representation of the knowledge within the organization. The knowledge graph is the abstraction of that knowledge across the data mesh - a logical representation on top of the data mesh nodes to help people make sense of the data mesh. Currently, Olivier is working with a client sharing their data in a data marketplace. They are working on implementing a knowledge graph on that but not on their internal data. If they are seeing value from applying a knowledge graph externally, they may apply to their internal usage. Olivier shared his view that it's easier to start with a data mesh than a knowledge graph - any first steps with a data mesh will bring you value. It is not the same with knowledge graphs - you have to do more work to get to value with knowledge graphs. Olivier previously worked on the operational side of software engineering. He realized they had lots of data sitting in databases but the data was just a consequence - it was state data, there was no temporal dimension. He wanted to apply machine learning to the data but he was suffering from low quality data, just like the data team. The producing team never really thought about the data to be shared with others. For Olivier, the way to fix this is to put data back at the center of the domain and flip the script - make the operational datastore the consequence of changes in the data. Olivier believes there is a need to think about the semantics of the data as it will be used by the rest of the organization - data can no longer be just an internal asset of the domain. If we believe we need the data to actually be used, we need that data to actually provide value. Make the data usable and useful. To get specific, Olivier shared a use case of a clothing retailer. They might be getting people's body type and measurements to help them choose clothes that fit better. But the company could also use the data to make better fitted clothes for a broader range of their customers. They might have more insights to change the way they tailor or design their clothing. Ellie asked about can we standardize how we capture and share data. Olivier is not sure if we can harmonize how we capture data. So instead, we need to harmonize on aggregation and integration. Olivier also talked about the challenges around finding the equilibrium between data consumer and data producer needs/wants. Per Olivier, adapt is better than adopt. There is no by-the-book way to implement data mesh because that would never be...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1ZCJGJU5jh5qqYN5wVrWMgTGRGOYF3PbkNO181I5PoKM/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) Knowledge Graph Conference website: https://www.knowledgegraph.tech/ (https://www.knowledgegraph.tech/) Free Ticket Raffle for Knowledge Graph Conference (submissions must be by April 18 at 11:59pm PST): https://forms.gle/Gy8KSMNDxbBfib2Z6 (Google Form) In this episode of the Knowledge Graph Conference takeover week, special guest host Ellie Young (https://www.linkedin.com/in/sellieyoung/ (Link)) interviewed Veronika Haderlein-Høgberg, PhD. Veronika was employed at Fraunhofer-Gesellschaft at the time of recording but was representing only her own view and experiences. She was invited for her special mix of both data mesh and knowledge graph know-how. At Fraunhofer-Gesellschaft, Veronika's employer up until recently, she and team were currently implementing a knowledge graph to help with decision support for the organization. And previously, Veronika worked on a data mesh-like implementation as part of the Norwegian public sector at the Norwegian tax authority before the data mesh concept was really congealed into a singular form by Zhamak. Veronika and Ellie wrapped the conversation with a few key insights: to share data, groups need to agree on common standards to represent it, and they also need to be able to share information with each other about that data into the future. To develop these initial data standards, and to build the relationships to coordinate around that data long term, different departments in the enterprise have to converse with each other. Building conversations across departments requires also building trust, and for this curiosity is a crucial ingredient, both on the individual level, but also at the domain and organizational levels. If people don't feel comfortable asking questions, they can’t understand each other’s perspectives well enough to contribute to that shared context. What does this look like in practice? Different departments coming to discuss the difference definitions they have of different terms, and finding out what data they need from each other, therefore, what data they must collect and protocols they must develop. And, computer scientists discussing data with business people—understanding both what business requirements are, and conveying the needs of data systems in order to provide organized, quality data. Veronika's recent organization, Fraunhofer, is using a knowledge graph as they need to make their investment decisions much more data driven. They need to do analysis across many different sources - they have some slight control over internal data sources but essentially none over external sources. They are repeatedly doing harmonization across these sources, often the same harmonizations. Veronika believes they shouldn't have to do the harmonization manually, so they needed a translation layer - the knowledge graph. To build out their knowledge graph, they need business experts to work with the ontology experts - however, it is a struggle for time and attention from the business experts and they need to learn the importance and how to do ontologies. This is when Ellie mentioned that by centralizing the integration, it might cost a lot of effort up front, but it's necessary if you only want to do the harmonization work once. For Veronika, thinking in the data as a product mindset and having data owners is crucial...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Knowledge Graph Conference website: https://www.knowledgegraph.tech/ (https://www.knowledgegraph.tech/) Free Ticket Raffle for Knowledge Graph Conference (submissions must be by April 18 at 11:59pm PST): https://forms.gle/Gy8KSMNDxbBfib2Z6 (Google Form) Thank you to our contributors! You can find additional introductory resources below. Contributors and their contact info: Karen Passmore, CEO and Founder of Predictive UX Karen LinkedIn: https://www.linkedin.com/in/karenpassmore/ (https://www.linkedin.com/in/karenpassmore/) https://daappod.com/data-mesh-radio/dont-sleep-on-ux-karen-passmore-steve-stesney/ (DMR Episode 30) Xhensila Poda, Machine Learning Engineer at CARDO AI Xhensila's LinkedIn: https://www.linkedin.com/in/xhensilapoda/ (https://www.linkedin.com/in/xhensilapoda/) Jens Scheidtmann, Lead Architect at Bayer Jens' LinkedIn: https://www.linkedin.com/in/jens-scheidtmann/ (https://www.linkedin.com/in/jens-scheidtmann/)
Juan Sequeda, Principal Scientist at Data.world Juan's LinkedIn: https://www.linkedin.com/in/juansequeda/ (https://www.linkedin.com/in/juansequeda/) https://daappod.com/data-mesh-radio/knowledge-first-juan-sequeda/ (DMR Episode 14)
Steve Stesney, Senior Product Lead and Data Practice Lead at Predictive UX Steve's LinkedIn: https://www.linkedin.com/in/stephenstesney/ (https://www.linkedin.com/in/stephenstesney/) https://daappod.com/data-mesh-radio/dont-sleep-on-ux-karen-passmore-steve-stesney/ (DMR Episode 30)
Tim Tischler, Principal Engineer at Wayfair Tim's LinkedIn: https://www.linkedin.com/in/timtischler/ (https://www.linkedin.com/in/timtischler/) https://daappod.com/data-mesh-radio/data-mesh-resilience-tim-tischler/ (DMR Episode 43)
Further Introductory Resources (provided by Ellie Young and Juan Sequeda): What is a Knowledge Graph video by Martin Keen: https://www.youtube.com/watch?v=y7sXDpffzQQ (https://www.youtube.com/watch?v=y7sXDpffzQQ) Knowledge graphs: Introduction, history, and perspectives (paper): https://onlinelibrary.wiley.com/doi/10.1002/aaai.12033 (https://onlinelibrary.wiley.com/doi/10.1002/aaai.12033) Knowledge Graphs book (free): https://kgbook.org/ (https://kgbook.org/) Knowledge Graphs paper by Juan Sequeda and Claudio Gutierrez: https://cacm.acm.org/magazines/2021/3/250711-knowledge-graphs/fulltext (https://cacm.acm.org/magazines/2021/3/250711-knowledge-graphs/fulltext) Juan's book "http://www.morganclaypoolpublishers.com/catalog_Orig/product_info.php?products_id=1658 (Designing and Building Enterprise Knowledge Graphs)", 25% discount using code DATAWORLD
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Knowledge Graph Conference website: https://www.knowledgegraph.tech/ (https://www.knowledgegraph.tech/) Free Ticket Raffle for Knowledge Graph Conference (submissions must be by April 18 at 11:59pm PST): https://forms.gle/Gy8KSMNDxbBfib2Z6 (Google Form) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1ieh2pwKIu2bk6uw5A2a1huADDVosHUioTkT_75KmlxI/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed two people from the European data consultancy GoDataDriven - Steven Nooijen, Head of Strategy, and Guillermo Sánchez, Analytics Engineering Tech Lead. Guillermo started off by talking about how for the last ~3 years, he was seeing the data engineering team as the bottleneck before data mesh came onto the scene. For Steven, they were seeing lots of companies that were building out data platforms, especially data lakes, and then not really getting the promised benefits so data mesh made sense. All agreed data mesh is not right for every company and then mentioned some good signs that an organization should consider data mesh. Guillermo pointed to a lot of the usual suspects: size of company, size of data team, how many data consuming teams do you have, how many data sources do you have, etc. He then gave a specific example: if you have a data analyst in a consuming domain that has to wait more than 1 week for data, there is a bottleneck somewhere. Is it centralization? Not sure but time to investigate and that might be where you start to consider data mesh. Steven gave the example of an even earlier indicator that bottlenecks are occurring: teams start to hire their own data people rather than leverage the central team. Guillermo also pointed to the rise of consuming teams getting direct data access from producing teams instead of going through the data team. Guillermo made a very crucial point: data mesh is really about interfaces. People talk about data products being the technical communication interfaces between teams. So as long as people trust the data product and teams adhere to interface standards, data mesh can work. But that trust is earned, not given. Both Guillermo and Steven talked about the need for an easy path, a golden path. With that golden path, it can be relatively simple and low friction for domain teams to share their data for the vast majority of use cases. There is also a need for some quality control to make sure data products are trustworthy, have reasonable SLAs, etc. Per Steven, what they've seen to date at customers is most domains are happy to follow that golden path. But, you certainly need that path to start out where they are - trying to make them use foreign tooling/workflows is going to be a tough sell. They are also leveraging the concept of a "certified" data product label. The core data mesh implementation team awards this to high-quality data products and it's easy for teams using the happy path get - it is harder to get certified if you use your own tooling. This acts as a quality control measure. Steven talked about when evaluating where on the centralization/decentralization slider decisions should fall, you should really look at domain maturity and organization maturity. Centralization provides more support, if being more of a bottleneck, in many cases so there is a trade-off. Are domains really ready to take on the full responsibility of whatever task you are considering? It's okay to not have a one-size-fits-all approach. Somewhat atypically, Guillermo and Steven recommended starting your proof of concept, if you need to do one, with a single domain. They recommended starting with an eager domain and to make sure they have support from the platform team. Pick a consumer-aligned domain
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released A mesh musings asking if we should embed data/analytics engineers into the domains to serve as the data product developers or not. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/15mp8YWgtYQxaOW9Rn_oltKF1oYYOPGYeSHmVJkmQuWw/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Sarita Bakst, Managing Director and a leader in Firmwide Data Management at JP Morgan Chase. Sarita was previously on a Data Mesh Learning meetup and is helping lead the firm's data mesh governance charge. Per Sarita, historically, data governance has meant controls and gatekeeping to most people - typically getting in way of innovation. So there needs to be a focus on changing that narrative, not just through words but actions to show it's not the case. In data mesh, you need to ensure that domains can make good decisions on governance and seek out subject matter experts when it makes sense. Sarita covered that one of the key issues in the way governance has been done is the people making decisions - the central governance team - don't have the real understanding of the data. When those decisions are put in the hands of the people who really know the data, but with guardrails and guidance, the fear is lifted about can we actually use this data and how. This opens up lots of new opportunities to leverage your data. Sarita strongly recommends starting with purpose-built data products. Find a use case and build data products to serve that specific. And data products MUST be about unlocking business value. You don't need to serve up all of a domain's data on day one, in version one of that domain's first data product - make it extendible and reusable so you can find additional consumers and expand over time. Get out of your own way on data governance in data mesh. You are going to learn and your approach will evolve as you learn. It's okay to not know everything upfront, set yourself up to not get in trouble - put the proper guardrails in place - but you won't know everything. Think of designing your risk controls as toll-gates and make sure they aren't bottlenecks. Have standards (not standardization) so people don't have to invent things from scratch. Standards for interoperability, naming, etc. - think of them as guiding principles instead of rules. Make sure domain owners know who to contact and when on governance subject matter expertise. Data Mesh Learning Meetup Presentation: https://www.youtube.com/watch?v=7iazNKG8XQo (https://www.youtube.com/watch?v=7iazNKG8XQo) Sarita's LinkedIn: https://www.linkedin.com/in/saritabakst/ (https://www.linkedin.com/in/saritabakst/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1PblP9tgQcqJp5ljyIH1MjNsSZmq9ot8pY3LzDiDBG1A/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Chris Riccomini, a Software Engineer, Author, and Investor. Chris led the infrastructure team at WePay when they embarked on a data mesh journey and made a well-written post on thinking about data mesh in DevOps terms. Like a number of people/organizations that have come on the podcast, at WePay, Chris was pursuing the general goals of data mesh and was applying some of the approaches as well - but it was not nearly as cohesive as Zhamak laid things out. Their initial setup had two teams managing the pipeline/transformation infrastructure. Chris's team was mostly handing the extracting and loading and then there was a team of analytics engineers handling the transformations. The Transformation team saw a major increase in demand and quickly became overloaded -> a bottleneck. Chris' team also started to get overloaded so they knew they had to evolve. One way the team started to address the bottlenecks was by decentralizing the pipelines. Teams could make a request and a scalable and reliable pipeline would essentially get automatically set up for them. WePay is in the financial services space so as part of those pipelines, to prevent risk, teams could mark their sensitive/PII columns and the infra team also put in some autodetection capabilities to make sure they didn't miss any. WePay created a "canonical data representation" or CDR, which is pretty analogous to a data product in data mesh. Chris really liked WePay's use of the embedded analytics engineer to serve as a data product developer. One key innovation for WePay was tooling to enable safe application schema evolution. They looked for things like dropped columns and had more comprehensive data contract checking mechanisms. It allowed developers to test changes pre-commit. 80-90% of data breakages were things the developers had no idea would cause an issue and they reverted those changes. 10-20% of the time, the developers still wanted to go through with the changes and that kicked off negotiations with data consumers. That forced conversation was very helpful for a few reasons. Chris talked about standardizing around technologies for the platform but allowing teams to roll their own if they wanted. But they were super clear with those teams that the infra team wouldn't support those other technologies, even if it was okay to use it. He also sees a major need for an API gateway concept for data. Currently, everything around versioning, auto-documentation, etc. is way too manual and high friction. Chris talked about taking the learning from DevOps and applying them to data mesh. One good one to look at, per Chris, is the embedded SRE concept - should you do the same with a data or analytics engineer? There is also a need for many standards and replicable patterns. They launched a data review committee, similar to a design review committee, that helped come up with standard data models and other standards. Sherpas not gatekeepers - build out your review functions as councils to guide and disseminate knowledge. The team's role should be about assisting where they can, being a trusted partner. And what WePay saw was as people went through more reviews and similar, they saw there was less of a need for them as people learned what good/best practices were. Lastly, Chris...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Weekly summaries quickly discuss the week's upcoming episodes and share the bottom line upfront summaries of the 3 interview episodes for the week. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1A4r9SkQP_x6ppIph_WGQxiKh3dDoY8f4hQ-OBJ0w0Xc/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Marius Ingjer, Co-founder and Senior Consultant at Knirkefritt AS. Marius has been working with a few clients on implementing data mesh. They covered 3 distinct topics: evaluating if data mesh is a fit for your organization, team structure challenges in data mesh or data mesh-like implementations, and a simplified definition of federated computational governance. Evaluating if data mesh is right for you: To start, Marius provided a list of evaluation questions to help you determine if data mesh might be right for you: How many data sets are you producing? What is the lead time to creating a new dataset? How well are your datasets serving your data needs? How many domains do you have? How complex are your domains? How does the team respond to new data requirements? How usable in general is your data?
Every company wants to share data well but the centralized data team isn't the bottleneck yet for many. Centralization can add a lot of value until it starts to become more hurtful than helpful and yes, figuring out that point is easier said than done. Centralization of data fights Conway's Law and can become way too much cognitive load so it will eventually become an issue for many organizations. A key question in evaluating if data mesh is a fit: what is the cost of allowing your data processes to fail? Per Marius, the business consequence of failed reports has historically not been that high. But if you are driving business decisions, whether that is ML or just crucial day-to-day decisions on your data, data mesh might become more attractive. Team structures and challenges in data mesh: In general, it's important to understand that implementing data mesh will cause cultural challenges - Marius believes developers generally don't want to ALSO share their data. It's additional work so you have to align incentives, which is far easier said than done. That additional cognitive load on developer plates is very crucial. We need to make we address that load to not burn them out. That means realigning incentives but also having extra help with things like grooming the work backlog. Providing extra resources helps but that is more about tackling the work, not handling the increased cognitive load. And learning about how to do data well is a pretty big learning task. Marius recommends giving teams the extra resources but also reshaping the team and business structure, such as the KPIs, to effectively prioritize and shape the requirements. He also recommends having a stick, not just carrots, or teams will try to just opt out. Simplified definition of federated computational governance: When talking about this topic with developers right now, it feels far too complex. To make it less complex, reduce the friction to developer decisions but not add much to cognitive load. An example might be providing easy data masking tooling for PII or extensible data APIs so developers can focus on the value-add. Marius' LinkedIn: https://www.linkedin.com/in/marius-ingjer-a155313/ (https://www.linkedin.com/in/marius-ingjer-a155313/) Marius' Medium: https://medium.com/@marius.ingjer Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn:...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released An episode about how to figure out what is good to reuse from your existing data approach and what to toss out when it comes to data mesh. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1ZwBSGVWypHJA9LYaqJjqM3jteaKdhLVfPhJhljCYwwY/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Scott Hawkins, Principal Data Architect at ITV. Scott views data mesh as a mechanism for change. Your company culture and your understanding of it are crucial to establishing data mesh well, driving that buy-in. Ask yourself: what challenges does data mesh actually address and hopefully solve, how will it impact the business not the tech, and what does it change. Your organization might not be ready for data mesh. Or a specific domain might not be ready. And that's okay! As ITV moved forward, they found a "good enough" solution via a global ID. It's not perfect, there might be some overlap - such as one person might have a different global ID for their online subscription versus their broadcast subscription - but it is far better than what they were doing. And it allows for interoperability/joins across the data. This is a big improvement - don't let perfect be the enemy of good or done. One thing working for ITV is deploying a "team-in-a-box" to help domains move forward - similar to an internal consulting team. Each situation is different so each box they are given is different. The team-in-a-box concept also means it is somewhat easier to build common best practices internally. Coming to the table with defaults has really helped ITV. Per Scott, there are 3 good ways to drive buy-in for the domain teams: At the senior level - so it trickles down as the management for the team is bought in. Via a strong carrot - solve a problem for them as a kind of quid quo pro / mutually beneficial solution. Trying to solve an unrelated problem will drive lower buy-in. Work on realigning the team KPIs/OKRs with the senior leaders to actually realign incentives.
Continuing on the driving buy-in, Scott recommends working with the domain managers to generate a viable/valuable carrot for the entire team. Explain to those leaders why it matters, work with the leaders to revamp the KPIs if the KPIs are getting in the way of delivering a good data product. This is why exec-level buy-in is so crucial - it is pretty hard to start modifying team KPIs/OKRs without it! Talking to teams to understand who they (the individual and the group) operate is crucial to developing the right path for them. Scott also talked about making failure an option. You can try to work with a domain and if it isn't working, it's okay to move on. You don't need to get everyone onboard on day one or sharing their data on day one. If you design incentives well, people will want to participate eventually. Until then, it's okay to walk away from that team. Scott Hawkins' LinkedIn: https://www.linkedin.com/in/scott-hawkins-8934393/ (https://www.linkedin.com/in/scott-hawkins-8934393/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/1hUFtLxvZf8OMK-ptJP-GEgNcaVi3LGFESb2-4DAiCO0/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Lorenzo Nicora, Principal Data Consultant at data mesh and AI focused consultancy Mesh-AI. Scott asked Lorenzo to be on to continue the series of interviews on domain driven design (or DDD) for data. It is a topic that many are struggling with so having lots of perspectives on it is crucial. On the episode title, a key output was explicit permission to skip a lot of the tactical patterns of DDD. Others have also said similar things but I wanted to make sure it was explicit. Before we jump into the DDD parts, Lorenzo made a good point on your data mesh Proof of Concept / starting your journey. You need to start with manageable problems. Start with a consumer-driven problem but a source/producer-aligned data product. There is a lot of nuance in the interview on why this matters. Per Lorenzo, identifying the domains is crucial but it is the hardest part of DDD. That shouldn't scare you because you can start with things being a bit blurry. It's important to understand your high-level domains but you can get moving without mapping out all of your domains. A key theme from Lorenzo: the language is at the center of everything in DDD. It is part of the data modeling and it goes all the way down to the code. Per Lorenzo, DDD is all about communication, knowledge capture, and knowledge sharing. Knowledge capture is about extracting knowledge and then writing it down. Knowledge sharing is about finding scalable ways to share context. Some advice/pointers from Lorenzo: Teams have to truly understand the language of their own domain - remove the ambiguities, even if that feels like it's putting in too much work. Event storming is a great way to approach tackling DDD for Data. Event sourcing is crucial for modeling the problem of the domain. Terminology is very key - identify the domain experts who can find/choose the right name for each concept. Keep a live document of terms and meanings - keep it updated! Encourage everyone to use the identified terminology when naming and in the code directly. Find your high-level domain first instead of your granular sub-domains. Ask your consumers for their specific data asks and then back into what would be a good data product or set of data products to start with. Again, look for high return, low effort/investment to get some wins under your belt and build your muscle memory.
Some key things to understand: Language changes - it changes across time and across the organization. The same words mean different things or different words mean the same thing. A major weakness of the central/enterprise data warehouse is the inability to easily deal with changes through time or nuance across the organization. When you first identify your domains, the boundaries might be blurry and that's okay! Data contracts are really crucial and the semantic issues, not the schema, are the most important - and hardest - part. And you can't just break contracts, there has to be a reason or no one will trust it is an actual contract instead of just a pub/sub model. Study up on and really think about your data on the inside versus data on the outside. If you aren't familiar, there is a link in the show notes to Pat Helland's work on the concept.
Lorenzo's LinkedIn: https://www.linkedin.com/in/nicus/...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcripts for this week's episodes are sponsored by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) Release schedule: Monday: Skipping the Fluff of Domain Driven Design for Data Mesh - Interview w/ Lorenzo Nicora Tuesday: Overcoming Obstinate Organizational Obstacles in Data Mesh - Interview w/ Scott Hawkins Wednesday: Differentiating the Baby and the Bathwater - Tooling Reuse in Data Mesh - a Mesh Musings Friday: Evaluating if Data Mesh is a Fit; Team Structures; and WTF is Federated Computational Governance - Interview w/ Marius Ingjer Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Transcript for this episode (https://docs.google.com/document/d/16EBMvfIqyEnf_0bdEjsZNj9yCM4NoLZIvQnUzRPLwtk/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Dan Sullivan, Principal Data Architect at 4 Mile Analytics. A key point Dan brought up is tech debt around data. Taking on tech debt should ALWAYS be a very conscious choice. But the way most organizations work with data, it is much more of an unconscious choice, especially by data producers, who are taking on debt that the data engineering teams will have to pay down. We need to find ways to deliver value quickly but with discipline. Zhamak has mentioned in a few talks that data engineers soon may not exist in orgs deploying data mesh. Dan actually somewhat agrees that data engineering will change a lot as right now, there is a big rush to build out the initial iterations of data products (the industry definition). Going forward, Dan thinks there will be a need for data engineers that can really understand consumer needs and build the interactions, e.g. the SDKs, to leverage data. Dan has 3 key pillars for driving data literacy for data engineers are domain knowledge, learning, and collaboration. Data engineers should pair with business people to acquire domain knowledge, they should be given the opportunity to spend time doing things like online training to learn, and they should collaborate across the organization instead of just being ticket tacklers. Per Dan, not all data engineers are the same depending on background - some come from a data analyst/data science background but many come from a software engineering background. So we can't treat training all data engineers as if it's the same. But we do need them to have a well-rounded background. A big need is for them to understand more about the data consumers and/or the producers so embedding them in the domains can really help. For driving buy-in with data engineers, Dan points to the problems typically being around incentives. Data engineering is often hampered by organizational issues and a lack of clear direction. So if you can tackle those, you can often win over DEs. In any organization but especially in one implementing data mesh, standards, protocols, and contracts are all very important. However, most data engineering teams are not given the time to create them. They take a lot of effort and are hard to get right! Dan talked about how data can take a lot of useful practices from Agile, especially the fast-cycle feedback loop. And that data people really need to think more about the user experience (UX) for data. Dan's LinkedIn: https://www.linkedin.com/in/dansullivanpdx/ (https://www.linkedin.com/in/dansullivanpdx/) Dan's Email: dan.sullivan at 4mile.io Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman):...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviewed Mohammad Syed, Lead Strategist - Data at Caruthers and Jackson about how data mesh governance has to be different from what we've done historically. Per Mohammad, data governance in data mesh is very different to doing governance for either a data lake or a data warehouse. The warehouse has a focus on high-level quality and usability but at the expense of context and agility. Data lake is about metadata and lineage but at the severe expense of usability - schema on query is not fun for consumers - and often quality. For most data organizations, governance has been very macro focused - governing the data warehouse or lake as a whole. That is part of why data governance has become a major bottleneck - the focus is on the macro but the individual requests are the micro. In data mesh, governance can shift to being about maximizing the value of the data instead of mostly preventing risk. Of course, there is a balance between local maximization - the value of each data product - and global maximization - the value at the overall data mesh level. A key focus to data mesh data governance is enabling - especially enabling the domains to govern their data products. Mohammad made the point that you need to enable your domains by creating the technical and business definitions of a "good" data product. Then the governance team needs to teach teams about the quality definitions, e.g. data product consumability. There is a need for policies of course but mostly focus on frameworks to enable policy creation and enforcement - decentralize! A key point Mohammad made was: governance only works with informed governors - you must teach domains to govern properly. Transparency is key to make data governance work. Mohammad emphasized the "good" data product definition leads to the separation of data quality and data product quality. A data product might be more valuable for other reasons - or less costly - by having relaxed data quality standards. In a data warehouse implementation, there is really only a single definition of "good" quality, but that just won't work in data mesh. We really need to develop better frameworks for what data quality means at the micro level. To get data governance right, strategy and maturity are crucial - what are you actually trying to accomplish? Data mesh for the sake of data mesh is worthless, just like any other paradigm. Mohammad's Latest Data Governance article: https://www.linkedin.com/pulse/some-thoughts-data-governance-mohammad-syed/ (https://www.linkedin.com/pulse/some-thoughts-data-governance-mohammad-syed/) Book mentioned: https://www.amazon.com/Disrupting-Data-Governance-Call-Action/dp/1634626532 Mohammad's LinkedIn: https://www.linkedin.com/in/mohammadsyed1509/ (Disrupting Data Governance: A Call to Action by Laura Madsen) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"):...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviewed Khanh Chau, Lead Architect for the Data Mesh Initiative at Northern Trust. Khanh believes you have to be passionate about making data better to do a good job implementing data mesh. And it is DEFINITELY a journey so you need patience and vision. Also, each journey is unique, you can't just copy/paste from another organization. You need to make failure okay - but you should look to make it easy to fail fast, measure, and adjust. Khanh talked about the need for exec buy-in before heading down the data mesh path. They got that exec buy-in by proving that the total cost of ownership of data was quite high as the consumers had to do a LOT of work to get the data to usable. When speaking internally, the business people were very excited to participate if it meant they could get quality data. Some of the IT/data engineering folks were harder to convince. It was especially hard to get them to shed layers of not-useful technology. Some IT teams were easier to convince - they had felt the impact of a few too many middle-of-the-night data downtime incidents. Other teams hadn't felt that pain so there were harder to win over. There was also the incentive of additional possibilities - data mesh meant they could do things they couldn't do before. Khanh talked about making the platform the easy and right path for 80% of use cases. They focused on making things easy to configure; basically: what transformations do you want to do and then it automatically provisions the pipelines. Their goal was to make it easy to make good progress quickly; their time to initial deploy went from 2-3 months per data service to 2-3 weeks per data product and they hope to drive it down further. Northern Trust has been moving forward with data mesh for about 7 months as part of their high-level digital transformation initiative. On the data side, they had previously focused on data virtualization and data federation but it was not delivering the results they wanted. It was not as scalable as they wanted - it was taking 2-3 months to launch each new data service. They also did not have great information on who was consuming the data and why. For their data mesh proof of concept, Khanh and team set a timeline of 9 weeks. They needed to prove value by then or data mesh would be a very tough sell internally. Khanh talked about the need to sell data mesh as a paradigm shift in order to get people out of technology-focused thinking. Northern Trust decided to take a pragmatic approach e.g. not pushing all aspects of data ownership fully left. Khanh and team were focused on finding a "happy balance" on data product SLAs and quality - improvement was necessary but the team preferred done to perfect. A big focus and a key driver for Northern Trust has been building muscle and learning/evolving along the way. It's important to evolve quickly and not build muscle in the wrong way. Northern Trust is still in the early days on figuring out interoperability between data products. It's more of an art than a science. Khanh believes bi-temporality is more important right now than interoperability. There are a lot of great learnings to takeaway from the Northern Trust journey. Khanh's LinkedIn: https://www.linkedin.com/in/khanhnchau/ (https://www.linkedin.com/in/khanhnchau/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Patreon) In this episode, Scott interviewed Tim Tischler, Principal Engineer at Wayfair. Prior to Wayfair, Tim worked as a Site Reliability Champion at New Relic and is well known in the "human factors" and resilience engineering space. Per Tim, our current work culture is overly action-item driven - every meeting must have a set of agenda items generated from it. This prevents people from having learning-focused meetings exclusively designed for context sharing. Humans' brains work differently between learning and fixing mode and we ask totally different questions. To be able to scale our knowledge sharing, we need to have the space to have learning-focused meetings. A good way to center learning-focused meetings, be they "show and tell" or event storming sessions, is via sharing stories - human communication is founded on story sharing through the millennia. Tim's "show and tell" and event storming sessions at Wayfair have had extremely positive reviews so far. Tim sees ticket-based interactions - just throwing requirements on someone's JIRA backlog or similar - as fundamentally flawed. If Team A gives Team B requirements, Team B just looks to close the ticket versus getting both sides in the room to exchange context and have a negotiation. Tim prefers two modes of interactions over ticket systems: #1 - no human-touch, automated interactions, e.g. an API; and #2 - high touch, high context sharing interactions. For resilience engineering specifically, you should apply learnings to each data product AND the mesh as a whole. Part of that is a broad acceptance that you are in a highly dynamic and highly changing org - there will be changes! A few anti-patterns to resilience engineering that apply to data mesh are: 1) a hub and spoke relationship model where one person is the key glue - this is bad at a human level and even worse at a technical level :); 2) business leaders pushing for metrics without sharing the specific context as the results end up as completely empty and useless things you are tracking; and 3) not embedding people building platforms into the teams they are building the platform for - they must really understand the workflows. Books/posts/papers mentioned: Blameless PostMortems and a Just Culture by John Allspaw - https://www.etsy.com/codeascraft/blameless-postmortems/ (Link) The Theory of Graceful Extensibility: Basic rules that govern adaptive systems by David D Woods - https://www.researchgate.net/publication/327427067_The_Theory_of_Graceful_Extensibility_Basic_rules_that_govern_adaptive_systems (Link) The Field Guide to Understanding 'Human Error' by Sidney Dekker - https://www.amazon.com/Field-Guide-Understanding-Human-Error/dp/1472439058 (Link) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Patreon) Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviewed Ust Oldfield, Principal Consultant at Advancing Analytics. They covered the concept of self-serve from a consumer standpoint in data mesh, and some ideas around how to get it right. According to Ust, the overall data and analytics industry is just starting to move from data consumers only consuming what others have prepared towards self-serve data consumption. But, it is important to still provide those prepared reports so 1) people are working from the same info / on the same page and 2) you give people an easy - and maintained - path to important business information. Ust also mentioned one key to getting self-serve right is to not just enable consumers to get to the data they want, they really need to be able to understand what they are seeing so documentation, sample queries, and other similar tactics are very crucial. Consumers also need training on how to use the platform - in general, training for self-serve data consumption is very lacking across the industry right now. How do we share information at scale? Forums? "Show and Tell"? Office hours? Neither Ust nor Scott had great answers just yet. Time will tell. Ust finished with a recommendation for those building out their self-serve platforms for data consumption: spend a lot of time interviewing your data consumers to figure out what will empower them rather than just trying to deliver what you would want. Also, make sure to enable those who just want to consume data as prepared - those who want to be spoon-fed the info, that's fine, allow them to self-select as that is a valid approach to leveraging data. Ust's LinkedIn: https://www.linkedin.com/in/ust-oldfield/ (https://www.linkedin.com/in/ust-oldfield/) Ust's Twitter: @UstDoesTech / https://twitter.com/UstDoesTech (https://twitter.com/UstDoesTech) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Patreon) Transcript for this episode (https://docs.google.com/document/d/1HwzX5uyqH-B5Wm6QotiZDnpsKWmLnQdGOF_r1nBqnew/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviewed Jessica Kerr more widely known as Jessitron who is a Principal Developer Advocate at Honeycomb.io. Or as many affectionally refer to her, the Empress of Software. You're probably gonna enjoy this one :) Scott asked Jessitron on because she had a very pithy tweet about data mesh being "conscious design for unexpected use" and because she knows the developer mindset extremely well and many folks are having trouble working with the developers to get them bought in to sharing their data well. Jessitron started off by discussing one of the big issues with application development: despite the tooling and process advancements of the last 20 years, it's all somehow only made application development harder. So the starting advice is don't just add more to their plate. It probably won't go well! Her biggest point was giving application developers agency in how they share their data is key. Autonomy is just passing the responsibility over without the help. Application developers want guidance/direction to the target outcome, but they want to make the choices on how to achieve said target outcome while being given the resources to do so. The information and capability to do their job is key. For driving buy-in, start with the why, not the ask. Let them know why their data is valuable. And BE SPECIFIC! The conversation should be about their potential impact, not just the negative of "you changed this and it broke" but the aspirational. You want them to start thinking about how can you work together to enable them to share their data in a high context / highly meaningful way. Data mesh is going to be a big culture shift for application developers - you need to not just put something high priority on the backlog. You need to give them the space - meaning that they have enough points or whatever on their backlog - to really understand and learn how to share their data well. You have to focus on teaching them how - possibly via an internal hackathon to start building that muscle - or possibly even a cross functional pair programming-like initiative to show each other your ways of working and share knowledge. Also, show them the impact they are having along the way as they get going, that will motivate them to do more. Be very conscious of language. The interview with the NAV team building the application platform talked about this a lot. Application developers and data people don't speak the same language. Work with them to put it into their language. Jessitron pithy tweet: https://twitter.com/jessitron/status/1471352190149087233 (https://twitter.com/jessitron/status/1471352190149087233) Jessitron Twitter: @jessitron / https://twitter.com/jessitron (https://twitter.com/jessitron) Jessitron emails: jessitron @ gmail.com OR jessitron @ honeycomb.io Jessitron Presentation: Principles of Collaborative Automation: https://jessitron.com/2020/08/03/talk-principles-of-collaborative-automation/ (https://jessitron.com/2020/08/03/talk-principles-of-collaborative-automation/) Jessitron blog post: To share the work, share the decisions: https://jessitron.com/2022/02/01/to-share-the-work-share/ (https://jessitron.com/2022/02/01/to-share-the-work-share/) Book mentioned: https://www.amazon.com/Why-Information-Grows-Evolution-Economies-ebook/dp/B00TT1VLAO (Why Information Grows: The Evolution of Order, from Atoms to Economies by César A. Hidalgo) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) https://www.patreon.com/datameshradio (Patreon) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/14kCvtE28g6HROYU_6Lzlemz9NjFhX6QaC9O1H7ryFE0/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) Scott interviewed Xavier Gumara Rigol who has been helping lead Adevinta's data mesh implementation as Area Manager for Experimentation and Analytics Enablement. The discussed the data as a product concept and learnings from Adevinta's journey thus far. Xavi has put out some great articles and did a Data Mesh Learning meetup that are linked below. One key aspect to data as a product is to understand the need for data product evolution, both relative to maturity and to what is consumed. This is a common theme in many data mesh conversations as historically, data consumption has resisted evolution and change. Consumers need to really understand that the business is evolving so what they consume will too. If you manage data products well, it won't be a sudden change but if we are trying to share insights into a domain, those insights will change. When thinking about data product maturity, it's totally okay to start by thinking of a data product as a single table or view. Xavi also mentioned some pitfalls to forced data product evolution - e.g. getting it wrong as changes can be quite costly to backfill. Adding new attributes is easy but computing something for 3 to 6 months in hindsight can cost a lot of compute charges. To do do evolution right, versioning and deprecation plans are key. To get data as a product right, Xavi recommends start by prioritizing which data you want to make available; this is a process, not a switch to flip. You should figure out which data is important for each domain and at the broader organization level. Applying data as a product thinking to your data sets is easier said than done. While data mesh is a leading proponent, companies not doing data mesh can also use data as a product thinking - Adevinta started down this path before embarking on their data mesh journey. Of course, data as a product is far easier said than done. For Adevinta's data mesh journey, they started with every data product being a single table. Data was originally centrally managed so interoperability was already established. However, the documentation was lacking and the general usability wasn't great. They spent their first few quarters just focusing on splitting their monolithic data production into separate pipelines for each domain instead of one giant cluster. The giant cluster was becoming a major bottleneck as changes were hard and maintainability was getting harder every day. Now, each domain essentially has one data product but with multiple dimensions/tables. Each product is layered and each layer has different granularity and SLAs. A few other notable points: Xavi believes all data products should be accessible via SQL but definitely not only SQL. Template/blueprints for data products are incredibly useful and important. The tooling/practices to prevent application changes from breaking the data are just very lacking - Adevinta uses data model reviews but it's still not perfect.
Xavier's Twitter: @xgumara / https://twitter.com/xgumara (https://twitter.com/xgumara) Xavier's LinkedIn: https://www.linkedin.com/in/xgumara/ (https://www.linkedin.com/in/xgumara/) Adevinta meetup presentation: https://www.youtube.com/watch?v=av6cT_r4orQ (https://www.youtube.com/watch?v=av6cT_r4orQ) Xavier's Medium Articles: https://medium.com/adevinta-tech-blog/building-a-data-mesh-to-support-an-ecosystem-of-data-products-at-adevinta-4c057d06824d (https://medium.com/adevinta-tech-blog/building-a-data-mesh-to-support-an-ecosystem-of-data-products-at-adevinta-4c057d06824d)...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript provided by Scott Hirleman https://docs.google.com/document/d/1riEZtawcGpwNvmq6N81iw16-zf_nSlFSTuAKE_ablHk/edit?usp=sharing (here) An episode all about driving buy-in by different personas for data mesh. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1OMy5r7dhPd5Wp_zIbeqgRp7MdmutMltdPK6g-BhZqAs/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviews Eric Broda, an Executive Consultant in the financial services space. Eric shared his learnings from aiding a large financial services firm to implement data mesh from the infancy of the project. Eric's big thesis for companies looking to be data-driven is to think of themselves as a platform for connecting supply and demand. The internal company may be the supplier, e.g. if a bank is lending money directly, but more often it is about being the platform - so here to match investors to consumers looking to take out a loan. Lowering the friction between both sides of your constituency is crucial and getting really good at data can help there. To Eric, the technology is like plumbing - you expect it to work but most businesses at the high level don't care about it as long as it works. You don't buy a house for the plumbing. Eric's big point of advice is that you shouldn't underestimate the organizational change required to do something like data mesh right. Plan for the change and don't try to skip the necessary change, that will lead to disaster. Speaking of organizational structure, Eric firmly believes that centralization of data ownership fails Conway's Law. While companies can overcome that with a LOT of effort, most don't get there due to fatigue. When developing a new data product, Eric recommends to first start with expected usage patterns pretty explicitly via a 1:1 relationship model; at least early in a data product's life, the data produced needs to explicitly match the needs of the first target data consumer. This is a departure from data mesh recommended practice but seems to be somewhat of a common emerging pattern, at least in financial services. Eric also stated his belief that master data management - or MDM - "is dead", especially in data mesh. It hasn't ever really worked and it's not worth trying to do it with data mesh. Time will tell on that one. Eric's LinkedIn: https://www.linkedin.com/in/ericbroda/ (https://www.linkedin.com/in/ericbroda/) Eric's Twitter: @ericmbroda / https://twitter.com/ericmbroda (https://twitter.com/ericmbroda) Eric's Medium (multiple posts on data mesh): https://medium.com/@ericbroda (https://medium.com/@ericbroda) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1TNj5nUeZNytw300emQQSilMK0ctg4sf4HH-GwFX77g8/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviews Audun Fauchald Strand and Gøran Berntsen of NAV. Audun is the Principal Engineer and Gøran is the Product Manager for NAV's https://nais.io/ (NAIS application platform) as well as their emerging self-serve data platform for data mesh called NADA. They covered a lot of different topics including: 1) building out the platform; 2) working with consumers to set expectations for common data products; 3) definition of a data product - and how it will evolve; 4) setting the frameworks for producer teams and allowing them to own the production; 5) communicating across teams; and 6) Cake! No really, a secret to success is cake. While NAV is early days in building out their data platform for data mesh, they are taking an interesting approach: work with the developers to set data product expectations and then see how the developers would go about creating those data products. Then, the data platform team will build the platform out to make developer workflows much easier. While Gøran, with a background as a data person, feels the pull to make the self-serve platform as data-centric as possible, he understands the need to make it developer friendly from his time building the application platform with those from a developer background like Audun. They both talked about reducing friction, including via sensible defaults, as a big part of their path forward. Stop trying to make developers come up with everything themselves. While they are still early days on developing those defaults, they are comfortable in their process to get there. And working with developers along the way is key. To start, NAV's definition of a data product is a single table or view. It will probably evolve to be more of a data set focus but they don't see a need to prematurely optimize or overcomplicate. Gøran emphasized the need to have empathy for data producers, to build that into the platform. Teams, whatever the strategic direction, can choose where they focus their time. Don't try to force them to spend it on data, spend the time to really work with them. As Brian McMillan said: find the opportunistic data folks. NAV tried to put analytics or data engineers into the domains but saw them sitting next to the team, not as part of the team. So they decided to rethink. Those data product developers were likely to become overly crucial to serving the data and thus were a likely single point of failure if they moved on. Okay, the most important aspect: Cake! For each team that puts a data product onto the mesh, they give that team cake. As in, an actual cake. It might seem silly but it really does work. It makes it feel less daunting to publish a data product and a bit like you are just having fun. It also means the team can show off a bit when they get their picture out there with their cake in the company Slack. And then people can use that cake picture as a jumping off point for learning more about the data product they just shared. It really is a fun community-building hack. This is a must-listen for anyone involved in building a self-serve platform for the application developers/data product developers. Gøran's Twitter: @gorzan / https://twitter.com/gorzan (https://twitter.com/gorzan) Audun's Twitter: @audunstrand / https://twitter.com/audunstrand (https://twitter.com/audunstrand) Gøran's LinkedIn: https://www.linkedin.com/in/g%C3%B8ran-berntsen-66066517/ (https://www.linkedin.com/in/g%C3%B8ran-berntsen-66066517/) Audun's LinkedIn: https://www.linkedin.com/in/audunstrand/...
Patreon is live. See it here: https://www.patreon.com/datameshradio (https://www.patreon.com/datameshradio) On Monday, we will have Building a Self-Serve Platform Developers Will Actually Use - Interview w/ Audun Fauchald Strand and Gøran Berntsen from NAV - BLUF starts at 5:24 On Tuesday we will have Pursuing "Platform Thinking" Through Data Mesh - Interview w/ Eric Broda - BLUF starts at 12:13 On Wednesday we will have a Mesh Musings on driving buy-in On Friday, we will have Getting Data-as-a-Product Right and Other Learnings From Adevinta's Data Mesh - Interview w/ Xavier Gumara Rigol - BLUF starts at 16:13 All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1RMcPjLNAakaj6RACEp4VIAofqSN1e7s2oTxbeBb_eeY/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviewed Dave McComb, the President and Co-Founder of Semantic Arts. Scott asked Dave on as part of the continuing deep dive into Domain Driven Design for Data and Data-Centric Application Development as Dave wrote the book on Data-Centric Application Development - https://www.semanticarts.com/software-wasteland/ (literally). Dave's overall argument is that most businesses really have very few "business events", ~500-2000 for even the largest companies. Those large enterprises may have 10K+ applications, each with their own data model and application model, leading to possibly 100M+ data attributes. All that leads to far more complexity than is necessary if companies just focused on building applications from the business events side. They discussed the amount of work an application developer would need to learn to be able to do data-centric application development; while it is mostly about learning data modeling, especially for graph databases, Dave has seen the application developers really not want to move to this model. This has meant a slower roll-out at a number of clients than if they were embracing it. Scott asked about the user experience (UX) in data-centric application development, both for the data producer and data consumer. Per Dave, the UX is pretty lacking, especially on the data producer side so there seems to be a need for better developer tooling for graph databases. Despite the "crude" UX, Dave says he sees data consumers really loving consuming data from a graph. The overall goal of data-centric application development is to provide simplicity and flexibility to organizations as most applications are too rigid for Dave and the system integration is even worse. As mentioned, the first 3 people who fill out a Contact Us on the Semantic Arts website and mention Data Mesh Radio will get a free copy of Dave's book. Semantic Arts website: https://www.semanticarts.com/ (https://www.semanticarts.com/) Dave's book: https://www.semanticarts.com/software-wasteland/ (https://www.semanticarts.com/software-wasteland/) Contact Dave: https://www.semanticarts.com/contact-us/ (https://www.semanticarts.com/contact-us/) Dave's LinkedIn: https://www.linkedin.com/in/davemccomb/ (https://www.linkedin.com/in/davemccomb/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Scott shares his emerging data mesh anti-patterns. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1Ei1x3pCjrrbBzrmTmCr3a3c7fWUYuQAzSFYutLVuY-E/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviewed José Cabeda, Data Engineer at Call-center-as-a-service provider Talkdesk. They talked about Talkdesk's start to their data mesh journey and progress so far. When José came across Zhamak's original post, it spoke to a number of the challenges Talkdesk was facing, checking many of the boxes to where they wanted to head. The team started from a single data product and iterated from there. While they are still relatively early in their journey, like every company, they have advanced far past their initial use case. At Talkdesk, a data product is typically a single table or view in Snowflake but the company's North Star is event streaming as their key information storage and sharing mechanism. However, it was sometimes difficult to train people to understand the difference between a business event - something that occurred in the real world - and an event streaming event. José had a few key takeaways and recommendations for those implementing data mesh: 1. Change will be constant in a data mesh implementation so it is best to standardize the way people and systems will interact as much as possible. Define expectations! 2. Be open to new ideas, there are many challenges ahead so it's best to face them together. 3. Use a single universal ID for major concepts like account or business events to make interoperability easier / possible. 4. Don't be afraid to slice your data in different ways to serve different use cases. 5. To drive buy-in, start with a single use case, whether that is a data product or multiple data products - most people recommend 2-3 data products in your PoC - so you can show why data mesh is a good idea. José's LinkedIn: https://www.linkedin.com/in/jecabeda/ (https://www.linkedin.com/in/jecabeda/) José's Twitter: @jecabeda / https://twitter.com/jecabeda (https://twitter.com/jecabeda) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1ls5QawrOffb3VGIfZCmPYHG0v729Ye7Gu7oKRaOYi8c/edit (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviewed Thinh Ha, Strategic Cloud Engineer at Google Cloud Professional Services. To be clear, Thinh was only representing his own views and was not representing Google/GCP in any way. Scott had asked Thinh to be on after Thinh wrote a post on Medium called 10 Reasons Why You Should Not Adopt Data Mesh (later changed to 10 Reasons Why You Are Not Ready to Adopt Data Mesh). While Thinh is a self-professed "believer" in data mesh, he brings up a number of very reasonable checklist/self-check reasons you wouldn't be ready to move towards data mesh yet. Scott and Thinh go down each of the 10 objections/reasons through the episode and it is advisable to read the article before proceeding. You can see the high level reasons below. There are a lot of very valuable insights into each of the reasons that could make this a 5 page summary so just listen to the episode instead ;) 1. You are not operating at a scale where decentralization makes sense 2. You do not have a strong business-case for how adopting Data Mesh will deliver business value for individual business units 3. You treat Data Mesh as a technical solution with a fixed target rather than an operating model that continuously evolves over time 4. Your organizational culture does not empower bottom-up decision-making 5. You do not have clearly established roles & responsibilities and incentive structure for distributed data teams 6. You do not have a critical mass of data talent 7. Your data teams have low engineering maturity 8. You expect to find off-the-shelf software to help you adopt Data Mesh 9. You do not have buy-in to “shift-left” security, privacy, and compliance 10. You do not consider Data Governance to be a core activity to be prioritized against other activities in every data team’s backlog Twitter: @thinh_ha / https://twitter.com/thinh_ha (https://twitter.com/thinh_ha) LinkedIn: https://www.linkedin.com/in/%E2%98%81%EF%B8%8F-thinh-ha-58945969/ (https://www.linkedin.com/in/%E2%98%81%EF%B8%8F-thinh-ha-58945969/) Medium post: https://medium.com/google-cloud/10-reasons-why-you-should-not-adopt-data-mesh-7a0b045ea40f (https://medium.com/google-cloud/10-reasons-why-you-should-not-adopt-data-mesh-7a0b045ea40f) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) 4 quick things: 4 episodes this week; interviews with Thinh Ha (re 10 Reasons You Aren't Ready for Data Mesh article), José Cabeda (Talkdesk's User Journey), and Dave McComb (more on Data-centric Application Design) and one mesh musing on anti-patterns. Still having issues getting transcripts done for episodes. If you are at a vendor that wants to sponsor transcripts, please let me know. Patreon to launch soon, probably Friday. Starting work on the getting started / proof of concept guide very soon.
As always, please rate and review the podcast on your favorite app. It helps to get it to show up in search results. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/133fTlVue-K-hUwpabjsYDj-adm6IGeKd0KVkkQtxo9M/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova?datameshradio (here) (info gated) In this episode, Scott interviews Azmath Pasha, member of the Forbes Technology Council, who has 25+ years in implementing large-scale IT projects including at CapGemini and Paradigm Technology. Azmath gave his 3 key measures for data value: cost savings, business value (e.g. driving new initiatives), and data reuse. For data mesh, the long-term value is in the second two but for Azmath, a PoC could be better served focusing on cost savings as it is easier to track and faster to realize. They dove into the concept of data discovery with human interaction, not purely an online experience. Similar to event storming for discovering your domain events (see DDD for Data episodes), discovery as a purely tool-based experience is always likely to be somewhat lacking. Scott was intrigued about this as that aspect of data discovery hasn't been widely discussed. To Azmath, the data product experience, part of what Zhamak calls 'the experience plane', is crucial. It is much harder to drive buy-in if your product is hard to use / has a bad user experience. Azmath's other crucial aspects to getting a data mesh (or any large scale data project) implementation right included: staying tool agnostic so you can remain "future proof"; supporting data producers to reduce time to delivery, especially initial delivery; and looking at your architecture and tool investments over a 5 year time horizon, not just for the short to medium-term. Azmath wrapped up by saying we are entering a new era of using data, we must democratize the data and also look to new metrics for evaluating the business value of data. Azmath's LinkedIn: https://www.linkedin.com/in/azmathpasha/ (https://www.linkedin.com/in/azmathpasha/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Kurt Gardiner, Engineering Manager of Data Engineering at Australian Insurance company nib Group. Kurt shared some insights into nib's journey so far, including the search for something like data mesh before Zhamak published, tool choices (Snowflake, dbt, Fivetran, EventBridge, Kinesis), the slow-role approach to replacing legacy implementation (the "application strangler" pattern mentioned), how they got started, and much more. Much of nib's approach is the small-scale tactical while building incrementally for the bigger strategic focus. E.g. helping teams to design their data products somewhat manually while building the reusable tooling to be far less manual going forward. Along their journey, there was some internal pushback from data consumers, especially those used to consuming from the data warehouse. To do data mesh right, Kurt and Scott both emphasized the need to set things up so they can evolve. That will frustrate or scare some people and it's important to work with them to see why that matters. There also needs to be a high tolerance for failure - you will NOT get everything right on your first go. Kurt also waxed poetic (said nice things) about event streaming patterns, especially CQRS - see link below for more info -, for a useful and scalable pattern that is good for both application development and creating a scalable and useful domain data model. But it requires a complete redesign so it is probably something to slowly introduce where it makes sense, if at all. Some pithy nuggets of wisdom from Kurt that are highly applicable to data mesh: "The single biggest problem in communication is the illusion that it has taken place" "Nobody cares what you know until they know that you care" Application Strangler pattern (recently renamed Strangler Fig Application pattern): https://martinfowler.com/bliki/StranglerFigApplication.html (https://martinfowler.com/bliki/StranglerFigApplication.html) CQRS: https://www.martinfowler.com/bliki/CQRS.html (https://www.martinfowler.com/bliki/CQRS.html) Kurt's LinkedIn: https://www.linkedin.com/in/kugardiner/ (https://www.linkedin.com/in/kugardiner/) nib Group careers page: https://nib.wd3.myworkdayjobs.com/careers (https://nib.wd3.myworkdayjobs.com/careers) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviewed Karen Passmore (CEO) and Steve Stesney (Data Product Lead) at consulting firm PredictiveUX. They touched on a lot of different topics - a key theme throughout is the importance of the user experience in data mesh, for both data product producers and consumers. Karen highlighted some parallels between data mesh and content management projects and how to take some key learnings from the past and apply them to data mesh implementations. They discussed the importance of providing your internal people with the right content at their current point in their learning journey - a successful implementation of data mesh requires making it far easier and more scalable to share incremental work artifacts and knowledge - it also means your crucial company knowledge actually gets documented properly. Steve talked about some historic challenges he had personally with decentralized teams - if you don't manage the cross domain collaboration, both at the business and the technical implementation levels, it is a major pain to stich your data together from all those sources. So there needs to be good alignment on interoperability. Basically, data mesh without a good interoperability strategy is just high quality data silos. Karen and Steve both emphasized the importance of UX (user experience) for driving adoption. You can have the best solution in the world but if the users don't want it, it's not going to be successful. So working with them throughout the process is crucial to get to a successful implementation, whether that is data mesh or not. Karen wrapped up by emphasizing the need to be patient and to not expect the same results or try to copy the exact path of other organizations implementing a data mesh. Every organization is very unique and you need to figure out what might work for your organization. Take learnings, not exact blueprints. The last key point to extract is the need for multiple communication methods, especially for data requests. There may be some overlap but it's a great way to ensure reliability and scalability of your business processes. PredictiveUX website: https://www.predictiveux.com/ (https://www.predictiveux.com/) PredictiveUX partnered meetup: https://www.meetup.com/hexagon-ux-dc-chapter/ (https://www.meetup.com/hexagon-ux-dc-chapter/) Karen LinkedIn: https://www.linkedin.com/in/karenpassmore/ (https://www.linkedin.com/in/karenpassmore/) Karen Email: karen at predictiveux.com Steve LinkedIn: https://www.linkedin.com/in/stephenstesney/ (https://www.linkedin.com/in/stephenstesney/) Steve Email: sstesney at predictiveux.com Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Andrew Jones, Tech Lead of the Data Infrastructure Team at GoCardless. Andrew shares the story of how operational system changes kept breaking downstream data consumption (sound familiar?), especially using CDC. The software engineers couldn't easily use the CDC tooling and the data engineers could easily use the data from it either as CDC didn't structure the data for easy consumption. Andrew wasn't really sure how other people were handling taking the API contract concept and leveraging it for data but started building out some generic simple tooling to let consumers and producers feel somewhat comfortable with their data contracts. A big revelation was in helping data consumers make better asks for data. The data consumers weren't used to asking the producers for data, especially in a reliable and scalable way (sound familiar?). GoCardless now has an actual standard form for data consumers to use to request data and that is working quite well. GoCardless plans to completely remove their CDC architecture by 3Q of this year to replace with data contracts. They are focusing on providing tooling to give domains the autonomy to serve data to consumers in their own way. While it isn't data mesh, especially with the lack of interoperability between data products and lack of source/producer-aligned data products, it seems to be working for GoCardless thus far. Andrew's Medium post called Improving Data Quality with Data Contracts: https://medium.com/gocardless-tech/improving-data-quality-with-data-contracts-238041e35698 (https://medium.com/gocardless-tech/improving-data-quality-with-data-contracts-238041e35698) LinkedIn: https://www.linkedin.com/in/andrewrhysjones/ (https://www.linkedin.com/in/andrewrhysjones/) Twitter: @andrewrjones / https://twitter.com/andrewrjones (https://twitter.com/andrewrjones) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews two Domain Driven Design (DDD) experts from Thoughtworks - Danilo Sato Director and the Head of Data & AI (part of Office of the CTO) and Andrew Harmel-Law, Technical Principal. This was a further delve into Domain Driven Design for Data after the conversations with Paolo Platter and Piethein Strengholt. Danilo and Andrew gave a lot of great information about Event Storming, domain definitions and boundaries, ubiquitous language, and so much more but the main theme was "just get people to talk to each other". DDD is about bridging the gap between how the tech people talk and how the business/business people talk; if you are doing it right, both sides can understand each other and then the engineers can implement those business process learnings as part of the code. For an initial PoC, Danilo recommends starting with 2-3 data products. It is better if you can do the PoC across multiple domains but it isn't necessary. Validate value and do it quickly. As Andrew mentions, the earlier you can show value, the less pressure there is overall. Look for the initial quick wins while also building for the long-term. One key thing to remember, per Danilo, when doing DDD for data and data mesh in general: it is always an iterative process. Andrew briefly discussed a way to do DDD in more of a guerilla style than the blue/red books (well known DDD guides). Don't get ahead of yourself as Max Schultze mentioned in his episode. Do not let the size of the eventual task throw you into analysis paralysis. Andrew talked a lot about how normalization and strong abstractions on the application side make it very difficult to re-add the context lost when you normalize. Both Andrew and Danilo talked about the need to embrace complexity. If you want context, you have to accept there will be complexity. In the pursuit of simplification, you lose the richness, and that is VERY hard to reconstruct afterwards. Some practical advice for boundary definition is that the boundaries need to be very clear but malleable. Build everything with an eye that it will evolve. Before you start splitting into many 2 pizza teams, look at the big picture and select some coarse-grained boundaries. It is MUCH easier to split later than it is to glue things back together. Danilo's Webinar with Zhamak called "Data mesh and domain ownership": https://www.thoughtworks.com/en-us/about-us/events/webinars/core-principles-of-data-mesh/data-mesh-and-domain-ownership (https://www.thoughtworks.com/en-us/about-us/events/webinars/core-principles-of-data-mesh/data-mesh-and-domain-ownership) Vladik Khononov, 7 Years of DDD: https://www.youtube.com/watch?v=h_HjtYAH0AI (https://www.youtube.com/watch?v=h_HjtYAH0AI) Danilo LinkedIn: https://www.linkedin.com/in/danilosato/ (https://www.linkedin.com/in/danilosato/) Danilo Twitter: @dtsato / https://twitter.com/dtsato (https://twitter.com/dtsato) Andrew LinkedIn: https://www.linkedin.com/in/andrewharmellaw/ (https://www.linkedin.com/in/andrewharmellaw/) Andrew Twitter: @al94781 / https://twitter.com/al94781 (https://twitter.com/al94781) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Angelo Martelli, Group Leader of Data Services at logistics company Vanderlande. Angelo laid out his framework for driving data mesh buy-in internally at Vanderlande which helped them take the idea from a small group to a company-wide initiative: Start with proving there is a problem that you are trying to solve - if everything is functioning well, why focus your efforts on that area instead of trying to fix another? Your proof should be as fact-based as possible, e.g. how long does it take to make a change to your data warehouse. Focus on proving that your incremental investments are driving sub-linear returns. Other areas to look to prove problems: how many people are involved in a change to your data warehouse, percent of time spent on regression testing versus development, mean time to resolution of challenges, etc. Once you have some proof, you need to work towards understanding the problem you are trying to solve. It's not "deploying a data mesh", it's scaling the organization to be agile relative to data and be able to make more (and better) data-informed decisions. Next, you need to understand your organization. Who are the right people that can help you? How does your organization work relative to culture and process? Which domains are struggling and how? Tie the implementation goals to the actual business challenges. Then, you need to demystify data mesh, make it easy to understand for people not well versed in data - what are we actually trying to accomplish and why? Last, make it concrete / prove it out. Make a few data products, make a simple platform for folks to use. Angelo then recommends that once you have momentum, sharing a very clear vision is crucial. Not just sharing in a document but actually having conversations to really make sure the context and vision is understood. Data mesh is about collaboration, you must work together so it is imperative to make expectations very clear. Similar to Abhi Sivasailam, Angelo also stressed the importance of the domain data model and abstracting that away from the application model(s). The business model is what matters for data. All of that and so much more. Also, Angelo gives a shout out to the usefulness of the Data Mesh Learning community. 😎 Angelo's LinkedIn: https://www.linkedin.com/in/angelomartelli/ (https://www.linkedin.com/in/angelomartelli/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) 4 quick things: Please tell Starburst you want transcripts. Do so on https://twitter.com/starburstdata (Twitter) or https://www.linkedin.com/company/starburstdata/ (LinkedIn). Please rate and review the podcast on your favorite app. It helps to get it to show up in search results. Going from 2 interview episodes a week to 3 for the foreseeable future. Still working on Patreon stuff. I am not looking to gate off content but it will include episode early releases (e.g. there are 8 interview episodes done at time of this release).
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Scott interviews Brian McMillan a former Enterprise Architect who took time off to write a book called 'Building Data Products: Introduction to Data and Analytics Engineering for Non-Programmers'. You can learn more about the book - and get a free copy, see below - here: https://www.minimumviablearchitecture.com/ (https://www.minimumviablearchitecture.com/) Brian's book lays out a path for the people who are doing the most with data in domains to elevate their skill sets and produce small-scale data products. They do this through a slow ramp from understanding SQL queries to learning data modeling to learning how to publish their data and use simple orchestration tooling. It isn't magic, it will take time and training, but it means you have more people with strong domain knowledge becoming part of the data and analytics engineering process, sharing their business context in scalable and repeatable ways. Brian's approach can also be used for a pretty easy path to an exploratory platform. There isn't a lot of pre-build to get going so teams can much more easily test out a hypothesis or two rather than it being a lengthy and costly approval and build cycle. There is also an easy path once someone finds a "there" there, to move it to something far more scalable and reliable in the cloud. There is a lot from the book and interview that can be adapted to help level up your teams' data literacy. Brian is giving away 10 copies of his book for free to those who sign up for a chat to share more about your current challenges related to the book topic or shadow/domain-based IT. To claim your free book, fill out a contact form https://www.minimumviablearchitecture.com/contact.html (here) and mention "Data Mesh Radio" in the comments. Brian's contact info: Email: brian at minimumviablearchitecture.com LinkedIn: https://www.linkedin.com/in/brianmcmillan01/ (https://www.linkedin.com/in/brianmcmillan01/) Website: https://www.minimumviablearchitecture.com/ (https://www.minimumviablearchitecture.com/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Somewhat short episode. Scott emphasizes that you aren't the only one facing challenges with data mesh. You aren't behind the curve, you are ahead of it - data mesh is bleeding edge. But with a bleeding edge, there is some blood... Also, you need to think about how others' implementations works for them. Trying to copy-paste another organization's implementation is going to lead to a failed implementation in yours. Take the learnings away and think about how to apply them to your organization. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Sheetal Pratik, Director of Engineering, Data Integration at Adidas. If Sheetal's name sounds familiar to data mesh officianados, she presented at the Data Mesh Learning meetup in August 2021. Sheetal is passionate about giving companies the permission AND a workable plan for getting started with data mesh. She covered a wide range of things regarding getting starting but a few really stood out: Don't try to tackle tomorrow's challenges today Break your implementation into phases: development, adoption, and scaling Start your initial data mesh MVP with a simple data product with a simple schema - your goal is to develop the "muscles" around creating and deploying data products rather than shooting for a high-value product first Keep to a reasonable budget and prove viability and value
Sheetal also covered how much you really have to have in place to create and evaluate your MVP. There will always be evolution and change and your organization has to be ready for that. That can be frightening or inspirational. Sheetal chooses it as inspirational - it gives you the freedom to move quickly as long as the organization understands that things will change in the future. She wrapped up with saying that data mesh shouldn't be scary, you should be excited about this journey and what it can mean for your organization. Get an MVP out the door, it will take time, don't get ahead of yourself. Data mesh success can happen if you let it. Sheetal's (and Divya and Madhu from Thoughtworks) Data Mesh Learning Meetup presentation: https://www.youtube.com/watch?v=5btUaLPdaNk (https://www.youtube.com/watch?v=5btUaLPdaNk) Sheetal's post re data mesh and using DataHub: https://blog.datahubproject.io/enabling-data-discovery-in-a-data-mesh-the-saxo-journey-451b06969c8f (https://blog.datahubproject.io/enabling-data-discovery-in-a-data-mesh-the-saxo-journey-451b06969c8f) Sheetal's LinkedIn: https://www.linkedin.com/in/sheetalpratik/ (https://www.linkedin.com/in/sheetalpratik/) Data Governance using Data Mesh paper: https://easychair.org/publications/preprint/qZ3m (https://easychair.org/publications/preprint/qZ3m) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Don't forget to catch Dr. Abadi at https://starburst.io/info/datanova2022?utm_source=DataMeshRadio (Datanova - the Data Mesh Summit on Feb 9-10th). Thanks to Starburst for sponsoring the transcripts for Data Mesh Radio, check out the transcript https://docs.google.com/document/d/1CkXYjUNMRLl1okJJFmEbzTgTybAAFa-_VC4sOvnFH8s/edit?usp=sharing (here). And check out Starburst's other free data mesh resources https://www.starburst.io/info/data-mesh-resource-center?utm_source=DataMeshRadio (here). In this episode, Scott interviewed Dr. Daniel Abadi, the Darnell-Kanal Computer Science Professor at the University of Maryland with a focus on scalable data management research. Dr. Abadi will be presenting next week at the Data Mesh Summit on Data Fabric and Data Mesh alongside Zhamak and Sanjeev Mohan. This was a pretty wide ranging and free wheeling conversation about data virtualization in general and how it can be used in data mesh. Both agreed that there are many places where data virtualization can play in data mesh, whether in extracting information from operational systems, stitching together a data product once data processing has been done, or at the mesh experience plane re combining data across multiple data products. Dr. Abadi specifically mentions something like a query fabric that makes use of a data virtualization approach, not just tools that only do data virtualization. There is a natural side effect of having multiple different technologies in use - when you give the domains the ability to use what they choose, the difficulty of combining data from multiple sources needs to be solved. There is always a balance between how much you just copy data and how much you can access in the source system and data virtualization can give a few more options rather than all or nothing. As data virtualization has been around as a concept for 30+ years, there is a lot of baggage with the term but Dr. Abadi sees there being recent advancements that mean more people should take a second look at where they can be useful. But warns to do your homework and really think through whether they fit your use case. A query fabric can make your user experience much more pleasant. Trying to create data products entirely within a data virtualization platform probably won't be, at least according to Scott. Additional topics included retransmitting or reprocessing data, versioning, the importance of denormalizing data for analytics and how that plays with data virtualization, and much more. It is a really fascinating deep dive into the history of computing and how it impacts what we are trying to do today. Dr. Abadi's blog post on data federalization and data virtualization: https://blog.starburst.io/data-federation-and-data-virtualization-never-worked-in-the-past-but-now-its-different (https://blog.starburst.io/data-federation-and-data-virtualization-never-worked-in-the-past-but-now-its-different) Dr. Abadi's contact info: LinkedIn: Twitter: @daniel_abadi / https://twitter.com/daniel_abadi (https://twitter.com/daniel_abadi) Starburst blog posts: https://blog.starburst.io/author/daniel-abadi (https://blog.starburst.io/author/daniel-abadi) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) Music used this episode created by Lesfm (intro includes slight edits by Scott...
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Don't forget to catch Colleen at https://starburst.io/info/datanova2022?utm_source=DataMeshRadio (Datanova - the Data Mesh Summit on Feb 9-10th). Thanks to Starburst for sponsoring the transcripts for Data Mesh Radio, check out the transcript https://docs.google.com/document/d/1-Jd5f79HfIXs_Sl93Z79Sdn-wbgvVe6pE_LMkywOOaw/edit?usp=sharing (here). And check out Starburst's other free data mesh resources https://www.starburst.io/info/data-mesh-resource-center?utm_source=DataMeshRadio (here). In this episode, Scott interviews Dr. Colleen Tartow, Director of Engineering at Starburst. They chatted about the fun and usefulness of food-related analogies to data mesh - you want it to work like a brunch buffet as that is basically a perfect Saturday in Colleen's eyes. And Scott shared the concept of a grocery store - intentional food preparation with different degrees of ingredient and meal preparedness. Colleen and Scott then covered the "Modern Data Stack" and its relevance to data mesh - in Colleen's eyes, the Modern Data Stack isn't all that modern, it is just the same old paradigm of data teams trying to do their best to work with the output of application/operational data stores and systems that aren't designed with the data in mind. Scott somewhat agreed with his own spin. Colleen shared her 4 S paradigm re doing data well: speed, simplicity, scalability, and SQL. Yes, slightly Starburst self-serving (4 more S words!) but still relevant and interesting. They wrapped up with trying to figure out how much of data mesh can companies get away with not doing, which seems to be the topic on everyone's mind. More questions than answers there... Colleen's contact info: Email: colleen at starburst.io LinkedIn: https://www.linkedin.com/in/colleen-tartow-phd/ (https://www.linkedin.com/in/colleen-tartow-phd/) Twitter: @CTartow / https://twitter.com/CTartow (https://twitter.com/CTartow) Colleen's blog: https://thesequel.substack.com/ (https://thesequel.substack.com/) The Data Buffet blog post: https://thesequel.substack.com/p/the-endless-data-buffet (https://thesequel.substack.com/p/the-endless-data-buffet) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) Music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript (https://docs.google.com/document/d/1ZX73_DajFy9hSKUOo7JBKFpWModLSAniPDnLkbEqLyw/edit?usp=sharing (link)) courtesy of Starburst; check out their other data mesh resources https://www.starburst.io/info/data-mesh-resource-center?utm_source=DataMeshRadio (here) Part of Starburst's Datanova Data Mesh Summit takeover week. Max and Arif's upcoming (Feb 9th) Datanova/Starburst Data Mesh Summit Presentation: https://starburst.io/info/datanova2022?utm_source=DataMeshRadio (https://www.starburst.io/info/datanova2022/) In this episode, Scott interviews one of the most prolific content producers in data mesh, Max Schultze, Data Engineering Manager at Fashion E-Tailer Zalando. Max shares a LOT of very valuable advice while reflecting on Zalando's data mesh journey so far two years in. Max recommends starting with empathy building and knowledge sharing but looking for ways to scale versus only 1:1 or small group conversations. He had a clever way to force the hands of data consumers to speak to data producers via an https://blog.octo.com/en/how-to-deal-with-an-inverse-conway-maneuver-a-talk-by-romain-vailleux-at-duck-conf-2021/ (Inverse Conway Maneuver) that might work for your org too. Overall, 5 main points/themes emerged: Data mesh is a journey Empathy is crucial - have it for those you are working with and work towards building it between teams Technology is not the most important aspect of an implementation - no matter how cool it might sound or be to focus on it Start from knowledge sharing, getting people to understand each others' roles and contexts (see empathy!) Try not to get ahead of yourself
Highly recommend giving this one a listen, maybe twice to pick up the nuggets in there. Max's contact info and related links: LinkedIn: https://www.linkedin.com/in/max-schultze-b11996110/ (https://www.linkedin.com/in/max-schultze-b11996110/) Twitter: @mcs1408 / https://twitter.com/mcs1408 (https://twitter.com/mcs1408) Presentations: Similar to content of the book (Data Innovation Summit): https://www.youtube.com/watch?v=rqYFqtztWi4 (https://www.youtube.com/watch?v=rqYFqtztWi4) Zalando's Story (Spark AI Summit): https://www.youtube.com/watch?v=UrM8yCjmzzw (https://www.youtube.com/watch?v=UrM8yCjmzzw) Max and Arif's 'Data Mesh in Practice' book links: O'Reilly: https://www.oreilly.com/library/view/data-mesh-in/9781098108502/ (https://www.oreilly.com/library/view/data-mesh-in/9781098108502/) Starburst (free gated download): https://www.starburst.io/info/data-mesh-in-practice-ebook?utm_source=DataMeshRadio (https://www.starburst.io/info/data-mesh-in-practice-ebook/) Max and Arif's upcoming training (Feb 14th): https://www.oreilly.com/live-events/data-mesh-in-practice/0636920508816/0636920068685/ (https://www.oreilly.com/live-events/data-mesh-in-practice/0636920508816/0636920068685/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"):...
Very quick one! The Data Mesh Summit (aka Datanova) is taking over Data Mesh Radio for the week! Sign up for an awesome two-day event (Feb 9th and 10th) https://starburst.io/info/datanova2022?utm_source=DataMeshRadio (here). We will have 4 (scheduling willing) amazing guests from the Data Mesh Summit on over the next week: Max Schultze (Zalando), Dr. Colleen Tartow (Starburst), Dr. Daniel Abadi (University of Maryland), and Dr. Teresa Tung (Accenture). In exchange, Starburst is doing a beta, sponsoring transcripts. So please let them know you want more transcripts and again, use the link to sign up to show your support! Again, https://starburst.io/info/datanova2022?utm_source=DataMeshRadio (sign up here). To get Max Schultze + Dr. Arif Wider's 'Data Mesh In Practice' book, https://www.starburst.io/info/data-mesh-in-practice-ebook?utm_source=DataMeshRadio (click through here)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Scott interviews Piethein Strengholt, Senior Cloud Solution Architect at Microsoft and author of the O'Reilly book Data Management at Scale. Piethein shares his tips and tricks for how to approach Domain Driven Design (DDD) for data. There are puts and takes to each approach so unfortunately for those looking for an easy button, there is a lot to consider. When starting, Piethein recommends looking at your applications and deciding if there is a logical mapping to a single domain or if the application is shared across domains. As you learn more about DDD you can start to approach your domains from multiple other angles to find the best solution for your org. There is a ton of really great advice, too much to sum up well here. Scott highly recommends reading https://towardsdatascience.com/data-domains-where-do-i-start-a6d52fef95d1 (this article) on data domains by Piethein before jumping in to the podcast episode. Mentioned articles and books: Data Management at Scale book: https://www.oreilly.com/library/view/data-management-at/9781492054771/ (https://www.oreilly.com/library/view/data-management-at/9781492054771/) Data Domains — Where do I start? https://towardsdatascience.com/data-domains-where-do-i-start-a6d52fef95d1 (https://towardsdatascience.com/data-domains-where-do-i-start-a6d52fef95d1) Implementing Data Mesh on Azure: https://towardsdatascience.com/implementing-data-mesh-on-azure-c01ee94306cd (https://towardsdatascience.com/implementing-data-mesh-on-azure-c01ee94306cd) Data Domains and Data Products: https://towardsdatascience.com/data-domains-and-data-products-64cc9d28283e (https://towardsdatascience.com/data-domains-and-data-products-64cc9d28283e) "The Blue Book" AKA Domain-Driven Design: Tackling Complexity in the Heart of Software: https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/ (https://www.oreilly.com/library/view/domain-driven-design-tackling/0321125215/) Find Piethein online: LinkedIn: https://www.linkedin.com/in/pietheinstrengholt/ (https://www.linkedin.com/in/pietheinstrengholt/) Twitter: @phstrengholt / https://twitter.com/phstrengholt (https://twitter.com/phstrengholt) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Scott interviews Bente Busch, Director of Teams Service Platform - essentially the applications platformm design system, and data platform - at Norwegian government entity NAV (Norwegian Labor and Welfare Department). Bente talked about some really interesting aspects of NAV's data mesh implementation including: The unique challenges of having producers more excited to share the data than consumers demanding more data Changing the culture from doing "projects" to building products How the on-prem enterprise data warehouse just didn't offer them the same agility they need Where they are in their data mesh journey so far And much more
It's a very interesting look into how data mesh could be applied in government with eventual plans to share information outside of the organization. One very interesting insight is that including those building the application platform in the data mesh self-serve platform build out has been a big win - they are already familiar with application developer workflows and how they think so they are better able to anticipate needs when it comes to building the data platform. Bente's contact info: https://www.linkedin.com/in/bentebusch/ (https://www.linkedin.com/in/bentebusch/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Molly Vorwerck, head of content and communications at data observability vendor Monte Carlo. Scott asked Molly to be on as she is a great community member in general and as Monte Carlo is well known in the data space for putting out high quality content on hot-button issues. The idea is to take how Molly and team get their ideas and apply that process to figuring out major pain points internally and creating a cohesive strategy to at least start discussing them. Molly recommends to constantly be interviewing stakeholders. She talks about interviewing people from multiple sides of a challenge, e.g. not just the data consumers but the data producers and the data engineering teams re data challenges. She tries to give people the space to tell their story and asks open-ended questions to truly get their perspective, not arrive at a pre-specified answer. Molly talks about ways to make the other person feel valued by active listening and making the conversation mutually beneficial. It may be a person you want to interview again so building the relationship is crucial, not just extracting info in a one-time manner. Scott and Molly dig a bit into the idea of blameless post mortems and how valuable they can be for doing data mesh, especially to figure out what happened to cause data downtime and how to prevent the same issue in the future. Molly's LinkedIn: https://www.linkedin.com/in/vorwerck/ (https://www.linkedin.com/in/vorwerck/) Monte Carlo's blog: https://www.montecarlodata.com/blog/ (https://www.montecarlodata.com/blog/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this somewhat controversial episode, part 1 of 2, Scott covers topics re is data mesh right for your organization. This should only be used as a jumping off point for discussions. The 4 segment titles are: Data quality challenges don't necessarily mean it's time for a data mesh Data Mesh Lite Building on a solid foundation How 'bout now, how 'bout right now? ("Patience you must have, my young Padawan") Take it with a grain of salt! Mentioned content links: https://www.thoughtworks.com/about-us/events/webinars/core-principles-of-data-mesh/lessons-from-the-trenches-in-data-mesh (Webinar with Zhamak and Sina Jahan - lessons from the trenches with data mesh) https://barryoreilly.com/explore/podcast/decentralizing-data-zhamak-dehghani/ (Zhamak podcast interview with Barry O'Reilly) https://www.youtube.com/watch?v=-POiudR2_R0 (Flexport Data Mesh Learning meetup) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Chad Sanderson, Head of Product: Data Platform at Convoy. This episode is part of our continuing series on data contracts and related topics. Chad covers a lot of the challenges relative to data quality, both in maintaining quality and in the challenges poor quality data can cause a company that is heavily reliant on data. Chad also shares his tale of trying to implement data mesh at Convoy via a large-scale https://blog.octo.com/en/how-to-deal-with-an-inverse-conway-maneuver-a-talk-by-romain-vailleux-at-duck-conf-2021/ (inverse Conway Maneuver). Chad covered 4 categories of data quality pain, which he calls the "4 Horseman of Data Quality" in https://www.linkedin.com/posts/chad-sanderson_data-metadata-analytics-activity-6886713385065566208-fNNy (this post): Omission - metadata is missing; no tool out today that solves the omission problem, so users have to bounce between too many tools to try to figure out data specifics like where it came from, the specific meaning, what it's trying to convey, etc. Waste: growth of unused, unmaintained, or duplicated data; waste happens when the cost of creating new data is less than using something already created Divergence: the growing divide between what's going on in "the real world" and what's happening in your data warehouse; your business logic, unless it is constantly maintained and updated, starts to diverge from what is happening to your business so what you show on dashboards and reports no longer matches business reality Downtime: periods of time where the data is missing, wrong, late, etc.; traditionally what most people think of regarding data quality issues Chad's contact info: LinkedIn: https://www.linkedin.com/in/chad-sanderson/ (https://www.linkedin.com/in/chad-sanderson/) csanderson.data at gmail.com Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Quick "episode" EXPLICIT PERMISSION: skip episodes you don't think will be useful - that is why the BLUFs exist In general, we will publish interview episodes on Tuesday and Friday morning. If there is a mesh musing, it will be published Wednesday morning Starting "takeover weeks" with conferences. Look for one soon. If you have a conference that should do one, get in touch: scott at datastax.com Be a guest. Seriously. This is not an "experts only" podcast! It's a learning out loud one. Data contracts deep dive continues. 3-4 more episodes lined up by end of Feb. Next 2 deep dives are 1) Domain Driven Design for Data + Data-Centric Application Development and 2) Your Data Mesh Proof of Concept I need guests and I need topics. Stop being shy. Do not expect an explicit invitation, step up. New podcast from me coming in early to mid-Feb. This Week I Learned in Data. Be a contributor sharing what you learned (not looking for contributions every week!) or suggest contributors.
P.S. Please do rate and review the podcast. It honestly really does help for visibility. I don't care about the ratings, I just want to make sure if it's a resource, people can find it! Scott's contact info: scott at datastax.com If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode Scott interviewed Benn Stancil, co-founder and Chief Analytics Officer at Mode about his emerging concept of global definitions that he is calling an "entity" - essentially, how are people defining terms like customer and what does that mean in each instance so people are on the same page. Data mesh also requires some clear collaboration on definitions, whether centralized or decentralized, in the federated governance pillar. This interview was generated from a Twitter conversation https://twitter.com/data_mesh_learn/status/1474622183116730372 (here). Benn then went into some details about how he views the importance of having visibility into how data "breaks" so it is much easier to identify and fix and that limiting custom integration is crucial so when you fix at the source, it properly propagates downstream. They wrapped up discussing the need to make data producers' lives easier while simultaneously doing the same with data consumers. Overall, there are some agreements and disagreements and Scott came out thinking more about what are the real causes of pain that would make a full journey to data mesh make sense. It's a good episode to see some of the challenges people are trying to tackle outside of the data mesh community. Links from show: Benn's Twitter: @bennstancil / https://twitter.com/bennstancil (https://twitter.com/bennstancil) Benn's Substack: https://benn.substack.com/ (https://benn.substack.com/) Benn's post on entities: https://benn.substack.com/p/metadata-money-corporation (https://benn.substack.com/p/metadata-money-corporation) Starship Technologies Data Mesh Post: https://medium.com/starshiptechnologies/dodging-the-data-bottleneck-data-mesh-at-starship-5925a2de45e6 (https://medium.com/starshiptechnologies/dodging-the-data-bottleneck-data-mesh-at-starship-5925a2de45e6) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Juan Sequeda, Principal Scientist at data.world and co-host of the Catalog and Cocktails podcast. They discussed Juan's knowledge first approach: putting the meaning and value of the data first instead of focusing on the amount of data we are handling/producing. Knowledge first has 3 components, 1) context, 2) people, and 3) relationships. Juan is a big proponent of knowledge graphs and the relationships side is one many people miss. Juan also gave some thoughts on what his approach to data mesh hinges on: treating data as a product and finding a balance between centralization and decentralization for all the aspects of building out an implementation. Juan mentioned Intuit's approach of fixed, flexible/extensible, or customizable as a good general tool and to look for (and embrace) what he calls intellectual friction. Lastly, Juan and Scott talked about the general drive to reduce toil, of reinventing the wheel re data interoperability and standard schemas in data mesh. Juan points to a lot of existing research and standards - e.g. RDF, OWL, and many more (see below) - as a starting point. Juan's contact info and related links: Email: juan at data.world Twitter: @juansequeda / https://twitter.com/juansequeda (https://twitter.com/juansequeda) LinkedIn: https://www.linkedin.com/in/juansequeda/ (https://www.linkedin.com/in/juansequeda/) Catalog & Cocktails Podcast: https://data.world/podcasts/ (https://data.world/podcasts/) Juan's post about Zhamak's appearance on the Data Engineering Podcast: https://www.linkedin.com/pulse/my-takeaways-data-engineering-podcast-episode-mesh-zhamak-sequeda/ (https://www.linkedin.com/pulse/my-takeaways-data-engineering-podcast-episode-mesh-zhamak-sequeda/) Juan's post about knowledge first: https://www.linkedin.com/feed/update/urn:li:activity:6884179569277059072/ (https://www.linkedin.com/feed/update/urn:li:activity:6884179569277059072/) Standards related links: Dublin Core Metadata Initiative: https://dublincore.org/ (https://dublincore.org/) RDF (Resoruce Description Framework): https://www.w3.org/2001/sw/wiki/RDF (https://www.w3.org/2001/sw/wiki/RDF) OWL (Web Ontology Language): https://www.w3.org/OWL/ (https://www.w3.org/OWL/) PROV-O: The PROV Ontology: https://www.w3.org/TR/prov-o/ (https://www.w3.org/TR/prov-o/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode created by Lesfm (intro includes slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (https://pixabay.com/users/lesfm-22579021/) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Peter Hanssens, Founder and Solutions Architect at cloud/serverless consulting company Cloud Shuttle. Peter also runs a few large data engineering focused communities (meetups, Slack, conferences) in Australia. Peter had reached out about how can startups and SMEs (small to medium enterprises) get the benefits of data mesh without the costs of building a solution fit for a 10,000+ employee company. Peter does a great job seeking information on behalf of his constituents :) We covered a number of topics including: The needs for cultural change as technology will only get you so far The beauty of pay-per-use solutions for startups/SMEs - and where there are still gaps in the market The challenges of not having a large pool of data engineers to build and manage a platform - or to help domains model their data The benefits of centralization until it starts to cause bottlenecks The importance of tracking lineage for upstream producers/domains - to see who is using your data and why/how And much, much more
I think you will really enjoy Peter's perspectives and there are some useful conclusions if not a perfect blueprint for startups. Peter's contact info and relevant links: Email: peter at cloudshuttle.com.au LinkedIn: https://www.linkedin.com/in/peterhanssens/ (https://www.linkedin.com/in/peterhanssens/) Twitter: @petehanssens / https://twitter.com/petehanssens (https://twitter.com/petehanssens) Cloud Shuttle website: https://www.cloudshuttle.com.au/ (https://www.cloudshuttle.com.au/) Sydney Data Engineering Community: https://sydneydataengineers.github.io/ (https://sydneydataengineers.github.io/) DataEngBytes Conference: https://dataengconf.com.au/ (https://dataengconf.com.au/) Article mentioned by Francois Nguyen (CIO of L'Oreal): https://francois-nguyen.blog/2021/03/07/towards-a-data-mesh-part-1-data-domains-and-teams-topologies/ (https://francois-nguyen.blog/2021/03/07/towards-a-data-mesh-part-1-data-domains-and-teams-topologies/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Dan DeMers, Co-Founder and CEO of Cinchy, a dataware platform / data fabric provider. Dan shares his thoughts on why data-centric application design is the best way to deal with the challenges of applications and analytics needing the same data for different purposes - the current approach is to let the application schema evolve whenever and however necessary and the underlying data applications suffer. Dan's view is "share access to data, not copies." The interview ties loosely with previous interviews re schema/data contracts and data testing as Dan argues using a dataware approach will prevent the issues of an evolving application schema breaking data consumption downstream. It isn't all rosy as this will take a fair bit of work for an organization to move to this approach. Food for thought and the first of a series of interviews re DDD (domain-driven design) for data and data-centric application development. Dan's LinkedIn: https://www.linkedin.com/in/demersdan/ (https://www.linkedin.com/in/demersdan/) Cloud Information Model Standard: https://cloudinformationmodel.org/ (https://cloudinformationmodel.org/) Cinchy website: https://cinchy.com/ (https://cinchy.com/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott discusses three concepts that are at best a concern. Consider it a late Grinch-inspired present for Xmas :) Reverse ETL meets a real need for analytical data being pushed into CRM, marketing, and other similar systems. But treating another pipeline as a first order concern is fraught with the same issues of most similar data pipeline treatment: who owns it, how does it evolve, who is observing/monitoring it for uptime and semantic drift, etc.? Should we look to create data products on the mesh to serve those needs instead of another ETL tool? Some organizations implementing data mesh are forcing their domains to consume any analytics from their own data products on the mesh. The good of this is that it aligns the domain with creating a high-quality data products. But will those data products be designed to fit the general organizational needs or specifically the domain's needs? There is an emerging push for software engineers to also own the data modeling. To get to a place where this is even feasible, don't we need far better abstractions for domains to _do_ the data modeling? And will this overload software engineers that are already dealing with a metric buttload of technologies and requirements already? Where would a junior engineer fit in that kind of organization? Does this mean more software engineers on the team -> 2 pizza teams now 3? 4? 5? 10? Maybe we pump the brakes on this for now? Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Jesse Paquette, Chief Science Officer and Co-founder at Tag.bio - a data platform vendor in the life sciences space, and Scott dive a bit deeper into data quality in general, especially data testing and versioning. You can see the LinkedIn post that sparked this discussion https://www.linkedin.com/feed/update/urn:li:activity:6881406997678424064?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A6881406997678424064%2C6881548651487907840%29 (here) Jesse recommends a number of things to ensure data quality, especially data testing and versioning. This includes versioning of 1) the code used to create the data (generally the ETL code), 2 the schema, 3) the business logic layer, and 4) timestamping / temporality based versioning. Jesse's general calls to action are 1) make data testing frameworks so testing is much less tedious and time consuming; 2) work with stakeholders to gain trust in the data and then continue the dialogue to keep said trust; and 3) create schema/domain model blueprints so that domains have a starting point - whether they use it is irrelevant but shortening the path to a working domain model is crucial. Jesse's contact info: Email: jesse at tag.bio LinkedIn: https://www.linkedin.com/in/jessepaquette/ (https://www.linkedin.com/in/jessepaquette/) Twitter: @bzdyelnik / https://twitter.com/bzdyelnik (https://twitter.com/bzdyelnik) Website: https://tag.bio/ (https://tag.bio/) Tag.bio vendor interview for Data Mesh Learning: https://www.youtube.com/watch?v=acQADu7ttqQ (https://www.youtube.com/watch?v=acQADu7ttqQ) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Abhi Sivasailam, Head of Growth and Analytics at unicorn startup Flexport, and Scott deep dive into data contracts. They covered a LOT of ground including: What is a data contract and how does it relate to API contracts and schema contracts Why data contracts are so crucial to treating data like a product and keeping data usable by consumers Abhi's rules for a viable data contract The importance of the analytics engineer to overall data usefulness, especially re domain data ownership in data mesh Why of the socio of the socio-technical approach to data ownership is the most crucial aspect The lack of proper tooling to monitor and execute data contracts How to minimize disruptive changes to downstream data consumers The issues with domains sharing their data as it is persisted/stored in the database instead of sharing the context via the domain model Much, much more
Abhi's contact info: Twitter: https://twitter.com/_abhisivasailam (https://twitter.com/_abhisivasailam) LinkedIn: https://www.linkedin.com/in/abhi-sivasailam/ (https://www.linkedin.com/in/abhi-sivasailam/) Abhi's Data Mesh Learning Meetup: https://www.youtube.com/watch?v=-POiudR2_R0 (https://www.youtube.com/watch?v=-POiudR2_R0) Debezium blog post mentioned by Abhi: https://debezium.io/blog/2019/02/19/reliable-microservices-data-exchange-with-the-outbox-pattern/ (https://debezium.io/blog/2019/02/19/reliable-microservices-data-exchange-with-the-outbox-pattern/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Matthew Darwin, Principal Data Engineer at Slalom Consulting, and Scott cover a wide range of topics including: Re an article Matt had posted on data platform re-use and shifting data ownership left: You can re-use the technologies you already know and love when building a data platform - there is literally no good reason to toss those out the window If your existing setup enables domains to easily transform, serve, and store their data, sure use your existing configuration; if not, there will need to be changes Your data platform will need to evolve and it is okay to start with a bit of an underwhelming data platform
Other nuggets and interesting topics: Insights from direct client engagements doing data mesh What makes for a good data mesh PoC The usefulness of data product blueprints How data mesh is still bleeding edge and is therefore not for everyone Slowing down to move faster / the long-term negatives of always looking for quick wins Why you can't just expose your operational data model as a data product The importance of data product interoperability - even at the PoC phase How crucial the organizational aspects of data mesh really are Much more
Matt's contact info: LinkedIn: https://www.linkedin.com/in/matthewdarwindba/ (https://www.linkedin.com/in/matthewdarwindba/) Twitter: @EvoDBA / https://twitter.com/EvoDBA (https://twitter.com/EvoDBA) Matt's post on platform reuse: https://medium.com/slalom-data-analytics/data-mesh-is-the-argument-a-strawman-3cffaf55ce5e (https://medium.com/slalom-data-analytics/data-mesh-is-the-argument-a-strawman-3cffaf55ce5e) Matt's LinkedIn poll on testing data pipelines: https://www.linkedin.com/feed/update/urn%3Ali%3Aactivity%3A6877216459458719744/ (https://www.linkedin.com/feed/update/urn%3Ali%3Aactivity%3A6877216459458719744/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Olivier Wulveryck, Senior Consultant at consulting company OCTO Technology, and Scott discuss all things schema contracts. The episode will probably leave you with more questions than answers at this point as the concept/practice is still emerging. But you will get a good sense for how to do it right - and how not to do it too! - from the conversation. Olivier's contact info LinkedIn: https://www.linkedin.com/in/olivierwulveryck/ (https://www.linkedin.com/in/olivierwulveryck/) Email: olivier.wulveryck at octo.com Twitter: @owulveryck / https://twitter.com/owulveryck (https://twitter.com/owulveryck) Olivier's post on schema contracts: https://blog.octo.com/en/pov-a-streaming-communication-platform-for-the-data-mesh/ (https://blog.octo.com/en/pov-a-streaming-communication-platform-for-the-data-mesh/) (mentioned in BLUF) Sofia Tania presentation on Data Mesh Testing: https://www.youtube.com/watch?v=stNZQESndAA (https://www.youtube.com/watch?v=stNZQESndAA) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Wannes Rosiers, CTO at Golazo Group, who previously led DPG Media's data mesh transition. They discussed a lot of different sub-topics around data products in data mesh including: Wannes' framework re his 3 types of data products The importance of a specific purpose for each data product Data product evolution from single purpose to multi-purpose The owner roles/functions related to data products in data mesh How to get moving with an initial data mesh implementation Data product interoperability and global definitions and identifiers
Wannes' LinkedIn: https://www.linkedin.com/in/wannes-rosiers/ (https://www.linkedin.com/in/wannes-rosiers/) Wannes' data mesh related content: Meetup presentation: https://www.youtube.com/watch?v=l_5fkpweQwM (https://www.youtube.com/watch?v=l_5fkpweQwM) Domain Driven Design for Data at DPG Media: https://dpgmedia-engineering.medium.com/ddd-data-area-at-dpg-media-f0130e4d9766 (https://dpgmedia-engineering.medium.com/ddd-data-area-at-dpg-media-f0130e4d9766) A specific run down of data mesh at DPG: https://dpgmedia-engineering.medium.com/data-mesh-at-dpg-media-dfebdd612087 (https://dpgmedia-engineering.medium.com/data-mesh-at-dpg-media-dfebdd612087) A co-authored post with Snowflake: https://levelup.gitconnected.com/data-mesh-a-self-service-infrastructure-at-dpg-media-with-snowflake-566f108a98db (https://levelup.gitconnected.com/data-mesh-a-self-service-infrastructure-at-dpg-media-with-snowflake-566f108a98db) Short video describing DPG's overall data journey, not just data mesh: https://itexecutive.nl/leveraging-the-power-of-data-technology/data-strategy/van-goed-op-weg-zijn-met-data-naar-data-rock-star-in-twee-jaar/ (https://itexecutive.nl/leveraging-the-power-of-data-technology/data-strategy/van-goed-op-weg-zijn-met-data-naar-data-rock-star-in-twee-jaar/) Golazo Careers Page: https://www.golazo.com/jobs/ (https://www.golazo.com/jobs/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Host Scott Hirleman gives some helpful context around 3 separate data mesh topics: Why Data Virtualization is not a good fit for creating and managing your data products in data mesh A reply to a wonderful article by Matthew Darwin (link below) asking if we can keep our existing data platform / architecture and only decentralize our data ownership - surprise, surprise, Scott argues that it isn't that simple Introducing the concept of a speculative data product to solicit feedback in a bit of internal data product marketing
Matthew Darwin article: https://medium.com/slalom-data-analytics/data-mesh-is-the-argument-a-strawman-3cffaf55ce5e (https://medium.com/slalom-data-analytics/data-mesh-is-the-argument-a-strawman-3cffaf55ce5e) Intuit article referenced in spec data product segment: https://medium.com/intuit-engineering/intuits-data-mesh-strategy-778e3edaa017 (https://medium.com/intuit-engineering/intuits-data-mesh-strategy-778e3edaa017) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Another short one, basically asking people to be on the show and to provide feedback. Let's work together as a broader community to provide content that is useful and helpful. So please let Scott know if you want to be a guest or if you have feedback about existing content or what content you want to see. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) In this episode, Scott interviews Paolo Platter, CTO at Agile Lab, about data product flow or 1) how to identify what data products you already have that you need to migrate to the data mesh and 2) what data products you need to create to further the business processes that will move your organization forward. Paolo has been trialing this methodology on a number of Agile Lab's consulting customers. Paolo's contact info and his great blog post on this topic: Email: paolo.platter AT agilelab.it LinkedIn: https://www.linkedin.com/in/paoloplatter/ (https://www.linkedin.com/in/paoloplatter/) Blog post: https://www.agilelab.it/how-to-identify-data-products-welcome-data-product-flow/ (https://www.agilelab.it/how-to-identify-data-products-welcome-data-product-flow/) Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Host Scott Hirleman gives some helpful context around how to think about data mesh including The onion/parfait model: a layered look at data mesh - core is the data products, 2nd layer is the platform and enablement of domains, and 3rd is the data culture aspects Supermarket/Grocery Store analogy: rather than having a bunch of raw ingredients or having to forage/hunt for food, supermarkets have sourced the food and included a ton of metadata around it for you. They have organized it into domains (departments) even. It makes for a better and more reliable shopping experience - are you looking for a data shopping experience? :D Maximizing context AND maximizing usability: historically, we have had to choose between context or usability - accessibility, understandability, quality, interoperability, etc. - at the broader organization level. Data mesh is attempting to do both and then also lowering the bar to using the data for consumers. Stretching or Training analogy: athletes prepare ahead of time to be able to compete and react on the field as circumstances change. If we prepare the data in such a way that we are prepared for anything, not just for what we thought was going to happen, how might that lead to better outcomes?
Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)
https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) An episode telling you about what the podcast will be. Data Mesh Radio is hosted by Scott Hirleman. If you want to connect with Scott, reach out to him at community at datameshlearning.com or on LinkedIn: https://www.linkedin.com/in/scotthirleman/ (https://www.linkedin.com/in/scotthirleman/) If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/ (https://datameshlearning.com/community/) If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see https://docs.google.com/document/d/1WkXLhSH7mnbjfTChD0uuYeIF5Tj0UBLUP4Jvl20Ym10/edit?usp=sharing (here) All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): https://pixabay.com/users/lesfm-22579021/ (Lesfm), https://pixabay.com/users/mondayhopes-22948862/?tab=audio (MondayHopes), https://pixabay.com/users/sergequadrado-24990007/ (SergeQuadrado), and/or https://pixabay.com/users/nevesf-5724572/ (nevesf) Data Mesh Radio is brought to you as a community resource by DataStax. Check out their high-scale, multi-region database offering (w/ lots of great APIs) and use code DAAP500 for a free $500 credit (apply under "add payment"): https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB)