https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1WRNYqsgfM-S4jt4ExOxnW8kuhvcaX1sh2CIGJFzvlHE/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Immanuel Schweizer, the Data Officer for EMD Electronics. Some interesting thoughts and questions from the conversation: Good governance starts at data collection - what are ethical and compliant ways to collect data from the beginning? This points to intentionality around data use stretching into the application - what should you collect that might not be part of the day-to-day application function but that might might lead to generating insights that will be used to generate a better user experience? And what are the ethical concerns? Should we initially create data products to serve specific use cases or should we focus on sharing data first and then shaping what people consume most into data products? EMD is approaching data products from a different angle than most, using the second approach. When looking at data mesh, should you start with the high data maturity teams or work to pull everyone up to at least a decent baseline maturity level? If you work with the most mature teams, will their challenges really be applicable to the not-so-mature domains? Can you find good reuse patterns to scale your mesh implementation? Domain owners are much more willing to share data if they understand use cases for how their data will be used and maintain control to prevent misuse. Reluctance comes from an incomplete picture causing concerns - the more visibility into how data can be and is being used, the more willing domain owners are to share. But understanding your end-to-end data supply chain is tough, especially to start. How do you evaluate when to spend the time with a domain to get them data mesh ready? If you need a high value use case to justify spending time with that domain, are you leaving many domains behind? This ties to #2 and #3. Set your target picture but be ready to adjust your target picture along the way. The world is ever changing, don't lock in to an expected target outcome. Good data governance is about speeding up 1) access to and 2) usage of data. EMD launched a data literacy program where the employees spend the majority of a 10 week timeframe learning about data and how to make use of it. For Immanuel, making things tangible relative to data makes people much less hesitant to explore and use data. You should make using data a "part of the job" so it is tracked and part of the review process. Otherwise, you are missing out on a key incentive to leverage data. How many people in your organization wish they could be leveraging data more often to make decisions? What's holding them back? Is it tooling, knowledge, incentivization, access, etc.? How can we democratize insights? So much of insight generation is one-off, how do we make that scalable, shareable, and repeatable?
Per Immanuel, EMD's data mesh journey is not that typical in that they are still getting their arms around centralizing data in a constructive way. It was previously locked away in the domains. So, they are starting their data mesh - or decentralization - journey by centralizing data in a certain sense. Wannes Rosiers mentioned this at DPG