https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1hNNILAXWGAZ8xEyjUxe7oiek88R0C_PcFIKo3EBgRfQ/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Jay Sen, Data Platforms & Domain Expert/Builder and OSS Committer. While Jay currently works at PayPal, he was only representing his own view points. Some key takeaways/thoughts from Jay's view: When you get to a certain scale, any central team should focus on, as Jay said, "Empower people, don't try do their jobs." That's how you build towards scale and maintain flexibility - your centralized team likely won't become a bottleneck if they aren't making decisions on behalf of other teams. To actually empower other teams, dig into the actual business need and work backwards to a solution that can solve that. If there is a solution already in place that isn't working any more, look to find ways to augment that rather than trying to replace or reinvent the wheel. Self-service is a slippery slope - it often solves the immediate problem of time to market but also creates next level challenges. A big issue is that when you remove the friction to data access, you are throwing challenge of finding right data on consumers plate. Data contracts are great when everybody aligns on a single contract and there are enough tools to support the contracts. But they also create a proliferation of data to enforce the contracts required by multiple consumers - thus, they often don't survive the real world. The data catalog space is finally getting some needed attention. But there are still a myriad of issues that need solving. Will those be solved by technology or by leveraging a "data concierge" remains to be seen. It's insanely easy to overspend in the cloud. Everyone is vaguely aware but cost should be part of every important architectural discussion. You can drive business value but it absolutely must also be focused on the cost as return on investment is far more important than simply return.

Jay took a few lessons from working on a central services team in a company of ~200 people. Having a centralized team was doable at first but as the org scaled, it quickly got complicated. As a centralized team, it's very easy to become a bottleneck but Jay learned a lesson that has continued to help in his subsequent roles: "empower people, don't do their jobs." Focus on reducing the friction to others doing their work instead of doing it for them.

Easier said than done so how do you empower people? Per Jay, you must understand the business aspect and what the requestor actually needs. That isn't really going to get communicated well in a ticket most times so you should have a high context information exchange to take what they need and convert it into a workable solution. And often, there is already a solution in place but it's just not handling the job anymore. So you want to consider if you should solve the same issues in a better way. It's much easier to do a greenfield deploy but brownfield is an inevitable facet of enterprise data work.

Per Jay, a few good things to remember: 1) frameworks and technology come and go but the concepts are the things that stick around. Focus on solving issues by leveraging technology and frameworks, not relying