https://www.patreon.com/datameshradio (Data Mesh Radio Patreon) - get access to interviews well before they are released Episode list and links to all available episode transcripts (most interviews from #32 on) https://docs.google.com/spreadsheets/d/1ZmCIinVgIm0xjIVFpL9jMtCiOlBQ7LbvLmtmb0FKcQc/edit?usp=sharing (here) Provided as a free resource by DataStax https://www.datastax.com/products/datastax-astra?utm_source=DataMeshRadio (AstraDB) Transcript for this episode (https://docs.google.com/document/d/1C8KjRoqwiwoZaSQ9LcuZ4Xew_TCx1A69VnEOIeZkLp0/edit?usp=sharing (link)) provided by Starburst. See their Data Mesh Summit recordings https://www.starburst.io/learn/events-webinars/datanova-on-demand/?datameshradio (here) and their great data mesh resource center https://www.starburst.io/info/distributed-data-mesh-resource-center/?datameshradio (here) In this episode, Scott interviewed Björn Smedman, Engineering Manager at Communication Platform-as-a-Service (CPaaS) company Sinch.

Some interesting thoughts or takeaways: A good indicator for when decentralizing your data team might make sense is the cognitive load of a centralized data team. How many systems - including a measure of how complex - are they managing? How much of their time is spent in meetings, especially trying to understand context/requests? Is there starting to be combative prioritization from multiple domains? It can be very beneficial and scalable to apply data mesh principles to non analytical use cases, especially sharing data for application purposes. It is still often difficult to prioritize creating a data product for machine learning without knowing the business value of the ML model. But the ML team needs the data first before they can figure out the business value of the ML model. You have to make speculative bets. If you see the data platform team start to dig into the semantics of a use case, that's a red flag that people are trying to leverage them as a data team. And while you want a centralized data platform team, you probably don't want them to become a centralized data team.

Since December 2020, Sinch raised nearly $2 billion USD. With this funding, they have made a number of sizeable acquisitions, with the company growing from 500 employees to over 3,000 in about a year. This has led to some interesting challenges in sharing data in a hyper-scaling environment.

Per Björn, data is a very key part of Sinch's plans for growth. Sinch's operational systems are often very transactional, as some product lines can process tens of thousands of monetary transactions a second, so data that might be typically shared on the operational plane in other companies is shared on the data plane lest the operational data stores deal with billions of events, making the data challenges even more complex than for most organizations. Then add in the regulatory requirements of telecom.

Björn helped lead the move to decentralizing the data team. When Björn joined, the central data team organization was 4 teams and 25 people. The data function was previously centralized and that was becoming a bottleneck, even for the legacy business. Now that the company had acquired a number of other sizeable companies, that central data team setup clearly wasn't going to scale. The company reorganized around business units and started to build data and analytics teams inside each BU.

For Björn, who started in December 2021 just as Sinch started acquiring new businesses, the central data team was clearly not going to be able to meet the needs of this new organization that was about 8x larger than a year earlier. There was too much cognitive load on the team, especially trying to understand the product lines of five distinct business units, many of which were entirely new to the company.

Björn gave a few good indicators of what to look for when considering if you should decentralize your data team. A big one is team cognitive load. Cognitive...