Brian O’Neill (Founder and Principal @ Designing for Analytics) interviews enterprise data product managers, data scientists, and analytics leaders and discusses the roles that design and UX have in creating useful, usable, and valuable decision support software. From advanced analytics and machine learning applications to AI, BI, and SAAS products, Brian shares the stories of software leaders in their journeys to create meaningful value from data, simpler tools and analytics, and better user experiences with data. The show also features special episodes on the role data is playing in the world of music.
Jim Psota is the Co-Founder and CTO of Panjiva, which was named one of the top 10 Most Innovative Data Science Companies in the World by Fast Company in 2018. Panjiva has mapped the global supply chain using a combination of over 1B shipping transactions and machine learning, and recently the company was acquired by S&P Global.
Jim has spoken about artificial intelligence and entrepreneurship at Harvard Business School, MIT, and The White House, and at numerous academic and industry conferences. He also serves on the World Economic Forum’s Working Group for Artificial Intelligence and has done Ph.D. research in computer science at MIT. Some of the topics we discuss in today’s episode include:
What Jim learned from starting Panjiva from a data-first approach
Brian and Jim’s thoughts on problem solving driven by use cases and people vs. data and AI
3 things Jim wants teams to get right with data products
Jim and Brian’s thoughts on “blackbox” analytics that try to mask complex underlying data to make the UX easier
How Jim balances the messiness of 20+ third-party data sources, designing a good product, and billions of data points
Resources and Links:
Jim Psota
Jim Psota on Twitter
Panjiva
Quotes from Jim Psota
“When you’re dealing with essentially resolving 1.5 billion records, you could think of that you need to compute 1.5 billion squared pairs of potential similarities.”
“It’s much more fulfilling to be building for a person or a set of people that you’ve actually talked to… The engineers are going to develop a much better product and frankly be much happier having that connection to the user.”
“We have crossed a pretty cool threshold where a lot of value can be created now because we have this nice combination of data availability, strong algorithms, and compute power.”
“In our case and many other company’s cases, taking third-party data, no matter where you’re getting your data, there’s going to be issues with it, there’s going to be delays, format changes, granularity differences.”
“As much as possible, we try to use the tools of data science to actually correct the data deficiency or impute or whatever technique is actually going to be better than nothing, but then say this was imputed or this is reported versus imputed…then over time, the user starts to understand if it’s gray italics [the data] was imputed, and if it’s black regular text, that’s reported data, for example.”
Episode Transcript Brian: Hey everyone, it's Brian here with Experiencing Data. On this episode, I’m going to interview Jim Psota of Panjiva. Panjiva is an enterprise data product SaaS company that provides analytics and insights on the global supply chain. Jim’s going to talk to us about three things to get right when you’re creating a data product. He’s also going to go into some of the lessons that he’s learned from the early days starting from a data perspective versus starting from a customer perspective. They’ve recently been acquired as well so I’m going to let Jim talk about his journey with Panjiva. Here’s my interview with Jim Psota.
Welcome back to Experiencing Data. This is Brian O’Neill. I’m really excited to have my friend Jim Psota on the line today. Jim is the cofounder and CTO of Panjiva which is a software company that has mapped the global supply chain using a culmination of over 1 billion shipping transactions and machine learning.
Jim and I have known each other for a couple of years and they’re actually celebrating two recent achievements. One of them being acquired by S&P Global—which is a very old data company—and they were also just named one of the 10 most innovative data science companies in the world by Fast Company which is awesome. Can you tell us what’s going to happen now that Panjiva’s a part of S&P? Is the product going to exist on its own or are you going to fold it into their software? Tell us about what that’s going to mean for your customers and the experience.
Jim: We’re very excited to be part of S&P Global. S&P, just by way of background is an old and one of the first—if not the first—data company founded in 1860 as standard and force rating railroad companies has evolved today into a company that has only the best data if of the best data on companies out there. It’s got a lot of very rich and esoteric but useful data. We’re very excited to complement the offering and also have the brand behind us since everyone knows S&P through S&P 500 etc.
Panjiva is still running on its own right now and over time, we’ll be sort of grafted into their flagship data product which is called S&P Global Market Intelligence as a new supply chain offering. The Panjiva Supply Chain Graft has been concorded or linked up to the S&P Company Graft. It’s actually something that’s already done and we’re about to roll out a new product in the market intelligence product line which essentially gives the first taste of Panjiva and over time the more advanced analytics and data visualization will also be rolled in. The medium term, panjiva.com, will go away. But I actually think that’s probably a good thing because we’re now going to go from having 20,000 customers to over 10x that many users.
Also, a lot of the data that will be used to enrich the offering within S&P but also the techniques and data science pipelines that we’ve built are actually pretty generic at this point because we’ve created over 30 datasets within Panjiva–lots of different types of datasets. We’re talking a lot of that shipping data, but we’ve also pulled a lot of company-level data as well. We’re excited to also leverage a lot of those data science pipelines and other software tools that we built for shipping data and use that sort of more generically across the S&P datasets.
Brian: That sounds like it’s going to be a big project to merge all your analytics.
Jim: For sure.
Brian: That sounds like a lot of work to get right and to do carefully and to make sure the value still is evident. Does S&P see that as you are filling gaps that they didn’t have or is it more like you’re adding on features, or insights, or additional layer? For example, they have every company that you have but they don’t have data XY and Z or analytic insight XY and Z. Is it more the enhancement or is it more of a gap fill?
Jim: It’s actually both. It’s sort of the bread and butter of S&P has been more public companies and banks inside more developed countries. There’s a lot of Panjiva data in those areas as well but Panjiva really shines in places where a lot of data providers have stumbled which is in the longtail of smaller companies, a lot of private companies. Alternative private data is very popular right now and very unique. Panjiva’s really going to be adding a lot there.
Panjiva currently has profile from the 1.5 billion shipping records that maps about 10 million unique company entities and that’s going to significantly increase the company level coverage that S&P has. They’re excited about that but as you mentioned, there’s sort of the data science and analytics capabilities are on top of that that we’re going to be working with them to fold a lot of that in as well to do sort of on-the-fly reporting, allow the product to be really customized for a particular user and use case as opposed to pre-packaging reports–that’s going to be another piece.
Finally, the team, I’m very proud of the team. They’re pretty amazing and trust me, every day, they get me excited to work closely with the Panjiva team to develop the next generation data and AI products.
Brian: You’ve also done PhD work at MIT, so you’ve got a lot of background in computer science, data products, and I’m really excited to talk to you about what you guys started at Panjiva, where you guys went. Obviously, you were doing something valuable because a: customers and b: you’ve just been acquired. Welcome to the show! Tell us a little bit more about Panjiva and what you’ve been up to over there.
Jim: Yes, Brian. Great to be here and see data products,products in general, analytics products similarly,so great to talk through this. I can say that we’ve had a lot of achievements recently, but it’s taken definitely a long time to get here and a lot of meandering of back and forth, a lot of mistakes. We can talk about all that as well.
We’re AI company focused on supply chain, helping companies engaged in global trades make better decisions when they’re confronted with different aspects of the global supply chain. We realized that global supply chain is incredibly complex, incredibly opaque, and we sought out to use data and technology to help people make better decisions. Our bread and butter is transactional level shipping data. We’ll talk about that more but suffice it to say, for now, it’s really large data and it’s really messy on structured data. That’s where the technology comes in to essentially map that data, turn it into structured data that is actually amenable to analysis. There is a lot of noise that we split out and then we essentially package that data in a SaaS product—very visual and intuitive—to essentially allow non-technical users to ask questions and gain insights from Panjiva. That’s kind of the broad strokes.
I can give you a few specific examples to make it concrete. We have over 3000 customers and those customers span across a variety of industry. Physical good manufacturers are our bread and butter, folks that are actually importing goods. We have customers like Walmart, Home Depot, but also folks that are companies that are analyzing global trade sometimes at a macro level looking at an industry or looking at a company perhaps for investment purposes. We have hedge funds and asset managers using Panjiva. Also, shipping companies use Panjiva to optimize their shipping roots.
To just give you a concrete example, imagine if you run a shipping company and you have one of your big container ships going from the West edge of the Panama Canal up to San Francisco. That boat specializes in refrigerated containers, carrying goods like vegetables. Let’s say that boat is only 60% full and you want to find customers that are shipping on that route that are shipping a particular product that you care about and in particular volumes and frequencies. We have 10 million companies in Panjiva and with a few clicks in the interface, you can very quickly find a shortlist of companies you can go talk to potentially to partner with them and get them into your shipping route. That’s a concrete example.
Another concrete example, importers such as Home Depot who’s been one of our long standing customers, they use Panjiva to find new manufacturers. If they’re coming out with a new product line or ramping up volume on a particular product line, they can look for manufacturers all over the world. It’s interesting because in the beginning of Panjiva we thought that our sort of go-to target market would be small, medium-sized customers, but we learned very quickly that even large companies—even ones with offshore offices in places like China and Vietnam—they also subscribe to Panjiva because things are constantly shifting with shifting tariff rates and product lines and companies going in and out of business. These folks are really hungry for data.
Importers use Panjiva to find new manufacturers and also to keep an eye on their competition. And then finally, exporters, we have a many customers in places like China and Vietnam. They will use Panjiva essentially to find companies that buy the goods that they make. You can go in to Panjiva and look at a groups of buyers, so to speak, categorize them, importers, and essentially find a shortlist of companies to reach out to and market to.
Brian: It’s almost like, not necessarily lead generation, but you’re almost connecting a buyer and seller–like you gave the Home Depot example. If I have that right, that might be an example such as, I’ve a vegetable garden actually in my backyard so I use this drip irrigation system and I got it at Home Depot. Let’s say Home Depot is like, “We want to have our own line of drip irrigation.” They might go and look up, “Who’s shipping from this area that specializes in plastics or something?” Or they might look at competitor and type in their name, is it something like, “Oh, let’s go look at whatever rainwater.com,” whatever the competitor is and see who’s their supplier and that type of thing? Is that how they might use the product?
Jim: Yes, exactly. Lead gen is one of our few key use cases and then in the end verse, sort of vendor generation and competitive intelligence. You hit the nail on the head, that will be a very popular case to look at as a competitor or a future competitor and say, “Were they getting their goods made?” The reason that—and this gets a little bit into the secret sauce—the reason that that’s so straightforward in Panjiva but difficult with sort of the raw material data which is the shipping data is all the structuring that we do.
We’ve essentially mapped a graph for the network of the global supply chain. The Panjiva supply chain graph, as you call it, is the largest and most complete supply chain grain in existence because we have essentially pieced together shipping data from about 20 different transaction level data sources. These data sources tend to be government agencies. In the USA,it’s Department of Homeland Security. Because we’re kind of piecing these together, you can see the primary nodes in the graph are essentially companies. The links between those nodes are supply chain transaction level relationships. You can see customers or vendors of relationships. You can very quickly navigate that graph versus kind of see how you’re connected to a particular company much like you can do so on LinkedIn or Facebook with people.
Brian: Are some of the data science aspects ,in terms of what you’re doing, have to do with resolving entities that have names that are listed Dell Computer, Dell EMC, Dell, Dell Limited, and it’s like, “Well, this is actually all one company?” Even though their manifest may have different company names on that and then understanding that that’s one logical entity, so that you can look at all of Dell’s suppliers for example. Is that one of the things you guys do in terms of cleaning the data and actually being able to provide this graph that’s accurate?
Jim: Exactly. The entity resolution technology and clustering of the data is something that we spent over a personal decade of work on…
Brian: Wow.
Jim: …and is definitely a great [...] fourth generation. There’s a couple of reasons why doing it at this scale that we’ve had to do it at is more difficult than de-duplicating your Salesforce account which is another example of this but at a much lower scale. When you’re dealing with essentially resolving 1.5 billion records, you could think of that you need to compute 1.5 billion squared pairs of potential similarities. If you just run the naïve calculation, that could end up taking many, many years.
Basically, the techniques we’ve developed handle both the scale but also the sort of dirtiness of the data for a given company. We may have, you gave a Dell example, we may have tens of thousands of name variations and in different languages. Right now, the Panjiva data comes in six or seven different languages, character encoding, Chinese, etc. Handling all that gracefully is difficult. But as you said, one of the key value add from the data science side, we haven’t talked about the product yet, but on the data science side that’s one of the handful key enhancement that we make to the raw data to get it into essentially a new dataset which is this combined dataset and [..ellipses.] graph that is I think the foundation of the whole company and product.
Brian: Here in the Cambridge, Massachusetts area where I am there’s a company in town called Basis Technology. I’m not sure if you are you familiar. I think they do some of the same work like text analytics and more on the Homeland Security, the immigrations, and border patrol—the software that the agents use when people come into the country and the same issue with multiple languages, anglicization of for example a Russian name that may end in OFF versus OV. It’s this federal law for something like this and knowing that that’s logically the same person. It’s interesting in that space how much clean-up you have to do. Sort of a theme on this show with a lot of the people I talk to is how much work has to be done just to get to the point where you can do some fun stuff and you can actually solve some problems because there’s so much of the data engineering and data clean-up and just getting to that point where you can do the analysis and stuff is hard.
Jim: Absolutely. First of all, I love Basis Technology. I’m friends with Carl, the founder and CEO, but you’re right it certainly is a similar problem. To your second point, yes, there’s sort of a dark reality of data science or data projects is really spending a percent of your time dealing with a, understanding the data. I mean, one of the big lessons that I’ve had and one of the things that we did not do very well at all in the early days of Panjiva was really deeply understand the semantics and definitions of the dataset before really trying to get value out of it. We kind of just loaded it and just had at it from the computer science standpoint. But we really skipped over one of the most foundational and fundamental precursor tasks to that I think which is doing a research project and really understanding, what does the data mean, what does the columns mean, what are the distributions of the data.
The data that we’re using, and this is often the case, was collected for other purposes; it was collected for collecting tariffs as goods cross borders. The data was meant for that purpose and nothing else. They happened to find a resell use cases for it but because of the data—again this is often the case—the data is not really meant to be nice, clean, normalized, error-checked data. We had to spend a lot of time and currently spend a lot of time on every new data that we’re adding and we’re continuing to add data; really understanding the semantics of the data before we even get into it.
We have had a whole team called our data science analysis team, they are sort of former academic researchers, economics type researchers ,who have some technical skills, and essentially do a deep dive for weeks per dataset to really deeply understand what the data says and means before we try to start develop a product on top of that data.
Brian: I understand that side of it which is understanding the materials that you’re going to be working with and the information that you have. Maybe this has changed over time but it’s not like you took the data-centric approach in your early days…
Jim: Early days, for sure, yes.
Brian: … and started surfacing, “Let’s get it on the screen. Now, let’s filter. Now let’s have some controls to do things with it.” Has that changed over time? How does that user, this person that’s looking for those irrigation supplies or like, “I want to change the coffee beans",whoever it may be, how did they fit into that story if you bring in a new dataset or something like that, maybe it’s in a domain, maybe it’s new columns or new information that you don’t currently have, is there something that you guys do differently now to make sure that user actually is going to get some value out of it before it goes into the product?
Jim: This is something that you and I have spoken about a decent amount, Brian. I think we really frankly wasted a lot of time doing some kind of science experiment in the beginning phase of the company. I think there’s a place and a time for that, sure. Just kind of getting your head around the data and potential use cases but where we’ve evolved to is a very user-centric mindset about how to actually build and deliver viable products.
I think this is something that’s IT, when I’m inviting companies or helping friends, again this is what we did in the early days, it’s very common to see companies that are data source-centric, AI-centric, and I really think every company needs to be user-centric and use case-centric, and then needs to have an arsenal of tools from the data science world, from the visualization world, etc., etc., and have those tools to solve the user’s problems and tap into them as we do. Know what the user wants so you can go gather the data that the user cares about and really package the product to deliver into particular use cases.
The early days, we were a little more of a hammer looking for a nail kind of scenario. We attack the problem quite differently now although I would say it’s still not perfect and we still have a number of challenges, but we really try to focus on the used case. We have a whole product team now that we didn’t have before that really deeply understand that and then we do the technology and data piece after that.
Brian: I imagine you guys probably talk to customers at times to inform this process. Is that right?
Jim: Absolutely. We have a number of different ways to engage with our customers. Some are very high-touch and small number of data points and some are low-touch but thicker on the number of data points from the customers. We have a customer success team that is very close to our 20,000 paying users, the closest you can be to 20,000 paying users ,obviously. Closer to some more than others. We have an internal system of course to constantly be tracking feedback, that’s one way.
Our product team also has developed a small cohort of users, called our VIP Users, that essentially get early access to beta features and we have in-product feedback mechanisms and phone calls, screen shares, etc, with these users especially for the users who are really engaged and really excited about Panjiva developing product etc. That’s been a fantastic resource. Finally, there are a number of hooks that we have that are constantly measuring data that we’re able to get a very broad view about usage and discoverability, etc. Obviously, we need to interpret that data very carefully, it’s very easy to draw poor conclusions from skewed or thin data. But if you kind of put all these pieces together then we try to get pretty close to the user.
I’ll also say that in addition to, we’re talking about customer success and product, but engineering and data science is, I would say, very close to the customer of Panjiva. It’s always been very important to me because I think that when you’re developing a product, the engineers often have to make micro decisions about how exactly to implement a feature, how to lay the groundwork for a potentially future features, and just to give out accountability of the software. The engineers are going to develop a much better product and frankly be much happier having that connection to the user, it’s much more fulfilling to be building for a person or a set of people that you’ve actually talked to. We sort of touch the user in all of those ways.
Brian: That’s great. It sounds like you have some type of Google Analytics equivalent or an internal metrics on page views ,maybe a timeline page, stuff like that, but it sounds like you’ve also got some qualitative. I don’t know, is it like a chat window where I can say, “Hey, what does this button mean?” or these types of feedback mechanisms. Is that correct? You have some interface like that where they can send an email right from the interface about issues?
Jim: Yes.
Brian: Got it.
Jim: Exactly. It’s sort of an ever-persistent feedback mechanism and that’s for everybody. And then on top of that where the VIP users, that beta group, what we do is we essentially build a little extra user interface element to allow all our users to type in all their feedback directly. That feedback actually goes to both the project management team or the subset of that team working on the product as well as to the specific engineers that have worked on that product. They’re able to get direct feedback. I just think you want to lower the friction as much as possible so that at that moment, when they’re feeling the annoyance or have the inspiration for a way to make it better, you just want that box to be staring them in the face. It’s not even a link that you click and then open the box, it’s just a box that then you type and click. We just want to lower that barrier.
Brian: From having to put these processes into place, is there a story or a particular anecdote that comes to mind? Something that you learned from having these customer touchpoints on a regular basis? Like we never would’ve known X had we not either asked for that passive feedback or maybe in an interview or a screen share session or something. Any particular nugget, story?
Jim: Yes. I’ll play you kind of a broad story which is a little bit more of a broad learning and maybe I can give you a specific example as well. But the broader thing is when we started the company, we basically tried to over simplify and smooth over too many details when it came to distilling this mountain of shipping data into insights. We did that by developing what was called the Panjiva Rating. We’ve actually [...] that feature. But we essentially came up with our own metrics to look at the data and boil it all down to a number between 0 and 100 that assessed the “goodness” of a particular supplier.
What we learned from users—again and again and again—was that that just was oversimplifying and not really appreciating the nuance of each particular user’s questions and use case. They were all using the data for a particular use case and we started the company focusing on finding vendors or suppliers’ use case that we talked about a few minutes ago even though we were focused on the used case, the users, they didn’t have enough fidelity in the data that we are offering with a single a number. They often didn’t trust it. They kept saying, “Hey, give us the data.” At that point we were a couple years old as a company, no one had ever heard of us, we just didn’t have the trust from the users.
Frankly, the product that we were offering wasn’t very good. The number didn’t even work that well so as the metrics. That sort of taught us—and this is sort of by direct customer conversation—that at that point, they were asking for the raw data then they got the raw data. We built another piece to give our raw data and then they said, “That’s too much. I need to have a machine learning background to make sense of this.” And the pendulum kind of swung back again and we kind of ended-up where we are now, which is somewhere in the middle. But if it wasn’t for sort of getting kicked in the Ts by a lot of customers saying, “This is not what we need to answer our question,” we would not have gotten to where we are today.
Brian: You created this Panjiva Index. I’ve seen this with other products as well. Do you think the issue was really that they didn’t trust it? Was it that the data wasn’t available in the tool for them to unpack the 72 score that they saw for whatever this lead was or was it there, but they couldn’t understand how you guys packaged it up and made the index? Do you think it was that or was it a lack of belief that you guys could really boil the word down to a single number? It’s kind of like, “This 72, this is supposed to be right for me and all the other that are not really doing exactly what I’m doing so how much should I really believe in the 72?” I’m sure it’s about that thinking there.
Because theoretically, it’s on the right track. You’re trying to reduce the amount of input required to get some insight like how much tool time has to be extended in order to get some kind of value from the tool quickly because your goal is not to spend time in Panjiva probably; the goal is to get some insight out of it. I can understand the desire to go into that tactic. Could you unpack that a little bit more?
Jim: You’re right. Actually, I think it’s a bit of both, but I will unpack it and try to assign some weight to each of these. I really think that the primary problem was that it just wasn’t enough fidelity and it wasn’t enough for the users to really sink their teeth into to actually answer their question. It was on the right track, but it was too far, it was too extreme. To this day I believe that even though we have a brand now, we have a very well-known parent company that lends us a lot of credibility and grace, I still don’t think of the right level of fidelity. I think it’s too blunt of an instrument for allowing users to get insights that they really want them to get. I think that was a primary issue.
I think sort of the extra factor was they also didn’t trust us. I think there’s this, “If you want to go in that direction, you want to give users as much insight as possible while limiting the amount of time they’re spending getting that insight.” But I do think there is a point where it’s too far. It could either be too far because you’re not allowing the user to fully articulate what they want out of the tool, maybe the user interface is sort of the querying mechanism instead of the search type of interface is oversimplified. A lot of people want to emulate Google for example but sometimes Google is actually too simple, so I think you could have oversimplification on the input side, and then you can also have oversimplification on the output side in terms of how the data for the insights are ultimately presented.
Brian: Do you guys have any type of internal benchmark use cases or some kind of way of testing the product with users that you repeat over time to see if like, “Are we still delivering the quality and the value we want to give”? I imagine you probably bringing on new data or maybe you’re creating new reports, or that the product is changing over time. How do you make sure that you’re not making it worse? Obviously, when we add in information, we potentially add value but we also add noise potentially and friction. That’s where the design tier comes in. I’m just curious if you have some way of real-time monitoring, having some sense of a benchmark where the tiJme it should take to pull up a company, find all of its suppliers, and do X, we want that to take eight minutes or something. Can you talk about that a little bit? Do you have any type of process you guys use for studying that?
Jim: Yeah. I like the benchmark you just talked about. We don’t have anything that tight on the user on the flow or the UX side of things, but something that’s a good idea. I tell you what we do do. We basically compartmentalize the quality assessment. What I mean by compartmentalize is we look at it in a few different key areas. I think it’s actually important and useful to look at it both in compartments as well as holistically.
The first one is on the data side. Essentially, the quality of the data science algorithms and the machinery that processes the data. Do that independently and you can essentially use different techniques to make sure that if the data format changes, we have infrastructure in place to automatically notice that, to make sure that that doesn’t leak into the product. There are often these breakpoints or thresholds that sometimes get triggered or maybe the machine learning model was trained on one type of data. Ultimately, that model ends up being stale over time. We have so much data and so many different models that we have automated ways of checking that certain quality metrics, sort of standard data science data quality metrics are upheld. I think that’s the foundation to Panjiva is the data itself and the matching graphs so that we have lots of mechanisms in place. Over 5000 automated tests if you include the data side of the product side that are constantly running multiple times a day.
On top of that, there’s performance. That’s just like, “Is the website fast enough?” Especially for a company like ours where the amount of data is only ever increasing, often you have performance needs where all of the sudden, the data that used to fit in memory and now it’s spilling over to disk or in particular, database index for example, and then you’ll hit a performance [...need]. That will show itself in the product. We have ways of essentially monitoring our key products for our key use cases to make sure that searches are fast enough, company profiles and some reports are fast. That’s compartment number two.
Compartment number three is more on the product side. There I would say that the best that we do is kind of the stuff we talked about earlier with vetting a product quality is make sure it doesn’t break. But that’s different than what this sort of holistic flow that you mentioned that I think we should be doing and we’re not.
Brian: Since it’s a process, the qualitative types of research activities and design is not binary. It’s not like you are or you aren’t. Most people are somewhere on a continuum of the habits and activities and routines that they go through to be customer-focused. Some are doing a lot of stuff, some aren’t, but that benchmarking thing is something that I think helps companies.
With my clients, especially with analytics progs, the problem space is indeed a space. It’s usually not just like, “Here, there are five things we need to do " and that’s it. If we get that right, then we’ll sell the company and go to an island and party or whatever.” It’s never that simple.
But , actually you could use something along those lines as a benchmark to kind of check yourselves especially as products grow. I like to try to encourage clients usually. I mean, it depends on what the problem is in that particular situation. But having some kind of idea of, we have to put a stake in the ground somewhere if we’re going to evaluate the quality of the product from the user experience standpoint, we need to be able to go and run tasks with customers and ask them to perform activities that are realistic based on what their job is, so let’s pick a handful of these to get going. They may change over time but instead of trying to solve an amorphous global supply chain, in general, like we’re going to solve their problem. Well, what problem?
At some point, ink is going to go on a screen, buttons will be created, workflows will happen, and that can either be a very deliberate process or you can fall into it. My thing is like, “Well, let’s try to model it around problems that people actually have, pick a subset of those,” and then over time you’ll probably learn whether those are even the right benchmarks or not, but it will at least guide the processing, keep it from being a mediocre product for everybody instead of you actually have a really great product for a smaller set of people. That might mean some users have to suffer a little bit. They’re not going to get the A+ experience. They make it a B-, a C+ experience but you’ve decide that, “Hey, we’re not going to sacrifice the jolly Jane persona or the supply chain Eddie.” You come up with these people that kind of your models or you really need to satisfy the most and you say, “We’re not going to put that guy’s job at risk because they’re our top client, they’re our top customer, they’re the ones that actually spend two hours per day in the product. They’re not the people that check-in once a month to download this report. It’s okay if that reporting interface isn’t as great. Let’s not screw with the features and the tool that we know that Eddie is using two hours a day every morning when he gets his coffee and he sits down. The first thing he does is open email and he opens Panjiva and does X, Y, Z.
Jim: Exactly. So easy to get overwhelmed with the degrees of freedom that you have when starting a company and thinking about a new product. One of the things we struggled with, in the early days especially, is the blessing and curse of Panjiva. We have so many different types of users and use cases. It’s actually about 10 if you enumerate them, but only a few are the key ones. Eventually where we landed is we have mechanisms for all 10 of those use cases to get the value out of the product, but we really try to nail top three use cases and that is special flows and special nomenclatures, etc. for these particular users.
Brian: I’m curious. I don’t know if you call your company an engineering-driven company or not, but for ones that are, do you get the people that are always wearing the exception hat? I could see in a space like yours where there’s always going to be someone that’s going to say, “Yeah, but someone might want X,” or, “we taught that one guy that wants to do it like Y.” How do you solve that tension? Do your guys run into this problem where maybe a squeaky wheel and early customer’s been with you for a long time, has a really strong feeling about something, and I’m sure you guys have some internal debates about, “Are we going to satisfy this guy or gal with what they’re asking for?” or “Nope, we’re not going to go there because of this.” How do you handle the competing requests?
Jim: As an engineering-driven company, it’s fun to build things. You have folks join the company because they like the entrepreneurial freedom they have, to talk to customers and to develop products. Actually, right now, as we are speaking, there’s a hackathon going on and everyone’s working on projects that they came up with on their own. It could be product features, it could be data science features, but only the ones that are truly deemed to be valuable to customers are actually going to be worked on for real after the hackathon.
Back to your question, I would say in the beginning, it’s very exciting to have any customer that will listen to you at the beginning of any company, very excited to build something that does an okay job at meeting their needs and solving their pain. I think we hastily built products in the early days, oversimplified the needs of the customer, and the ultimate product that we developed. But that’s okay because we move fast and we built the next one as we learned.
I think there’s a couple of ways the tension that you’re describing manifest itself. One is, as you described, a squeaky-wheel customer that gets very upset. This manifest itself especially sort of perditiously when a salesperson is on the line with a customer that is about to close or they’re already a customer, they’re not sure if they want to renew or not, and you have to make a call. Are we going to keep this feature? Are we going to build this feature for this one potential customer who’s somewhat valuable but it’s not a crazy amount of value? That comes up all the time and that’s just the top judgment call that we constantly have to manage.
Ideally, you have a strategic direction for the product. It’s easier to make the call to build if that thing is in line with where you’re going anyway of course, but it gets really hard when you’re going to hit your monthly goal or quarterly revenue goal. This thing is a little bit out there and I think it’s great to have the freedom to have discipline to say no.
There’s another way that manifests itself beyond customers, though, which is just internal. There’s just a mix of personality that any group of people, and it’s very common for maybe a little bit more extroverted or strong-minded team member. It could be another engineer, it could be a salesperson, it could be business developers who essentially is getting an oversized share of sway in a company. It gets really important to try to acknowledge that sort of natural distribution of different personality types within any team and try to bring out of different team members and pull out of different team members the opinion that they have, and then try to have more of a cross-cutting look at what are we hearing in general from a broader set of customers, a broader set of people.
The final thing I’ll say is in another bias, which was recently bias, is very common with the way we’ve operated and has worked pretty well, is did our major planning of larger projects on kind of a quarterly basis, obviously re-evaluating every couple of weeks with our strong style iterations. But it’s very common and very easy to intentionally prioritize things that happened to have come up within a couple of weeks of the planning this meeting which maybe around beginning or end of the quarter. It’s important to try to take a little bit more of a longitudinal view of the feedback we’ve been getting over time and documenting that so that when you do do the planning, you’re smoothing over the recently bias, you’re smoothing over the strong personalities and the irate customers, and just trying to do what’s best for the company more in the medium and the long term.
Brian: I know that you have recently given a talk at a conference on this. If I recall, the title of the talk was Three Things To Get Right In Data Science. Is that correct? Could you share with us a brief version of what those things were? If that’s not quite the right title, tell us what it was?
Jim: It’s Avoiding Data Disillusionment: Three Things To Get Right When Building Data Products. First, a little bit of a preamble. There’s a reason I think we are headed towards disillusionment or at least a lot of data science projects are. There’s just a ton of hype and excitement around data, data science, machine learning, AI, and I think for good reason. We have crossed a pretty cool threshold where a lot of values are able to be created now because we have this nice combination of data availability, strong algorithms, and compute power. That combination is certainly powerful. But there are a lot of folks out there with the data set hunting for a problem to solve and aren’t necessarily going about it with a user-centric and use case focus.
Gardner put up hype cycles for different technology. If you look at the hype cycle that was put out a couple of months ago by Gardner for data science in general, pretty much all of the technology, except for a few, are in that peak. If history repeats itself for data science like it did for enterprise software, internet technologies, mobile technology, et cetera, a lot of folks are going to be disappointed over the next 3-5 years.
It’s all about the decade that we’ve spend building Panjiva and a few of the key learnings, I think for me at least, you apply these learnings,it will reduce the risk of disappointment. The three areas to get right: one is deliver what’s valuable for the user, two is demystify the technology, and three is democratize the data science talent.
On the first one, this is so obvious and sounds [...contrite]. It almost feels silly to say, but we talked about this a lot on this particular conversation. I just think this is worth repeating and this is worth keeping front and center all day everyday. Just really focus on delivering something that’s valuable to the user no matter what.
When we started Panjiva, we actually were not an AI company focusing on shipping data. We were actually focused on deobfuscating the supply chain in a very different way, which is building more of a platform for a rating with subjective reviews and helping companies learn about other companies in far away places like China or India because the reviews from other people. We thought that was going to work. We saw a lot of inspirational examples. After about six months, we learned that that was just not going to work. There are a lot of incentives in play that a lot of people do not want to share their good suppliers and people want to tarnish their competitors. There are a lot of incentives in play that made that business model not viable.
But in the process of building that product, we stumbled upon some data that we were planning on using to assess the veracity of the ratings themselves. That data was shipping data. We were planning on using the shipping data to figure out if the ratings are coming from real customers. We get a rating and look up, with the data in the background, kind of manually, is this a real customer customer of this giant manufacturer?
In the process, we realized, is that this rating thing is not working but the data is actually really interesting. At first, we thought there’s just too much to get any value out of, but looking at it from a machine learning and data science angle, we realized that if we worked really hard at this, we can actually turn it into something valuable.
This is one thing we’ve got right, there was a lot of luck here, but all along we were focused on helping the user get insight about these companies around the world at a distance and developing trust at a distance, which is an age-old problem. No matter if it was a platform business or a rating business or an AI data business, we were focused on solving that problem. That’s the first lesson delivered,what’s valuable to the user.
The next one is demystify the technology. We talked about this to some degree earlier with the Panjiva rating. We’re trying to have this magical black box that took all this data, which could be tens of thousands or hundreds of thousands of shipments associated with a given company and boil all of that down to one number. In a way, we were wrapping up all of this technology behind this black box. We talked about before, the user just didn’t trust it, they don’t understand it, and they couldn’t contextualize it with this particular business problem. Demystifying the technology is really important.
This is coming into play a lot in data science and AI in particular now, where the reality is the models are quite complicated and quite sophisticated. But that doesn’t mean that we could just let it be a black box and spit out an answer. I think it’s so important to wrap that technology and give users hooks into the technology so that they can, a: trust it, and, b: take the insight, contextualize it for their particular problem or their particular business use case, and make it as user-friendly as possible. You have to demystify technology.
Three is democratize the data science talent. This is a little bit more about tactical approach that is just necessary, given the scarcity of data science talent in the world today. I don’t know the statistics but there are more job openings for [data scientists than data scientists ]out there and that’s going to be the case for quite some time. Carnegie-Mellon just came out with the first undergraduate major in AI but it does take some time.
I think it goes beyond that. I think this is beyond just leveraging the scarce data science talent to actually get products built. I think the reason that it’s important to democratize the data science talent is also because helping product managers, other engineers, and even folks doing business development understand the capabilities of data science is going to fundamentally shape how they think about developing a product and mapping the user’s need into the technology domain.
Thinking about building a model, thinking about features of the data, developing a training set, and assessing error rates and communicating how the black box works is really a fundamentally different approach to developing products. I think it’s really important to educate the broader team in the broad strokes of data science so that they understand how to leverage out the tool even if they don’t know how to write the code or think about it from a mathematical perspective. The third thing to get right when building data product is democratizing the data science talent.
Brian: In terms of the second one, in terms of demystifying the text, I’m curious how you see that line between either [user sit down and I have some task or job ]that I need to perform. With analytics, it’s usually some kind of decision support that I’m looking to get or in near could be a lead, something like that. How does the need to understand how Panjiva generated the response that I got? Where’s that line between, “Woah, I just want to go to the grocery store. I don’t need to know on the screen of my car, ‘The fuel is now being injected into the whatever,’ and you can see every part of how the engine is going to move the wheels, etc.”? Where’s that line between noise and not needing to really understand all of it versus it sounds like maybe are you saying you need to expose enough to get some trust? Is it about building the trust that’s important and then exposing the right amount of the magic sauce? Can you unpack that a little bit?
Jim: Yeah. There’s a couple of aspects here. I think the actual form that the product can take that provides a nice, happy medium, to not overwhelm the user but also give them the hooks they need to actually demystify the technology is I’ll use the onion analogy. You give them a high-level view of the insight and ideally package the product or the insight in language that maps the way they think the problem. Make them very simple at the outset but provide a set of drill and mechanisms to actually go deeper if they want to. Start with the simple thing but don’t stop there from a product standpoint.
I’ll say that that is a tactic of building a product itself. I would say that the goals though are two-fold. One is to build trust, as you said, and second is to provide context for the particular use case. In Panjiva, there are many use cases, and different use cases may need more or less fidelity and nuance, and maybe fields of data or types of customization. The ideal solution is to build a specialized product for every use case, when the user just goes in and make it exactly what they need for their use case. And maybe an advanced user peel the onion a bit and go a little bit deeper if they need to, if they want to really understand it. That’s the ideal scenario.
In our case, we kind of have a hybrid solution where a few of our use cases have that level of tightness, where we’re mapping product to use case or use case to product. But we have sort of long tail of use cases and for those, we provide a little bit more of a generic advanced interface, give folks the onion approach where they’re drilling in if they need to, give them enough hooks into the data so that they can understand how to map this insight into the particular problem that they’re solving, be it for optimizing their shipping lane or finding out information about their competitor.
Brian: I think you outlined the sort of a framework for products that have discrete conclusions or insights that you know there’s going to be repetitive need to go in and get answer A for question B, as it may be. I’ve seen that process where the framework for a design work well where you’re trying to prevent what I would call ‘presenting a conclusion first, not the evidence’ but you kind of provide that, “The answer is it’s the 72 index,” if you’re going to go something like an index, for example, but then you need to provide the right amount of supporting evidence to back that up.
One thing I’ve seen work well in this space, too, is sometimes that includes information on what you didn’t do, like, “What information did we not include"? Or, "Hey, our supply chain data is a year old, so what we’re giving you is actually from 2017 not 2018,” or, “We did not adjust for inflation,” or, “We did not do X.”
It’s not list out everything that you didn’t do because that could be a mile long, but as you get to know your customer and the questions that may be going through their head, which you only can learn by talking to them, you may be able to answer that in the interface explicitly on the “evidence page.” It may not be a single page but the place where you backed up some of those analytical conclusions.
You might give them an idea about, “Here’s what we checked, here’s what we looked at, we cross-referenced it with this, but we did not do these things,” and then they can start to believe the trust that they have a little bit more idea how the sauce works.
I’ve even seen to the point where at some point, they may even stop looking at that and now they start to really trust the conclusions because they know what goes into the recipe and they don’t need to know the recipe anymore, they just want the pizza, like, “I know it’s good, I know what kind of flour you use,I was interested the first time because it’s so good but now it’s just...”
Jim: Now I just eat it.
Brian: Yeah.
Jim: Absolutely. We saw actually that exact pattern play out at Panjiva where a lot of users in their first couple of months will use the more advanced interfaces, where one exports the data to Excel, double-checks the math doing pivot tables, etc. As time goes on, they’re just using the repackaged reports because hopefully, they’re better than the manual analysis over time because we kind of consider these exceptional data cases.
Any data product that’s done well or any non-trivial data product that’s done well, sometimes is going to have dozens of these things. No matter if you’re mining first party data of companies like Ballinteer, Tamer, and other companies that are essentially taking companies’ internal data and working with it. In our case and many other companies cases, taking third-party data, no matter where you’re getting your data, there’s going to be issues with it, there’s going to be delays, format changes, granularity differences. They can get overwhelming to the user, too, essentially list out every contingency, every deficiency. I think there’s a real balance here, a balancing act.
As much as possible, we try to use the tools of data science to actually correct the data deficiency or impute or whatever technique is actually going to be better than nothing, but then say this was imputed or this is reported versus imputed, I think that can lean on some common design language in the product basically, and then over time the user starts to understand if it’s gray italics that was imputed and if it’s black regular text that’s reported data, for example. I think that can help the user just sort of intuitively grasp, “Okay, you need to be a little bit more careful with this, but I [..get the jest.],” and depending on their use cases, as an example of communicating just enough to get the user who need the level of fidelity or maybe a tool tip can say imputed methodology, if they need that or if they just [ need the broad stroketo get the product ...], they don’t take in.
Brian: I like that idea of the design, if you can build those kinds of things in and again over time by minimizing. You don’t have to hit people over the head with it. It’s knowing the questions that your customer might ask but are the ones that you might need to fill in. It’s not every deficiency. It’s just the ones that may be a friction point for them like, “Hey, am I really going to pull the trigger based on this insight?” Is there some answer you can give them or some information that you can give them to make them feel comfortable with the decision support that the tool is generating? Like in your case using imputed data, for example, I think that’s good especially if we can balance the subtlety in the interface there such that it’s not noise, you’re adding additional noise as well. I like that thinking a lot.
Last question. We’re getting close to the clock here on our time. I’m curious to know without getting into the hype of data science and all that, but is there something that you’re excited about in your space that you think, because of the climate we’re in, with the compute power being available, you guys obviously are dealing with a ton of information, you’ve cleaned a lot of it, is there a new place you guys can go with this technology in terms of simplifying the experience for your customers? Whether it’s a new feature or, “Hey, we’re going to be able to cut out a whole section of the product because this technology is going to allow us to do X now and we could never do that before.” Just curious. What’s the next journey look like? Is there something new that’s going to be enabled or is it more of a slow crawl and the toolsets get better along the way? It’s still going to be a house but we don’t erect the walls the same way. They may not see how we erect the walls but for us internally it’s easier. Can you just kind of speak openly to that?
Jim: I’ll go back to the user again. I think the really, really tough nut to crack when building a product is finding the product market fit and really finding the key nugget or nuggets that are going to answer questions and solve pain for your user. All the stuff that we’re working with, data, different data science frameworks, et cetera, these are great tools and these tools are getting better. Not essentially accelerating the pace of development, but I don’t actually think it’s unlocking, say for a few key use cases like autonomous vehicles and some other use cases that really were very difficult before but now are actually possible.
For those use cases, I think the fundamental challenge that we all face is really just how do we use this ever increasing and improving arsenal of tools that we have at our disposal from AI, data visualization, and really fast parallel analytic capabilities on the backend? How do we cobble together a product that really meets the user’s needs? I see that as we gain more inches, not feet. It’s going to be such a fundamental problem, I think it’s not going to go away, maybe ever.
Brian: You hit on a topic I’ve talked about on this show many times before. It’s not a magic bullet. All these technologies, machine learning, and whether it’s faster computing power, whatever it may be, most of these things are not the just take the pill, swallow it, and then bam, instant new business value. Now, run out and find a place to go use this new tool. That’s not necessarily going to save you.
You really got to understand the problem space and know how to deploy the technology properly because at the end of the day, they’re still going to log into this interface that’s running in a browser, they’re going to pull up some information, somehow they’re going to receive pixels and ink on a screen telling them something. How you guys did all that in the backend and the technology that went into it, that may change over time, it may get better, it could be accuracy-improving, but it’s not usually a magic bullet. You kind of reiterated that. Typically, I hear it. I think it’s important with all the hype that’s going on that’s it not like, “Oh my God, we don’t even need an interface anymore.”
Jim: To be fair, I am a technologist at heart, I love this stuff, and I think it’s super exciting. I just think that it’s very easy to get wrapped up in the technology and have that overshadow what the most important thing we’re here for.
Brian: Cool. This has been super fun. We’ve been talking to Jim Psota from Panjiva. It’s now part of S&P. Can you tell people where to find you online? Are you on Twitter or LinkedIn? Is there a place where people can learn more about you and the company?
Jim: Yes. The website is panjiva.com. You can find me on Twitter @jimpsota and drop me a line LinkedIn as well.
Brian: Great. I will put the links in the show notes. It’s been great to talk to you again. Again, congratulations on the acquisition and props from Fast Company about data science, that’s really great. Hope we get the chance to talk again soon.
Jim: Always a pleasure, Brian. Thank you.
[bws_google_captcha]
Subscribe for Podcast Updates Get updates on new episodes of Experiencing Data plus my occasional insights on design and UX for custom enterprise data products and apps. Email Address [text-blocks id="eu-consent-checkbox-textblock" plain="1"] .
We’re back with a special music-related analytics episode! Following Next Big Sound’s acquisition by Pandora, Julien Benatar moved from engineering into product management and is now responsible for the company’s analytics applications in the Creator Tools division. He and his team of engineers, data scientists and designers provide insights on how artists are performing on Pandora and how they can effectively grow their audience. This was a particularly fun interview for me since I have music playing on Pandora and occasionally use Next Big Sound’s analytics myself. Julien and I discussed:
How Julien’s team accounts for designing for a huge range of customers (artists) that have wildly different popularity, song plays, and followers
How the service generates benchmark values in order to make analytics more useful to artists
How email notifications can be useful or counter-productive in analytics services How Julien thinks about the Data Pyramid when building out their platform
Having a “North Star” and driving analytics toward customer action
The types of predictive analytics Next Big Sound is doing
Resources and Links: Julien Benatar on Twitter
Next Big Sound website
Next Big Sound blog
The Data Pyramid model
Quotes from Julien Benatar "I really hope we get to a point where people don’t need to be data analysts to look at data."
"People don’t just want to look at numbers anymore, they want to be able to use numbers to make decisions."
"One of our goals was to basically check every artist in the world and give them access to these tools and by checking millions of artists, it allows us to do some very good and very specific benchmarks"
“The way it works is you can thumb up or thumb down songs. If you thumb up a song, you’re giving us a signal that this is something that you like and something you want to listen to more. That’s data that we give back to artists.”
“I think the great thing today is that, compared to when Next Big Sound started in 2009, we don’t need to make a point for people to care about data. Everyone cares about data today.”
Episode Transcript Brian: I’m really excited today for this episode. We have Julien Benatar on the show and he’s from a company that I’m sure a lot of people here know. You probably have had headphones on at your desk, at home, or wherever you are listening to Pandora for music. Julien , correct me if I’m wrong, you were the product manager for artist tools and insights at Next Big Sound, which is a type of data product that provides information on music listening stats to, I assume, artists’ labels as well to help them understand where their fans are and social media engagement.
I love this topic. I’m also a musician, I have a profile on Next Big Sound and I feel music’s a fun way to talk about analytics and design as well because everybody can relate to the content and the domain. Welcome to the show. Did I get all that correct?
Julien: Yeah, it was perfect.
Brian: Cool. Tell us a little about your background. You’re from France originally?
Julien: Yes, exactly. I grew up next to Paris, in Versailles more specifically, and moved to New York in 2014 to join Next Big Sound.
Brian: Cool, nice. You’ve been there for about four years, something like that. You have a software engineering background and then now you’re on the product side, is that right?
Julien: Exactly yes. I joined the company back when we were a startup. Software engineering was perfect, there was so much to do. To our move to Pandora, I moved to a product manager role around a year ago.
Brian: Next Big Sound was independent and then they were acquired by Pandora. I assume there is good stuff about your data. Why did Pandora acquire you and how did they see you guys improving their service?
Julien: We got acquired in 2015. The thing is, Next Big Sound was already really involved in the music industry. We already had clients like the three major labels and a lot of artists were using us to get access to their social data. I think it was a very natural move for Pandora as they wanted to get closer to creators and provide better analytics tools.
Brian: For people that aren’t on the service, I always like to know who are the actual end users, the people logging in, not necessarily the management, but who sits down and what are some of the things that they would do? Who would log in to Next Big Sound and why?
Julien: Honestly, it’s really anyone having any involvement into the music industry, so that can be an artist, obviously, try looking to try their socials and their audience on Pandora. But you can also be a booker trying to book artists in their town. We have a product that can really be used by many different user personas. But our core right now is really artists and labels, having contents on Pandora and trying to tell them the most compelling story about what they’re doing on the platform.
Brian: When you think about designs, it’s hard to design and we talk about this on the mailing list sometimes but it’s really hard to design one great thing that’s perfect for everybody so usually you have to make some choices. Do you guys favor the artist, or the label, or as you call them,the bookers or whom I know as presenters,in the performing arts industry? Do you have a sweet spot, like you favor one of those in terms of experience?
Julien: I think it’s something we’re moving towards, but it hasn’t always been this way. Like I told you, we used to be a startup or grow us to make a product that could work for as many people as possible.
What is funny is we used to have an entity on Next Big Sound called Next Big Book where we used to provide the same type of service for the book industry. If anything, it’s been great to join Pandora because then we could really refocus on creators and it really allowed us to, I believe, create much better and more targeted analytics tools to really fulfill needs for specific people like artists and labels.
Brian: I would assume individual artists are your biggest audience or is it really heavily used by the labels or who tends to...
Julien: I think it’s pretty much the same honestly. I think the great thing today is that, compared to when Next Big Sound started in 2009, we don’t need to make a point for people to care about data. Everyone cares about data today. I think that everyone has reasons to look at their dashboards and especially for a platform like Pandora with millions of users every month. Our goal is really just telling them a story about what does it mean to be spinning on the platform and the opportunities it opens.
Brian: You talked about opportunities, do you have any stories about a particular artist or a label that may have learned something from your data and maybe they wrote to you or you found out like in an interview how they reacted like, “Hey, we changed our tool routing,” or, “Hey, we decided to focus on this area instead of that area.” Do you know anything about how it’s been put into use in the wild?
Julien: Yeah, it’s used for so many different reasons. For the people who don’t use Pandora, something I really like about the platform is it’s really about quality. As you use Pandora, you have the opportunity to thumb up or thumb down songs and as you do, you’re going to get recommended more songs like the ones you like. It’s really about making sure that you get the best songs at all times.
The reality then is that for artists, their top songs on Pandora can be pretty different than their top songs on other platforms because sometimes their friends are going to be just reacting more to some part of their catalog than another one. I’ve heard many times of artists changing their playlists in looking at which songs where their fans thumbing up the most on Pandora.
Brian: Could you go through that again? How would they adjust their playlist?
Julien: Usually, people use Pandora as a radio service. While we already have internet today, most people are listening to the radio because they’re usually are very targeted and it just works really well. The way it works is you can thumb up or thumb down songs. If you thumb up a song, you’re giving us a signal that this is something that you like and something you want to listen to more. That’s data that we give back to artists. We tell them, “This are your most thumbed songs on Pandora. These are the songs that people engage with the most on the platform.” Looking at this data, you can actually inform them songs that they believe they should be playing more on the store.
Brian: I see. A lot of it has to do with the favoriting aspect to give them idea what’s resonating with their audiences.
Julien: Qualitative feedback, yes.
Brian: Got it. Actually, it’s funny you mentioned the qualitative feedback. In preparation for this, I was reading an article that you guys put out back in March about a new feature called weekly performance insights, which is really cool and this actually reminds me of something that I talked about in the Designing for Analytics mailing list, which is the act of providing qualitative guides with your analytics. A lot of times they analyze for turnout quantitative data and whenever there’s an opportunity to put stuff into context or provide qualifiers, I think that’s a really good thing and you guys look like you’ve have done some really nice things here.
I’ll paraphrase it and then you can jump in and maybe give us some backstory on it. One of the things that I think is really cool is there're concepts of normalcy in here so that, if I’m an artist and I look at my numbers, I have an idea. For your Twitter mentions, for example, you say, “For artists with 26,000 followers, we expect you to get around 44 mentions.” When you show me that I have 146 mentions, I can tell that I’m substantially higher than what my social group would be.
I think that’s a really fantastic concept that people not in music could try to apply as well which is, are there normalcy bans where you’d want to sit? Is there some other type of group, maybe, an industry, or apparent group, or another business unit, whatever it may be to provide some context for what these out of the blue numbers mean that don’t have any context?
How did you guys come up with that and can you tell us a bit about the design process of going from maybe just showing, “You’re at 826 apples,” as compared to what? How did you move from just a number into this these kind of logical groupings where you provide the comparisons?
Julien: I think what’s really fascinating is, we really live in an age of data. As an artist, you need to be on social media for the most part. There still a lot of artists I listen to but just decide not to. It’s part of things but at the same time, real big success in the music industry didn’t change. It’s still being on the Billboard chart, getting a Grammy and all these things. But as we see this, we have millions of artists looking at their data every day and just are not able to understand, like is it good or is it not good. Everyone starts at zero.
We have a strong belief that data can only be useful when put in context. Looking at the number on its own can give you a sense of how things are doing but that can also be dismissive. An example is, a very common way to look at data is to look at a number and look at the percent changing comparison to the previous week. You’ve got a bunch of tables and you look at, am I growing or am I not growing.
The reality is it’s actually impossible to always have a positive percent change. There’s no artist in the world that always does better week by week. Even Beyonce, I can assure you that the week she released Lemonade, she had more engagement on Twitter than the week after. With that in mind, we really try to give a way for artists to understand how are they doing for who they are and where they are currently in their career.
Next Big Sound started in 2009. One of our goals was to basically check every artist in the world and give them access to these tools and by checking millions of artists, it allows us to do some very good and very specific benchmarks. For an artist, like the example you said, for instance an artist with a thousand Twitter mentions in a week, is it good or bad in comparison to their audience size? This feature comes because that’s just the question we’re asked. Artists want to know is it any good? What does this number actually mean for me? That’s why we really wanted to, in some ways, get out of being a content aggregator platform and really be a data analytics platform. How can we actually give information that can help artist make better decisions?
Brian: I remember the first time I got what I would call an anomaly detection email from your service and it was about some spike in YouTube views or something like that. I thought it’s fantastic in two reasons. First of all, you identify an anomalous change and I think in this case it’s a positive anomalous change. That tells me that I should log in the tool. Secondly, you proactively delivered that to me. On the Designing for Analytics mailing list, we talk about is that user experience does not necessarily live inside your web browser interface or your hard client or whatever you’re using to show your analytics. Email and notifications are a big part of that. Can you tell me about how you guys also arrived at when you pushed these things out and maybe talk about this little anomaly detection service that you have?
Julien: It all started when we got acquired by Pandora. We decided to just invite a bunch of users and just talk to them, understand how to use our product and what did they think about it. We had artists, managers, and label people come over and we just talk to them and basically they all said, “We love it.” But then, by looking at their actual usage, they don’t use it that much. I guess one of their questions was when should I be looking at my data? Everyone is very busy. As you’re an artist, you need to perform, you need to write music, you need to engage with your fans and same goes with everyone.
When should I look at data? The reality is by being a data company, we do get all the data, we have all the numbers. We have ways to know when things are supposed to be known, when artists should be acting on something. We just turn this into this email notifications. Anytime we notice that an artist is doing better than expected, we just let them know right away.
Brian: That’s great. Do you do it on the opposite end too? If there’s an unexpected drop or maybe like, “Oh, you put a new track out and your socials dropped,” or something like that, do you look at the negative side too or do you tend to only promote the positive changes?
Julien: As far as pushes, we decided to only do push for positive. But as you mentioned weekly performance, weekly performance can give you some negative insights, like, “You’re not doing as well as artists with the same size of audience as yours.” The reason we didn’t do it for our notification is, anomalies are really hard to completely control. A reason, for instance, is Twitter removing bots. Basically, every single artist would have had an email telling them, “You lost Twitter followers this week.”
It was a lot of work to really tune our anomaly factor to actually only send emails when something legitimate happens. That’s the reason we only decided so far to do it for positive but we actually have been thinking about doing the same for negative but that’s another type of work.
Brian: Yeah, you’re right. You have to mature these things over time. You don’t want to be a noise generator.
Julien: Exactly.
Brian: Too many, then people start to ignore you. I’ve seen that with other data products I’ve worked on which just have really dumb alerting mechanisms that are very binary or they’re set at a hard threshold and just shootout noise and people just tune it out.
Julien: I’m glad you mentioned this because this feature was in beta for a year for that specific reason.
Brian: Got it.
Julien: We had to learn the hard way. We had like a hundred beta users. We’ve got way too many emails because anytime there were an anomaly anywhere, they would just get an email. For the most part, it was things that were supposed to help them. If a notification becomes noise, then that’s absolutely against its purpose.
Brian: I don’t know if everybody knows how the music business works, at least from the popular music side, but just to summarize. You have individual artists that are actually performers. They may or may not have an artist manager which takes care of their business affairs, represents them like negotiations with people that book shows. Then you have labels which are sort of like an artist manager except they’re really focused on the recording assets that the artist makes and they actually tend to own the recordings outright at the beginning and then over time, the artist may recoup through sales they make it the ownership act and the sound recordings they make. Of those kinds of three major groups, is there a one that’s particularly hungry or you’re the squeaky wheel that is most interested in what you’re doing?
Julien: I really think that into these three groups, we have a subset of users that are really into the data and into the actionability of it. I don’t think it’s one specific group of user. It could be all around the industry like we have the data-savvy, they really want to know. We have some users that actually would rather get more notifications even if they need to on their end to figure what is right from what is wrong. But since we have such a wide user base of different type of people, we decided to go on the conservative side and make sure to only share things that we thoroughly validated through all of our filters.
Brian: I assume that your group reports into some division of Pandora, I’m not sure of that. Are you reporting into a technology, like an IT, or a business unit, or marketing? Where do you guys fit in the Pandora world?
Julien: We’re part of the creator’s tools. I don’t really have a perfect answer to this.
Brian: Okay. I guess my main question being, because when we talk about designing services, we talk about both user experience, which is the end user thing and about business success or organization success. I’m curious, how does Pandora measure that Next Big Sound as delivering value? I can understand, I’m sure our artist can understand how the artists value it through understanding how is my music moving my audiences, et cetera. Is there a way that Pandora looks at it? Are they interested in just time spent? The analytics on the analytics, so to speak, is what I’m asking about. How do you guys look at it like, “Hey, this is really doing a good job,” or whatever? Do you know how that’s looked at?
Julien: To be honest, I think you said it right. Our goal is to help artists make their decisions through data and having artists use the platform is currently the way Pandora sees us doing a good job. Actually, it hasn’t changed that much since our acquisition.
One of our main KPI for the past and couple of years is something I would call insights consumes. Just making sure that our users, artists, anyone using Next Big Sound are consuming data. That can be them logging into the website or that can be them opening one of our notifications. But so far that was our main KPI. We’re trying to work on some more targeted KPI, potentially like actions taken, that would be the North Star, but we're still working on how to do that right.
Brian: Do you guys facilitate actions, so to speak, directly in the tool or are there things people can do with those actions really take place outside of the context of Next Big Sound?
Julien: There are actions that artists can take to the other creator’s tools provided by Pandora. For instance, artists have the ability to send audio messages to anyone listening to them. If they go on tour into the US, they can have targeted messages in every single song they’re going to play. If anyone listens to them there, they can just click and buy a ticket.
We’re working to make sure that artists are aware of these tools because they are free and they’re generally helping them grow at their careers. But regarding external actions, so far we don’t have any one-click way to tweet at the right time to the right people or with the right content or anything like this.
Brian: Sure and that’s understood. Not every analytics product is going to have a direct actionable insight that comes right out of it. You guys may be feeling a longer term picture about trending and maybe for a certain artist to get an idea if they’re releasing music fairly frequently, what stuff is working and resonating, and what stuff is not. I can understand that. There may not be a button to click as a result immediately.
Julien: That’s the goal though. Everything we do right now is going towards this objective. Maybe I can tell you a little about the way we think about data and that can give more sense to it.
In order to work on any new feature, we follow this concept called the data pyramid. It’s something that you can Google. There’s a Wikipedia page for it. Let me explain to you how it works. The data pyramid, it’s a pyramid formed of four layers. It could be upon each other and each representing an exquisitely useful application of data. At the bottom of the pyramid we have the data layer. Any sort of data that we may have. For our case, Android data, Twitter, Facebook just getting the numbers, getting the raw data.
On top of it, we have the information layer. The information layer is going to be ways you have to visualize this data. I guess it’s like the very broad sense of analytics. We’re going to give you tables, graphs, pie charts, you name it. We’re giving you ways to craft stories about this data but it’s on you to figure it out.
Then on top of it we have what we call the knowledge layer. That’s where things start to get interesting. The knowledge layer is the contextual part of it. It’s like, “What do this number actually mean?” It has industry expertise. For instance, the way we’re going to work about it for musicians and their true data may be different than any other industry. The knowledge layer goes like a weekly performance. It’s a perfect answer to it. It’s what does it mean for me as a musician with a hundred fans to get two mentions this week. Same for notifications. It’s telling you that you should be looking at your data right now because something is happening.
That’s how we get to the North Star and the last part of the data pyramid which is intelligence. The goal of intelligence is actionability. Now that I get to understand what does this number mean to the specific context, what should I be doing?
Following your question, everything we’re trying to do here is to get to a point where we can just send an email to an artist and tell them, “Hey, you should be doing this right now because, with all the data that we have, we believe that this is going to have the highest impact for you.”
Brian: It‘s really fascinating that you just outlined this data pyramid. I actually haven’t heard of this before. It made me think of one of the kind of, it’s not a joke but in the music community, I’m also a composer and when we write stuff, the kind of running joke is like nothing is new. Your ideas for this new song or this new melody I’m composing, it probably came before you. You heard it there before.
I wrote a post on my list that was pretty much exactly the same thing except the knowledge layer. I was calling that insight. Data have been this raw format and information being the first human-readable format that’s like say going from raw data to a chart, a histogram. Now I have a line on a chart and then the insight layer being, I have a line on the chart and another line comparing it to like you said, average, or my social group, or a parent group, or some taxonomy, or an index. Then the action or the prescription for what to do or the prediction those that kind of lead you in about action which would be that fourth state. You’re like, “Oh, is this really a new concept?” It’s like, “Nope. Someone else already thought of that.” I totally want to go read about this data pyramid.
Julien: That’s amazing.
Brian: I’ll find that link to the data pyramid and I’ll put that in the show notes for sure. I thought that was really funny.
Julien: It’s funny that you called it insight because that’s the way we call a lot of our features are working out. The way we define insight is bite-size, noteworthy, sharable content. How can we get into the noise of all of the data that only gives you exactly what you should be looking at. That’s how we got into notification and weekly performances. This is the one thing you should be looking at.
Brian: I understand what you’re getting at there. The insights are, like you said, bite-size chunks of interesting stats that someone can put some kind of context around. That’s great and it’s good. One of the things I liked, too, that you talked about was you said, “Oh we got like a hundred users, like a beta group and that kind of inspired some of this.” Your product response to how do we help people know when to come and look at our service. I think this is really good because one of the problems that I see with clients and people on the list, I think is low engagement. This is especially true for internal analytics companies. Low engagement can be a symptom of a difficult product, it doesn’t provide the right information at the right time, it may not have a lot of utility, or it’s a resistance to change. People have done something the old way and they don't want to do it the new way.
One of the recipes you can follow if you’re trying to do a redesign or increase engagement is to involve the people that are going to use the service in the design process, both the stakeholders as well as the end customers. This is especially true again for the internal analytics people. Your customers or other employees and your colleagues. By engaging them in the design process, they’re much more likely to want to change whatever they’re doing now.
I loved how you guys did some research. Now I want to ask, do you frequently do either usability testing or interviews? Is that an ongoing thing at your company or is it really just in front of a big feature release or something like that? How do you guys do this research? Can you tell me about that?
Julien: Of course. It’s consent. We haven’t released any major feature without doing some heavy user testing. I’m very lucky to be working with two designers, Justin and Anabelle who are very user-focused. Honestly, if you come to our office, at least every week we’re going to have some user interview and just talking to them, showing them prototypes, and just see how do they play with it.
Brian: So you’re doing a lot of testing it sounds like. That’s fantastic.
Julien: At the same time it’s always to find the right balance because you could be overtesting things too. We really are focusing on user testing for new things and make sure that the future that we are working on actually answers their user story that we intended.
Brian: I don’t know how involved you get participating in these, but do you have any interesting stories or anecdotes that you got from one of those that you could share?
Julien: Let me think. I do participate into a lot of them but I’m not sure I have an example right now.
Brian: Are most of the people you interview, are they current users of Next Big Sound or do you tend to focus on maybe artists that haven’t experienced the service yet or you mix it up?
Julien: We mix it up. We mostly engage with users that we already have but then we can decide to go with users that haven’t used the platform for a while, or more active users if you want to understand how we’re useful into their day to day. What I would say is that, surprisingly, it’s very easy to get users to chat about their experience with the product. I didn’t assume that we would get so many responses when we tried to have people come over or just hop on the zoom to check a new feature.
Brian: I’m glad you actually mentioned that because I think in some places, recruiting is perceived to be difficult and it probably isn’t. Maybe you haven’t done it before but as I tell a lot of my clients, a lot of people love to have someone listen to them talk, tell them all about their life and what’s wrong with it, and how it could be better with their tools. They love having someone listen to them and especially if they know that their feedback is going to influence a tool or a service that they’re using. They tend to be pretty engaged with it. I find it’s really rare that I do an interview with a client’s customer and they don’t want to be included in the future round like, “Hey, when we redesign the service, can we come back to you and show you what we’ve done?” “Oh, I love to do that!” Everybody wants to get engaged with it.
There are places where recruiting can be difficult when it’s hard to access the users, some of the enterprise software space that can be an issue sometimes. But generally, if you can get access to them, they tend to be pretty willing to participate. I’m glad you mentioned that.
Julien: I think the great part about testing with current users on the platform is to actually show them prototypes with real data, not just show them an abstract idea that we want to work on. As soon as they can see what we’re working on apply to their own career as musicians, for instance, that can lead to fascinating discussions.
Brian: You made a really good point on the real data thing. I remember as far back as 10 years ago or whenever, I use to work at Fidelity Investments, we would see this issue when we’re working on the retail site for investors. When you show a portfolio that, for example, has Apple stock trading at $22 in it, you’re not really there to test what is the price of Apple stock but you might be testing something entirely different and the customer cannot bear what is going on? They’re so stuck on this thing. It’s all fake seed data in the prototype.
The story here being if you’re a listener, when you test it’s important to have at least realistic data. You don’t want to have noise in the test or whatever your studying or else you can end up on this tangent. Try to make the numbers looks somewhat realistic if you’re using quantitative data.
In some cases, people can be taught to roleplay. Pretend you’re Drake or pretend you’re some big artist and then they can get their head around why they have billions of streams instead of thousands which they’re used to.
Julien: Absolutely. That also helps us just build better products because the reality is we have a lot of artists with maybe 10 plays in a month. As we build visualizations like something that we built a line of looking at Drake’s data, it’s not going to work as intended for a smaller artist sometimes. Having real data involved as soon as possible into the design process has been such a game changer for us. We really have a multidisciplinary team involved into the research and design of everything we do. I’m working with a data scientist, data engineer, a web engineer, and designer on a daily basis.
Obviously, we all have our things to do. But as we get into creating something new, we just make sure to have someone helping us get the real data, interview the right user, and just create prototypes as soon as possible. Working with prototypes is essential into building useful data analytics tools.
Brian: Yes, you do learn a lot more with a working prototype. It’s not to say you can’t test with lower fidelity goods, especially early on but for a service like yours when the range of possible use both the personas and also you’ve got the Drakes of the world, big major label artist and then down to really small independents, it’s really important to have an idea how your charts are going to scale, and what’s going to happen with data. Even just small stuff like how many decimal points should you be showing on a mobile device, some of the numbers might cram up.
Julien: Exactly.
Brian: All this stuff that you never think, if you only look at one version of everything, you can end up with a mess. I’m glad that you brought that up.
Julien: I couldn’t say better. The decimal is actually something that we’ve had to discover through real data.
Brian: To all of you in the technical people out there, I will say this. If I’ve seen one trend with engineers, is they love precision and there’s a lot of times when there’s very unnecessary precision being added to numbers. Such as charts and histograms. Histograms are usually about the trend, they’re not about identifying what was the precise value on this date at this time. It’s about the change over time. Showing what’s my portfolio worth down to three digits of micro-cents or something like that is just unnecessary detail. You can probably just round up to the dollar or even hundreds of dollars or even thousands of dollars in some cases.
It actually is worse. The reason it’s worse is that adds unnecessary noise to the interface, you’re providing all these inks that someone has to mentally process, and it’s actually not really meaningful ink because the change is what’s important. Think about precision when you’re printing values.
Julien: This concept of noise is so essential today for any data analytics tools. There is so much data today. There is data for everything. I think it’s our responsibility as a data analytics company to make sure what are we actually trying to help our user with this data set is not just about adding new metrics. Adding new metrics usually is just going to add noise and not be helpful in comparison to fairing what do they need to make the right decision.
Brian: Right. Complexity obviously goes up. The single verb, ‘add,’ as soon as you do that, you’re generally adding complexity. One of the design tools that is not used a lot, and this is something I try to help clients with is, what can we take away? If we're not going to cut it out entirely, can we move this feature, maybe this comparison to a different level of detail? Maybe it’s hidden behind a button click, or it’s not the default. But removing some stuff is a way to obviously simplify as well, especially if you do need to add new things. Your only weapon is not the pencil, you’ve got the eraser as well in the battle so to speak.
Julien: I couldn’t agree more. On Next Big Sound we have this concept of artist stages. It’s a way for us to put artist into buckets and by looking at their social instrument data. It goes from undiscovered to epic. We do that by looking at all of the data we have and looking at it in context.
I don’t have the numbers right now because they update on a daily basis but every artist starts undiscovered. For instance, as they get 1000 Facebook likes, maybe they’re going to get to a promising stage. We have all of these thresholds moving everyday looking at trends among social services. But what is interesting is that for instance, for a booker, a booker doesn’t need to look at the exact number of Twitter followers for an artist. He needs to know that he’s booking for a midsized venue in the city he’s in and he’s probably going to be looking for promising to established artists and not looking for the mainstream to epic artists. It’s always about figuring a way to use the numbers to tell the story.
Brian: I’m totally selfishly asking for myself here, but I was immediately curious. I live in Cambridge which is in the Boston area, and I am curious who are the big artists in our area and what is the concentration? I’m in a niche. I’m more in the performing arts market, in the jazz, in world music, and classical music but I’m just curious. Is there a way to look at it by the city and know what your artist community looks like? You guys do anything like that?
Julien: We don’t currently. But I think YouTube has actually a C-level chart available. It’s not part of something we do because I think the users it would benefit are not the users we specifically try to work on new features. It’s more something for bookers than artists ,specifically ,but it’s exactly the type of thing that we need to think about when we prioritize new features.
Brian: I’m curious just because the topic’s fairly hot. Everybody is trying to do machine learning projects these days. I don’t like the term AI because it tends to be a little bit overloaded but are you guys using machine learning to accomplish any particular problems or add any new value to your service right now? Is that on your horizon?
Julien: How do you think about machine learning?
Brian: A lot of times I associate it with predictive analytics or understanding where you might be running instead of just using statistics. I don’t know what kind of data you might have for your learning that you can feed in but maybe there’s aspects about artists that can predict. Especially, I would think like in the pop music world where there tends to be more commercialization of the music, I would say, where it’s like we need a two-minute dance track at this tempo specifically because DJs are going to play it. It’s a very commercial thing. It’s very different than what I’m used to.
So I’m curious if there’s a way to predict out how an artist may do or what kinds of tracks are performing well. Like these tempo songs, we predict over the next six months that tech house music at 160 beats for a minute is going to do really well based on the trending. I don’t know. I’m throwing stuff out there. The goal, obviously, is not to try to use like, “Oh Home Depot has this new hammer, let’s run out and get it. We don’t even know what it’s for but everyone else is buying it.” That’s how I joke about machine learning. It’s like you need to have a problem that necessitates that particular tool. I don’t ask such that, “Oh there should be some.” I’m more curious as to whether or not it’s a tool that you guys are leveraging at this time.
Julien: The Next Big Sound team doesn’t worked on features following the musical aspects of things. We really are focused on the user data.
Brian: Engagement and social.
Julien: Engagement data mostly, yes. But at the same time, I’m sure teams have worked on this because of the way that genome works. We have a lot of data about the way songs are made. Regarding machine learning, on the Next Big Song team, we actually have something that is called the prediction chart. You said predictions. We have this chart that is available every week.
Basically, it really goes back to having data for a long time. The fact that we’ve had data since 2009, we’ve been able to see artists actually get from starting to charting on the Billboard 200. By having all of these data, we’ve been able to see some trends, some things that usually happen for artists at specific times in their career up until they get into the Billboard 200.
We actually do have some algorithms that allow us to apply this learning to all of the artists on Next Big Sound right now and have a list every week of artist that we believe are most likely to appear on the Billboard 200 chart next year.
Brian: I see. Got it. Do you track your accuracy rate on that internally and change it over time? Do you adjust the model?
Julien: Yeah, we do.
Brian: Cool That’s really neat. Tell me, this chat has been super fun. I’ve selfishly got a little indulgent because being a musician, it’s fun to talk about these two worlds that I’m really passionate about so I could go on forever with you about this. But I’m curious. Do you have any advice for other product managers or analytics practitioners about how to design good data products and services? How to make either your own organization happy or your customers happy? Do you have any advice to them?
Julien: Yeah, of course. I guess it’s all about asking questions, honestly. What is very good with working at Next Big Sound is that it all started in 2009. Maybe actually I can go back and tell you the story about how it started and why it’s so different today.
It started in 2009. It was actually a project, a university project by the three co-founders. Basically, they were wondering about one thing. How many plays does a major artist get on the biggest music platform in the world? At that time, it was MySpace. The artist they picked was Akon. Basically, they just built a crawler, went to bed, woke up, and discovered that an artist like Akon was getting 500,000 plays on MySpace in one night in 2009.
The challenge in 2009 was to get the data. That’s why for the most part in Next Big Sound as it started was, I really think a data aggregation tool. Our goal was to get as many sources as possible and just make them easily accessible into the same place. We really are much into the information layer here. We’re giving you all the numbers and you can compare Tumblr to Vimeo, to YouTube, to Twitter, to Facebook, to Vine, to you name it into a table or a graph that you want to.
The reality is, today things change. We don't need to fight to get data anymore. We don’t need to hike our way into getting the numbers. Now, data is accessible to everyone in a very easy way. It’s kind of a contract. You, by being an artist, you know you’re going to get access to your Spotify, YouTube, Pandora, Apple Music or any other platform data very easily just by signing up and authenticating as an artist. That’s where our goal changes. Thankfully, we don’t need to convince people to care about data, we know they do already. But now the challenge is different. Now, the challenge is to make them understand what does their data mean and how can they turn it into getting even more data, getting into having even more engagement, and having even more plays.
I think that’s something that is very interesting because it really resonates into the question we’ve been asked in the past few years like, “What does my data mean and when should I be looking at my data?” If anything, these two things correlated pretty well. People don’t just want to look at numbers anymore, they want to be able to use numbers to make decisions. That’s the core of what we’re trying to achieve today. We couldn’t be there if we didn’t have users that ask us the right questions.
Brian: Cool that’s really insightful. Just to maybe tie it off at the end and maybe you can’t share this but what’s your home run? What is your holy grail look like? Is there a place you guys know you want to get? Maybe it’s the lack of data or you don’t have access to the data in order to provide that service. Do you guys have kind of a picture of where it is you want to take the service?
Julien: What is very noble about our goal at Next Big Sound specifically is we’re here to help artists. The North Star would be to make sure that any artist at any time in their career is doing everything they can do to play more shows, to reach to more people, and to make sure their music is heard.
Brian: Nice. I guess it’s like you’re already there, just maybe the level of quality and improving that experience over time, that’s your goal. It’s not so much that there’s so much unobtainable thing at this moment. Is that kind of how you see it?
Julien: I think the more we don’t feel just a data analytics tool, the more we’re getting to that goal. I really hope we get to a point where people don’t need to be data analysts to look at data. We’re always going to provide a very customizable tool for the data-savvy because they know what they need more than we can ever do it for them. We want to make sure that for everyone else, we can just make it very easy and as simple as a click for them to do something that’s going to impact them positively.
Brian: Cool, man. This has been really exciting to have you on the show. Julien, can you tell the listeners where can they find you on the interwebs? Are you on Twitter or LinkedIn? How do they find you?
Julien: For sure. @julienbenatar on Twitter, nextbigsound.com is free for everyone. Actually, we made our data public recently, so if you ever want to learn more about what we do, please check it out. We try to post on our blog about what we learn through data science, through design, and share more about why we build what we build. I recommend to just check blog and do some commitment to learn more about what we do.
Brian: I definitely recommend people check out the site. The fun thing is again, as you said, it’s public. If there’s a band you like or whatever, you can type in any group that you like to listen to and you can get access to those insights. Just kind of get a flavor of what the service does. I’ll put those links in the show notes as well as the data pyramid.
Julien, cool. Thanks for coming on. Is there anything else do you like to add before we wrap it up?
Julien: No, thank you so much. I love reading your newsletters and I’m very happy to be here.
Brian: Cool. Thank you so much. Let’s do it again.
Julien: Cool.
Brian: Cool. Thank you.
We hope you enjoyed this episode of Experiencing Data with Brian O’Neill. If you did enjoy it, please consider sharing it with #experiencingdata. To get future podcast updates or to subscribe to Brian’s mailing list where he shares his insights on designing valuable enterprise data products and applications, visit designingforanalytics.com/podcast.
Never forget to look up the online HTML CheatSheet when you forget how to write an image, a table or an iframe or any other tag in HTML!
.
Jason Krantz is the Director of Business Analytics & Insights for the 135-year old company, Weil McLain and Marley Engineered Products. While the company is responsible for helping keeping homes and businesses warm, Jason is responsible for the creation and growth of analytical capabilities at Weil McLain, and was recognized in 2017 as a “Top 40 Under 40” in the HVAC industry. I'm not surprised given his posts on LinkedIn; Jason seems very focused on satisfying his internal customers and ensuring that there is practical business value anchoring their analytics initiatives. We talked about:
How Jason’s team keeps their data accessible and relevant to the issue they need to solve for their customer.
How Jason strives to keep the information simple and clean for the customer.
How does Jason help drive analytics in a company culture with a lot of legacy (from its people to its parts)
The importance of focusing on context
How Jason drives his team to be business partners, and not report generators
Resources and Links:
Jason Krantz on LinkedIn
Quotes from Jason Krantz: "You realize that small quick wins are very effective because, at its core, it’s really important to get executive buy-in."
"I’m a huge fan of simplicity. As analytics pros, we could very easily make very complex, very intricate models, and just, 'Oh, look at how smart we are.' It doesn’t help our customers. …we only use about two or three different visual types and we use mostly the exact same visual set-up. I can train a sales rep for probably five minutes on all of our reporting because if you understand one, you’re going to understand everything. That gets to the theme again of just simplicity. Don’t over complicate, keep it simple, keep it clean.”
"…To get buy-in, you really got to have your business case, even to your internal customers, really dialed in. If you just bring them a bunch of crap, that’s how you’re going to lose credibility. They’re going to be like, “I don’t have the time to waste with you,” even though we’re trying to be helpful.”
"What my team and I do is we really help companies weaponize their data assets."
Episode Transcript Brian: Jason, are you there?
Jason: I’m here my friend.
Brian: Sweet. How’s it going?
Jason: It’s going very well today. How’s your Friday going?
Brian: I’m doing awesome. We’re going to talk a little bit about analytics. Is it Wile McLain or Weil McLain?
Jason: I say Weil McLain. If I’ve been saying it wrong, I’ve been saying it wrong for a while.
Brian: As I recall from my musical training, I think in German, the second syllable is the one that says its name. I guess it would be Wile McLain, like if it was W-I-E-L it would be ‘Weil.’ But I don’t know. Its anglicized as they come over the pond.
Jason: I’m going to go with you on that when you sound like an expert.
Brian: Nice. Well, you sound like an expert in analytics at Weil McLain. Tell us about what you’re doing over there. We met on LinkedIn, I’ve been enjoying your postings on the social feed about your approach. You seem really passionate about what you’re doing and I’m like, “I don’t know who this guy is, but that was really interesting.” I just have. Tell us about the company, what they do. I know they’re in heating, right?
Jason: Yes, absolutely. The company I work for, and I work in the HVAC space, we’re a 135-year-old boiler manufacturer. Whether you realize it or not, you probably have one of our products in your house or building or very close to where you live. What my team and I do is we really help companies weaponize their data assets. As you know, a lot of companies are very skilled at acquiring data since the Big Data Movement. But the reality is that a lot of these companies don’t know what to do with all this data. That’s where we really come in.
What I always tell my team and our business partners that we work with internally and externally is that our focus is on solving business problems. In order to do that, you have to identify what is the business problem that you’re trying to solve or strategic agenda that you’re trying to address. In order to do that, you really have to be anchored in the biz. Again, that’s just my perspective, but if you’re in the business day in, day out, you develop this very keen stand of what the business would need to accomplish its objectives.
Just like right now, we are based in the marketing group and it’s a great spot to be. I’m a firm believer that every analytics team should be based in the business for a reason that I just talked about. But what that does being business-first is that gives us a great lens to look at data from. Sometimes analytics people would be IT-centric and they can do a lot of academic work against the data set or different data sets. But the business might look at the output and be like, “Yeah, that doesn’t help us.” We always, always, always start with, “What is the business problem we’re trying to solve or strategy we’re looking to address?” It also helps us when it comes to curating data also. That’s one of our primary response [00:03:21] this too, is to look for different data sets both internal and external that can help us identify strategic opportunities.
It sounds really unsexy, I’m not going to lie. I think some of my LinkedIn post just say that data is boring. It really is. It’s mind-numbing, too, about 85% of my customers. But that’s the important part is understanding what do our customers need and that’s really the lens that we look at this through. We are a service provider, our customers are internal and external, we have customers just like any other business. We have to take this really boring, but really potent product in data and make it accessible to them. That’s really where we use design to really try to make that magic happen.
Brian: I love that you said, “Trying to understand what the problem is.” This is something we talk about on the mailing list quite a bit. In fact, falling in love with the problem is a good basis for doing good work instead of kind of jumping to solutions or feeling […]. As I tell my clients sometimes like, “Our job is not to go and visualize the data. It’s not […] available for someone to put into another tool or whatever the heck it is. The job is to find an insight that already is used. Probably they’re already in your job and you’re there to make […] if you’re doing internal analytics. Help them do a better job at what they’re doing, offer more value. You need to figure out how to work that into their life.” For example, for you guys then, your customer, I assume is it primarily sales people that you’re working with? Who are your customers and your […]?
Jason: Yes. Great question. One of our biggest customers is sales. Sales has been one of my biggest customers for the past 10 years of my career. I’m very intimately involved with the sales team, sales operation, sales optimization, insight gathering, pricing, things like that, but also marketing. We do a lot in terms of competitive intelligence gathering, market research. We also do a lot of operations in finance obviously related to the prices, that sort of thing. We really touch all areas of the business, but without question, our biggest customers are going to be sales and marketing.
Brian: If you were to bring a new initiative like, “Hey, we have access to…” I don’t know what it might be but for you maybe your point, [might be a line 00:05:49] of data that could actually give them more leverage. We know what the negotiation brings, better than […], we know we kind of have an idea now from what the industry is doing for their sales such that we can now tell the CRM like, “This is your […] or something.” When do you get that sales person involved? Do you deliver a solution and get feedback? Do you bring [...] early and say, “Hey, we think we can tell you more about how to do better pricing on the spot with this thing.” Do you bring them in or when do they fit into your process?
Jason: Great question. A lot of times because we spend so much time actually in the trenches, that’s one of things I think is unique about the way that I design my teams to do analytics. It’s not like hand off product and we’re like, “Godspeed. Good luck.” Once we deliver a solution, we’re actually in the trenches with the business trying to implement what we’re talking about because it just works better. The team work is just more effective and they know that they’ve got back up, they know they’ve got air support.
Really, a lot of times when we come up with something new, a lot of times we will frame it from the lens like, “Hey, we know that we’ve got opportunity A or issue B, or whatever it is. This has been an issue or an opportunity for months or years or whatever.” We think that we’ve identified something that could help us in solving that issue or realizing the potential of that opportunity and then it becomes, “Okay, let’s sit down and talk about, do you agree that this might actually help us in this process?” Because the one thing that I’ve learned is, in order to get buy-in, you really, really got to have your business case, even to your internal customers, really dialed in.
If you just bring them a bunch of crap, that’s how you’re going to lose credibility. They’re going to be like, “I don’t have the time to waste with you,” even though we’re trying to be helpful. What we found out is if you really dial in what are we trying to address with this, just as you would with any business case, and you bring that to them, I have found that they tend to be much more receptive. It’s not to say there’s not going to be resistance—resistance comes with any change—but we found that typically framing it from that lens and saying, “We’re trying to solve a problem that you have, we think that this data will help,” that’s a great starting point.
Brian: Do you have an example of a before/after with that? I don’t want you to get into proprietary stuff you can’t talk about but is there like a, “Before they did it this way,” and then we brought them in and said, “Hey, we think we can get […].” and how you went [00:08:19].
Jason: Yeah. What I can talk about is just the manner in which we distribute sales information, specifically insights. I think that, for your listeners, this is going to ring true to a lot of sales forces. I know for all them that I’ve been in or worked with, this case was true 100% of the time. But one of the things that, again, keeping the customer-centric focus, that if you look at your sales reps, a lot of time is you’re going to be what I call casual data consumers. By that, I mean that these are guys and gals that aren’t really into data day in and day out like guys like you and I or some of the listeners maybe. What we have to do is, as I always encourage my team to take empathetic lens and look at, “Okay, if we give them what our first […] is going to look like, how are they going to interpret this?” A lot of times, to be honest, it’s not very good. Now that’s where we have to look at internally and kind of rationalize and say, “Okay, let’s find this. One of us will find [00:09:14].”
But one example of that is traditionally, sales reps and sales teams will get the information in a flat Excel table. Just lots of rows and columns and just gibberish everywhere. That’s a very financial-centric view of sales data. But the reality is—I don’t know about the rest of mankind but I know for myself—I can’t remember much more than 10 numbers. The mental computational cost of extracting insights is just gargantuan. What happens is, I just don’t even bother to do it. I’m just like, “Yeah, whatever.”
An equivalent of that is, you know if you get a big block of text in email? Even though if you took that same block of text and broke it up into two paragraphs or two sentence segments which is very easy to read when you put the effort in, but for me, if I get a big block of text, I’m not even going to read that. It’s kind of one of the same things that we see on the sales side. What we do is just say, “You know what? There’s a lot of really good information here and we need to make it digestible for our customers.” That’s where we found traditionally, visualization can be an incredibly effective tool to communicate insights to this casual data audience, to this casual data consumer.
Brian: Do you have to work through the visuals with them? Do they tend to get it the first time? Is it a process of you share, “Here’s a report or here’s some new view on X.” How do you know if the visualization is actually allowing them to pull the insight out of what other [00:10:46] broad data? How do you know they’re actually “getting it”?
Jason: That’s a great question. I’m a huge fan of simplicity. As analytics pros, we could very easily make very complex, very intricate models, and just, “Oh, look at how smart we are.” It doesn’t help our customers. It doesn’t help anything. Really what we do—this is going to get to the theme of simplicity—is we only use about two or three different visual types and we use mostly the exact same visual set-up. Just to kind of frame it, what I’m a big fan of is a simple bar chart. There’s more details attached to it but to the right of the bar chart, we’ll typically put a tabular data set. What we do is, as you think in US at least, we start in that left-hand side of the page or we […]. What we do is we look at the visual real estate. We say, “Our customers are going to start in the left-hand side. We want them to look at the bar chart because it allows them to very rapidly assimilate it at a high-level what’s going on.”
It’s great at communicating at top-level churn very quickly but the trade-off is, is this horrible imprecision. You have no precision at all. What we like to do is then we address that issue by putting a simple table, very clean, very simple table over to the right. What that does is that then provides the precision that the customers are seeing in most financial-centric tables. What we found that does is that we have to train our sales team on one set-up and then that set-up is used virtually universally on all of our solutions.
As an example, I can train a sales rep for probably five minutes on all of our reporting because if you understand one, you’re going to understand everything. That gets to the theme again of just simplicity. Don’t over complicate, keep it simple, keep it clean.
Brian: I think those are good. A lot of times, when I work with engineering clients, they fall in love with consistency. I guess one point to maybe just the contrary of this is that, I think consistency is generally a good rule with design. We want to minimize unnecessary change but at the same time, I would recommend to listeners is to always look at context first, and context should always come in.
Let’s say Jason comes up with report number 12 and they have 11 now or whatever, and it doesn’t feel right for number 11. That’s a place where a designer would probably push for, “Well, no. The 12th one actually needs to be different because it’s not […] 11th and even though it’s not consistent, in this context, we don’t need it to win. This version will deliver the usability and the utility that we’re looking for better than the other 11 will.”
In general, I think it’s smart to not get creative unnecessarily with meaningless ink on the screen like, “Let’s try it this way. Let’s change the color palette. I’m tired of this.” Those are not good reasons for […], you’re just introducing noise and it’s unnecessary. But I like that you guys are thinking about simplicity and trying to reuse templates and not looking at it as a creative tableau. Ironically, people think it’s a creative “design” tool, but at the same time with all those weapons, you have a lot of different weapons you can use in that toolkit and part of that is knowing how to use this. It’s the same thing with Photoshop, a million buttons and all this stuff. The Photoshop doesn’t make you a designer. It’s being aware of your customer’s pain and the problems they need and knowing when to use all those filters and all those different things that it can do. I like that you guys are looking into that simplicity and reusing templates when it’s meaningful to do so.
Jason: You bring up some great points and I 100% agree. My team that’s listening there, they’ll laugh because I beat it in their heads, “Context. Context, context, context.” Both in design as you’ve talked but especially with numbers in general. Like, “If I give you a number, a billion, that doesn’t mean anything, you got to have context.” I’d say the same is true for design just as you articulated. Great point.
Brian: Where does the impetus for “everybody is a data company, everybody wants to do analytics”? But then there’s operationalizing that, there’s getting buy-in, leadership behind it. Where does that come from in your org? Where is the interest in taking a 130-year-old company and getting it to care about this? Where does that come from, your influence and all of that?
Jason: That was driven by our current president because he saw it as part of a digital transformation. Obviously, this was an essential component of that. Obviously, we do a lot with analytics, but we’re also involved in a lot of other digital components that lead to that overall digital maturation. Analytics is a very, very big part of what we do but it’s not all that we do.
We serve as kind of that quarterback for a lot of the digital initiatives to help basically, guide them through the process. Because even though some of the nuances of each of this project, each one will have its own nuances, they all come back to data. Data is the currency. We found out pretty quickly that if you want to stay relevant in this day and age, you need to be digitally evolved but more importantly, as you look at it, do you [compare the 00:16:02] advantage that you can derive from analytics?
I would argue that gap is slowly closing known certain industries like manufacturing, but we probably have a little bit more runway [00:16:10] it. But for a lot of industries, analytics is becoming table stakes. It’s one of those where you can certainly expect incremental value and competitive advantage, but the question becomes how much longer. That was kind of the impetus of saying, “Hey, we got to get this going sooner rather than later.”
Brian: Do you have people in sales that are resistant to using the reporting or taking advantage of your information or is it pretty ingrained in the company culture that it’s like, “This is a tool. Why would you not want to use it?” Or did you guys have a […] getting adoption?
Jason: Yeah. I would say anytime you’re going through a transformation of this magnitude, it’s hard and I would say especially for other manufacturers. I found in general, manufacturing in general, tends to be one of the laggards industry-wise in analytical maturity. Unquestionably, it’s tough for no other reason than change is tough. You’re taking legacy plants, legacy steer pieces, legacy process, and some people has been around the company for decades potentially, and we’re asking them to change almost on a dime on their time scale how they do business.
It’s not that it’s right or wrong but what we try to point out is that, as I always say, we have to acknowledge the past. We’ve been where we’ve been, we’ve been successful at where we been. But there’s been more change in the past two or three years than maybe you’ve seen in the past 15-20 years. In order to stay relevant, you really have to be ready to evolve, not only evolve but evolve quickly. But I have to openly acknowledge that that’s hard. It’s a hard proposition for a lot of people.
Again, it comes down to change management and managing not only expectations but supporting that change. Change doesn’t happen by itself, we have to support that. That’s really what we try to coach through. The way that we try to do that is by developing a product with our customers. I’m sure as you can […], if you force something upon somebody, it doesn’t get received too well. But if you develop it in conjunction with them and do tie it around their needs, it tends to get better adaptation.
Brian: You used the word product in there and I’m interested, do you see the outputs of your efforts? Primarily, it’s BI reporting as I understand it. Do you look at that as the product that you offer to sales? Is that kind of how you see it?
Jason: Yeah. We offer a product in the form of the insight packages but it’s also the service. Service that goes with it where again, we serve as essentially internal consultant to help them along. If you take just the product-centric approach, you just deliver an insight package and you’re like, “Good luck. It’s [00:19:35]. Have at it.” What we do is we deliver the product and then we partner with them and say, “Okay, here’s what we see. Now, remember you’re talking about this going on in the channel last year and our note show that there’s been a lot of competitive activity in this area. Here’s some of the question that we have. You’re the expert, so what do you think?” What we found is that working together like that, we tend to get pretty good results versus just leaving these guys on an island to kind of figure it out themselves because they virtually always know the answer but sometimes it’s up to us using these products and then offering the service is to ask question that maybe aren’t getting asked. A lot of times, we find out that they know the answer it’s just that you kind of have to ask the question.
Brian: Is that often like, “I was using XYZ report. Could you break this down by county instead of just by whatever because I feel there’s more people living in the East side of town and the average is here or […] the whole county. I really just need this one county because that’s where everyone lives. Is that really underserved? Blah, blah, blah,” that kind of stuff and then you guys will go off and work with them for more of that detail then maybe you release that back into the product as a feature if it seems like a one-off or something. Is that how it works?
Jason: It’s actually a very fluid process. An example of what you just described is exactly what happens if hey come to us with questions. But we also do it where we flip it around because a lot of products that we create are more aggregate discussion tools. We don’t design a lot within our primary visualization package. To really get into the weeds and everything just becomes overwhelming. We have other tools like your traditional [00:21:22] pivot table to kind of dig into that stuff.
But the exact example that you just gave, they will ask us those questions, but we will also flip the script and say, “Hey, we saw that the mechanical chain in the Northeast is up 50%,” I’m just making up a number, “and at a higher level, you can see that but when we segment it out, here’s what we see. Not only when we break this down to this level, we see that’s specifically being driven by A, B and C.” That gets to where I push heavier at my team to do root cause analysis. That’s really where we provide value is by digging into it and asking questions like that. Again, operating from the lens of trying to solve a problem or answer a question or root cause something in conjunction with the business. A lot of times, we will ask those questions and at the same time, they will ask us, which is great. It’s amazing because you get the better solution faster.
Brian: I think that’s great. I’ve worked on several different tools that have varying sophisticated means of doing root cause analysis and I think it’s a really powerful way to bring some why to a what that has happened in the past. Most of the time, why is really where the money is at. The value comes in being able to understand why. A lot of times, we don’t have all the data. You can’t know for sure but a lot of times I tend to say, “Our guess, if they’re just going to make a WAG—a wild ass guess—then our guess, as long as we qualify what ingredients went into the pie, our guess may be better than any WAG.” They’re going to make one already.
If they’re going to make a decision here and go off gut, there is maybe a chance they’re right and their experience will say something. But maybe our elementary root cause analysis, which we can improve over time, will actually be better and we can get out of the total guessing game and start with something that’s kind of a macro ballpark thing. Then overtime, you can improve that analysis as new data becomes available or maybe learn about how two variables are related in the business and you can bring that knowledge into the system.
I totally hear what you’re saying. It’s a nice mix of internal product plus services and also, it sounds like it gets you guys do good discovery work as well. You guys are not just responding to questions but you’re maybe asking them questions together as a group. You kind of work through what opportunities maybe latent that no one’s talking about by asking questions using data to do that.
Jason: Yeah. In the lens that we’ve been talking through, this is really sales-centric, but this applies to any group that we interact with. We have the same level of proactive discussion with any group that we interact with. In some of these, in our market research side, it’s 100% proactive. We’re going out there scouring for information and trying to see the other things that we see. That one it’s completely proactive and now we bring insights to the business and say, “What do you guys think?” The sales one is the most fun because, let’s be honest, there’s no business if you’re not selling anything and nothing happens until a sale is made.
Brian: Right. I get that. You talked about other clients, do you work at all with the actual hardware, is there any IoT type of analytics going on with the boilers and machinery that you guys create?
Jason: We’re early in that process. We actually are getting ready to go down that task very soon. On the hardware side, we tend to not have as much involvement. That’s really more on the engineering group. I think for any manufacturer product or engineering groups probably going to be the most involved in that. But obviously, we get involved into the discussions of answering the fundamental question. What are we actually going to do with this data when we collect it? Because as you can imagine, IoT can spit out a lot of data real quick. They can become incredibly burdensome very quickly if you don’t have a plan on how to manage it. But then, if you’re going to go through the effort of managing that, you got to be able to say, “What are we going to do with this?”
Brian: Yeah. I guess the first thing that would come to mind for me would be predictive maintenance, like, “Is it going to break down soon?” I worked on a cooling company that does cooling and really as the guy told me is like, “We’re not selling refrigeration. We’re selling consistent temperature to our clients. It’s not really about coolers and all of that, so we need to deliver consistent temperature. If we don’t do that, they lose products, they can lose whatever is being stored in cold storage.” That is significant business. I’m sure for you guys, it’s heat, you want to sell heat so how do you get in front if there’s a maintenance plan or whatever, how do you stay on top of that kind of stuff?
Jason: Absolutely.
Brian: [00:26:13] IoT. One of my clients used this word one time, which I now use all the time which is like, “We don’t want a metrics toilet.” An example of you can get to a metrics toilet really quickly with every stat under the gun and how many ounces of water per minute through this pipe, that’s great because that’ll help me do, as a sales guy or as a technician, how am I going to use that information just because there’s a sensor on that pipe. It’s working something around like, “Oh, there’s a sensor. Put the data in the grid.”
Jason: I’m going to have to borrow that. I’ll give credit whenever I use ‘metrics toilet,’ that’s a pretty good one. I may actually [00:26:56].
Brian: Nice. Tell me, where does it go from here? You had mentioned like, “Oh, the competitive edge, maybe it’s closing.” Or maybe you guys feel your competitors are all kind of maybe they’re doing the same thing that you guys are doing and we are all aware of where the data can be used to drive the business. Are there other places where you see design or technology like predictive analytics or machine learning and some of these other new technologies that are out there to help drive predictions and things like that? Are you guys leveraging any of that or have plans to look to the future? What does that look like? I know you probably can’t talk about everything but maybe broadly.
Jason: Absolutely. I would say that that’s content that’s definitely, if it’s not already being done then it’s on our radar. We’ve got a pretty talented team here that goes a lot of your traditional data science turf. As you can probably surmise in this conversation, is in addition to having all skills, we’re probably the most heavily focused on the business side. As we say, we explore opportunities for a lot of this. We always look at it, again, like machine learning. Great, but we got to make sure it’s very powerful stuff. We got to make sure that whatever we’re embarking upon, because we have finite work capacity, if you pursue something, machine learning, it means we’re not doing something else. It’s not to say that it’s not important, but we really have to be able to answer to that question. Again, come back to, “This is our anchor. What are we going to do with it?”
I love this stuff. I love the stats. I love machine learning, AI, all that stuff. If you’re not careful, you can really quickly get into an academic exercise that we think is really cool. “Oh wow, look at this. We’ve got this awesome algorithm here. It does all this magical stuff,” and then the business looks at it and goes, “Yeah, so what? I don’t care. How does that generate revenue? How does that improve our margins? How does that reduce our cost? How does that enable to build the sales pipeline?” If we can’t answer those base questions and we don’t get alignment, that’s probably the most important thing is executive buy-in on exactly what we’re going to be working on, why it’s important. No, we don’t pursue it but those things are most definitely, as with any analytics teams today, I think that that content is definitely being done and/or on your radar.
Brian: You make a really good point. Sometimes I almost hesitate to ask the question. But I think it’s an exciting space in terms of predictive capability and removing viable analysis and what we call time tool time in the design world, there is there. But at the same time, you make a really good point which is again, these are tools that need to be leveraged to service an opportunity or a problem. The goal is not to go do the machine learning, the goal is to solve a business problem by which machine learning maybe applied a better […] do it, reduce cost or reduce effort, speed, something like that. I completely respect that.
I’m glad to hear that you guy are looking that as not a leading step. I know there’s conflicting signals out there. I’ve been talking to people in the International Institute for Analytics about this and at the same time you hear a lot of stuff which is, “If AI is not part of your strategy, you’re going to be missing out,” and boards just want to hear that people are doing AI. At the same time, you’ve got academic exercises going on, you’ve got people trying to take on massive like, “We’re going to shoot for the moon,” and it’s like, “You don’t even have an airplane and you’re trying to go to the moon with this thing. Show us a small win if you’re going to do an investment in AI.” It’s okay to go try it out and say, “Let’s do a small thing but let’s try to solve a business problem or have some definable output that we’re looking to do here such that we’re not just writing code and doing experiments.”
I hear there’s a problem with people putting this on their resume. It’s like people just want to have machine learning. Everyone’s a data scientist now that used to stay in analytics. [00:30:47] It’s scary in the sense of just wasting opportunity and wasting money because at some point, your smarter competitors are going to be saying, “This is a new hammer. Let’s find some nails that we can use for it. But we think […] right nails and it needs to be the right application before we whack at it. It’s not just […].”
Jason: I really like your point because again, if my peers were listening to this they will laugh because they say, “We are professionals of this trade and the tools that we might want to use might not be the right tool to use for a specific job.” I couldn’t agree more with that sentiment. It’s one of those core philosophies that I have and share with my team. Also to it is with the AI. I think that you truly made a very astute observation here and comment in that, I think a lot of companies do feel compelled to have to make significant investment in AI like today. It’s not to say that there’s not merit. There clearly is plenty of merit and plenty of potential there, but kind of your point, I really believe that it’s much more beneficial when you really minimize the risk of project and budget flow and minimize overall project risk.
You take that small bite and try a little bit, then try a little bit more. When you get to win, socialize the win, and your executives feel comfortable because I’ve done it on the analytics side. I went for a big bang approach and after nine months they were like, “Hey, man. Where’s the output?” All you need is to get bit by that once and then you realize that small quick wins are very effective because at its core, it’s really important to get executive buy-in. A lot of executives are not willing to wait nine months or a year for something when they’re expecting to see at three months. I totally agree with your sentiment.
Brian: When you talked about the wins, I totally understand if you’re close to it and maybe hard to remember those, but is there a particular story or time where something in the product and the insights that you guys put out to your customers that it was like a real win, like a sales guy said something to you or maybe an executive said something to you about how this moved the needle, like this was a memorable moment for us. Like, “I changed a customer’s mind with this,” or, “We closed the sale that we never would have been looking over here if we didn’t do it.” Do you have any anecdotes like that that you can share?
Jason: One that we had recently, again, just for confidentiality purposes I can’t get too deep.
Brian: Sure.
Jason: We did have one recently where we just basically revamped our insights packages that we distribute to our internal team. We really, really gathered feedback. We had version one, we gathered ton of feedback, kind of refined, iterated, got the feedback without making it a major release. Got feedback, refined it, refined it, and then what we did was, with a small group, we got that beta in their hand, they look at it and they’re like, “This is great. This is exactly what we need.” Because what we were doing, what we found was—I’m sure you’ve experienced this—everybody wants their own part of things. Everybody wants certain view of a report or they want certain insights or whatever it is, and it’s great. But if you have limited resources, really high-powered resources like an analytics team or data science team, you’re going to look at the opportunity cost of trying to do one of these one-offs, we were getting a ton of report flow.
Again, what I tell my team, I don’t mean to be derogatory to the DI guys in this comment, but my team’s side, I always tell them, “We don’t create value if we’re just creating reports. We create value when we’re actually partnering with business to extract insights, identify opportunities amidst all that stuff that goes well with it.” What we realized though is that, what started out as a nice, clean, three- or four-page insights package and blow it up to like 20 and [34:21] doesn’t that meet our original criteria?
Essentially, what we do is once we have the rationalization enough to say, “Okay, we’ve got all these stuffs right across 20 pages. We can actually distill it down to four pages.” It will give you the exact same information, but it might not look the exact way that you wanted it to look. The question becomes, are you willing to deal with less stuff and maybe have it look a little different, but you’ll get it in a much more concise package that you’re actually able to use and process?
What we found out is that a lot of people were doing these packages and getting the reports that they want but they weren’t actually using them to drive decision-making because they can’t see the paragraph or the block of text story before. They look at it and they’re like, “I don’t know what the hell to do with this.” We would dial that in and it just been a screaming success. It’s really nice to have it where something like that you see the evolution of it. This is just one of those things that we had, and this was kind of a side package or wasn’t a primary, but it’s become a primary now because it’s so effective.
Brian: Would you say that when you talk about reducing this, is the report like a PDF or do they access it through a browser the insights package?
Jason: Yeah, we have the options to do both. We distribute it initially via PDF, sometimes along with our comments if there’s really, really big stuff in there. We’ll say, “Hey, we see this. Here’s a driver. Here’s a supplemental package.” A lot of times it’s PDF first and then if they want to go on the web, start interacting with it, they can do that. Those are nice, but the reality is a lot of them don’t do that which is understandable.
Brian: You took it from 20 pages down to 4, is that what you’re saying?
Jason: Yeah. Same information.
Brian: This is a really good point. I’ve frequently had clients come in and they’re with data products and their concern is information overload. We’ve heard this a lot of times and the irony is that, the issue is usually not information overload. It’s usually a design problem that the information is not presented properly because sometimes, it can increase the density and increase the utility and usability, not the other way around. In fact, removing data can actually make it worse.
A basic example of that is when you’re trying to compare A and B. If A and B are not on the same, what we call a viewport like in a browser world, it would be within the browser window there. When you require someone to toggle between two screens, they have to change context and visually, your eye can process the information a lot better when it’s within proximity. Sometimes, increasing the density actually will give you a better design. It takes more care in how you do it, but it’s not always about information overload, “Oh, it’s too crazy.” They may not get it on the first time but your sales people, if they’re looking at this stuff weekly or monthly, at some point they’re going to be pretty comfortable with this.
I always tell my clients, “You need to look at the switch frequency as well because if it’s going to be used a lot, you can actually get more detailed and you can really push the, what you might see as complexity or the information density, can go up because they’re going to get familiar with the formatting. Typically, the density is actually going to probably improve the utility as long as care is given to the choices. But having that eyeball comparison without having to change pages and all of that, typically you’re going to give a better story as a broad rule. I like hearing that you guys went down in page count, up in density and in turn a better user experience at the end so that’s great.
I think we’re about done here. I don’t have too many questions for you, but this is super great. One of the reasons I contacted Jason is because I remember seeing this quote, “Jason is like a category five hurricane in the data analytics world.” I’m like, “Who the hell is this guy? No one talks like that.” I started reading your stuff and I enjoyed watching your LinkedIn social posts and things like that. Where can people find out more about you? You’re obviously on LinkedIn, I can put LinkedIn in the show notes and stuff, but are you on Twitter, any social media places they can follow you?
Jason: No, actually, I’m not on Twitter. But the best place unquestionably is going to be LinkedIn. I’m pretty involved there. I do like to engage. If you want to direct message me with questions, just talk, meetup, connect, whatever it is, I welcome that. I love the platform, it’s a great family. I just really started using it maybe nine months ago, really getting into it. It’s been great meeting guys like yourself. It’s actually phenomenal.
Brian: Cool. I’ll put a link to Jason’s LinkedIn profile on there and you guys can find him. I recommend, especially if you’re in an internal analytics type of role at your company, to follow Jason and then check out what he has to say on there. This has been great. Thanks for coming on the show. I look forward to meeting you at some point in person.
Jason: Dude, thank you for having me on here. I really appreciate it.
We hope you enjoyed this episode of Experiencing Data with Brian O’Neill. If you did enjoy it, please consider sharing it with #experiencingdata. To get future podcast updates or to subscribe to Brian’s mailing list where he shares his insights on designing valuable enterprise data products and applications, visit designingforanalytics.com/podcast.
Never forget to look up the online HTML CheatSheet when you forget how to write an image, a table or an iframe or any other tag in HTML!
[bws_google_captcha]
Subscribe for Podcast Updates Get updates on new episodes of Experiencing Data plus my occasional insights on design and UX for custom enterprise data products and apps. Email Address [text-blocks id="eu-consent-checkbox-textblock" plain="1"]
.
Vinay Seth Mohta is Managing Director at Manifold, an artificial intelligence engineering services firm with offices in Boston and Silicon Valley. Vinay has helped develop Manifold’s Lean AI process to build useful and accurate machine learning apps for a wide variety of customers.
During today’s episode, Vinay and I discuss common misconceptions about machine learning. Some of the other topics we cover are:
The 3 buckets of machine learning problems and applications.
Differences between traditional product development and developing apps with machine learning from Vinay’s perspective.
Vinay’s opinion of what will change as a result of growth in the machine learning industry
Maintaining a vision of a product while building it
Resources and Links:
CRISP-DM
Ways to Think About Machine Learning by Benedict Evans
The Lean AI process
Vinay Seth Mohta on LinkedIn
Big Data, Big Dupe: A little book about a big bunch of nonsense by Stephen Few
Quotes from Vinay on today’s episode: “We want to try and get them to dial back a little bit on the enthusiasm and the pixie dust aspect of AI and really, start thinking about it, more like a tool, or set of tools, or set of ideas that enable them with some new capabilities.”
“We have a process we called Lean AI and what we’ve incorporated into that is this idea of a feedback loop between a business understanding, a data understanding, then doing some engineering – so this is the data engineering, and then doing some modeling and then putting something in front of users.”
“Usually, team members who have domain knowledge [also] have pretty good intuition of what the data should show. And that is a good way to normalize everybody’s expectations.”
“You can really bring in some of the intuition that [clients] already have around their data and bring that into the conversation and that becomes an almost shared decision about what to do [with the data].”
Episode Transcript Brian: We got Vinay Seth Mohta on the show today. I’m excited to have you here. Vinay’s maybe a little outside the normal parameters of who we planned to have as a guest on designing for analytics but not entirely. He has an engineering background but he’s done a lot of stuff in the product management space as an executive. Correct me if I’m wrong. You’ve been at MathWorks before, you worked on search at Endeca Technologies, and you were at Kayak, which is one of my favorite sites, actually, for booking travels. I’m sure everybody listening has probably touched Kayak at some point, and you were a product manager there, correct?
Vinay: That’s correct, yup.
Brian: Okay, and I know you did some healthcare. You were a CTO at Kyruus, and now, you are a Managing Director of Data Platforms at manifold.ai, which is a services company that works on data science, machine learning projects, and artificial intelligence. Is that correct?
Vinay: That’s right, yup.
Brian: Tell us a bit about what Manifold’s doing and what you’re doing there.
Vinay: Sure thing. Manifold, as an organization, is an AI consulting company, as you mentioned. More importantly, we unpack AI into [...] really focus on data engineering, data platforms, getting your data ready, and then also building machine learning models and getting all of that put together into either an internal-facing or an external-facing product. So, I’m looking forward to talking a lot more about that.
As a company, we largely work with Global 500 organizations and also a spectrum of organizations. Sometimes, I actually get down to fairly early stage startups, where they’re looking for very specialized help in a particular area like Computer Vision, for example. We are largely a team of experienced product folks and engineering folks who’ve worked at both large organizations like Google and Fullcom as well as venture-backed startups like some of the companies you’ve mentioned in my background.
Brian: What kinds of projects are people coming to you guys with? Obviously, the whole AI machine learning thing is a pretty active space right now. Everyone’s trying to jump on to that and you got to invest in this. What kinds of projects are you guys doing?
Vinay: That’s a great question in terms of the different places and the different motivations people have when they come to us. I try to demystify AI right from the first conversation. Particularly, when we’re talking to executives, which we often do, we want to try and get them to dial back a little bit on the enthusiasm and the pixie dust aspect of AI, and really start thinking about it more like a tool, or set of tools, or set of ideas that really enable them with some new capabilities that also can be thought of, and what I at least see as some more traditional product development spectrum.
That’s really what I like to use to frame where customers are when they come to us. By the product development spectrum, I mean there is a starting point of what are the right questions to ask and what are the right types of business strategy questions I should think about, go to market-type questions that might be relevant to consider.
Some customers that we've talked to are starting all the way back there. There are folks who’ve answered that question for themselves, and now, they’re actually starting to think more actively about what are the product-related areas I want to invest in based on my overall business strategy, what are some of the technology approaches I can take. Machine learning is not always the right answer for a pretty business problem and then really getting into more of the actual design and architecture pieces, and then the hands-on keyboard of actually building, and then deploying data engineering, related data pipelines, or machine learning models, for example.
We’ve really seen clients come to us at all different phases. The parts we generally like to focus on start from the product strategy, technology strategy-type conversations, going all the way to building and delivering software and machine learning models that are going to get deployed into production. So, that’s really our zone of focus.
Brian: If I could take it back for one second, you said pixie dust and I thought that was funny. But I also get what you’re saying in there. Do you think, as consultants and service providers working in the space—I work on the design side, you’re working a lot on the engineering side and the data science side—are we propagating the wrong thing when we say artificial intelligence and in the analytic space, the term big data?
Stephen Few just wrote a book, I think last year, they called Big Data, Big Dupe. I tend to agree with it. There’s a lot of marketing hype surrounding the term. No one can really even define what makes it big versus regular. Do you think we have to stop using that as that? Does it matter what we call it? I feel kind of silly every time I say “AI” because it has such a loaded meaning to people that maybe don’t know as much about it. What do you think about that?
Vinay: I generally agree with the spirit of your question, which is, it’s just good to use words all of us understand that map to things that we can touch when we type with our keyboards and things like that. So, it’s very helpful to talk about software engineering as oppose to AI for example or a machine learning model.
I’ve also come to terms with the fact that there is a massive marketing wave that is much larger than what you or I choose to do and I think that creates the context that someone is coming into a conversation with us. When they enter the conversation, they already have some of that context. So, what is more important for us to focus on, as opposed to the specific choice of words, is really taking where people are starting in a known context and then walking them into either a world where we feel we can have a much more real conversation with the types of things that are grounded and the actual work that we do. A lot of people are uncomfortable with terms they don’t understand but they believe they’re supposed to continue using them and they should understand them, et cetera. I also find the other thing that’s nice about taking in marketing term but then really almost using it as an educational opportunity when you’re unpacking those terms.
People start to feel more comfortable that, “Oh, okay. These things can be mapped into things I understand,” and then being able to use some much more effectively. At least, in our conversations with them, we have a shared vocabulary. I often bucket those conversations under recognizing that this is a marketing term. “Let’s talk about what you mean by AI and let me unpack what I mean and make sure we have a shared vocabulary.” I think there’s some nice ways to undo the marketing hype in more intimate settings, but at a larger scale, I had found that anytime I try to fight the marketing, the five-year macro trend marketing term, people mostly say, “Oh, you don’t do anything related to that and you do this after-effect.” And it’s like, “What? No, no, no. That’s not what I meant.” I think we have to pick our battles.
The other thing which I always have mixed feelings about but it does feel like—and I’ve seen this with several of the major technology trends over the last two to three decades—is that it does motivate organizations that traditionally wouldn’t look at technology as enabling components of their business strategy. It does force them to at least take a look, revisit new ideas that may have been scary before. But now they feel like, “Oh, well, let’s at least take a look because it seems everybody else is getting some value from it.” It does at least stir up things inside organizations where you get some creativity going and people are willing to at least step out of their day-to-day and take a look. I’m definitely not a hype person in general, but it does seem to serve at least some positive purpose in that sense.
Brian: I kind of see it—we’ve joked about this in the past offline—like there’s a new hammer at Home Depot and everyone’s racing out to go buy this tool but not everyone knows what it does. It’s just, “I got to have one like everyone else. It does everything.” On that thought, of the ten people, ten clients that come in, what role would your typical client be? And of ten of those, how many of them have either unrealistic expectations of like, “Hey, we want to do this grand project with AI and machine learning to do X,” versus, “Hey, we want to really optimize this one part of our supply chain,” or, “We want to do…” something very specific that’s been thought of in terms of either products or service offering or an internal analytics thing where they want to actually apply an optimization or something like that. How many had fallen to the “educated versus maybe less educated,” in terms of what they’re asking for from you?
Vinay: I would probably say order 20% to 30% of folks are in that bucket of, “I have a very targeted need. I know exactly what I want to get out of this state of pipeline. I have this other data pipeline I’d like you to work with to put the whole thing together,” or, “I need a specialized machine learning model that will help me segment some of my customers into more fine grain way for this very particular use case,” things like that. Those tend to be organizations that already have a software engineering capability. There’s some data for other business problems already and they either need more help than they have in house or they need some kind of specialized help. So maybe, they have largely done more structured data marketing-related use cases and now, they want to do more natural language-related or in a different area.
They generally have a fairly good feel of the landscape and they know how our work would plug into their work. There is probably roughly 50% of what we get as more where we get people who are VPs of Technology, VPs of Product. They understand operations in a pretty meaningful way. A line of business leader who has a meaningful business case in mind, so they already have one or more business problems in mind that they think will be compelling. They want to know, is this a good fit for a machine learning or not? What would be required to actually get to even trying out machine learning?
I would put those folks in the bucket that they have thought through some of the business strategy related, sort of going back to that spectrum idea of starting from business strategy all the way to shipping something to production. I would say they are more in the product and technology strategy bucket where they want to figure out, “I don’t know what I have in the rest of my organization, but I know we have some software, we have some data based on running a website for the last four years, whatever else, or some other kind of operational system. I’d like to figure out if we could use machine learning in some way to do something predictive, for example to improve how a call center handles inbound calls and prioritizes some of the tasks.”
There are cases where people have much more thought through use cases in mind, but they don’t have the expertise on: What is the data pipeline? What data do I actually need from machine learning? Have I actually ever built and deployed a model before? They've usually not have done that. There’re a lot of folks in that bucket. And then, the third bucket is the remainder, which is really people are starting more in the business strategy side, where they’re saying, “Oh, we’d really like to have an open-ended conversation. Our CEO has a five or ten-year vision around transforming our core business and how we service our customers.” I’ve talked to folks that are in much more traditionally industrial businesses like paper processing, for example, or staffing, or more instrument manufacturing, or other types of manufacturing.
Those kinds of areas, there is really this historical model of hardware or some other service that gets provided as opposed to Software as a Service. I think everybody is interested in some kind of move to a subscription model and also some understanding of what is the relevance of these technologies. But they are not at the stage where they’ve identified a particular business case or a use case.
Brian: If I’m a product manager or someone that’s in charge of bringing ROI to data within my company, say I’m not a technology company, should I be looking to make an investment in a place where maybe it’s more of a traditional analytics thing or maybe I have humans doing eyeball analysis, making decisions about insights from the data, and then saying, “Okay, what we’d like to do is actually see if we can automate this existing process. So, it’s like A, B, C, D, E, F. We want to swap out stage D with a machine learning solution to free those people to do other work”? Or is more like, “We have this data we’re sitting on. Hey, we could train it and do something with it. We’re not doing anything with it right now.” Is there a strategy or some thinking around one of those maybe being a more successful project to take on, any thoughts?
Vinay: I think that’s a great way to pose the question because one of the things I would think about as with any new effort in an organization, is that you want to be successful as the person who’s bringing in some new technology or new approach, whether it’s process or people or technology. I think really having a lower risk, a smaller bite at the apple in some sense to get your first success on the board, and then starting to build on that nucleus would definitely be the way I would think about get it going.
There may be different situations where, as a leader of a large organization, you really have a directive to be more transformative and that can be a different type of conversation. But as I’d think about somebody who’s in a product role at—let’s call it just for the sake of brevity—a non-tech organization, I think starting with a smaller project where you can get people used to the idea that you could do more with data, it’s not that scary, it’s like another tool, it’s like buying another piece of software and doing some training around it and those kinds of things, then it gives you a success that you can build on and people around you start to have some familiarity with it, where you get less resistance the next time you go and do some things. I think of the overall change management challenge would frame the choice of project in some ways than not.
One of the other frameworks I would use also, Ben Evans from Andreessen Horowitz, recently wrote a really nice blog post about how people can organize their thinking around applications of machine learning. The core of the framework is, there are three buckets in which you can think of the problems and potential applicability of machine learning. The first one, actually, falls very much into exactly the example you gave where I might have an analyst working with existing data, etcetera. That’s ‘a known data, known questions’ bucket. So, you have a set of data already available. You have a set of questions your analysts ask every day. Maybe they’re eyeballing it. Maybe they’re running a simple linear regression or something.
What’s nice about applying machine learning in that case is it’s literally like, “Oh, you have a mallet. Here I have a stainless steel hammer. Let’s see what happens if I apply my stainless steel hammer.” It’s relatively easy to get set up to do it. Our organization who knows roughly what’s already involved with that data, the semantics of the data. It’s clean enough that you could probably start working with it. It gives you a relatively easy pathway into trying out machine learning. Just saying like, “Oh, we got 50-basis point lift just by applying this new tool, without really changing anything else.” That’s one bucket.
The other two buckets, I definitely encourage folks to read the article, to put in the show notes or something. The other two buckets are ‘unknown data, new questions,’ and then the last one is ‘new data, new questions.’ Just to give you a placeholder for what the last bucket is, those are opportunities that you might be able to apply computer vision or put new sensors in a particular environment. So, gathering entirely novel data streams, unmasking new questions. There’s a handful of organizing ideas like this. We generally suggest a few different articles and I am definitely happy to offer those for the show notes as well, if [you’re looking for 00:17:27] different ways to organize their thinking around approaching machine learning problems.
Brian: Great. Yes, I’ll definitely put those links into the show notes. Thanks for sharing those. Also, a follow up to that. Once you’re into a project, what are some of the challenges around for projects that have user interface or some kind of user experience that’s directly accessed? Are there challenges that you see your clients having with getting the design right? Are there challenges about getting the model and the data science part right or getting it into production? I heard a lot about this at Strata Conference that I was at in London, that they’re talking a lot about you can do all this magic stuff with your data sciences in the PhDs. But if they don’t know how to either help the engineers or themselves get that code into a production environment, it’s just sitting in a closet somewhere and it’s never going to really return value. Can you talk about some of the design and the engineering challenges that you might be seeing?
Vinay: I’m assuming most people listening to the podcast are familiar with traditional product development processes, design iteration, and so forth. What I’ll offer here is the difference when you start thinking about data and machine learning. We have a process we call Lean AI and what we’ve incorporated into that is this idea of a feedback loop between a business understanding, a data understanding, then doing some engineering—this is the data engineering—then doing some modeling, and then putting something in front of users.
The major part here is that, you may have a particular idea around what the ideal user experience might be. But then as we start to get into the data, as we start trying different modeling techniques, we might either surface additional opportunities that there may be something compelling that the user could do in their workflow using what the model has surfaced. Or it may be that the original experience as envisioned is going to have to change because there is not enough predictive power in the data, or a data source that you thought you’d be able to get your hands on is just not going to be available, or things like that.
So, there is an additional component to the [iteration 00:19:46] loop that you have to rely on, which is just what is in the data, how much can I get access to, and then some of the more traditional software engineering constraints. If it’s going to take six months to get that particular piece of data cleaned up enough such that we can actually use it, is there something lighter weight that we could at least get started with at something in front of users first, and then continue to refine and iterate over time? That’s probably the big difference in terms of traditional product development that just involves software engineering in apps versus working with the data and machine learning. There’s a little bit of just this science of what is possible inside of the data given the signal inside [00:20:27] them.
The engineering part is definitely, as you said, something that is talked less about historically and it sounds like, based on some of the things you’ve heard at Strata, that is something that is starting to change. What I’ve seen is that a lot of the tutorials, a lot of the content out there has historically been focused on, “Get your first model going,” or, “Take this particular data set and try out building a model or tweaking this or that.” In that sense, there’re also a lot of tools available for doing data science and data science exploration.
It’s great that, exactly like you said, Brian, that somebody’s built a model that’s interesting. But one, if we haven’t built the rest of the product around it and then if we haven’t actually got that model to production; as I like to say, if at the end of the day somebody’s not pushing a button differently because of your model or pulling the leverage differently because of your model, it really doesn’t matter that you built it in the first place. That actually goes back to requiring engineering and product development type expertise as opposed to data science type expertise, which I feel a little bit more like traditional on science type disciplines where you’re doing experimentation.
Brian: Do you get to the point where you’re midway through a project and just kind of like, “We’re not sure if we can do this,” or “The predictive power is not there”? I imagine you probably try to prevent getting into a situation where that happens. Is there a client training that has to go on if they’re coming to you too early? Like, “We’re ready to build this thing. We want to put a model to do X,” and you’re like, “Whoa.” How do you take them on like, “Come back to us in two months or when you guys have figured this out”? How do you take them on to make sure that doesn’t happen and they don’t spend all this money on hiring data scientist internally to work with you or on their own, or just you and not getting an ROI? How do you educate on that?
Vinay: That’s again what we have incorporated into this Lean AI process where we’ve taken the spirit of Agile and some of the ideas around Lean startup, for example. There’s actually an old framework from the late 90s called CRISP-DM—it’s from the data mining community—and really, the idea in all of these things is tackling your big risks early and surfacing them. We take a similar approach where anybody can do this. But it’s getting an understanding of what is the business problem you want to go after and what is the data you have available. We call it a business understanding phase and a data understanding phase.
During that phase of the data understanding, it’s really doing a data audit. Particularly, it’s an issue on large organizations. People think they have access to certain data but it may be that somebody in a different organization owns the data and they’re not going to give it to you. You sort of have the human problems that we’ve always had. Then there’s other parts which are, “Is there a predictive power in the data? Is the data clean?”
Generally, the first thing we do is just apply a suite of tools that will characterize the data, profile the data, and help us get an understanding of what do we think is there. Usually, we found working with clients, team members who have domain knowledge. They generally have pretty good intuition of what should the data show and that oftentimes is a good way to normalize everybody’s expectations.
As an example, we’re working on one with an industrial client last year. In addition to sensor data coming off their devices, they also had field notes that people had entered when they were servicing some of the equipment. As we were working with their experts during the data understanding phase, the experts actually said, “You know what? I wouldn’t trust the field notes. People sometimes put them in and sometimes they don’t. The quality varies a lot across who put those notes in and what they put in there. So, let’s just not use that data source.” You can really bring in some of the intuition that people already have around their data and bring that into the conversation. That becomes an almost shared decision about what do we think we can try and get out of this data, what’s in the data, and do you guys agree that this data actually is saying what you think it should say? Those kinds of things.
I would say, tackling big risks early is one of the major themes of what we do. The other part really comes from, again, the engineering approach that a lot of us have taken historically from our past experiences. [Probably 00:25:48] the best analogy I can do from their product management days is this idea of just doing mockups and doing paper mocks and those kinds of things before you get to higher fidelity mocks. There’s a similar idea in machine learning where we have this idea like, “Okay, get some basic data through your data pipeline. It doesn’t have to be perfect.” Then we build this thing called the baseline model, which is, “Yes, there are 45 different techniques you can use to build a machine learning model. Let’s take one of the simplest ones. Something like random forest where we know that’s not the best performing model for every use case, but it’s really easy to build. It’s really easy to understand at least out of your first version what the model is doing.”
You can get some baseline of performance pretty quickly, which is, does it perform at 60% or does it perform at 80%? From there, you can start to have a discussion about, how much more investment do we want to make? Do we need to get more data in here to clean the existing data and transform it in different ways, explore different modeling techniques? Those kinds of things. I draw the analogy to some of the product development processes that we would follow if we were just doing software engineering project, which is, let’s get something built end-to-end then add more functionality over time, things like that and then take it from there.
Brian: Most of the time, the projects you work on, are your clients the actual end users of this service or the direct beneficiaries, or typically, are they building something internally that will be used by other employees or vendors or their customers? How close to that person that’s going to benefit from or use the service that you’re building?
Vinay: I’m definitely not aware of all of our projects, but the projects I’m aware of and the ones I’m working on right now, they all have enterprise users. None of them are applications that are going to go out to end users. But nonetheless, the enterprise users are folks who are not technology people or not particularly specialized in data or anything like that. They are more folks who are executing on processes as part of a broader workflow. For example, it might be a health coach that is at a particular company, or it might be a call center employee, or it might be the maintenance and repairs center at an industrials company. It’s more internal users or if it’s external users, it’s still again enterprise users who are using a larger product.
Brian: Do you ever get direct access to those when you’re working with your clients or typically, is your client the interface to them? How involved do you get with some of these like a call center rep or something like that?
Vinay: It actually depends on the type of expertise that our client has. If they have a product owner and a product manager who’s fairly confident about their ability to interface with the end user, we might. Instead of them being part of the user feedback sections, as some of these models go in front of users, there may be at the beginning of a project, having a few conversations to understand the context in which particular operational data was gathered, or the workflow that might surround the model that we’re building, or the data pipeline that we’re building. We might have a few conversations.
But again, if they have a strong product function already, we would probably be more isolated from that. If, on the other hand, there isn’t that much of a product function that is familiar with software engineering and product developments, some of these non-tech organizations, product managers, they are maybe much more hardware-oriented or they may not even have a product to roll, depending on the type of operation. There, we would be much closer to the end users understanding the use cases.
We also want to partner with whoever is doing the product design and some of the other UX components as well. I would imagine that there’d generally be another partner of some sort. We’re interested in talking to the end users. But we’re definitely not the experts on product design and so forth. We’d expect somebody else to play that role. Either somebody like you where the client is partnered with another organization or individual, or they have capability internally.
Brian: One place we think of lots of data, obviously, is in the traditional analytics space for internal companies or even information like SaaS products and information products. Do you see the capabilities of data science and machine learning that have really been enabled in the last few years?
Primarily what I understand is there’s more data availability. There’s more compute power availability. It’s not so much that the science is new. A lot of the science I hear is quite old. The formulas and algorithms have been around. It hasn’t been as feasible to implement them. Now that it is, do you see that traditional analytics deployments over time will start to leverage more and more like predictive capabilities or prescriptive analytics where there’s less report generation, less eyeball analysis?
Say, in the next five years, 20% of traditional analytics capabilities will be replaced by more prescriptive and predictive capabilities because of this? Or is it really just it’s going to take a lot longer to do that? I imagine some of it’s just at the mercy of the data you have available. You can’t solve every problem with this, but do you see an evolution happening in that data? Is that making sense?
Vinay: Yeah, absolutely. You’ve hit upon a really important idea. I’ll start my answer though taking a slightly different view, which is what is going to stay constant, and then we can talk about what is going to change. The part I found most exciting about business intelligence, analytics reporting, pick your category name, is when you can get it embedded into a workflow. The folks who are actually on the front lines making, running through a workflow, or going through a customer interaction or whatever, they actually have access to that data and they’re able to drive decision-making as part of their process.
What we’ve seen in the last order of 20 years, is this continued increase of this notion of a data-driven organization, that people should have more access to data when they’re in these workflows and decision-making. Everything from things you’ve probably heard about, like insurance companies or telco companies, call center folks being able to offer you something if you’re going to turn, for example. An offer pop ups on their screen and they’ll able to give that to you. That’s a nice example where somebody’s actually using the decision-making as part of their production workflow. We’re just generally seeing more of that. So, no matter what, whether it’s prescriptive or descriptive, whatever else, I broadly see continued adoption of analytics and data in more workflows across a whole range of software products.
I’m generally excited about that. I wish it would take less time but at least we’re continuing to make progress on that. I think what you hit upon is what’s going to change. I firmly believe we’re seeing this in name today but we’ll see this more in actual. The nature of the work itself in the future, there’s a lot of people who have the business analyst role today and organizations in their supporting different functions. Largely, I think of them as people who have a fairly deep understanding of the business. They generally live in Excel. They’re complete masters of Excel. They can build what-if models, they can do scenario-solving, they can do VLOOKUPs, and do all of those kinds of things in Excel. I think they’re going to get a whole additional set of tools.
I tell people this and I’m going to go on the record here and suggest that, I’m almost imagining Excel 2020 is going to have a button that you can hit and you can say, “Here’s my data. Go try out 50 different models or 500 different models.” Excel will go off, ship your data to Azure, it’ll run a whole bunch of different models and come back and tell you, “Here’s the three that seem to fit your data best.”
Really, the skill that you need at the end of the day, which is the skill you need today, is understanding the statistics of the data, having some intuition around the business and what’s going on around you, and then really being able to swap ends and these other statistical methods that we group under machine learning, being able to swap those in once those tools are mature enough for broader use in deployment. Because of that, I think yes, in the five-year timeframe, we’ll see the leading edge of more prescriptive analytics entering product workflows just like we’re now. I’d be curious about your opinion on this but I feel like we’re past the earlier doctrine more now in the mainstream phase of descriptive analytics entering some of the different products.
Brian: Yeah, maybe it’s fed Microsoft a little tip for how to improve their office lead down a couple of years from now. This has been really informative. Thanks for coming on. Do you have any single message or advice you’d give to data product managers or analytics leaders in businesses in terms of how they can design and/or deploy better data products in their organization or for their customers if they are like a SaaS or information provider? Any general tips you’ve seen or something you can offer them?
Vinay: Maybe a handful of things just to run through it with different levels of applicability. One of them is that having a good business case, as the way we talked about earlier and taking on something small is definitely very helpful to build some success. Also, maybe squelch some of the visionary enthusiasm that people might have. In general, trying to feed some of the vision component while you’re trying to get a great concrete success on the board, is something just to keep in mind to get people excited about the potential and the future. That’s one bucket.
If you have a vision in mind, one of the things your technology teams and your machine learning teams can do, and is something we definitely ask for when we do our engagements, while you’re solving a specific business case and a specific problem, you can do the work in a way that lays the foundation for longer term leverage on the work. So, if we build the data pipeline, we know that you have a specific two-year vision. We can actually start to lay some of the pieces even as part of that project to make investment towards that vision. While you should execute on smaller opportunities, you should also dream big. I think that’s one general thought.
Another thing I’ve been starting to form an opinion around is that, to execute successfully on a product and execute data and machine learning component of a product, you have to have a ‘what’ in mind, like, “What is this product going to do with the data?” You need to have a product direction, product sense, product vision, whatever you want to call it to know what’s going to happen in the context of that product.
Longer term, when you start to think about the context for these kinds of capabilities you need to think about organizational vision. For this product it may be that you did it with a couple of folks from another team that sat down the hall just to get something out the door. But then, really having an idea in the 18-month timeframe, do you want to build a software engineering organization? Do you want to build a data engineering capability? Do you want to have a data science team? Do you want to work with the finance team to maybe get a couple of business analysts over to a new team? I think really starting to contextualize your product vision with what’s your organizational vision, is important for the longer term picture and having clarity around that even as you tackle on the shorter term opportunities. Those are probably a couple of things that hopefully people find helpful.
Brian: Yeah. I definitely did. I was actually going to follow this up but it may be an unnecessary question. But one of the services that I’m often asked to come in with clients is to help them either envision a new product, something that they’re working on, and it’s what I call getting from the nothing to something phase where it’s a Word document of requirements or capabilities, features, what have you and getting to that first visual something. It sounds like you still think that that step, even if you don’t bite off the whole thing from an engineering standpoint, having an idea of your goal post about where a service might go that could incorporate some machine learning or AI technology, still is helpful and deploying a small increment of utility into the organization. Would you agree with that still?
Vinay: Yeah, absolutely. Even for the folks building their models or building your data pipeline to get the data cleaned up and usable, whether it’s for analytics or for your models, it’s really helpful to have that broader context as opposed to having a very narrow window into, “Oh, I need these three fields to be cleaned up and available.” If you can’t provide that broader context, I feel you end up with a lot of disjointed pieces as opposed to something that feels good when you’re done. I would definitely agree with that.
Brian: Well, Vinay, thank you so much for coming on. This has been super educational for me and I’m sure for people listening as well. Where can people learn more about what you’re doing? I’ll definitely put the Ben Evans link and your Lean AI process that you talked about. So, send me those links. But where can people learn more about what you do?
Vinay: Our website manifold.ai is definitely in the best place to start. We have a few things about the type of work we do and some case studies as well as some background of our team. That would be helpful. In terms of my own time, I actually don’t spend that much time on social media. LinkedIn is probably the easiest place to find me. Generally, I post things there occasionally and definitely participate in some conversations there. It would be great to chat with folks there.
Brian: All right, great. Well, thanks again and I hope to talk to you soon.
Vinay: Thank you, Brian. I really appreciate it. It’s great conversation.
[bws_google_captcha]
Subscribe for Podcast Updates Get updates on new episodes of Experiencing Data plus my occasional insights on design and UX for custom enterprise data products and apps. Email Address [text-blocks id="eu-consent-checkbox-textblock" plain="1"] .
In Episode #003, I talked to Mark Madsen of Teradata on the common interests of analytics software architecture and product design. Mark spent most of the past 25 years working in the analytics field, and he is currently the global head of architecture for Teradata Consulting. He is a true analytics pioneer and a regular international speaker who also chairs several conferences and is on the O’Reilly Strata, Accelerate, and TDWI conference committees. If I only looked at job titles, Mark would be an odd fit for Experiencing Data, but the reality is that Mark has many of the traits of a good design thinker including a good sense of empathy about what users need in the world of analytics and decision support software. It's a rare combination in my experience, so I hope you enjoy the interview. Besides, Mark is also highly entertaining ;-)
Topics we discussed include clay tablets and:
Why Mark doesn’t include his business title on his business card
How Mark's video game development background influenced his approach to analytics
Data-centered approaches (“what was the crop yield last year”?) vs. problem-centered approaches (“how much can I charge for the crops this year?”)
Why piling all the data together and then building a feature is exactly the wrong approach
Mark and Brian’s takes on designing a system that needs to scale as well as support specific tasks
Resources:
Mark on LinkedIn
Research Gate: Big Ball of Mud
Machine Learning The High-Interest Credit Card of Technical Debt
The Theory of Fun (book)
Mark on Twitter
Quotes from today’s episode: "The reason that these things fail is that people think they need to build an intergalactic data system to solve that problem." -- Mark Madsen "You don't build libraries by stacking books and hoping to find order in them. You figure out orders and then impose those orders in order to solve the problem." -- Mark Madsen "Open-ended problems and broad problems tend to not lend themselves to traditional engineering design solutions and that's where you really hit back again on UX as a starting point." -- Mark Madsen "The interesting thing to me is the knack for software developers and the educational program we have for software development is all based on function, “What it is you need to do?” -- Mark Madsen "We used to call it decision support. We didn't call it business intelligence or analytics or anything like that. I still like that old term." -- Mark Madsen
Episode Transcript Brian: Alright. Mark Madsen, are you there? Mark: I am here. Brian: Sweet. And where is here? Where are we talking to you from? Mark: You’re talking to me from Mount Tabor in Portland, Oregon. The only volcano inside the city limits of the city in US. Brian: Fun facts, alright. We’re already into fun facts. Mark: We are. Brian: Exactly. We have Mark Madsen who’s the—correct me if I’m wrong—you’re the chief architect of Teradata. Although, as I recall from when we met in London at the O’Reilly Strata Conference, your business card is null. There is no title. Can you tell us why there’s no title on it? Mark: I can tell you two things. One, a chief architect for Teradata services, not for all of Teradata, not the Teradata mixed products, so the services side. The cards are null because the conversation the chief architect has might be with very detailed developers, or they might be with IT people in management, or they might be with executives. You don’t want to set people’s expectation based on a title when you talk to them, so you leave it blank and then you just talk about what you do instead based on what they are interested in. Brian: That sounds like some of your consulting background at play. As I recall, you were a consultant, independent for quite some time in this whole BI space and analytics space for a long time. Can you tell us a little bit about your background? If I recall correctly, you started in game design for Apple computers like 8-bit and then you did some AI projects about 20, 30, 40 years too early. You did some mobile robotics work a little bit too early. At least too early in the sense of prior to when these technologies are more like daily topics and not academic topics. Can you give us some background right now and where you came from? Mark: Yeah, actually that’s a really good point what you just said. A lot of things went from academic projects when I was playing around with stuff to commercial now. The thing about academia is that they are not commercially viable much of the time. I was doing all these stuff when it wasn’t viable to do it I guess. But yeah, the AI work with expert systems was the final stage of the death of AI back in the late 80s. That was funded by me in my spare time, writing 8-bit video games that went out on diskettes for Apple II computers. That’s how I paid in part for college. Not a whole lot of relationship then although I was always interested in, “Well, if I have an expert system that understands this, could I apply it within the context of the game and make smarter opponent?” which has come full circle now because they’re not expert systems anymore. They kind of are, but that’s how those things came together, and that AI stuff led to the robotics stuff because if you’re trying to do autonomous robots there’s this intersection thing there. That was in the early 90s and that was a bit too early as well. That’s how I got that start but all the psychology, behavioral economics, and AI stuff led me commercially to data. When I left academia I was like, “Well, you can either make a fraction of a normal income or you can apply your skills to business.” Business is much easier in the intellectual rigor sense, but it’s much harder in the complexity sense. You trade one set of puzzles for another and it works out. Brian: That’s an interesting perspective on it. I want to get into your business insight on this whole world of data and analytics, and now we’re talking predictive analytics, machine learning. There’s all this stuff going on in this space, and of course, the theme of our discussion on this podcast is obviously the tie into user experience as helping both drive customer value and experience, but also ensuring tools and enterprise products that we build actually get used. Ultimately, they actually provide decision support and they actually either make money, or reduce cost, or whatever those end goals are which sometimes are not clearly defined. I’m curious about why do you care about user experience. The vibe I got when we met in London, I think it was at the Strata Conference, I kind of have this thing when I meet people that are not title designers, that there are natural design thinkers out there. They might have a technical title or something and usually my radar for that is strong. I’m like, “Okay, this guy has that bone in him,” because you’re talking to a room full of the tech people about experience. One of the quotes I saw in one of your slides was, “The right tool is the one people will use and not the one that you want them to use.” Tell me a little bit about your interest in that last mile. If the whole data pipe from data ingestion all the way through to some experience at the end, if that’s the marathon, you seem to be aware the value and the importance of the last mile. Where did that come from? Why do you care about experience and what has it done being aware of that in your career? Mark: That’s a good question. I mean this is at the crux of a lot of product design. I was just going through a product design problem today with a company that makes the service request system that we use to fulfill our IT service request, which has a problem on one of its user submission forms. Things that are so basic and are infuriatingly frustrating because they prevent you from doing what you need to do. I have a lot of empathy for people, but I think really, professionally for years, I was a programmer and you’re just sitting there writing the gut of systems. When I started doing data projects, it was early on because first we were applying behavioral economics and things to decision making. It was all decision theory stuff that I was working on and trying to incorporate context or AI assist. We used to call it decision support. We didn’t call it business intelligence or analytics or anything like that. I still like that old term. In order to do that, you were trying to put a computer with a person who is a not-technical person and make apparent the information that they need or focus them on the important information either by letting them find it or guiding them to it. Those things are very high touch. If you get them wrong, then these systems do not get used or the results are not what you hope for. That’s what’s led me down there and the other aspect is totally unrelated which was that if you write a video game, a video game has to draw you in. You have to be engaged in whatever that world is that the video game creates, and the aesthetic of that world, and the rules that one operates in that world. That requires that you approach the problem user-first or player-first. Brian: Do you think that the fact that now maybe we have an oversupply of data and then under supply of good decision support, has that created more of the aware? I think that the average executive or a VP or someone that I talked to as a client these days, they know what you access now. It’s no longer even explaining that and they understand. They don’t maybe understand it fully but they understand its value. Do you think that came out of the fact that there’s a glut of data now and now we have this problem of making it accessible? Did the glut create the awareness of, “Oh, this is a thing we actually do need to care about this.” Does that make sense? Mark: Yeah, it makes a lot of sense. I think there’s multiple things going on. I think that’s a big one, what you just said. I think there is a deeper reality to it. I think one of the things is that my career spanned a period when nobody had a computer on their desk at work to the period where everybody does. One of the things about getting it right early on was business managers whose sole experience was VisiCalc running on an IBM PC or an Apple II, the very first spreadsheet, and that broke the IT monopoly which was green phosphor screens. That was the state of the art and cryptic incantations and IT in control of everything and that put things into consumer hands. That created first an ability to do stuff like spreadsheets. But then when they started figuring out that if you took some of your business data and jammed it in the spreadsheets, we started to try to link the two things together—the mainframe and the spreadsheet—and that’s what got everything off and running was more things more useful. But along that path, there was a lot of design work that’s uncredited in that history that relates to the oversupply of data and the undersupply of usable systems for it, which is what you just said. I think that what you put your finger on is key. You go through various periods of history and data gets made available, but we don’t know how to make it usable or findable or whatever, and every system has a pivot point where at first there’s not enough or just enough stuff but eventually there’s too much stuff. You hit one of the keynote topics I did for a Strata Conference early on back in 2012 or so. Just on a history of information explosions and that history of data now is kind of the same. We’ve got lots of data and it’s distributed across silos and systems and repositories and website. You’re trying to find all these things that are applicable to your situation and use them. It used to be that most of that data got jammed into a warehouse. It got that because there were a bunch of different mainframe and minicomputer applications that had pieces of it. You put it into a warehouse to get a hold of stick view, so that you could find and dupe things. That whole design paradigm took a decade or more to develop and then 25 years to mature to today’s state which is not supporting today’s need. But this one all the way back to Clay Tablets. I had a research question once which was just how does one manage large collections of information when they exceed the capacity of the technology? Too many files begets databases, that sort of thing, and in Clay Tablet land the question I had was, “Gosh, if you’re recording taxes on Clay Tablets, how do you manage them? What does it look like to have your tax records on big hunks of unbaked clay?” That led to a lot of digging around and Mesopotamian architecture and the information architecture of libraries like Ashurbanipal which is being or has been, I think partly reconstructed, at the British museum now. The artifacts that they came up with to tag, essentially build metadata, to organize and structure for findability—because if you want to look at last year and compare it to this year to see whether the harvests are better or worse so you know what levy to place on the goods—all of that stuff requires information retrieval and it turns out that the techniques that were used 7,000 years ago and the techniques that we use today from an information theoretic perspective are exactly the same. But we keep forgetting that, and we build things, and then the technology becomes the view of the problem. And so instead of thinking from principles, you think from technology, and you end up where we are now. You have this oversupply of information but everybody’s viewed it through a technical lens. The BI stuff, for example, is crap tooling for today’s information landscape which is a glut, but it was perfect for yesterday’s landscape as the solution to the previous glut. Brian: One thing I want to stop you on that I really liked was when you talked about the tax levy, what was the crop yield the previous year. This is a great example of focusing on the end user problems and this is something I see with clients. If I’m talking to usually someone on the engineering side and they’re thinking implementation, they’re thinking how do we aggregate all of the previous crop data that we have? And the actual user question is, “How much tax can I charge this year?” Probably they want to charge as much as they can without going too high. That’s actually what the problem is but you might need previous crop history data to make that decision. If you don’t know that and you look at it as a, “We need to visualize the crop history data,” then your chances of striking out are higher. Do you agree that that’s a gap that we see a lot in this space is the people building the services don’t always know what those tasks are? Granted some things are exploratory, but I find that a lot of times there’s an 80/20 rule especially with tools that are designed for repetitive use. You need to support those repetitive tasks that people are going to do. If you know that the goal is to charge tax, “I need to do this every March I’m going to go and calculate the next year’s tax or whatever it is,” this system should be designed to do that. Mark: Yes. You’ve got a knack for finding some of the fundamental problems. I think you put your finger on a couple of them in that statement. I think I can only remember one now. But the point of focusing on that decision, that’s key because that’s what we build data or analysis or analytics systems for. Whether it’s base information retrieval, “What was the level of inventory in some warehouse?” answering the question, “Do I have enough or do I need more?” or something much more complex like levying taxes in a kingdom where you really need to know what’s enough. Is it to maintain the roads? Is it that you have to deal with the neighboring kingdom and so you’ve got to pay a bunch of soldiers to go invade them, in which case you need more money. There is always the extended context to those things. But that is the starting point. The interesting thing to me is the knack for software developers and the educational program we have for software development is all based on function. What it is you need to do? The problem is that building systems, building applications were typically building things that collect data that do stuff on forms, “Fill in this user registration form to download this white paper,” that sort of thing, or something complicated like an inventory management system. They’re very functional. You get functional specifications to do functional tasks, with very narrow task-based context, and that task is embedded in a larger process which is the end-to-end of say, inventory management. But inventory management in a business is one process that is part of a larger logistics problem. It’s also part of the, say, retail merchandising problem because that feeds into the stuff that’s on the shelves, which stock should be on the shelves, and which stuff shouldn’t we sell anymore. All of these things get entangled in this bigger enterprise organizational workflow and that is not a functional problem. That is a data- and decision-oriented problem. The decision making that goes with it is interesting. That means that your functional solution has to be focused on decision-making or aid in context. At a narrow level, there’s one set of things that are on a betting and that the wide ranging level, it’s completely different. Your approach to solving that is not what you learned, it’s not what you’re taught. All of the methodologies that tend to support this tend to be very different than the agile methods that everybody applies today. It’s a very interesting difficult problem to address. I think when you describe it the way that you did, it throws it in there because data problems came to be broader than a single system and open-ended. Open-ended problems and broad problems tend to not lend themselves to traditional engineering design solutions and that’s where you really hit back again on UX is a starting point. If you focus on the person and how and what they do in a much larger context and functional requirements, it drives you to think about the problem differently and more holistically. That open endedness is something that a lot of us as developers had to be trained out of in order to work on data systems. That was kind of a long-winded, wasn’t it? Brian: No, that’s okay. I think you hit on a lot of good points there and I agree. Some of this stuff is squishy when you get into the difference between getting a team aligned around a scenario versus the functional requirements. I see this happen in Agile, too, where sometimes when teams are doing Scrum, they’re really taking old-fashioned requirements and they’re just backing them into. As a whatever business analyst, I need to do X so that I can do Y. While I understand the spirit is there, they’re following the template of Scrum and writing stories, sometimes, they’ve never gone through the process of looking at the bigger like, “Where does this guy do his work? How often does he do this? What’s his life like and why does he hate doing this? Does he love doing this?” They don’t know what that experience is like. Either he wants to have or she wants or does not want to have, the task repetition that might be involved, so much of that context is lost. I think, again, that falling in love with the problem and getting your head really around the problem is critical. Otherwise, it’s just really falling into getting to big architectural decisions and all this stuff about how you’re going to suck all this data in and then spit it out the other end and it could be a total fail. Mark: If you look at the industry survey that a lot of the recent market attempt, has been a total fail. Gartner, Forrester, McKinsey, these analyst firms at various levels either in IT or business, are saying that in the analytics and sort of big data realm, the project success rate is somewhere in the order—depending on who you look at—of 10% to 20%. The standard, the baseline that has run through the software industry since the 70s is 50%. It’s about a 50% failure rate plus or minus five and has been since the first paper I read on the subject of giant project failures which was written in 1970. You touched on Agile and things like that. Agile is a great methodology when you already know your architecture, when you know your fundamental architecture. If your problem is web application, or let’s say you’re Etsy or somebody like that, there’s a pretty well-understood framework within which you operate and your Agile supports the exploratory work to build a feature. What I liked about it was that it got us away from a development model of know your requirements, builds to those things, heavy upfront engineering, because websites and mobile applications are high-touch and user-dependent. All this AB testing and things that’s supported by that very method, and along with that, of course, trying to take some of the operational components of, say, the DevOps world, melding all of that together, and that is great when you have that framework. The problem is when you have to deal with a deeper information systems and problems that people are trying to solve. Data problems are just viewed as, “We’ll pile all the data together and then I will build a feature for it.” That is exactly the wrong approach. You don’t build libraries by stacking books and hoping to find order in them. You figure out orders and then impose those orders in order to solve the problem because the problem is one of something like say, findability which requires certain things, but there’s a lot more than that obviously. Brian: You touched on the failure rates for these analytics and data projects. I actually wrote an article trying to gather up all of these surveys, as many as I could find. I think I only found about six. The sad part being, the November 2017 Gartner one was 85%. They actually put out a funny tweet like, “60% of all big data projects fail,” and then cross out, “oh, we meant 85%.” It was so funny. It’s been bad for a long time and something is wrong here with these big enterprise systems. This actually gets to my next question for you. You might be a really good person to answer this or at least have a perspective on it. It even touches on the whole Agile thing. A lot of times, when I’m working on a new product or a new application, if they want to do Agile, I don’t think Agile is always the right choice for what we would call a design sprint or sprint zero. I still feel like a more traditional design process needs to happen. You need to build a runway of some design work before the Agile is going to deliver the returns for the business that it is supposed to do. I don’t think necessarily you just start coding and building day one without any, especially for a data product, you need to have some idea of where you’re going. You want technical people involved in the design process with the product manager or whoever that’s playing that data product manager role and the designer. How do you think the right way to build if you’re building and a custom enterprise data product or application and you have a nice clean slate? There’s a ton of data out there, but you don’t want to just build another tool to go visualize all this data that’s in the warehouse or wherever it’s located from a technical standpoint. How do you build a small increment of value when it might require a tremendous amount of plumbing just to get to step one like, “Oh my gosh, we actually spit something out in a browser.” The amount of work it require just to start getting data on the screen was huge because I know that’s an engineering problem that happens on these large enterprise projects. It might take a while just before you get something on a screen, so how do we do a small amount of value without focusing, getting too lost in big architectural discussions? Do you have any suggestions for how to do that with design and business in mind? Mark: There’s a bunch of questions in there. I think you hit on some interesting things just in your choice of words. Like when you said, “The accessing a bunch of data and visualizing it,” it’s a presumption that all I need to do is see the data and then my problem is solved. When the data under glass is the departure point for the end-user to actually do something. The focus if you’re designing any kind of data system is, what is the action that is intended at the end of it? And that action could be, “I’m using Tableau and I’m trying to understand a problem so that I can figure out what to do.” There the action is, to inform or understand versus something a bit more dashboardy where you’re working out what do you need to know, to measure the health of this business process and its operational status, and what do you need to know to diagnose problems within that so that you can make decisions. Do this, do that. Change this, change that. Or data products in the sense of something I used to work on for a bit was recommendations. Recommendations are very different depending on the type of thing you’re doing in the context. So you can’t just say, “I’m going to apply the same system or techniques that I used for music recommendations as I did for retail recommendations.” And that goes to the context. The way that you approach that actually looks—this is sort of surprising—but it looks sort of waterfally. It doesn’t look very agile because of what you said. I don’t know which data I need. Your core root of your problem is, “What information do you need in what frame,” I use frame to mean sort of the mental frame or the frame of reference, “for what kind of problem?” What I see people doing repeatedly is actually succeeding first before failing. They do one siloed problem, and they build a thing that gets data from five different places. One of them was data we never used before. Fifteen years ago it might have been Clicks, now it’s something else. Blend that together to either produce a service or deliver information to somebody or to actually embed as an analytic bit that then feeds back into a system. That is successful because the bounding on it was narrow. The goal was fairly well-understood. The reason that these things fail is that people think they need to build an intergalactic data system to solve that problem. Step one, install a Hadoop cluster. Step two, feed massive amounts of data into it. Step three, build that data product or data pipeline or whatever it is, and people look at it like, “Wow, this works. This is fantastic.” Now, when you want to start another project, they say, “Well, we did this for department A. Let’s try this for this other problem over here.” You realize that you built a siloed, hyper-optimized, functionally-oriented system that solves exactly one problem. The problem in our market is that, handcrafting data pipelines to support individual things is exactly the pattern that we broke in the late 1980s with the data warehouse because every single process in a mainframe basically took files, built pipelines, and produced output files that were the information that was needed. It’s a human-driven, human engineering problem which builds no smarts into it, because you didn’t get enough context to solve more than one or two problems at a time. That leads you to, “Oh this is successful.” You do it a second time, you do it a third time, and then the fourth time, you start to look for these commonalities and you realized that, “No, 50% of the data is overlapping between these things, but the way we process them is different,” and you build tangles. You end up with the ball of mud architecture, to refer back to that famous paper. I think that, if I’m thinking academically, of course, if you’d want two great references, one is big ball of mud architecture, and the other is, I think it was AI or machine learning is the high-interest credit card of technical debt, the paper that was written. They outlined this in much more technical terms. You have to do that thing you don’t want to do, which is get a broad enough view to establish the level of infrastructure support that you need, essentially to define the architecture. There’s a part which is Agile, which is the upfront exploratory pieces and the contextual construction of application and data product, and there is a part which is foundational infrastructure, which is the data components that live underneath this. The fatal mistake that is made is thinking of it as a technology problem. “We can’t use databases because X. Their cost of storage is too expensive.” I hear that all the time and it’s the stupidest thing I’ve heard. Cost of storage doesn’t matter. Your cheapest cost of storage is /dev/null. Just, “Hey write once, forget about it.” If you want to retrieve it, that’s what really matters. The reason some classes of system content management repositories, data warehouses are so expensive is the labor that goes into making retrieval fast and efficient, and it comes at the expense of making new information available, slow and inefficient. This is the actual problem that the Dewey Decimal System solved for books 100 years ago. That is what we need now. If you don’t think about that problem, and the fact that you are not building a custom functional solution, you are making information available so that it can be remixed quickly to build the next one and the next one. You need Agile and [dween 00:30:30] and exploratory above the line, carefully curated, and fast enough to support the accretion of new information and cataloguing of it below the line. If you don’t divide the problem, you are screwed. That’s why I think there was five years of successes and excitement around a lot of analytics, followed by the last couple of years of, “Gosh, this is expensive and things aren’t working out the way we expected.” Brian: I’m going to sound like my engineering leads, my clients, and stuff, “We can’t afford to rebuild it again in the second iteration.” The general sense is that there is a tremendous amount of lift in the first version to get to anything, and then after that you can make it better. We have this kind of ping-pong back and forth about like, “Well yes, you can develop something, but if it doesn’t generate any usage, and the usage doesn’t generate any decisions support, and that doesn’t generate any value, then you just wrote code, and you built a software application that may not have a problem that it solves.” But I can see the alternate point which is, “Okay. We solved one or two problems here, maybe we did get an idea…” I’m trying to think of an example to put this in concrete, so let’s take a fictitious example. Let’s just presume in 10 years you walk outside, there’s hundreds of drones circling your house, delivering packages and doing all this stuff, and maybe you’re like a third-party drone service provider. We swap out propeller blades at the right time or something, I don’t know, and they want to develop a service. We all know there’s probably tons of IoT telemetry available about every working part on the drone, the towers, the communications, and all these kind of stuff. You could say, “Well, our first problem is we want to predict when the propellers need to be changed out, and I don’t know what it may be, there’s a couple of handful of tasks there.” The engineering person is going to argue that there’s a ton of data you’ll need to gather just to get to that point where we can start doing that one prediction. But the fear is going to be, “Well, this needs to turn into a product we can charge money for to these drone operators or whatever, so we need to have more than just that one thing, or we don’t have a commercially viable product, so that means we need to think bigger about the whole architecture at the beginning,” and the next thing you know is we’re spending all this time clumming for all the possible drone, and the tower data, the whole system before we’ve even solved that first problem, which is just propeller replacement, or whatever it may be. Are you seeing what I’m saying? You can make the argument about the need to go understand these scenarios and what the usage scenarios are, and the actual problems, and what the decision flow might look like for these users, to inform the initial engineering sprints, but there’s still that lift. Do you think it’s like, “Yes, it start with individual problems, solve those, and rework the architecture over time even if by the fourth strike it’s a big lift?” Is that the way to go? Mark: That is exactly the wrong way to go because if you try to do that, that’s basically the solve one problem at a time, focusing on the functionality of the problem rather than what is the aggregate set of things that you need to do in the bigger picture. This is complex system stuff. You need different sets of thinking tools around it. Just applying systems dynamics, systems modeling things to think about, that forces you into the broader context. You start and you’re like, “Okay, I need this information, I’ll slap it out here. And I need this information, I’ll slap it out here.” You don’t have a framework for the information architecture. You end up with a big pile of data, which is a big part of what happened to a lot of people. One of the big vendors in this space advocate building one system at a time but using these big clusters and just keep piling projects into it, and somehow magically, all of the information you piled into it is completely reusable. That’s a programmer-centric view of the world because as a programmer, “I see XYZ, and figure out what I need, I build my thing.” When you try to put that into the hands of a user, or you try to expand the scope of that across an organization, you end up with a giant collection of single-purpose things. It’s sort of like trying to build a 100-storey office building and refactoring every few floors. Eventually, the technical debt that accrues, unless you figure out what you’re needing to do, that will kill you. You have to understand, “Well, this is a 100-storey building, we’ve got to use steel girder construction and concrete, so we have to put that in place first.” I think, though, construction analogies are bad analogies. I think it’s better to think about infrastructure systems, municipal water, where does the water come from, where does it go to, how is it being used, because I look at data systems and I see the two parts, and I try to partition them. One part is the data collection and provisioning infrastructure, which is common to all at various levels of capacity. 100-storey building needs big pipes, single-family home needs little pipes. Then the second part of it is the application, and that application is where things like the Agile methods, the exploratory stuff to build data products comes from. The infrastructure piece is something you’ve got to get right. What happens instead is that people go out to the lake and build a pipe that runs all the way to their house or their building, as opposed to investing in the water system and then breaking apart the problem of water consumption. Changing BI tools is sort of like changing the faucets or the fixtures in your kitchen. It’s at the end of a very long chain of dependencies, and that is just like the data dependencies. If you’re problem is kitchen sink faucet is one thing, fire hydrant is a different thing, and bottled water is yet another thing. We tend to focus on that system like bottled water, and then work everything in the enterprise backwards to the data. That is what you don’t want to do, that’s where I said the waterfall piece kind of comes in, or we’ll call it Sprint-Zero. You have to survey the organization and look at both what you’ve got, what you don’t have, what you need, what the problems are around that. You have to focus on the business uses, the business cases, what’s feasible, what’s possible, and that gives you a pretty good grasp of the overall. Where I see a lot of data products stuff go wrong, whether it’s in the startup world or in the enterprise, it’s not doing that. That first discovery phase that leaves out that context and landscapes so that you see where you’re headed and what information you need, because there’s going to be 50% overlap on a core set of information, and then there’s going to be things that only one piece needs. Going in a warehouse, picking things, and putting them into boxes for order processing. Those pick events, they’re probably only useful to somebody who’s worried about efficiency of picking operations inside warehouses. That’s a not usable piece of data, but it’s tied in with all the products data, and the order data and the other things, and that information is probably common across three quarters of the organization. Understanding these aspects of information overlap and how one builds a framework around making it possible to supply both sets of needs simultaneously, that’s the kind of thinking where you have to sit back. It’s like that old Alan Kay quote, “You don’t Agile your way into a compiler. You have to know your methods and know when you need to gather requirements and when you can skip the requirements because you’re exploring.” Brian: There’s two things here. I guess I’d push back on one thing and I would totally grant the other. I think that discovery phase is so often lost. Some of the people that need to be involved with that, to develop empathy, to understand who’s going to be using this stuff like day in the life, what’s it like to be this person that ultimately is going to end up using that thing you’re going to work on, are not always present in that. They’re very decoupled, they don’t have that empathy, I love that. We would call that UX research typically, but it’s going to discover what the needs are before you’ve done anything, but if you can get, especially the engineering people or the data people, whoever those are, the SMEs about the data and the analytics, get them involved with some of that process so they can understand that world a little bit before any code has been written, I think that’s a good insurance for the project. You really have two choices. For my clients, we can design on assumption or we can design on fact. Now, you may not have all the facts, but it’s a choice. One is higher risk. Designing on assumption, or just using some designer’s opinion about what it should be based on them talking to you, you might get lucky. That’s probably better than just taking a wild-ass guess on your own. But it’s not as good as going out and spending some time, “Oh, we don’t have time to do research.” It’s like, “How can you afford not to do it?” You’re about to spend millions of bucks on this thing. So I totally agree with that. But one thing that doesn’t scare me, but that I get concerned about is that you do some of that stuff, you then get into the weeds, and the next thing you know, “Tell me how much tax to charge for the crops in the coming year,” that kind of got lost, and it’s still really hard to do that by the time the product comes out. There’s not a black-and-white answer to this, so it’s not like, “Mark Madsen, tell us the…” We’re having a discussion here, but I think that’s the fear, is that we can sometimes lose sight of what I would call the benchmark success criteria. Maybe you have these 8-10 problems, as you call them, like the pipelines of the older computer systems, which shoot out a file that had just what you need in it. I would say we need to shoot out experience in the tool that’s good for each one of those. It doesn’t mean there’s a wizard for every single one of these necessarily, but you do need to have some kind of criteria by which you are going to qualitatively measure the success the user experienced, the usability of the system, or else, again, you risk just writing code and having this big platform. But at the end of the day, at the last mile, it’s like the faucets don’t work well, or, “Well yes, water comes out, but it drips, and you need to fill your gallon container, and it takes an hour to make a pot of coffee,” or whatever. I don’t know, any comments on my rant, I guess that was rant. Mark: I like that we might get lucky. Designing in the absence of any requirement works if you can be the proxy for the person who is on the receiving end of it. But if you can’t, that’s where it gets totally random. That’s where all those discovery sessions and understanding the context. I love day-in-the-life kind of exercises, they’re my fave. I just did one yesterday. In our world, somebody who is going to be on the receiving end of, I don’t know, some data product or data system. Let’s take recommendation. What’s the context of which they’re doing things? There’s much more of a passive recipient side to that. But if you said I want to build an environment for a data scientist to do their work, and it’s a very complex environment, and there’s all these other people that are involved because the task crosses many domains, so day-in-the-life exercises I like because they show you, “I wanted to do this, and in order to do this, I had to do that, but in order to do that…” That was the old developer meme of a few years ago about yak shaving. You’re sitting there staring at the yak, wondering why you’re shaving the yak, and thinking back on the long chain of terrible consequences of things that had to be done in order to do the thing you really wanted to do. That’s actually where a lot of users are in organizations. I find that a lot of times, people like to blame the developers, but nobody ever educated the developer or the data guy in how to approach these kinds of problems. Universities have failed at this, everybody focuses on computer science stuff, or now with data science, they focus on math. Nobody focuses on the whole problem. When you go out and you do these things, if you do them appropriately, and that’s the trick, talking to somebody about the problem they’re solving, “Well, why are you worried about raising taxes this year? Why don’t you just do what you did last year?” “Well, because, we have a war coming” “Okay, well if you have a war coming, and you need to make spears, then you’ve got a bunch of things you need to think about.” So you’re asking these sort of what next, what before, what after, all of those things flesh out in understanding. I think that puts the understanding into the developer to make better design decisions. That’s why I feel that UX stuff and starting exercises from the complete end-user, and removing a lot of technical aspects out of the conversation helps so much. The last mile being, the key to the success or failure of a lot of information-driven systems. That right there, that’s what you started with, the right tool is the one that people wants to use, not the one you want them to use, which is how IT thinks of their role. You use what we built and bought for you. Eat your vegetable, they’re good for you. Brian: Totally. You’re spot-on. I’ll tell you, there’s nothing like a developer or a stakeholder who has seen the light, and they’d either watch someone suffer through the crappy thing that they made, or someone else’s crappy tool, or they’ve simply just spent some time and a light went on. I don’t know if maybe, did you have any particular thing that was illuminated from the session, the discovery you did yesterday? That’s one of the most exciting things for me, I think, about being a designer is when you find this nugget of stuff that no one has talked about, and you’re like, “Wow, I had no idea that you have a team of…” there’s like four people that you need to talk to to do this, and we think we’re building a self-service tool for this, and there’s an approval chain, and you have to send this data, this other thing and it comes back, and you got to share it with this other person. Wow. Just head exploded but in a good way, like, “Oh my gosh, and we can totally solve all of this but we never knew that this was even a problem.” Did you have any moment like that yesterday in your session? Mark: I think there were a couple, probably. It happens almost every time because there’s always some bit of context. Maybe one person on your team knows it, the other five don’t, and it’s just assumed. Everybody just sort of assumes, or they’re unaware, so fostering that, sometimes it calls into question assumptions. The discovery that, “Wait a minute, we’re building all of this stuff into our product to do X. But do we really, actually need to do X? Because most of the time in this context, it’s going to be done over here, not over there.” That completely changes what the product ought to do or what the data should be, or whatever it is you’re building. That’s the sort of thing that comes up. That changes your engineering efforts. Everybody talks about self-service data integration in order to do things like build data products. You have a data engineering team and they work on this stuff but you want self-service so that analyst types and data scientists could do a lot of that themselves. Then you build a system which is only amenable to developers rather than those guys, which happens all the time. There’s all these assumptions about resources, where you can do things, how you can do things, what skill levels people have, and where they view the value of their time. I think one of the big things for me years ago was I did an internal user survey. I was running two teams: a business intelligence team and an analytics team. The analytics guys were doing consumer research, so digging into people’s behavior and what they do, and the other was just the core business intelligence in the organization. I was really struggling with some of the contextual aspects of this. At context, I just, I talk to all of these people so I kind of knew, but I didn’t know how much time they spent every day in various tasks. We did some task studies, nothing really formal, in fact, driven by interns. Armed interns with a piece of paper and a pencil, or a spreadsheet, sent them out to talk to people, look at what they do, how often and how frequently, how long they do these things, and you find that the average business intelligence tool, or tableau dashboard, or whatever it is that you’re using, 15 minutes. Our organization, the bulk of it, the median was at 15 minutes. If you only use a tool for 15 minutes a day, that’s not enough time to really become proficient or learn how to do most things, and so you better design the experience around that much more tightly than the small number of people who spend a lot more time in it. But developers, and I include in this professional analysts, tend to presume that other people do that a lot more than they do. So, ”Oh it’s easy to use this tool. Just do X-Y-Z.” It’s the curse of knowledge right there. They’re so familiar and they spend four hours a day doing it. It should be obvious to someone who does it, you know, 10 minutes, every other day. That’s the sort of the last mile problem to me was figuring out what those environments needed to look like based on one single fact, which was, “How much time do you spend interacting with this system? If 10 minutes out of 8 hours in a day is all you do, do you view that as important?” That is one of the most interesting things that leads to the success and failure of just basic and give data to end-user systems. Unless they see the value of the information, the KPIs, whatever it is you’re delivering to them, they view it on the basis of, I need to get to this meeting, and I need to do this stuff–that is unimportant, that is probably one of the least important things, and what they want you to do is minimize that time from 15 minutes to five. Brian: Yeah, I’d say, broadly speaking about any time-based stuff in the UX, to take it with a grain of salt because sometimes more time spent can be good, more time spent can be bad, and less time spent could be good or bad as well. You need the qualitative side of it, you need to understand the context, “Are they in and out to solve a specific problem? What is the current value of x for this report? Okay, got it, it’s 92.6.” Or is it like, “I need to tell them which department we should spend more money on next year to get more whatever it may be,” that’s a different thing. It sounds like you guys did it, a diary study, so you had the, what we call a diary study in the UX world, but they’re self-recording their usage of the tool in this type of thing, is that what that was? Mark: We did two things. One, we instrumented the system so that we could see how long they spent doing which types of activities, like look at a dashboard, drilled down into some metrics, run a report, run a query, so you could see what they did, and what they did most frequently. The other was qualitative, it was sending people out saying, “Okay, you looked at this, what were you doing, why were you looking at that?” to get a more complete picture. This is a really good point about the time aspect of it, because sometimes the answer is, they need to spend more time but it’s too hard. It’s like playing that video game where you get stuck at the same point every single time. Years back, I worked on a video games system, where we were instrumenting the games to understand what was going on, which is common practice today in multi-user games. One of the things is find these places where people get frustrated and quit. They get stuck, it’s too hard, but in a multi-user game driven off of servers as opposed to one installed via CDs on a PC, you can change these things and you can make it slightly easier. In one case, we are looking at a racing game and there was a particular sequence of things at one point where people just really couldn’t get through. If you didn’t master that, you got frustrated, and then you could see it, because you take all the user activity and you map cohorts of people, and you look at the pattern of gameplay. In a lot of these systems you can of course create skinner boxes, but the idea is to try and maximize that play time and keep them playing and if it’s too frustrating, they quit. We did that right around that time that Raph Koster wrote one of my favorite books called The Theory of Fun. It’s really called The Theory of Fun for Game Design. But we actually used that as part of the design bible for how one does end-user data delivery. The idea of skill plateaus is not something that most application developers think of, because they think they’re building an application with a specific function. A lot of data systems are really tools for people to accomplish the goal, not the system that embodies the goal itself. That book has just tons of great design and experience advice and how to build systems that successively reveal complexity so that as you get better, the experience becomes richer, but you’re capable of working in that environment. That’s a very hard nonfunctional requirement for people to design towards. Brian: That’s a great example, about the game analytics there. This is again something that sometimes my clients have trouble with, or there’s pie in the sky, what I would call like very-few-word business goals like, “Drive revenue,” and they’re so non-design-actionable. This is a great example. One could be, we want to increase gameplay, or more specifically maybe it’s we want to increase gameplay by 5%-10% in terms of time. The scenario for the tool or the service that the internal tool you need to build may be, “We need to find points in the tool where gameplay is too difficult and people abandon.” That’s your service and then the tasks might be like, “You log in, and I want to see, has anything changed since last time? Are the sticky points still in the same place?” Theoretically, you might have made some changes so you probably want to understand, “Is it still hard or not? Do we need to go revisit more time or not? Where are those places?” Then the next question may be, “Why is it hard? What is going? What does the data say?” Ideally, the system could generate some conclusions for you and provide evidence as back up, but maybe your MVP is, you just provide some kind of evidence and they have to conclude the why part of it. But that’s how you take this pie in the sky goal of, like, “Increase gameplay,” down to something really specific there, and understand what’s going to go into that from an experience perspective. That’s really cool, I didn’t [...] they did that, but I don’t work on games, but that’s pretty cool. I didn’t know they’re all doing that now, that’s neat. Mark: They’re pretty much all doing that now. Any multi-user, or even mobile games, you’ll see it in mobile app design, too. I didn’t work on it, but I would presume that one of those popular games like Candy Crush or Angry Birds would dial in that stuff. I know the guys who’re building Angry Birds had a lot of telemetry on these things, but I don’t know what they did. It’s the collection of those things that you just described. Well, first of all, you described that the thread from the top to the bottom, which is how you figure out all the information for a particular use case. The other is that the collection of 20 or 30 of those across different points in the organization gives you the shape of the data space and the kinds of things in terms of capacity and capability that you need. One part of it feeds the application line or which is what are you trying to do and enable to support, and the other part of it the infrastructure component, and you now have enough information to guide the lower levels, the platforming space for the data work. Those two things go hand in hand. Brian: Wow, man, this has been fun, we could probably go on for more hours and stuff, but I don’t want to take up too much more of your time. But is there any concluding thoughts if someone was to walk away here? You have a lot of experience in this space. If I wanted to get better at designing good enterprise data products, is there any particular advice you might give to a data product manager or an analytics leader, a data science manager. Everyone’s intentions are good. They’re all trying to develop better services, but is there a core message you would give to the listeners? Mark: Really, over the years, we put different labels on things, but I think the key point is that starting point of goal. Looking back on the conversation about failure, there are companies that have spent hundreds of millions of dollars putting in big data and analytics infrastructure. They spent all that money without really knowing what the goal was. In data science, the problem is exactly the same. A bad data scientists doesn’t take what you described as a problem and turn into something actionable. Increase margins. Well, we could increase margins by decreasing costs or increasing sales, which way do you want to go, and you play that game of chasing it down, but the art of building an analytic model is very similar. You have to come up with an evaluation criteria for the model, and it has to be concrete and explicit. If you don’t know of the usage that you are heading towards, then it’s sort of like you don’t know where you’re going, any road will do. That’s the challenge. That’s why I thought it would be interesting talking with you, and just from sort of UX perspective, even though these days most of my time is based on building plumbing. The reason it’s all based on plumbing is because everybody wants fixtures, and then starts with a fixture, and runs it all the way back. It’s like rewiring your house every time you buy a new toaster. Brian: I’m still picturing a four-foot sewer pipe running from the pond down to my house. I hope that doesn’t break or get clogged. Mark: In IT we’ve got all of that. That, plus the rewiring of the house, plus we rebuild the house every year because we’re adding another floor. Brian: This is super fun. Mark Madsen, where can people find you online? Are you on any of the Twitters and the social medias and the interwebs? Where are you out there? Mark: These days I’m not out there that much because I’m not really doing much that’s public anymore but Twitter is one place, which is mainly just random things I find interesting. Conferences, there’s always the O’Reilly and [...] conferences, because I like doing conferences. Brian: Your Twitter is @MarkMadsen? Mark: @MarkMadsen, yeah, and the other thing is whenever I do something that’s speak-worthy, I post it to Slideshare. Brian: Well, thanks so much for coming on here, it’s been great. I’ll put some of our links, maybe we’ll have a Clay Tablet link and a big ball of mud, there’s a lot of mud and dirt themes going on in this episode! I’ll try to put links to those, the credit card AIs, the high-interest credit card, there’s some good stuff here. I got to check out that design bible to The Theory of Fun for Game Design, that sounds cool, so thanks for those recommendations. Thanks for coming on. Mark: You bet! Thank you for having me. Brian: Yeah, alright. See you later. .
Julie Yoo is the co-founder of Kyruus, a medical technology company that is the developer of ProviderMatch. One of the most frustrating things about the healthcare system is the tendency for patients to be sent to the wrong type of doctor for their health issue. The industry term for this problem is patient access paradox.
ProviderMatch is software that directs patients to the proper medical specialist for their specific needs.
During today’s episode, Julie and I discuss the components that make ProviderMatch an effective tool. Some of the topics we touch on are:
How ProviderMatch has changed the customer service side of healthcare.
How ProviderMatch helps combat physician burnout.
The 3 major user bases served by the application.
The 3 types of tests Kyruus uses to test new and upgraded product features.
The 3 levels of analytics that Kyruus uses to measure RIO and value.
Resources and Links:
Kyruus
Kyruus on Facebook
Kyruus on LinkedIn
Kyruus on Twitter
Julie Yoo on Twitter
Julie Yoo on LinkedIn
Thank you for joining us for today’s episode of Experiencing Data. Keep coming back for more episodes with great conversations about the world through the lens of analytics and design. .
Kathy Koontz is the Executive Director of the Analytics Leadership Consortium at the International Institute for Analytics and my guest for today’s episode. The International Institute of Analytics is a research and advisory firm that discusses the latest trends and the best practices within the analytics field. We touch on how these strategies are used to build accurate and useful custom data products for businesses. Kathy breaks down the steps of making analytics more accessible, especially since data products and analytics applications are more frequently being utilized by front-line workers and not PhDs and analytics experts. She uses her experience with a large property and casualty insurance company to illustrate her point about shifting your company’s approach to analytics to make it more accessible. Small adjustments to a data application make the process effective and comprehensible. Kathy brings some great insights to today’s show about incorporating analytic techniques and user feedback to get the most value from your analytics and the data products you build for the information.
Conversation highlights:
What is The International Institution of Analytics?
What is the analytics leadership consortium?
The “squishy” parts of analytics and how to compensate for them.
The real value of analytics and how to use it on all levels of a company.
How beta testers give perspective on data.
The 3 steps to finding the ideal beta tester.
Learning from the feedback and implementing it.
How to keep ROI in mind during your project.
Kathy’s parting advice for the audience.
Resources and Links: The International Institute of Analytics Kathy Koontz on LinkedIn Surf Camps Thank you for joining us for today’s episode of Experiencing Data. Keep coming back for more episodes with great conversations about the world of analytics and data.
Quotes from today’s episode: "Oftentimes data scientists see the world through data and algorithms and predictions and they get enamored with the complexity of the model and the strength of its predictions and not so much with how easy it is for somebody to use it." -- Kathy Koontz "You are not fully deployed until after you have received this [user] feedback and implemented the needed changes in the application." -- Kathy Koontz "Analytics especially being deployed pervasively is maybe not a project but more of a transformation program." -- Kathy Koontz “Go out and watch your user group that you want to utilize this data or this analytics to improve the performance.” -- Kathy Koontz “Obviously, it’s always cheaper to adjust things in pixels, and pencils than it is to adjust it in working code.” -- Kathy Koo
Transcript Brian: I’m excited to have Kathy Koontz of the line. She is the Executive Director of the Analytics Leadership Consortium at the International Institute for Analytics. That’s quite a mouthful. Did I get that totally right? Kathy: Yeah. You did. I tell people, “Yes, it is a real job.” Yeah, you got it right, you nailed it. Brian: Tell our listeners, what does that mean? What do you do? Kathy: The International Institute for Analytics is an organization that was founded by Tom Davenport, one of the early readers and using analytics for business performance. We are research and advisory firm that helps companies realize value from analytics. Brian: How do you work with them in your leadership capacity there? What’s your specific role there? Kathy: I lead the product line that’s called the Analytics Leadership Consortium. What that is, is a group of analytics executives from different companies who are in non-competing industries that meet of a regular basis to better understand trends and best practices in analytics to vet ideas with one another and ensure that they’re doing the best that they can to deliver analytics value for their organization. A really great opportunity for leaders in this field of analytics that’s changing a lot and has a lot of emerging practices, have regular time to get together in a confidential setting to understand what’s working and what’s not, and how they can improve analytic values for their company. Brian: You mentioned trends, I obviously try to stay on top of what’s going on in that industry and that’s actually how I came across you, originally, I think was you guys had put on a webinar on five trends and analytics going on right now. At the end of that, you had mentioned that one of the things that's starting to change now is the importance of design and user experience as we move beyond designing reports which is one of the most difficult deliverables, so to speak, of analytics—as we move into the user experience it's becoming more important. That's what I was like, "Oh, this could be really interesting to hear what you have to say about what’s changing. Why is UX now relevant? How is capital D design relevant to the world of analytics?" that's what I was curious to learn about. Can you talk about that a little bit? Kathy: I think as analytics mature in organizations, the need for design is what’s going to drive adoption and utilization of those analytics. In companies that are just starting their analytics journey, it’s a little bit easier to realize analytic value by doing a couple of really big data science projects that don’t really require a lot of design thinking. It’s a powerpoint to an executive group that has some higher level of organizational level thinking that the executive lean in, understand the analytics, understand the value of making their decisions in line with the analytics and then move on and do their regular roles. But as the organizations try to make analytics more pervasive, particularly into front-line associates or individual contributors who are really using analytics to make a lot of small decisions within the execution of their work or as they try to integrate analytics into processes that are monitored, that design thinking taking an approach as user experience, can really break down a lot of barriers that organizations encounter when trying to have people who are not necessarily used to using analytics in their decision-making process, use them so that they can make a better decision for the organization. Brian: Got it. Is the trend that people are recognizing the problem because they went through some pain of maybe they delivered some big multimillion dollar platform, and like you said the front line associates didn’t use it. As I always use the example, it's like, “They’re out driving a truck. They don’t have a laptop. They’re not going to go download a report and change the columns or whatever.” It’s like the wrong mechanism, you didn’t fit it into their job; you try to get them to change their job to accommodate your tool. Is the awareness because they went through a failure or is it like, "We already know this isn’t going to work if we don’t get the UX right," just because companies are a little bit more design aware these days? What drove that to change? Why is it now? Kathy: I think there’s two things. First of all, they’ve failed, they invested millions of dollars in some sort of decision-support sort of application that may have millions of dollars and data integration work that we’ve done to build it and purchased a lot of really advanced software that can help users slice and dice the data and dive in and understand it. But they really took more of a data-centric approach rather than user-centric approach. They really didn’t take that extra step to say, "How does this person go through their normal day in their decision making. Where do they need this information in that decision making, and how should it be best presented so that they don’t have to do any cognitive task switching, that it just fits into how they go about making the decisions." I think those failures are one big driver. But then I also think as organizations move up the analytics maturity curve and move from BI reports to predictive analytics to prescriptive analytics. Those prescriptive analytics are going to be much more pervasive across the organization. I think it’s this analytics maturity that’s also driving this need to put more design thinking into this creation of analytic products. Brian: Do you have any specific examples of a company that may have like, "Version one was X, and we didn’t get what we wanted. Version Y or X.2—we then went back and tried to fit this in better with that employee, and we saw some kind of change of result." Can you cite any examples of how design allowed the data to actually be insightful and create a meaningful change? Kathy: One was a project that I was involved in at a large property and casualty insurance company. We were trying to alert our franchise insurance agents if they had somebody in their book of business who was likely to not renew. It was based on really good data science and model scores with different, clear [...] that showed significant likelihood to a trait between them. It was originally deployed in a separate application with some general groupings of how likely they were to leave. The utilization of it was more of a curiosity at that point. There were some adopters, but not as broad as it would be hoped to really be able to substantially move the needle on this. As the design was redone and integrated into their overall CRM system—was legit on their CRM system that came up when they logged in in the morning, they didn’t have to go look to see, it was there for them. As we moved from groupings and scores, into one star to five stars, five star is they're likely to go–you really need to work those. So just very simple little changes that drove a significant change in the utilization of that capability. Brian: Do you know what the blockers were such that it took a redesign? If someone was listening to this and they’re like, “I’m that person right now. We don’t want to go through the learning experience that your company went through.” What do you do to prevent that? Like, "Here are some roadblocks to watch out for." Kathy: The thing is, these products are usually developed by folks within a data science group. We have the data, and we know this is a business problem, and we've done the analytics. That is expected but not sufficient. Then the next step is to really use just basic research techniques from consumer companies. Go out and watch your user group that you want to utilize this data or this analytics to improve the performance. See how they do their job, what tools do they use, where would this information be most relevant, how can it be presented in the context of their general activities, where it’s not a separate thing, where it’s integrated into the stuff that shows up of their performance evaluation. That’s the way to really avoid some of the deployments that may have really great data and science behind them, but don’t get user adoption that’s needed. Just take that extra step and really understanding the user and where this information is going to be most relevant for them. Brian: Do you think the appetite from executives and people that are at the top of the reporting chain for these things support the time and the effort to go out and do that type of research or to try to fit it in and not so much focus on "When are we releasing code? Show me some progress." They want to see a glee. "Just show me some proof that these millions of dollars is doing something," versus this kind of squishy. It can be squishy—at least in my experience with certain executives, especially qualitative stuff. “How many people did you talk to?” It’s like, “Oh, we talked to eight so far.” It's like, "We have 10,000 employees." It’s hard for certain ones to understand the value of qualitative research in these things. Do you have any experience or thoughts about that? Kathy: It can be squishy, but I think really analytics especially being deployed pervasively is maybe not a project but more of a transformation program and you have to take the transformation program perspective to it which includes such squishy things that change management and business process redesigned. Somebody using those analytics within their decision-making process is really where the organization gets value from. That code that gets released and deployed, it is an interim step. But it is not the final step to that organization getting value, maybe you’re just a data scientist of the company, trying to deploy a great app that could help a group of marketing folks better invest marketing, or supply chain folks better manage costs to suppliers. It doesn’t have to be this big concerted professional effort. One way to do that in a very agile, low-cost way is to find a couple of folks that seem to really jazzed about getting this and use them as some beta testers. Maybe you start out with just the information on a spreadsheet and say, “Hey, does this make sense just from this information that you would use?" And then have them walk you through, "Where would you use this information? How would it be best presented to you?" And then think about working within the constraints that you have of maybe you can’t change the screen for this digital marketing [...] that somebody uses, but how can you make that information as accessible as possible, where it really is more of a push of information at the right time, as opposed to pull of information that a human being has to remember to go get when they’re in the process of executing their “day job”. Brian: You hit on some great stuff there. I talk about this to my list frequently which is understanding tasks and workflows and people’s day-to-day jobs that the goal is to fit your solution into their existing behavior as much as possible. It’s really hard to change behavior. “Oh, I got to go remember not to load this other screen and pull out those number of page two and paste this into the other screen and then hit enter and then it does some analytics.” These are the kind of stuff why people don’t bother to do it. It’s like, "Well, my guess is good enough. I’ve been doing this for 20 years. I'm this good. Whatever. No one’s going to know if I did that or not. They don’t know where that number came from.” They don’t want to do that. First of all, it's to understand their pain, and then another good technique there is if you’re going to improve a process and we want people to use this tool in order to realize this new value, first of all, understand that benchmark of what their existing workflow, their playbook is and then ask them to do the same playbook using the new platform. This is a great way to uncover the things that you don’t know to ask about necessarily because if you first get that model of how they do it now, it’ll help elucidate those gaps that we can’t see. I just read an article, there is an example like, “Let’s imagine designing a new hotel room.” and you might take for granted that that hotel room has a bathroom. [...] is talking about this. You’d be really surprised if you walked in and there was no bathroom because no one stated that was a requirement, and because everyone just took it for granted. It’s those kinds of things that you won’t see as a data scientist or product manager, perhaps because you don’t know what these front-line workers are necessarily doing all day. But I love that you talked about getting into the minds of what people are doing and fitting your solution into them. One problem I see is that in some places, it’s really hard to get access to customers. I think sometimes, this is more on enterprise companies that are selling a product that has a data-specific part of it, it can be hard to access them. You would think companies doing internal analytics that this would be easier. Do you think that comes from leadership? Do think that comes from the ground up? Any comments on how do you get access to the right people? What if they’re not being like, “Hey, I’m a call center. I get paid hourly by the number of calls I do, and you want me to go work on your new software?” How do you incentivize that so that they become a design partner which ultimately, is really where you want to go is to have a team of partners from subject matter experts, and product managers, the data science people, whoever it is that is going to affect the solution that comes out. What needs to change in order to incentivize that participation? It takes time to get these things right. Kathy: Yeah, for sure. We see that a lot just because as you noted, the breath of the organization or the different incentives and priorities that other groups might have. I hear companies all the time, senior leaders, we are going to be an analytics competitor, we’re going to do analytics. They think if they hire 200 Ph.D., data scientists that they are now an analytics company. I do think you’ve got to have somebody at a leadership role at least over the data science group say, “Look, this company gets value from this when we are making better decisions because of this analytics.” And those better decisions are going to come from people accessing the analytics. When they’re in their decision-making process, and really working more leader-to-leader. But that can’t be where it stops. You've got to create a peer relationship that all levels, number one. And then number two, if you get the leadership of the other group engaged early on—as this is a problem that they feel they have ownership in and that they are a co-creator with you, and starts at the leader level, and then works its way down—then, I think you’re going to have greater access and more ability to do that. Finally, in the absence of that, if you don’t get it right, then I would at least, as you're deploying the new app, say, "Our next step is to watch the staff members use it and identify how to better integrate it." So that at least you’re showing some leadership in your thinking that, "I know this is going to be a problem because I didn’t have access to them. I’m already teeing it up." We’re going to have to come back and see how people are using it and how we can utilize some user experience and design thinking in the subsequent phase after [...] is reached. Brian: Sure, sure. I think some of my clients in the past, they have to go through a failure first in order to decide that they don’t want to do it. We don’t want to do the build-first, design-second, process on the next one. There’s a certain level of convincing that can happen, and then at some point, you got to move forward because typically, IT or the business, they have been tasked with, "Deploy this new model and this software into the survey," and they’re going to do that no matter what. They’re not going to stop and wait for design if that organization doesn’t have a matured design practice. So you might have to go through that. I would say, obviously, it’s always cheaper to adjust things in pixels, and pencils than it is to adjust it in working code. The more you can get in front of these decisions, and then form the engineering and the data science prior to deploying a large application, it’s a lot cheaper, and it’s a lot easier, and you don’t have all the change costs associated with that, both time, money, labor. No one likes to do it twice, most people want to work on new problems, they don’t want to redo the same ones. Engineers usually don’t like doing this, so getting in front of it is good. Kathy: I think what you’re touching on there is, maybe if you haven’t done your user design work, then you think about this release as a beta release. You plan this release, and you plan this app, and you manage your code knowing it’s going to change at some point, and knowing it’s going to change at some point, and knowing it’s going to change soon. The second thing is, key point is to decouple the analytics insight from the application. There is an analytic insight, whether it’s a number, or a recommendation, or some score, or something, is it is just loosely coupled to the delivery mechanism than it is easier from a code engineering perspective to have to it delivered in a different way. Brian: Right. No, that’s great. A part of that is that the attitude and culture of change and not falling in love with our first versions of stuff. To me, that has to be ingrained in both the engineering–all of the teams that are touching the product that the service that’s being put out there. A lot of companies say they’re doing agile. A lot of them are skipping one of the most important parts which is getting some customer representation involved, and actually, iterating and not just doing incremental design where you keep adding more, "Add another feature, add another data point," that’s incremental design, that’s not iterating and changing as you get feedback. That’s something to watch out for as well. At some point, you need to get some stuff out there, and it costs to have come down obviously, to deploy software. It’s overall cheaper to get something out there quickly and to start getting feedback of it, but it’s very easy to no get the feedback and to let the working code feel like success. Until someone at the top is like, "Where is all the bang we’re supposed to get for this?" Kathy: I think the approach there is, get it deployed, so that’s great; it’s out there, and it’s working. But again, build into your deployment plan, use your feedback, and change from there. If you are not fully deployed, until after you have received this feedback and implemented the needed changes within the application, then I think that’s another way to prevent that falling in love with your first design. You know that this is just deployed so that people can use it, and that getting user feedback and making those needed changes is an expected part of the deployment process and is not a failure that the application was not sufficient. Brian: I think that should be part of any software development process to have a loop of test design, refine, deploy in the circle. For the most part, for larger applications, you’re never really done with it. It’s a process of getting better and figuring out the ROI like, "Have we hit the market? Is it worth spending more time and money on this?" but yeah, those are great insights. I’m curious in terms of the roles, I feel like someone that’s at the top of the responsibility chain here needs to have a healthy dose of skepticism about their own stuff, especially when it comes to prescriptive and predictive analytics or any service where it’s custom software they’re deploying into the organization have a healthy dose of skepticism about how great it really is. Maybe you deployed on time, bug-free, you can see the stats and all of these. But is that responsibility primarily falling into the data science realm because companies are investing in that area right now and so they become the de facto, what we would call a product manager more in the SaaS world? Did they become that and is that the right place for that responsibility to be? Kathy: I see that happen really, not necessarily through design and intention. But just because if a data scientist wants to take this great science that he’s discovered and make it accessible, in a lot of organization they have to do it, and so it’s put upon them when they’re probably not well prepared for that. I do think that that’s the problem. And then as a lot of different tool and capabilities make it easier for data scientist to deploy an application, if there's not a really good user design construct within whatever application they’re deploying, in their data science application, then they need to be the one to take the extra step. As I said earlier, often times, the data scientist sees the world through data and algorithms and prediction. They get enamored with the complexity of the model and the strength of its prediction and not so much with how easy it is for somebody to use it. I think that’s probably a reorientation with training that should happen within data science groups around design, best practices, and user experience design. They’re not going to evolve into great user experience designers, but they are going to be able to recognize something that’s really bad and perhaps ask for help and guide to make it better. Brian: Do you think that’ll stick then that that role will continue to live there? This person that needs to understand the business value that needs to be obtained, the user experience side of it and the technology side is trifecta there. Do you think that’ll stay in that data science world? Kathy: I don’t think so. I do think there will be a separation between the data application design and the creation of the data science that informs that [...] in that application. I think as we talked about as some of these companies get more companies get more mature and need to have more pervasive deployment of data science insights, they are going to realize that they need to take a different approach. I often said that between data and analytics, I see the industry mature along the lines of software development. That design focus was not a big part of software development early on. I think it’s just going to have to happen within data science. Brian: I look at it this way, product owners can take many different titles. I have had all kinds of different clients, but ultimately, the box stops with everybody that’s working on it. But it helps to have someone that’s at that intersection of, "What do we need to do? What’s the overall picture of this?" and they understand the tech, the business, and the user experience side. Whatever the title of that person is, that role to me is really critical so that technology doesn’t run with everything. You have to have all three of those ingredients to deploy successfully, at least in my experience. It seems like that’s pretty critical to have that. Kathy: Yeah, for sure. A number of organizations that ask us about how do we show the value of our data science investments, how do we demonstrate our data science ROI. I think if more data science groups really looked at how their data science products were being consumed and started quantifying those, and then using some of the business metrics that are involved with that business process, whether it’s optimizing a spend or reducing average handle time at the call center, that would give them a great task to be able to validate their ROI to the organization, but I don’t see a lot of data science groups doing that. Brian: I’m curious, who ask that question? Is that the data scientist themselves or is it the business stakeholder who’s hiring the data scientist? What role asks that? Kathy: Usually, it’s a data science leader who has a large organization in large enterprise organizations. If you have a large organization with a lot of expensive resources, there is that continuing need to show the value that this organization brings to the overall company, especially when the output of that organization is not well understood or has not then a traditional part of that company. A lot of senior executives, C-Suite executives at large organizations, didn’t have data science when they were coming up through the ranks. This is something new that they don’t really understand, they know how much money they spend on it. A lot of data science organizations will need to demonstrate, “Here’s the hardline benefit that this company has gotten for investing in data science capabilities.” Brian: This blends nicely into my next question which is about obviously, AI, and machine learning are hot topics right now in technology. I hear this from people I talk to frequently which is, “Oh, the board knows we’re supposed to be doing some AI." They’re asking me, "How many sensors do we have installed? Do we have digital transformation?” they ask these really high-end questions, and they want to go spend some money on it because they’re so afraid they’re going to miss the boat on that. From a design standpoint, we would say, "That smell, it reeks of possibly putting a cart before the horse." The tool comes out, "We got to go buy this hammer because everyone else is buying this hammer. We have no idea what you hit with it, but we got to have it. We got to go spend some money on it." How do you ensure that you don’t waste money? You want to invest in this. You don’t want to miss the boat. Maybe there’s potential for a project to deploy machine learning. It's just pretty much what a lot of these companies are doing in terms of AI right now. How do you make sure that the desired investment from the business is actually going to have some ROI? They have heard this tool is hot. Kathy: It’s what I call the hype cycle mandate. Whatever is the new thing, it was a big data, it was AI and machine learning, with data sciences right, we have to have them, we have to tell the board we’re doing this. I think that is where executives earn their money, is being able to manage the message to the senior leaders who may not understand what’s needed and how you use it so that they can say, “Yes, we’re using it.” But the leadership than the executive leadership or the ones who has to go out and figure out where are the business problems and what is actually needed within our operating environment and our company to really deliver value from this capability. I will say, oftentimes, I see executives making the right call in that way. I have seen cases where folks have gone out and bought a lot of a software, and hardware and stuff that they have no idea of how they’re going to use, and that’s the shame. Brian: Do you think the right step there is to take on a small project, find a small win, show a small value and you can at least satisfy the, "Are we doing something?" "Yes, we’re doing some machine learning or whatever." Do you think that’s the way it starts? Kathy: I am big just in any data [...] investment, a big fan of used case-based development. Come up with a used case that requires this type of capability to execute it at the scale and precision that’s needed and then do some pilots to prove out the value, show the value, and then that then builds up a business case for the larger deployment that way. Yeah, I totally agree with that, Brian. Start with a used case, start small, understand, have an eye towards scaling as you start small, but get a small win, show that you’ve done it, the CEO can, in all honesty, say, “Yeah, we’re doing machine learning, and we’re going about it in a way that is physically sound but will also put us in a position to be able to compete with this capability in a quick amount of time." Brian: That’s great. I think that I don't know if I've said it, but this concept of falling in love with the problem and if you and your team, the people I work for you, or whoever it may be can fall in love with the problem and then weaponize your machine learning against that. That’s always a great thing, it’s to get everyone jazzed about the problem so that you know, especially if you can line it up with that technology not that the goal is to do that, but that’s when big wins happen to me, at least in my experience. This have been awesome. Do you have any advice, overall what will we talk about in terms of data product managers, data science managers, data science leaders, analytics leaders in terms of design, experience, what they should be looking for, going forward and just, in general, bringing more value to the customers. Is there a theme or something on your mind right now that needs to get mitigated to them? Kathy: My theme is a data science is about big complex data and a lot of technology and really advanced math, but they’re still human beings who have to use it. Don’t forget the humans. Focus in of how great this data set is that you created or how advanced this analytics technique that you’re using, but remember, the humans are your last line to realizing value from all of the stuff you’ve done before. Make sure you keep them in mind as you go through all of those other stuff as well. Brian: I think that’s great advice. And it sucks, those pesky humans. Kathy: I know, I know! They just get in the way of all those greatness! Brian: This is awesome. I have one last question for you and that’s have you ever surfed on a river? Kathy: I have not. Brian: I’m just curious. That came up in the webinar. I just saw this article on The Times about river surfing and I’m like, “I got to ask Kathy about this.” Kathy: Yeah. I've seen a lot of videos around that. I think it’s like this bore tide where the tidal action creates like this perfect wave that you can surf like forever. There’s some really good videos of folks doing that down in Brazil with some of the rivers that are draining out into the Amazon and other rivers that are draining out. It looks really awesome, but there’s this surf ranch in California that has this man-made wave. Kelly Slater, I think worked to create the engineering for this technology. That’s exactly what it looks like for those river-bore waves. But I’ve seen actually somebody in Munich surfing one of those [...]. It’s on my bucket list, Brian. Brian: Sweet. I’m going to put a link to the surf camp, but where can we put some links to you? Where can people find you on the interwebs? Kathy: You can find me on LinkedIn, I am, Kathy Koontz. You can also find me at the International Institute for Analytics, it’s iianalytics.com, that’s our website. My LinkedIn name is customerjourneykoontz because that’s always been my passion is using data and analytics for customer journey. That’s where you can find me on LinkedIn. Brian: Cool. I will put those links in the show notes. This was super awesome. Thanks for coming and talking to me today. I’m sure people are going to enjoy listening to this. Thank you. Kathy: Thanks for having me, Brian. Have a good day. Brian: Alright, see you. .
Hey, everyone. I’m Brian O’Neill and I’m excited to share my new podcast with you called Experiencing Data. I’m a consultant specializing in design and user experience for custom enterprise data products and apps. I’m also the founder and principal of Designing for Analytics. My goal with this podcast is to expose you to you or rather to other professionals like you. Who is you? Like any good designer, I had a persona in mind when I started designing this podcast. This persona is basically, modeled on my past clients, conversations at data and analytics conferences that I’ve spoken at, and email exchanges with subscribers on my mailing list. My guest and I assume my listeners are usually going to be data product managers, engineering and analytics leaders, data scientists, and executives. Regardless of the title though, Experiencing Data is really a podcast for business leaders responsible for turning data into useful, usable, and valuable decision support via custom software applications. Maybe you’re wondering why I’m doing this and I am too a little bit. But here is why, I believe the success of analytics software and data products intended for people, since some of them obviously, don’t have interfaces as many of you probably know is, products that are intended for people are only as good as the experiences that they afford, sometimes I refer to that as kind of the last mile of this large technology projects and products that we put out. Because not all companies have trained designers and UX professionals on staff, I was curious to learn how my guests consider user experience as they design these enterprise data products and software tools. On this podcast, we’re not going to go deep on design implementation topics such as data viz and user interface design, some of these things are inherently visual, and I think reading about them and seeing examples is more relevant. But more importantly, I want to look more broadly at what I sometimes call Capital D Design. Capital D Design looks more at defining business objectives, user needs, the problem spaces especially, and the success criteria for new products and services. We’re also going to stay clear off heavy technology discussions since there’s already plenty of that kind of stuff out there and that’s not my area of expertise. Also, on occasion, I may record some solo episodes and share some of my insights on designs that you can put them into play in your daily work. If you’re looking for this kind of insight on a regular basis, you can head over to my Insights mailing list which is at designingforanalytics.com. I write pretty regularly to my list. Feel free to subscribe there if you’re interested in learning more about designing UX. I’m also a professional percussionist. I’m a professional musician and performing artist. In addition to my design consulting work that I do, I wanted to find a way to bring my two worlds together. I’m going to have occasional episodes with music technologies when it’s relevant to Experiencing Data. To kick that off, we’re going to have an upcoming episode featuring a guest who’s a product manager, and his name is Julien Benatar, he’s over at Pandora which I’m sure many of you know. He’s going to come in and talk about how Pandora has gone about designing their services analytics platform which is called Next Big Sound, so looking forward to that one. I hope you will be too. One of the things about podcasting, in general, is ironically how few analytics, we, the publishers and the producers and hosts, receive about our listeners. As those of you on my mailing list already know, I routinely going out and interviewing your customers on a one-on-one fashion; customers, users, whether they’re paying for your software or using an internal tool, I really advocate going out to uncover latent problems they’re having and latent needs that may not be necessarily expressed. But since the podcast environment though doesn’t let me eat my own dog food and do this type of research since we’re kind of in a one-way broadcast modality, with me speaking and you listening, I hope you’ll leave me feedback, either in iTunes or via email. You can reach me at brian@designingforanalytics.com. This show is my MVP, and I’m sure this show may change over time. If you don’t know what an MVP is, well, stay tuned because we will probably cover that as well. If this show sounds interesting to you, please head over to iTunes or your favorite podcast app, and click the subscribe button, and then you can join my mailing list at designingforanalytics.com/podcast. That page will be the homepage for this show. Thanks again. I’m Brian O’Neill and welcome to Experiencing Data. .