Unlocking Product Growth with Big Data, Data Science and Machine Learning

23 Apr17:30 – 18:00 UTCStage: Main StageTalk

Checking session availability…

Hang tight while we load the latest updates.

In today's competitive and fast-changing market, product managers need to leverage the power of data to drive product growth and innovation. Big data, data science and machine learning are key tools that can help product managers understand customer needs, identify market opportunities, optimize product features, and measure product performance.

Unlocking Product Growth with Big Data, Data Science and Machine Learning

Poornima Muthukumar at UXDX Community: Product Growth: Data, CX, and Research Strategies. Video: https://youtu.be/yaVGgHVpu3c

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

How companies use data to drive growth

[00:00:00] Poornima: Hi everyone, thank you so much for joining today's presentation. I'm super excited to speak to all of you today on how to unlock product growth with big data, data science and machine learning. I want to add that I'm not speaking on behalf of Microsoft, but rather sharing the knowledge and experiences that I've gained in my journey along the way, and the themes and messages that I bring forward are my own. So, without further ado, let's get started.

[00:00:25] Here I have five companies. I just want to drive home how these companies are using data to drive their product growth. First, we have Netflix. What you see is all driven by data. They also use data to decide what kind of movies they want to make, how they want to invest their money, and what kind of content resonates with users. They also use data to decide which movie to store in which CDN location that is closer to the user, so that the movie is streamed efficiently for users and they have a better movie-watching experience.

[00:01:03] Next up we have Tesla. Tesla has all of these sensors and cameras that are constantly sending data back to Tesla to drive their autonomous driving system, and they also use data in general for improving their autonomous driving system and their self-driving cars. Next we have Amazon [?]... growth. It uses that data for when you search for something. They use that data for their pricing strategy; they use that data for warehouse optimization and for their inventory management.

[00:01:39] [?]... similar users. It is using your data, your interaction data, your engagement data, all of it, to decide what kind of content to show you, in order to keep you really engaged on the platform, in order to drive their product growth, and in order to make customers engage with the advertisements that they show, to bring revenue. That is their business model.

[00:02:18] [?]... that is data to build the [?] experience, to optimize your productivity suite experience. Now we also have Copilot that is integrated into Microsoft's [?]... again, driven by data.

Organizing data to extract insight

[00:02:47] What I want to drive home here is that all of these companies have a huge customer base that generates a huge amount of data, and today storage and compute and processing have become so cheap that you can store all of this data, extract insight from it and then in turn use that to drive product growth and business growth. Let's say you join as a product manager at any of these companies. As a product manager owning the [?] cycle, you're getting data from disparate sources, from various channels: the engagement channel, revenue, the sales channel, the feedback channel, the usage channel, telemetry. All these different data are coming in.

[00:03:28] What you want to do as a product manager, in order to drive product growth efficiently, is organize it in a very clever way, such that you can extract insights from it and use that in turn to drive growth. This is where your data science algorithms and data science techniques come into play: how you organize data in a clever way, such that you're able to extract insights that in turn give you information to improve your product. Which is why I say that today, the melding of data science and product management is the future of technical product management. In this talk, what you'll see is how to build data-driven products backed by insightful analysis, and how to utilize big data, data science and machine learning to inform complex decisions.

[00:04:19] Here you have the product life cycle. I've broken down the different kinds of data that I use on a day-to-day basis to drive product growth. Some of them are leading indicators, in the sense that you get this data very early on in the product life cycle and can use it to drive your growth. For others, you need your product to have been deployed, you need customers to engage, you need customers using your product for a certain duration of time, and then you can run analysis on which you can extract insights. So you have funnel analysis, segmentation analysis and engagement analysis. These are some leading indicators. And then you have experimentation, machine learning, feedback analysis and retention analysis, which are lagging indicators, for which you need at least a month's time to get these signals back in order to efficiently make a conclusion. Since we have only 30 minutes today, I might not have time to go into all of these different analysis methods, but we will see a few of them in detail.

Funnel analysis

[00:05:21] First is funnel analysis. Funnel analysis is nothing but a method used to analyze the sequence of events leading up to a point of conversion. Let's say I'm a product manager for [?]... a certain percentage of users will enter the payment details and complete the whole purchase flow. If you see here, ideally you want all customers to take all of these actions, but that will not happen. Customers keep dropping off at different stages of the funnel, which is why the funnel keeps getting shorter and shorter. The ideal journey is the whole thing, but customers keep dropping off. So you see 100% of it; 60 went to the product details page, 30 added to cart, 10 went to checkout, and 3 went on.

[00:06:27] As a product manager, let's say you own an email signup flow. If you had this data in the form of a funnel, you're able to visualize where exactly the majority of your customers are dropping off in the funnel. Once you have this data, you can hypothesize what exactly is going wrong, and then you can run different experiments, validate your hypothesis and, in turn, improve your product. Let's say in this case I see that the majority of customers are dropping off at the home screen itself. Then maybe I can hypothesize that the page is too slow; customers are losing interest, and that's why they're dropping off.

[00:07:05] If you see that the majority of customers are dropping off at the checkout flow and at the payment stage, then you can hypothesize that maybe the price is too expensive, or maybe my competitor is offering a better price. Once you have this data and these different hypotheses, you can run experiments, and I'll talk in the latter part of the presentation about how you run these experiments and how you can use that data to actually validate your hypothesis and, in turn, improve the overall product experience.

Segmentation analysis

[00:07:34] Next is segmentation analysis. Segmentation analysis is nothing but how you slice and dice your customers into different segments. Here I have a 2D scatter plot, and I am organizing my users on two different dimensions: one is income and one is age. You see here I've marked three circles, which you can characterize as outliers. But then you see four distinct clusters of users when you segment your users based on these two dimensions. One could be students, because they are younger in age and have lower income. Then you have techies here, middle age, middle income. Then you have executives, and then finally you have retired people.

[00:08:19] What I'm trying to drive home is that if you are able to segment your users based on different metrics, which could be income, gender, age, different things, it will give you an insight into the landscape of your product, which in turn will help you make different data strategies to target your product differently for different users. So segmentation analysis is important; it helps you identify [?]... based on different segments.

[00:08:50] If you don't have a lot of different dimensions that you can cluster your data on, you can always start with customer data, and you can always start with demographic data, and then use other kinds of data to further cluster your users. Customer data could be something like products, purchase amount, time spent, purchase frequency. Demographic data could be something like age, income, gender, marital status, education, geographical data and things like that. So that's what segmentation analysis is.

Engagement analysis

[00:09:27] Next, we have engagement analysis. Engagement analysis is dependent upon your product; the engagement metric that you track could differ from product to product. Let's look at some different engagement metrics. Let's say you're a product manager for a social media company like Instagram. If you're looking at [?]... number of posts, number of comments, time spent, these are some of the metrics that will denote how well customers are engaging with your product, how deeply they are engaging with your product, and how much they really like your product.

[00:10:10] Let's say you're a product manager for a search company like Bing, and [?]... which part of your product is lacking engagement. Let's say Bing has now introduced Copilot, we've introduced chat, GPT, all of that. If you're a media company like [?]... Netflix, or a company like Amazon, cart, products, purchases.

[00:11:21] What I'm trying to drive home here is that if you want to really understand the engagement of your product, you want to have a dashboard that tracks these metrics daily, weekly or monthly. That will really give you insight into the engagement of your product. These are some different engagement metrics based on the product, but you can also have some very generic metrics, like daily active users and monthly active users. User engagement indicates how often customers use your product and how passionate they feel about your product.

[00:12:00] Again, here I'm just talking about some generic metrics that you can use, like daily active users and monthly active users. These are some metrics that I track on a daily or weekly basis for a product, because I want to really understand, as we launch new features, as we bring new features to customers, how these metrics are changing, and how I can use that data to understand what is resonating with users versus what is not resonating with users.

A/B experimentation

[00:12:23] Next is A/B experimentation. Let's say I'm a product manager and I'm sending a Christmas greeting card to my users. I have one variation on the right and another one on the left. Maybe with the Christmas greeting on the left, customers are more likely to click on it, open it and see it, versus the one on the right. Here it's a trivial example; in this case it's a Christmas greeting, and it's not really impacting business metrics and business revenue on a large scale, which is fine. But say you have an open house website. You want users to click on this open house button, register for the open house and come see the house. If one variation of the button is driving more engagement with users, then I want to know in a very data-driven way what is working well with this customer and what is resulting in a better conversion rate.

[00:13:22] Here I have [?]... result on the left, whereas there is a different search algorithm powering the search results on the right. Maybe the algorithm on the left results in more customers clicking on the shoes, eventually adding the product to the cart, eventually purchasing, and eventually generating higher revenue. So A/B experimentation is not limited to visual elements on the page. You can also use it to test and validate different back-end algorithms, different search results, different machine learning models, different APIs and different web services, to understand how that in turn will drive your product metrics and product growth.

[00:14:11] What exactly is A/B testing? It's also called split testing, bucket testing, or a randomized controlled experiment. It's typically used to compare different variations of a web page, and you can test anything from the color of a button, to the back-end algorithm powering search results, to the layout of a page. You have one group, which is called the control group, and the other one, which is called the test group, and you want to keep all the elements constant except for the one thing that you really want to test. Then you measure the outcome. You have a metric that you test for at the end of the experiment, and at the end of it you see which one results in a higher conversion rate.

[00:14:51] A/B experimentation is the best scientific way to establish results with very high probability. What I mean is that as a product manager, you are not making decisions with your gut, you are not making decisions based on your instinct, saying, "I think changing the color of this button to green will result in higher revenue." No: you are using experiments and you are using data. You run the experiment for two months, and at the end of the experiment you see which variation results in better revenue, and then you launch the change to all users. That way you are able to take decisions scientifically, as opposed to relying on your gut and relying on instinct.

Problem statement and hypothesis

[00:15:29] What are the different stages of A/B experimentation? You always start off with a problem, and then you define the hypothesis, you design the experiment, you run the experiment, then you interpret the results, and then you launch the change to all users. The problem statement depends upon your product, and your problem statement will vary. Let's say you join as a technical product manager at a booking company or a travel company like Expedia or Booking. The kind of experiments you run will be very much driven by the company's goals, by the business objectives, what the company is aiming for.

[00:16:06] Let's say Expedia wants to increase the number of bookings, they want to increase their loyalty program participation, they want to increase the number of searches. Your A/B experimentation will be very much driven by the problem statement that your organization and your leadership team is dealing with, because you want to make sure that the results you're driving ladder up to the goals of the company and your leadership team. So you want to make sure that you are very crisp and clear on the problem statement before you even design and run experiments.

[00:16:44] Next is defining the hypothesis. The hypothesis in A/B experimentation is a testable statement that predicts how changing something will affect a certain metric and user behavior. This is the three-statement framework that I use for defining the hypothesis, even before I design my experiment. You start off with a problem, based on some evidence. Next, you believe doing something will impact a certain outcome and will improve a certain problem. And you know that you have achieved the outcome when you see some metric change. You want to have these three statements defined for your experiment even before you start.

[00:17:21] Let's take one example of how you would define this. Let's say you are a product manager for a retail [?] website. You are seeing fewer units sold, through sales data. So you have a problem, that fewer units are sold, and you have evidence based on sales data. What is the next step? You believe as a product manager that incorporating social proof, as in showing "X number of people purchased this in the last 24 hours," is what influences visitors to make the purchase. It's like a psychological trigger, and customers feel that they're going to miss out on the product. So your hypothesis is that if you apply it to all users, this would increase conversion rates, and you'll know based on some metric change. In this case, the metrics you'll track are revenue and units sold.

Designing and running the experiment

[00:18:15] Okay, the next step is designing the experiment. When you design the experiment, you want to be really clear on what you're going to track for the course of the experiment. You'll always have one primary metric. In this case, let's say I'm going to track revenue per user per month. You could also have a secondary metric, and other metrics that you track, but you always want to start with one key metric that will denote the success or failure of the experiment.

[00:18:44] Next, you want to determine the population that you're going to run the experiment for. Are you going to validate the experiment with all users, or certain sections of the users, say in the USA and Europe? You want to be very clear on your target population and scope your experiment to them before you launch it to all users. Next, you want to determine the sample size. Here's a formula that you can use to determine the sample size. The industry standard is to use an alpha value of 0.05 and a power of 0.8. That will determine the sample size that you need in order to have statistics on the results before you draw conclusions on whether the data is significant or not.

[00:19:25] Similarly, how long do you run the experiment? The general industry standard is that you should run the experiment for at least two weeks, but you also want to factor in seasonality, days of the week, holidays. You obviously don't want to run an experiment like an email campaign during the holiday season, when customers aren't even opening their emails or not checking them so much, because then the data would be biased and wouldn't be indicative of the regular business rhythm and business flow.

[00:19:55] Next is running the experiment. You want to randomly assign users to both the test group and the control group. You want to randomly split them in order to make sure that the data isn't biased; the randomization will help. You want to make sure that the data is something that you can use to make conclusions and drive results. Finally, you want to work with the dev team to instrument logging. If you're tracking revenue per user per month, you want to make sure that you have a dashboard collecting and surfacing this data that you can see at the end of the experiment in order to draw conclusions, and avoid peeking at the results and assuming that you have enough data to drive results.

Interpreting the results

[00:20:41] Finally, when interpreting the results, you want to make sure that the data is reliable. In some cases when running A/B experimentation, in my case in the past, it has happened that the data got corrupted, in which case you want to discard the results, rerun the experiment and do sanity checks. You also want to decide the tradeoff between different metrics. Let's say, as a product manager, you're tracking revenue and engagement. You introduced a new change, and you saw that with this new change, engagement is going up. Customers are engaging more with the platform, but revenue not so much.

[00:21:17] So you want to decide the tradeoff between the different metrics, and how the different metrics are doing. In most cases, when you introduce some change, one metric could go up and another metric could go down, and maybe both metrics are doing really great. So you really want to track two different, conflicting metrics, to make sure that your change is effectively launching something that's going to resonate well with customers. And then you launch the change when the results are statistically significant [?].

When to use machine learning

[00:21:44] Finally, machine learning techniques. Machine learning is not a magic wand, but it's an application of AI that gives systems the ability to learn and improve from experience without being explicitly programmed. In some cases, classic programming can solve the case, and maybe you don't need machine learning. You just write a bunch of if conditions, you write your classic code, and that can solve the case. But when do you use machine learning? Machine learning is a great thing to use when you have lots of data, when you have lots of good, ethical, unbiased data. Then machine learning fits the bill.

[00:22:25] Next is when you have really complex logic, like I said with the search query. That is not something that can be solved with classic programming, because you can't anticipate what customers will search, and you can't code up all these different cases: if the customer is searching this, show them this; if searching this, show that. This is where you need machine learning, because you want the system to learn with time, based on lots of data and based on customer patterns and behaviors.

[00:22:51] Next is when you want some sort of personalization. Let's say, for Spotify, [?]... you want to personalize the experience for users based on their characteristics, like [?]... YouTube's trending on [?]... Twitter's trending. What's trending on Spotify today might not be trending tomorrow, because as new songs come, as new artists come, the system keeps evolving based on user patterns, based on all of this data.

[00:23:35] Machine learning is typically split into three different types. One is supervised machine learning, where the machine learns from training data that is labeled. Then you have unsupervised machine learning, where the system learns from unlabeled training data. And finally you have reinforcement learning, where the machine learns on its own.

Ranking, recommendation, classification, regression and clustering

[00:23:54] Here I have a few basic examples of machine learning that can help you think about how these different algorithms and techniques can be used. First is ranking. We have seen a great example where machine learning can be used to power search results. You have ranking algorithms that are driven by machine learning models, based on what users are searching for. Next you have recommendation models that use collaborative filtering. All of these things show you recommendations based on your search queries, your patterns, and also based on what other users, similar users like you, are doing on the platform. The great thing about recommendations is that they don't have to be perfect; as long as it's close to what customers like, it's great. You can give something, and the system can learn over time.

[00:24:38] Next is classification. [?]... machine learning works great because you have all of this data. It can learn, it can predict in real time, and it can automatically assign tags to products based on what the machine learning model is already trained on. Next you have regression. Let's say you have a customer support situation, and you want to predict something. You can use machine learning, a linear regression model, to predict how much support volume you will get today, or how many customer calls you will get tomorrow, based on historical patterns and other things that you have seen in the past. You can use regression to predict what your support volume will look like. This is a great example where machine learning can solve the case for you.

[00:25:34] Similarly, clustering. We already talked about how Spotify uses [?]... to tag fraudulent transactions, and automatically tag them, saying that this transaction looks very different from usual patterns. This is where machine learning can help, because you can't have a human manually looking at every transaction and doing so. So these are some examples of where you can use machine learning to really identify [?]...

Q&A

[00:26:28] [?]... to take any questions.

[00:27:04] Poornima: [?]... and if you understand, then the point you're making is, let's say, in the case of Amazon, I was part of a project that involved analyzing petabytes of data. We did some analysis, and at the end of it we saw some data and made some judgment calls based on that. Then along the way we realized it looked like we hadn't taken some variables into our calculations. At that point, we really had to go back and redo our analysis to take two new data sets, two new variables, into play, because we wanted to really make sure that the results we were using to drive decisions were real, in fact authentic, because these are impacting millions of users. It's impacting so many users on such a large scale.

[00:28:17] In that case the leadership is already aligned, because you want the analysis to be correct. But in some cases, when you run an experiment and you know that the data is corrupted, it's really unfortunate. You want to be true and authentic to what you're driving for your customers, so you go back and convince them: this is something that has happened, and we really need to redo it. But now that we already have all the systems in place, the pipeline and all of this, it shouldn't be as time intensive to do it again. But we have to do it.

[00:29:28] Host: Now, this is being recorded, so yes, we will be posting each of the talks separately on our channel, so do follow us on LinkedIn and on [?]...

Speaker

Poornima Muthukumar

Poornima Muthukumar

Senior Technical Product Manager

Microsoft