How to Build Data Products in a Fast-Paced Environment

10 Oct9:00 am – 9:35 amTalk

Checking session availability…

Hang tight while we load the latest updates.

How do you meet expectations from different stakeholders while being truly data-driven? From changes in requirements and ad-hoc requests to international macroeconomic demands and constraints, it has become even harder to build scalable products in a fast-paced environment.
Monique shares learnings for the past year on how Trustpilot built data products in a cross-functional team to help their global customers overcome these challenges.

How to Build Data Products in a Fast-Paced Environment

Monique Marins at UXDX EMEA. Video: https://youtu.be/QMipGgiaw9w

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Who I am and what Trustpilot does

[00:00:00] Hi everyone, my name is Monique. I hope you're enjoying the conference. I'm happy to be here and talk about how to build data products in a fast-paced environment. Today I'm going to share my experience working in cross-functional teams, and some of the lessons and learnings that I had building data products while needing to quickly adapt to changes in the requirements and the demands.

[00:00:32] Before we jump into the presentation, I would like to share some information about me. My background is actually in finance, and I've been working my way through to data science, mostly because of my passion about how to use data in different ways. I've been working with marketing consultancies and on projects with blue chip companies on how to apply AI techniques to solve real problems. More recently, in the last years, I have also experienced shifts in how we work, and I would like to share that with you today. I'm working as a senior data scientist at Trustpilot, building and designing data products and improving customer experience.

[00:01:38] More about Trustpilot. Trustpilot was founded in 2007 with a vision to create an independent currency of trust. It's a digital platform that connects consumers and businesses, so we help people shop with more confidence, and at the same time we provide actionable insights to businesses. We are in 10 different locations, we have more than 167 million reviews, more than 700 websites reviewed on the platform, and more than 50 nationalities working at Trustpilot.

[00:02:26] To give more context about my role and the team: I work in the data science team, across all the products, serving both the business and the consumer side. We work in cross-functional teams to improve and to design those features and data products. My experience and some of the learnings that I'm going to share come from the business side, and can actually be easily applied to any scenario and any team, so I hope you enjoy it.

Data products and why most projects fail

[00:03:10] To start off, let's talk about data products and some of the elements that we need to consider. Let's start with the customer pain, so we think about the customer's problem, what we wanted to solve. We also have our data characteristics, so the data that we produce and the raw data. And then we also have the value that we can add to this data, so basically how we can enrich this data to make it actionable.

[00:03:46] Before we jump into more details about the workflow and how the development process actually takes place, I would like to share some statistics about those projects.

[00:04:02] According to Gartner, only 20 percent of data and analytics insights will deliver business outcomes, which is quite alarming. And it's actually a little bit worse, as 85 percent of all big data projects fail. You are probably wondering why. There's so much investment in time and resources, so what are the reasons why most of these big data projects fail? I highlighted the ones that are, for me, the most interesting reasons for that.

[00:04:44] In most cases we're trying to solve, or we do not identify, the right problem. We are not actually adding value. Another interesting reason: we think about deployment as the last step, in the last phase of the process. Or simply, it's a problem of the process itself. There is no process, or we are just doing a wrong process, and that's quite interesting.

[00:05:17] Today I would like to focus on giving more insight on the process.

The development workflow and what slows it down

[00:05:24] If we look at the data product development workflow in a general way, first we have the first steps of talking to customers, understanding the problem, identifying our audience, our target audience. And then we quickly try to work on a solution, create a prototype, get feedback from the team and from stakeholders, and implement the solution. And then we iterate a little bit more until we are happy with that, and then finally we launch the feature, the product, and then we start over.

[00:06:13] What can possibly go wrong there? I would like to highlight a couple of things. In the first phase, talking to customers, it is quite interesting that we sometimes assume that we know what the problem is and we just jump into the next phase, building the solution, without really understanding what the problem is. Or we have some assumptions about what the customers want and we just jump into the development phase, the prototype phase, too soon.

[00:06:58] There are other reasons in the prototype and development phase, and I would like to give an example. Also as a result of not really understanding what the problem is, there is a lack of requirements for the prototype phase, or very custom requirements. As well as that, when implementing the solution, we end up with technical problems, integration issues, or we figure out that it is actually more complex than we estimated in the beginning. Common issues are also related to dependencies on other teams. And all these problems cause the project to actually take longer than estimated.

[00:07:56] If we look into the goals and the people involved in this process, usually what we see is that customers are represented by project managers who talk to customers, UX designers, the marketing team, sales, customer success teams. They are heavily involved in these phases. And then in the prototype phase you see more development teams, so engineering, data science, data analysts, machine learning engineers as well. And then also in the implementation phase, the development team. And then finally, when we launch the product, then we have the marketing, all the communications, the release process.

A new way of working: everyone involved in every phase

[00:08:49] But if we think about the challenges that we had, and how the roles are traditionally involved in this process, can we make something out of it? Maybe can we change the way that we work so we solve some of those challenges? So I proposed a new way of working.

[00:09:12] If we think about all the people involved initially, we have customers and identifying the problem. It is not really a big shift, but in this slide what I'm proposing is that actually we are all involved in the process, but we contribute in different ways. In some of the steps we contribute more, and in others we contribute in different ways, as for example giving input.

[00:09:52] Now you might be wondering, okay, so in the prototype phase, the development phase, how can the other roles actually contribute? I'd like to give some examples, because it's actually a very difficult question. For example, in this phase, when data science is working on the prototype, it is quite helpful to have some feedback, even in terms of whether we are choosing the right metrics to measure success, or to keep track of any integration issues, so we can make sure we reach our timeline and launch the product, launch the new feature.

[00:10:39] And in the same way for engineering and the data teams: how can they contribute to the other phases that usually they are not so active in? In the first phase, for example, data science and data analysis can help to explore data, validating for example personas, or identifying market trends, and providing more information to stakeholders so we can actually understand the business and the customer problem.

[00:11:24] As well as the last phase, data science and data analysts, and also the team itself, can be involved in listening and helping with some questions, or understanding why customers are having questions and what exactly they have questions about. So this is actually a way to be involved through the whole process.

How data science contributes to understanding the problem

[00:12:00] As I mentioned before, data science can actually be quite active across the whole process, across the whole development process. But I would like to go a little bit deeper into the details of how this works in practice. I think it is also a good opportunity for professionals outside the development, technical area to understand how they can collaborate with data science, and how data science is used in the different steps.

[00:12:46] I would like to start with the problem statement. If we think more about Trustpilot, businesses receive tons of reviews on a daily basis, so how can we translate all those reviews, all that data, into actionable insights? We have a given problem to solve.

[00:13:08] The initial phase, as I mentioned, is understanding the problem, understanding who this customer is. The traditional methodology to understand is that we interview customers, we create personas to understand their journeys, which are the reasons why the customers have that problem, and we do that through customer surveys and through customer interviews. And finally we try to come up with a solution for how to solve that problem. What I would like to share is how data science can contribute to that.

[00:13:57] From the data side perspective, we can actually help by looking at the data, exploring the data, understanding if there are any patterns in the data, if there's any relationship with location, languages, customer base segmentation. Establishing these correlations and patterns, and helping UX researchers, and the team, to understand the problem.

[00:14:38] Also in the ideation session, usually with the team, trying to come up with solutions, data science can be quite helpful in providing stakeholders with early assessments and insights on the data that we actually have available, and in identifying limitations, or red flags, for some solutions at the very beginning.

[00:15:05] As a side note, especially for technical members of the team: we usually jump into, and I can include myself in that, we usually jump into solution mode, thinking about how to solve the problem, and sometimes going too much into the solution and forgetting about the problem. It is quite important that we focus on the problem instead of just jumping to the solution, to how to actually build a solution.

The prototype phase and four elements to consider

[00:15:53] Going to the next step, that is the prototype phase. A prototype is an early sample, model or release of a product that is built to test a concept that you are trying to build, and it is typically used to validate the product design or functionality and gather end user feedback before deploying it into production. That is usually the common definition of a prototype in software development, and I would like to challenge that today by how we actually see it in data science, building data products.

[00:16:35] When designing the solution and exploring approaches, I would like to highlight four elements. First, it is scalability. During the prototype phase we usually use a sample of the data, or of our customers, so how does it actually work if we wanted to scale that to all our customers? When thinking about the solution, the model, the approach, it is quite important to think about scalability.

[00:17:22] Also explainability, and I really like this one. It is basically the bridge between the technical part and the non-technical audience. How can you explain your model, how can you explain the outputs of your model to, for example, a non-technical audience?

[00:17:42] Another one, the third one, is implementation. Here you should consider the challenges of the approach you have chosen when implementing it into the product. You can think about the technology stack, how difficult, how complex this approach is to implement into the current architecture.

[00:18:08] Flexibility. Flexibility is about changes. Changes will come, we don't know when, but they will come, and then you need to have that in mind when choosing your approach. I will go through some examples in the next slide for all these four elements.

[00:18:30] But before that I would like to highlight, also in this phase of creating a prototype, that it is quite important that we have in mind how to validate our approach. I like to think about that in two different ways, two different categories of metrics. We have the technical metrics, that can be related to the machine learning metrics that you use, like precision and recall, more technically. And there are also business metrics: how your model, how the output of your model or approach, will impact the customers. So we can talk about coverage, how much of the customer base your feature impacts, and the changes that it will promote when deployed to production.

The four elements in practice

[00:19:32] Some examples of the four elements when choosing your approach in data science. When creating a new prototype, I would like to go through some examples of how this works in practice. When it comes to scalability, as I mentioned before: is your approach, is your model, is your algorithm able to scale when deployed in production? Is it able to scale to all the customer base?

[00:20:02] When it comes to explainability, the chosen approach should be able to explain the model outputs to non-technical colleagues. Remember that we are building products for different audiences and we need to understand your audience, so your model should be able to speak to them. And also from the data science perspective, the model output should be easy to understand, so data scientists have a good intuition on how to interpret them, how to improve them, how to correct them when something is wrong.

[00:20:47] On the implementation side, we need to consider the current architecture, as I mentioned before, when choosing the machine learning approach: how the model outputs will be consumed, for example the data structure, how it relates to the current technology, and latency. Those are a couple of examples that you should consider.

[00:21:14] Flexibility: consider a machine learning approach that allows you as a data scientist to iterate quite easily over time. Part of that is the approach and the model, and part of that is also how mature your development process and your current tech stack are. For example, your machine learning engineers are helping you to create a pipeline that can automate the process so you can deploy experiments faster, more quickly. All that goes into flexibility. So flexibility on how you build, and flexibility also on the algorithm behind it: how easy is it to change something?

Customer feedback and how to act on it

[00:22:18] Let's say that we are finishing our prototype, we implemented the model into the product stack, we will launch a new release. That is great, so am I finished? Well, no. Usually we then have a phase for customer feedback.

[00:22:46] We expect to have feedback from different customers, so different opinions about the same feature. Some customers can be quite positive and actually validate what you built. Other customers can have different reasons why they don't like the feature, that it's not useful, that it's not exactly what they want, or they have additional requests. From the feedback, and a couple of examples that I put in, it's quite easy to see that it is not only one reason. We can see that customer feedback goes from user interface issues to communications, through the data science model, to engineering problems.

[00:23:48] So how to proceed? It is important that we understand the feedback, so we ask why. Why, for example, is a customer asking for this new feature, why is it important, why is it urgent? Trying to understand the reason behind the feedback, and then working as a team, prototyping and improving the product.

[00:24:23] Now you may think, yes, but how can data science and the engineering team work together in the same sprints? We work quite differently. Retraining a model is not the same as fixing a bug in the UI, so there are different requirements for that. There is no right answer here, but one of the learnings that we had is that for the short term, for what we consider fixes from a data science perspective, two of the elements that I mentioned, flexibility and how we implemented it, actually allow us to quickly make some changes, so we can keep up with the demands and the critical things for the short term.

[00:25:32] In the medium term, for example when we think about the next release, it is important to have in mind as a data scientist that we can identify areas of improvement from the data, so from your metrics, but it is also important to listen to stakeholders. They have the domain knowledge, and it's important to consider that. But it's also important to validate that through data. So we receive feedback, it is important to the customer, and we can also find it in the data, we can also validate it through data, and then we can implement it.

[00:26:18] Also, as I mentioned before on scalability, it's also important to have pipelines in place. Being able, for scalability and also for flexibility, to have a retraining process and pipeline in place is quite important in a production setting.

[00:26:43] And it is important to say no sometimes to very custom requests. This is important also for engineering teams, that we cannot just implement every single feature that a customer requests. But also for data science I would say that's quite important, so that we don't prioritize or favor some segments or businesses by implementing custom requests and custom features.

Takeaways

[00:27:29] I would like to recap with some takeaways. As a team, it is important that we all work together and focus on delivering value to customers. It is easy to get distracted with the bugs, with the different requests, but it is important to have in mind the customer pain and the customer problem.

[00:28:02] Also, get yourself involved in all the steps. As I mentioned before, of course we are more active in some steps than others, but it's important to understand that what you do as a data scientist, as an engineer, as a UXer, or in marketing promoting the product, will be impacted by what your colleague is doing. So it is important that we work well together.

[00:28:29] For data science I would like to highlight a couple of things. Ask for feedback in all phases. Early feedback is quite important, so if you have an opportunity to ask for feedback from your colleagues, from engineers, from other data scientists, ask for feedback.

[00:28:54] Prioritize the approach that delivers value for customers. It's quite easy to be seduced by state of the art models and complex models, but sometimes they are difficult to implement or to understand, so again the explainability element comes in here. Keep in mind scalability, explainability, implementation and flexibility when choosing your approach. And be aware of the trade-off between performance and time. We tend to aim for perfection, improving and improving, and we can actually spend a lot of time on that, so it is important that we draw the line when it's enough, so we can actually focus on what is important to customers.

[00:29:48] I would like to thank you for your time. I hope that some of the learnings I shared today are useful to you and your team, whether you're in data science or in any other position. So thank you for today, and I hope you enjoy the conference.

Speaker

Monique Marins

Monique Marins

Sr. Data Scientist

Trustpilot