Data Driven Product Development Approaches

28 Apr16:25 – 16:55Stage: Main StageTalk
Slides

Checking session availability…

Hang tight while we load the latest updates.

Developing new systems and features we have to make decisions that might cause the whole project to fail. How can we take these decisions with confidence? How do we plan for related work? How do we measure success?
In this talk I want to give several examples of approaches that we took and reflect on how well they worked given the requirements we faced. I will cover topics such as identifying key metrics, developing towards a vision and decision making based on data.

Data Driven Product Development Approaches

Florian Biesinger at UXDX Community: Germany. Video: https://youtu.be/86cV0R_3vfY

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Why even be data-driven

[00:00:05] Florian: I'll talk about data-driven product development. I will start with a short introduction about myself. As already mentioned, my name is Flo. I'm working as a staff engineer at Spotify, dealing with authentication, authorization and user lifecycle.

[00:00:23] Host: Florian, I'll just jump in. If you could share your slides as well. Sorry, just to jump in there. Perfect, thank you.

[00:00:38] Florian: I'm sorry. Yes, exactly. I'm working as a staff engineer working with authentication, authorization and user lifecycle. Before we jump into data-driven approaches, I want to outline some key concepts that I personally think are super important to get right. The first thing is why even data-driven? Did you ever find yourself in a planning session where you planned certain features or certain systems that are more than, let's say, three months ahead, and then ask yourself how did that go? I've been in planning sessions where I've personally planned systems, or had to size stories and create stories, that were more than 12 months ahead, and it went horribly. Once we reached that state, 12 months ahead, our product looked completely different from what we envisioned it to look like back then.

[00:02:07] Therefore the question, if you encounter these situations: did you end up building what your customers really wanted, or did you end up building what you in the beginning thought would be useful? And if you still think that you've built what your customers really wanted, how did you measure this? Because in my experience it's incredibly hard to measure customer success, and there's no one size fits all. It really needs to be defined from system to system, and probably adjusted on the way as you continue development.

The Comet and the black box

[00:02:55] Let's follow up with a short story. I'm usually working remotely with a team in Stockholm, but I'm flying once a month to Stockholm to see my teammates, meet them in person, have some one-on-one meetings, build up relationships and build my network within the company. I'm usually flying from Berlin to Stockholm without thinking a lot; I mostly watch Netflix on the planes. One day I asked myself, though, why are the windows round in an airplane? I did a bit of research, and it turns out that in the 1950s there were Comet 1 airplanes, which were some of the first passenger jets, and they quite frequently fell apart while being in the air. No one really realized what the issue was, and so they started banning these airplanes and said, "We just cannot let them fly, because they usually have disasters when they fly."

[00:04:16] So they banned the airplane, and after two weeks of investigations they couldn't find a reason why it was happening. So they said, "Okay, maybe we should let them fly again." They let them fly again, and directly after that the first plane fell apart again, directly after the stop. After a while, of course, that led to a permanent ban on them flying. But after a while they found out that the problems were in the rectangular windows. There were small cracks going from the corners all across the body of the airplane, and once the pressure increased, it fell apart.

[00:05:12] But the way more important finding, and way more important development after they analyzed that, was that they introduced a black box. For all of you who don't know what a black box is, it is a small component that is really hard to destroy, which gathers data while the airplane is in the air, while the airplane is operating. Now they can use this data and constantly analyze what is going on while flying, even analyze that data offline, and make constant improvements to the airplanes. I think if we had just started planning an airplane from the beginning, saying, "We want to be in the air, and these are the exact steps that we need to follow," and then just executed on it, we never would have ended up where we are today. We would probably never have been as fast. I would even make a more bold statement: we would probably still not be in the air if we had not experimented.

The importance of a vision

[00:06:50] In the next part I want to talk about the importance of a vision. We've already heard it quite some times today, that you need to have a vision, but I want to be a bit more concrete here and say what a vision is. In my point of view, in order to formulate a vision, you should ask yourself: what do I see this system doing, or this feature doing, this product doing, whatever you're designing at the moment, in the future? Let's say two, three, four years from now. It is not about an exact feature definition. It can be fluffy, or it should even be fluffy, because if you're trying to be too precise you're going to fail, most likely. We're going to see why in a second.

[00:07:50] Once you formulate that vision, and you usually do that in a group of people who are building a system or a feature together, you're going to share it with everyone who has a need to know, and you've already got the first round of feedback, and you adjust your vision. So why should we even do this, if it's not too precise and it's fluffy? The reason is it is great for aligning teams. For example, if there had just been a vision for a person to check off, "Yeah, we should fly," some people would have maybe envisioned an airplane that does not transport passengers but rather goods, and others would have maybe envisioned an airplane with, I don't know, six wings. So aligning these thoughts made super much sense.

[00:08:59] Another example, from my day-to-day development, is where we had an account activity service. We thought, "Okay, what could this system do?" I will go into that later on in more detail, because this is exactly one of the examples I'm going to discuss later on.

[00:09:31] When you are developing with a vision, this is where you start: you're somewhere today and you envision a future. What will actually happen is that you will most certainly never end up exactly here. You will start iterating on your problems, you will start working, and you will notice that you either pretty quickly derail from your vision and you will end up somewhere here, or you will stay on it quite a long time and then derail from your vision and end up somewhere here. But most certainly you will not end up at the vision, and I think that is not the real task of the vision. A vision is to really align teams up front on what you envision a system to do, and therefore set a rough timeline of what should be done.

The importance of data

[00:10:33] Now let's talk about the importance of data. I already mentioned it: you're going to iterate, you're going to measure. Very often when I talk to friends in the industry, or other people in the industry, I hear, "Yeah, but I collect data." And my question most of the time is: really? Or are you just having some log file somewhere that no one really analyzes? Are you just logging a huge amount of stuff into some file, or do you have an actual plan? Planning how to analyze the system is in my opinion even more important than the actual functionality.

[00:11:27] That immediately leads to the question: what kind of data do you need to collect? I think the minimum data for each system needs to be system health, but that just tells you whether your system is working. What you need to think about for each system that you're planning, or for each feature, is how do I measure customer success, how do I measure customer experience? It is super hard to measure, and it's individual to each company or team or service.

[00:12:04] For example, imagine this account activity system. Let's assume we envision a feature that should send out an email on login from new devices. So on some suspicious account activity, we send out an email. How can we measure the success of this? It is super hard, because you are triggering an event, and somehow in other systems there might be an action, meaning a user does not recognize this login, so they are opening the password reset flow and resetting their password. So this is one example of how you could measure customer success: how long is the time right now between suspicious activity happening and a login to the account, until the user resets the password? That is, where a user initiates a password reset via your flow, and it is not some fraud prevention system or anything else that could lead to a password reset. In that moment you can measure this lead time, and that's a customer experience metric, for example.

Data science and hypotheses

[00:13:43] Another big part in this space is data science. I think in the last five years, in my jobs at least, data scientists became more and more important. Also my understanding, and that's probably just me growing, I don't think that the space changed that much, but my understanding around it I think changed a lot. In order to start: you have data, you probably have a goal, you have the metrics. Now you need to start to formulate a hypothesis. For example, we've already set the example with the email and the lead time. We say if we enable this feature of sending emails on logins, we will decrease that lead time. You should phrase it in a way that it's measurable. For example, right now it's 48 hours, or whatever time you're measuring, and in the future we plan to get this below 24 hours, meaning a 2x improvement.

[00:15:02] Then you can start developing this and you can start measuring. The most important part here is measure, measure, measure, and there can't be enough hypotheses around a system. I think the more you have, the more you want to adjust, the better. And also, of course, if you see hypotheses fail, or you fail to verify them, reflect on why and what you can do better next time.

A self-contained system: account activity

[00:15:39] Now I want to outline two different approaches that we took in our day-to-day life. One is a really self-contained system, which I already mentioned: this account activity system. As we own the logins domain, it is only called by our system. We can identify what devices we have seen so far, what devices we already know, what devices are new. We can define all of this within our systems. There is no one affected around us, except maybe the password reset people, as they will maybe get an increased load. We had no clear definition of what it should look like. We even started here with a name, where we said, "This could be account activity, where we saw activity on the device." Then we started with a vision of where this could go. We envisioned the system that will send out emails on suspicious activity. We also envisioned this data to probably be available to the user, maybe, and we formulated hypotheses, as I've already mentioned.

[00:17:15] What happened was we then started formulating high-level tasks, like: the system should have a database, the system should be integrated into our emailing service. Tasks on this level, not really low-level tasks. Then we just started executing. As my girlfriend always says, the best way of doing something is to simply start doing it, and that's exactly what we did here. Within three hours we had a blank service in production, before any logic, that had 300 requests per second. It was not doing anything, but we already saw how many of these requests the system should expect. So we were already gathering data at that point in time.

[00:18:13] Then we started to formulate a hypothesis again and said, "How many of these do we already know?" We formulated our hypotheses and slowly added the logic that we needed. For example, we usually start to test features on employees, and this was one of them. In the end we just measured and adjusted, and adjusted here doesn't just mean adjusting the code, but adjusting the vision, adjusting the high-level tasks that we set. Then, three or four weeks in, we decided the system now looked solid enough to be rolled out to employees. What are the features that we definitely need to have in place before we roll this out to employees? What are features that we need before we roll this out to our first market, and what are features that we need before we roll this out fully? The system itself is still in ongoing development today, but it has already, as mentioned, drifted away from the initial vision.

A system with external requirements: OpenID Connect

[00:19:37] The second approach is a system with external requirements. The only knowledge that we had was a rough deadline for when we need to support this feature, and that it would be based on OpenID Connect, which is a protocol in the identity space. It is basically to exchange identities between two internet services. For example, Google has an OpenID Connect endpoint, so you can implement an OpenID Connect client, and then you will know who is calling, and therefore create an account for that person and let that account use your platform. The requirements were really fluffy, because we only knew it's going to be based on OpenID Connect, but it's also not really following all the standards, and there was still a lot of development going on on the other side. And we knew it's a new login mechanism for Spotify.

[00:21:01] What we did is we asked ourselves: where do we see this, which teams are involved, how do they see this? Then we started with an aligned vision. I don't really remember concrete numbers, but it's probably around 10 to 20 teams within Spotify, so a lot of people. We brought them all together and said, "Where do we envision this system, or this feature, in the future?" Then we started with clean APIs to our systems and already deployed them, so that people could start developing from the other side, and we measured and adjusted again, as always.

Discovery work and RFCs

[00:21:55] One thing that I missed the whole talk now is that sometimes you have a vision but you don't know where to actually start, and this is where discovery work comes in. Usually we try to spend one or two engineers doing discovery work, meaning checking which teams are involved, what their requirements are, what their limitations are, and then formulate these requirements in an internal request for comments, meaning an RFC document. Then everybody can share their opinion and their thoughts on a specific implementation. But again, if you already have a vision, then you will naturally reduce the comments on an RFC, because most of the teams are already aligned: "This is roughly where we want to go, this is where this team roughly wants to go, and that's why this system probably looks the way it does."

[00:23:06] We worked on this, we measured, we adjusted the vision, we adjusted the discovery work. We probably did some more discovery work during the project as it continued. And then we finally released it. The success story here is that today we can add a new login mechanism to Spotify within a single day. This is just on the backend side, of course, but given that you have so many teams involved, doing such a big change in one day was, for me, something where I was quite proud of my team.

Speaker