Defining & Adopting Your Delivery Metrics
Checking session availability…
Hang tight while we load the latest updates.
Building great products starts with building high performing teams. In this session Patricia will talk through that high-performing starts with clear and tangible metrics.
She will touch on how metrics should be defined, measured and adopted by the team to ensure not only product success but also happy teams.
<i>If you have any feedback for Patty, please share your thoughts through this link: https://docs.google.com/forms/d/e/1FAIpQLSci860OH8kZq0E_jxhRhSGRpUR7KHmKMadH2zBDuyi4x33YtQ/viewform</i>
Defining & Adopting Your Delivery Metrics
Patricia Trejo at UXDX Community: LatAm. Video: https://youtu.be/wTL3uCTokEU
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
What flow metrics are and why they matter
[00:00:00] Hi everyone, hi people, how are you? I'm Patricia Trejo, you can call me Patty. I'm engineering director and head of software development at Cornershop by Uber, and today I'm going to talk about delivery flow metrics in software development.
[00:00:14] First of all, what the heck are flow metrics? In this case, flow metrics are tools to constantly measure the flow of the development process. That's to measure it, to learn from it, to optimize the flow, and thus transform our development process and its results, hopefully, into predictable ones.
[00:00:43] Take into account that one of the things that most projects have in their first days, regardless of their nature, whether it's a waterfall project, an agile one, I don't know, any kind of project, is that stakeholders start to ask these kinds of questions. For example, how long does it take to deliver? Well, it depends. I don't have a crystal ball to see the future and to say, okay, we're going to deliver at this exact moment. Another typical question is, when are we going to finish the project? And the same thing: I don't know. Maybe you have some roadmaps, some planning, some Gantt, and you have some desirable due times, but it's not that that is going to happen exactly.
[00:01:45] That's also why, instead of talking about projects, we want to talk more about products and how products evolve. In that sense, based on the flow metrics information that we have from our process, we can have some metrics that will help to answer some questions. For example, how long does it take us to deliver value to our customers? That is a super simple metric, which is called lead time. Lead time is the time that occurs between a card leaving the backlog and reaching production. On the other hand, a question like how much customer value are we delivering at any given time is answered by another super simple metric, which is throughput. Throughput is the number of cards or tasks or features delivered in a period of time.
[00:02:46] Based on these kinds of flow metrics, I want to share with you some types of information that you can start to have in mind to make decisions based on metrics and objective data. I have split them into six types of information. The first one is monitoring and risk analysis, then motivation and awareness, backlog analysis and prioritization, behavior analysis, projection, and basis for execution of experiments. I will talk a little bit about each one of them.
Monitoring and risk analysis
[00:03:30] The first one is monitoring and risk analysis. First of all, let me explain this graph. This graph shows how much time a card spends in each flow state. We take the flow from issue tracking tools such as Jira, Rally, Trello, Shortcut, Pivotal Tracker, among others. That is the source of our information, our issue tracker. The vertical axis shows the cards' IDs and the horizontal axis shows the number of working days. That's super important: working days. It doesn't have any holidays, it doesn't have any weekends, for example, to have a clear view related to when the team is working.
[00:04:22] For example, card 6915 was one day in analysis, one day in ready to go, three days in progress, two days in QA and three days in review. If a color is not shown, as you can see in some of the bars for some of the cards or tasks, it is because the card spent less than one working day in that state.
[00:04:49] Let's see an example. Here we have this graph, and we can easily see that there are some strange behaviors. For example, these first cards: why did the cards located up here take so long to be delivered? We are talking about more than 100 working days, and the other ones not. So this is a super rare behavior here. The cards that come further down have much more normal behaviors. They have a life that does not exceed five days in general, except for some that reach up to 20 working days, but they're much smaller than the other ones in terms of their lead time, the time from when they left the backlog until they reached production.
[00:05:56] Here is also something interesting. The card IDs on the vertical axis are ordered from smallest to largest, and in general issue trackers assign the cards' IDs consecutively. So we can conclude that the cards that took a long time are older, because they have smaller card IDs. Indeed, in this particular case, the team didn't know how to use the issue tracker well at first, and therefore they did not move the cards from one state to another when it was appropriate, just for example. These unwanted behaviors were generated and are otherwise unreal, because the cards had already finished.
[00:06:42] In any case, there are still process things that the team had to work on: make blockages due to dependencies visible, break the work into smaller tasks and not a single task, carry out the PR reviews as soon as the development finishes, among others. That is the kind of thing a team has to start asking themselves: what is happening, and how to improve to make the flow, first of all, real, but then more efficient. So the team is going to deliver value and deliver part of the product constantly, in an efficient way.
[00:07:26] Then here we have another case. What do you see that looks strange in this graphic, just at first sight? Well, this one: we have just one day in development, or in progress, and then we have a lot of days in QA. So why is that the case for these cards? Why does the QA process take so much time? What happened, or what is happening? In case the card has not been delivered, we can have these graphs for cards that are in progress and not yet delivered, for example.
[00:08:11] Those kinds of things are related to monitoring and risk analysis, because here you can see risks: what is happening along the development process, along the product delivery or along the team. Maybe the team is demotivated, I don't know. It could be a bunch of things that could be happening. And monitoring, to see if our flow is efficient and if we are delivering at the right time, et cetera.
Motivation and awareness
[00:08:42] The second type of information is motivation and awareness. Here we have, for example, a super simple graph that shows the throughput of cards or tasks delivered each week of a year. Why does throughput vary so much from week to week? Here at the beginning it was around seven or eight cards per week, then it started to grow, and now it's decreasing a little bit. Here there were a lot of features, final features for our end users, that were delivered, but in the later weeks we don't have many features, and chores, technical tasks, have grown. And here in the last week a lot of bugs were delivered.
[00:09:45] What is happening? Is our development process not appropriate? For example, it could be that we are not doing some technical things very well that are creating more bugs, or the specification of our tasks is not the best, so we are creating and then delivering more bugs. What things happened such that some weeks delivered many bugs instead of features, for example? Those kinds of things are questions you can ask yourselves, to start wondering what is happening, what could be done better and what we can do as a team.
[00:10:20] This is another example. This shows working versus waiting time. Working time is the time that someone is effectively working on a card or a task. For example, when you have your issue tracker you have some columns like in development, in QA, in review, and that means that someone is working on that card, doing the development, doing the QA, doing the review. That's the working time. Then there are others that are waiting times, or waste times if we're talking more lean vocabulary. Those are the states where cards are waiting for someone to take them, for example ready for development, ready for review, ready for QA, ready for deploy. Those are the kinds of waiting stages.
[00:11:10] Here you can see, again, each one of these cards, the card IDs and the working days, and you can see how much time there was someone working on that task and how much time the card was waiting for someone to take it. You can see here there are a lot of things that were waiting a lot of time, and the later ones were much more efficient: they didn't have much waiting time and had very little working time.
[00:11:42] It talks about motivation and awareness. Awareness because, again, we see which things we can change. Motivation because sometimes when we are working so hard on a task and we do not deliver that task, we get demotivated: where is the work I have done? So if we are waiting too much to reach production, for example, that could demotivate the team, and it's super important to have those kinds of things in mind.
[00:12:17] A final example here: again, maybe you can easily see some strange or weird behaviors. Why, if we have cards that take very few days in development, for example here one day, one day, one day, two days, three days, would it take so long to review their PRs? Seeing the graph, this seems to be a behavior that we should discuss in the team, to see how we improve our development process. Maybe we are not doing the PR reviews in a good way, or maybe we are iterating too much because there were a lot of undefined things when the card was developed, and now we are iterating between review and development. Or maybe the ones who are doing the reviews are on vacation, I don't know.
[00:13:24] It could be a lot of things, and each team has its own context, so each team has to start asking themselves what is happening here. But usually a PR review shouldn't take longer than the days that were spent developing, so these are super weird behaviors. Those kinds of things are also for awareness, but also for the motivation of the team. Maybe the team doesn't know very well how to do PR reviews, I don't know, it could be a lot of things. That is the kind of thing for this second type of information.
Backlog analysis and prioritization
[00:14:02] The third one is backlog analysis and prioritization. In this case, this is a new graph I'm presenting here. This graph is called a burn-up. It shows, in a cumulative way, the delivery of cards by the team and the cumulative growth of the backlog. The delivery is the red curve and the backlog is the blue one. It is necessary that we are constantly monitoring the behavior of the delivery by the team versus the growth of the backlog.
[00:14:37] In this graph we have two important insights. On the one hand, the delivery behavior of the team has decreased. It was, oh, super cool, and now it's a flatter curve. On the other hand, the backlog has started to grow a lot. Both effects mean that it will take a very long time for the team to finish the delivery. And also, let's face it, the backlog always keeps growing, so if we look at it from the perspective of a project, it's possible that it will never finish.
[00:15:16] So what to do? On one hand, we have to see what happens with the constant growth of the backlog. Maybe we should do some review of the backlog, some backlog grooming, redefine which release each card belongs to, realign development expectations with stakeholders, remove cards, archive cards, I don't know. On the other hand, the delivery has dropped, and it's necessary to be clear about why, what has happened. If this burn-up is built only on the basis of number of cards, for example, not points (but we can also do it with points), could it be that the most recent cards are bigger, so they take longer?
[00:16:05] Could it be that it's vacation time and we have less development capacity in the team? Could it be that we have requirements, features to develop, that are less defined and therefore iterate a lot before they can be delivered to production? Could it be that we have reached a point where our cards are simply not well written or defined? It could be, again, a bunch of different things, and that will change depending on the context of the team. Unfortunately, there is no one-size-fits-all answer for this, because it's always going to depend on the team's context and reality.
[00:16:48] This is a second example here. In this case, cards are shown that haven't been finalized yet. The red color in this case represents the time spent in the backlog. So why do we have so many cards in the backlog, for example here up to 40 working days? Maybe we should do some backlog grooming again and see if it still applies to have those cards in the backlog. Maybe yes, maybe not. Maybe some have already been solved by other cards and we must archive them, or maybe it's okay to have them there in the backlog. Those kinds of things are also necessary to do constantly.
Behavior analysis with lead time distributions
[00:17:41] Now the fourth type of information is related to behavior analysis. Here I'm going to show you a new graph based on the lead time concept. Remember the lead time concept I told you about at the beginning of this presentation: lead time is the time that occurs between a card leaving the backlog and reaching production. This graph shows how lead time behaves: how many cards, on the vertical axis, have a specific lead time, on the horizontal axis.
[00:18:18] For high-performing teams, we would love most cards to live for a short time, which means the lead time is small. This is after the first four weeks of a team, and we don't have much information, so we're not able to say anything at all at this point. But if we start tracking this behavior around the fifth week, tenth week, 15th week and the 20th week, we can see that there is some behavior. We can see that most of ours used to have small lead times, which is awesome, because that means we are delivering value fast for our product. Also, it's similar to a beta distribution, but that's more of a statistical thing that we can talk about in another presentation, at another moment.
[00:19:21] If over time we start to see that there are some peaks, some local maximums, we can have them in mind to do, for example, estimations. We can analyze it and improve our process. Or, if we see that a local maximum starts to grow at a very late moment, a very big lead time, maybe we should make some improvements in our process and start to ask ourselves what is happening. Maybe we are creating tasks that are too big, and maybe we can break them down or split them into smaller cards that could bring value faster, reach production and bring value faster.
[00:20:21] I have already said that we can see several local maximums in this aggregated curve, the dotted black one. We can see here that several cards are delivered at four days, several at six days, then at eight, then at 11 and then at 14 days. This could be useful if the team estimates, for example. Maybe instead of estimating we could do some t-shirt sizing using this info, saying, okay, a small t-shirt is four days, a medium t-shirt is six days, a large is eight days, extra large 11 days, and extra extra large, I don't know, is 14 days, for example. And remember, it's always working days.
Projection with Monte Carlo simulations
[00:21:17] The fifth type of information is projection. In this case I will show you the burn-up graph again, but now we have two new curves, the green and yellow ones. The green and yellow are curves that produce projections based on Monte Carlo statistical simulations, which are based on the team's delivery behavior, the red curve, and the existing backlog, which is the blue curve. The yellow and green curves show when, based on that information, this backlog could be consumed, delivered.
[00:22:10] Now, first of all, we have some problems here. First, the backlog will probably continue to grow. Second, the behavior of the team may vary its throughput over time. And third, there is little information about the delivery yet, and that's why the yellow and green curves do not even manage to intercept the blue curve on this graph. This graph shows the burn-up of a team that has been working on a project, or product, for a short time, but we already know that the backlog was big from the beginning. This is the number of cards, almost 500 cards. That sounds more like a traditional project than a leaner project, for example.
[00:23:01] This one is the same graph, but it shows the burn-up of a team that is already ending the delivery of a product. The team was constantly, over a long time, reviewing, cleaning and prioritizing the backlog. So when the backlog grew too much, we started to prioritize and realign the expectations, and the backlog kept growing, but not in a super heavy way. That meant that the difference between the delivery, the red curve, and the backlog remained more or less constant, and there came a point where the backlog stopped growing, so we can see that the end is near.
[00:23:46] In addition, the delivery information has already been fed into the forecasting algorithm, so the yellow and green curves are almost the same, and that means that the uncertainty of finishing the backlog is simply within a small range of possible dates. These two different curves are based on different percentages of confidence. The green one has 50% confidence and the yellow one has 70% confidence, and that's why, based on this simulation, they are different.
[00:24:24] Regarding this, be careful. We cannot make projections with 100% confidence, because otherwise it would be like having a crystal ball that predicts the future, and reality does not work like that. We all know that it doesn't work that way.
Flow metrics as a basis for experiments
[00:24:44] The final type of information is using flow metrics as a basis for the execution of experiments. Why, and what for? First of all, we can always run experiments in our teams to see which things fit better than others, or how behaviors improve with some things rather than others. The idea is to evaluate the state before and after an experiment, and we can do it with data, and this is super objective data. From that I start to see new ways of doing things and opportunities for improvement inside the team.
[00:25:33] Here is a real example. I didn't say it at the beginning, but all the graphs shown here are real. They're all graphs from different kinds of teams and different kinds of products that we have been developing over the years. Everything is real, and this example is also real. There was a hypothesis, once upon a time, where someone said: why have pairs of developers working on a single card? If each of the developers took a card, we would go twice as fast, instead of doing pair programming.
[00:26:22] We started to do a lot of analysis. I will show very briefly here just two little examples. For example, this graph shows a person who in the first and third weeks was working alone, while in the second and fourth weeks they were pair programming with someone else. The average development time of the cards that person worked on that week was in some cases shorter when working alone than when pair programming, but in other cases it was longer, and the same happens when working in pairs. So there isn't any pattern to see here of how I deliver features over time, whether I work in pairs or not.
[00:27:22] The second scenario is for a person different from the previous one, who always worked pair programming. But the same thing happens again: there is no pattern in the average time that a task takes. The tasks are different. Furthermore, there is no pattern that says whether pairing takes you longer or not, so you cannot conclude that the initial hypothesis is true.
[00:27:49] I have already told you this is a limited real, but not complete, example, because there were a lot of cases that we were analyzing. In fact, we repeated this same analysis at the beginning of the team and of the product development, when they had already been working together for some months, and then almost at the end of this project, of this team delivering a product. We never, ever found any pattern at all. There is no pattern. We tried to compare just little tasks with other ones, just bigger tasks with other ones, and we never found any pattern.
Recap
[00:28:38] At the end of the day: "If you can make decisions based on fact rather than forecasts, you get results that are more predictable. Lean development is the art and discipline of basing commitments on facts rather than forecasts." This is a quote from Mary Poppendieck, from a paper she wrote called "Lean Development and the Predictability Paradox." If you have the chance to take a look at it, I fully recommend this paper.
[00:29:08] Just as a final recap, some important points. First of all, make decisions based on facts rather than forecasts, please. If you have facts, if you have data, it's much better than saying, I don't know, I think this could work, or I think we are not doing things right. So please, if you have data, use it. It's a super powerful tool. Almost all issue trackers have an API, so you can integrate with the API and extract the data to analyze it, see what behaviors are present, and start doing the analysis itself.
[00:29:54] Third, start with the simplest flow metrics, such as throughput and lead time, as I have already told you. Then continue by splitting the simplest flow metrics for different scenarios. For example, I'm taking a look at the macro throughput, but now I want to see it for each team or for each kind of card or task, I don't know. Lead time and throughput are super simple, and you can get a lot of info from them.
[00:30:28] Show your team the metrics you obtain and all the interpretations, so together you can make decisions to make improvements. Don't leave the decision to just one person, to the team leader for example. Try to involve the team, and the team will also gain more ownership. And finally, train your stakeholders on the meaning of those metrics, behaviors and analyses, to get aligned, for them to understand what you are doing and the information you're going to show them, for example related to projections. That will also calm them down regarding when we are going to end the project and those kinds of things.
[00:31:14] So please, again, have in mind all the data you have. There's a lot of information in there. Thank you all. Before you go, please take one minute to tell me how it went. I will share the link to the feedback form, and here is also my LinkedIn handle if you want to talk even more in the next days. Thank you.
