Right-First-Time Code: Early Defect Detection for Product Teams
Checking session availability…
Hang tight while we load the latest updates.
Join Fabrice, CTO of Theodo and author of ‘The Lean Tech Manifesto: Learn the Secrets of Tech Leaders to Grasp the Full Benefits of Agile at Scale’, for an open conversation on software delivery, just-in-time and right-first-time. Fabrice will share his innovative “Dantotsu” approach to detecting defects early in the process and systematically analysing them, effectively combining lean strategy with the speed and scale of digital.
Right-First-Time Code: Early Defect Detection for Product Teams
Fabrice Bernhard at UXDX EMEA. Video: https://youtu.be/3eTrSZI8DcA
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Software is eating the world, and so are bugs
[00:00:08] Hello everyone. I'm very happy to be here to talk about a topic that I really care about, which is quality in tech products. And I think it's a very critical topic nowadays, because tech products and software are everywhere. Marc was lucky enough to coin the sentence that has now become very famous, "software is eating the world," and it was true. It's been true for about 20, 25 years. It really started when all these American startups disrupted entire sectors using software, and of course the amazing examples are Netflix, which killed Blockbuster, or Amazon, which has grown to be much bigger than Walmart, and much faster, or Airbnb. And this is not about to stop any time soon. When you look at the Gartner projections on IT spending, they're still growing, even this year.
[00:01:07] Which in a way we could say is good news for us, because we work in tech, we build tech products. So this is all basically saying that we have a bright journey in front of us. "To the moon," as some people would say, probably. Until bugs start to have expensive consequences. Because of course, if software is eating the world, we end up with software everywhere, and we might end up with bugs in places we had never thought of.
[00:01:34] So this illustration is the Ariane 5 space rocket. The explosion cost $500 million and was caused by this line of code. Somebody forgot to cast a 32-bit integer before passing it to a function that was expecting a 16-bit integer. This is not the kind of code we usually work with. This is quite low-level code. But when you think of the price of that one line, it's of course very painful.
[00:02:10] And that's in an industry that is aware of the price of defects. Of course there are extreme quality processes in the space industry. To give you a very good illustration of what that means: when the space shuttle software team added GPS to the navigation system, they added 6,000 lines of code, but they had to write 2,500 pages of specs before that, to make sure they had taken into account all the possible edge cases and really thought through every source of potential bugs. So that means that for the whole program, the whole code that powers the shuttle, we're talking about 40,000 pages.
[00:03:08] So I wanted to see what 40,000 pages were, and I found this beautiful photo. That's a very old photo of Margaret Hamilton, the first software engineer actually, who worked on the guidance software for the Apollo program, and that's her proudly posing next to the code that she wrote, because at the time we were still printing code. But that is not 40,000 pages. That is only 11,000. So if we actually want to see what 40,000 means, well, basically we have to pile it up about four times, and we end up with this huge pile of paper. This is me just copy-pasting in Google Slides, so this is not extremely scientific, but I think it's still a good illustration.
Not scalable versus not acceptable
[00:03:56] And so what is very interesting is: what do we do in the future now, and what is our responsibility, working in tech and working on tech products, to avoid quality issues? Because on one side we can decide to write four pages of specs for every 10 lines of code. Or we could continue the way things have been done. This is a scene from The Social Network, where they represent software engineers at Facebook at the time having shots, in case you don't see it on this small picture.
[00:04:32] And that ends up with two situations. One where we have very few defects: we're talking about one defect per 400,000 lines of code. Another situation, which is very hard to estimate, but the few people who've made the effort to try to estimate the number of defects, we're talking about 5,000 defects per 400,000 lines of code. And that means that at the moment, when we look at our industry, we have a solution that is not scalable, because I'm not sure any one of you is ready to write four pages of quality specifications for every 10 lines of code produced. And on the other side we have a solution which is not acceptable, because the number of defects we're talking about is way too big and not acceptable for a world where software becomes more and more critical.
[00:05:23] So this is something that we've explored at Theodo, and I'd love to talk to you about that. But today it's not just about me, it's also about you. So here I am, switching to the ping [?] system, so that we can start having a shared discussion around these topics of high-quality software that doesn't require writing tens of thousands of pages of quality specifications. So this is not the end, this is the beginning.
Who is responsible for finding defects
[00:06:16] So yes, the first question is really about measuring defects. And what I found interesting is, before you even measure defects: who in your organization is responsible for identifying defects in the products? There's this very funny option: end users. Somebody chose end users. Thank you.
[00:06:49] So yes, that's a very interesting topic, and maybe I'll start by sharing my experience, and then I have this mic that I'm very happy to throw to you, to also share your own experiences. Our inspiration on these kinds of quality topics is very much Toyota, because it's this incredible company that's been able to have tremendously higher quality than their competitors, despite being of course in a very complex environment. And their main learning is that if you really want to have superior quality at the end, you need to start really at the earliest possible point and make everyone responsible, so in that case even factory workers, for the quality. The very famous term around that is jidoka: the idea that any factory worker is empowered to stop the production line if they feel that some defect is moving forward to another team.
[00:08:01] And so of course the question is, how can we make sure that defects are identified less by the QA and more by potentially product managers or developers before that? Does anyone want to share some experiences they have around that? Oh, there's Kevin over here. You can throw it. Maybe we'll encourage more people to throw.
[00:08:29] Kevin: I'll happily take a thrown box. So I worked at a company that was about 5,000 people that had no QA, ever. And then I wanted QA, and they fought against it for this exact purpose. Every single person was required to report, and then we did dogfooding up to the CEO level. So we had to be the courier, we had to be a delivery driver, we had to be customer service. And it was a huge help for finding problems, because then you felt bad if you saw something and didn't say something.
[00:09:03] Fabrice: So what you're saying, which is very interesting, is that this company on purpose didn't have any QA team, to make sure that everyone in the organization felt responsible for quality. Thank you very much for that amazing shared experience. It is indeed like that. We've worked on hundreds of projects, and something we've identified is that when you introduce QA in an organization, you get this really weird side effect that a lot of people stop caring about quality. And so three, four, six months down the line, you have the same quality issues and a much slower organization. So it does require courage, as in your example, but in our experience, when you make sure there's no one officially solely responsible for quality, then you have much, much better results. Any other experience around that? Someone here.
[00:10:05] Audience: I just want to second what you just said, that when getting rid of the QA team you get better quality, because you're making everyone responsible for the quality. There was no option there, but everyone should have been an option, because of course, depending on what kind of defect you're talking about, if it's a bad design or something like that, that can also be a defect, and then of course it's the product manager, maybe a designer. And as a product manager you should constantly be testing your own solution, and if there's something wrong, it's your responsibility to figure out what's going on there. But day to day, on the release cycle level, it's the developers, and they have to test everything, and if they break something in production, they have to fix it.
Classifying defects by stage of detection
[00:10:57] Fabrice: That's a really good point. Sorry for not putting everyone on that graph. Two things I can add to that. One thing I found very useful is to classify defects by stages of detection, which means that you can actually visualize whether your organization is improving on that. The way we do it is A, B, C, D, E. A is dev. B is the team, so continuous integration. C would be product management, QA. D would be monitoring in production. And E would be the customer complaining.
[00:11:34] And the idea is you can have this management push that I think is healthy. Because if you say, okay, we need to have fewer defects, the organizational answer to that, at some scale, is of course just to measure them less, and you have fewer defects. That works, it's super easy. But if you say, okay, we want to have fewer D and E, but we want to see more A, B and C, or more A and B, then you get this healthy thing where you can see that people are actually putting in effort, identifying more defects, but early in the process. So yes, any other question before we move on to the next question? There are three questions, I think.
[00:12:16] And one last thing. What is very interesting is, if you make teams own the quality, there is still the need for some transverse quality team. For me, that would be the role that makes sense for a quality team: to be responsible for transverse bugs, where there's no one clearly owning the resolution and the analysis, and you need someone to coordinate multiple teams.
Shifting detection left
[00:12:46] All right, the next question goes a bit in that direction: what is the furthest upstream, or left, you have shifted detection of defects? So yeah, it's a beautiful graph. It's very interesting to see that. I don't know, because a lot of you are more on the product side and UX side, but it's very interesting to see that on the tech side, senior devs most of the time believe that test-driven development is a good thing, and rarely do it. So that's a bit the reality of our industry.
[00:13:57] And so yes, I think if you really care about quality, there's a huge need to bring it as far left, as far upstream, as possible. I say left because the industry term that is often used at the moment is "shift left" of quality. Then it depends on whether you read from left to right or right to left. And there's not just TDD. There are linters you can use, so static detection of defects in your code while you're writing. You can use languages. That's a very important thing, the choice of language. There are a lot of studies that show that when you use TypeScript instead of JavaScript, you reduce the number of defects by around, I don't remember the exact number, but 20, 30%. You can use frameworks that, just because they reduce the complexity of your apps, also reduce the number of defects that the teams will introduce.
The right-first-time star experiment
[00:14:52] And one very interesting experience that I'm sharing here before sending you the microphone: because we introduced this classification of A, B, C, D, E, the hard one is really A. How do you measure the number of defects that were detected by the person actually coding them? But some teams found it fun to actually try and gamify it. The way they did it is, when they would start working on a small feature, they would put it in Notion, and then they would code it, and then they would check the number of times they refreshed the browser or the phone simulator to check if the feature was actually working as intended. And if they had needed more than one refresh, they considered that as not a right-first-time piece of work, and they would actually analyze what were the wrong assumptions they made that made it not work the first time. And every time it would work the first time, they would just give themselves a star. And so at the end of the week, in Notion, they would count who had the most stars in the team, and they would celebrate that.
[00:16:01] So that's a very interesting experiment, and the crazy thing about that team is they actually ended up having a rate of defects of, I don't remember the exact numbers, but about 10 times less than the typical team in software. So that was huge. I mean, it's one piece of anecdotal evidence that shifting the quality to that level, the level of each dev owning their own quality, has a huge impact downstream. And what is also very interesting is we never introduced the words test-driven development. But of course, everyone in that team started doing test-driven development, for the sake of collecting a star every time they would test their feature.
[00:16:45] All right, any shared experiences around shifting left? Maybe a question, because a lot of you are in product: can someone share an experience where they were struggling to convince the devs to do quality checks themselves? Or, on the contrary, a successful implementation? Yeah, we have a question. Thank you so much.
[00:17:33] Audience: We have a QA team, and basically what the developers usually are saying is: we have a QA team, we write the code, we do the code reviews, and by our conscience we are doing everything right, and that's why we have a QA team, and they should check for the defects. They are not creating defects on purpose, but if defects are detected, they are not taking it as their responsibility. They are fixing them, but they are not thinking that they did anything wrong or that they should have checked that. So basically they put all the responsibility on the QAs to find the defect.
[00:18:12] Fabrice: And I guess everyone is on the same page? Have you seen some people convince themselves that they should do more about it, or is that the current status quo, without any change?
[00:18:27] Audience: From the dev side, all the dev team are on the side of "we have QAs and they will ensure the quality of our product." We are trying to make this shift, to give more responsibility to the developers to check their code. But that's the current challenge I'm dealing with in my team: how to convince the developers that they are as responsible for finding the defects as QA.
[00:18:54] Fabrice: Okay, one related question: does anyone have QAs that actually code tests? Someone here in front who doesn't want to share. That's an interesting way. If you have a QA team, one solution we mentioned is just to get rid of it. That takes a bit of management planning, I would say. One solution that's probably more progressive is to get them to write the tests. I've seen that in quite a few organizations. That at least gets you to the continuous integration step much more, and I think it also creates a better dynamic, because potentially the devs can also contribute to the tests that are coded. Here, someone's raising their hand.
What counts as a defect, and selling TDD
[00:19:53] Audience: Yeah, two things. When you say defect, that's questionable. What's a defect? Is a defect when software doesn't work as specified? Is a defect when something doesn't work as you, the user, expect it to work? Is a defect something which worked before and doesn't work now? And I think the definition of a defect can reduce defects. That's one thing. We should not reduce defects through what is probably the wrong definition. And on the shift-left thing, I'm actually seeing a bit of difficulty the other way. When developers are very happy to get into TDD, but because they were not doing it beforehand, it requires a bit of thinking, refactoring, and then modifying the code ways and the processes. And how the hell do you sell it to management, that basically we are rewriting code to do the same thing which it was doing before, but also writing tests to prove it? And then your business value from that is the promise that we will be faster. But that promise is sometimes hard to sell.
[00:21:12] Fabrice: Two really good points. And maybe on the first point, I went a bit too fast. We had huge debates around the definition of defects. The definition we decided to settle on is any behavior that is unexpected by the user. The reason it got into such heated debate, I'll give you a very concrete example. We had a user complaining because there's a button that says "login." It's in a French context, and so they ask, "What does login mean?" And the devs are like, "Come on, this is not my problem that the guy doesn't speak English and the content team put an English word there." But we said no, now we will say this is a defect.
[00:22:03] However, to not add another debugging responsibility on the devs, we said we need to identify where the defect came from. Typically, if it's a wording issue, like using an English word, then you would say, okay, that's part of the content team, so it's the content team that needs to analyze what they could do better next time to avoid having these kinds of issues. So yeah, a very large definition of defects, but again also making sure everyone is involved. There can be UX defects, there can be content defects, and therefore the content team and the UX team and the product team can also take their part in that. So that's part one.
[00:22:49] Part two: yes, it's very interesting. It's like every investment. I think TDD is a muscle. The issue is, once you are strong, it takes no time. The problem is who invests in the initial training. The easy answer, which sometimes people give, is to say the devs should not mention the tests. That's part of the way they write code. And I would say, if you take the example of this team, because I think it's a very interesting small experiment we had: because they were very focused on code that works the first time, I'm pretty sure that from the product point of view or client point of view, there was no difference in time spent. But that's because, I guess, the contract at the dev level was very clear. They were not doing TDD to add tests, they were doing TDD to get done faster. So it was an initial investment, but at the day level, to get the day done faster. So I think the muscle was probably trained in one, two weeks, and probably not really visible as an investment.
[00:24:02] Now, when you're on a very large code base, that's a very different problem, because then you really need to invest a lot of time, and that's a CTO decision, or an architect decision. And that's another level, which is how do you negotiate quality with shareholders or stakeholders. Well, we'll skip to the next question before I start delving into that. So, final question. No, I don't. So we need to go to the next question. Sorry, can we use the next slide? All right, okay, it's coming. So this time there are multiple options possible.
Post-mortems and defect analysis
[00:25:42] And maybe if anyone wants to share the kind of post-mortem or defect analysis process they have. Ah, someone here, thank you very much. So we need a microphone for the lady in the front.
[00:26:06] Audience: Yeah, for example, in the company where I work, I work in an automotive company, there was a problem with the design of the handles. They were a too innovative design, but not practical at all. The shape was not intuitive, and this is also, in terms of safety, a risk. And we had a lot of complaints from customers. The designers are the ones that have learned that this is actually a defect, because if you push your design to the extreme without taking care of the use, it's probably something that the customers don't expect. So in this case it was a big defect. We had to redesign basically all the handles of all our cars.
[00:26:57] Fabrice: Oh wow. So one advantage we have in software is that it's usually less costly to repair defects. So a bit of empathy for you, who has to recall cars everywhere whenever you make a bug. Thank you very much for that example. I think it's a great example. We don't know if it's Tesla and their weird handles. It could be. Anyone else who wants to share a bit about post-mortem culture? Here at the back. Ah, finally someone throwing, thank you.
[00:27:41] Audience: Well, I'm a developer. In my team we have a rule that every time a defect is detected, we analyze it, and then when it's getting fixed, we ensure that we cover it extra, either with unit tests or with end-to-end tests. That's one process we follow.
[00:28:04] Fabrice: But who else is involved? Is product involved in the analysis? Do you share that with more people outside the team?
[00:28:12] Audience: No. I mean, it depends what kind of defect it is. If it's something that's coming from some external package we're using, then we obviously share the knowledge with the other teams. But if it's something that's very specific to the product, then no.
Prevent and detect, and negotiating for quality
[00:28:30] Fabrice: Okay, well, thank you very much. I guess I need to wrap up, but really quickly, some key learnings for us on that, which are very interesting. We had a lot of debates around: should you learn about how to prevent it? Typically, how can I train better, how can I refactor the system so that you can't even introduce the defect? Or should I learn about how to detect it earlier: what kind of additional test could I have added that would have spotted the defect earlier? And one thing we learned from this amazing book called Toyota radical way of quality [?], from which we've taken so many ideas and applied them to software, is that you should actually always do the two analyses, because they're two very different ways of thinking, and both have a lot of value. How could I not introduce the defect in the first place, and how could I have detected it earlier? So two key learnings for the organization.
[00:29:25] We've also done a lot of work around standardizing the way we present the analysis, in a way that fits in about one or two pages, so that you can actually print it, so that managers can have access to it. So you don't need access to GitHub, you don't need access to different tools, which would mean that your analysis will never be able to travel outside of the team. And to go back to the question of negotiating with stakeholders, I think that a very good defect analysis, that you can export as a PDF to make it simple, is definitely a tool that you can use to justify investment in quality. Because everyone in the leadership team that cares about the product will find it fascinating to actually read those things and learn themselves. And that can be, I think, a really good way to convince them around investment in quality.
[00:30:31] And by the way, when I talk to leaders in a lot of organizations, they are all very interested in quality. It's more that they don't see a process or a system they trust enough, and that's their main issue. The main issue is not that they don't care about quality. The main issue is they don't see the system that would actually help them. So bringing a good system, and that's what we found in this dantu [?] book, which of course, this is the time to plug my book, we mention in The Lean Tech Manifesto. I think that's a really key element in the negotiation. Oh, thank you so much. Well, thank you everyone.
