A Whole Team Approach To Testing In Continuous Delivery: One Tester's Journey
Checking session availability…
Hang tight while we load the latest updates.
It’s not uncommon for teams moving towards continuous delivery to face a growing backlog of customer-reported bugs. It’s a challenge to maintain the cadence of frequent deployments. If a team has testers, they’re often expected to continue to do all the testing activities, without any thought as to how those fit into continuous delivery (CD) or continuous deployment (also CD). Teams without testing specialists often struggle with insufficient automated test coverage and inadequate exploratory testing.
In this session, Lisa shares her experiences with a team striving to deploy smaller changes more frequently. Her team’s challenges are common ones. They came up with some innovative experiments to find ways to deploy more frequently with confidence. Lisa will introduce some techniques to consider trying, including:
- Visualizing deployment pipelines to shorten feedback loops and fit in testing activities
- Using a test suite canvas to determine the minimum automate test
- Ways to analyze risks and determine next steps to mitigate them
- Ways testing specialists contribute to team success
A Whole Team Approach To Testing In Continuous Delivery: One Tester's Journey
Lisa Crispin at UXDX Europe. Video: https://www.youtube.com/watch?v=RcTTGqZKS58
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Introduction
[00:00:00] Hi, I'm Lisa Crispin. It's great to be able to join you here at UXDX 2020. Wonderful program, wonderful lineup of speakers and great participants, and I'm honored to be here. I'm going to share a story today of one of my journeys with a team who was trying to get to continuous delivery. I've worked with quite a few people over the last couple of years to develop some of this material.
[00:00:31] I've been lucky to work mostly on small cross functional teams where we all pitched in on all those activities along that DevOps infinite loop for the past 20 years. I've also been lucky to partner up with Janet Gregory to write some books about testing and agile development, and also create video courses or live courses that can be offered in person or virtually. I've also done a course on test automation for DevOps for Test Automation University. If you haven't checked out Test Automation University yet, it's all free content, wonderful stuff. If you're interested in the course Janet and I provide, you can go to agiletestingfellow.com and check that out.
[00:01:17] I work remotely for OutSystems. My team is based in Portugal. I am in Vermont on my farm with my donkeys, so in my spare time I do a lot of driving and hiking with the donkeys.
[00:01:32] I'd like to share some of my team's experiences in learning how to build quality into our products, and to be able to fit in all the testing activities and to succeed with delivering small changes to production frequently at a sustainable pace, which is really the very definition of continuous delivery.
A high performing team that got into trouble
[00:01:50] Now, in my experience, you learn more from failure than from success, and so today I'm going to talk about a team that I was on a few years ago that for several years anybody would call a high performing team. Lots of awesome people on our team. Pairing a hundred percent of the time, doing test driven development, embracing all of the extreme programming practices, automating tests at all levels.
[00:02:24] There was really a great culture of quality. Quality was highly valued. Testing was highly valued. And as we got the infrastructure in place and we were able to move our application to be hosted in the cloud and get all our deployment pipelines going in the cloud, having a blue-green deploy where we could have two production instances and easily flip between them, so that we had more safety nets when we deployed to production.
[00:02:55] The managers wanted to start trying to deploy more often. At the time we were deploying just once every two weeks, I think, and sometimes that wasn't always going well, for various reasons. It was a really complex code base, a single page application with JavaScript and Ruby on Rails in the background, and being able to support lots of concurrent users. There's a lot of business logic in the UI. It was a pretty tough nut for automated regression tests and other things.
[00:03:29] Let's try releasing twice a week so that we have smaller changes. They'd be less risky. We can easily revert our deploys now. And that sounded really good, but I was kind of worried, given that I knew our history and some of the problems we had. But we did forge ahead.
[00:03:46] We had literally thousands and thousands of automated regression tests at all levels. However, we still had a checklist of manual release regression checks that we were supposed to do before every deploy to production. Of course, the two testers on the team out of 25 or 30 developers always ended up getting stuck with this manual regression checking.
[00:04:13] I don't know if you've ever gone through a checklist like that, but as you're doing it you start to notice things. "Oh, that doesn't look right now. I have to go look and see if that's already on production or if it's a new bug." Usually the things we found had nothing to do with the checklist itself, but other things that nobody had noticed. And it was a huge time suck. And of course, if we found something, we had to stop everything, get that fixed, redeploy, start all over. It was terrible, and we had no time for really important human-centric testing activities like exploratory testing.
[00:04:49] We were really starting to get into trouble, and we were in fact not deploying twice a week. We were really having problems even deploying once a week or every two weeks still. Kind of a mess.
Dropping the manual checklist and spreading exploratory testing
[00:05:00] But this is a great team, and we were really good at reflecting on our problems together in retrospectives and experimenting with ways to fix those problems. We decided the biggest problem, because we were having a lot of problems in our release candidates and sometimes they were even getting out to production, was that we were not doing enough exploratory testing. Having the automated regression checks was nice, but it doesn't really test your new changes. And obviously having two testers for that many developers wasn't scaling.
[00:05:42] One of the actions that we first did was, we testers said, you know what, we're not doing this manual regression checking any more. That's out. Because nobody else wants to do it and we don't want to do it any more. Guess what? After we did that, no problems ever happened that would have been found on that checklist. Just saying.
[00:06:01] The development managers realized how important the exploratory testing was, and so they actually added exploratory testing skills to the skills matrix for developers, so that in order to proceed on their career path they had to become competent at different exploratory testing skills. And then they asked us as testers to have workshops and teach the whole team, including developers, product owners, customer support, designers, everybody, how to do exploratory testing.
[00:06:35] And also, we started pairing with the developers as they wrote production code, and we could help them not only with their test design but with doing some manual exploratory testing before they decided to check in their code. Another developer and I created a little manual, a little exploratory testing checklist for developers. And again, this is just at the story level. At the story level, where we're working on a very small story, we're doing test driven development, or automating tests for it at the unit level, at the API level, and if it was a UI story at the UI level as well, and doing this exploratory testing.
[00:07:16] The developers started to feel comfortable with these skills. And then of course, this is at a very narrow level, because our stories only took one or two days to finish. Each epic or new feature set was a whole lot of stories. What we started doing also was writing exploratory testing charters at the epic level or feature level, and putting those charters in the backlog along with the feature stories. We used Elisabeth Hendrickson's template for testing charters in her awesome book Explore It! If you haven't read it yet, I would highly recommend it.
[00:07:55] We had all these charters in the backlog. Anybody on the team could start pitching in on doing these charters as we had enough stories done to make it possible to test at a more end to end level across the epic or feature. That was one piece of the puzzle, and as we did that, we started seeing these crazy problems that would be found right before we release or after release — unexpected impacts on other parts of the system from the new changes we were making — we started seeing those go down. We saw immediate changes with that.
Make the problem visible: mapping the path to production
[00:08:28] One of the lessons that I like to emphasize, and this is something I learned from Janet Gregory: when you have a problem, make it visible, so that you can talk about it. And I'm going to give a couple of different ways to help talk about fitting testing activities into continuous delivery using some visual models and frameworks.
[00:08:44] We have lots and lots of types of testing: security, accessibility, performance, exploratory, you name it, and testing it at different levels as well. So how do we get all that done if we're trying to release every day or several times a day? I think it helps to start looking at the path to production that your new changes take, to see where that testing fits in.
[00:09:15] This is a technique I picked up from Abby Bangser and Ashley Hunsberger a couple of years ago. These visuals — back in the days when we could be physically co-located, we could just use index cards on a table or stickies on a wall, but this is really easy to do with collaborative tools like Google Jamboard or Mural or any of these online collaborative tools if your teams are still remote.
[00:09:40] So this is just an example of a visual to start making improvements. Start mapping out your pipeline. And even if you don't have anything automated yet, your code is still going through a bunch of steps and stages to get to production. Whether they're automated or not, things have to happen.
[00:10:01] Visualizing that is basically like a value stream map. You could just do a value stream map for it. Where are the bottlenecks? Where are things waiting on handoffs? Where are the dependencies? Where do we see things slowing down? How can we start, step by step, addressing those things?
[00:10:20] In this example I show the manual steps in yellow, because yes, those manual steps, or human centered testing as I like to call them, they're still part of our path to production. They're still part of our deployment pipeline. They're just not automated. We have to take it into account. And so what we may be able to do is have feature flags that let us keep features hidden in production until we finish the testing and feel confident about it. So with feature toggles, at least, we can do dark launches, progressive rollout, testing in production by turning it on just for ourselves.
[00:10:59] There are lots and lots of options here, because these types of testing in many domains are still really important. There are a lot of big enterprise companies that, because of their industry domain, still need to do user acceptance testing and things like accessibility testing. We don't really still have the tools to automate all of that testing. Even security testing sometimes, and definitely exploratory testing, is going to have a human component to it.
[00:11:30] Start laying these out. Have a meeting, let's take an hour, get your team together, see what your pipeline looks like. Now, I'm telling you to do this, but of course I could never get my teams to get together and do this. I know, it's sad. But I could do it myself, or I could do it with another tester or just one other person on my team. And by doing that I start gathering a whole lot of questions, and they're really good questions, and now I can take those questions to my team. "Hey, this test suite is always flaky." Or, "hey, this stage doesn't really need to be in this production deployment pipeline, because we've already tested that in another pipeline to a test environment."
[00:12:15] I could take all these questions, and we did do a lot of improvements. A lot of our problem was that our pipeline was too slow, so we were always looking for ways to speed it up, speed up that feedback loop, be able to deploy to production faster in case we do have a problem and we want to get a hotfix out.
Flaky tests and the test suite canvas
[00:12:36] Of course, I mentioned flaky automated tests. Think about your own team's automated test suites, if you have some. Our team at that time that I've been talking about relied heavily on automated regression tests and performance tests, and we had thousands and thousands of tests, and then more problems with exploratory testing. But we were still getting some regression failures in existing functionality in production. New changes were breaking features that customers were already using. And we started looking at our tests and we really couldn't trust them. We had a lot of flaky tests. We clearly didn't have the test coverage that we needed.
[00:13:18] One of the things that helped me here was using Ashley Hunsberger's test suite canvas, which she modeled on Katrina Clokie's canvas. And if you haven't read Katrina's book, A Practical Guide to Testing in DevOps, which is available on Leanpub, that's another one I highly recommend.
[00:13:36] This canvas is just a framework to help your team talk about your automated test suites, or perhaps ones you don't have automated yet, and ask important questions about them, to make sure that these tests run reliably, run fast, and that when there are failures somebody looks at those failures. Make sure they're tests that you can trust. I know you probably can't read this, but you can certainly download it from Ashley's GitHub repo and print it out — or not print it out, use it on a sticky, on Google Drive or some other online collaborative document while you're talking.
[00:14:13] Some of my questions for this are: what information does this test suite provide, and to whom? How do they get that information? Did they get a message on Slack? Did they get an email? How does that work? How will we know when a test fails, and who's going to take responsibility for making sure that failure gets addressed? Are you doing pairing on test automation? Are you doing code reviews on your automated test code?
[00:14:48] A lot of important questions here that we may not think about otherwise. So again, even if you can't do this, talk about this together with your team. Thinking about it on your own or with a subset of your team is still going to bring up important questions that you can address, and start step by step improving those automated test suites.
Finding the risky areas in the UI layer
[00:15:10] Now, I keep mentioning how the team had thousands of automated tests from the unit level up to the UI level, but we were still having regression failures. And we started looking at the UI level regression tests, and guess what, they had been created by the developers who were writing the production code. And they were doing a great job of writing the production code most of the time, but the test code, not so much. Some of these tests didn't actually have assertions.
[00:15:42] So, like, what were they testing? They were mostly happy paths, which in some domains could be okay, to just have happy path tests at the UI level because you've covered it so well lower down. But in this case we had a lot of complex logic in the UI level, so we couldn't really test all that lower down.
[00:16:01] What did we do? We got a cross functional group of people from the team together. It would've been too much to try to get the whole team together. There were about 10 people: developers, product owner, designer, manager, testers, customer support. One of the senior developers who knew the front-end architecture drew it on the board, and we started identifying what are the risky areas and writing those out in red. Of course, this is something easy to do also on an online collaborative tool. And this let us prioritize what we needed to make sure that our code covered.
[00:16:39] What absolutely do we need automated regression testing for? And then we could go back to the existing tests, see what was covered already and see what we needed to add. And so we made a commitment to start using a new framework for our tests, use a page object pattern to make better UI level tests so that they could be easier to maintain. And we agreed that as people wrote new code or went back and changed code, we would refactor the old tests or write new tests in a better way, so they were more maintainable and more dependable. And again, we started seeing results as we refactored our tests and added new ones. We had visibility into what we needed.
Surfacing risks, and observability in production
[00:17:24] Now, there are a lot of techniques to surface risks and assumptions that people are not really thinking about, and these are just a few examples of ways you can do it. I like mind maps. I like good old fashioned risk analysis, probability times impact. Risk storming is available online, you can Google that, and that's an awesome thing. One thing we use where I work: OutSystems has led that effort, so before big efforts we do risk storming with the whole team.
[00:18:00] These are really important things to help you think about how you are going to mitigate those risks. Automated tests might be one way. There are a lot of other ways, you just have to plan for that. We can do all this wonderful testing in advance, and we should do all the testing we can do in advance. But what if we have an unexpected load on the system, somebody accidentally drops a table in production, or releases a configuration change and it's not good, or we're using some external API that suddenly starts returning errors?
[00:18:26] Teams now that have complex distributed systems in production, we can't even replicate those in a test environment. We can't test all those scenarios, even if we think of them, and we don't think of a lot of them. We can't put logging and an alert in if we don't know what's going to happen.
[00:18:52] So what do we do? This is where my team — and again, this is three or four years back — we were just starting to learn about new things, and at that time observability was a pretty new thing. We read this blog post from Cindy Sridharan, Copyconstruct on Twitter, and it was like, whoa, eye opening. We need to make sure that we capture the information we need. Because what we ended up having to do all the time is, oh, we had a 500 error, something crashed, we don't know why, customers are complaining, but we can't tell in Splunk what happened. Let's add some log data and wait for it to happen again. And so these things ended up taking days or weeks.
[00:19:35] It's crazy. But what if we instrumented our code in a really smart way, and structured events with high cardinality and high-quality data, so that we can go in and ask questions of our system in production about things we didn't expect to happen in advance? I am still just learning about observability, I am no expert, but we started to immediately see: data, we need more data, we need to be looking at all kinds of data. Even looking at analytics data in Mixpanel was helpful to us. So we started really paying a lot more attention to what was happening in production and finding ways to be able to assess our work quickly.
Building quality in is a whole team responsibility
[00:20:17] Building quality in is easy for me to say, and it really has to be a whole team responsibility. I think this is really what resonated so much with me when I first joined my first extreme programming team back in 2000. It was about the whole team being concerned for quality and testing, and involving the customer in that discussion as well, and in that responsibility as well. It's really easy for me to say that. It's like mom and apple pie.
[00:20:45] How does it really work? Over the past several years I've been really interested in unconscious bias, for a whole range of reasons, and we do have scientific data that shows that companies with more diversity are more innovative and they make more money. I have an unscientific theory that when we have a diverse group of people with different backgrounds, different skills, specialties, different experiences, that maybe it helps us offset our unconscious bias. Maybe we have different ones, or just by getting together we can help each other notice more.
[00:21:17] And so getting this diverse group together and saying, okay, here's the level of quality we want, we're going to make an absolute commitment to that, and whatever gets in our way we're going to find a way around it. I found that making that commitment, that's the thing you have to do. If you're just waving your hands and saying, oh yeah, we want quality, then as soon as you run into a roadblock you're going to say, well, let's just go ahead and hack that in, and we'll look later at how to automate tests for it, or we'll look later at how to instrument that code, or whatever it is.
[00:21:50] It's really, really important to get your team together, have these conversations, make that commitment, make it mean something.
Transferring testing skills: pairing and ensembles
[00:21:54] If you're going to have a whole team responsible for testing and quality, well, what about the people who don't have a lot of testing experience or testing skills? This is where we can transfer our skills in a lot of different ways. As testers we need to focus on being consultants. Modern testing principles: if you look at modern testing, moderntesting.org, Alan Page and Brent Jensen's modern testing principles. Janet had been kind of saying this for years too, but they've really found a good way to say it: testers really need to step up and kind of coach the team, be a consultant for the team.
[00:22:32] And this is what I've tried to do within the past several years. I can't test everything for the team, but I can teach them how to test by doing pairing. I really like strong style pairing for this, where we have the driver-navigator role and switch that off every few minutes, or ramping that up into mob programming and mob testing. I like the term ensemble programming and ensemble testing for that better nowadays.
[00:23:00] But these are great ways, because like with an ensemble you can get everybody you need: product owner, designer, tester, customer support, developers, architect, whoever. Get those people together working on something, and then whenever you have a question you've probably got the person right there to answer that question right away. It's the quickest feedback loop that you can have. No waiting around to find somebody or get on Slack and get them to answer you with this question.
[00:23:27] Actually, one of the ways I used this in this team I've been talking about was, as developers got enough, four or five stories for a feature, together — and we were trying to do really small learning releases and MVPs, getting just a thin slice out and getting feedback on it — so when we had a little piece ready we would get the product owner, the designer, a customer support person, a tester and the developers in a room, and for 30 minutes we would go through all those stories. We didn't do traditional driver-navigator. We switched on each story, and so everybody's making suggestions on what to test.
[00:24:10] And then if we had a question, the person was right there in the room to answer. Like, oh, is this icon in the right place? Oh, is this the error handling we really want? The product owner, the designer, everybody's there. Or the customer support person is there to say, hey, red flag, there's something here that customers are really not going to like, or there's something they really need. And so within 30 minutes we could accept or reject all these stories. That's a great example of the power of collaboration on different types of testing.
Relationships, psychological safety and small steps
[00:24:44] How do you do this? Well, we have to build relationships, and remember we're humans first. It's all about people, getting good people and helping each other do good work. Most of us are still remote and so it's a little harder to do, but my team at my current job, we have a WhatsApp channel where we just share what we're doing on vacation, or some great meal we just made, or some crafty thing we just made. And it helps, because it gets us bonding at a more human level.
[00:25:16] Offer help, ask for help. Asking for help is one of the best ways to make a friend. Try to connect with people. And again, I started this new job back at the end of March, just as everything was in chaos from the pandemic and everybody's working from home and they weren't used to that, and so I've taken advantage of every opportunity I can to build relationships. Having virtual coffees with people. My team has a virtual lunch every week. We just had a party last week on our monitoring and observability team because we finished our rollout of the last alarm for the monitoring. We have an alarm party. These things are all really important, and when we make these relationships we can help have a better environment.
[00:26:05] We know, again from science, that psychological safety is a prerequisite to team success. A lot of people think that's just unicorn land, and I know there are lots of companies that don't have this, the vast majority, but there are more and more who do. When we have leaders with a vision who know how to serve and support people, get out of their way when they should, let them do their best work, which is what agile has always been about — that's when the magic happens.
[00:26:33] If you've never been on one of these unicorn teams, I feel bad, and I hope you will be able to be on one. Because I think if you haven't been on one you don't really know what it feels like. And it isn't easy to get there. It's not going to happen overnight. It's going to be a journey of months or years. But if you keep committing as a team you're going to get there, and you have to focus on the quality first, doing things the right way. That'll pay off later. You'll be able to go faster later. But there are a lot of things you need to learn, like the business domain, so that you can always be doing the amount you need to do.
[00:27:11] I really like to use the principles of continuous delivery to guide our work, and I think they're really in line with the principles of a lot of other things. It's a good way to mitigate risks and put more joy into your work. Those are things you can't do overnight. Don't try to boil the ocean, just do one small step at a time. Have your retrospectives, identify your biggest problem, design some small experiment with some way to measure it, that can be done in a short period of time, to try to make that problem a little bit better. That's what my team had done, and we eventually did get to where we could pretty comfortably deploy code to production twice a week, without any fires in production.
[00:27:59] I've given you some visual models to use and a visual conversation framework to use. Look for more of those that work for your team, and have these conversations. Talk about quality. I'm always available for questions on Twitter and email. I look forward to maybe getting to talk with some of you, and I will make my slides available. I've got some resources that you can use to learn more. Thank you very much.

