Mixing Methodologies: Qualitative & Quantitative Testing for User Behavior Analysis

Apr 1919:30 – 20:00 UTCStage: Main StageTalk

Checking session availability…

Hang tight while we load the latest updates.

Help software product development team unlock the power of layered quantitative and qualitative testing to better understand their users and deliver effective product solutions. Improve KPIs with techniques such as user interviews, AB testing, unmoderated tasks, and more!

Mixing Methodologies: Qualitative & Quantitative Testing for User Behavior Analysis

Charlotte Cunningham, TJ Bowen, Aaron Knoll at UXDX Community: Testing and Scaling the User Journey. Video: https://youtu.be/-hNRAoQGr9w

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Introductions and why testing pays off

[00:00:00] Charlotte: Hi everyone. Thanks for having us, Rory. We're really excited to speak about some user testing methods today. Like you mentioned, we're speaking about mixing methodologies. We all know there are a lot of different ways to go about user testing. We think that all of them have their value and place in building a really high-quality product. We're going to look at what they are, why each type of testing is valuable, and then how you implement that, both at the design team level and at the organizational level. I'm Charlotte Cunningham, I'm a product designer with a company called Crafted.

[00:00:37] Aaron: Aaron Knoll, product designer, also at Crafted.

[00:00:41] TJ: Great. TJ Bowen, 10 years of product development experience, so I come from the product side. The three of us all work at Crafted, a product development consultancy where we believe that balanced teams of designers, engineers and product, working in tight collaboration and quick iterations, representing the UX perspective, feasibility and business viability, are the best way to build great products.

[00:01:12] Charlotte: All right, jumping into the content. Why should you listen to us? I'm sure we've all got experience with user testing, so what insights are we bringing to the table? Being from a consultancy, we bring a really unique perspective, in that we've had the chance to bring positive impacts across a really wide range of industries. We've seen everything from startups to Fortune 200s, and how implementing user testing goes at all the different levels, from legacy tech stacks to greenfield products. Across all of those we've been able to deliver quality results with measurable positive impacts on whatever project we're working on.

[00:01:50] Some examples of this: we had the chance to work with a really extensive A/B testing system at one client, and we saw the chance to deliver up to a 10% increase in purchase or conversion rate through a single test. That was a really great example of quantitative testing at the enterprise level and how that brings value. At another, smaller company, they did job onboarding and recruitment, and through user feedback we were able to identify a completely overlooked pain point in the user journey that was the main contributor to user drop-off, which is one of the primary KPIs. Through that we were able, early on, to steer the project in a new direction that we knew would bring the results we were hoping to find.

[00:02:39] And combining contextual studies and analytics tracking, we've also reduced abandonment for a non-profit application that initially took three-plus hours and required multiple sessions, and we brought it down to 25[?] compared to the previous year. Those are just some examples of the great results we've seen user testing, both qualitative and quantitative, bring across our experiences, and we're hoping to share that with you all today. Anything to add from you two before we move to the next slide? Perfect. I'll let TJ speak to why your business should listen to us. Sometimes the hardest thing to get buy-in for user testing is from the larger organization.

[00:03:21] TJ: Great. Yeah, sometimes the perspective of a business might be: hey, why should we invest money and resources into testing, qualitative or quantitative or both? What's in it for us? The perspective here is that at the beginning of an initiative, the risk and uncertainty are very high, so decisions made early on can really make a big impact, positive or negative, on the direction and outcome of your undertaking. The better information you have, especially in the form of quantitative and qualitative research, the better decisions you're going to make early on, before you're investing a lot of money, like development time, into the initiative. It's just really essential for the organization to invest appropriately in research.

[00:04:14] Aaron: People often ask where quantitative and qualitative fit into the process. Many of us have probably pitched the British Design Council's famous Double Diamond to our stakeholders and our clients. Oftentimes there's a misconception that maybe qualitative especially only has a place on the left-hand side of things. As you're going to see in our case studies, my pitch, our pitch, is that qualitative and quantitative loops appear throughout the process. They're not just for problem discovery; they're also for vetting your solutions and continuing to test your product once it's out there. If you've shown someone this diamond, you can say: testing everywhere, every part of it.

Defining quantitative and qualitative

[00:05:02] Charlotte: All right, now you've seen the high-level value of testing. Getting down into the methodologies of qualitative versus quantitative: how are we defining the difference? Everyone has their own approach here, but for us quantitative is data-driven testing. It's whenever we use numbers to make design or business decisions. You can collect quantitative data about a lot of different areas: your populations, user behaviors, key performance indicators. Some examples of quantitative testing are traditional A/B testing, when you have a green button and a red button and you want to see which one people click more. It can even be multivariate testing: you've got a lot of different colors, you want to see what people click more, and maybe you see people click warm-tone colors more because it aligns with your brand.

[00:05:56] Some tools for quantitative testing are Optimizely, which integrates with your website and performs web analytics, and FullStory, another web analytics tool that also helps with visualization of that data. But you can even go down to the granular level of surveys or impression testing, and do really affordable, quick methods of getting some data on your concepts.

[00:06:20] Qualitative is the flip side of that. It's anything that is not number-driven. This can be feedback directly from actual or potential users. Sometimes you have to approach a theoretical user base, sometimes you have access to your actual user base, and both can bring a lot of value. This can include a variety of methods to understand attitudes, beliefs, motivations. Those are the sorts of things you measure with qualitative testing. User interviews are usually what comes immediately to mind for qualitative testing, but those can often be expensive or time-intensive. Qualitative testing also includes things like surveys, diary studies, contextual studies, and even unmoderated user testing can bring a lot of great insights.

[00:07:09] Aaron: That is what they are. And it really is one plus the other equals success. We might know quantitatively that 50% of people are dropping out of the funnel at this moment in the process. But the why is really key for us to understand: is this a moment that is only happening here, where the intersection of a user's motivation plus the appearance of this page is causing this interaction? Or is there something more generalizable here that we could be applying elsewhere throughout the application? The what is informative, but the why helps us see opportunities for potentially reusing this learning elsewhere. That's why I suggest these two things really work in tandem. It's when you combine the two that you're really able to say this learning can be used elsewhere throughout the application, and we might be able to reduce drop-off in future funnels because we've done this.

[00:08:12] TJ: Great. And just to reinforce from the business perspective: having both the what and the why behind behavior is just the most efficient and effective way of driving decisions and reducing cost at the end of the day. It's also great to provide engineers and other stakeholders both perspectives, quant and qual. Sometimes stakeholders lean towards one or the other depending on their perspective. So having both is, in my opinion, not just the most effective way to see both sides of things from the product perspective; stakeholders also reveal different perspectives from utilizing both.

Case study: slimmer buttons that lost

[00:08:59] Charlotte: All right, jumping into some examples. How do these two methodologies combine at an actual company, in a real use case? Here is an example from a client that we were recently on. This is something we actually did and implemented and saw great success with. Like everything in design, it's a circle; it is continually feeding into itself and improving itself. Starting with some context on this client: they were a large corporation, so at this higher level, and they hosted personal blogs and placed their own e-commerce CTAs on those blogs, related to the content.

[00:09:38] Starting here, we got some qualitative user feedback from those blog owners that informed us of an incremental[?] improvement opportunity. Their complaint was that the e-commerce CTAs that were related to the blog content, but not owned by them, were taking up too much room. They were big, they were bulky, they could be slimmed down. With that feedback we jumped into a quantitative A/B test. We designed several different versions of slimmer, more modern buttons, and we wanted to see how those would feed into our primary KPI, which was increasing conversion rate from these CTAs. Hopefully, by slimming them down and meeting that user request, we would improve those conversion rates and modernize the platform.

[00:10:25] What happened, though, with our new modern buttons: we did not expect the results we got. We had the what of the data, and the winner was one that we didn't really think would win, and we were like, but why is that happening? That's when we jumped back into qualitative testing. We had a big why question that we needed answered. For this we did traditional user interviews, where we presented potential users with our designs, asked them which one they preferred, and learned why they preferred the one that the A/B testing had indicated. What was great is that we were getting the same feedback from the qualitative as we were in the quantitative: that one of the variants was the big preference, the overall winner. Everyone preferred it, for the reasons that it was easier to read and it was bigger. Then we were able to take those learnings and feed them into future quantitative A/B tests, where we were able to apply the button style that won across the platform and further increase conversion rate there.

[00:11:36] I might have given it away a little bit, but here is our test. On the left we had the current design, the big buttons. You can see they're blocking a lot of the content. And on the right we have all of the modern, slimmer buttons that follow a lot of existing concepts and existing practices, and they took up a lot less room on the page. Like I mentioned, people preferred the bigger buttons here. We saw drastic decreases in not only conversion rate but also click rate of the modern buttons, which surprised us all. They were smaller, they were more modern; we thought people would trust them more. And then when we did those user interviews, like I mentioned, people really liked the big buttons even though they were clunky, because they were easier to see and easier to read. I'll let TJ speak to how that then translated to some of the business initiatives following this test.

[00:12:35] TJ: Great. Yeah, understanding why this test failed was really, really helpful from the business perspective, because if we didn't know why, we probably would have leaned into the icon approach. One of the compromises in the design being shown is that you have big buttons that are easy to read and easy to press, or you have icons, and we really thought icons would make a difference in the performance. Knowing why, knowing that the icons weren't helpful and that the users preferred the larger buttons, prevented us from investing a lot more resources into testing different icons, or icons in different places. We were able to cease that train of effort and switch gears onto something else. And by the way, since this user testing was conducted, several times, every couple of months, someone will have an idea like: oh, why don't we put icons on these buttons? And we're able to confidently say: yeah, great idea, we tested this, and we now have information from the users that this is just not as good as the larger buttons with the larger text.

Case study: machine learning the recruiters would not use

[00:13:53] Aaron: I'm going to share another case study here. Hopefully my storytelling is top notch, because this deals with a lot of personal information, so I can't show the actual study, but I'll try to weave an interesting narrative through it. This study began with a tech impetus, which I think many of us are probably experiencing right now. AI and machine learning have become more affordable and more accessible. The company had a hypothesis that if we had people spending less time manually researching people's backgrounds, and more time sitting with them face to face, we could increase our conversion rate. We already knew that face-to-face interactions were critical for this company, so we wanted to get our people in front of more people faster, and we hypothesized that machine learning could help us get there. Maybe ChatGPT could.

[00:14:52] We began with general user interviews, getting to know people's day-to-day. One of the interesting insights that came out of that was one of the ultimate low points when they're doing research: if I find out the wrong thing about TJ, and then I say it to him in our face-to-face interaction, that is ultimately a very embarrassing interaction for me, the person who wanted to meet with you. These people were very risk-averse to that. They were very nervous about it, because it would likely result in not being able to convert this potential lead.

[00:15:29] But there was a tech impetus, so we took this insight and we were still going to forge ahead with a pilot and see: could we provide some value and help these people spend less time doing research? We leveraged our machine learning tool to begin surfacing data points collected on people, and we asked these people to continue doing their job as they always did. They would research this person, and we tracked quantitatively how often they were correcting data points that we were suggesting. And the answer was actually quite often. It's really interesting: from a machine learning model, you say 80, 90%, that is stellar, statistically, that we're able to get this information. But for these face-to-face interactions, people's threshold was much higher. They needed to be 99% sure, and when the computer said it was sure, that didn't align with their level of certainty.

[00:16:32] The next step in this was that we took the information from the quantitative study and set up a qualitative diary study. We would check in with folks. They had to do this every day as part of their job, and we would just send them a study over Slack and say: hey, did you use the tool today? How was the data? We used a metaphor, we gave it a cute name: would you hire the tool to join your team today? And people would say yes. They were enthusiastic. It was doing a good enough job; they perceived that they would bring it onto the team.

[00:17:10] This led to further tech investment, since we saw some positive data here. We rolled out the tool, an enhanced, integrated tool for surfacing these findings, and we tracked how people would use it. We gave them two options: you could create a blank slate and start gathering the data individually, or you could start from this template that you said you would hire, that we thought was really good, that was quantitatively pretty accurate from a machine learning standpoint. And over time, people just trended back to the manual.

[00:17:50] To get back to what TJ was saying about the value of these cycles: we saved the business an immense amount of money and investment in this tool, because we found that these recruiters would not use it. The ultimate decider was: I want to ensure that I have a positive interaction and that I'm confident in the information I'm bringing to that face-to-face engagement. The business invested in other ways to increase the amount of face-to-face connections rather than using machine learning. Maybe ChatGPT is not going to solve the world, I don't know. But it's an interesting case study of how even a negative finding through this process can positively affect your bottom line.

Qualitative outputs and research ops

[00:18:43] Charlotte: Awesome, thanks Aaron. All right, now we've shown you what they are, the value they hold and how they're implemented at the actual organization level. Hopefully we've convinced everyone that doing both qualitative and quantitative testing is incredibly important and useful to any organization. But how can you actually implement this at your organization? No matter what level of user testing you have, especially as a designer, I like to say you can always do more. We're going to look at some of the outputs and implementations of these types of testing.

[00:19:17] Starting with qualitative. This one is always hard if you don't have existing practices. How do you measure qualitative outputs? You have all these different opinions and motivations and abstract information; how do you turn those into actionable, targetable insights? It's not that easy all the time. Oftentimes you have to adapt what you want to know to what you have learned. Some tools we have found really helpful: Miro. You can see we're using Miro here. We took a bunch of insights from some user interviews and affinity-bucketed them, and then by grouping them based on some of our key questions, we were able to identify patterns across all of these different opinions. That led us to: okay, for this behavior we were trying to observe, people are more often finding the thing we want them to find, so that's been successful. Or sometimes we'll see: hey, people are oftentimes getting distracted by this thing over here, and we don't want that to happen. How could we go fix that? That becomes a next step to put into your backlog.

[00:20:25] It's a lot of insights. You can see positives, negatives, neutrals. Sometimes people just really don't care about something you're really excited about; with negatives, sometimes people are performing an action you really don't want them to do. But how you format this really depends on what you're hoping to learn from it, and then identifying the patterns that answer the questions and problems you're hoping to solve.

[00:20:51] Another example of that is down there. I just used a Google Sheet to do some heat mapping. We were trying to target really specific behaviors, and seeing those colors helped us identify which behaviors were being performed and which weren't. That is a great support, especially if you have an extensive A/B testing system: a great way to look at the why, or at where people aren't clicking on something, which we don't have numbers on. That is some of the literal outputs you can get there.

[00:21:22] Aaron: I think Myron put it very succinctly: think about the tool, and think about it early. Qualitative research ops, I've found, is honestly one of the largest barriers to adoption, I think, in organizations that aren't already using qualitative research or aren't sharing it effectively. If you haven't used a tool like Dovetail, it's one I've used and I really enjoy; there are other options out there. What it allows you to do is share the entirety of a qualitative study, such as an interview, a video or a usability test, and you're able to tag it and share it with the larger organization.

[00:22:03] I fundamentally think that if we are more effective at sharing our qualitative studies, it enables the entire organization to more readily see the value of the work we're doing. You've conducted a study with someone, you've interviewed a person, you've derived your insights, but that might actually be helpful for the problem the sales team is trying to solve down the line. Again, qualitative research ops are a really effective way to help your organization reuse and see the value of these sorts of studies.

Quantitative outputs and research ops

[00:22:49] Charlotte: All right, and then jumping to quantitative, the other side of the coin. Quantitative outputs are a lot easier to understand and visualize, because they are numbers. We are definitely not experts in data analysis. There are some great teams out there that specialize and focus on that, and some great tools out there that focus on that. But here's a great example of some quantitative output that we didn't just get the winner from; we were able to gain several other insights from these numbers. If you're looking at this and confused by some of the language here, what we were testing was literally tree images. We were trying to increase the donations for a tree-planting organization, and we were seeing what sort of image on the purchase page would help motivate people to contribute to this organization most.

[00:23:45] Here we can see the Sapling B image is the one people preferred the most, so that's our clear winner. But again, we don't know why, and that's maybe something we could go and do some qualitative testing or impression testing on. And at the very bottom, Sapling D underperformed relative to all the other options. What we learned from that is we've got a winner, we've got a loser, and there are no specific trends. We have saplings at the top and bottom of the chart. We were testing baby trees, people planting trees, families with trees, and we didn't see any preference in terms of what the content was, which told us that it was the specific image we had found that was driving the most conversion.

[00:24:36] This is just one example of quantitative output. You can get very in the weeds with all of these numbers, but it's a lot easier to measure and point to a winner, and bring it to your stakeholders saying: here's what we learned, here are the numbers. This is what a quantitative output looks like.

[00:24:53] Aaron: Yeah. And quantitative research ops, I would say, are far more mature at the average organization. Business intelligence tools: if you're working in an organization that has products, you might have even the most basic thing like Google Analytics, which is a solid foundation for this. This space is a little more mature, but there are other ways to scale your quantitative research operations, like Mixpanel, where you can configure funnels and create simplified charts so people can see insights in real time. Scaling this is a little bit easier just because it is more mature. And if you're not even using Google Analytics, I bet your organization has Excel spreadsheets with some sort of data. Don't discount those as foundations for your quantitative research, to maybe help kick off your next qualitative study.

Starting where you are

[00:25:49] Charlotte: All right, we've seen some ops advice, but how can you start doing this wherever your organization is at? Maybe you have extensive user testing and you're just motivated to go and do more. Maybe you haven't had the chance to do any user testing at your organization. How can you do this? Step one is always: what questions are you trying to answer? Identifying what hypotheses you are testing will inform what methods you use. Are you trying to find the why of something, the what of something, a little bit of both? That'll really tell you what sort of testing you need to start looking into.

[00:26:28] Then, to narrow in on that, know what your budget is, for both time and cost. We've mentioned a few times that stuff like user interviews can be expensive, both in time investment and cost investment, depending on how much access to your user base you have. But there are definitely ways to do user testing, both qualitatively and quantitatively, quickly and affordably. One of my favorite tools to use is UserBob. It's often overlooked; not many people know about it, but it is incredibly affordable, and they have a great range of testers to give you feedback. It's unmoderated, so you just submit a prototype, let it out into the world, and then people go through, try to perform tasks and answer questions. You could use that to do even one-minute impression testing, comparing version A to version B and seeing which one people perform a task on faster or better, or whatever you happen to be looking for.

[00:27:24] Aaron: Yeah. And shortly, the best advice for building up an ops program around this is to start where you are. Find out what analytics you might already have, and use that to inspire a single conversation with a customer. Start a couple, build up the habit, make it a more recurring thing. It's okay to start small. It doesn't have to be a fully scaled program with a beautiful Dovetail and Mixpanel to be effective in helping to improve your product. Start where you are.

[00:27:53] TJ: Great. And definitely find an ally in product. As I've stated many times, it's definitely worth the investment to get both the qual and the quant supporting decisions; there's a strong ROI. Another note is to not give up. Sometimes, if you just want to start talking with one or two customers, it could take a while to get hold of customers, because of some organizational barriers that might exist if it's not something that you regularly do.

[00:28:22] Charlotte: Yeah, design should not exist in a vacuum. TJ's helped us get a lot of user testing done: performing the testing, getting buy-in for it from stakeholders, as well as getting those insights out into the larger organization. This is a full team effort, for sure. It looks like we don't have much time for questions, but we're going to throw this up here. Feel free to reach out to Crafted, or to any of us on LinkedIn or other social media. We're always happy to do a virtual coffee, or an in-person coffee if you happen to be in the Denver area. We love chatting with everyone and getting to know your stories at your companies. Thank you all.

Q&A

[00:29:03] Rory: Thank you very much. That was really good. I think it tied in perfectly with Myron's talk as well, so it was a nice connection there. Thank you for sharing.

[00:29:14] Charlotte: Appreciate it. Thank you for having us.

[00:29:16] Rory: Unfortunately we are out of time for questions, but I did have one, and I'm going to post it into the Slack channel. It comes back to something you mentioned a few times: we've already tested this and this is the answer. But sometimes technology changes, so I'm sure now, Aaron, if you try it with ChatGPT it will work this time. I guess the question I'll post in there is: when is the right time to revisit something, and how do you know when your research has gone stale?

[00:29:51] Aaron: I'm on UXDX, so I'd be glad to. I'll answer it on the channel there.

[00:29:56] Rory: Great. All right, well, thank you very much. If anybody has any questions, just send them in and we'll copy those onto the Slack group as well, and I'll invite Aaron and the rest of the team to make sure they answer them for you. But thanks once again for sharing.

[00:30:13] Charlotte: Thank you all for having us. Cheers.

Speakers

TJ Bowen

TJ Bowen

Product Manager

Aaron Knoll

Aaron Knoll

Product Designer