The UX Scorecard: Transforming Research into Measurable Impact Across the Org

May 1411:55 am – 12:30 pmStage: Main StageTalk
Slides

Checking session availability…

Hang tight while we load the latest updates.

How do we ensure that UX research truly drives decision-making and delights customers? Join Amit and Marina for an eye-opening exploration of our journey to creating a unified, impactful measurement framework that transformed our entire organization.

  • How to create a measurement culture that prioritizes customer impact and product quality in every launch
  • Building and applying an Experience Scorecard to assess and improve product usefulness, satisfaction, and overall performance across stages, from concept to live release
  • Setting clear launch criteria with a "passing score" to ensure readiness and alignment with customer expectations
  • How to integrate measurement into OKRs, allowing product, design, and research teams to track progress and continuously enhance user experience
  • Empowering teams through enablement and education, so they can use measurement tools to make data-informed decisions that drive product success

The UX Scorecard: Transforming Research into Measurable Impact Across the Org

Amit Sathe, Marina Lin at UXDX USA. Video: https://youtu.be/JqlcMi0U-yI

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Warm-up poll and the pain points

[00:00:12] Amit: All right, last talk before lunch. How many researchers do we have in the room? Good. Looks like some honorary researchers, incidental researchers. Great. As we get started, we're going to do a warm-up activity. There's a poll coming up on the screen. Pull out your phones and answer the question. There are a couple more questions coming, but we'll leave the first question for about 10 seconds and then we'll move to the next question. Of course we have to survey you guys. We're researchers. Still more responses coming in. That's good. Can we move on to the next question, please?

[00:00:54] All right. Welcome. I'm Amit. This is Marina. Over the next half hour we are going to talk to you about the problem we solved, the process we followed, the rationale for why we did it, and some learnings from having successfully implemented this program.

[00:01:11] The first pain point, which you saw in the poll earlier, I just wanted to reiterate. Our experience over the years has been that researchers are brought in as a tiebreaker to validate competing opinions. That in itself isn't a bad thing; resolving differing opinions with data is a good thing. But what it actually creates is busy work for researchers, while at the same time they could be working further upstream on generative research, or trying to uncover user needs. So in a team that's already strapped for research resources, if all the researchers are doing is evaluative research as a tiebreak, maybe as a confirmation bias medium for the team, that's not the best use of researchers. That's a pain point we faced several times throughout our careers.

[00:01:58] Marina: Another common pain point that we faced, and some of you have as well, judging by the polls, is that the product has already been developed, and then the users aren't adopting, and so research is called in after the fact to figure out why. There's a term for this that we started using at work, called expensive guessing, meaning that the product has been released purely on a hypothesis or a guess, and the end result is expensive for everyone from a time and effort perspective.

[00:02:31] Amit: And the last one. Another pain point that we've noticed through the years is that qualitative research doesn't always get a seat at the table. There were several different sessions yesterday, and I think the one earlier today as well, where research having a seat at the table has come up, and it's been varying over the years. Sometimes, depending on organizational buy-in and organizational philosophy, research is tightly integrated, but more often than not it's not. There could be several reasons why that happens. We were looking to break that cycle.

What gets measured gets improved

[00:03:04] Marina: While we were dealing with all of those pain points, we also had a VP of design at the time that asked us, are we doing what it takes to delight our customers? In other words, what gets measured gets improved. So if we aren't measuring, what are we improving? And so we set out to change this.

[00:03:26] Amit: Before we even took any actions, we did some introspection. We did some reflection on our internal processes, our internal situation, and we realized that there were several broken links in our product development life cycle process. I'll go over some. We had difficulty quantifying success metrics, even determining or defining what success means. Our researchers were running studies using disjointed methodologies. There was less consistency in how those findings and insights were being triangulated and socialized with the rest of the team. And we didn't have standardized benchmarking mechanisms across the team. Depending on the team and the researcher, they were using different methods, different tactics, and collecting different measures. As a result, each researcher was left to fend for themselves when it came to study selection and method selection, as well as how the data was collected and how it was socialized. So we set out to change all of that.

[00:04:24] Marina: What we did at the time was form a small task group consisting of a couple of user experience researchers and our resident researcher data scientist. We met once a week and tried to figure out how we were going to get to a culture of measurement.

Choosing what to measure: usability, usefulness, satisfaction

[00:04:44] Amit: The first thing in front of us was deciding what to measure. Since we all had collective years of research experience, we first started with the frameworks and approaches that we knew about. We explored the HEART framework from Google. We explored SUM, the Single Usability Metric. If anyone's done research in the early 2000s or the first part of the 2010s, you might have been familiar with SUM, the Single Usability Metric. Personally, I've spent many months and years trying to get that institutionalized, trying to get it integrated into the process. We also looked at the QX score, the quality of experience score, that we knew about. But either some of these frameworks were too broad, or sometimes they were too narrow, too restrictive. We were looking for something that was just right.

[00:05:31] So we decided to create our own framework, and we decided that there were three things that were critical to measure. First of all was usability: is the product that we are designing usable? The second one was usefulness: is it solving an actual need? Is it actually solving a problem that the customers and the users have? And lastly, satisfaction: is the overall experience satisfactory? So we decided to formulate our framework, as well as our scorecard measurement program, on these three dimensions.

When to measure: leading, lagging, observed and self-reported

[00:06:03] Marina: The other thing that we were trying to decide is when to take the measurement. Where we ultimately landed was drawing boundaries around a product or a feature, so that the experience is a standalone product or feature that makes sense to a user in a usability test, for example. And it becomes a snapshot in time for that feature or product.

[00:06:27] We also gave thought to how we consider leading versus lagging metrics. If you think of the release date as a point in time, anything leading up to that would be when you're designing concepts and testing prototypes. Those are your leading indicators, and this is where you're measuring task success, and Kano, for example, and CSAT satisfaction. And then once the release date happens and the product goes live, from that point on everything is a look-back metric. That is a lagging indicator, and at that point you can look at the conversion funnel and revenue, for example.

[00:07:15] Amit: In addition to leading and lagging metrics, we also wanted to give special attention to observed and self-reported metrics. It seems like there are several researchers in the room; there's familiarity with research. How many of us have noticed that what users say during a usability test and what they do is sometimes not the same? More often than not, they completely contradict each other. Most recently, we were in a usability test where the participant was clearly struggling to complete the task, but later, when we offered them a self-reported score, they rated it the highest possible score. So clearly we wanted our overall metric program to not skew towards just self-reported or just observed. We wanted a mix of both. That's the reason that, when we did our entire inventory and selection of metrics, we used a combination of both observed and self-reported measures.

Busting the first myth: you can quantify qualitative

[00:08:08] Marina: During the time that we were having all of these conversations, as you can imagine, we had a lot of discussions and internal debates, and we had myths that we had to contend with and ultimately dispel. We'll show some of these to you today, starting with "you can't quantify qualitative measures." What we realized ultimately is that you can and should, and you probably already do. For example, you can code behaviors such as positive and negative reactions. You can then do affinity mapping to see how patterns emerge. And best of all, you don't need a large sample size to do that.

[00:08:46] Amit: Speaking of sample sizes, the research organization that we lead uses primarily mixed methods, but trending more towards qualitative. So a majority of our studies are small sample studies. But even within small sample studies, we figured out the right mix of observed and self-reported metrics. We heard about time on task and task success earlier; those would be a couple of examples that we use as part of our observed measures. As part of self-reported, we have Kano, we have the SEQ, or Single Ease Question, and we also use CSAT. So regardless of whether it's a large sample or a small sample, the framework that we are going to share works in both situations, but it works especially well in small sample situations.

One overall score and the 70% pass mark

[00:09:33] Marina: What we ultimately decided on was having one overall score, depicted as a percentage, and that score comprises three buckets: usability, usefulness and satisfaction. Each one of these buckets is weighted equally. The researchers have the flexibility to choose whichever metric from each dimension, and we encourage a mix of observed and self-reported. Once those metrics are collected, our data scientist created a simple normalized equation for us, so that we can create that overall score and average out percentages and Likert scales from a study, for example. We'll show you how that translates to an experience scorecard in just a bit.

[00:10:20] Amit: Our passing score is 70%, and in a few minutes we'll get to why it is 70%. But here we also wanted to map the numerical score to an expectation scale. Going back to Marina's comment, our VP challenged us with what it takes to delight our customers. Our hypothesis was that if we can meet and exceed user expectations in enterprise software, in business software, that could translate to delight. That's why we mapped our success scores to a meeting or exceeding expectations scale as well. A score from 72 to 89 would be deemed to pass, or meet expectations. Anything 90 and above would be exceeding expectations, and anything 70 and below would be failing to meet expectations.

[00:11:11] Marina: Now, why 70%? We did an audit of all the metrics that we were using at the time, and you can see the screenshot of that audit behind me. What we started to see is that right around the 70% mark is where a lot of these metrics either meet expectations, or pass, or are acceptable, for example. This was true of SUS, and this was true of CES and CSAT [?], et cetera. So we settled on 70% for passing.

[00:11:42] Amit: Next, we also chose a threshold for the minimum number of participants, or minimum number of observations and data points, required to generate a score, and we arrived at a minimum threshold of six. Again, we are largely talking about qualitative studies, and generally five to eight users has long been considered the industry standard or norm. We arrived at six in consultation with our resident data scientist, and some of the secret sauce that they applied to the overall calculations. But as far as our scorecard is concerned, six is the minimum threshold. I'm repeating the words minimum threshold. That doesn't mean that we aim to have six participants in each and every study, but a minimum of six participants is required to generate a score using our calculator.

The process and the scorecard itself

[00:12:37] Marina: Ultimately, the process that we have now gives our researchers flexibility. They choose the methods and metrics that make sense for their study, and they have a buffet of metrics that they can choose from. As long as there is one from each dimension, the overall score will calculate. They then conduct a study and populate their metrics in a calculator, which we set up in a spreadsheet. Once they have their overall score, they can port that to a visual scorecard, which is depicted as a slide in a deck that they then present to their stakeholders during a readout.

[00:13:16] Amit: We've said scorecard enough times that everybody might be wondering what it actually looks like. Here is a sample scorecard. I'm just going to quickly walk through the different parts of it. We start with the project name, or the overall name of the experience that we are measuring, followed by a timestamp for when that measurement was taken. Above that you see three pills: concept, prototype, live code. In a few minutes we'll cover what those mean. Over on the extreme right, we have the single overall score, the combined score, and a grade for whether the experience passed or failed.

[00:13:50] And then we have the three component scores, or the three dimension scores. For the purposes of this example, we can see that this particular experience is passing when it comes to usability and satisfaction, but usefulness not as much. That gives the team an area to investigate further in the next iteration. If usability and satisfaction are okay, maybe usefulness is something that needs to be explored. Over on the left at the bottom, we have a section to add qualitative insights, because again, we are talking about primarily qualitative studies, and it's not purely about the number. Sure, the number plays a part in it, but it's about the qualitative insights, so we have a placeholder area to add those. And then over on the right, we have a KPI placeholder. That's where our partnership with our product data science team and any other quant teams within DocuSign comes in: any studies, and any results coming from those explorations, are referenced over here as well.

Piloting the scorecard on the workflow tool

[00:14:46] Marina: At this point we'd like to share an example of what the scorecard looks like in action. While we were figuring all of this out, there was a project going on at the time in DocuSign called the workflow tool, and this was to automate the agreement process. This particular project had all the symptoms that we discussed in the beginning. The research was being done, but the insights just were not landing. The release date was quickly approaching, and so we decided to use this as a pilot to test out the experience scorecard.

[00:15:21] The first time that we ran a test and then calculated the score, it was an eye-opening experience for everyone involved. This time the team had a shared language to understand that this was not good. It was kind of like being in school and getting a test back with a lot of red ink on the paper. We were able to have real conversations about our prospects with real users, were this to go live in its current shape. One of the outcomes of that was that the beta phase of the project was extended. The product team and above, everybody agreed. Usability fixes were prioritized, and the team agreed to hold back, make those fixes, and then retest again.

Busting the second myth: usability testing is expensive

[00:16:07] Amit: That brings us to our next myth, which is that usability testing is expensive and time-consuming. I have heard it several times in my career. I'm sure most of the researchers here have heard it. But look at how usability testing has evolved over the years. I started doing this about 22, 23 years ago, and back then usability testing was done in fixed labs. If your company didn't have its own lab, you had to rent facilities and conduct usability testing there. Fast forward a few years, and my usability lab became a laptop, and I literally traveled the world using that laptop, and I could meet customers where they were. Fast forward even further, and now all we need is a computer browser, an internet connection and some kind of meeting software. So I personally and strongly disagree: it's not as expensive as it used to be. It's not as time-consuming as it used to be.

[00:16:58] But there's another way to look at how we break that myth. If you want to take a picture, I would take a picture of this slide, because again, I feel really strongly about this. The most expensive usability test that you'll ever do is the one that you don't do. We are all researchers, or we work very closely with researchers; we know the importance of it. But if you have product, or other cross-functional team members, trying to take a shortcut, trying to release a product by skipping usability testing, you can tell them, sure, you can save some money today, but you'll probably end up spending tenfold after the product is released. Because guess what? There is a high likelihood that the product is going to be released with bugs. So there's going to be bug fixing, there are going to be increased support calls, and there's going to be a next iteration that will have to be rolled out pretty soon.

[00:17:47] If you add up all those collective costs, you're not really saving that much by skipping the usability testing. It would be a case of penny wise, pound foolish, or penny wise, dollar foolish, seeing that we are in the US. That's the overall myth, and I think we've done a strong job at busting that myth, not only in our organization, but you can use this in your organizations as well.

Measuring at concept, prototype and live code

[00:18:11] Marina: In that spirit, we proposed three points at which we take those measurements: during the concept phase, during the prototype phase, and then once the product launches and goes into live code. This guidance actually mirrors our product life cycle, so it was easy to work in. During the concept phase we are only measuring usefulness and satisfaction. There is no usability dimension yet, because there is no interaction. Then as it moves to the prototype phase and gets ready for the beta phase, we take another measurement. And then once all the recommendations are hopefully implemented, the fixes are included, and the product goes live, we take the final measurement in the live code, and that becomes our benchmark for any future improvements going forward.

[00:19:08] By testing during these moments, we're able to measure and track over time, and we're able to influence go-to-market decisions within our company and most of our teams. And best of all, we don't need large samples in order to drive these decisions forward.

Busting the third myth: a culture of measurement in a large organization

[00:19:27] Amit: That brings us to our third and final myth: creating a culture of measurement is hard in a large organization. Again, this is something that we highly disagree with. DocuSign is a fairly large organization. Our product portfolio is broad as well as deep. If we're able to make it work in an organization the size of DocuSign, I'm sure everybody here, regardless of whether it's a small, large or anywhere-in-between type of organization, has enough learning from here that you can incorporate in your own organizations. What you really need is a process, some tools, and a whole lot of enablement, adoption and various kinds of training. It came up yesterday in several sessions, and it came up earlier this morning as well, but it can be done. The size of the organization doesn't really matter.

[00:20:14] Just to show some highlights about where we are today with our scorecard program: our scorecard program has been integrated into our product development release methodology. A scorecard, or a measurement, is required before a product exits the beta phase. So that's the big win. Scorecards are discussed at the highest levels in our organization, right up to the C-suite, our CPO, our CEO. It has come up in several conversations even at that level. And finally, all teams, not just UX, are enabled to conduct measurements. Our research team is limited in scope, but we've enabled the non-research teams as well to conduct some of these measurements.

[00:20:56] Marina: We got here by doing a few things. One, after we spent some time trying to figure out what the methodologies and the calculations would be, we formed two teams of researchers that tested these calculations with their own real projects. Then they gave us feedback, and so we tweaked, revised, and landed on the scorecard that we shared with you all today. And then we launched a pilot with the larger team, so that our cross-functional partners were able to see it in action as well, and that is where the workflow project came in.

[00:21:32] Speaking of that workflow project, not all was lost. A few months after the usability fixes were prioritized, we ran another study, and this time the score markedly improved. Usability and satisfaction were passing, and in fact usability was now exceeding expectations. Usefulness was still below pass, but the overall score was at 70, so we considered it a pass, and the product was able to launch successfully. But the good thing is that we now had a shared understanding of where we need to focus and what we need to do to increase usefulness for our customers, and that is what the team agreed to work on.

[00:22:17] Amit: As an example of this demonstrated success, we moved forward with adoption and rollout of the experience scorecard across multiple product teams and multiple product releases. Our teams now have a common language. We've enabled and trained the rest of the teams to a level where they understand the difference between leading and lagging metrics, between observed and self-reported, when a concept scorecard comes in, when a prototype scorecard comes in, when a live release scorecard comes in. All that has happened given the successful implementation of the case study for the workflow project, and now we are increasing rollout and adoption to even broader teams at DocuSign.

[00:22:58] Marina: Now we hear things in meetings like, "We can't release yet with this usability score," and "What are we doing to increase usefulness for our customers?" and "When is the updated scorecard coming?" It's really good to hear, because this is the language that comes from our product partners, all the way up to the C-suite, which is pretty impressive in a large company.

Considerations to take back to your teams

[00:23:18] Amit: Before we wrap up, we would like to leave you with some considerations that you can take back to each of your teams and organizations. The first one is that you're already measuring experience one way or the other, so start with those. You don't have to wait several months or several iterations to come up with the perfect mix of metrics and measures. Start with what you already have and build your way up to incorporating other metrics over time. But more importantly, start with what you have.

[00:23:48] As we've spoken about earlier, consider a mix of observed and self-reported. Consider a mix of leading and lagging, so your measurements are not skewed towards one or the other of the dimensions that we discussed. The third one is to align your measurements with the key phases in your product life cycle process, because ultimately the research team, or the team that's bringing the experience scorecard, needs to align and adapt itself to the product development methodology. The product development methodology is what it is; they are not going to change. So if we can figure out a way to align and blend our techniques into the overall PDLC or SDLC, that's a sure recipe for success.

[00:24:27] And then finally, enable your cross-functional partners to measure experiences. Researchers don't have to take it all on themselves. I know democratize is probably a slightly controversial term. We use the term enablement rather than democratization. The concept is the same, though: researchers don't have to do it all by themselves. If you include sufficient guardrails, process and tools, other teams can do a little bit of lightweight usability testing by themselves as well. And that's it. Thank you.

Q&A

[00:25:07] Host: Firstly, congratulations on your tremendous success. What you just described is really hard to do. I've been a researcher for 15 years, and you can have the scorecard, but actually getting it embedded in the product team, and those overheards, is phenomenal. So congratulations. Let's chat a little bit more about what you just discussed. We've got some great questions coming in. How did you manage difficult conversations with product owners that had low scorecards? Tell us a little bit about some of those conversations. You created the scorecard, and it possibly led to a low or unsatisfying number. What did those conversations sound like?

[00:25:44] Marina: It was a little hard, I would admit. I was on that team, and I would say it wasn't a surprise, because even though some might bury their heads in the sand, you always have an inkling that something's not right. A couple of researchers went through this same project and were having the same challenges. They were doing studies, and the concepts and the prototypes weren't landing. It was just too difficult for the persona that we were going after. So it didn't come out of nowhere, is one thing. And then we did a roadshow ahead of time, saying that we're rolling out this scorecard thing, and it was nebulous at first. But then when it applied to their project, all of a sudden it was like, wait, this is a real thing. It's all red. It's not usable. People are not thinking it's useful for their work. So that was a real red flag, and at that point, people got on board.

[00:26:49] Host: Absolutely. And you have these three measures, right? Usability, usefulness, satisfaction. Did you weight them all the same? How did you choose that weighting, and why?

[00:26:58] Amit: We did decide to weight them equally there. We wanted to start bringing the first iteration of the scorecard out, so we decided to go with weighting them equally. We wanted to avoid analysis paralysis, where we are spending multiple cycles just trying to debate whether usability is weighted higher than usefulness or satisfaction. And actually, I do strongly believe that they all play an equal role in it. Think about it: if usability is high but usefulness and satisfaction are low, it's going to result in several customer support calls, or errors while using the product. Conversely, if usefulness is high but the other two are low, it's going to result in low adoption, or in loss of revenue. So I think all of these dimensions play an equal role, and that's the reason why we weighted them equally. Being researchers, the natural tendency would have been to go into deep thinking mode and figure out whether one is greater than the other. We said, okay, let's try to weight them all equally and see where it goes, and it turned out to be the right decision.

[00:28:06] Marina: And I will just add one thing. During the concept phase, it's actually not weighted equally. It is 70% towards usefulness and 30% satisfaction. That's because we are trying to bake in a little bit of that strategic piece, because it's still early on. As the designers are giving us concepts to test, we can also convey whether or not our customers will find the product useful, and so it's weighted a little bit higher.

[00:28:36] Host: Absolutely. And you've got some curious audience members about the percentages and how they're calculated specifically. What scale are you using, and is it a top-box method or something else?

[00:28:44] Marina: Each metric is on a scale of 1 to 5, and it could actually be 1 to 7, depending on the preference of the researcher. But if they are using, for example, SUS or CES or CSAT or whatever, each one is calculated with its own metric. It just gets normalized when it gets put in the spreadsheet. And if it's a percentage, like the percentage of task success, it is already a percentage point, so it just gets put into the calculator.

[00:29:15] Amit: Yeah. And the normalization was something that our data scientist, who's an expert in all these topics, was able to figure out: the way to normalize different scales into a single scale.

[00:29:25] Host: Another great partnership for you. Okay, wonderful. That's it for our questions. Thank you both so much. That was wonderful.

Speakers

Amit Sathe

Amit Sathe

User Research Director

Marina Lin

Marina Lin

Sr UX Research Manager