Translating Experience Value to Revenue: An Autonomous Vehicle Case Study

28 Apr16:30 – 17:00 UTCStage: Main StageTalk

Checking session availability…

Hang tight while we load the latest updates.

We all believe intuitively that customer experience translates to revenue, but the value of experience can be difficult to quantify. So what if you could trace the link from experience to revenue in a straight line?
I'll discuss how a measure called Behavioral Intention (BI) can rigorously quantify the impact of experience on business goals, using a case study of on-road data reflecting the experience of autonomous vehicle (AV) behavior changes--before these changes were implemented or even defined. BI is simple to deploy and has a validated relationship to future behavior across a variety of domains, so you, too, can speak confidently to the revenue impact of experience changes--when those changes are still in the future!

Translating Experience Value to Revenue: An Autonomous Vehicle Case Study

Krysta Chauncey at UXDX Community: The Experience Advantage: From Journey Health to Revenue. Video: https://youtu.be/bGEqF3OYxow

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Why quantify experience before you build it

[00:00:08] I'm really excited to be here and to talk about this work that I did about a year ago. It's very exciting to be able to translate experience to revenue directly, since as many of us know, the changes in experience themselves can be hard to argue for at the highest level of decision-making.

[00:00:27] Just a word about myself. My background is experimental psychology. I've been doing technology development research for over 10 years with finance, transit, autonomy, AI, all kinds of unexpected places. I often get involved when we need to measure how people are going to use technology we haven't made yet. I have recently started my own solo shop, Action Potential LLC, focusing on dragon mapping, which is what I call identifying behavioral derailers in products and systems, as well as their impact, and on my favorite, measuring how people use technology that hasn't been made yet, which is one of the things that is involved in today's case study.

[00:01:16] The question of why we should value experience is not a new one. The ROI of UX work has been calculated, validated, generally, specifically, by a wide variety of sources over the last two decades. But most of that is done either at an organizational scale or retroactively. So the question is, how can you quantify the value of experience, or the difference in value between experiences, in a specific product before you've made it, so you can know if you're making the right thing or not? The specific question of what is this experience work worth? If we give people this experience, then what happens in terms of adoption, revenue, metrics that C-suite executives live in?

Robo taxis and the remote assistance pause

[00:02:09] I'm going to tell you about a case study where we did exactly that. Robo taxis are an interesting business model for autonomous vehicles from an experience standpoint because you have to make that decision continuously. Private vehicle ownership, you buy one once every seven years and then you're done. With robo taxis, you have to decide that you're going to take a robo taxi every time you use one. And so the technical function of autonomous vehicles is irrelevant if the experience is so unpleasant that passengers never want to do that again. So the question is, how much disruption does it take to change the experience enough that it moves the needle on whether a passenger is likely to return or not?

[00:02:54] As another note of context, all robo taxi companies have a remote human vehicle assistance function. So this means when the car gets itself into a situation that it can't get out of, or encounters a situation that it can't get out of, there are off-site humans who can be called upon by the car to make suggestions about what they should do. These operators, like I said, they make suggestions. They're not remote driving the car.

[00:03:25] The example that seems to help the most is thinking about being in the passenger seat, the front passenger seat, for a teenage driver. With a teenage driver, you might say, there is a construction zone coming up. The cop who is directing traffic in front of it is going to expect you to slow down before you get to that tree and then stop about at that tree. So make sure that you are slowing down and ready to stop. Then you would never lean over and take the wheel. That would be a terrible idea. So that's typically how these remote human agents support the vehicles.

[00:04:04] However, this does take a second. These agents are not monitoring the cars synchronously. That wouldn't scale. That wouldn't make sense. So it does take them some amount of time to orient themselves in the situation, start to understand what the car is confronting, what the car's capabilities are in this moment, and what the best solution will be. During that pause, the car is typically sitting still in traffic.

[00:04:29] So the question is how much of that becomes a problem? That's a safety question, of course, but you could absolutely imagine a degree of traffic stop that would not be a safety problem, or that would be adequately handled by safety engineering, that would nonetheless be enough of an experience problem that it creates repeat customer problems, creates organizational-level adoption problems.

[00:05:01] So the question I was confronted with was how long can these in-traffic stops last? How does that change the experience? And how do we measure this before we make implementation choices that are very expensive to reverse?

Designing a simulation of an experience that doesn't exist yet

[00:05:18] As I talk about this case study, periodically I will stop and talk about how you might generalize this process to a different problem, and what kind of processes we went through in order to arrive at each stage of planning and analysis.

[00:05:35] So the questions that are important to think about in terms of being able to measure behavior before you make the technology. The first one is what's critical in forming the real experience, not the simulated experience, not your best effort, but the real experience that you're trying to cast ahead to. What's really important there? In this case, obviously, an unexplained stop in an autonomous vehicle in real traffic is very different from a stop in a human-driven car in traffic, or even an autonomous vehicle on a track with other cars. There is a level of uncertainty, a level of communication that is assumed, that is just really different.

[00:06:17] So then, breaking that down a little more, we asked what elements that create this experience are really essential, and secondarily, what artificial elements of our simulation might undermine the experience and how can we minimize them. The elements in this case that were essential were a moving car without a visible driver. You can't have a visible human in the driver's seat. On a populated public road: if it's a controlled circumstance, or there aren't enough independent, semi-unpredictable actors around, it's just not going to be the same.

[00:06:53] In terms of the artificial elements that undermine the experience, one is a researcher presence. Sitting in the back of an autonomous vehicle in traffic by yourself is very different from sitting in the back of an autonomous vehicle in traffic with a friendly researcher sitting right next to you who's presumably not freaking out. And any clear control, either remote or in car, of the car's motion. It really needs to look and act like an autonomous vehicle.

[00:07:21] We had an added wrinkle in this case because at the time the AV stack was not ready for on-road fault injection. So that is failing in specific ways on demand. That's not normally something you would ask an AV to do.

Hiding a driver, guiding participants, controlling the environment

[00:07:36] So we had to hide a driver. This meant we had some additional work to do around simulating what's important. I already said it's really important that participants believe that they're in an autonomously operated vehicle. And so we had to use a hidden human to get the vehicle behavior that we wanted to produce instead of automation. We also needed to guide participant behavior, so directing them to do what we want without necessarily telling why. And we needed to control the environment to minimize any issues that might come up that would interrupt our simulation.

[00:08:18] So in this case, we asked the drivers to be as still and quiet as possible and avoid really human acts like waving other drivers out. AVs might do that at some point, but they're definitely not doing it now. So it would be a giveaway. On their own initiative, the drivers, who were vehicle operators, imitated the driving manner of an autonomous vehicle very effectively. They had spent hundreds and hundreds of hours in the car when it was being autonomously driven. They knew exactly what kind of things it did and they were able to reproduce that in their own driving, to a slightly comical extent.

[00:08:53] In terms of guiding participant behavior, just on a physical level. If the participants walked around the front of the car looking at it, there was no way we could hide the driver in the front seat from them. So we had to give them very specific directions about what would happen next and what we wanted them to do. As an example, when we were done with the data collection and heading back to the office, we gave participants very specific directions on how to get in and out of the car. We'll be back at the office soon. When the car stops there, you can open the door and head towards the door you entered the building through. I'll get out, stop the camera, and be right with you. When given instructions like that, everybody did what they were asked. When not given instructions like that, people had a tendency to wander around and look at the car.

[00:09:44] We also had to be careful to control the environment not to interrupt our simulation. So what we did was we modified a seat partition, like in a taxi, to be maximally obstructive between the passenger and the driver specifically, while maintaining the driver's field of view and lines of sight. So here you can see what that looks like. We put some reflective material on the back and we added material to the interstices to make sure that they couldn't see through. And it worked pretty well.

The study: stops of 30, 60, 90 and 130 seconds

[00:10:19] So what did we actually do once we had figured out how to simulate this? We put passengers in a robo taxi and simulated in-traffic stops that lasted 30, 60, 90 or 130 seconds, on routes that had pre-existing municipal approval for AVs doing testing operations. We used an in-car survey to measure stress, estimated pause length, and intention to take a robo taxi again at each stop. So you can see the intention question at the bottom here circled in green. The next time I have a chance, I intend to use an autonomous vehicle again. And this was our proxy measure for future behavior that I'll talk more about in a minute.

[00:11:04] We also had a standardized communication acknowledging the fact that the car had stopped. So after 20 seconds there was an automated announcement. It seems like we're not moving. Please hold tight while we look into it. And at specific intervals there would be later announcements. There was never any chance we were going to allow people to encounter these stops without any communication in real life. That's not under consideration, for obvious reasons. So this was our best estimate at the kind of communication that we would be likely to implement.

Choosing a proxy metric: intention versus expectation

[00:11:40] So heading back to the generalization question, how might you do this in a different setting? The choice of a proxy metric is really important in this case. You want to cast forward to revenue. You want to cast forward to behavior. We're not in the future yet and you haven't made the technology. So you've got to choose a proxy metric. There are a bunch of options, but I'm going to focus on two specifically.

[00:12:06] The first one is the one we used, behavioral intention. The next time I have a chance, I intend to use an autonomous vehicle again. And one that we considered but didn't end up using in this case was behavioral expectation, which is a simpler statement of, I expect to use an autonomous vehicle again. And in both cases we asked, or would have asked, participants to tell us how much they agreed with that statement using the visual analog scale you see here.

[00:12:37] There are both some fairly heavy-hitting differences between these measures as well as some relatively subtle differences. First of all, behavioral intention does not include a participant's estimation of whether they will have a chance to do whatever this thing is. It assumes that they will have a chance. Very specifically tells them to assume that they will have this chance. Its strongest link with future behavior is when the action in question is close in time to this response, or the behavior is familiar.

[00:13:10] Behavioral expectation, on the other hand, is modulated. It does include a participant's estimation of whether they'll have the chance to do this thing. It's just expectation. It doesn't say anything about assuming or not assuming a chance. Its strongest link with future behavior is when the action in question is farther in time from the response, on the scale of days or weeks, and the behavior is unfamiliar.

[00:13:36] So in this case we chose behavioral intention because, both then and now, if we had included people's estimate of whether or not they're going to have a chance to use an autonomous vehicle, that would wash out everything else. You either probably do have a chance to use these, if you live in one of the places that they've been deployed, or you almost certainly don't. We don't need to ask people that. We don't need to know about their estimate of that. What we need to know is, if they did, would they? So despite the fact that the strongest link with future behavior is when the action is close in time and the behavior is familiar, which is not a great fit for this situation, we used behavioral intention because behavioral expectation would just absolutely wash out everything we needed to see and just tell us the extent to which people thought they would be able to do this again, which doesn't help us.

[00:14:32] There is some pretty good literature on this. These are some of the papers that I used to understand this, both at the time and since then. I particularly want to direct you to the role of time in self-prediction of behavior. This doesn't sound super relevant or interesting to this use case, or the use of this method in an industrial setting, but it is a very good paper distinguishing these measures. In addition, I want to call out Jessica Fishman, who provided the validation data set that I'll talk more about later.

What the data said: we had more time than we thought

[00:15:05] Okay, so we chose this measure, we used this measure, we collected this data. What did it actually tell us? I was pretty shocked by these results. We had a lot more time than we thought. An unexplained 30-second stop in traffic in an autonomous vehicle. That sounds substantial, alarming, non-trivial. I thought there was a chance that we were already measuring everything too long. However, a stop of 30 seconds at any point in the ride did not increase overall stress substantially. You can see here participants' stress ratings by the pause length. And it did increase variance. So it increased stress in people specifically prone to that for a variety of reasons, but at a population level, it didn't move the most frequent range.

[00:16:02] In addition, with the communication level that we tested, behavioral intention did not decrease until between 90 and 130 seconds. This was truly shocking to me because this sounds long. But this is very consistent with what we saw both at an individual level in people's reported data, but also on a qualitative level with their responses and reactions in the car, qualitatively, just in terms of what they were saying, how they were reacting. It was not uncommon that they didn't notice the 30-second stop. And they didn't seem super bothered by it.

[00:16:44] This was very surprising to me. I think it's possible that this is related to a short-term window in terms of where autonomous vehicles sit in the technology acceptance process culturally. However, at this point, this was fine.

From behavioral intention to returning riders to revenue

[00:17:04] So then how do we get from that to adoption, to revenue, to things that non-experience-minded people care about immediately? We used the validation data set from Jessica Fishman that I mentioned. It comes from public health and it linked the self-reported level of behavioral intention on a specific action with the likelihood of taking that action in the future. So we used this to extrapolate the difference in returning riders and revenue for each delay length.

[00:17:39] The data that we collected fundamentally was a distribution of behavioral intention after each stop length. The validation data set told us that the bottom third of behavioral intention has about a 25% likelihood of taking the action. The middle third of behavioral intention has about a 55% likelihood, and the top third has about a 65% likelihood.

[00:18:04] So given that we know how many people were in each of those thirds at each delay length, we can just do the math to get to how many people out of a thousand who had this experience would come back. Which let us say that out of a thousand people having the 30-second stop experience, we might expect 608 to come back. And out of a thousand people having the 130-second experience, we might expect 595 to come back. The odd thing about this data set is that really the top line was that even though a 30-second stop in traffic and a 130-second stop in traffic sound very different, in terms of returning riders, not that different. So we can use other factors to choose between how we prioritize work and how we allow the car to behave.

[00:19:02] So in that context, we used internal financial assumptions to get to a revenue projection from this. Here I've used notional assumptions to get to how much returning revenue we might expect based on the different experiences. And that part of using internal financial assumptions, so that we were speaking the same language and in the same numeric range as everybody in strategy and corporate was already used to functioning in, that was really important.

Use the numbers your organization already runs on

[00:19:40] So the question here, if you want to generalize something like this, is what metrics of success does your organization use internally to make top-level decisions? If they use revenue, you've got to use revenue. If they use lifetime customer value, you've got to use lifetime customer value. It does not matter a huge amount which metric you do, but it's got to be the same that people are already used to functioning in.

[00:20:08] I was personally surprised by how much running around and question asking it required to get hold of these numbers. I had to talk to multiple people in operations, multiple people in finance, multiple people in strategy to get all of the numbers that we needed that they were using for revenue planning, so that we could operate within this framework. So in this case we needed per-ride revenue. We needed assumed rides per week, and we needed new customers per week, to get to a sensible projection that would give people a sense of how much of a difference in revenue we were looking at.

[00:20:49] And this allowed us to start from a position of speaking the same language, instead of having to first argue for the value of experience by itself and then explain how the experiences are different.

[00:21:06] So I appreciate everyone's time and coming. If this sounds interesting to you, or if this sounds like a problem that you're facing, I'd love to hear from you, and I'd love to take questions now as well.

Q&A

[00:21:18] Host: Excellent. Thank you very much, Krysta. That was really an interesting case study, because it's one of those things that I didn't even think about with autonomous cars, but it's definitely very important, I guess, particularly if you're sitting there, nothing's happening for a few minutes, you might start getting a bit freaked out.

[00:21:37] Krysta: Yep.

[00:21:39] Host: Brilliant. So just a reminder, you can write in on whichever platform you're watching this on if you have any questions for Krysta, because I know I certainly do. If you write in those questions, I'll put those questions to Krysta. The first thing that I want to ask is, one of the things that was just so counterintuitive to me, the wait time just didn't really seem to have that much of an impact. How did you ask the people if they spotted the driver or not?

[00:22:14] Krysta: Yes. So I was obnoxiously smug about this, perhaps. But none of them spotted the driver. All of them believed that the car was autonomously driven. We asked them about this in a funnel way, where we first asked them what did you notice about the front seat of the car, and what did you notice about the manner of driving of the car, and did you notice anything moving in the front seat, and gave them a little more breadcrumb of information to try and get them to say, yes, I saw the human in the front seat.

[00:22:48] Krysta: We had one person spot the front seat driver because we drove past a building that was all mirrored glass and they saw the driver in the front seat. It's like, okay, you caught us. There's nothing we could do about that. So we excluded that person from the results and moved on. But none of the people in this data saw the driver. They all thought that the car was autonomously driven.

[00:23:16] Host: Okay. And the other one, I know it sounds like I'm trying to pick holes, but I'm just trying to understand, because you said that you gave a warning at 20 seconds, or some kind of, rather than a warning, just an update, and then periodically after that. Is that what actually happens in the cars when they stop?

[00:23:37] Krysta: So in terms of the real experience, no, but it is not the plan for that to be true ever. Developing that communication was actually happening in parallel to this study, because it was clear that most of the time people don't notice, but if they do notice, they care, so we need to communicate about it. So at no point was anyone in product or engineering considering just not saying anything about these delays. So we used the current thinking on what would be useful, that was already in development, for how this communication should look.

[00:24:20] Host: And because you mentioned you have to really use this as your trigger metric. Did you try running anything without communication? Because I guess that would be, if the cars are out there in public right now without the notifications.

[00:24:35] Krysta: Yeah. So we did run a couple of sessions without any communication. And what we found is that people largely didn't notice these stops.

[00:24:47] Host: I guess it's no different to just stopping in traffic or at the lights.

[00:24:51] Krysta: Yeah, exactly. And if you think about stopping in a taxi in a highly dense neighborhood like the Seaport in Boston or the Strip in Vegas, if a car stops, you're not going to assume that everything is broken forever. You're just going to assume that something is happening that you're not looking at. And in a taxi, because of that seat partition, you can't see that well. You definitely can't see stop lights. You're probably not paying attention. So it was surprising and very interesting how quickly people were willing to check out of monitoring their environment.

[00:25:32] Host: Yeah. Particularly when you're trusting a machine that...

[00:25:36] Krysta: Yep.

[00:25:37] Host: Hopefully the airbags are good quality in the cars as well. What's the, in life, I don't know if you have the statistics or you're able to share them, but what are the frequency of issues and the average duration of issue?

[00:25:55] Krysta: So the reason that we chose 30, 60, 90 and 130 seconds as a duration was based on the distribution of actual resolution times that we were seeing in the field at the moment. 30 seconds was the ideal case. Everything works immediately, solves quickly, nothing is really a problem. Just boom, boom, boom, there we go. Great. 30 seconds, done. The majority of stops were between 60 and 90 seconds. And the 130-second was the mean plus two and a half standard deviations of the overall data point. So we used the fastest ideal case, two bookends of the most common case, and something that was an outlier.

[00:26:48] Host: Okay. I'm unfortunately in one of those cities where I don't get to test these out. So I'm also thinking, because I'm from Ireland, and if you go outside of the big cities, a lot of the roads, particularly rural roads, you can have grass down in the middle of the road. If you meet another car, somebody has to reverse. So I'm just trying to figure out how they're going to deal with that.

[00:27:12] Krysta: Wouldn't you love to see an autonomous vehicle try and navigate that?

[00:27:15] Host: Yeah, particularly because you kind of have to go off the road to let the other...

[00:27:19] Krysta: Yeah, you have to go into a layby.

[00:27:23] Host: I guess the last question I have is just around research methodology. Often when you ask people to gauge their future behavior, it's just not very predictable, as in what people say they will do and then what they actually do are often quite distinct from each other. How did you try to factor that in?

[00:27:48] Krysta: Yeah. So the way we factored that in is basically in the validation data set, because you notice that the top third of behavioral intentions still had only a 65% likelihood of doing that. And so that's not that much higher than the middle third, which is 55% likelihood. So that validation, right there, bears out what you're saying, because the top third was people who said, yeah, definitely I'm absolutely going to do this, put in the top third of the scale. And they didn't even do it three-quarters of the time. So that was factored into our expectations, that 65% of the people who strongly said that they would do something would actually do it.

[00:28:36] Host: So you actually tracked their future behavior to see if they did in fact rebook.

[00:28:42] Krysta: So that's what the validation data set was. It came from a different domain, but it actually collected behavioral intention and then collected whether or not people had actually done the thing that they said they would do.