Designing, Implementing, And Analysing Product Experiments

07 Oct19:00 – 19:30 UTCTalk
Slides

Checking session availability…

Hang tight while we load the latest updates.

People say that it's difficult to come up with ideas worth testing, that implementing a randomised controlled experiment is complicated, and that you need math skills to analyse the results. These are dirty filthy lies!
In this talk, Cian will give you the knowledge and tools required to quickly experiment, allowing you to build a demonstrably more useful product.

  • How Hubspot do quick-iteration product experimentation
  • How to bring these methods to your organisation

Designing, Implementing, And Analysing Product Experiments

Cian Mac Mahon at UXDX Europe. Video: https://youtu.be/TuNpBdl-s7c

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Three eras of using data at HubSpot

[00:00:01] Hi, my name is Cian and I'm a technical lead at HubSpot, where I work to build delightful onboarding experiences for our free users as part of our growth team. I've been working on growth projects at HubSpot for over five years, quite a long time, including building the very first version of our free marketing tools, and also building infrastructure for running and analyzing our product experiments.

[00:00:19] During all of that time, our understanding and usage of data to drive product decisions has changed pretty dramatically. I like to think about it in terms of three eras. Back in 2014, when we were first starting to build out our free marketing and sales tools, we built what we felt was right. We have fantastic product managers and designers, so this pretty frequently worked out, but every so often it backfired a little. We got a feature wrong, and the experience didn't quite match with what our users were looking for, and we didn't really know why.

[00:00:48] So by 2017 we understood the importance of data in the product development cycle. So we started tracking absolutely everything. Every single user interaction from every single user was collected, catalogued and then never looked at again. Sometimes we'd build charts which would show us how usage was changing over time, but this was generally either unscientific or retrospective, rather than specific and real time.

[00:01:15] Now in 2020 we have a process which lets us make accurate predictions about how the changes we're making are going to impact user behavior. We mix qualitative, which is interviews, and quantitative data, which is user tracking I guess, with product experimentation, to deeply understand how to best serve our users. Most of our experiments boil down to something pretty simple: we just give slightly different experiences to different groups or cohorts of our users, and then we see what happens. We've used this process to drive meaningful increases in revenue. And because it's been so successful, we often get asked to consult with internal and external teams on how they too can build experiments. That's what we're going to talk about today.

[00:01:56] Before we get going, though, because I've done all this consultation and have spoken to a few places about building experiments, I frequently hear a few problems that people perceive. I hear that it's difficult to come up with ideas that are worth experimenting with. I also hear that all the work involved is just far too much: I know my users, I know what I should build, I'm just going to go ahead and do that. And I've also heard that all of this data requires a deep understanding of maths, and personally I can't remember how to do long division, so how can I be expected to do product experimentation?

[00:02:28] Well, I'm just going to quickly rearrange the heading of this slide and add a few more letters, because these are all dirty, filthy lies. Coming up with ideas is really easy, and I'm going to tell you about a framework that I use to do that. You're also probably already doing the required research, and the engineering work can be as simple as a few lines, and tools do exist which do the analysis for you. I'm going to tell you about some of those as well. By the end of this talk you're going to walk away with the knowledge and tools required to build a demonstrably more useful product.

[00:02:55] We are going to talk about how we come up with ideas, how we implement ideas as experiments, and how we analyze the results. But because I wanted to give you practical examples, I'm going to tell you about two experiments my team has run in the last year. And in order to do that it's going to be useful for you to understand what HubSpot is, what we do, so I'm going to give you a quick primer. If you use HubSpot, if you've heard of us, feel free to tune out for about 30 seconds, but then please do come back.

[00:03:19] At HubSpot we build software to help companies grow in better, more sustainable ways. Our suite includes free CRM tools, sales tools, marketing tools, customer service tools, as well as even more powerful premium versions of each. Well, every HubSpot account comes with all of this functionality built in, and we know that everyone who comes to us, who uses our software, is using it to at least initially perform one specific job. So our onboarding teams split across user missions. The mission of my team is to help teams who are focused on marketing decide that HubSpot is a good solution for them, and if so, how they can make use of the rest of the suite as well.

[00:03:57] On the onboarding teams we measure our success in terms of user success. Are the users generating new leads? Are they making more sales? Are they inviting other team members who are also making use of these features? We find that all of these are solid predictors of whether a user will grow with our software and eventually upgrade. So now you know a little bit about HubSpot, it's going to help out with the experimentation section a little bit later.

Guide rails for picking experiment ideas

[00:04:19] But first, let's talk about how we come up with ideas. It can often seem really difficult to brainstorm ideas for experiments. There are so many things that we could change in our apps that it can just be hard to choose one or two. But, like everything, picking experiment ideas gets easier if you introduce some guide rails. So when we're trying to come up with experiments on our onboarding teams, we set ourselves the following limitations.

[00:04:38] Firstly, we say that experiments obviously should always be done with the aim of building a better user experience. I say obviously here because our success is driven by user success. Secondly, we should use experiments to answer risky questions, especially where using existing data to predict the result is a little bit difficult. And thirdly, experiments should reach statistical significance. We'll talk about that later, so don't be worrying your head about it just yet.

[00:05:08] But from the top, let's get this one out of the way first. At HubSpot, the customer comes first. It's our first and most important product value, and as I mentioned a minute ago, when a user is successful in growing their business, they're more likely to grow within our suite of tools. So when we're designing a product experiment, each one needs to have the aim of improving the user experience, whether that means how they understand our tools, how they interact with them, or how they meet their own goals. Easy, it's obvious.

[00:05:32] Secondly, let's talk about answering risky questions. Before we start coming up with ideas, it is really helpful to realize and understand that only a subset of ideas are worth experimenting on. Running experiments is not free. It takes product manager time, design time, definitely engineering time. At HubSpot we also involve product analysts, we involve content designers, and we also of course run the risk of shipping a sub-optimal experience for a subgroup of our users. So let's be sure that we are limiting what we experiment on to the most impactful and high value of ideas.

[00:06:04] I've built a chart here. On the X axis I've put how risky an idea is. Here I define high risk as likely to impact either the users or our ability to do business. On the Y axis we have the amount of existing data we have, and that's user interviews, usage analytics and so on. In the lower left hand corner, if an idea is low risk and we're pretty sure we know what's going on, we're just wasting our time doing a product experiment here. That's just build that feature, give it to our users. And in the lower right hand corner, if an idea is high risk but we're fairly certain we know what's going to happen, as well, just kind of do it, I think. And keeping my eyes on the chart, there is also some case here for a product experiment too.

[00:06:44] In the top left hand corner we have a low risk idea and we have no real idea what's going to happen. We should probably consider doing some user research here before we make any moves, and we could also just try, depending on the risk. But in the top right hand corner, we have an idea which is high risk and we have no idea what's going to happen. That is where an experiment will be really useful. So with experiments we should answer high risk questions where we don't have a huge amount of existing data.

A gut feeling for statistical significance

[00:07:14] Finally, the big scary one, the tough one: reaching statistical significance. You'll be really glad to know that I'm not going to get into any maths here. We just don't have the time. Also, to be honest, I'm not certain I'd be able to explain it. Although, we're computer people, so we can let computers do the hard work for us. But let's at least get a bit of a gut feeling as to what statistical significance is.

[00:07:31] Let's imagine that our app has a user base of 10,000 people and we want to perform an experiment on them. So we pick 1,000 of our users and we show them something different, leaving 9,000 people in our control group, our control cohort. We get a result and we are happy, but how can we be sure that we have gotten the full story? Maybe those randomly chosen one thousand users are special in some way. Maybe a greater than average number of them signed up on a weekend, for example, or maybe they use slower computers than the rest of our users, which can take a whole lot of time to load stuff.

[00:08:07] Let's take another example. Let's imagine that I suspect that one of my coins is more likely to land heads than tails when flipped. So I decided to run an experiment. I've got a control coin which I definitely trust, and when I flip it 10 times it comes up heads four times, which seems roughly right, real life, not always 50/50. I'm going to flip my suspicious coin, and it comes up heads seven times out of 10, which to me feels a little bit like something is up.

[00:08:31] But if I run the same experiment for 100 times, I see that the coins are landing on heads roughly the same number of times that they land on tails. We see there isn't actually a thing, and we see this because we took a larger sample size, we checked more times. The technical term for the probability of our results being reflective of reality is P value. At HubSpot we generally require a P value of less than 0.05, which is a less than 5% chance that the results we're seeing in our experiment don't reflect reality, don't reflect what would happen if we show this experience to everyone.

[00:09:07] By the way, the P value of this coin flip task, before you do the maths, is 0.23, which means there's a 23% chance here that my hypothesis, that my suspicious coin, is actually a little bit dodgy. This is far too high a chance for me to make a call either way. I think, though, what is good now is I suspect you have a good feel, or a good instinct, as to what statistical significance is. But as I mentioned, I'm not going to show you how to calculate it. There are many tools online which do this for you for free. If you just Google statistical significance calculator you'll find some really, really simple ones, and I'm also, near the end of the talk, going to make a few recommendations based on my experience of how we do it.

The Experiment Afternoon

[00:09:47] But how do we come up with experiment ideas? We've never done this before, I hear you ask. Something we do on my team is called an Experiment Afternoon. It's called that because it's in the afternoon. I was jet lagged at the time, I wasn't feeling creative, and it just kind of stuck. So all the engineers on the team lock ourselves in a room for three hours, and we'd have two or three roughly written Experiment Documents — I'm going to talk about those later — which were then perfected over the course of the next week or so. Think of this as much a learning experience in how to do experimentation as a way to improve life for your users.

[00:10:14] This is the layout. First, we choose what we would like to achieve. It's really important here that we pick two or three metrics to move, hopefully metrics that are going to move fast. So, for example, daily active users rather than monthly active users, or clicks on a specific link rather than retention over a period of weeks. The reason that we want to pick a metric which is going to move fast is it closes the feedback loop really, really fast, considering we're using this kind of as a learning opportunity as well as an experimentation generation exercise.

[00:10:45] Once you've chosen the metrics, let's talk about them. We need to be really, really scrappy, and we spend 15 minutes just chatting about each of these metrics and coming up with hypotheses that we might have as to how we're going to improve them. We write every single idea we have for each metric, no matter how off the wall, onto a whiteboard, or just write it down somewhere. I guess we're not in person anymore, so put it in a Google Doc, I don't know.

[00:11:07] Third, it's time to start tearing those ideas apart. Be really, really ruthless, hunting every piece of evidence that you can find that each of these suggestions is wrong. We spend about 30 minutes doing this. You can use both qualitative and quantitative data here. So if you have user interviews, go get those. If you have usage tracking, dig into that, see what you can find. Finally, it's time to write some Experiment Docs.

Experiment one: importing contacts

[00:11:32] But before we do that, I want to tell you about a quick experiment that we ran at HubSpot. We gave it the not very catchy yet pretty descriptive name Importing Contacts to HubSpot. Personally I prefer descriptive over catchy, but some people don't, whatever. This isn't catchy, but I know exactly what it is.

[00:11:52] A bit of background context. Once we send users out of our onboarding experience that we've built, into the big bad world of the free HubSpot tools, most of the ways that they can quickly see value require them to have at least some data in our CRM. When we were doing user research as to what our users were looking to do first, the vast majority of them told us the first thing they wanted to do, the very first thing they wanted to do, was to import contacts. To give a direct quote: I want you to organize my life, I want you to get my clients in first, that's the main thing.

[00:12:19] We've got this checklist of things that we'll walk you through when you sign up with HubSpot in order to get started, and one of the first things on that checklist is importing contacts. Despite 95% of our interviewees telling us that they want to do it, only 10% actually go through with it. So there's a disconnect here. People want to import contacts but they just don't. Importing contacts is like going to the gym. It's not a fun thing to do, but you need to do it to see results. The second you read the words import contacts on a button, you immediately just tune out, you start checking Twitter, honestly. I'm pretty sorry I brought it up, it's probably lost your attention already, and if it hasn't, I'm about to definitely lose it, because I'm about to throw up a chart.

[00:12:59] Because we do a great job of tracking user interactions, we can see exactly where in the import flow we are losing users. There are eight steps to importing contacts, and the first ones, I'm not going to lie, are tough. I'll show you them later. But what I want you to know is that from the point where the user starts the task, they show us the intent that they want to do something, to the point that we ask them to upload a file, we lose like 67%. 67% of these people who told us they wanted to do this, they're done. And honestly, these steps aren't all that hard, and once we get them to that step, by the way, they pretty much just finish the process, even though there are way more difficult steps ahead of them.

The Experiment Doc

[00:13:40] So that's the problem we were tackling. We found the problem, we quantified it, it's time to build an experiment. And the first part of building an experiment is writing an Experiment Doc. An Experiment Doc serves as the history of your experiments, and also helps keep you honest about your methods and how you were going to measure success. Also, just as humans, we can tend not to want to remember the things that don't work. But if what you're doing is innovative and risky, most experiments you run are going to show your hypothesis to be flawed. Like 90% of them are not going to pan out. So Experiment Docs help us remember and tell others about previous experiments that we run.

[00:14:15] Not going to lie, you're about to see a whole load of text, because we are using an experiment that I just told you about as an example. I've also messed around with the numbers a little bit for legal reasons. If you want to revisit the slides, I believe they're going to be online later. But an Experiment Doc lays out five things. Firstly it lays out, what is your hypothesis? Next it says, what evidence do you have to support it? Next we say, what change are you going to be making? How are we going to measure success with this experiment? And then, what is the minimum improvement you'd accept in order to consider the experiment a success? And that's where statistical significance comes in.

[00:14:48] We're going to attack each of these one by one, but taking it from the top, we need to say exactly what our hypothesis is. This lays out why we are running the experiment, and gives the reader a quick overview as to what the experiment is about. Somebody outside your team should be able to read your hypothesis and have a general idea of what is being tested. You should also, by the way, here call out any assumptions you are making, just to get them on the table. We know that importing contacts is the best way to get started using HubSpot, and we know that users want to get their contacts into the system, but we believe that it is too hard, which we believe discourages users from doing so. We've laid out our hypothesis: what we know and what we believe.

[00:15:29] Next up is trying to explain why we're running the experiment. What is the evidence that our hypothesis is correct? Gut feeling here does not count. We do a lot of user interviews. Kelly, who's a fantastic researcher on my team, spoke to over a dozen users as we were coming up with this hypothesis, and almost every single one of them said the first thing they wanted to do was import contacts into the system. But only one out of every five users who start importing finish. We also know that if we can just get users to the bit where they choose a file, they're pretty likely to finish the import. So what we need to do is get them to there, the point where they have selected a file to import.

[00:16:10] Next, time to lay out the change that we are going to be making. At HubSpot we do this in two ways: we use both language and a table, just so everything is super clear. And to use our example here again, we will split our signup cohort from July 1st to July 14th in three. Control see the existing onboarding experience. Variant one will see a quick import banner, I'm going to show you that in a sec. And variant two will have the new quick import checklist item. In language, and you can see in the table, they have pretty much the same: control, no change, 33%; variant one gets the in-your-face variant, 33%; and again, variant two, they have this integrated variant.

[00:16:46] What does that actually look like? In the control, this is our getting started checklist. In our control we can see the regular import your contacts task in the red square to the far left, and this is the one I just showed you that eight step chart for, by the way, with the 67% drop. Variant one[?] is this huge big in-your-face banner which tells the benefits of importing your contacts, as well as offering this new super speedy import flow. Variant two linked to that same super speedy import flow, but back in the old interface.

What the import flows look like

[00:17:12] So what actually are these import flows, this thing that we're trying to change, that we're trying to encourage users to complete? I'm going to show it to you. It is important that I say right now that this is actually a good flow. It's long, it's a bit convoluted, but if you follow it through step by step we know that you're going to get what you need to do done, and it also handles a lot more cases than just importing contacts, which kind of explains the extra complication around it.

[00:17:42] But first of all, we say, do we want to import contacts to use now, or a list of people who've opted out of your marketing? Secondly we say, do you want to import one file or multiple files, each from the other? Next we say, how many types of things? Just contacts, or just one type of thing, or maybe you're importing contacts, companies, deals, tickets, all these things all at once. I, by the way, chose just contacts. We then say, okay, just one type of thing. I said contacts here.

[00:18:12] Now, upload. Here is the point where, effectively, we've lost 67% of them. By the time they get here, 67% of them are gone. If we can just get them to choose that file, we're home free. Next we ask them to match the columns in their import to how HubSpot thinks of contact properties. We have some machine learning stuff in here which makes this really easy. And next we ask them to give the import a name.

[00:18:39] So this, as I said, is not fantastic, but it is meticulous. If somebody follows this flow they will get their stuff imported into HubSpot. But we know we're asking them to import contacts, we know we're asking them to import a single file of contacts, so we can take those first five steps and just kind of mash them into one, which is why you can see here, this is our quick contact import flow. Step one is just right into that import contacts section, like the select file section. We've also put a bit of educational content there. And step two, match columns to properties. Step three, give it a name. So this is a quick import flow.

Success metrics and minimum improvement

[00:19:18] Next step: what is your success metric? Holding yourself accountable to a specific metric makes it really easy to decide if an experiment is a success or not. It also helps us avoid confirmation bias, this perfectly human tendency to look for successes where they don't exist. Say, sure, the number of upgrades are the same, but Android users in Canada went up by 3%, success? No, that's probably just chance. We will judge success on the percentage of users who complete an import from the Getting Started checklist. It's pretty fine, it's pretty straight, direct to the point.

[00:19:52] Finally, as I mentioned before, running experiments is not risk-free. It takes product manager time, design time and also engineering time. Of course, as I said, we at HubSpot also involve product analysts, content designers, and we run the risk of shipping a sub-optimal experience for a subgroup of our users. So when we're writing our doc, let's make sure that the improvement we're chasing is actually worth all of this time that we're putting into it.

[00:20:18] Our minimum improvement is a 25% improvement on the existing metric. What I mean when I say 25% improvement is like a 25% improvement on four is five. Cool. In order to detect a minimum improvement of 25% with 95% significance, that's that P value of 0.05 I mentioned earlier, we need to run our experiment for 14 days. We calculated this using one of those calculators. We have an internal one, but the ones you'll find on Google are just as good, and by feeding in the number of users you expect to be passing into your experiment every single day, you'll find out how long you need to run it for.

[00:20:51] So that is a huge amount of data, I am really sorry. Let's run through it one last time. An Experiment Doc says what our hypothesis is. It tells us what evidence you have to support it. It says what the change you'll be making is. We define what our success metric is. And finally we state what the minimum improvement we'd accept is, in order to consider the experiment a success. I have made available a kind of toned down version of what an Experiment Doc template looks like at HubSpot. You can head to that bit.ly link[?] or scan that QR code. I hope it's really useful for you and your organization when you decide to run experiments.

Building the experiment

[00:21:31] We've made it this far, we've identified a risky experiment to run, we've written our Experiment Document, and now it's time to build the experiment itself. Other than building the experience, there are three important things you have to do when you're building an experiment. The first is assigning users into cohorts, or saying, what's this user going to see? Cohort assignment should be random, stateless and functional. That's a whole load of engineering speak for saying you should be able to run the code which assigns the user a cohort, given a user ID, and it should always return the same cohort for that user and experiment every single time.

[00:22:06] You also, by the way, have to consider here if you want to assign based on user, or if you're like us and you have multiple users, whether you want to assign the account to the cohort rather than just the individual user. Once you've assigned a user a cohort, you need to let your analytics tool know, by sending an event with the assigned cohort and experiment as a property, so you can build your charts later. And finally, once your user performs the actions you're trying to impact, send that to the analytics tool too. If you're already on board with user interaction tracking you're probably already doing this bit, which is nice.

[00:22:39] Thankfully there's a whole load of tools and libraries to help you out in doing this. One that we like is PlanOut, which is built by Facebook and imported by us into JavaScript. They built it in PHP. It's a complete experimentation library which at its most basic will functionally sort users into cohorts. It has hooks that can be used for logging experimental exposures, so you know which users are seeing what things, and allows you to namespace experiments so that you don't show competing experiences to a single user. Say you're running an experiment on your signup flow, or two experiments I should say, on your signup flow, using PlanOut you can say a user should only ever be in one of these two things.

[00:23:18] It's also been ported to multiple languages, and we wrote a blog post about how to use it with React. Once again I have a QR code there, so you can check out that blog post. There's also a link at the end of that QR code, on the blog post, to where you can get PlanOut itself.

[00:23:35] Okay. At HubSpot we're at the point where we're running many, many experiments at once, and managing them all through code was just becoming a bit of a chore, to be honest. So we spun up an infrastructure team as part of growth, and they built us an experiment service which is still based in the background on PlanOut, but which is fully controllable by all members of our team through a UI. Now the front end on our app makes a request that says, tell me what experiment cohorts this user is in, and then handles the response and shows the correct experience to the user.

[00:24:04] Everything else is configured through the UI by either engineers or designers or PMs, and the UI can also do automatic analysis and builds a library of our Experiment Docs, which is absolutely fantastic. In the last, say, 14 days across the onboarding group we switched on, I think, 11 experiments. So not having to manage all of those through code is a lifesaver. It makes life really, really easy in comparison to what we were doing before.

Analyzing the results

[00:24:34] So we have waited our experiment duration, as per our statistical significance calculator, and we're going to want to analyze the data. If you're using an analytics platform like Mixpanel or Amplitude, this should be pretty easy, they pretty much just do the heavy lifting for you. Google Analytics does a pretty decent job too. I've never used it, and as far as I know you're going to need to manually calculate if you've hit statistical significance. There might be a plugin to help you there.

[00:24:57] So let's take a look at the results for the experiment that we were playing with, our import flow. Before I show you the chart, though, I should note that HubSpot is a public company, so I have not been able to share the actual results here, but I'll show you something which are almost the actual results. Two weeks have passed, and looking at our results we can see variant one, which is the big banner, has absolutely crushed the control and variant two, beating the control by 27%. 27%, we've got 27% better chance of the user uploading their contacts. That's pretty good.

[00:25:33] That's statistically significant, by the way, with our P value of 0.05. So we are happy that the results reflect reality and we have a winner. We've productized this variant, variant one, for several of our user groups, and are slowly rolling it out across our other cohorts as we learn more about how our users interact with it. So that is how we do experiments at HubSpot.

Experiment two: telling users the tools are free

[00:25:55] I want to give you another example, tell you about another experiment, one that I find kind of fun, kind of silly, so you might find interesting. When somebody signs up for a free HubSpot account, we throw them into a sandbox-like environment that my team built to help them understand how HubSpot works. We bring them on this whirlwind tour through the products, tailored specifically towards their needs based on what they told us when they signed up. And in this tour a user might, for example, create and send a marketing email, they might dig into our reporting feature, and they may play around with our ad tracking[?] tools. They do all of this in dummy data[?], so they're not going to mess up their actual account.

[00:26:29] In speaking with users after they've experienced this, we have learned that at the end of the tour they assume they're going to be asked to pay for this functionality, which is not the case. Everything we show in this tour is free. We found that this causes anxiety around whether we're going to suddenly start asking them for a credit card, for money, for tools we've just shown them, and so they just don't try the tools.

[00:26:45] So let's try and see if we can change this behavior. Our hypothesis is, we believe that if we reiterate throughout the demo tour that all functionality is free, users will be more motivated to try those tools. We have a huge amount of user research to suggest there is confusion here. Thanks to Kelly, user research indicates that after completing the demo tour some users are uncertain what free tools they have access to. And user research also indicates that after completing the demo, some users say they assumed the tools shown required payment to use. So that's our evidence.

[00:27:20] Let's take a look at the change we are going to be making. Once again, that's the language and the table. We will split our signup cohort for two weeks into control, who will see the existing demo tour copy, and the variant, who will see updated demo tour copy which mentions throughout that all the demonstrated functionality is free. And I really do mean throughout. Every time we show the users something, we hammer home, this is free, you will never be asked to pay for this. So in the table, 50% no change, 50% reiterate through the demo that everything's free. Cool. Okay.

[00:27:55] Our success metric is, we're going to judge success based on activation of our free marketing tools. At HubSpot we define, or on my team I should say we define, activation as a user seeing success. So generating a new contact, making a lead, and that kind of thing.

[00:28:13] Now, when we ran this experiment we used to have much stricter requirements, I should say, on the P value. We used to require a P value of 0.02, that is a 2% chance the results don't reflect reality. We've since reduced that to a 5% chance, because we had this sudden realization that our experiments don't quite need to be medical grade. So our minimum improvement is 0.4 percentage points over the control on activation. We will need to run the experiment for 14 days in order to reach a P value of 0.02.

[00:28:38] So, it's an interesting experiment. I just think it's kind of funny that our users think our tools cost money when they don't. And once again, I'm about to show you a chart. I'm not even going to be able to show you the numbers on this chart, because this is an activation metric, but it is reflective of reality. We saw a very small improvement in activation on our variant, but nowhere near enough to suggest that the results reflected reality. We didn't consider this experiment a success and we have reverted to the control for all our users.

[00:29:09] This honestly just felt like a home run to us. People were telling us, we think your tools cost money. So to discover that we can't just tell them our tools don't cost money, and that doesn't resonate, this was just a huge surprise. We do know that nine in ten[?] experiments we try don't lead to a metric improvement. So we have attacked this quite a few other ways since, with follow-on experiments, and we were eventually able to reduce the number of people telling us they thought our free tools were paid. I just wanted to tell you about that experiment. I think it's a good example of something going wrong despite seeming to be a slam dunk.

[00:29:46] Thank you very much. Despite everything going on, we're still hiring like the absolute clappers. We're hiring across the board, every single one of the roles you heard me mention today. We are particularly interested in chatting to senior level UX experts. I'm going to make my Twitter unprivate for the duration of UXDX, so come say hello. Thank you.