Improving Quality Of Response Times: Application Performance Measurement
Checking session availability…
Hang tight while we load the latest updates.
- Collecting application performance metrics at Sky
- Understanding how your service is performing to speed up production
- Using histogram metrics to monitor response times
- Good and bad practices when monitoring response times
Improving Quality Of Response Times: Application Performance Measurement
Daniel Rolls at UXDX EMEA. Video: https://youtu.be/AwVH46_Lsp8
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
The questions we ask of a web service
[00:00:00] At Sky, we build services to coordinate the delivery of video content over the internet. The brands you might be familiar with are NOW TV and Sky Go. Our services run 24/7. We see usage peak in the evening, when everyone comes home and starts watching their favorite TV programs, and we see usage bottom out in the middle of the night, when everyone's in bed. Naturally, we want to understand our web services.
[00:00:29] For this talk, a short talk, let's concentrate just on service time, and let's concentrate just on a single endpoint on a web service. But imagine whatever you're monitoring in your software: there are similar lessons to learn. If we concentrate on that, there are many questions you can ask, such as: how much traffic is hitting that endpoint? How quickly are we responding to a typical customer, whatever a typical customer is? Well, maybe one customer complained we respond too slowly. How slow do we get? That is to say, what is our maximum response time?
[00:01:08] Now, the perception of response time is related to the consistency of response times. So what's the difference between our fastest response time and our slowest response time, i.e. what is the range? Well, maybe the slowest response time occurs one in a million times. Maybe we don't care about such anomalies. So how slow are our slowest one percent of response times? That is to say, what is our 99th percentile response time?
Dropwizard Metrics and the dashboards we depend on
[00:01:42] Now, at Sky our applications are predominantly written in Java. We make heavy use of the Dropwizard framework, and it comes with a fantastic library called Dropwizard Metrics. It does a lot of this for us. We got lucky: it does it on the application hosts. So we plugged that in and started using it, and with very little effort we were able to produce dashboards like this. Maybe you have something similar in your working environment, maybe you're trying to get something similar. For monitoring, we typically graph time on the x-axis and response time on the y-axis, and we typically use color coding to encode different percentile response times.
[00:02:26] We got lucky. We used this framework with very little configuration. We just pointed our web service at Graphite, which was filling up with useful statistics, and we pointed dashboards at the statistics. With very little effort, we could get all sorts of information on, say, JVM memory usage in Java. Just by adding timer annotations to our classes, we got information on the rate of requests and on percentile response times for our applications. All very easy. Now, this library has ports, at least for Go and Scala. This talk is not trying to sell this particular library. I hope that what we learn applies no matter what you're using.
[00:03:04] What we are using was dreamt up by our architects, and developers make heavy use of these dashboards. It's not unusual at Sky to see developers gathered round dashboards, looking at them, trying to understand application software. We feel responsible not just for writing software, but for knowing about the health of our applications in production. And when something goes wrong, managers want numbers, and the numbers they get, they get from these dashboards. So we trust these dashboards, and we all depend on them, but we rarely understand them, and we have a tendency to lie to ourselves, and to lie to our managers, with them.
[00:03:49] It's my belief that when that graph shoots up to 500 gigaflops, we should know whether that 500 is a number we want to give to our manager. Is that a number you want against your name? The graph shooting up may be indicative of a problem that needs to be investigated, but the 500 is not a number that should be reported in itself.
What is the 99th percentile?
[00:04:12] Bearing that in mind, the goals of this talk are to understand how we can measure service time latencies, to make sure we give meaningful statistics to our managers, and lastly to learn how to use appropriate dashboards for monitoring and for alerting. I'd like to start by posing a question, and the question I'd like to pose is: what is the 99th percentile response time?
[00:04:35] We typically graph the median response time, the 50th percentile, and the 99th percentile. And we typically do a pretty good job of testing, so being on support for most teams is pretty boring, which is great, we like that. The median response time typically doesn't budge. It's really boring, it just stays flat, nobody cares. The 99th percentile, on the other hand, jumps about a bit. It moves, it changes characteristics all the time, sometimes depending on the time of day, and it naturally gets criticism. It's not uncommon for developers to say the 99th percentile is broken, it's not working. But to say such a thing suggests you know what it is.
[00:05:14] So what is it? Well, didn't I just give you the answer? It's the top 1% of response times. On the next slide, maybe you're picturing some sort of distribution, and it's at the top end of this distribution. If you had 100 response times in order from fastest to slowest, it would be the 99th. If you had 1,000, it would be the 990th. But we're plotting this percentile against time, and when a manager says, "What's the 99th percentile?", they typically mean the 99th percentile now. But "now" doesn't really mean anything mathematically, does it? "Now" doesn't help me pull out some data that I can draw such a distribution from and pull out the top one percent. So which values do we use?
Reservoir sampling: sliding windows and exponential decay
[00:06:01] As it turns out, the values we use we call our reservoir, and the techniques that we use we call reservoir sampling. We need to decide on a scheme for what goes into our reservoir and what leaves our reservoir. Fortunately, our library gives us some options, on the next slide.
[00:06:18] The first option it gave us is a simple sliding window. Let's just keep the last 1,000 values. Simple, right? We can all understand that. Unfortunately for us, in the evening, when everyone starts watching TV, those 1,000 values might represent a few seconds' worth of data. In the middle of the night, when everyone goes to sleep, maybe they represent the last few hours' worth of data. The graphs behave very differently at different times of the day, and at night they don't seem to budge, they don't seem to react. Criticism.
[00:06:53] Okay, a simple solution to that is a time-based sliding window. Let's just keep the last 5 minutes' worth of data. That's great if you're on support, because if you see a problem in the graph and you think you've fixed it, wait five minutes, look for the next data point, and you immediately know whether you've fixed it or not. Unfortunately, we never really know how much data we're going to get in five minutes. Maybe we're going to get more than we can store in the application instances themselves. This is a really inefficient option, and not one I would recommend.
[00:07:24] The last bullet point is an exponentially decaying reservoir. Exponentially decaying reservoirs are great, because they weight more recent data exponentially higher than older data. That's kind of cool, it's kind of what we want. It's the default: if you're using this library, it's probably what you're using. It's what we're using. So let's take a closer look at it on the next slide, please.
[00:07:49] Imagine a timeline, and imagine events on this timeline, which are representing response times. Let's pick an arbitrary point on this timeline. On the next slide, we pick an arbitrary point, called a landmark, and on the next slide let's measure the distances between the landmark and the response times. We call these X. So X is high for more recent data and low for less recent data. On the next slide, you can see that the formula for weight is simply e raised to alpha X, for some constant alpha. So basically we have weights that are exponentially higher for more recent data, and we have these pairs, the V's and W's, response times and weights.
[00:08:33] On the next slide, let's sort these pairs from fastest response time to slowest response time. Just like in school, when you learned to do interquartile ranges, you put all the numbers in order. So let's do that. But now, on the next slide, let's do something slightly different to what we did at school: let's normalize the weights. Let's add up all the weights in the table and divide each weight by the total, giving us weights that sum to one.
[00:09:01] Now, when we want to look up a percentile, let's do something slightly different. Say we're looking up the median, the 50th percentile: 50% is 0.5. Let's start adding up the weights from the left of the table until we get to a value that's larger than or equal to 0.5. When you get that value, read the corresponding V value. That is your response time, the median response time in this case. So if you imagine you had 500 very slow events in the past, a gap in time, and then 500 very fast events, because the fast events are more recent they're going to have higher weights, so the median is going to shift towards that more recent data, which is kind of cool, and kind of what you want.
[00:09:41] What I haven't mentioned, on the next slide, is data retention. The data is stored in a sorted map, indexed by the weight multiplied by a random number. The random number just guarantees uniqueness. And we remove smaller indices first. Smaller indices are exponentially more likely to be older data. That's great, it's kind of what we want.
Experiments with a dummy application
[00:10:02] I found this out by looking in the code. I thought I had it in my head, but the problem is, when you want to test your assumptions by looking at the graphs of production software, it's hard to pull apart the production software and the quirks of the graphing solution. So what I did was build a dummy application, an application that I could tell to behave any way I wanted, and I pointed our dashboards at it.
[00:10:28] This graph is experiment number one. I'm graphing percentile response times ranging from the median, the 50th percentile, up to the maximum, on the graph. I told my dummy application to start responding in 50 milliseconds, persistently, and all the graphs were about 50 milliseconds. Previous slide, please. Then I told it to start responding in 300 milliseconds, and the upper percentiles quickly shot up to 300 milliseconds, followed by lower and lower percentiles. I then told the application to start responding in 50 milliseconds again, and the same thing happened in reverse: the lower percentiles came down, followed by higher and higher percentiles.
[00:11:05] This makes sense when you think about it. Your application is consistently responding in 50 milliseconds, and one stray request takes 300 milliseconds. If someone asks what your max was, you have to be honest and say it's 300 milliseconds. As more and more requests take this time, you have to admit that lower and lower percentiles take this time.
[00:11:21] Okay, next slide. I did another experiment here. I told my demo application to make one percent of the requests, one in every 100, take 500 milliseconds. Pretty quickly, the maximum, the purple line on the graph, jumped up to 500 milliseconds, and then eventually the 99th percentile came up, and then came down, and up, and down. Well, it was one in every 100 requests, right on the line, so we shouldn't be surprised. But note that the 99th percentile jumps all the way up to 500 and all the way back down to 20 milliseconds. Well, this makes sense once we know how the reservoirs work. We know the values in the reservoir are real measured values. These numbers on the graph are numbers that we can carefully report to our managers.
Averaging percentiles makes no sense
[00:12:06] Now, at Sky we have different clients coming in to watch video. You can watch content on your phone, on a PlayStation, on the PC, and some clients are more popular than others. So let's shift this graph up: this is graph A, and graph B below is the same again, a 1% rise from 20 milliseconds to 500 milliseconds. Graph B gets 10% of the traffic; it represents a less popular client. You see the same thing happening, except it takes longer. That shouldn't be a surprise: it's got 10% of the traffic.
[00:12:41] Now, typically, developers graph all of these clients together, and they get one big mess on the graph, and they can't understand anything, and they think, "Oh, I'm going to have to do something clever here." And typically what they do is take the arithmetic mean of corresponding percentiles. So that's what I did. On the next graph, if you look at the purple line, the maximum, it starts from 20 milliseconds as before, and it ends up at 500 milliseconds, as before. But in the middle you've got this strange value. What is that, like 260 [?] milliseconds? What does that mean? Can I give that mad number to my manager? Should I put that number against my name? If we've said that we should take action if response times go above 300, should I do anything?
[00:13:26] As it turns out, taking the arithmetic mean of corresponding percentiles makes absolutely no sense. But not only did that happen at Sky, I've since learned, doing these talks, that it's happened at many other companies as well. It's a very common thing. We also learned the hard way that it's a really bad idea to break your data down into separate reservoirs and then try to pull it back together again, because you will never get the same value, and you'll have all sorts of problems. So don't do that.
Lessons learned
[00:13:51] On the next slide, let's talk about lessons learned. As developers at Sky, we're really proud of being autonomous teams. We're really proud of the fact that we can know our applications, and the performance of our applications, far better than one person or team responsible for the performance of all the applications at Sky could possibly know. In the next bullet point, I'd like to make the case that developers aren't perfect. We do make mistakes, we get things wrong.
[00:14:19] We used similar techniques to graph the relative performance between data centers. The same mistakes were made, and the numbers made absolutely no sense at all. We were happy to find it. We told our manager, we told our fellow developers, and we told our architects. We made one really silly mistake: we didn't take the dashboards down. A week later, traffic was shifted between data centers, and all three groups stood around the screen and started commenting on the relative performance of the data centers, just like, six years earlier, academics had been looking at tables comparing universities [?]. So, as it turns out, we won't just make mistakes, we'll ignore the mistakes as well.
[00:15:04] Next slide, please. So what did we learn? Well, if you want fast alerting, use the maximum. And if you hide the maximum, what is it that you're hiding? If your 99th percentile is high, that means at least 1% of your requests are that bad, maybe considerably worse. Use one reservoir for each thing you want on your graph, and you'll get far more accurate results. Now, in the web services world you still need to aggregate across application instances, and when you want to aggregate, you're interested in the worst case. Had we been taking the maximum, the line would have shot up to 500 and immediately signaled problems. The worst case is the maximum, so use the maximum.
[00:15:45] Don't immediately assume the numbers on your graph are meaningful. Try to understand what it is you are graphing. Developers love playing around with tools. If you give us tools, we will happily play with them, we'll happily plug them in, but we won't necessarily get it right. We'll make mistakes. So how can we make sure that we don't make these mistakes? How can we stop these things happening? Well, like most things in our discipline, in the end I believe it comes down to the KISS principle. Wherever possible, keep things simple, and try to make sure that everyone in your team knows what the values on your y-axis actually represent. Only then do we stand a fair chance of knowing which numbers are lies and which are real. Thank you very much.
[00:16:36] And first of all, I'd like to apologize...
