Checking session availability…
Hang tight while we load the latest updates.
Modern websites love JavaScript but it is the number one performance killer today. In this talk, Taylor will go through a case study on improving the performance of the Kroger.com website (one of America's largest supermarket chains) by 11x from an order time of 3 minutes 44 seconds to 20 seconds.
- The value of performance
- Javascript budgets
- Perceived performance
- Design tradeoffs
- SPA versus MPA
- Lessons learned
Improving Website Performance by 11x
Taylor Hunt at UXDX EMEA. Video: https://www.youtube.com/watch?v=uINCmQqrC-g
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
An 11 times faster Kroger.com
[00:00:00] Hello. This talk is about a drop-in replacement front end for Kroger.com, where the time from first load to checkout was 11 times faster than the existing front end. Here it is on the left, with the existing site on the right. These are screen recordings of a $40 Android phone over simulated in-store cellular data; more on that later. The front ends are racing to search for eggs, add them to cart and then check out. The 11 times speedup is that whole journey, not just initial load. My demo is a prod build in front of real APIs, through Akamai, over the real internet, and it was also responsive.
[00:00:32] In contrast, this is actually one of the flattering recordings of the existing site. I had a lot of takes where I interacted too quickly and broke the interface before it was ready. My front end never shipped, but it got close enough that UXDX thought its story was valuable to learn from. I certainly learned a lot from it. In fact, I learned enough that this talk will have to go quickly, so don't be afraid to use that pause button.
[00:00:55] Website performance is like cars: to be 11 times faster, you're going to need to optimize many things. Cramming a V12 in a minivan won't make it raceworthy, and swapping a site's JavaScript framework won't magically make it fast. You need to scrutinize everything, from kilobytes hitting the network to pixels hitting the glass.
Why speed matters and why we're bad at it
[00:01:12] The core idea is simple: don't make humans wait for websites. That's truly all performance is. When you try quantifying that, it all goes to hell. You've probably seen stats like these; there are hundreds of them. This is the tech industry: we can't just make software fast because it respects users or makes the world better. We need to prove to managers that performance isn't another thing developers want at the expense of making money, like refactoring or 100% types or other code virtue signals. Which is odd, because the math says improving performance is one of the surest returns on investment a website can get. Speed is the king feature, because it improves all your other features. No feature exists until it finishes loading.
[00:01:46] At least, that's what the math says. But for such a certain investment, how come we're so bad at making fast websites? Why do we keep getting slower, out of proportion to most users' networks and devices? Why do we set new devs up to fail with React's buzz [?] by default, when keeping React fast is not a beginner topic? These are simple questions, but they don't have simple answers, because if they did, we'd have solved our performance problems by now. So instead, let's answer an easier question: what happens when I tried changing buying from Kroger.com to take seconds instead of minutes?
[00:02:17] Kroger's the biggest grocery chain you've never heard of, because it has a ton of sub-brands, but it's 17th on the US Fortune 500. So Kroger.com was in the bigger leagues of web dev, where nobody shuts up about scale. It was a React single-page app that was really slow, and still is, according to PageSpeed Insights. To be fair to my former employer, Walmart's and Target's sites are also React and also really slow.
[00:02:39] I started learning web development from influences like Rachel Andrew and Jeremy Keith, who championed reach as the web's greatest strength. As far as I was concerned, an ideal website turns away as few users as possible, regardless of their browser, device, disabilities or network. As you might imagine, there was culture shock when I was hired to work on a giant React single-page app. I frequently disagreed with co-workers.
Putting a dollar figure on a millisecond
[00:03:00] One of those times was when Tailwind CSS slowed our desktop time to first paint by half a second. I complained that kind of impact meant we shouldn't use it. I was countered with a fair question: half a second sounded bad, but what did it actually mean to us? I didn't know. So a co-worker and I founded a performance team to find out. We'd seen those stats about how faster sites improve business metrics, and we wanted those numbers, but for our site.
[00:03:22] In theory, you can find them by improving performance, then measuring metrics changes. Unfortunately, that's a catch-22. You need a large speed improvement to get statistical significance, but you need statistically significant numbers to get approval to work on large speed improvements. So instead, use others' homework to get a ballpark number. Substitute your own average order value or whatever, then extrapolate. You won't get precise numbers, but that's fine; they constantly change anyway.
[00:03:45] Our Kroger.com estimate was $40,000 of yearly revenue per millisecond of load time. Then the pandemic surged our usage, since buying groceries in public was suddenly a bad idea. Whatever this number is today, it's probably way higher. You should recalculate after changes like that. But even then, don't consider them the limits of what performance can achieve.
[00:04:02] Time doesn't have a linear relationship with dollars. The closer wait time gets to zero, the more users change behavior. The reason this graph puts quotes around "engagement" is because your fastest performance timings are the rarest ones, as shown by the other line, and often from users who aren't really using the product in the first place, like how an empty view loads much faster than a useful one. It may be better to represent that part as a weird quantum range, where the data can't predict how performance can improve metrics, especially if you change the red boundary of best possible.
[00:04:31] Unshackling speed for users not only changes their behavior, it opens your site up to new worlds. For example, YouTube once shipped a 90% smaller watch page, but their average performance metrics tanked, because users with terrible connections could finally start watching videos. Minor performance improvements pay for themselves, but meaningful speedups make your site become orders of magnitude more useful, by finding new users that do things you can't anticipate.
[00:04:56] But back at Kroger, we were having a hard enough time with incremental performance improvements. People said they cared, but the performance tickets always got out-prioritized. Speedups got hoarded in case they needed to ram a feature through the bundle check, and even the smallest developer experience gain always seemed to outweigh what it cost our users. Maybe proving that speed equaled money wasn't enough. We also had to convince people emotionally, to show everyone how much better our site could be if it were fast. So, in a fit of bad judgment, I vowed to make the fastest possible version of Kroger.com.
A 20 kilobyte budget
[00:05:25] So how fast is as fast as possible? This is the fastest web page. You may not like it, but this is what peak performance looks like. It's not a useful web page, but it does show that the fastest you can be is an HTTP response that fits in one round trip. I aimed to be as fast as that. How many bytes fit into a response's first round trip? I once thought that was 14 kilobytes, the number you may also have heard. It turns out the HTTPS handshake takes several round trips by itself, so that no longer helped. I needed a more real-world performance budget.
[00:05:53] Luckily, I found the post "Real-World Performance Budgets" from Google's former Chief Performance Mugwump. He advocated picking a target device and network, then finding out how much data they can handle in five seconds. For maximum relevance, I chose Kroger's best-selling phone, the Hot Pepper Poblano. Its specs might have been good once. My target network was cellular data as filtered through our metal buildings. Walking around with a network analyzer told me that resembled WebPageTest's slow 3G preset. The Poblano and the target network specs resulted in a budget of about 150 kilobytes, which was not bad.
[00:06:27] Problem: Kroger.com's third-party JavaScript totaled 367 kilobytes. My boss told me in no uncertain terms which scripts I couldn't get rid of. So after further bargaining, workarounds and compromises, my front-end code had to fit into 20 kilobytes, which is less than half the size of React. But Preact is famously small, so why not try that? A quick case limit [?] seemed promising. The Preact ecosystem had a client-side router and a state manager integration I could use, leaving about five kilobytes left over.
Why not a single-page app
[00:06:53] But is this all the JavaScript an SPA needs? Some more code I knew I would need eventually: translating UI to JSON and back again. Something like React Helmet, but not the Preact Helmet package, because that one is four kilobytes for some reason. Those "this app just updated, please refresh" notices. Something like the webpack module runtime, but hopefully not actually the webpack module runtime. Re-implementing browser features within the page navigation lifecycle. And if any of you have done analytics for SPAs, you know they don't work out of the box. I don't have size estimates for these, because I had already abandoned the single-page app approach.
[00:07:26] Why? I didn't want a toy site that was fast only because it ignored a real site's responsibilities. I see the responsibilities of grocery commerce as security even over access, access even over speed, and speed even over slickness. I refused to compromise on security or accessibility; I didn't want any speedups that conflicted with them.
[00:07:45] The first conflict was security. It's not fundamentally different between multi-page apps and single-page apps; both need to protect against known exploits. For multi-page apps, almost all that code lives on the server, but for single-page apps... For example, anti-cross-site request forgery, where authenticity tokens are attached to HTTP requests. In multi-page apps, that means a hidden input in form submissions, but single-page apps need additional JavaScript for client-side details. Not much, but repeat for authentication, escaping, session revocation and other security features. I did have five kilobytes left over, so I could spend them on security. Discouraging, but not impossible to work around.
[00:08:20] Speaking of impossible to work around: unlike security, client-side routing has accessibility problems exclusive to it. First, you must add code to restore the accessibility of built-in page navigation. Again, doable, but it means more JavaScript, usually a library, but downloaded JavaScript just the same. Worse, some SPA accessibility problems can't be fixed. There's a proposed standard to fix it, so once that's implemented and all assistive software has caught up supporting it, it won't be a problem anymore. Don't hold your breath.
[00:08:48] But let's say I'm going to get all that. Sure, it sounds difficult, but theoretically it can be done by adding client-side JavaScript. So we're back to my original problem. Beyond the inexorable gravity of client-side JavaScript, single-page apps have other performance downsides. Memory leaks are inevitable, but they rarely matter in multi-page apps. In single-page apps, one team's leak ruins the rest of the session. JavaScript-initiated requests have lower network priority than requests from links and forms, which even affects how the operating system prioritizes your app over other programs.
[00:09:18] Lastly, server code can be measured, scaled and optimized until you know it's fast enough. But client devices... This is a chart of processing speed across iPhones, flagship Androids, budget Androids and low-end Androids, in descending order. The web's diversity makes "fast enough" impossible to know for the client side. Devices are unboundedly bad, with decade-old chips and miserly RAM in new phones. The Poblano was Kroger's bestseller two years ago, and it still is. Even if the two cheaper categories at the bottom get significantly faster, would you bet against a cheaper, slower one popping up a third time? Trick question: the Poblano already scores below that fourth line.
HTML streaming and Marko
[00:09:54] All of that was enough for me to abandon a single-page app. And so I doomed my site to feel clunky and unappealing. Or did I? Chrome joined other browsers in 2019 with paint holding, which eliminates the dreaded white page flash between navigations. I figured if I could send pages quickly, interactions would seem smooth and not jarring. I figured if I inlined CSS and sent HTML as fast as possible, the overhead would be negligible compared to the network round trip. Concatenating strings on a server really shouldn't be the bottleneck; we've been able to do that in under a few milliseconds for 20 years.
[00:10:26] But there was one problem with quickly generating HTML. Like many large companies, Kroger.com's pages were made from multiple data sources, which could each have their own teams, speed and reliability. If these 10 data sources each take one API call, what are the odds my server can respond quickly? Odds are pretty bad. If 1% of all data responses are slow, then a page with 10 back-end sources will be slow 9.5% of the time. And sessions load more than just one view. If a session has eight pages, then that original 1% chance turns into near certainty for every user, which is even worse than it sounds. A one-time delay has a cooling effect on the rest of the session, even if everything afterward is fast.
[00:11:05] So I needed to prevent individual data sources from delaying the rest of the page. I suspect this problem alone may be why so many big sites choose single-page apps. But this cooling effect also argues that maybe a single-page app's slow load up front doesn't really fix the problem. We had fast websites from big companies before we had single-page apps, so there's no way this was a new problem. I vaguely remembered early performance pioneers saying browsers can display a page as it's generated, but I couldn't remember what that was called. It turns out that's because everyone calls it something different. Regardless of the name, this is a technique Google Search and Amazon have used since the 90s, and it's even more efficient in HTTP/2 and 3, so it's here to stay. Now let me show you what HTML streaming is.
[00:11:46] These pages both show search results in five seconds, but they sure don't feel the same. Beyond the obvious perceived improvement, this also lets browsers get a head start on downloading page assets, doesn't block interactivity like hydration does, and doesn't need to block or break when JavaScript does. It's also more efficient for servers to generate. Clearly I wanted HTML streaming, but how do you do it?
[00:12:07] Today, popular JS frameworks are buzzing about streaming, but at the time I could only find older platforms like PHP or Rails that mentioned it, none of which were approved technologies at Kroger. Eventually I found an old GitHub repo comparing templating languages, which had two streaming candidates: the doubly deprecated Dust, and something I had never heard of, but at least it wasn't Dust. Disclaimer: I now work on the Marko team, but the following predates that.
[00:12:29] Okay, so Marko could stream. That was a good start. Its client-side component runtime was half my budget, but it was zero JavaScript by default, so I didn't have to use it. Where's that streaming, though? Buried in the API docs, I found it. Code like this results in search results like before. It uses HTTP's built-in streaming to send a page in order as the server generates it. But await had even more tricks up its sleeve.
[00:13:02] Let's say fetching recommended products is usually fast, but sometimes it hiccups. If you know how much money those recommendations make, you can fine-tune a timeout so that their performance cost never exceeds that revenue. But does the user really have to get nothing if it was unlucky enough to take 51 milliseconds? The client-reorder attribute turns await into an HTML fragment that doesn't block the rest of the page and can render out of order. This requires JavaScript, so you can weigh the trade-offs of using it versus a timeout with no fallback. Client-reorder is probably a good idea on a product detail page, but you want it to always work for the user on the dedicated recommendations page.
[00:13:36] Await had already sold me, but Marko had another killer feature for my goal: automatic component islanding. Only the components that actually could dynamically rerender on the site would add their JavaScript to the bundle.
Designing for speed
[00:13:46] Right: I had my goal, my theory, and a framework designed to help me with both. Now I had to write the components, style the design and build the features. Even with a solid technical foundation, I still had to nail the details. At first I thought I'd copy the existing site's UI, but that UI and its interactions were designed with different priorities, so I had to suck it up and redesign. I'm no designer, but I had a secret weapon: users think faster sites are better designed and easier to use. My other secret weapon is that I like CSS. The bonus of me doing both is that I could rapidly weigh pros and cons to explore alternatives that better compromised between user experience and speed, or even alternatives that improved both.
[00:14:22] My design priorities from before still held. This website sells food; I will ruthlessly sacrifice delight for access. My login page on the left is no prize winner, but you've got to admit it is something over the real one on the right. Another problem was our product carousels. They didn't fit the Poblano's screen, and because they took up so much real estate, trying to scroll past them would scroll-trap me, like a bad Google Maps embed. Their dimensions took up so much space that they caused juddering from layout cost and GPU pressure. So I went boring and simple, relying on text's horizontal nature and some enticing "see more" links instead of infinite scrolling.
[00:14:55] Then an easy choice: no web fonts. They were 14 kilobytes; I couldn't afford them. Not the right choice for all sites, but remember, groceries. This industry historically prefers effective typography over beautiful typography.
[00:15:08] Unlike my decisions to follow a compromise, banning modals and their cousins was a user experience improvement. Do you like pop-ups and all their annoying friends? They also make less sense on small screens. In particular, modals take up nearly the entire page anyway, so they might as well be their own page. Lastly, they're hard to make accessible, and the code required to do so really adds up. Check out the size of some popular modules, each a reasonable choice for tooltips, modals and toasts respectively. The alternatives to those widgets weren't as flashy and were sometimes harder to design, but hey, great design is all about constraints, right? That's how I justified it.
[00:15:42] This led to a nice payoff: fast page loads can let you design better. Our existing checkout flow had expanding accordions, intertwingled error constraints, and tricky focus management because of the first two. To avoid all that, I broke checkout into a series of small, quick pages. It didn't take long to code, it was surprisingly easy, and the UX results were even better. With paint holding, a full page navigation doesn't have to feel heavyweight. I could do all this easily because I was simultaneously the designer and the developer, which used to be a lot more common when web designers were expected to know HTML and CSS. An intuitive understanding of what's easy versus what's hard in web code can go a long way.
The result, and what it could be worth
[00:16:18] So after all that, what did it get me? See for yourself. Now, is this comparison truly fair? No, because the Android app got to skip its download from the App Store. Amazon also keeps trying to move off jQuery, or somewhere here. Jokes aside, could my demo have kept this speed in the real world? I think so. It hadn't yet withstood ongoing feature development, but far bigger and more complex software successfully uses regression tracking to withstand that, such as web browsers themselves. Later features that could live on other pages wouldn't slow down the ones seen here, thanks to the nature of multi-page apps. And this recording also lacks performance tuning I dearly wanted, like edge rendering and faster HTTPS.
[00:16:57] Even if it got twice as slow, our $40,000 per millisecond figure from before would estimate this kind of speed would equal another $40 million of yearly revenue, assuming an 11 times speedup wouldn't change user behavior. And you know it would.
[00:17:11] But enough about me; let's talk about you. Do you compete with Amazon? Do you want the web to compete with native? Do you have the responsibility, or a business opportunity, to serve the most users possible, even on horrible networks and devices? At the very least, do you like the sound of adding millions in revenue? If you can say yes to any of that, then this video should represent something important. Because great performance is exceptional in this industry, you can become exceptional with great performance.
Three rules for exceptional performance
[00:17:39] And if you want to be that exceptional, here's how. There are only three rules, but each could use unraveling. First: aim to be noticeably faster in the ways that matter to humans. Fitting in 20 kilobytes for my first load was just to ensure the rest of my grocery-buying flow was set up for success; more code loaded later as it was needed. Metrics such as time to first byte and first meaningful paint are for diagnosis, not meaningful payoffs. Twenty percent sooner time to interactive is just maintenance.
[00:18:05] If you want the speed I showed, under the circumstances I targeted, I can guarantee Marko's Rollup integration has it. But your needs may differ. Don't pick the technology until it meets your goals on user hardware. I'm often asked if other technologies can be as fast as what I used, and the bad news is I don't know. The good news is that if you set important goals and follow through with them on real hardware, you'll inevitably find technologies that are fast enough. To minimize throwing away work, rough estimates like I've been showing here are useful to avoid dead ends. Looking at the performance characteristics of older, well-known technologies will also help guide what you can accomplish. We knew how to make fast websites decades ago. Many things have changed, but just as many haven't.
[00:18:42] You probably balked at my 20 kilobyte limit, but it's not as unrealistic as you might think. Modern websites can accomplish a lot in similar amounts; Google's suggested budget for feature phones is only 30 kilobytes, for example. In fact, I doubt you'll risk being too ambitious. Cars aren't engineered only for dry roads under good conditions. Ambitious performance goals benefit everyone all the time. Usable speeds in the worst-case scenario translate to even faster, more consistent experiences the rest of the time.
[00:19:07] Second: remember we're doing this for humans. Now is an unprecedented time for low-income connectivity. Getting online used to require stable housing, but nowadays you can be on the real internet with a real browser for almost nothing. Correspondingly, smart devices and the internet have become important tools no matter your status. If you're homeless, why wouldn't you want a way to reach important contacts, apply for jobs and talk to other people? Wouldn't you also appreciate at least browsing products online, just to cut down on time spent in a grocery store, especially in the midst of a contagious disease you can't afford to catch?
[00:19:39] That's why it's important to verify you can serve those people, by checking your assumptions on the devices they can use. Tools like Google Lighthouse are wonderful and you should use them, but if your choices aren't also informed by hardware that represents your users, you're just posturing. The Poblano's stated 1.1 gigahertz isn't fast to begin with, but it gets worse. If it ran at that speed all the time, it would produce the heat of a 40-watt incandescent bulb. Because phones don't have fans to dissipate that kind of heat, they throttle themselves when they start overheating. This happens with every phone, not just the cheap ones. And that's just the CPU. Your time to respond to user interaction may be fine in the abstract, but how does it feel combined with the delay from a cheap touch screen?
[00:20:19] Your company probably doesn't sell phones, so you can't do what I did. If you really don't have any hints for how low your user base can go, you can at least check what cheap Androids Amazon sells a lot of. Also, the network throttling in Lighthouse and browser dev tools is better than nothing, but they're doomed to be much more optimistic than actual slow networks. You'll need something that throttles at the packet level. Packet-level throttling is more work to set up, but it's the only way to be accurate. Here are some programs that can do it. Worst-case scenario, there's always some command-line program you can search Stack Overflow to figure out.
Taming third-party scripts
[00:20:50] Now, if you can avoid third parties, then perfect, do that. It would have made my life 130 kilobytes easier. But absolutism won't help when businesses insist. If we don't make third-party scripts our problem, they become the user's problem. It's tempting to view third-party scripts as not your department, or as a necessary evil you can't fight. But you have to fight to eke out even acceptable performance, as evidenced by me devoting 87% of my budget to code that did nothing for users.
[00:21:15] First, apply engineering principles. Marketing revenue is one thing, but decreasing its cost significantly is surprisingly low-hanging fruit. Marketing people aren't stupid, but they have different incentives, and they're missing the incentive to care about the partnership after the window ends. Write down what third parties are live on your site and for how long, and, like all accounting, track their actual payouts, not just the projected estimate they approach you with. This can be as fancy as automated alerts or as simple as a spreadsheet. Even a little bookkeeping can avoid money sinks.
[00:21:44] For example, TikTok's affiliate program. There are people with entire jobs dedicated to convincing you that their JavaScript snippet is free money, but remember that advertisers lie for a living. Loading their proposed JavaScript on the most powerful computer I had access to proved that the juice wasn't worth the squeeze. TikTok is an unusual network card [?]. Third-party code offers massive convenience, but in exchange for a lot of risk. With increased legal scrutiny of just how profitable selling other users' data is, weighing the risk is a skill you may want to practice.
[00:22:11] The ones you can't avoid, you can still compromise with, to find happier mediums between marketing income and user penalties. The image alternative is usually warned as only 90% accurate, and the payout is similarly cut, but a single non-blocking HTTP request has almost no overhead compared to yet more JavaScript. A lot of analytics providers have also started offering server-side usage, because ad blockers and other user privacy tools have warped their measurements. I've also heard good things about Partytown. It wasn't ready at the time, but that was two years ago.
Blind spots in the data
[00:22:40] Worse than the performance penalties of analytics is when their blind spots lead to plans based on a distorted view of the world. Measuring accurate, valid data is hard. People get degrees in statistics because statistics is tricky; much of the field is learning how to avoid lying to ourselves with numbers. When it comes to web performance, I worry that we grapple with one of the strongest cases of survivorship bias. Getting even a rough number of how much market there was for my prototype was bizarrely difficult. It wasn't that I rejected data that didn't support my hypothesis; it's that the data was either untrustworthy or completely absent.
[00:23:12] Our analytics' reported bounce rate for older Android versions varied by up to 10% between days. That kind of variance isn't just impossible to use; it indicated something was fundamentally wrong. Then I tried asking our Dynatrace [?] folks how many HTTP requests never made it to a fully JavaScript-initiated session. That discrepancy could show how many users left before our site got its act together. Unfortunately, a week later they told me Dynatrace intentionally discards that information. This sort of thing happened a lot.
[00:23:39] Now, these instances don't prove there were a huge number of folks trying to use our slow website and leaving frustrated. But a hypothesis that survives several attempts at refutation is at least an interesting possibility. It's wild to me how many decisions are made using analytics software that doesn't even estimate how much it can't tell you. Is it that hard to believe we might have a blind spot? Our industry trends rich, white and male, so there's already a bias to compensate for. Then, if our stats come from JavaScript that must download, parse and then upload to record a user, it's not surprising we'd miss users we're too slow to serve.
Normal web development, and what's stopping you
[00:24:12] As unusual and hardline as the rules of meaningful web performance may seem, maybe the strangest part is how normal accomplishing them can be. To make my 11 times faster front end, I picked a software platform, then wrote business and UI code inside it that mingled back-end data with links and buttons, aka normal web development. I didn't even touch the back end. The actual hard part seems to be consistent priorities across the platform, the component code, the design and the product decisions.
[00:24:38] As much performance knowledge as this project required, I wonder how much more knowledge of a different kind it would take to ship it intact, so it can actually improve users' lives. Luckily, if you're attending UXDX, that kind of knowledge is likely your job. I wish I had more useful advice and processes that produce web performance, but that would be disingenuous. After all, my thing never shipped; I hadn't found a process that works.
[00:24:59] Which raises the question: if web performance is so good, what stopped me? Truthfully, I don't know. None of our proposals got specific rejections. But even if I knew, there's a more important question: if the technology has been here, and users are out there, what's stopping you?
