Server Components, Out-of-Order Streaming and the Future of the Web
Checking session availability…
Hang tight while we load the latest updates.
Web development is always evolving, but recent movements might stand out as a critical pivot point when we look back in a few years' time.
When we introduced the idea of SPAs and the frameworks around them, we introduced the problem of hydration. Now that we’re moving more and more back to the server, let’s take a moment to look at how new ideas and concepts try to solve the problems of "classical" server side rendering?
How do server components tackle the problem of hydration? How does out-of-order streaming work? And what are some of the new frameworks that are here to help with all of that?
Server Components, Out-of-Order Streaming and the Future of the Web
Julian Burr at UXDX Community: Owning Customer Success & Shaping the Web’s Future. Video: https://youtu.be/XCptAIkRRKU
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Suspense: control over loading states
[00:00:08] Thank you. Yes, I want to talk about the future of web development, and for that I will basically touch on two major innovations that fairly recently popped up into that whole space when it comes to web applications.
[00:00:25] The first one is suspense. React introduced suspense a while ago. Other frameworks have caught up by now. They might call it something different, you might know it under a different name. But the idea of suspense on the client is basically to give us more control over loading states. By giving us the ability to wrap certain parts of our application in what we call suspense boundaries, we can then show a loading state, or a fallback UI, whenever anything within those boundaries is doing some asynchronous work. That's usually fetching data, but it could be other asynchronous stuff like lazy loading JavaScript chunks, or you name it. The idea is that we then don't have to have very granular loading states, like having a spinner in each of our components, and instead we can be much more deliberate about which section of our application should have a loading state and what that should look like. So that's already great.
A history of web rendering
[00:01:24] The second thing that I want to talk about is server components. To understand what problems server components are trying to solve, we need to take a look at the history of web rendering first. We need to go all the way back to the good old days, plain HTML and PHP. It could be any other server-side language. The idea is, plain HTML was great because it was literally just serving files from the server. That's really fast. But it's very static.
[00:01:56] In order to be able to serve dynamic content we started introducing those server-side languages like PHP, and the benefit of that is we can serve stuff that is connecting to a database or other outside sources. But the downside of it is, the more work we actually do on the server, the slower the server response time gets. Especially at scale that can become very noticeable very quickly. At those times that was not only a problem for the initial load but also every subsequent load. So if you navigate around the application, we would need to do a server request for every one of those navigations, and if those server requests take a long time, that's not a great user experience.
[00:02:42] So we started to introduce JavaScript into the whole thing. This is the jQuery era of, like, 2006ish, and the early days of Ajax. Google introduced that concept of being able to do server requests from the client after the page has already been loaded. That's great because that gives us the ability to be much more explicit about what we're actually fetching from the server. So instead of having to fetch the whole page every time the user navigates, we can just fetch certain sub-parts of the page that we know change, and we can keep the main navigation and stuff like that untouched. And because it's all happening in JavaScript on the client, we're in more control over what we actually want to do when we swap out the content. We can maintain client side state like scroll positions, etc.
[00:03:36] This is generally great, it definitely improved the user experience. The developer experience was pretty lacking at the time. It was mixing concerns of, like, where does the HTML actually live? What do we do with JavaScript? We're mixing those two worlds. It's also famously the era of spaghetti code, because at the time frameworks were pretty sparse, so people were usually rolling their own solutions and reinventing the wheel. So that wasn't great.
[00:04:03] As a solution we then just decided, okay, let's just go all in and do everything on the client. We tried to completely remove the server. It's the era of single page applications. These are the frameworks that you probably know of, like your Angular, your Vue, your React. The idea was that the server really just serves an empty shell, like an empty div, and then we do all of the work in JavaScript. That gives us full control over how we want to generate the DOM for the page and what we want to do with the DOM when certain events happen, and it allowed us to build much more dynamic applications, which is great. And the user experience after the initial load, when people navigate around and click around, also great.
[00:04:48] The big downside is the initial load, because it's just sending an empty shell and then doing a lot of work in JavaScript. Based on your device and based on your network speed, that initial load can become a real bottleneck. We realized that fairly quickly, so we started moving back again by introducing concepts like server-side rendering and static site generation.
[00:05:15] The idea of server-side rendering is to have a server that's still running, that gets the request and generates that initial HTML from your React app or whatever framework you're using, and sends that back additionally to the JavaScript bundle that then makes everything dynamic. So this is trying to get the best of both worlds. We still have everything in our frameworks that we love, but we're now doing the server-side rendering to get that quick initial render, and then everything subsequent still gives you the single page application benefits from a user experience perspective.
[00:05:54] This seems great, but it still comes with two problems. For one, we haven't dealt with the original problem of the HTML and PHP era yet: if we do a lot of work on the server, that server response is still going to be slow. The second problem is a new problem that we kind of introduced, which is called hydration. Because everything is written in JavaScript, and by default everything can be dynamic, we have to essentially render our app twice. Things like event listeners or local state don't exist on the server, we can only render a static representation of that on the server, and because we don't really know what's dynamic, we have to essentially, on the client, run a virtual copy of the whole app again just to be able to then inject those dynamic parts. So that's a big downside of frameworks like React, and again, we realized that that became the main bottleneck.
Enter server components
[00:06:49] Enter server components. This is one of the primary things server components are trying to solve. The idea of server components is that you can tell your bundler which of your components are static and which components are dynamic. Static components by definition just render HTML. They don't need any JavaScript, they can just be shipped as part of that HTML response. Nothing needs to be added to the JavaScript bundle. Dynamic components are components that have event listeners, local state, etc. So they do need to be sent to the client.
[00:07:20] But by being very explicit about it, and most importantly by having server components be the default and you have to opt into client components — at least that's how React does it — we're basically going back to what we used to have with HTML, PHP and maybe some jQuery sprinkled around. It's static by default, dynamic as opt-in, and you're much more deliberate about what you're actually sending to the client. So this massively reduces JavaScript bundle sizes. It massively reduces that hydration problem that we saw.
[00:07:52] What it doesn't do is do anything about that first problem of, if the server does a lot of work, that server response time still leads to a bad user experience. This is where the combination of server components and suspense can actually be really magical, by introducing a new concept of how we render our applications.
[00:08:11] To understand that, we're actually going to build a very simplified version ourselves. To do that I'm basically going to try to compare three different ways, historically, how we've rendered web applications. The first one I'm going to call the classical server-side rendering. Again, that has all the problems I just mentioned: when the server takes a long time, the server response is slow. We can improve that with streaming. Streaming has been around forever, essentially, so we're going to take a look at that, and then we compare that to what we call out of order streaming, and this is what suspense in combination with server components allows us to do.
The demo app and classical rendering
[00:08:57] All right, so let's write some code, hopefully. To do all that I basically built a very small demo application to show the difference between the different approaches, not only in code. If you're not a developer or you don't know React or JavaScript very well, don't worry. The frameworks or implementation details are not really the core of this. The main idea that I want to bring across is the concept that's more on a user experience level. So don't freak out when we switch to a code editor in a second.
[00:09:32] But this is the demo application. It's a very simple movie app, because that's what everyone seems to be building in their demos. We have our logo at the top. Then we have a title section with a poster, the movie name and some meta information. We have a detail section with a movie cast, and at the bottom we have a section with similar movies, so I can keep clicking around in that infinite loop of binge watching. Like I said, there's not much to that app.
[00:10:06] If we look at the code, it's an Express application, so a simple HTTP server. There's no fancy frameworks or anything on top of that. And then I basically exposed three different routes, and those are the different implementations that I just mentioned: the classical rendering, the streaming, and the out of order streaming using suspense.
[00:10:31] So let's take a look at the classical rendering. I'm using JSX here and I'm using React to render that JSX to HTML. That's mostly for convenience and because it makes it easier to read. Again, the core point that I'm trying to make is the user experience that we're introducing. We could do all of that with just writing plain HTML strings. So none of this is tied to any framework, this is really more a concept that you can apply to anything. But what we're rendering here is basically what we've just seen on the screen: the header, the title section with a poster and everything, details with the cast, and a similar movies section.
[00:11:16] If we look into one of those components, you can again see it's fairly simple. It just renders that HTML of the movie cast carousel. But it also fetches the data. This is the important and kind of very convenient thing of server components: because they only run on the server, we can make them asynchronous functions, and then they can do asynchronous stuff on the server. The reason why you want to do that is because it reduces the amount of back and forth between the server and the client. We don't create those waterfalls where components fetch data when they render, and then that causes another component to fetch something when that renders, and all that. The servers are already close to your data sources usually, so fetching stuff on the server makes sense usually.
[00:12:07] If we then look at the rest, all we're really doing in the classical render part is we start collecting the HTML. We start out with having the opening tags, then we loop through all of our components. This might look a little bit convoluted. Again, don't worry about the implementation detail itself, but all we're doing here is looping over each of those children. Because they're asynchronous functions, we await them, and then we render that as a string to HTML. At the end we close our HTML and then we send all of that through once we have it.
[00:12:41] And what that looks like is this. This is showing the classical render and it seems to do the job. But if, for example, we open this in a new tab, we see the big problem: the load time. This is taking a while to even appear, which isn't great. It's not only the initial load. Like I said in the very beginning, even if I just navigate to a new movie, so I already clicked it, it's loading, but it's not doing anything. The user doesn't have any visual feedback of what's actually happening. It's a pretty bad user experience.
Streaming
[00:13:13] All right, so I already kind of hinted that we can improve that by introducing streaming. Streaming as a concept, like I said, has been around for ages. It's basically the idea of you opening the document early on when you get the request, and then you just write as data becomes ready, and at the end you close the document. HTML is really good for that because HTML is a very forgiving language. If you don't have any closing tags, the browser will still interpret your stuff. So it's a perfect use case for streaming, as stuff becomes available.
[00:13:53] We're basically doing the same thing that we've done before. We still have our loop here where we render our components to a string, but instead of collecting the content and sending it all at the end, we're already starting to write immediately. That, in Express and other frameworks like that, allows us to immediately open that connection and start streaming. So we also write to that stream, and at the end we need to tell the browser that we're finished. That's all we're doing here with the res.end.
[00:14:24] And if we now go back and open that, we see that the logo shows up immediately, which is great. And then it's all a bit mixed up. Yeah, the sections don't really make sense anymore. Now that we're streaming, we need to make sure that we're actually resolving our promises in the order that they're actually meant to be. So let's just quickly replace this. This is not the most efficient way to do it, there's better ways to still run the promises in parallel, but for the sake of this demo this works. We just change it to a for loop, so the children will be dealt with now in the order that we feed them in.
[00:15:06] And if we now refresh this, that solves a problem. The title section is still at the top. It's a better user experience for sure, especially when you navigate around. So if you go back here, there's immediate user feedback, the visual feedback of something is happening, but we're still looking at a mostly blank screen. We're only seeing the logo. Imagine if the logo was rendered underneath the title section, then we would still literally be looking at a blank screen. So it's not ideal. It's better, but ideally I would want to show some form of loading state for each of these sections.
Suspense boundaries on the server
[00:15:46] If you remember, this is exactly what suspense allowed us to do, right? It basically introduced the ability to define suspense boundaries and then define your fallback UI for those boundaries whenever asynchronous stuff happens. On the client that's pretty straightforward, but we can also use it on the server, that same concept.
[00:16:07] If we look at it, this is basically still exactly the same implementation as the streaming one, but now I'm going to introduce those suspense boundaries. Going to wrap the title. This is what a suspense boundary looks like, again with a fallback UI that I want to render whenever this isn't ready yet, or whatever is inside isn't ready yet. And I'm in full control over what I'm actually wrapping in my boundaries, which is cool. Let's wrap all of these in their own boundaries. And one. Yes, cool.
[00:16:56] So how does this actually work then? If we look at the suspense components, this is a very simplified version of what suspense components look like. This is not the implementation React has, or any other framework. This is just for the sake of showing the conceptual work that these components do. It doesn't do much, it's very simple, it's 12 lines of code. All it's really doing is it's taking those children that we're passing into the boundary and, instead of rendering, it doesn't render them at all. It just stores them in an object, in a map, based on unique identifiers. And then what it is rendering is that fallback component, and it attaches an identifier to that div so we can later find the fallback UI based on that unique identifier again.
[00:17:45] Nice CSS tip on the side: we're rendering a div here, which can mess up layout stuff. So if you're rendering, like, a CSS grid around your suspense boundary, display contents prevents this div from messing with layout concerns. So you have your grid outside of the suspense boundary, within your grid you have the grid items, this will still work if we set the display to contents. Free tip on the side.
[00:18:16] But yeah, you can see basically what we're doing is we're building up this map and we're storing those potentially asynchronous children in that map. So if we look in here, or if we call that up, you see that's exactly what it's doing. It's only rendering the fallbacks, because we haven't told it to do anything with those children that we just store in the object. But this already looks pretty cool. So let's tell it to do something with the suspended object.
[00:18:52] After the loop where we rendered all the main children, we basically want to check, do we have anything in that suspended object, and if we do, deal with those individually. Basically do what we've done before: you render all of them to a string and then you write that to the stream afterwards. And if we do that, we still render all the loading states, but now the content starts popping in underneath it.
[00:19:23] Because, again, we're in an HTML stream. The main downside of a stream, as I described it before, is it is linear. You can only add to the bottom of the stream, because you're writing the document in one go. You cannot replace previously written content. So this is the main limitation of classical streaming.
A little bit of magic: out of order streaming
[00:19:49] So this is close, but what we want now is to use this content to replace the loading states that we previously rendered. To do this we need a little bit of magic. And when I say magic I of course mean JavaScript. There are a million different ways to actually do this. Some of them actually need less JavaScript, or even no JavaScript. I've seen recent implementations using template slots and stuff like that. I'm going to use some JavaScript, but you'll see it's very little.
[00:20:25] We need to do two things. For one, we can't just render the content straight as it is. We want to wrap the content in some tags, and I'll explain those in a second. So we have a template tag here, which is an HTML element that basically just tells the browser, ignore this, don't actually render this. But it allows us to still send it in the stream. And then we're rendering a custom element, and that custom element just gets the unique ID of the element of the map, the suspended map that we created in the suspense component. All this does is allow us to tie some custom JavaScript to that custom element.
[00:21:14] And this is now the little bit of magic part that we need. We only need this if we have suspended components, so I'm going to put it in here. What this JavaScript basically does is, we define this custom element that we call suspense content, and then we give it a connected callback. This connected callback is called whenever that element is rendered on the page, and then we can do stuff there. What we're doing here is we're selecting the previous sibling, which, because we're controlling all of this, we know it's going to be the template tag. Then we try to find the original element with a target ID, which is that suspense target ID. And then we just swap out the contents. That's it. It's like 10-ish lines of JavaScript.
[00:22:08] And with that, if we refresh the page, you see it still loads immediately with the loading states, but it now loads in the actual contents whenever they're ready, in any order, and it's leading to a much nicer user experience.
[00:22:30] The cool thing about this is all of that is server side. Yeah, we're shipping a little bit of JavaScript, in this case we're shipping it as part of the server HTML response. So if I navigate to a new movie again, the user experience is immediate with loading states and everything. But also, because it's all in that HTML that we send as a stream, when I go back the server still has the whole HTML cached, there's no loading states at all, it's just immediately there. It's all these kind of things where we can take advantage of the platform as we've had it for years now, by just doing it in these new and right ways.
[00:23:11] Like I said, we're in full control of these loading states. So instead of having suspense boundaries around each of the sections, we could have just one big boundary. Again, this really just depends on your application and how much you can actually control. You want to avoid layout shift, but this might be more what you want to be doing in most cases. But again, the idea is, out of order streaming allows us to give those dynamic user experiences with dynamic loading states while still doing most of the work on the server, and that gives us kind of the golden ratio of both the performance of server-side rendering and the dynamic and developer experience of the single page application that we're so used to.
[00:24:01] All right, I am running out of time. So I hope this kind of showed you conceptually where we're heading with web development in the future. I think this is a really exciting space to look for in the next year or two, especially when it comes to frameworks that start abstracting this out and helping us actually implement this without having to manually implement this. These are some resources if you want to look deeper, if I tickled some interest in anyone watching. I found these resources really helpful to not only explain the implementation details but the reasoning behind it, the motivations from a user experience perspective. And yeah, these are links to the slides, to the GitHub repo with the code if anyone is interested, and to my socials. Thank you.
Q&A
[00:24:57] Host: Excellent. Thank you very much, Julian. That was great actually, to do that hands-on coding. I know it's always a little bit scary doing a live event, but it was great to see it kind of just iteratively get better.
[00:25:08] Julian: Yeah, but you could tell that I was kind of cheating out with the copy pasting there. It is scary doing live.
[00:25:16] Host: I'm going to kick off with a question, but if anybody else has questions out there, please do write them in. Why would we go with streaming when we already have, like, optimistic rendering? Maybe you could create a static page for each of those movies with static site generation. Why complicate things with this streaming? Because as you said, it is a bit of effort to kind of be thinking about all the fallback states, all of the extra bits that you need to do.
[00:25:47] Julian: A very good question. You should absolutely not do it if you don't need to. It's basically the first principle of, use the dumbest — dumbest in quotation marks — form and technology that you can for the use case that you have. If those movies don't change and static pre-rendering is an option, you should absolutely do that, because then you're just serving static files, that's always going to be faster and, like you say, it's less complex.
[00:26:13] Julian: So this is really more solving for that use case where you do have dynamic content. Think even like a user dashboard where the data has to be fetched from the database on every request. But it's also sensitive because, you know, you have users on low-end devices and with not so good network connections and that kind of stuff. So this is really trying to solve the problems of those applications with a lot of overhead on the server potentially, where hopefully this direction can solve the three problems of developer experience, user experience and performance overall.
[00:27:02] Julian: What you're mentioning of complexity is still very true, and my hope is that within the next year or two frameworks will have evolved enough that you don't have to think about that complexity anymore. It's basically abstracted away for you, like we've done with single page applications, where you don't do the jQuery messy spaghetti code yourself anymore, you have a framework that essentially hides that.
[00:27:28] Host: Excellent. Well, that's actually really interesting. Somebody mentioned a point that AI won't be able to help with these new things because it's based on old paradigms. So if something new comes along, the AIs won't be very good at writing the code for it.
[00:27:44] Julian: The irony of that is, I kind of implied it with the diagram that I tried to show: a lot of this is inspired by old ideas. We learned that we went too far with single page applications and we're now learning back the best practice principles from the original server-side rendering, etc. So this is actually going to be really interesting to see, like, how much of those old ideas can we apply while still also applying new technologies, new ideas.
[00:28:10] Host: Brilliant. Well, that brings us to time. So thank you very much, Julian. I really enjoyed that.
