Checking session availability…
Hang tight while we load the latest updates.
During the past year we've explored a new approach to Minecraft release management, with concepts such as Minimum Lovable Product and four levels of Done.
In this talk I'll show how it works and what we've learned.
How We Manage Releases at Minecraft
Henrik Kniberg at UXDX Europe. Video: https://youtu.be/-5m2MHhtM3Q
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Minecraft and how it ships
[00:00:00] Hello, I'm Henrik. I work on the Minecraft gameplay design team and do a mix of gameplay design, feature development and team coaching. I'm going to talk a bit today about how we manage releases and a little bit about how we do design.
[00:00:17] First, a bit of context. Minecraft is a game. It's been around for 11 years or so, and it's been growing pretty much continuously ever since the start. Now we have more than 120 million active players, and they are of pretty much all ages. This game is actually two games, because technically there are two codebases behind it. There's the Java Edition of the game, which runs on Windows, Mac and Linux, and it's built in, well, Java. And then there is the Bedrock Edition of the game, which runs on Windows, mobile and consoles. These two are trying to pretend to be one game. We want it to feel like one single game, and we want as much feature parity as possible. But for historical reasons, it's two codebases.
[00:01:08] When we release the game, we tend to make larger releases about once or twice per year, and we used to give them a theme of some sort. For example, Update Aquatic was all about improving the oceans and making them more fun and interesting. In 2019 we did the Village and Pillage update, which was all about making villages more fun. And then the latest release, the Nether Update, was all about making the Nether dimension more fun. We tend to have this high-level focus and improve one aspect of the game for each release.
[00:01:40] However, we don't ship them as big-bang releases. Instead, we deliver small increments that we call snapshots or betas. We have three teams and we work in two-week sprints, and then every week we ship a snapshot or a beta. A snapshot is pretty much the latest and greatest. It's whatever we've built up until now. We ship it, and players can opt into that. What those players get is of course the latest, coolest stuff, but it might be a bit unstable and it's definitely not complete.
[00:02:13] It's a bit of a partnership there, because these players who are playing on the snapshot, and that's quite a lot of players, give us really useful feedback, and then we adapt to that feedback and improve the product. And of course, by the time we get to the actual release, we typically know that it's going to be a success, because we've already had so many chances to improve it. We aim for the clouds, we want to make a big impact with the game, but we ship in small increments to make sure we're always learning.
Our challenges
[00:02:38] Just like any organization, we have a lot of challenges, and here are just some of them. Scope management: we have fixed-ish dates with our big releases, and we want of course to maximize the amount of fun that we can put inside that release, so how do we manage that in practice? Dependencies between features: dependencies are sometimes pitched as something bad for development, and you want to minimize dependencies. But in our case dependencies are good. We thrive on dependencies, because you can have three different features. Let's say a new item, and a new mob (a mob is just game speak for a creature), and a new structure, like a hidden temple or something. If you have to go to the mob and trade with them to get the item, and that item unlocks the hidden temple, now one plus one plus one equals ten, because these three features together form a system, which makes the game really fun. We like dependencies, but of course we have to manage them in some way.
[00:03:34] Java and Bedrock parity is a challenge, because we want to ship on the same date on both platforms, and they should contain the same features, and they should work in the same way. And that's hard, because some features are easier to build on one platform and hard to build on another, so how do we manage that in practice?
[00:03:45] Then there's the thing about our players. Minecraft is more than just a game. It's almost like a game engine or a platform, because you have players building mods on top of it or building minigames. We have different play styles too. We have people who are adventuring in Minecraft, building fantastic architectural masterpieces, or just socializing, or perhaps building massive machines. People play in very different ways, and we want to cater to everybody and make sure we don't leave someone out. These different players all love Minecraft, but for different reasons. And of course that leads to a lot of change. As we get feedback from these different types of players, we want to adapt to that feedback. And how do we handle that, and how do we stay aligned in the face of all this change when we've got multiple teams? We have a lot of challenges, and I could go on and list more, of course.
Finding the awesome
[00:04:38] But I'd like to start talking about design and share some of my thinking around that. Ideas are never the bottleneck. We always have an infinite number of ideas. And we have this release, which is like a bucket, and we can only fit so much. The challenge is, out of all these billions of ideas, what are we going to put in the bucket, right? Because some ideas are better than others, some are more costly than others, and we have to figure this out.
[00:05:00] The challenge is really figuring out what not to build. I used to work at Lego, and they refer to their design process as an idea-killing process. Because again, ideas are free. They come from all over the place. The challenge is to figure out what not to build, or at least what not to build right now. Choose wisely.
[00:05:23] How do we do this? How do we find the awesome? We want to build features that are fun, but also technically feasible, and ideally that fit the theme as well. In the perfect case, we find the diamonds that fit all three. But we're okay with gold-like features that are fun and technically feasible but maybe don't fit the theme. We at least try to find features that, for the most part, match the high-level theme of that update.
[00:05:49] As a designer, we need to wear two hats, both gameplay and tech. In some organizations these are different roles, but in our organization, as a designer, I'm both designing and coding. I'm building the feature that I designed. Of course in collaboration with other people, but there's no handoff. It's not like, now you take over and now you build this feature properly. I think that's really useful, because it means that as a designer I have to take both aspects into account.
[00:06:13] Ideally, gameplay should not be limited by tech. We should adapt the technology to all the great ideas we have, right? Technology should serve our vision. But we also want to be smart. When we are prototyping, we are also looking for, "Oh, what's hard to do? What's easy to do?" And then we adapt the design to that, so we go both ways. And that's what lets us find great design, which is features that are really fun and aren't unnecessarily hard to build.
[00:06:46] If we take those two aspects and put them on a graph, we can put gameplay value vertically, and technical cost, risk and uncertainty horizontally. Then we have this line, and of course we want to be above the line. We want to find the features that live up to the left. Features down here, down to the left, could be useful. This could be low-hanging fruit, some little thing that is easy to do but that's going to make at least some player segment really happy. It might still be worth doing. Low gameplay value doesn't mean it's not fun. It could also mean that it's really fun, but for one small player segment.
[00:07:21] Similarly, sometimes it's worth making a major technical investment if the gameplay is that awesome, right? That's fine. But what we try to avoid is this land of meh down here, right? A meh feature like this one could be great, maybe players like it, but it just wasn't quite worth the cost. We could have spent that time building better features.
[00:07:45] When we are brainstorming a release and the features that are going to be in it, we have this very vague, high-level sense of where each feature is on this spectrum. Feature A might be up here, Feature B down there, and Feature C over there. If we take the bucket into account, then Feature A we're going to keep, Feature B of course we won't build, and Feature C, well, maybe, if it can fit in the bucket. We'll see. This is a tool for us to prioritize.
Exploring the design space with prototypes
[00:08:17] In practice, if you look at that journey, a lot of it is driven by prototyping. Prototypes are really useful, because when I have to make my feature actually run in the code itself (I'm not talking about paper prototypes, I'm talking about actual running code), that forces me to think about both the gameplay and the tech. Even though I'm taking all kinds of shortcuts and doing hack fixes because it's just a prototype, even so I am getting a sense of, "How hard is this to do properly? What is it going to take?" Again, two hats, right? Gameplay and tech, and prototypes let us wear both.
[00:08:57] Let's say I'm building a prototype for something. Let's say I want to build "Ride a Dragon", a new feature where you can ride dragons in Minecraft. Super hypothetical. I have an idea of how that might work, and I code a prototype for it. And here's my hypothesis. As part of that, maybe I'll learn that, "Oh, this was really tricky, especially the part of a player controlling a flying entity. That's really complicated." Then maybe I can make some compromise and say, "Well, what if the player can't control the dragon? I can sit on it, but the dragon decides where to go. Maybe that'll be fun too, and a lot easier technically."
[00:09:35] I make a new prototype, and that's my hypothesis: that it's going to be easier to do and still fun. And maybe that hypothesis gets mostly validated. Okay, this is promising. Now I might feel that it's time to put it inside a snapshot, and that means giving it to the players. And when I give it to the players, I will almost inevitably be surprised, because they try stuff and use it in different ways than I could have imagined. Maybe they set up automated transport systems using dragons, and I had no idea that was even possible. Put stuff in the hands of players, get surprised, learn from it, and almost inevitably get ideas for how to improve the feature.
[00:10:12] I have a hypothesis that if I make this small change, maybe I make it possible for dragons to carry items effectively. I can put bags on them or something, I don't know. A sled, like a Santa Claus dragon, I don't know. And my hypothesis is that this can be done without a lot more technology, so then I build another snapshot. Of course, now I'm not prototyping anymore, I'm writing production code, because snapshots are production code. They're just not polished and complete. And then I release that, et cetera. This is what I mean by exploring the design space of a feature. We're bouncing around here and trying to find the awesome.
[00:10:48] And sometimes it leads in another direction. Sometimes I'll be hypothesizing and prototyping, and then finally I'm like, "You know what? It's going to end up there." It's going to end up in the graveyard of darlings. Kill your darlings. And it's not alone. There's a bunch of darlings in that graveyard, features that once felt promising and didn't quite pan out. Maybe it's an okay feature, but there are other features that deserve that spot inside the bucket more than this one.
[00:11:18] As opposed to reality, these darlings are not permanently dead. They're like zombies. They may come up for some future release: "Henrik, you buried me last release and now I'm back again. Please put me in the next release." And I might be like, "Hmm. You know, maybe... eh, maybe not." Then I'll kill it again and put it back in the ground. But they're not permanently dead. Ideas can come back again, and that's perfectly fine. It can go both ways.
The Vanilla Dashboard
[00:11:48] All right, I'd like to share a little bit about the practicalities of how we manage this way of working as a team, or as three teams. The thing is, whenever you have multiple teams, it's very easy to get misaligned. What I mean by that is, here this team is trying to solve the problem of "we need to cross the river". They're building a bridge, because that's their hypothesis for a good solution. But what they don't know is that there's another team on the floor below, and they're also trying to solve the problem of "we need to cross the river", but their solution is different. And then maybe they don't match each other, and now we have a problem. We have misalignment and frustration and waste. This tends to happen quite a lot when you have many teams trying to work together.
[00:12:33] And there is a somewhat universal remedy. I find it works most of the time, which is to make everything really visual. Simple things like this: To Do, Doing, Done. Some would call that a portfolio board: features, priorities, which team is working on which, et cetera. In most cases this layout would be sufficient. However, this didn't quite suit us, because what does Done mean? Is it done when we've put it in a snapshot and started getting player feedback? Or is it only done when it's releasable? And what about the two editions? What if Java is done and Bedrock is not? It's hard for us to get an overview if we simplify it down to just one Done.
[00:13:20] We've introduced a model which we call the Vanilla Dashboard. Vanilla is just an internal term we use to refer to the game itself, Minecraft. It's the same high-level concept as what I just showed, but more adapted to our needs. This is what it looks like. There's my team lead, by the way. She is awesome. If you were to come to our office, before COVID shut it down, you would see this massive thing on the wall. In fact, this is only half of it. The other half is behind there. A huge, massive visualization of what's going on. And you'll see people gathering in front of that board, having animated discussions and making decisions and things like that. And you'll see feature cards, and weird concepts like a minimum lovable release, et cetera. I'm going to go through what this is, how it works and why we made the system.
[00:14:17] The purpose of the Vanilla Dashboard, and it is both a tool and a process, is to help us stay aligned. What it gives us is less stress, because people get stressed when they don't know what's going on, and that causes sub-optimization. It also gives us a better game, because if we are aligned we can move faster and make better decisions. And it does that by giving us realistic expectations. We can see what's going on, how fast we're moving. It triggers discussions and it reveals problems. Visualizations like this don't solve problems, but by visualizing them really clearly, they help us detect problems early, and then we can often fix them before it's too late. And it makes change easier, because if the plan and the status are super visible, it's easier to walk up to it and move something and change it because we learned something new.
[00:15:06] There it is in its full glory. And again, this is only half of it. Then COVID came along and of course we all went home. We became a distributed team with no access to the physical wall anymore, so of course we digitized it. Here's the digital version of our board. It looks pretty much the same. We use a really great tool called Miro to make this board.
Features, the minimum lovable product and the minimum lovable feature
[00:15:28] All right, but how does it work? What do the things mean? Let's decode it. At the top you see features. Each card is a feature, and they are prioritized from left to right, so leftmost is the highest priority. On the card we have the name of the feature, some concept art, and then the "Why". Why are we building this feature? In this case, we're building the respawn anchor because we want people to be able to live in another dimension. It's really important to keep that Why, because as soon as we start forgetting the Why of a feature, that's when we start falling into mediocre design. It's really important to always remember the Why.
[00:16:06] Each column is a feature, and they go from left to right. And they are grouped into what we call the MLP, the Minimum Lovable Product. There's a red line hanging on the wall, and everything to the left of it is included in the MLP. It means pretty much what it sounds like. It's the minimum content of the bucket. Without these things, we don't really have a release that we could stand for. We don't call it minimum viable, because we don't want minimum viable, we want minimum lovable: something we can stand for. If we were to ship this and get these things in and nothing else, we would still be proud of the release.
[00:16:44] However, we are not aiming to release the MLP. We're aiming to release more. This is really a minimum. Therefore, the MLP should be small enough that our gut feel is that it's ridiculously small: we can finish this in half the time. Obviously, we're going to be wrong. Things always take longer than you think, right? But by really, really, really making it minimum, we increase the likelihood of finishing the whole MLP plus more. And so far we've managed to do that.
[00:17:10] The guiding question for deciding if a feature is part of the MLP or not is: would we delay the release for this feature? If this feature is not done by the release, would I rather move the release date to get this feature in? If the answer is yes, then it's MLP. If no, then maybe it's beyond MLP. Again, we aim to ship more than the MLP, but we don't bother trying to predict how much more. Instead we just say, "At least this, and then as much as possible of this." And because we humans are seemingly incapable of making estimates that are correct, we don't bother trying to estimate how much of that we'll finish. Instead, we'll do as many as we can, and we'll see how many we end up with.
[00:17:52] Furthermore, each feature itself has what we call an MLF, a Minimum Lovable Feature. It's a subset of the feature that would make that feature lovable. It's like the MLP, but within one feature. Example: here's the piglin mob. The MLF for us was that you can barter with them, they can be distracted by gold, they hunt hoglins, and they attack players who aren't wearing gold. For various reasons, that's what we considered to be the minimum for them. If we took any of this out, the piglins would lose their core purpose in the game.
[00:18:28] The guiding question when trying to decide if something is MLF or not is, "Would we risk the next feature for this sub-feature?" Let's say, for example, baby piglins riding baby hoglins. That's a really nice thing. But if push comes to shove, I would rather we ship the next feature and not have baby piglins riding hoglins than risk delaying the next feature. With that in mind, we would put that kind of thing on a separate note called, let's say, Piglin Extras or whatever. Those can be put on the wall as well, but they go after the MLP, and then they get prioritized in comparison to other things that are also outside the MLP. This is how we manage scope creep. Scope creep can be okay, but it should be deliberate. We shouldn't be swelling the minimum lovable product side of it. We should have that after.
Four levels of done
[00:19:19] All right, back to the concept of Done. I mentioned that we didn't want to have just one definition of Done. Instead, we split it into four levels of doneness. The first level is Designed. When a feature is designed, it means we know what the MLF is. We've decided through prototyping, we've concluded, that this is what makes the feature awesome. These things would be great, but they can be done later. We also have a sense of the technical constraints or needs for this feature, and we've synchronized across both platforms: "Okay, is this realistic? Can we do this?" It's not a frozen design specification. It's not like we're ever completely done with design. But when we say that design is done, it really just means done enough for now, so we can move on. And Designed is shared across both platforms, because we design for both platforms.
[00:20:09] But the next three levels of Done are tracked separately for each platform, because we've got to do them on both codebases, and they might be at different states. Runnable means we have something running. I can look at it, I can play this feature. It might not be stable yet, it might not be complete, but I can essentially play the feature. We actually renamed that to Playable recently. Snapshotable means the feature is in a state where we could put it in a snapshot. When a feature is in that state, it essentially means it is in the hands of our players, or will be within a week. That's a really important milestone, because that's when we start getting real feedback.
[00:20:45] And Releasable, of course, means: if the release date were tomorrow, would I be okay with this feature being included in the actual final release? If the answer is yes, then okay, it's releasable. This is how a feature goes through these different stages, in a sense. And they're not super distinct. Even during the Releasable stage, we might go back and make improvements.
[00:21:03] A blue note is work in progress. In this case, we've already snapshotted the Crimson Forest, but we are in progress trying to make it releasable: fixing the last critical bugs, tweaking things. A green note means done, or done enough for now. And again, it doesn't mean done forever. And sometimes we annotate to mark problems: I'm blocked here, I need something from somebody, et cetera. That's one of the purposes of these visualizations: to make problems visible, so we can address them and do something about it.
[00:21:40] We do a little bit of estimation. We don't spend a lot of time on it at all. But sometimes it's useful to mark a feature as large, or maybe small or medium, just to get a sense of, "Okay, this is going to take a lot more time than that one." It helps with planning and prioritization, but our estimation is very, very lightweight.
How the green flows
[00:21:59] The nice thing about this is that you can take a step back, look at the board, and just look at the green. As it fills up with green, that is how we get a sense of progress. And we can also talk about, "Okay, where do we want to go next? Where do we want to see the green flow?" How does the green flow: is it from the bottom up, or from left to right? What's the sequence of things?
[00:22:19] In some ideal, lean-perfect, theoretical world, maybe we would do Feature 1 from beginning to end, ship it, and then do Feature 2. One thing at a time. In the lean world this is called one-piece flow. This could be useful if you're on an assembly line or something, or maybe if you're implementing support tickets and each feature is independent of the one before. Then this might be great. But for our type of work, this would not work. If we ship Feature A and get it all done and then design Feature B, we can't let the design of Feature B impact the design of Feature A. It's too late. We want this systemic thinking of how our features relate to each other. This will not really work for us.
[00:23:09] Instead, our world is a little messier, deliberately. We typically design a number of features in parallel, to prototype them together and see how they fit together and how they support each other. Then we move towards snapshotting maybe the first one, while we are making the other ones playable. Then maybe we learn something from that, and maybe we change the design of Feature C or D. Then maybe when we're snapshotting Features B, C and D, we start designing Features F and G, and let those impact each other. Oh, and we came up with an idea for how to improve the design of E. We're going a little bit back and forth, but in general we are trying to prioritize the first feature as much as possible, while still allowing some level of parallel design so that we can have features build on each other.
The weekly dashboard review
[00:24:03] A visualization like this tends to go stale quite quickly unless you have rituals in front of the board, where people look at the board and talk about it and update it. For us that's called the Dashboard Review, and we do it every week. It typically takes about half an hour. It's an open format; anyone is allowed to join. We normally get a fairly big crowd. It varies, but it's opt-in. Show up if you want to know what's going on, or if you want to influence what's going on. And most people do want to know what's going on and do want to influence it, so we get people showing up even though it's an optional meeting.
[00:24:33] That's where we go through what's going on. Is this view correct? Are there any problems people are having? Is anybody waiting for someone else? Do we have any major decisions or pivots we need to make? It's a really important meeting, and it happens every week, so we get a fairly good pulse, so to speak.
[00:24:52] To the left you see the Feature Dashboard. That's the board I showed you before with all the features. It focuses on the player-facing things, so it's a user perspective. The board at the back is the Tech Dashboard. It follows the same format, but the stuff up there is technical things and enabling technologies that we need in order to ship the gameplay features. Both are equally important, so we put them both up on the wall near each other so that we can make sure they're aligned. The MLP of technical infrastructure should align with the MLP of gameplay features, and sometimes we need to make trade-offs.
[00:25:29] Another purpose of the meeting is to get a sense of progress. How are we doing? Are we on track? Are we going to get the minimum lovable product done in time for the release? If not, what are we going to do about it? To get a realistic picture of how we're doing, we can do some simple math. We just count the green stickies every week, put it on a graph and do the math. A feature is essentially seven green sticky notes that need to get in there. That's how many sticky notes fit in one column. We call those feature points, and we can just count how many we get done per week. Sometimes it's more, sometimes less, but on average that line is fairly smooth. And we can do some basic forecasting: "Okay, there's the date, and here's how fast we're getting stuff done, and how are we doing?"
[00:26:14] That's a bit of data, but subjective data is also really important. Sometimes, maybe once per month or so, we do a quick survey and ask people, "How does it feel? What is your gut feeling looking at our current set of committed features?" And committed features, by the way: the minimum lovable product is by definition already committed. But as we get those done, we commit to more and more features. And committing means saying, "You know what? I think this is going to get in. Let's assume this is going to get in."
[00:26:41] The question then is, looking at the features that are committed, do we buy this? Does this feel realistic? Quick survey, scale of one to five, put it on a chart, and that gives us a signal. The goal is that it should look something like this. It should be fours and fives, because otherwise the plan is probably unrealistic and we should make adaptations. This is a warning system. It tells us if we're in trouble. And again, if we notice problems early, we can almost always adapt and make the most of it. Nowadays we tend to do this digitally, using a tool called Mentimeter. Just pull up your phone and punch in a number from one to five, and boom, we have this graph.
Key points
[00:27:24] All right, let's wrap this up. What are the key points that I'm trying to get across here? Product development is, to me, all about experimenting and learning. We are in the realm of complexity. We don't know exactly what our customers need or how hard things are to build. We need to see it as a journey of exploration.
[00:27:46] And transparency, I find, is fantastic. It makes everybody smarter somehow, and happier. Because if people know what's going on, they relax a little bit and they get happier, they get more focused, there's less frustration, less gossip, less confusion. Transparency is super useful, as long as the stuff you visualize is the stuff that matters, and that can take a bit of experimentation to find. Visualize the stuff that matters, and then people will get engaged. It can take a while to find exactly the right stuff to visualize and the right way to visualize it. What I showed you was the result of months of experimentation to find our way of visualizing what we're doing. And we still change it fairly regularly. I would say maybe every month or two we make some tweak to how we visualize stuff.
[00:28:29] It's domain specific, so you probably won't want to use our exact system, but some parts of it might be useful to you. Maybe especially, I would say, the MLP and MLF concepts. I think those concepts are not in any way specific to our domain. They're just generally useful. I can recommend trying that: minimum lovable product, minimum lovable feature.
[00:28:52] And finally, I want to emphasize the importance of shipping early and often. In our case, that means shipping stuff every week, getting real feedback from real players, all the way to real developers, with no intermediaries. That fast feedback, I think, is crucial to the success of this game. Without that real feedback, all these sticky notes on the walls would just be one big illusion of control, and we would actually be flying blind. Regardless of what you do to visualize what's going on, make sure there's some way of actually shipping stuff and getting real feedback. All right, that was it. Thanks for listening.

