Zero Downtime Releases: Migrating Systems To The Cloud
Checking session availability…
Hang tight while we load the latest updates.
Nik talks through the ugly stepsister of the shiny new codebase - monolithic legacy systems. Topics discussed include how to work with an older framework that has a host of technology, resource and attrition issues.
Zero Downtime Releases: Migrating Systems To The Cloud
Nix Crabtree at UXDX Europe. Video: https://youtu.be/kz8YD6K1Yp4
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Legacy and monolithic code bases
[00:00:07] Thanks, Olivia. Okay, let me sit up here. Hopefully not too many sore heads in the house this morning. Who was out partying? Liars, like two people. Thanks for coming in bright and early to see me kick off the execution stage here at UXDX in Dublin. My name is Nix Crabtree. I'm going to be talking about migrating systems to the cloud and zero downtime releases, and when I say that, I think we all know that I mean monolithic legacy systems, which is a problem that we're all going to face at some point or another.
[00:00:57] Let's just talk about terminology briefly. A legacy code base is essentially one to which you are no longer willingly and actively adding features. It's likely to use older frameworks, code styles, architectural patterns, hosting technologies, and is therefore by definition the ugly stepsister of the shiny new code base. Because of that, people are unlikely to want to work on it, and that's going to cause problems with motivation over time. That's likely to cause problems with attrition, and then you can't hire good people to come and work on it, and so on and so forth.
[00:01:40] A monolithic code base is almost certainly a legacy code base. The reverse is not true: a legacy code base doesn't have to be a monolithic code base, but a monolithic code base is probably legacy, simply because we now know that there are better ways of doing it. We may have known then as well, but we don't talk about that. It has to be built and deployed in its entirety, and it's probably composed of tightly coupled, interdependent components which are probably communicating in process. You might find that your presentation and your logic are all mixed up, so you can't make a change to them independently and deploy them that way. When you do deploy it, almost certainly it's going to involve a complex, sequential series of time-consuming or manual steps, and that often means that deploying it means downtime.
Leave it, bin it, or fix it
[00:02:55] What are our options in this case? Well, we could leave it. I think we all know that this is deferring the problem. The cost of maintenance is going to continue to rise, and not just in time and effort, not just in motivation and attrition. You'll find that with old technologies there is a rapidly diminishing pool of resource, because people no longer have the skills or the desire to work with those technologies. At some point you're going to have to do something with it.
[00:03:29] Option two: you could bin it and build a shiny new utopia of new frameworks and new technologies. You can go wild with all this great stuff that you've only read about. This is the stuff that developers' dreams are made of, but probably not commercial success. You can't just throw away the old system and do nothing for a while whilst you play with the new tech. It just seems very unlikely. It would be nice, right?
[00:04:09] That brings us to the third one, which is fix it. When we say fix it, what we really mean is we're going to keep the old system rolling whilst we change the wheels, probably also the engine and the chassis, and respray it, and we're not allowed to pull into motorway services unless we absolutely have to.
[00:04:32] How do we go about fixing the big ball of mud? Essentially, we need to break it up into smaller balls of mud, we need to scoop out the insides of those balls of mud, and then we need to clean up the mess. You could also look at this very much like the plot of the film Invasion of the Body Snatchers, where they send down small spores that grow into exact duplicates of the things that are already there, and then they dispose of the waste afterwards. It occurred to me that this film, from nineteen fifty-something, was definitely defining patterns for modern software architecture. Watch it again; you'll see what I mean.
Guiding principles: test and constrain scope
[00:05:18] Before we go into how we break it up, let's talk about some fundamental guiding principles. We need to understand what we're doing. We need to understand the impact of change, and therefore we don't make a change to anything that is not under test. Why? Because we just don't know if we've broken something until it's too late. It might seem like a good idea, it might seem that we're using the right tools for the job and we've done everything properly, but in actual fact, when it's too late, the whole thing comes off the rails. Breaking stuff means downtime, and again, that's not something that we want.
[00:06:08] Tests ideally assert the behavior of the system when compared to the requirements of the system. That's not always the case with a legacy system. You may have so few tests that you've no idea what the system really is supposed to do. You may have so many tests that you've absolutely no idea what the system is supposed to do. What we can do in this case is use something like characterization tests, where we retrospectively write automated tests that say, this is how the system currently behaves. That gives us a baseline to say, when we change something, we know we haven't broken it. The current behavior of the system might not be the desired state, but that must be our minimum baseline of what we're going to go with in the new system. We can't break stuff and make it worse.
[00:07:11] Ideally we'd have a properly shaped test pyramid and not some kind of abstract cubist nightmare, but that has to be balanced with the benefits, when you think about the effort it takes to build a proper test pyramid for a legacy system that ultimately you're moving away from. If you understand the behavior of the system right now, you probably have enough to move forward.
[00:07:45] Constrain your scope. This is really important when you're moving, especially from a monolithic system to a shiny new system. You can imagine that business and tech alike are likely to see this as an opportunity to shoehorn in everything on their Christmas list from this year, last year, maybe the year before. They've all been waiting for this moment, but it's pretty much a direct path to failure.
[00:08:15] We need to constrain both our new features and our emergent design. We need to make some big decisions up front about how we're going to go about it and what we're going to do, and we need to stick to them. We need to make some big decisions about our big-box architecture up front, and we need to stick to them. We need to lay out an appropriate but limited set of technologies and build the new system on them, and then we need to give the teams the autonomy to go and make something they're super proud of, but within those constraints. After that, you can start talking about continuously improving and adapting. You can change those small balls of mud into completely different languages, frameworks and whatever you want; they don't have to be in any way related. But during the carving up of the system, it's really important to keep that scope as tight as you possibly can.
Automate your pipeline
[00:09:27] Automate your pipeline. This man in this film embodies the automated pipeline in all its glory. This is Groundhog Day. This man wakes up every day and it's the same day, but it's not exactly the same day. Every day he makes a minor incremental improvement to the way the day goes, and over time the day becomes absolutely slick from start to finish. He knows where stuff's coming from, he doesn't even have to think about it, he just sails through the day. To me, that's what we want when we build an automated pipeline. We want it to just happen, we want it to be slick. We want to commit some code and we want it to go through to production, or maybe what we want is for it to go through to a state where it's ready for production. But ultimately we don't want to have to do that stuff. We want it to be as polished as it possibly can be.
[00:10:25] Automated tests are where we want to be. If you saw Scott's talk yesterday about the benefits of manual tests and the human element, that is absolutely a good way to get going. When we work in an agile fashion, we try to avoid doing big stuff up front, and the same is true for an automated pipeline and a set of automated tests. If we can get going, and we can do so with manual tests, then we absolutely should. We should build that pipeline as we go. If we can use exploratory tests to determine the behavior of the current system, then we should absolutely do that. Get going quickly, release in small increments: classic agile, XP-type stuff.
[00:11:09] An automated build gives us that repeatability and consistency every time we build. "It builds on my box" cannot be a thing, should not be a thing in this day and age, but we all know that it is. Automated deployment opens up our ability to run automated functional tests. At that point we can automate our tests of the behavior of the system. If we've written those characterization tests, that's a perfect time to do it, so we know that we haven't broken anything compared with the old system.
[00:11:45] Automated provisioning opens up our ability to spin up and also to spin down environments in a known state every time. You have no configuration drift, no patch drift. You don't have someone going onto the box and tweaking something, and now it works, brilliant, but when you move to another environment it doesn't, because no one remembers what changed. Automated provisioning, especially if you're able to do this in a cloud provider, also gives you the benefit of cost saving. You spin down those environments at night, you spin down those environments at the weekend, and the amount of money you save even doing that in a large environment is surprising. Motivating, let's call it.
[00:12:36] All of this is ultimately what we all call our CI/CD pipeline. It's really important to get this stuff in up front, or at least plan to do it, and plan to increment over it such that your feedback cycles get shorter and shorter as you go. You want to relentlessly shorten your feedback cycles, such that you can get change into production and through tests as quickly as possible.
Collaborate
[00:13:06] Collaborate. This is super important. Collaboration is your key to success in any big replatforming [?] or migration exercise. If you're breaking down that big ball of mud into smaller balls of mud, and those teams are going to start to have autonomy, you need to be able to talk to each other, you need to be able to exchange ideas. If you're having meetings, you're probably not collaborating well.
[00:13:36] It's like the difference between me cooking dinner, realizing to my horror that we have no lemon-soaked paper napkins in the house, so I just get on the phone and go, "Can you pick up some lemon-soaked paper napkins on your way home, Dad?" Brilliant. Versus me setting up a meeting to talk about the stocking and ordering times of lemon-soaked paper napkins, which means that I'm still not going to have one by the time I need it. It's not collaboration. If you're setting up meetings, it's probably too late. Go and talk to people. People like to be talked to. When someone comes and asks you a question, it's a sense of achievement to be able to help someone and answer that question. Share your knowledge. It should be that way.
Breaking it up with bounded contexts
[00:14:30] Let's talk a little bit about how we might break it up. Bounded context is something that was introduced to us in domain-driven design. A bounded context represents a purpose within your domain. Back in the day, we all thought it was a great idea to have a universal, unified, global data catalog or global data model that modeled the entire business. But anyone who worked with one of those, and I'm sure many of you have, will realize that it was a nightmare to set up, a nightmare to maintain, and a nightmare to implement.
[00:15:14] The idea of bounded context was that you can have the same thing in multiple smaller contexts within your business, and although it meant the same thing, it didn't have to be implemented in the same way. You would therefore provide mapping between the two, and each bounded context was master of its own destiny while still working as a cohesive whole. We use well-defined interfaces to make sure that other people know that when we talk about a person, it's the same kind of person they're talking about, and how we can map between the two, that kind of thing. When we separate those models into our bounded contexts, it also facilitates the separation of the classes, the components and so on that implement them, and that means that we then have control over the code in our bounded context, to do with as we please.
[00:16:13] Let's have a quick look at a bounded context we might have, for example, at a large online fashion retailer. A product, a customer, a payment: this is essentially a checkout process. You put a product in your bag as a customer and you pay for it. Cool. We also have visual merchandising, the people who make our products look nice on the site. They put them in the right place, just as when you walk into a shop, the thing that you want is right at the back, so that you have to walk past every other item to go and get it, so that you might buy more stuff from them. A product is slightly different for them. They don't necessarily care directly about the stock of that product or the price of that product. That might influence the visual merchandising, don't get me wrong, but they're not going to use those precise values at the product level as you would in the checkout.
[00:17:16] Identity, our ability for a customer to log on to our website and identify themselves, cares about the customer, but cares nothing about payments or the products or anything else. Product, customer and payment is also the bounded context of refunds, for example. Different things, completely different actors, completely different ends of the business, working in reverse, but a different bounded context.
[00:17:46] Once we start to draw these lines around the bounded contexts, it allows us to start to break up our big ball of mud into logical parts. We can then go from there to the more physical part, for example a microservices architecture. It doesn't have to be a microservices architecture, it could be any architecture, but ideally what we want is to be able to encapsulate parts of our business. That includes the way that we expose it to external consumers, which might be a web front end, or another API, or it might be directly to the customer. We want to be able to encapsulate the logic, the models, the data storage, the logging, monitoring, security, and we want to be able to release that stuff at will, independently of all the other small balls of mud, which hopefully by this time are no longer balls of mud but shiny, beautiful balls of shiny. We can take these small components and they all work together beautifully, if anyone remembers Terrahawks. If you don't, I'm sorry for you.
Strangler pattern and the anti-corruption layer
[00:19:10] Let's talk about some patterns that we might use to break down our monolith. The strangler pattern is based on the strangler fig, a plant which essentially grows from the top of a tree down into the ground, roots itself in, and then over time gradually kills the tree that's hosting it and takes over. Nasty. But essentially the strangler pattern approach is to add new stuff in the new way, and then the old part of the application essentially just dies. Over time your new application parts become the application whole and the old one dies off. Fairly high level, but it's probably easier than trying to rewrite in situ the older parts of the application.
[00:20:14] You might want to have an anti-corruption layer. An anti-corruption layer is just a bunch of stuff that stops your old, nasty monolith from contaminating or corrupting your new stuff. Your new stuff will likely use a set of interfaces defined between tech and business to do what you need to do. They're probably lightweight, they're probably arranged in such a way that you can do lightweight messaging, event sourcing, whatever it might be. Your old system is probably going to return you some gargantuan XML document that you then have to filter and parse and everything else. Chances are you're still going to need that in this interim stage. What you do is you put in a layer of components that take on board that translation, that mapping, whatever it might be, so your new system essentially has no idea what the old system looks like, which is a good thing.
[00:21:28] There are a few patterns that we can use within our anti-corruption layer. One is the adapter pattern. The adapter pattern essentially provides a new interface over an old interface. Quite simply, you can take one of your old classes in your old system, you have an adapter, and the new system can use that class as if it were a new class. The adapter basically says, "Yep, I'm one of these shiny new ones," and then goes off and does all the nasty mapping and whatever else it does at the back end.
[00:22:18] There's also the facade. The facade pattern is essentially presenting a subset of a larger interface. That might help you to start breaking things down, especially if you're moving into something like a CQRS architecture. You want to break up your commands and your queries, and you may use a facade to essentially pretend to be a command or a query, even though in the interim they're using the same monolith behind them. There's also a translator, which is essentially an adapter but for messaging. It will do some mapping. It will take in a message and say, "Yeah, I understand that, that's the nice new one," and then map it to the old one and pass it on.
Cleaning up the mess
[00:23:09] Then, once we have scooped out the insides of our little balls of mud, it all looks the same. We've taken the big ball of mud, we still have lots of small balls of mud, and our system is doing the same thing. It's performing in the same way that it was before in terms of behavior; hopefully not in terms of actual non-functional requirements. Hopefully it's faster, it's more easily monitorable, but it looks the same to external consumers. At that point we want to clear up the mess. We want to hose down our small balls of mud to reveal the new balls of shiny underneath.
[00:23:53] We want to decommission the strangled components. For example, if we've moved over to new components, the old components may still be there. We don't want them. We don't want them taking up money, we don't want them taking up support costs and time and so on and so forth. Cleaning them up is really important, and you need to be aggressive about that. With an old system, a system that maybe nobody fully understands, you will have those moments where you say, maybe I'll just leave it, because someone else might be using it. That's just not the way forward. You need to do this stuff. You need planning, you need courage, and you need to do this stuff aggressively. Pare down your system: once you've moved over some functionality, take out the old stuff.
[00:24:40] You would want to collapse your anti-corruption layer. If you've got an adapter that's pretending to be the old interface, hopefully by this point you no longer need it. The old system is gone, so take out those adapters, those interstitial components. Simplify your architecture. Simplify your routing. You may have been using some kind of blue-green deployment, which we'll talk about in a sec. Simplify your routing, your network layers, your calls, all that kind of stuff.
[00:25:19] Then you should hopefully be in a position where you can deploy at will. You can get features in front of your customers much, much faster, much more rapidly, and the really important thing is that you are confident, because of your automated pipeline and because of the smaller scope of your change, that when it goes into production it's going to work.
Methods for zero downtime
[00:25:43] Whilst we do all of this stuff, and beyond, there are some methods we can use to ensure that we have as little downtime as possible, ideally zero downtime. Immutable infrastructure: we talked about it. When you deploy, you provision, and that means that there's no configuration drift, there's no patch drift. Telemetry, telemetry, telemetry, telemetry, telemetry. You need to know what your system is doing at all times: logging, monitoring, alerting, anomaly detection. If your system goes wrong, if something happens and someone says to you, "What's wrong with it?", you need to have an answer. You need to have an answer, because if you don't, it's just sitting there in production and it's not working.
[00:26:38] Security. Don't underestimate security. If you were to put a new virtual machine on a public-facing IP, in 90 minutes the thing would be broken. It would be attacked, it would be starting to explore other virtual machines on your network, it would start to be finding vulnerabilities. It's a serious problem, especially when you're working at scale, especially when you're a big name, and even if you're not, even if it's your personal virtual machine you've put up there in your personal developer subscription. There are vulnerabilities, and there are attackers who are constantly scanning this stuff. If you've got a security breach, it's not only your downtime that's a problem; potentially it's your reputation as well.
[00:27:38] Blue-green deployments. Blue-green deployment is essentially taking two sets of architecture, deploying your new stuff to one, and then gradually feeding traffic over to it. In Microsoft Azure you've got Traffic Manager, which allows you to use a weighted round robin, or you can direct traffic specifically to different endpoints. You can do various things in Akamai, for example. When you're happy that it's working, you start to feed more traffic over, and then the environment that you've left behind is essentially your staging kit, and so you keep swapping.
[00:28:19] Canary deployments are a similar thing, but instead of separate architecture, it's a subset of your architecture. You might deploy to one or two servers, or you might deploy a few services, and again you use some kind of weighted DNS round robin, whatever it might be, to start feeding traffic over. It allows you to see the thing scale up, so you're not only testing that the code you've written works and that the infrastructure works, but also that it can scale properly. If it's not working, you simply switch off that traffic and go back.
[00:28:53] Then feature toggles, obviously. At the code level you can say, if you pass in this value, then you'll get this feature or you won't. Or you can use API versioning, resource versioning, all of these kinds of things, to limit the access to your new stuff until you're happy that it works and you can roll out that full release to all of your customers. Thank you very much.
[00:29:29] Host: Thank you, Nix. Thank you very much, that was really interesting.
