Continuous Delivery At Scale
Checking session availability…
Hang tight while we load the latest updates.
Haroon will share the story of how Glovo's engineering culture and mindset evolved over time with the introduction of continuous delivery practices that enabled the team to adopt zero downtime continuous deployment. Follow the journey from monolith architecture to now growing microservices and the benefits.
Continuous Delivery At Scale
Haroon Rasheed at UXDX Community: Spain. Video: https://youtu.be/oeAWb1UXX_A
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Introduction and Glovo
[00:00:04] Hello everyone, and thank you, UXDX, for providing this opportunity to present a case study that we went through in implementing continuous delivery at scale. The high-level agenda today is: we'll introduce the problem that we were trying to solve and why we chose continuous delivery for that specific problem, what the process was that we went through as a team, and then what the outcomes are and where we are heading next.
[00:00:39] Let me introduce myself. My name is Haroon Rasheed. I've joined recently as an engineering manager for Glovo. I'm new to Barcelona, and I'm looking forward to exploring the city; for now we haven't been able to because of the COVID situation. For the last few years my focus has been around promoting and adopting continuous delivery and engineering productivity practices in different organizations.
[00:01:09] Something about Glovo as an organization: our vision is to create a super app that makes everything in your city accessible to you at your fingertips. We have two mobile apps, one for Android and one for iOS, and we are a three-sided marketplace. Obviously we have users and customers. We have Glovers, who are our couriers that deliver, with the flexibility to choose their working hours on our platform and work on their own schedules. And we have partners: we enable partners to connect with new and existing customers, bringing incremental revenue to their business.
[00:01:51] You can see the growth that we went through in the last 18 months or so, and that's phenomenal. Last year we were opening cities at least every four days; we were opening up a new city. Two years ago we were operating in around three countries and eight cities, and now we are operating in around 22 countries and more than 300 cities.
The monolith problem
[00:02:16] The problem we had was around one of the key components in our technology stack, and we call it the monolith. The monolith is the back-end service that powers whatever we are doing at Glovo. To make you understand how critical this system is for us: if the monolith goes down, then Glovo stops. Our customers cannot make orders. Our couriers cannot see what orders they need to deliver. Our partners won't be able to receive orders and won't be able to prepare them. Imagine the situation if something goes down on the partner side: our couriers will be at the partner restaurants, but they won't be able to get what the customer might have ordered. These kinds of situations are very critical to our business.
[00:03:13] The monolith is a single code repository that is shared across many different teams. People who are developing features for couriers, for our partners, for our customers are all contributing to the same codebase, and the sheer volume of the code, 560K lines of code, is very huge as well. So if something goes wrong there, we lose a lot of money in revenue, which is quite impactful for the organization.
[00:03:51] The problem we had was with our deployments to production for the monolith. Until late 2019, our deployment frequency was once per day. Imagine a scenario where, as a business, you come to our engineering folks and say, okay, we want this feature to be developed, and our engineers have developed that feature, but we are not able to release because our deployment window has gone by. So the business and the development team have to wait 24 hours for the next deployment window to be available so that they can release the feature.
[00:04:27] And whenever we used to deploy, the key thing was that our change failure rate was around 10 to 20%. Out of a hundred deployments, ten to twenty deployments would result in rollbacks, and rollbacks are not a good thing to have in any engineering organization. When you have a rollback, you have to do the root cause analysis, you have to involve people from many different teams, you need to conduct post-mortems, and then come up with action plans and action items for team members to make sure this kind of situation does not happen again. So there was a lot of waste in our process, and it was preventing us from moving faster.
[00:05:10] Also, we had a strict policy that we don't deploy on Fridays and weekends, because Fridays and weekends are traditionally peak time for us, and if we deploy on Fridays and weekends we can potentially impact the business, and the impact can be a lot higher. So this was the main problem that we were trying to solve at Glovo.
What continuous delivery means to us
[00:05:34] We decided to adopt continuous delivery practices to resolve these specific issues. Most of the things that you will see on the following slides are not new to Glovo, and they're not very specific to Glovo either. What we have done is read about what continuous delivery is, and we are taking these practices from industry experience. There are a lot of smart people who have written books about these things, and we believe we can leverage their experiences.
[00:06:03] For us, continuous delivery is not just about CI/CD, and incidentally Dave Farley, who has written a book on continuous delivery with Jez Humble, agrees with us. This is a tweet that resonates with what we are trying to do at Glovo as well. As per Dave Farley, continuous delivery is much more than just automating your deployments, and it's actually broader than DevOps too, in his opinion, and we tend to agree with him. When we say continuous delivery, it's not just about automating your CI/CD. Implementing tools like Jenkins and Spinnaker is not going to mean that we have adopted continuous delivery practices.
[00:06:54] So we defined continuous delivery with this definition. The idea is our business can come to us at any given point in time and say they want this idea to be implemented. As an engineering team, we should make sure we are able to develop it and go through this whole build, test and deploy life cycle as quickly as possible, maybe in an incremental way, and we should be able to ship our software whenever our business is ready. Our main branch should be releasable at any time. That often requires not just technical changes; it might require organizational change, which luckily at Glovo we are going through right now, so we are ready to adopt these practices. The structure is there already. Technical changes are required, and development cultural changes might be required as well, and we'll go through what kind of changes we adopted when we went through this journey.
Why adopt continuous delivery
[00:07:53] The question we then asked ourselves was, why do we want to adopt continuous delivery? Obviously we want to fix the problem that we had around the monolith, which is reducing time to market and making sure our quality is not compromised when we go faster. But why do we want to do this, and what can we learn from the industry? This is quite key for us. We said we want to go faster, so we want to shorten the lead time, we want to improve the feedback loop, and we want to enable our business to experiment a lot more when it comes to new features on the platform.
[00:08:37] The idea is our business can come up with an idea, we can develop it, iterate and deploy as quickly as possible, and then roll it out to our customers, and our operations folks can get feedback from our customers, and based on that feedback they can iterate on their feature requirements. This is quite key, because without experimentation you cannot enable your operations or your business to get meaningful insights on a day-to-day basis within the platform.
[00:09:08] Jarnell Spa [?] has this famous diagram which highlights how slow delivery cycles can impact your business as well as your development teams, compared to fast delivery cycles. On the left-hand side, the graph shows you have large batch sizes going into production. If it is not daily, it might be weekly, it might be monthly, or it might be a quarterly release. The more you delay your release, the more lines of code you are releasing into production, and it will also take a long time to release. At the same time there will be much higher chances that something will go wrong, because the batch size is too big. On the other hand, if your change is small and you are deploying frequently, then the chances are your code changes will not cause a systems outage, your business will be happy, and in turn your development teams will be happy as well.
[00:10:16] One of the key things for us was that we looked at different surveys being done by companies, and one survey done by Puppet Labs was quite close to what we were trying to achieve. It highlights a number of KPIs or metrics on the left-hand side: deployment frequency, lead time for changes, mean time to recovery, and change failure rate. When they floated this survey to different organizations, depending on the answers, an organization could be deemed a low IT performer, a medium IT performer or a high IT performer.
[00:10:54] When we answered these questions, we were around a medium IT performer. We were releasing every day, our change failure rate was around 20 to 30%, and our mean time to recover was less than one day. Our motivation was that by adopting continuous delivery practices we could move to maybe a high-performer organization, where we can deploy multiple times per day. We should ideally be able to deploy in less than one hour, but that was not possible for the monolith, and I'll explain why. And our change failure rate should be between 0 and 15%.
Project Valkyrie: trunk-based development
[00:11:30] To resolve that problem, the approach we adopted was an initiative, a project, and the name of the project was Valkyrie. For project Valkyrie, we gathered different people that were contributing to the monolith: developers from the courier side, developers from the partner teams, developers from the customer teams. All these back-end developers were tasked to do one thing, and there was only one objective: to adopt continuous delivery practices for the monolith.
[00:12:03] What were those practices? One of the key things was to enable continuous deployment to production on each commit. We had to change how we manage our source code. We use GitHub as our source code repository, but the branching and merging strategy that we used to follow was a bit slow. Git flow, although it's a good branching strategy, did not enable us to have continuous deployment. So we moved to trunk-based development, which means every commit by a developer to master will go to production automatically, and there is only one source of truth for our code base, which is master. By adopting trunk-based development, you can say that whatever is in master is in production, so there is only one source of truth. If you want to read more about trunk-based development, there is a very good website that is managed by Paul Hammant, and he's spent a lot of time making sure all these strategies are defined properly and are clear, so the adoption path is very clear by following this website.
Testing as a whole-team activity
[00:13:10] Again, one of the key things for us was to define what testing means to us. Traditionally, what we have seen in the software industry, and in many organizations as well, is that testing is a phase in the software development life cycle. As a developer, if you develop something, you move your story from development to testing, and someone else, either in your team or in a separate department in your organization, picks up that story and starts testing. The risk of doing this is you go into this space where you throw your feature over the wall to some other department or some other team member, and then you get into the cycle where they test something, they report bugs, and then you fix, and that can go on for many iterations. We think that's counterproductive.
[00:13:54] We think testing is a cross-functional activity that involves the whole team, and that happens throughout the project. It starts when the project starts, but it never finishes; it's going on throughout the project. This diagram, which is from Dan Ashby, highlights that if you want to adopt either DevOps or continuous delivery practices, then testing is involved in each stage, or each phase, or each activity, whatever you want to call it: planning, branching, coding, merging, building, releasing. If you are deploying, then post-deployment verification is one thing that you want to test as well. When you release, you are testing there too. Your monitoring is also contributing to validation and verification activity. So all these things are testing in some form or other, and the whole team is responsible for that.
Automated deployments with Spinnaker
[00:14:47] When it comes to automated deployments or releases, we adopted an open-source tool, Spinnaker. It's a cloud-native continuous delivery tool. Netflix contributed a lot to Spinnaker, and it came out of Netflix as well. In 2015, I think, Google joined Netflix to develop this further, and in 2016 they open-sourced the tool, and since 2016 many companies have been contributing to Spinnaker and using it as a tool for their continuous deployment workflow orchestration.
[00:15:24] In terms of our deployment strategies, Spinnaker out of the box supports all these different types of strategies. We went for blue-green; in Spinnaker terms they call it red-black. I'll just explain why we went for blue-green and how this is helping us in terms of our continuous deployment efforts. We deploy a new cluster with the new changes, and we start rolling out the traffic to the new cluster. We do 50/50, so 50% of the traffic is going to the new cluster while 50% of the traffic is going to the old cluster, and we monitor for half an hour. If within that half an hour nothing goes wrong, the deployment is deemed successful, and then we destroy the old cluster.
[00:16:19] Obviously there are other strategies as well, and we are evaluating canary deployments, because we think by adopting canary releases we will be able to reduce the verification time window, and we will be able to do automatic rollbacks. Right now our rollback strategy is that if something goes wrong, developers can initiate a rollback using Spinnaker, which obviously we want to automate as soon as possible.
You build it, you run it
[00:16:50] One of the key things for us, and this is more of a philosophical point than a practice, I would say: continuous delivery says you build the software, and then you are responsible for running or operating the software in production as well. We wanted to adopt this practice, but to adopt it we had to develop some tooling, a lot of release process documentation, and a lot of runbooks as well, to make sure that when something goes wrong, developers know what they need to do. Runbooks are there to help them understand, if this specific situation happens, what the next steps are that they need to take to make sure either they can go back to a stable state or they can resolve that problem as quickly as possible.
[00:17:36] This slide highlights the tooling that we built and how developers are utilizing it. That specific bot, we call it internally the Valkyrie Police, and it basically sends developers notifications when their changes are about to go live. This highlights that Ray has made some changes and his feature is about to go live, and we have sent him a message on Slack that says, okay, hey Ray, your changes are about to go live, and this is the change set. Now Ray is responsible for making sure that the deployment is successful. He's also responsible for monitoring the systems when his changes have been deployed, for the half-an-hour window that we have, and if something goes wrong, he would initiate a rollback.
[00:18:27] All of these events we are sending to Datadog, which is our observability platform, and then we can monitor over time how many successful deployments we had and how many rollbacks we initiated, and this helps us to identify what the change failure rate is in terms of our deployments.
Outcomes
[00:18:47] Successfully, now our deployments are automated, and we are deploying at least 8.7 times per day. Obviously our goal was around 10 deployments per day when we started the project, but because the monolith is so huge, it takes a lot of time to build, test and make sure the quality is not compromised. The deployments are fast, but verification is still manual for that half-an-hour window. That's why our feature turnaround time is 2.5 hours. But if you compare it to before, from 24 hours to 2.5 hours is a big improvement.
[00:19:26] But the key thing is our change failure rate has gone down from 10 to 20% to 1.53%. So our business is much happier, our platform is much more stable, and we can release with much more confidence. Release now is a non-event. Previously at Glovo, release every day used to be an event. The platform team would gather, the developers would hand over their code, we would try to deploy it manually, and if something went wrong, the whole team would be stopped to look into issues and debug them.
[00:20:03] Now it's a non-event. It happens automatically. The only thing developers notice is a Slack message from the Valkyrie Police that their change has been rolled out, and if something goes wrong, they have the tooling available to initiate a rollback. Then we have processes in place to do post-mortems and these types of things, and to come up with action items if we need to improve the tooling or the process or whatever, but that's rarely happening right now. So that's a big win for us as an engineering organization, and for Glovo in general.
Project Turbine: from monolith to microservices
[00:20:38] What this has enabled us to do is plan our next project. I've shown you this figure before, so I won't go into detail, but the main thing here is that because the monolith is shared, nobody owns it. So there's an ownership problem. If something goes wrong, we scramble people and say, okay, what went wrong, what was the change, who is responsible, who needs to roll back, and all these types of things. Obviously, because it's so huge, it's brittle in nature. If you change something as a developer, you don't know what the impact is going to be on other teams or departments or the other features that we have in the app.
[00:21:23] That's why we have come up with a lot of processes: where you are making impactful changes, you need to initiate these processes and a lot of discussions, and we think that we can eliminate these kinds of inefficiencies in our processes. And obviously it's very complex. It's a monolith, it's a legacy system. It's very complex to operate; it's very complex to build, test and deploy.
[00:21:44] So we have initiated a new project, which is called Project Turbine. It's our journey from monolith to microservices. We believe that by having specific microservices which are owned by different teams based on their domains, we will solve the ownership problem. Also these will be more resilient, because they are small in nature. It's easy to reason about what the change is and how we are building, testing and deploying new changes. It will also help us to be more agile. If our business comes to us with a new feature, we should be able to respond to their request much more easily and much more quickly. At the same time, if something goes wrong in production, we should be able to isolate that problem to a specific microservice or set of microservices and resolve it quickly.
[00:22:37] Our first microservice went live in production, and again using the same mechanism, Spinnaker as a deployment tool, with a lot of automated testing in place, we were able to have 15 deployments per day for our first microservice. Our feature turnaround time for that specific microservice went down to 0.5 hours. So if a change is ready to be deployed, we can go through the whole automated pipeline within half an hour, and that's quite a big milestone for us.
Key takeaways
[00:23:11] The key takeaways from project Valkyrie and Project Turbine so far: it always helps to have focused teams which are delivering very focused objectives. Whatever project you are going to initiate, if it is a technical project, you need to define what the business KPIs are and what the technical KPIs are that you want to measure as the success criteria for that specific project.
[00:23:42] Then there are some key principles that we have adopted along the way. Immutable infrastructure is one of them. Immutable infrastructure means that you run some scripts and your infrastructure is there, and the source of truth for your infrastructure is those scripts. There is no config that is done manually, there is no setting that is done manually on any part of your infrastructure. The idea is you deploy a new version, you bring up a new set of servers or load balancers or whatever is required to run your piece of software, and you throw away the old infrastructure.
[00:24:21] It helps you to have repeatability in your process, and repeatability and consistency are key when something goes wrong. Obviously, when everything is successful, you don't find consistency and repeatability to be a big bonus point. But if something goes wrong and you can reliably reproduce that issue, then it's a big win, and we have seen this many times. If something goes wrong and you can run the same script and get the same results, then it's a lot easier to debug and find out what the root cause is. On the other hand, if you don't have this kind of repeatability and consistency in your infrastructure, or in anything you do, especially around testing, then you introduce flakiness, and flakiness can contribute to loss of confidence in your infrastructure and in your practices as well. If something goes wrong, and you run the same script and get a different result, then it's very difficult to find out what the problem might be.
[00:25:23] Again, think about what kind of deployment strategies you want to adopt. We went for blue-green because we thought it was well suited for us. Now, retrospectively, Spinnaker has been in production for the last three months, and we think by reducing that half-an-hour window, where a developer is proactively looking at the logs and the monitors to make sure nothing is wrong with our production system, by adopting canary releases with automated rollbacks, we can bring in a lot more efficiency in terms of the process. We can reduce this half an hour to maybe five to ten minutes.
[00:26:00] With automation in place, we can free up the developer. If something goes wrong, we should be able to automatically roll back and send a notification to the developer: hey, you deployed something, we had to roll back, these are the logs, these are the issues that you need to investigate. Rather than a developer proactively looking and thinking, okay, I have initiated my changes, now I need to monitor for half an hour. So automation and deployment strategies are key to make sure you are freeing up developer time to focus on innovation, and on new features going into production.
[00:26:35] Operational integration is quite important. All the data that you generate from your CI/CD processes and from your testing, we are sending to Datadog, which is our observability platform, and that in turn helps us to identify where the bottlenecks are. Again, the same example: that half-an-hour window that we have right now has been highlighted through the Datadog integration. Our time to market is 2.5 hours, and out of those 2.5 hours, for half an hour we just wait while the feature is being exercised in production. To reduce that time, what we can do is canary releases and automatic rollbacks.
[00:27:18] Again, a lot of tooling is required to make sure developers are informed of the changes going into production, and they have the documentation available. The release process is very clearly documented, especially for onboarding new people, and runbooks are there for people to understand, if something goes wrong, what they are supposed to do. All of these contribute to the very healthy engineering culture that we already have at Glovo, but we are striving to improve that a lot more, based on the feedback that we are getting from all these systems and all these initiatives that are going on at Glovo. Thank you.
