No Risk, No Reward: The Joys Of Testing In Production.
Checking session availability…
Hang tight while we load the latest updates.
Natalia and Fabio will share their story of how the HBC DevOps culture and mindset evolved. Follow their journey from moving to a distributed microservice architecture to enable simple, quick and frequent production roll-outs, through phased deployments to dark and live canary nodes, to a radical 'no staging environment' approach. Hear about the benefits and caveats they've experienced along the way.
No Risk, No Reward: The Joys Of Testing In Production.
Natalia Bartol, Fabio Cognigni at UXDX Community: Dublin. Video: https://youtu.be/e2x-IUSo3sI
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
A 348-year-old company
[00:00:00] Natalia: I'm Natalia Bartol, I'm the director of mobile engineering at Hudson's Bay Company. This is Fabio.
[00:00:19] Fabio: I'm Fabio Cognigni and I'm the principal engineer in the mobile team at Hudson's Bay Company.
[00:00:29] Natalia: We work for a company that is 348 years old. Hudson's Bay Company was once the largest landowner in the world, possessing the area around the Hudson Bay watershed in Canada, comprising an area of over one third of modern day Canada. And it was controlling the fur trade throughout North America for several centuries.
[00:01:12] Natalia: The logo of the company that you see here has a Latin motto on it. It says "Pro Pelle Cutem," and this means "skin for skin." It's often interpreted as human skin for animal skin. So as you can imagine, the early days of Hudson's Bay Company were rough and bloody. There is even a Netflix series called Frontier that tells the story of how HBC was established in Canada. Has anyone seen this? I'm not going to tell you if HBC people were the good guys or the bad guys, you have to watch yourself.
[00:02:02] Natalia: Even up to these days, the employees of the company are called adventurers and the high level executives are called governors. So the company is still living up to this tradition. Anyway, rest assured that despite being called adventurers, we no longer hunt wild animals for their skins. It's much more peaceful now, although we may be still involved in selling animal skins in the stores.
[00:02:37] Natalia: So with the decline of the fur trade, Hudson's Bay Company evolved into a mercantile business, delivering vital goods to the settlers in the Canadian West, and they're still in business to this day. Over the years HBC acquired other retail brands, including Saks Fifth Avenue, the luxury retail store from New York, as well as Lord & Taylor, or Galeria Kaufhof from Germany, or Gilt.com. And in addition to running and operating the physical retail stores, HBC is also actively expanding in the digital space. And this is where we come into the picture.
[00:03:28] Natalia: So we are the tech team behind the luxury online experiences of Saks Fifth Avenue, Saks Off 5th, Lord & Taylor, Gilt and thebay.com. And our team in particular is responsible for delivering the mobile shopping experience for our customers.
The project: a new iOS app on a 20-year-old backend
[00:03:49] Natalia: So let's start with a little bit of the story of our project. The goal of our team was to implement a new shiny iOS app for the customers of Saks Fifth Avenue. And there already existed an e-commerce website where you can shop shoes and handbags. This website has been around for a while, it was launched roughly 18 years ago. And this website is powered by a backend, and it's a Java-based e-commerce platform. And the assumption was that the new mobile app will be simply another front-end client in addition to the website that is powered by this backend.
[00:04:48] Natalia: So given that this backend, this box there, is roughly 20 years old, and it's a Java e-commerce platform, would anyone here dare to make a guess what is the internal architecture of this black box?
[00:05:04] Audience: Spaghetti.
[00:05:07] Natalia: Oh my god. You have a brilliant intuition[?]. First I was looking for pictures of the ball of mud, but we ended up actually putting a block of spaghetti here.
[00:05:22] Natalia: So that was a huge challenge for our team, that we were facing this ancient monolithic block of spaghetti code that was website-oriented. It was never designed to be used by the mobile apps. And we had to figure out how to build on top of this. And we also had to understand what is the life cycle of shipping features on top of a block of spaghetti.
The release process on a monolith
[00:05:52] Fabio: So basically, dealing with this kind of backend, with this kind of monolithic app, means having a release process that is roughly like this. Basically this consists of a single very large code base where dozens of contributors work on initiatives, new features, all happening in parallel at the same time.
[00:06:30] Fabio: So the way it works is that every developer makes code changes in a local dev environment. Then these changes need to be committed, pushed, and automatic tests are run. Finally this can be handed over to our QA team, that needs to test this first in a QA environment that is isolated, and just to iterate over these features with the fixing and everything. Then finally, when the QA team signs off on this environment, everything can go to a staging environment, that is the last step before production. Checking for regressions and retesting again these features. And finally, again handing over this process to yet another team, the operations team, to have this deployed in production and monitor how it performs.
[00:07:41] Fabio: So yeah, as you can imagine, this process is really incredibly slow and frustrating. It involves quite a huge system in four different environments at least, and ownership is shared between three different teams. And all these operations are very difficult and risky, because you find every time yourself dealing with deploying quite a lot of changes. And it takes a lot of time also to get to the position of saying "now I'm ready to deploy to production," because of regressions, because you have to check this. So you have a very, very low level of maneuverability. At best, the best was getting to the point of being able to deploy once a week, and missing this frequency very often. Shipping new features was on average taking months. And also troubleshooting production issues, or even issues in staging, were taking very, very long.
Isolating ourselves: the strangler pattern
[00:09:05] Fabio: So, what we've done: obviously the first step was to abstract and isolate ourselves from this mess. While there was some work in progress to include that, spinning up a new middleware[?] service between the client apps and this backend, that could expose the API that the mobile clients wished they had. And we literally started treating that backend as a third party black box, completely isolated from us.
[00:09:53] Fabio: And then, little by little, we started implementing new features that were not in the web experience yet, and started implementing them in new services that we spawned on our side. And also keep practicing this technique of proxying[?] other third party services where these were not the best for mobile experience. And keep going in also taking pieces of functionality or features from this monolith and extracting this into services, to add this to our fleet of microservices, that literally became now many dozens of services.
[00:10:49] Fabio: And literally, we really applied what in software architecture is called the strangler pattern, that actually is inspired by nature. These are the pictures of a strangler fig, as it is called, and it really embodies this principle of having these roots that completely strangle the original trunk, and eventually we completely replace that. And this is the technique that we applied as well in shrinking down this monolith and killing it gradually.
[00:11:40] Fabio: So, applying this process basically put ourselves in a place where we moved from having just one huge service to having to deal with n services, but still dealing with two environments, a staging environment and a production environment. But finally we were completely owning the whole process of implementing, releasing and deploying, and having a clear ownership of all the process.
[00:12:24] Fabio: And most importantly, probably the most powerful thing was being able to deploy very small changes, but very frequently. This gave us a lot more control on our production environment. So when you do that, it's much, much easier, if you start noticing issues, to find out and detect the causes for these issues, and iterate very fast on implementing. And also you really can see that you get a much higher quality of code, because you basically find yourself writing code that needs to get to production very soon, that you are responsible to deploy and monitor and do the full handling. And also, very importantly, we were finally able to ship new features independently, from either other features that were work in progress or other services that we were depending on.
Canary deployment
[00:13:47] Fabio: Once there, we took a step further in that direction, especially in the last step of deploying to production. We went through a phased rollout to production, applying what is called canary deployment. I'm pretty sure that many of you are familiar with that technique. The name comes from canary birds that were actually sent into new coal mines to detect if there were toxic gases, so that miners wouldn't go and get killed.
[00:14:42] Fabio: The fact is, for us it was very easy to go there, since now we were dealing with services that were fully deployed in our cloud, where it was very easy to apply patterns of this type. The first step is for us to deploy to a node, a dark canary node, that basically is a node that is exactly as any other node in production but that doesn't receive any customer traffic. It's just for us, for internal access, to either quickly validate a deploy that we are doing, or showcase.
[00:15:25] Fabio: And if everything is fine there, we go and deploy to a live canary node, that is basically a node of the production fleet that receives a portion of the real customer traffic. And this is very helpful to really see how your changes behave with real traffic, with real load, and monitor what's going on there for a while. And if everything is fine, deploy that new release to all the other nodes of the production cluster.
[00:16:10] Fabio: So basically we got to a much better place. But I guess we weren't happy yet.
The maintenance headache of a staging environment
[00:16:28] Natalia: Yes, so things got much, much easier for us to deploy. We had those nicely decoupled services. Everyone on the team could deploy them whenever they wanted, whenever they thought the functionality is ready. But as our ecosystem of microservices powering the app started growing, we realized that it's also quite a maintenance headache to have two environments, staging and production, for all those services.
[00:17:01] Natalia: So what is the maintenance headache? First of all, because you need to have two environments for every service, this adds to infrastructure, complexity and cost. We host everything in Amazon Web Services, so every service has its own EC2 instance. So that would be one instance for production, one instance for staging. With every new service you'll have two additional machines to keep an eye on and make sure they're healthy, up and running, and they all cost money.
[00:17:32] Natalia: On top of this, we realized that we've spent way too many hours troubleshooting and debugging issues that only happened in staging. Staging environments have this track record of being fragile and usually out of sync, and they're almost never maintained with the same accuracy as production environments, for many reasons. Mostly because it's not production. If something doesn't work in staging, nobody will wake up at 3am to the PagerDuty alert because staging went down, right? Developers usually wouldn't know that. However, this will then become a bottleneck in your deployment process. If the staging is actually not operational, you cannot deploy to production.
[00:18:20] Natalia: So it was getting all really, really complicated and we started thinking, okay, why are we using staging at all? The first two things that come to your mind when you think about why am I using staging is that staging gives me this safe place to test my code. So definitely we want to store test data separately from production data, and we want to apply the same separation for the test traffic, test load, versus production load. Okay, fair enough, those are very good assumptions.
[00:18:56] Natalia: In our case we also use staging because we are forced to do so by interacting with some third party services. In some cases there are features that can only be tested end-to-end using a staging environment. In the case of our app, that's the core feature of the app, which is the checkout flow. So to test that an order actually can be placed, and that the payment will be made, that the credit card will be charged, we can only test this in staging. But as Fabio mentioned, we are treating the website backend as a third party service, because we don't have control over it. So it's a third party service for us. And we just accept this, that okay, we will have to test this functionality against staging, because there is simply no other way.
[00:19:43] Natalia: And the second case is that sometimes testing in staging is the only way to catch the breaking changes before they show up in production. This is a really bad scenario, but it happened to us multiple times. And it's a serious design flaw between us and those third party services, that we don't have well-defined contracts of how the API should look, that the breaking changes should be versioned and so on. Anyway, that's our priority, it's tough. So this is why we still keep testing in staging, just to make sure that no changes that are rolled out to a third party production environment will break our app[?].
[00:20:25] Natalia: Anyway, taking aside those considerations around third party services, when it comes to the code that we own and we have control over, we decided that for those first two bullets it's actually okay to try and get rid of staging, and we can deal with this in production.
Testing in production: read-only first, then writes
[00:20:47] Fabio: Yeah, exactly. So at this point we said, okay, we have a lot of read-only operations, and that was really easy and not risky, to go and say, okay, let's remove staging and let's hit directly production for these cases. And it wasn't even a concern for us to say, okay, we are going to add some additional load to the production environment with our testing. Just because, as part of our business being in e-commerce, we are already exposed to lots of spikes in the traffic due to sales events that are sent out to customers and all the notifications. So it really happens every day that during that 10 minutes where notifications go out, you get 10 times the normal load that you usually have. So that wasn't a concern at all.
[00:21:58] Fabio: And anyway, it was fine. After we got confident about this, we started also saying, okay, even for services or operations that have any writing operation, or have to modify any kind of state... we tried the approach of saying, we already had designed all our services to be multi-tenant, in the sense that they need to serve different banners. All the banners that HBC currently owns are separate businesses, so we really designed it that way. So for us it was very natural to go and say, okay, we will just create test organizations, where in an isolated way we say, okay, write the data in a safe way in this sandbox organization, and just use that. Or we use test accounts for scenarios that involve login, or both, like saying, these specific reserved test accounts will then act on these test organizations.
[00:23:21] Fabio: And basically, this process was great for us in the sense that at that point production really becomes your friend. Because then the developer experience really is something like this, where, after the classical making all the code changes, pull requests reviewed and then merged, triggering the unit tests, then we have integration tests that actually run against the production versions of all the dependent services. So if everything succeeds at this point, we can have a new release of our service. And with the deploying process that we described earlier, we go with the canary deployment.
[00:24:28] Fabio: And one very nice effect of an approach like this is that, really, if you ship a feature today, getting your production system exercised this way really gives you the confidence that that same feature will work in one day, in one week, in one year. Obviously in addition to all the monitoring and analyzing that you already have. But it's a much more productive approach.
[00:25:13] Fabio: And then finally we got to the point of saying, okay, we ended up with a position where those numbers became having a fleet of services, but working only in one environment, and everything owned within a team. And really, now we are in the position of doing multiple deploys a day, shipping complete new features in a matter of weeks, and troubleshooting any kind of issues became really easy. When you deploy small changes but very frequently, it's very easy to understand what changed and act on that. And most of all, you have no friction to get and push your code into production, and a lot more independence.
What happened after we removed staging
[00:26:22] Natalia: So we got rid of staging environments, and then we failed in a spectacular way.
[00:26:35] Fabio: No we didn't. Actually nothing ever went wrong. Yeah, no. So it's been more than a year and a half that we are using this approach, and really not much can go wrong with an approach like this, just if you really make sure that you have good monitoring and easy and quick rollback. Nowadays it's very easy to get that with all the tooling that you have in AWS or any other cloud solution. Rollbacks are very quick just because you always deploy very small changes.
[00:27:28] Fabio: And also, at least for us, it turned out to be very valuable to get to production much earlier than what was expected, and taking just a small amount of risk, than going much later. And this is great even for business, in terms of revenue and under many aspects.
[00:28:02] Natalia: So to summarize, your takeaway from this talk should be a recipe for building this happy place for your developers. A happy place for your developers is a place where the development team owns every phase of the software development lifecycle, can release a product independently from other services, and can test with confidence without the need for staging or stubbed or mock data. And finally, where developers can deploy as well. This may take a little bit of effort to introduce, but the results are amazing. There is a boost in productivity, and the quality of the final product is much, much better. Thank you.
[00:28:59] Fabio: Thank you.

