Moving 65,000 Microsofties To DevOps

05 Oct3:40 pm – 4:05 pmStage: DX StageTalk

Checking session availability…

Hang tight while we load the latest updates.

  • Microsoft’s internal engineering systems: the pre DevOps days
  • Changing the company culture from the top and bottom up
  • The implementation approach
  • Unique challenges from Microsoft's structure
  • Lessons learned

Moving 65,000 Microsofties To DevOps

Martin Woodward at UXDX EMEA. Video: https://youtu.be/WLw_gssGbco

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

One engineering system for all of Microsoft

[00:00:00] Right, my name is Martin Woodward. martinwo at microsoft.com if you want to email me. At MartinWo if you want to send me some abuse on Twitter, or if you want to get me on GitHub or whatever. And visualstudio.com/devops is where I hang out.

[00:00:12] I'm fancy titled Principal Group Program Manager, DevOps. What does that mean? Principal Group Program Manager means I delete a lot of email. DevOps means I delete a lot of email with a bash script, basically.

[00:00:25] So I work on a thing called Visual Studio Team Services, which is our collaboration system for engineers. It's Git hosting, agile planning, Kanban boards, work item tracking, release management, your builds, you can do hosted and on-prem builds and all that sort of stuff. There's telemetry and all that thing. We have a cloud, as we all know, we love the cloud. That's about as much of an advert as you're going to get from me today, because I'm on the engineering side, but that's what I work on.

[00:00:58] When I joined Microsoft in 2009, a cartoon came out the year after of the top tech companies and their org charts. Coming into Microsoft from a cheeky little startup, this felt very familiar to what I was witnessing when I came in. It was a very hierarchical organization and not everybody talked to each other.

[00:01:21] One of the things that Satya Nadella did when he came into the org was break a lot of these boundaries down inside Microsoft and really try and get everybody to work as one Microsoft. And that means one engineering system in Microsoft as well. If I want to go into Notepad and add UNIX line endings support into Notepad, and believe me I've tried, okay, and if I was finally trying to do that, I can actually now go look at the Windows source code and find it, because it's in VSTS, it's in the same hosting systems we sell to everybody else. And I can send them a pull request, and they can reject it, believe me they will. But we use these inner sourced workflows inside the company because we're all in one engineering system, which is hosted on top of Visual Studio Team Services.

[00:02:13] So we now have in that engineering system, as I say, 75,713, sorry, there we go, that was the number at the end of August, so it keeps going up and up. That's everybody in engineering. That's Windows, Office, Azure. So we have this amazing thing where the engineering tools for the engineering tools are deployed to Azure using the engineering tools, because remember we have a build and release management system. Azure is deployed to Azure using the engineering tools hosted on Azure. Azure runs a lot of Windows, which is built using the engineering tools which are hosted on it. It's like Inception. I hope there's a backup plan if it goes wrong. It probably involves computers buried in mountains or something, but I'm not going to worry about that. I will get fired if it happens, so I'll be fine.

[00:03:09] The journey to DevOps has been an interesting one for us. When I joined in 2009 we were shipping shrink-wrap products. My first course I had to do when I joined the company was to approve the hologram that was used on the DVD that my product was now shipping on. I came from a startup and I'm in a course about approving holograms to go on DVDs that get made in Dublin and in Puerto Rico. What am I doing? This is going to be weird. But the whole company has moved closer to how I was working anyway in my little cheeky startup, so it's been a really great transition to be part of.

[00:03:46] We started the journey probably around 2010, from shrink-wrap to service, and we're now a fully online service. There's a couple of regular trains that ship every day, but we can patch production all the time, and then there's a big train that happens every three weeks as well, because we have a three-week sprint cadence for the big feature crews. So it's a big transition, great fun to be part of.

One: be completely customer-focused

[00:04:12] The five things that are trying to break us down, it's kind of five things we learned as we went along, to try and make it more applicable to everybody else. The first one was to be completely customer-focused, and this is easy to say, so I'll try and explain what I mean in practice.

[00:04:32] Obviously we do all the usual stuff you're trying to do when you're listening to customers. You have a UserVoice, check, we've got whatever. Stack Overflow tags, check, we sponsor ours because we're Microsoft, so we also give the Stack Overflow guys a bunch of money because they're awesome.

[00:04:50] And then we also do some things a bit more in-depth where we have direct relationships with a lot of our customers into the engineering team. We have our top customer accounts, customers who are our top not just in terms of usage, but in terms of they're doing interesting stuff that's maybe a bit different. We assign them a champ, who is a person who is on one of the engineering teams inside the product, and they go talk to them and they completely shortcut all the product support, they shortcut the salespeople, because he wants to talk to them. It's just engineer to engineer, helping them to be more successful. We also obviously have in-product feedback and all that sort of stuff, which we probably all do, and collect a stack of telemetry.

[00:05:42] Our definition of DevOps inside the company is this union of people, process and tools. It's about delivering value as quickly as we can to our customers. All of the things we build, we build with the mindset of how do we get this into the customer's hands as quickly as possible.

[00:06:02] This is how we define when a feature is done. In the old days, when I first started, it was done when it was checked in, and then you would go for these massive stabilization phases and you would fix bugs and all sorts of badness would happen. Then when we did agile, it became: done was potentially shippable. And then that's got a bit of a woolliness to it, what's potentially shippable, what's not? Nowadays, done is shipped. So everything at the end of the sprint has to be running in production, it has to be collecting telemetry, and that telemetry has to tell you if the hypothesis you're running is true or false.

[00:06:41] How are you proving or disproving your hypothesis? And your hypotheses are always in terms of customer value. I think that if we change the system in X way, then people will be able to use this faster, or people will be able to do that better, or we'll get more revenue, or we'll reduce the cost of running the system, those sorts of things.

Measuring what matters, and watching the outliers

[00:07:01] We collect tons of telemetry. It's just knocking on to petabytes of telemetry a day coming into our system, which we ingest using, we call it Kusto, but I guess Azure Application Insights Analytics is our catchy name, because Microsoft's awesome at naming things. But Kusto is what we all call it internally, and it's an awesome way of... we basically bring in a ton of data, we capture as much data as we can, obviously without the PII parts, we just ingest a bunch of data, and then that allows us to ask questions after the fact. But we try and measure only what we care about. We're very careful about what we measure.

[00:07:43] Here's an example, some of our dashboards. We measure things like, are you using your features? How quickly can I build something? How quickly can I get to self-test? What's my time to deploy? These are very, very important ones: how quickly do we detect live site issues, how quickly do we tell people about them, how quickly do we fix them, and then how quickly do we remediate, how quickly do we stop them from ever happening again?

[00:08:08] These are things we don't watch, which I always like to highlight. We don't ask people how close they were to their original estimate. We don't ask them what their team's velocity is, because as a manager I don't care. I just care what your results are, what you're delivering to your customers. I don't care how you got there. I honestly don't care, and I want the team to be self-empowered to do whatever it is they want to do to get that way. But if you're not delivering value to your customers, then I care. We also have some debt counts and things like that to make sure we don't build up too much debt.

[00:08:43] When we're looking at customers, one thing we like to do is focus on the outliers. When you've got a big service, we've got like millions and millions of customers now, which is awesome and it's great, it's actually reasonably easy to stay up most of the time when you're in multiple data centers with millions of customers. So those nines start not to mean as much in your SaaS. So what we do instead is focus on the experience each customer is having, and I mean that by a customer account. So, if you are a user of VSTS, what's your experience? Then we look at the outliers.

[00:09:22] I had a story I was going to tell here, but then a better one happened at the weekend, where a customer we noticed looked weird and our telemetry was picking up, this is odd, the machine learning models were, hey, this is weird, things going funny here, what's happening? A bunch of data was getting deleted by a job that was running with admin privileges at the API level with warnings switched off. And we're like, whoa. So you ring up the customer, and it turns out the admin for that customer was doing a migration job with warnings switched off that they'd written wrong and was deleting most of the data in their account. And so, hey, we probably don't want to do that. And so we ended up starting a restore for them and got them back up and running on Monday morning. So we still have a VSTS administrator in that organization that's not been fired, so that's good. Our customer champ knows who to talk to. I can't promise that level of service, but that was an example of things that can happen when you're looking at the outliers and trying to identify strange patterns.

Hypothesis-driven development and feature flags

[00:10:20] We do subscribe to build-measure-learn, so hypothesis-driven development. The user stories tend to be around, we believe this person wants this and we're going to test it to see if they do, and how we're going to write the tests, those sorts of things.

[00:10:37] We are always deploying into production, and we do this by using feature flags. A feature flag is basically a fancy name for an if statement in your code. Rather than having big long-running feature branches which you then merge when you're really sure it's ready, we're trying to always ship from master and always be continuously integrating as much as possible. So once we have a really high degree of confidence the code's good, we integrate it in very, very early, but we have it switched off. We'll put an if statement in, we'll start implementing the feature, and then we'll make it toggleable so you can switch that feature on or off.

[00:11:17] And then we do things like, we measure, for people that have opted into the feature, how quickly they opt out if they opt out, how many people opt out, because the opt-out is a sign that the feature is not good. It's a classic: if your UX is a bit like a joke, if you have to explain it, it's not working. And the same's true, that's how we measure things. We measure if UX or UI changes or product changes are working. If users, when they switch to it, if they stay with it, if they don't, then it's not working. I don't care how fancy you're telling me it works, it doesn't, people are opting out. So we do everything based on data. Data, sorry, I've been hanging out with Americans too much, and they're awesome.

[00:12:04] The problem with feature flags is it enables dark launch. So you can build a feature and have it running in production, and then you can have some fancy event where you get some executive to stand up on a stage, press the big red button and say, ta-da, the feature's now live, which is awesome. And so it goes wrong.

[00:12:21] We did one a few years ago. We switched it on on stage, got the exec to switch it on on stage, and then the demo bombed because the whole site started going down. Now, we'd rehearsed that loads and loads and loads with the feature flag just switched on for that particular demo, and it worked great. But once we switched it on, turns out everybody started using it and there was some problem with the live load coming in and it impacted us and it was a nightmare and we had a bunch of egg on the face. So don't do that, okay? We switch on feature flags gradually and we do it well before a big event, just to make sure everything's running, the system's in steady state.

A production-first mindset

[00:12:59] That takes me on to having a production-first mindset. So, live site incidents. We have a live site incident, basically a fancy name for a conference bridge. When there's a problem detected, a conference bridge gets fired up. Nowadays it gets fired up using AI, because we're Microsoft, you've got to try and throw some machine learning in there. So it fires in machine learning and it calls a conference bridge and gets what's called the DRI, a directly responsible individual, on the phone, and pages them.

[00:13:28] When you're in your team, you have five minutes during the working day to get on the phone bridge, or 15 minutes in the middle of the night, or if you're walking around Harry Potter World like I was the other week with my family. But never mind, that's the joys of being on call, as we all know, DevOps teams, it's awesome. So we want to try and minimize those, which is why we then go into self-healing and things like that, to try and minimize the number of alerts.

[00:13:51] When we do go wrong, we try and be incredibly transparent about it and give a lot of detail to prove we're not idiots. But the key is that our alerting systems needed to get really good and we were able to detect quicker, and this is why time to detection is an important metric for us. When we have a live site incident, we want to make sure we put systems in place so we can detect that problem again quicker, if we can't completely stop it from ever happening.

[00:14:19] One of the things we also learnt is an important lesson: there's no such thing as partial automation. I'll show you some PowerShell. In the old days we used to copy scripts around and things like that. We used to have a script that was built up that you ran when you did the deployments. Anybody see what the problem is with this script here? It's PowerShell, but apart from that, what's the problem? Okay, smart quotes. When you copy things around in emails it turns proper quotes into these crazy smart quotes and breaks everything. That happened and we broke production one day. Awesome.

[00:14:50] So, no such thing as partial automation. Automate everything. Nowadays we can't run scripts in production. You have to check them in and then run them from version control. Again, it's just little things like that, you learn from mistakes.

[00:15:06] When we go through production we go through five rings. We're the first ring, and then we go to our smallest data center, and we go to our biggest data center, we go to our furthest away data center, and then we do everybody else. That's kind of how we roll stuff out, and that helps us detect early if things are going to go wrong, and it gives us a high degree of confidence before we deploy to the rest. We do staged deployments. When we are doing this as well, it's important, instilling this live site culture was the biggest change when we went from shrink-wrap to service.

[00:15:36] We also think a lot about security. Whenever I get an email that has a link in it nowadays, I'm assuming it's my own security team trying to phish me to get into a system. It's no holds barred. We have our own internal pentesting teams, our own breakers[?], but more importantly we also have our own internal teams that are trying to spot the pentesters. So you assume a breach mentality and you're trying to do wargaming to try and make sure you can detect who's in and who's not.

Team autonomy with alignment

[00:16:05] I'm quickly going to whip through this because we already talked a bit about this in the keynote this morning. We are also a big believer in team autonomy. That book I mentioned, a cheat's way is, go look up RSA Animate on YouTube and watch the video, takes five minutes, saves you reading the book. We have autonomy in the plan and the practices, but we have some alignment within the business about how we operate. Sorry, I'll pull that one back up. Slides will be up on SlideShare afterwards, don't worry about it.

[00:16:34] The reason why you need some alignment is, if you give people too much autonomy, bad things can happen. Here's a case when we brought Windows over. We brought them in, we had a migration plan, we were going to do it a bunch of ways, and then Satya came on board and said, no, they're going to go use VSTS and we want them there in two weeks. Oh, okay, awesome, two weeks.

[00:16:51] When you get a deadline like that from an executive, you take a lot of shortcuts. So we said yes to a bunch of stuff which we later regretted. We got to their bug form when we did this. This is how many fields they had in their bug form, because we kept saying yes to them. Sorry, let me scroll down. This is how many fields they had in their bug form. It was nuts. So then we had to do a process where we simplified that.

[00:17:12] Here's how we have our teams structured. We brought engineering and test together into an engineering team, and then we have a single feature team that's got program managers, people like me, engineers, and ops people, all in one team, all delivering the value to the customers. They all live in one team room, they all have a clear charter, they stay together. We might change the team's charter, but we keep the team together for 18 months, and that's very important for team dynamics. We do allow teams to move around, we allow mobility. We keep the teams, it's a vertical slice, so it's not like a back-end team and a front-end team, it's you own the entire customer feature.

[00:17:52] Every 18 months we do this thing where we allow engineers to self-select where they go work. So as a manager I pitch my team, what we're going to work on, what my vision is, and then engineers come along and decide if they want to work for me or not. Incredibly scary. As a good manager you kind of pitch to make sure you're going to get your right team in place. But it allows mobility within our org without people deciding the only way to have career prospects is to leave. So we try and encourage mobility within the teams, and then we lock the teams down for a year or so, two years, and keep them together.

[00:18:24] When we plan, we have these 18-month big goals. They're literally just bullet points. The six-month plan is the point of, we have screenshots, wireframes kind of thing. Three sprints is the only way we can actually sensibly plan something, is the only thing we actually believe in. And then a sprint is the only thing... I know what's going to happen this sprint, I mean, we're in sprint 125 right now, I have a good idea what's going to happen. I kind of know what's going to happen in sprint 127, and my six-month plan is just some wireframes where I'm kind of going, we're going to go do this, boss, is that all right? That's how we do our planning.

[00:19:03] So leadership look at the six-month plan, and then the feature teams have complete control over how they deliver everything else. Complete autonomy. Doing this, it's awesome and everybody smiles and puts post-it notes on whiteboards. I think that's what that slide says.

Shifting left, and infrastructure as a flexible resource

[00:19:20] Okay, then the next one is around shifting left. We had a lot of tests that were UI tests and we needed to move that to proper integration tests that we could run everywhere. That's great to say, and everybody says that in manuals. I just wanted to make you feel better by how long it took us to do that. Because we had a big legacy code base, it took us years. Literally, you can see, from sprint 78 to sprint 120, so times it up by three weeks, a long time. But we got there. So it can be done. Have faith, gently chip away at it.

[00:19:53] We also do a lot of stuff in our pull requests. This is what a pull request looks like in VSTS. We do a lot of pre-merge validation, so we do security testing and a bunch of build validation before we merge, because we want to continuously integrate.

[00:20:10] And then the final thing, so I've been rushing through this, the final thing that we learned was trying to learn how to treat infrastructure as a flexible resource. We went from a big monolithic application that was shrink-wrapped to running in the cloud, and we want to have the benefits of the cloud. So we started by making use of the dev test environment, because it's easy to do dev test even if you're still shipping on prem. So we went from VMs under a desk, to VMs in a lab, to using Azure Dev Test Labs, which allows us to easily fire up dev test VMs in Azure, which is still connected to our corporate network. And that's awesome, and even though we're not shipping out there, we're still doing this sort of thing. So, also using the cloud, we ship the same code on prem as we do to the cloud, which is important.

[00:21:00] This is how we structure our data centers. We had one code base which we've now split into two services, and that was painful, okay, because it was a big monolithic code base. That gave us geo, that gave us the ability to scale across data centers, which was key. Once we had that capability, we did that. Now we're taking this code base and actually containerizing a bunch of stuff. But again, it's not easy. You can't just containerize overnight. So we're taking individual new services being built as containers and containerizing those and running them as microservices, but we've still got this big monolithic code base that we're gradually pulling things out of. And to be honest, that's how most normal business applications end up happening. You can't just rewrite everything.

[00:21:49] And then finally, containers are awesome because it allows us to do better cost analysis, basically. But you all know why containers are cool.

DevOps isn't unicorns and rainbows

[00:22:00] All right, last thing. DevOps isn't this magic unicorns and rainbows, despite everything. It's hard work. The technology wasn't the hardest part. The hardest part was the people part. It was building a live site culture within the team and actually transforming how we did our business. But we knew we needed to worry about how quickly we shipped, not the mean time. So mean time to release, not the mean time between failures, was what we wanted to focus on, because that allowed us to innovate quicker, allowed us to build quicker and test things with customers. And then we can take those same learnings and put them in the shrink-wrap product, which we do ship quarterly, which we still do for people who don't want to be on the cloud. So we were able to transform our business to allow us to still be doing DevOps, to be agile, but also still keep our on-premise products as well.

[00:22:53] We did a bunch of stuff, and then finally: if it hurts, just do it again. When you're automating, if it's painful, then do it again so you get good at it. If it's good for you, keep doing it until it stops hurting. And make it your engineers' responsibility, because then they'll automate it, especially if you call them up at three in the morning. Okay, I think that's about it. Thank you very much for your time. Any emails, just give me a shout. Thank you.

Speaker

Martin Woodward

Martin Woodward

Director of Developer Relations

GitHub