Data Driven Engineering Team - Changing The Culture

05 Oct19:00 – 19:30 UTCTalk
Slides

Checking session availability…

Hang tight while we load the latest updates.

Since joining Hootsuite, Greg Bell has been leading the change of his engineering teams, not only between themselves but how they are working with other teams to focus on the customers needs.
This session Greg will look at how he's changed the culture over the past few years within this team to become a data driven engineering team.

Data Driven Engineering Team - Changing The Culture

Greg Bell at UXDX Europe. Video: https://youtu.be/2K2F-JahUDQ

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

Cultural change is not a linear path

[00:00:00] After reading most books on strategic leadership, organizational change or cultural change, you'll likely have a mental model that looks something like this. Along the time axis, you'll build a vision and mission and create some values, and then you'll collaboratively work with your teams to build goals and assign some measurable metrics, and voila: leadership complete, clear sailing from here on out. It will be results, results, results, up and to the right. My experience is that it feels a lot more like this when we're in it. While we're deep in it, progress is not always clear. It may or may not feel like we're heading in the right direction, the whirlwind of every day is very real, and progress feels very slow. Are we making progress on strategic objectives? Who knows? I need to deal with this specific issue right here, right now. Cultural change is not a linear path.

[00:00:54] Hi, I'm Greg Bell. I'm the VP of Software Development at Hootsuite, where I lead a team of 200 engineers across six offices to deliver on a mission to make social the highest performing customer engagement channel. Our software delivers 30 million posts a month and processes over 50 million social events a day for large and small businesses around the world.

[00:01:16] Cultural change, data-driven engineering: these are some lofty topics. They're loaded terms, full of different meanings to each of us, and I can't possibly do them justice in the next 25 minutes. I do, however, want to share our actual journey, not just the glossy, successful version, but some of the challenges that we encountered, where we made change and what worked. At Hootsuite we have a core value of working out loud, and today I'd like to do just that.

[00:01:53] For me, cultural change is about changing behavior. In my role leading software development, I have peers like our Head of Design and our Head of Product Management that I collaborate with every day, and the culture of our engineering team is deeply connected to the culture of our design and product management teams. We work together to build the product, and our culture is really just one. I've learned that we must change behaviors together.

[00:02:17] And data-driven engineering: what do I mean by that? For me, I'm thinking about running an entire department, not just specific teams. When I'm talking about data-driven engineering, I'm worried about what measures and metrics we need across 200 engineers and a hundred designers and product managers to increase the quality and value of our products in market over the long term. At Hootsuite we're split into teams of six or eight people, fairly standard in our industry now, with representation from product, design and development. We have about 30 of these teams that make up our entire product and development organization, and like many organizations, especially right now, we're either split across offices or working from home a lot. Each office ends up having its own culture that has evolved over the years.

[00:03:10] Today's journey, by design, is from the perspective of a VP of Engineering. Hopefully the story is helpful no matter where you sit within an organization.

The desk drawer strategy

[00:03:20] To get us started, I'd like to rewind to January 2018, which seems like a long time ago now. Many of you, I'm sure, have kickoffs at the beginning of the year, and Hootsuite is no different. I've been at the company for over five years now, and we've done it every year. Yearly strategic planning ends up taking a lot of time, even with a company of our size; we're about a thousand people worldwide. One of the biggest challenges is what I call the desk drawer strategy.

[00:03:49] Desk drawer strategy plans are the ones that our leadership team puts together at the beginning of the year to rally the troops at some type of kickoff event, without really doing the hard work of planning how the organization will actually deliver on them and the results they expect. The plan is put together, then put in a drawer, and everyone goes back to their regularly scheduled work. At the end of the year, in preparation for the next year's planning cycle, the leadership team opens that drawer again and finds the plan from last year. At this point it's: what? Remember what we said last year? How did we actually do on our plans? As you can imagine, and maybe some of you have actually experienced this, it is not the most effective way to make change in an organization.

Thirty key results and no single bullseye

[00:04:34] Circa 2017, this was generally the style of strategic planning I saw in our organization. When 2018 planning came about, we decided to build goals for the year that we could look back at and decide if we'd actually delivered on them, for better or worse. We were looking for a behavior change. We needed all 30 of our teams to build a cadence of delivering on their team roadmaps and departmental or corporate initiatives, and we needed to hold ourselves accountable to the results we desired, not just a set of projects that we wanted to do.

[00:05:08] We made the decision to structure these projects as objectives and key results. Of course, OKRs were all the rage, and they continue to be, and we really liked them. The key property we liked about OKRs was that they forced us both to clarify what we were trying to accomplish and to create measurable results that we would be able to track to know how we were doing against them.

[00:05:28] We went away as an engineering leadership team and ran a process to figure out what we wanted to accomplish in 2018. We did the hard work to prioritize the objectives down to three. I'm glossing over it here, but I don't want to underestimate how challenging it is to do this. You'll probably hear me mention it a few times through this presentation: prioritization is hard work. To do it well, and create a sense of direction and alignment for a large team, takes a lot of effort. But in the end we had our three department objectives, along with three key results for each. We were incredibly excited that finally we would be able to measure the success of our yearly plans.

[00:06:10] Our product management team went away and did the same thing. Instead of three objectives, they prioritized four, which in and of itself wasn't that big of a deal. Then our operations team did the same thing, and we found ourselves with 10 objectives and 30 key results across three departments that all work together to deliver our product to our customers.

[00:06:33] On the positive side, we did the hard work to come up with measures that we could hold ourselves accountable for, and we built communication plans to ensure that everyone knew the objectives we were trying to accomplish for the year and how they could help. Everyone in our department could read about them and get clarity about what we were trying to accomplish. The process actually created a game for our teams, and it was a game we could actually win. Individual team members could plan how they could contribute to these goals, so for some it was incredibly engaging.

[00:07:07] And I'm certain, from the slide you're looking at, that you could already guess some of the challenges we ran into. We went from zero measurable targets in 2017 to 30 in 2018, and while we did our best to build a cadence of reporting, we had none of the organizational discipline, structures or tools in place to really measure ourselves well. The big challenge we ran into is that we didn't have a single bullseye. Instead we had 30 bullseyes, all competing for mindshare and attention. Teams didn't know if they should be working on departmental objectives or their own teams' roadmaps.

[00:07:42] Teams did a great job with what they had, but by midway through the year it became clear we were missing the boat. We were making progress against some of our OKRs, but teams were confused about what they should be working on, and the net result of all of this was that quality in our product started to go down. This is a chart of our bugs over the last half of 2017 and the first half of 2018. I'm not particularly proud of this, but this is what being more strategic actually got us. We were super clear on the major things we wanted to accomplish; however, we did it at the cost of everything else. The blue line is new bugs in our system, and you can see a big backlog of bugs growing.

The Quality-a-Thon

[00:08:29] Fast forward to July 2018. We're seeing this situation where bugs were growing, production incidents were a problem for us, and the performance of the platform was getting worse. On top of this, Cambridge Analytica happened, which meant that the APIs we rely on to build our base business were in a state of rapid change. We knew we had to act, and we had to act quickly.

[00:08:52] On July 18th we had a hackathon planned. Now, hackathons are a big deal at Hootsuite. Every six months we spend three days working collaboratively across dev and design, product and ops, to build something innovative or fun, or just to learn something new. Three weeks before the hackathon, I made the very hard decision to rebrand the hackathon as the Quality-a-Thon. There was a collaborative group that got together to make the call, our product managers and our operations leaders also. To this day I have the sticker on my laptop, and it's a reminder to me and our teams to be humble, never to forget that we had to rebrand one of our hackathons, which was supposed to be nothing but fun, as a Quality-a-Thon.

[00:09:44] For the week-long Quality-a-Thon we clarified three goals: reduce the number of open bugs, increase the number of automated smoke tests, and increase the observability of our microservices architecture. We spent the last couple of weeks before the Quality-a-Thon building a scoreboard that would allow us to know if and when we were making progress during the five-day Quality-a-Thon. To all of our surprise, the Quality-a-Thon was actually a huge hit. We had a ton of fun, we gave out amazing prizes, and we made a huge difference for our customers. There's a lot of pride in building great stuff, and that week we built a stronger culture and cultivated some of that pride.

[00:10:32] Another benefit that we saw was that we built a bunch of dashboards to run the Quality-a-Thon, in particular a daily dashboard for the number of escaped bugs. It was an integration between Zendesk and Jira, two systems that didn't talk to each other before, and these were tools that we could use over the long run.

A prioritization framework and a health scorecard

[00:10:53] However, we knew that the scoreboard alone wasn't enough. We needed to give teams a system to prioritize work, something that would allow local teams to make the right choice based on their context. Again, we had about 30 scrum teams, so we couldn't micromanage all of them. We needed them to feel empowered to make the right choice, to work on new features or to fix bugs.

[00:11:18] We designed a very simple tool, a prioritization framework that went along with the health scorecard. The prioritization framework gave teams a decision-making tool, and it's incredibly simple. It says that if your health scorecard is red, then that's priority number one; that's at the very top. Second is your committed roadmap items: what have we committed to our internal or external customers that we're actually going to build? Third is longer-term or strategic projects.

[00:11:46] After rolling this out along with the health scorecard, I received immediate feedback on how valuable it was, and it was a real learning experience for me. This prioritization framework seems simple and obvious, but it removed a lot of guesswork for teams. They had a tool that they could point to for priorities, where they didn't have to go and speak to managers and directors or other teams to try to find out what they should actually be working on. They could point to the prioritization framework, and away they go.

[00:12:26] The prioritization framework is only as useful as the set of metrics that we had for that first layer of it, which was the health scorecard. We designed a very simple health scorecard that was visible to everybody. It was reported on monthly, and at first it was almost completely manual. To get started, it was very simple, just a Google Sheet that drove this reporting. The health scorecard and the prioritization framework meant that we had data and tools in place to allow the teams themselves to prioritize and be aligned with where we were trying to go as an organization. This step was gold, and we have essentially been iterating on this same concept ever since.

[00:13:13] I won't go through each of the different metrics here, but I'm certain that this will be available after the fact, and you can certainly review each of them. I'm happy to chat about them more afterwards. So what was the result? The result of the Quality-a-Thon and then the prioritization framework is that we had a massive decline in the number of escaped bugs month over month. During this time we tried a bunch of other tactics to get better quality, but the most impactful thing that really stuck was this idea of clear metrics, visible scorecards, and systems to prioritize and keep teams accountable. These are the tools that helped us change the behavior of our teams in a repeatable way.

2019: the year of reliability

[00:13:55] January 2019. We were able to speak directly to how we performed in 2018, because we had the OKRs that we set at the beginning of the year, and it was humbling. Out of the 30 key results that we designed at the beginning of 2018, we only hit a few, and we came to realize that the OKRs we'd set out weren't actually the most important work for us to be doing. We did, however, see just how important our health scorecard became in the last half of the year. We made amazing progress against our bug numbers, and we were making some progress against our reliability targets. For us, 2019 would become the year where we would focus on reliability.

[00:14:38] We set out, along with product and design, to create a system like we had for bugs, but for reliability (think availability, correctness and latency), that would allow us to create the same behavioral change on teams as we had with bugs. We realized that we couldn't centrally manage a set of reliability projects, although there were hundreds on the table to consider. Instead we wanted every team to own the reliability of their own products in production. We didn't know how, but we knew that we needed a way to encode metrics within a system of accountability that enabled the teams to prioritize reliability work against all the other work they could possibly be doing, of which, just like in your day to day, there are a ton of different types.

SLIs, SLOs and error budgets

[00:15:32] Enter SLOs, SLIs and error budgets. Like many other organizations, we have taken inspiration from the site reliability engineering practices that Google has shared with our industry, and these terms are all from the SRE world, as it's known. I'll briefly define these terms for you, but for a thorough description I highly recommend reading both of the SRE books. They're available for free online, they're very well written, and they give a ton of insight into how Google does SRE. I'll use one of our actual examples to define them quickly, so you get an idea of how these all work together.

[00:16:11] First are the SLIs. SLIs are service level indicators, and you can think of these as the actual metric. The important aspect of these metrics is that they're expressed as ratios. In this example, we're saying that the SLI is the proportion of requests that were sufficiently fast, with sufficiently fast defined as less than 1.5 seconds. It's very easily readable language. Anybody on a team, not just engineers, but program managers, product managers and designers, all participate in creating these.

[00:16:43] The SLOs are service level objectives, and they define our targets or objectives for any given SLI. On the right, you can see the SLO we actually agreed on. We agreed that we want our system to respond to 99.9% of requests sufficiently fast. You can think of the error budget as the inverse of the SLO. We want 99.9% of our responses to be sufficiently fast, which means that during any given 30-day period we have 0.1% of our traffic to play with, to try out new stuff. We can be slow on those, and that's okay. We've all agreed that we can be slow on those requests.

[00:17:28] Okay, now that we have those high-level concepts out of the way, let's talk about what attracted us to them in the first place, and it's the contract. The thing that caught our attention most was this idea of a contract. The SLIs, the SLOs and the error budgets all work together to create a contract that we can all point to. It's another system of prioritization and accountability. Yes, it's based on data, but it also touches on many other aspects of how to shift a culture; it's not just about the data. And most importantly, to the point of this entire conference, SLOs are not an engineering decision tool. They are agreed upon by design, product and engineering leadership teams together.

[00:18:11] Here's an actual example of the top matter of an actual SLO at Hootsuite, and you can see that there's a list of reviewers and approvers. This list is made up of the product manager, the design lead, the engineering manager and a staff engineer. We generally call that the triad, and the triad actually owns the SLO. They get to decide what the expectations of our customers are, and this allows us to connect the reliability of our systems all the way back to our customers, agreed upon across product, engineering and design.

[00:18:55] At the start of 2019 we had zero of these SLOs defined, and our Q1 OKR was to have one SLO in production per portfolio. It's not that important, but at Hootsuite our teams are split up into portfolios; usually three to five teams are grouped into a portfolio. At the time this felt like a ridiculously small goal. I knew that our organization needed to learn how to do this stuff, but I was also pretty sad that we were only committing to building out under 10 SLOs. It's a good reminder for me of just how incremental this process feels while you're in it. In hindsight, lots of progress can be made, but while you're in it, it can feel very slow.

[00:19:41] Leaving 2019, we actually had over 120 documented SLOs, and we're up closer to 150 now, that had been negotiated across product, design and engineering, and they documented their error budgets and policies for what we would do if they tripped. This has been incredibly useful for our teams to proactively decide on nonfunctional requirements, and then to prioritize reliability work against all the other work they could have on their plates. Again, it's a tool that helps us systematically change behavior.

[00:20:17] In tandem, in 2019 we updated our health scorecard, and you can see that SLOs became a part of it, down at the bottom right. We modified some of the different measures that were on our health scorecard and continued to iterate on it throughout the year. This is basically what it looked like. Again, we won't go through all the details, but I'm happy to talk about any of these specific measures later. So what were the results? In 2019 we saw a 56% decrease year over year in the total hours that our system was in a degraded state, and we did this by focusing on clear metrics and scoreboards that supported the behavior change we all wanted to see across product, design and engineering to drive the quality of our product for our customers.

The technology master plan

[00:21:10] Now we find ourselves in 2020, and we're expanding on these ideas. We spent the last half of 2019 building out a three-year product and technology vision for Hootsuite. Hundreds of people contributed, and it was a huge effort. We did one of these look-under-every-stone activities where we considered everything, and in the end we had a really compelling plan for the next few years. We wanted to make it extremely accessible to everyone across all of our teams. Our goal was for an engineer, designer or product manager to be able to read a single page and understand immediately what our long-term technology vision was, and we did just that. We built what we call the technology master plan.

[00:21:59] The technology master plan is a one-page document, but in it we define 10 different metrics that we're accountable for delivering in the next three years. On the left-hand side you can see reliability-type metrics, the performance and availability of our systems, and on the right-hand side you have more internal metrics that we're holding ourselves accountable to. These 10 metrics, along with the strategy to deliver them, will guide our decision making over the next three years, and I do hope to be able to share more on what's working and not working on this journey as we make progress. Today, just for context, we've seen another 60% decrease year over year in the total hours that our system was in a degraded state. All signs are pointing in the right direction that using these different tools is really paying off.

What worked

[00:22:55] The culture has shifted. We've become a more customer-obsessed team, and we have better tools to help us prioritize, using data, all the different types of work we need to do on a large system. When I look back over the past four years, here is a summary of what I think has worked for us. We've done a lot of hard work around prioritization. We've built compelling scorecards. We've created systems of accountability that create alignment and create the opportunity for teams to be empowered to deliver on their own missions while being aligned to the organization. And we've measured both health and strategic metrics at the same time.

[00:23:39] Probably one of the biggest things that has come through is that it's not actually just about having one piece of this pie. It's about having the entire pie, and the only way to make progress towards the entire system is to do it very iteratively. It's not dissimilar to how we build software. We have to iteratively shift the culture, quarter after quarter and year after year.

[00:24:04] I'll end where I began, which is that this stuff is not easy work. Even in telling the story of the past few years at Hootsuite, it puts a bit of a glossy tinge on it and makes it seem like we knew the path more than we actually did. The truth is that we had a vague idea of where we wanted to go, but we didn't have the roadmap to get there. And I'd expect that many of you are on similar but slightly different journeys within your organizations. My hope in sharing this is that it has inspired you to iterate on your internal processes and systems to create more direction, alignment and commitment: in the end, doing the hard work to shift the culture. Thank you.

Speaker

Greg Bell

Greg Bell

VP, Software Development

Hootsuite