Evolution of Grammarly's Platform: From Developer Operations to Developer Experience
Checking session availability…
Hang tight while we load the latest updates.
In this talk, Serhii shares the journey of Grammarly's platform transformation.
It all began with a "you build it, you run it" culture that required developers to be software engineers and operations experts. While empowering, this approach created a significant cognitive load and slowed the onboarding of new team members. Grammarly has since significantly evolved its platform to boost developers’ productivity and improve their experience.
The evolution of the platform can be traced through several key stages:
- Initial Platform Team Intervention: Segregated environments with self-service provisioning and other automation were introduced, though new developers still needed up to a month to become productive.
- Platform University: Creation of structured online training courses, reducing onboarding time to one week.
- Golden State and Golden Path: Implementing high-level abstractions like one-button service deployments, simplifying complex operations.
- Current Infrastructure Abstraction: Moving away from "you build it, you run it" to specialized Infrastructure teams and a whole new tech platform.
Looking ahead, Serhii explores how the platform might continue to evolve, incorporating AI-assisted operations, enhanced developer experience features, and a more robust internal platform ecosystem.
Evolution of Grammarly's Platform: From Developer Operations to Developer Experience
Serhii Vasylenko at UXDX EMEA. Video: https://youtu.be/h-ihAlCjVCo
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Grammarly's engineering platform and its mission
[00:00:07] Hey everyone, my name is Serhii Vasylenko and I'm a staff engineer at Grammarly. Bear with me. Today I'm going to be talking about AI and stuff, but you already got used to it. We did it a long time before it became a hype, so I can allow it to myself. For those of you who don't know, Grammarly is an AI-powered writing assistant that works on desktop and mobile and helps you to communicate effectively through clear and personalized communication. But today's talk is not about Grammarly as a company, but about the engineering platform team that helps our developers to make Grammarly better, and to help you all communicate better.
[00:00:53] The mission of Grammarly platform is to help engineers to own their services with minimum burden caused by infrastructure, platform tooling, day-to-day operations. To focus on code. And the scale is quite significant. We have more than 500 engineers working on hundreds of backend services across hundreds of repositories. All this stuff is deployed on more than 14,000 server nodes, and our CI/CD system runs around 6,000 CI/CD pipelines daily. Our monitoring system withholds about 20 million requests per second, to help the whole engineering organization keep an eye on all Grammarly services, backend, frontend, everything.
[00:01:46] However, the core challenge is not about the scale, or the tooling, or CI/CD. The core challenge is about finding this sort of balance between developer autonomy and standardization. Because on the one hand, who knows better than developers what the project technical needs are? If a developer team owns their services end to end, they can make it super cool, super scalable, super everything. On the other hand, this doesn't quite scale, because developers are overwhelmed by maintenance tasks and a lot of other tasks, and they spend less time on coding and innovating.
Four pillars, and where we are today
[00:02:35] This is about the "you build it, you run it", or "you build it, you own it" principle. Even without a comprehensive strategy or understanding of what we should do, back in 2020 we realized that our understanding of "you build it, you run it" needed a fresh look, and I'm going to talk about how we changed this in the rest of the presentation.
[00:03:01] We understood that we need to change, and we sort of formalized four key principles, I would say, or four pillars that we need to uphold. Unified developer experience, to minimize the steps required for developers to accomplish their tasks. Then there is standardized implementation, standardized templates and patterns. The industry also calls it golden path, the way that you create new services, do stuff, in a paved road kind of thinking. Then self-service and abstracted infrastructure, again to spend less time on infra stuff and more time on coding, because of automations, developer portals, self-service stuff like this.
[00:03:55] And last but not least, technical compliance, because when we have everything up and running, when we create stuff, we also need to make sure that it corresponds to our standards and remains as we want it to be. The other name for this is maintaining the golden state. I don't know why the industry uses gold for this, but whatever, gold is money maybe.
[00:04:20] Before exploring this evolution in detail, the path we moved here, let me quickly give you an understanding of what we've got as of today. Engineering onboarding time now takes just a week instead of four weeks. Self-service tooling for the platform has decreased request times for a lot of tasks from days to hours, especially for backbone provisioning and stuff like this. Templated project creation, golden paths for Java or Python, reduced bootstrapping for projects from hours to minutes. Just imagine that you can click some buttons and you have boilerplate and everything. I'll tell you more later.
[00:05:05] Implementation of AI coding tools and code search saved approximately 200 engineering hours daily across the engineering organization. That's a lot. And not only this, because the organization itself also gained a lot of cool wins here, such as automated operating system security patches, which reduced our patch cycle time for the whole fleet from a month to just a few days. This, and all the other security implementations we did, helped us win a lot of important security certifications, to sell better for enterprises.
[00:05:47] Altogether this evolution has transformed our engineering platform from a support role to something that is actually valued as an important asset by our leadership.
Where we started: 2020 and the 2021 initiatives
[00:06:03] And now, closer to our journey. I will be covering the last five years. Actually Grammarly is 16 years old and the engineering platform is about 12 years or so, but today I will be speaking about the story of the last five years. We will be going through 2021, when we started to see the need for change. Then 2022, when we did some foundational architectural changes. Then 2023, which was a year of developer experience, I would say. And then 2024, where we faced a sort of familiar yet sudden challenge and we had to rethink completely how the platform is built and how it operates. And landing in 2025 with the stuff we have now, and forward plans.
[00:06:57] But before we dive into this, let's actually zoom back to 2020, to see where we started. In 2020 the ownership model of "you build it, you run it" was looking like this. Developer teams were maintaining comprehensive ownership around their services: CI/CD, code, infrastructure, observability, monitoring, logging, all the stuff. A platform team also existed and provided fundamental tooling like CI/CD systems, durability systems, backbone architecture, network, cluster, things like this.
[00:07:35] This operational model worked fine back then, because Grammarly was smaller, and maybe which is even more important, the market was different, at least for our niche, I guess for the whole industry. Developer teams could afford to spend more time on infrastructure and stuff. But we started to see some signs of something going wrong with this, because it wouldn't scale nicely with the company's goals and the ambitious growing we had in mind.
[00:08:10] So we started to explore, and our core initiatives in 2021 were directed toward the onboarding and observability stuff. We identified that it took about four weeks for an engineer to be onboarded, for a newcomer to get well onboarded onto the project and everything. I experienced this myself. I joined Grammarly in November 2020 and it took me like 10 meetings with different engineers just to get familiar with the stuff we have. That's too much, and that clearly wouldn't scale.
[00:08:52] Therefore we implemented several solutions here. First of all, the thing called platform university. This is a set of tutorials, text-based, with tasks and everything, aimed to let engineers understand in practice what they should do, how it works and everything. And this reduced onboarding drastically.
[00:09:18] The second thing is the infra service, because provisioning an AWS account for a new service took a lot of time, so we had to automate this stuff, and we did, reducing this provisioning from a couple of days to just a couple of hours or something. And the service catalog. Another cool thing we did, we used a tool called Cortex, if any of you are familiar with that. Basically it's a place where you can put all the information about your backend services, libraries, whatever, and have the owners, connections between them, dependencies, all the stuff. This helps a lot, because first it helps you navigate the organization better, but especially it helps a lot when it's about finding what's going wrong when you have some on-call alert or something like this.
[00:10:06] Altogether this lowered the barriers for developers and put some sort of foundation for further platform development.
2022: abstraction, golden paths and a CI/CD rethink
[00:10:19] But of course that was just the beginning, because next year we realized that, yeah, okay, we have all that, but infrastructure was still super complex for developers. They had self-services, they could provision accounts, they could provision some stuff, networks, that's nice, but still, even with this, they spent a disproportionate amount of time managing the platform, managing the infra. Because when you bootstrap a new project, you need not just an AWS account, you also need a codebase, you need CI/CD setup, you need some Docker files, you need some alerting configuration, logging, all the stuff, and you spend a lot of time on that. It just drains your time, and we had to do something with that.
[00:11:11] So we focused on abstraction as much as we could, and we focused on things like the golden path, as I already mentioned. For abstraction, the way is that, yeah, we had Terraform, we had a lot of automations throughout Terraform and everything to manage the infra, but we put as much as we could back then behind self-service portals using Jira help desk, to provision not only the backbone infra but also service-to-service communication, or setting up CI/CD for existing projects for example, or new projects, or whatsoever.
[00:11:48] We also focused, as I was saying, on the golden path. For the golden path, for example for Java and Python projects that we heavily use, it's code boilerplate, it's CI/CD configurations, it's Docker files, all the things. Using Cortex and its web interface, so we had to be a little bit of web designers, web something like this, not so good, but still it was usable. And yeah, you could, and still can actually, our engineers can bootstrap their projects in minutes, which is really really cool and saves a lot of time for them.
[00:12:26] We also had to rethink how our CI/CD infrastructure is built, and we completely redesigned it, and basically allocated the whole CI resources in one account instead of multiple accounts, and connected it with all other accounts properly, aligning with security network policies. CI provision time reduced drastically, because we also put a lot of custom logic on top of that, for provisioning and for bootstrapping new projects with CI configs, with CI runners included, and stuff like this.
2023: shared responsibility, the LLM proxy and AI tooling
[00:12:55] In 2023, everything was good and nice, but we still saw that it could be even better, and we saw that many teams use sort of inconsistent tooling. This team uses something for deployment, this team uses something else for deployment. It works on our CI/CD platform, yeah, but it's different, and it's not so easy to jump from project to project, because sometimes it's necessary, quite often actually, and to contribute to somebody.
[00:13:30] We tried to solve this, and the way we tried to solve this was our sort of shared responsibility model. We welcomed developer teams to help us build platform tooling with us, and to put their domain expertise in and merge it with our expertise of platform stuff. This especially was handy in things like the golden path, because with code boilerplate and templates and everything, that helps really well.
[00:14:05] Then it resulted in our next project, which we called LLM proxy, in 2023. It was a key initiative for the whole platform, to enable developer teams to integrate LLM features, providers, tooling, because we use not only external providers, we have a lot of internal LLM stuff, including our own models. So it helps developers integrate all that into their projects in a day instead of days, or maybe even weeks.
[00:14:36] Also we integrated AI coding assistants and code search, things like GitHub Copilot or Sourcegraph, to speed up the way how developers do their daily routines: searching repositories, writing some code, all the stuff. And we also enabled things like automated security patches for our repositories. Because again, when the security team finds that, hey, there is a vulnerability in the code and you need to fix this, this is nice, but you need to fix this, you need to spend some time. There are tools on the market, plenty of them. For example we used Renovate, and of course put some extra logic on top of this to orchestrate.
[00:15:22] But in a nutshell, now when there is a security patch for your source code, this tool creates a pull request, and if the CI tests pass, the pull request is merged automatically, without you even having to check in and see how it works. It requires properly set up tests and everything, if you don't want anything to be broken, but that's a different topic of course.
2024: the platform becomes the bottleneck, and the replatform project
[00:15:50] Now going further, in 2024 we faced a familiar but sudden problem, because the company was scaling heavily and fast, and the platform team and all our cool tooling and everything became a bottleneck, because we had to deal with more teams, more projects, more complex use cases than ever before. Our centralized process struggled to keep pace with all these cross-team or parallel collaborations, many teams, many developers, everything.
[00:16:25] We realized that we cannot solve this by just adding more self-services, or more AI tools, or whatever. We did that actually, but still it's a bit different. The whole concept should have been changed, and we needed a fundamental change again, not just for the tooling but also for the team and how it works.
[00:16:52] So what we did, we started this replatform project. And fun fact here, it actually started from a proof of concept project for a small team, I also was a part of it. We tried to set up Grammarly in Kubernetes to make it portable, the Grammarly backend, because you know, it's nice to have something portable. And so we did, and then we realized, hey, this actually can be scaled. Kubernetes became the foundation of the technical changes, but of course not only one tool, because it's a massive tool of course, but still it's just a tool.
[00:17:28] The core thing is that we had several key changes in terms of infrastructure and developer experience. For infrastructure we got a service mesh, to simplify and streamline service-to-service communication and fortify the security of this communication. We got uniform configurations for autoscaling and deployments, because again, when everything is managed by the Kubernetes platform or similar, you can do things in one way across the whole engineering org.
[00:18:01] And this unified foundation also allowed us, because Grammarly is now sort of portable, we can sell better, because some enterprises require on-premise installations. They say, "Hey, we trust you, but we trust ourselves more, so please host Grammarly in our cloud." Now it's possible. And finally, with tools like Karpenter or CloudZero, we got ourselves predictable resource allocation and super clear visibility around the prices and how much we spend on all the cloud cost.
[00:18:32] Of course we had a lot of developer experience gains here. For example, standardized deployments. The way that we do deployments now is that there's a special set of tooling based on Argo CD and stuff, that deploys services in a canary way, with of course configurable measures for how much to deploy first, second and everything. All this happens seamlessly, and which is more important, it's not the developer teams who manage this, it's the platform team who does this for them. They have an interface to configure this for the service specifically, but the core thing is managed by us.
[00:19:11] Then there is the internal developer ecosystem, as we call it. Basically we are integrating a lot of internal and of course external tools all together to make this DevX ecosystem. For instance, we have a service called Hacker Mango[?], which is used like this: when a pull request is created, it automatically finds and assigns a proper reviewer, not just random people, for this pull request. It's a small change, but on the scale of the whole organization it's a massive gain, because it saves a lot of time, and for some projects it reduces the time for a merge request, for a pull request, to be merged from hours to just minutes. A sort of game changer from a small thing.
[00:19:58] And of course there are ephemeral environments as well. This is a thing also related to pull requests. When you have a pull request and you have this code change you want to deploy, you want to test it somehow, probably. These ephemeral environments allow you to host your code change separately from the rest, but also add the dependencies of your service, like other services at Grammarly, in a way that it doesn't block owners of these dependencies or your teammates from doing their work. So basically we have this isolated environment with all the stuff needed for your service, to run tests, integration and end-to-end, whatsoever, things like this.
[00:20:43] And of course we also spent a lot of time on monitoring and analyzing productivity metrics. We use a tool called DX, which allows us to integrate Slack, GitLab, the deployment system, all together in one centralized sort of dashboards and panels and everything, where we can see how all the things are running, and it allows us to make informed, numbers-based decisions on what to work on next, instead of just guessing.
The new ownership model and three platform teams
[00:21:17] So the ownership model now looks like this. Developer teams still own their services, but it's less of platform and infra stuff. We took away a lot of infrastructure and process management from them. We use opinionated ways of doing things. This doesn't mean that we do not want teams to experiment. They can go off road and do it the way they want, but they don't do it too often, because generally it's easier to use this paved road and just make it faster to deploy or something.
[00:21:59] However, as I was saying earlier, this also transformed the way how our platform team works and looks like, because before that we had one team, but now we have three, with different areas of responsibility. There is a developer experience team, and their core meaning is this toolbox that developers touch every day. IDEs, service inventory, golden path templates, all this stuff. Basically, make a happy path from the IDE to a pull request.
[00:22:38] Then when the pull request gets into the CI/CD system, this is where our release engineering team steps in, and they own the whole CI/CD stack. GitLab, Artifactory, artifacts management, deployments, releases, all the things. Their goal is to make the deployments uniform and fast and stable. When the deployment is green in production, the production engineering team steps in, and the goal of this team is to make operations reliable and as less stressful as possible, by provisioning core infrastructure, clusters, autoscaling, all the things.
[00:23:16] The production engineering team then reports its metrics, from SLO breaches or something else, basically all the way how your services are operating, back to the developer experience team and their dashboard for developers. And this kind of closes the loop of this whole uniform idea of managing the platform.
Success factors, challenges and what's next
[00:23:36] Now a few of the success factors and challenges we've had. We learned that co-designing with developers is super important, and treating platform as a product also helps a lot. Basically we involve developers to design solutions with them, for them. We use internal surveys, we use internal research, we use sort of even marketing campaigns internally, because what we produce is a product, but for a small amount of people. Actually maybe not so small, about 500. Anyway, it's a product we create.
[00:24:13] And we also learned that overcommunication is key, because it's better to be super fluid in Slack and meetings and everything, to make sure that all engineering teams know what you're bringing to them, what you're giving to them, or what you're changing maybe, instead of just humble small announcements here and there. And of course, we learned that it's better to start sooner than later. What I mean is that some bold decisions might be frightening, and you may say, hey, it requires a lot of effort and time to implement. Yes, but they often outweigh all the small incremental changes.
[00:24:57] Talking about our future plans. Now we are concentrated around this idea of a developer portal. Basically, infrastructure and platform should melt even further, a step further into the background for developers, allowing them not to lose the focus on daily tasks. Things like Backstage, products like Backstage or the earlier mentioned Cortex, we are inspired by these products and use some of them. But the idea is to make this portal where developer teams can see their services' status, cost, CI pipeline statuses, bootstrap new services, all in one place, just to stay in the zone and not to lose focus on all the other stuff.
[00:25:46] In the end, what does this mean for you, for the different roles in this audience? For platform engineers, for developers: champion developer experience. Analyze what your customers are using, make weekly meetings to find the pain points or bottlenecks, prioritize them for your work. For managers and leads, embrace platform as a product, because it helps you to deliver faster, in a more reliable and sometimes even cheaper way. And then for product managers, cooperate with your platform engineering teams. Explain to them the product vision, how you can help them understand how to make the company more successful, and help promote their own products that they create throughout the company, and make them better. Thank you for your attention, and I'll be glad to answer any questions if you have. Thank you.
Q&A
[00:26:48] Host: Absolutely amazing, and an area that's very, very close to my heart actually. I cut my teeth, my first UX research role was in a team we called engineering productivity, but it was in platform as a product. So this is like coming home, seeing this talk. Lovely stuff. We've got a bunch of questions online, which is great. Obviously Grammarly, huge scale, big team. What is your advice to smaller dev teams trying to create this platform as a product approach?
[00:27:19] Serhii: Oh yes, that's nice. I would say don't overengineer at first. And if you see that you have a dev team that sort of needs support, allocate somebody to be this foundational platform engineer, and ask that person to work closely with the team to identify at least the top one pain point. Solve this and just move gradually. Yeah, that would be my first advice here. Small, small, and go.
[00:27:51] Host: Yeah, I love it. A competition question, less about the platform stack, but ChatGPT fixing grammar. Do you see that as a threat to Grammarly? Is that something you guys are worried about?
[00:28:03] Serhii: Oh well, obviously it's a competitor, but it's not a replacement, I would say. I'm from Grammarly, of course I can say this, but the point is that if you use ChatGPT you can see that it's sort of complicated to make it produce text in a way that you sound. Correct me if I'm wrong though, maybe it's about prompt engineering, I don't know, but I use it a lot and it does the job well. But the Grammarly feature, the main feature, I guess one of the main features, is that it actually learns how you type, learns how you express yourself, and can create similar sounding text. If that's important to you, it's nice. If not, well, also nice. But this is the difference, I would say.
[00:28:51] Host: Yeah, ChatGPT is a fairly holistic thing as well, right? It does an awful lot of stuff. Grammarly is a very targeted tool to do specifically that. So always ask the experts, everyone. I promise this isn't a question I put down in the hope of getting a contract with you later, but did you in any way involve UX designers, UX researchers, in improving developer experience as you moved to platform as a product?
[00:29:12] Serhii: Oh yes, for these self-services and the web we created for Cortex. Long story short, yes we do, not as much as we could though, because they have a lot of tasks to be done, but it depends on the OKRs, priorities and everything. We try a lot to involve web designers, web developers, to help us create interfaces for other developers, because we are not super cool in web coding, unfortunately. We're good in other stuff.
[00:29:41] Host: Yeah, nice, I like it. A great question here about metrics. I actually think the number of metrics you put up about the change in onboarding was phenomenally impressive. Shifting those numbers down to get people onto the platform and up and running as devs was absolutely amazing. Quick question here: what other examples do you have of productivity metrics you use beyond just onboarding? And how are you balancing productivity versus the outcome?
[00:30:06] Serhii: In terms of metrics, for starters, just standard DORA things. You measure the time for a merge request to be merged to the master branch. You measure things like how long does it actually take for a pull request to be reviewed, or things like this. And a lot of things, from then you also measure the resolution time of incidents, or deployment frequency.
[00:30:35] Serhii: It doesn't matter much what you track. It's more about why you track this, because we have a hell of a lot of metrics available in this tool I mentioned, DX. We also use custom ones. But we sort of balance between standardized metrics for all, and then fine-grained, tailored for the teams, and also allow these teams to contribute to our metrics codebase and configurations, to make it easier for them to use. So that would be my answer, probably.
[00:31:08] Host: That's a brilliant way to finish things. Fantastic. Thank you. Thank you so much for your time, everyone. Huge round of applause again for Serhii, thank you.
