Building a Recommendation Engine at ESPN Without Slowing Delivery: AI as Your Team's Multiplier
Checking session availability…
Hang tight while we load the latest updates.
How do you modernize a recommendation system while the product keeps shipping? In this talk, Katerina will share how the team at ESPN evolved the app's personalization stack from heuristics to an ML-first recommender system, where each stage delivered value on its own while building toward the next.
She'll show where AI tooling acted as a force multiplier, compressing iteration cycles across documentation, evaluation, and debugging, and helping them spot impact, so they could focus on the right levers.
You'll leave with key takeaways on incremental modernization and where AI can truly accelerate your team's velocity.
Building a Recommendation Engine at ESPN Without Slowing Delivery: AI as Your Team's Multiplier
Katerina Zanos at UXDX USA. Video: https://youtu.be/njbm74zmPnk
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Everyone uses AI, few ship faster
[00:00:08] I'm Katerina, and I lead the ML platform for recommendations at ESPN, which means I spend most of my days thinking about how to surface the right sports content to the right person at the right moment. Today I want to tell you a story about what happened when my team set out to rebuild that platform without stopping the world to do it.
[00:00:31] But before I get into what we built, I would like to ask you all a question. Is there anybody in this room who is not using AI in their workflow, in their work or in their team? Can you please raise your hands? And there's no shame in not using AI or using AI. I just want to see who is using it and not using it yet. Okay, I think the light is a little glaring, but I don't see any hands in the audience. So I'm going to assume that all of you here are using AI, in your teams or on your own.
[00:01:10] So now, second question. Hand on your heart, you have to be truthful. I'm assuming all of you are using AI. Who here can say that AI has actually made your team measurably faster at shipping things that matter? Can you raise your hand? 1, 2, 3, 4, 5, 6, maybe a 7. Okay, so I think it's seven people in the audience. I'm sorry, the ones over there I can't really see because of the light. But if we assume that everybody here is using AI and there are only seven people who see measurable impact, I think there's a real gap here.
[00:01:57] And that's what I want to talk about today. Because every team I talk to, whether it is within Disney or outside of Disney, is actually using AI, but not every team is getting faster. And the difference between those two groups is what this talk is about.
[00:02:18] Here's what I see happening across the industry, and honestly, what we saw happening on our own team before we got our act together. AI generates code at speed, but the fact that something runs is a really low bar. Running code can still be wrong code. Teams adopt AI overnight, but the architecture decisions, the choices that actually determine whether you can ship something successfully in 3 months or in 3 years, those don't change just because you have a faster coder. More code gets written, but delivery doesn't actually accelerate.
[00:03:02] Now, I want to let you in on a little secret. When it comes to building software, the bottleneck was never about coding speed. It was always strategy.
[00:03:18] A few months ago, my team set out to rebuild ESPN's recommendation platform so they can better personalize the user experience on the apps. We were five engineers, myself included. No pausing delivery, and no big bang rewrite until we shipped impact. The thing that made this possible was AI, but not in the way most people frame it. AI didn't write the strategy. AI executed against it. Once we had the right plan, AI became the engine that let a small team do what would normally take a much bigger one. That's the story I want to share with you all today.
What makes sports recommendations different
[00:04:02] But before I get into the story, I should provide some context here, because I'm assuming most of you don't work in sports media. In the sports domain, there are some weird properties that shape the decisions we made in our team. So, three things that you should know.
[00:04:16] First, live events. When the Knicks are playing, content relevance changes by the minute. The clip from 2 minutes ago might be the most important thing on the platform for our users, and a few hours later... Whoops.
[00:04:33] [laughter]
[00:04:33] And a few hours later, it could be completely irrelevant. Second, fandoms are deep and identity-level. People don't casually like teams. They are fans. So personalization isn't a nice-to-have feature here. It's the entire product.
[00:04:54] And third, editorial voice matters enormously. Sports media, and especially ESPN, has decades of expertise in surfacing the right story at the right time. In this case, heuristics, which include domain rules, editorial boosts, manual curation, actually go really far in this space.
The system the team inherited
[00:05:15] Which brings me to the system my team inherited. This is roughly what the recommendation system looked like when I joined ESPN last September. On the left, you can see a bunch of signals, some personalization-driven, like a heuristic user embedding based on how the user navigated and engaged with content on the app. User favorites, which is basically the teams, the athletes, the sports that a user follows. Implicit user favorites based on the user engagement, again. And some signals that were not personalization-driven, like popularity, recency, trendingness, and those editorial boosts that I talked about.
[00:05:57] All of those signals fed into a single Elasticsearch query with weighted combinations. For those of you who are not familiar with Elasticsearch, it's basically an index store that allows you to retrieve a lot of documents real fast with some features, some attributes. So all of those signals were fed into that query, and out the other side came a ranked slate of content our users consumed on the app.
[00:06:29] And here's the thing. This worked. It actually shipped real value to users. The editorial team also had powerful levers. The heuristic logic was interpretable. You could reason about why a piece of content surfaced. But like every system, it had a ceiling. And we hit it.
[00:06:52] There were three problems with this system, and these are going to sound familiar to anybody who has worked on a recommendation system that relied on a heuristic ranker. The first problem was impression skew. An impression is basically when somebody sees a piece of content on the product. The system gave most of its impressions to a small subset of our inventory. There was a long tail of content, and we have a lot of long-tail content at ESPN, which was getting starved, was never shown to users.
[00:07:30] The second problem was the universal rules. All of these editorial boosts that the editorial team put together, every filter, every heuristic, were applied to every user the same way. Whether you're a hardcore football fan or a casual fan, basically, you got the same overrides. So we were missing what each user actually wanted when they came to our products.
[00:07:57] And the third and last one was that content retrieval and content ranking were coupled together in that Elasticsearch query. So if you wanted to improve what type of content we retrieved for a user, you also had to impact ranking, which made it difficult to address each layer separately. The system worked. It just couldn't learn what each user wanted, and it couldn't be improved incrementally. So when I joined, the first thing I did was to look at what the team had been working on to fix it.
The two-tower model that ran but wasn't right
[00:08:34] Okay. Now, I want to be honest with you about a project that was already underway when I arrived, because it's one of the most useful lessons that I can provide during this talk. One of our engineers had been working on a two-tower embedding model. For those of you who are not working on recommender systems, or not familiar, the two-tower model is a neural network model that's basically the industry standard for content retrieval. So it's the right model for this kind of problem.
[00:09:05] And the engineer worked on it for about 3 months. They used AI coding assistants extensively to build out this project. And let me be clear, this is a complex project. It has several parts. It has data you need to pull, features you need to create, the model architecture you need to build, to design; train the model, figure out the loss function, and then also figure out how you're going to do inference later, how you're going to generate the recommendations for each user. So it was a very complicated project.
[00:09:39] The engineer generated a really impressive amount of code using those AI assistants, really fast. And the model ran, the model trained, it learned, and on the surface it looked like the right thing. But when we actually looked under the hood, about 80% of the code wasn't necessary for what the two-tower model was supposed to do. The AI had scaffolded a much bigger system than what was needed, which meant there was generated tech debt we had to maintain from the get-go.
[00:10:13] And more importantly, feature generation was coupled directly into model training. Those of you who are engineers or work with machine learning models probably already know that this breaks scalability. It makes future iterations painful, and it's basically the kind of thing you don't want sitting at the foundation of your platform.
[00:10:37] But I also want to be clear about something else here. This is not a story about a bad engineer. It's a story about how AI behaves when you point it at a goal without giving it architectural constraints. It will produce something plausible that runs. And the closer it gets to done, the harder it is for us to step back and ask whether it should exist at all.
[00:11:06] So we made the decision to pause this project. And that decision came from one realization. AI without a strategy doesn't slow you down. It actually speeds you up in the wrong direction.
Learning the landscape in days
[00:11:25] So what does AI as a multiplier actually look like? Let me show you. As I said, I joined ESPN in September. New team, new code base, new domain. I've basically worked on recommender systems for all of my career. I used to work at The New York Times building out their personalization platform. I also worked at Meta. But I had never done sports recommendations.
[00:11:50] That's where AI was really transformative for me. I was able to explore the new code base of my team really, really fast. I used AI to navigate parts of the system I had never seen. I asked it to explain components, trace data flows, summarize what certain files were doing. This is the kind of work that used to take weeks of pairing with other senior engineers and asking them lots of questions. So AI saved their time and my time.
[00:12:21] It also allowed me to do fast data exploration. I was able to put together SQL queries against unfamiliar telemetry data. I was able to understand the data schemas we were using. I found patterns in the data, and I started stitching together the story of how our users use the ESPN app.
[00:12:39] And finally, it helped me and other team members onboard new tools. One of the tools that we agreed to use in our organization was Databricks, and that was a tool that was new to literally everyone on the team, including myself. So instead of weeks of struggling, AI was an always-available expert that showed us how to deploy pipelines to Databricks effortlessly.
[00:13:04] And I'm speaking for myself now when I tell you how this boosted my productivity, but that was the deal for other people in my team, who also explored data, also explored unknown code, and also onboarded onto new tools. So this learn-the-landscape phase that usually takes an engineer weeks or even months, we compressed it to days. And that compressed timeline is what led me to spot the right place to act first.
Thompson sampling as the first win
[00:13:33] So in month two after I had joined the team, one of the first things that I wanted to ship was a Thompson sampling layer for content exploration. For folks here who are not familiar with Thompson sampling, it's basically a simple model for controlled content exploration. It gives content that we don't know much about, because it hasn't been seen by users, some impressions, so we can learn in a controlled way whether this content is good to show to more users or not. So it's basically a way to break out of that impression skew problem that we had, which I discussed earlier.
[00:14:15] The result was a meaningful lift in watch minutes. So this was real product impact, shipped in weeks. Now, the reason I'm walking you through this isn't because Thompson sampling is a fancy model. It actually isn't. It's because of how we shipped it. We had a clear problem, that impression skew, which I was able to spot using AI while investigating our data. We had a scoped solution, and we followed a principled approach. AI accelerated the prototyping and evaluation phases, but we owned the decision to ship Thompson sampling first.
An incremental roadmap
[00:14:58] Now, the Thompson sampling win was an important one, but it was a small thing in the grand scheme of things. The big thing was the architectural rebuild we needed to do underneath, and that's where the multiplier really showed up. The rule we followed was that every piece had to be independently valuable. We weren't going to ask the business to wait for that big bang rewrite of the platform. We wanted to keep shipping value. And as we shipped new things, those would unlock the next thing to build.
[00:15:36] So this incremental roadmap looked something like this here on the slide. The Thompson sampling algo was the first step. We then focused on developing data pipelines that allowed us to have training data to train real models. And while we did that, we also experimented with and started testing an XGBoost ranker. Then we also worked on decoupling the content retrieval and content ranking from that single Elasticsearch query, plus adding a feature store and inference server to enable robust model deployment and serving at scale. All that was done by roughly one engineer per project, in parallel.
[00:16:19] And finally, the two-tower model came back. I'm going to talk about that later. So as you can see, each block in this roadmap ships independent value. Each step unlocks the next. That sequencing, that's the part that took strategy.
Data pipelines and an XGBoost ranker
[00:16:36] So now let me walk you through how AI enabled us to do that. For our data foundations, we developed medallion architecture data pipelines. We had raw telemetry data coming in, and clean data sets coming out, with feature and training data ready for downstream models. This is normally a separate team's job. If we didn't have AI as an accelerator, we would have had to either pause the machine learning work we wanted to do and build the data pipelines as a team, or wait until another team in our org was ready to pick up that work stream. Instead, we did it with two machine learning engineers and one data engineer. Two months from nothing to fully productionized data pipelines. In parallel, we kept working on other projects and shipping value.
[00:17:27] Now, the next thing: the fact that we now had data allowed us to focus on building a ranker model. And the ranker here was the first real ML model we put into the platform. The ranker we decided to work on was an XGBoost model. Again, if you're not familiar, XGBoost is a simpler model. It's not a neural network. It's basically using decision trees to make a prediction about something. In our case, whether the user will engage with content or not.
[00:17:58] Now, here you might legitimately ask me, "Well, Katerina, why did you choose XGBoost? Why didn't you choose a fancy deep neural network? You worked with these models in the past, and you have AI as your accelerator here." And the answer I'm going to provide you with is that, again, it's about incrementality. As we're building out our new platform, it is much easier and more straightforward to ship a simpler model first to replace the heuristics we had in place than to try to deploy a much more complex model. And along the way, testing XGBoost, we also learned a ton about the nature of our features and how they contribute to predicting user engagement. This is the kind of decision AI is not going to make for you.
[00:18:53] We picked XGBoost deliberately because it could deploy inside our existing recommendation API. We also didn't have an inference server or a feature store yet. A neural ranker would have required all of that before we could ship. XGBoost let us ship ranking impact first and invest in serving infrastructure next. So, the same two machine learning engineers, three months from zero to testing and production, including building the features through the new pipelines I showed you. AI helped us write the model code and build evaluation harnesses, but the choice of what to build and when, that was a sequencing decision a human had to make.
Decoupling retrieval, feature store and inference server
[00:19:36] While that was happening, we also did two more things. First, we broke that monolithic Elasticsearch query apart. Now content retrieval and content ranking were separate. Content retrieval was independent sources, each contributing candidates to a separate ranking layer, which is the industry standard in recommender systems, basically a two-stage architecture.
[00:20:03] Second, we built the feature store and the inference server I talked about. These systems were developed by two engineers in two months. This was the infrastructure that unlocked everything sophisticated downstream that we wanted to do. If you remember the two-tower model anti-pattern from earlier, this is the foundation that this model actually needed to be able to scale in production. Instead of feature generation tangled up in training, we had a clean, scalable place for it to live.
[00:20:33] So across every step, AI removed blockers. Nobody on our team knew Databricks. With AI, that became a non-issue, not in three months, but in days. DevOps tasks that would normally block an ML engineer for days: getting continuous integration right, continuous deployment working, debugging a Kubernetes config, writing a Terraform module, or figuring out what framework to use. This was possible to tackle in hours, because AI is great at well-defined infrastructure work nobody actually wants to do.
[00:21:10] Code reviews on code bases we didn't write: AI was an always-available context provider. Offline evaluation tooling built in days, not in weeks. Understanding massive telemetry data, being able to craft SQL queries to explore the data at 10 times the pace. And what I want you to notice here is the following pattern. What AI did was remove the friction that slows small teams down. It was not used to define the strategy or make decisions for us.
The two-tower model returns
[00:21:42] Which brings me back to the model we paused. Remember that two-tower model from earlier? We brought it back. We stripped out about 80% of the unnecessary code, decoupled feature generation from training, and we re-platformed it on top of the feature store and the inference server we built. Same model. The only thing that changed was that this time it was sitting on a real foundation with a real plan.
[00:22:10] So, 8 months in, here's where we ended up. Thompson sampling shipped; we had a sustained lift in engagement. Data pipelines productionized. XGBoost ranking in production. Two-stage architecture in place, content retrieval separate from content ranking. Feature store and inference server operational. The two-tower model re-platformed on real foundations.
Compress the left side, don't outsource the right
[00:22:33] So, real quick, let's name what AI actually did and what it didn't. On the left side is the work that AI compressed for us. I'm not going to go over it again, because I talked about all of these items during the talk. What I want to focus on is the right side, which shows you where humans were irreplaceable. Deciding what to build and in what order. Spotting structural flaws in AI-generated code. Scoping problems before solving them. Picking the right metrics and evaluation design. The decision to pause the two-tower project and build foundations first. Knowing the difference between "it runs" and "it's right."
[00:23:20] So if you walk away with one mental model from today, I would love it to be this one. Use AI to compress the left side. Don't outsource the right side. AI is a multiplier of your engineering strategy. Feed it a clear plan, and your team ships in months what used to take years. Feed it a vague goal, and you can build the wrong thing faster than ever.
[00:23:50] So if you go back to your teams on Monday and want to do something about it, I would suggest three things. One, audit how your team is actually using AI right now. Is it multiplying a clear plan, or is it generating code without one? The two-tower anti-pattern could be happening on your team somewhere. Go find it before it costs you 3 months.
[00:24:19] Two, find your Thompson sampling. What's the smallest, most principled change you can ship this month that proves your direction? Don't underestimate small wins for building credibility around the bigger architectural calls.
[00:24:37] And third, name the foundation you are skipping. Every team has the unsexy infrastructure work it keeps deferring. That's almost always your real bottleneck. AI is going to make the temptation to skip it stronger. So resist that, and instead use AI to address it.
[00:25:00] And I will leave you with this. As you saw at the beginning of the talk, every team is using AI. I'm sure you already know that. But the teams that are getting faster are the ones using it to execute against a real plan, not the ones using it to generate more code. Thank you.
[00:25:18] [applause]
Q&A
[00:25:20] Host: Thank you. It's either the temperature in here or the concept of all the wasted time and energy that is giving me the chills. So you're basically saying I can't just throw AI at any problem and code a bunch of stuff and then maybe I'll come out with the right answer. And 80% of the code just useless.
[00:25:44] Katerina: Could happen. Yeah.
[00:25:45] Host: Okay, so this is great. I think all of us have this feeling about AI creating slop, but that's usually the slop we see. And the slop we don't see is terrifying. And this also brings the whole junior engineer, junior designer, junior product manager problem to the forefront. We'll go beyond that. Let's go to the audience questions.
[00:26:10] Katerina: Yeah.
[00:26:12] Host: Let's see. This is a very specific question. "What do you do when you have a results page with ads? Advertisers pay for spots. How do you blend this in terms of content?" Are you responsible for that, or are they separate areas, ads and content?
[00:26:28] Katerina: Well, I'm not responsible for the ads, but right now in some parts of the product we do serve ads, and in other ones, no. But we take it into context, let me just say that. So the question says, "Advertisers pay for the first spot, so that means the best results are hidden below the advertisement. How would you approach that?" I see. Good question.
[00:26:56] Katerina: Well, I would say first that ads are also content, and we should surface the right ad to the right user. If that happens, it's actually a great user experience, and it's not disruptive at all for the user. It shouldn't be. So the ads team is responsible for great personalization as well. I strive to surface the editorial content, the content we produce, to the right users and make that a great experience for them, but that should work with the personalized ads that the other team is surfacing. Of course, that's our business. We want to surface ads, but that should be a great mixed experience with what our personalization produces. So I think it's a balancing act between the two. I guess that's the short answer.
[00:27:48] Host: Okay. "How did you get your team the authority, the autonomy, to bypass DevOps processes in launch?"
[00:27:56] [sighs and gasps]
[00:27:57] Katerina: We didn't bypass DevOps processes. Okay. Yeah, no. We didn't. AI just helps us to do that faster. I've been working as an engineer for more than 10 years now, and DevOps for me was a major pain point, especially because I just wanted to work on building cool machine learning models, and then you just have to deploy this thing. Some people outsourced that in the past. I've seen that. They were like, "Okay, I built my model. Now you, software engineer who is the DevOps expert, go ship it to production. Not my problem." Well, now with AI, that's not the case. It's like, "Okay, you can own the problem end to end." And guess what? DevOps work is actually really straightforward. So you can have an AI, you tell it this is how I want it to be deployed, you prompt it, and it helps you put it in production. So that's how we addressed that problem.
[00:28:52] Host: I think you have a very smart engineer here. I think DevOps is kind of hard.
[00:28:55] Katerina: It is. [laughter] It is. It is.
[00:28:57] Host: Because with those models you have to have the data, the model, the back end, and the front end. It's so much more to ship. Okay, great. There are a couple of aspects of strategy here. One of the questions was, did you use AI to come up with the strategy at all, or to help you, or to be a thought partner? And the other question that goes with it is selling a simpler version first, because obviously they went to this two-tower because somebody said, "Let's just do this." How did you bring them off a ledge? And then how did you figure out that strategy?
[00:29:34] Katerina: Yeah, absolutely. I definitely use AI as a way to brainstorm. But primarily, you know how you would go to Google and find a lot of sources, and then you would have to go through them and compile it. Now with AI, you can still pull all of that stuff and then summarize some of the sources and see, "Okay, is there anything interesting here for me?" And then it's a more guided way for you to do research and decide, "Okay, where should I go?" But I always treat AI as a brainstorm partner with caution, because as you probably know, it can be very affirmative. It can say, "Yeah, that's a great idea. You should go do that." And then you try it, and it's actually not a good idea. So that's my two cents. You can always try and brainstorm with it.
[00:30:22] Katerina: Then the other part: I definitely see why the engineer went for the two-tower model. It made total sense. As I said in the talk, it's the industry standard. It's just a matter of laying out a plan and basically deciding to do the hard part first, to get this productionized, rather than just building the model without the other important components. Which is the third takeaway that I hope you remember from the talk. Don't skip the unsexy stuff just because you can build the cool stuff faster. Use AI to help you accelerate that process, but it's still important and you still need to do it.
[00:31:05] Host: Was there a part of convincing people to move this way that came from just your authority of having done this before?
[00:31:12] Katerina: You think? Yeah, could be.
[00:31:14] Host: Because I think that's what everyone here is trying to figure out: how to convince stakeholders to go in a different direction sometimes, when stakeholders are just plowing forward with this AI, AI, AI, and you want to come out and bring maybe a smaller piece of value faster. Where do you get that ability to convince those folks?
[00:31:28] Katerina: Yeah, sure. Definitely when you have a title, you're like, "Okay, I'm going to lead the team." But honestly, when you work with people, it's not like people will follow your suggestions and your plan just because "Oh, I came here to lead the team, do this." For me, the main way to do this and convince stakeholders and your team is with data. For instance, I had one of the more junior engineers, while we were developing the XGBoost model, and we had a difference in opinion. I was like, "We should do it this way." She was like, "No, no, no, we should do it this other way." I'm like, "Okay, I think this is a difference of opinion. Prove me wrong. Go test this thing, collect the data, show me one is better than the other, and I'm totally happy with you shipping your solution." So I think it always comes down to the data. When you're not sure, just go back to the data and prove it's better.
[00:32:23] Host: Yeah. "When people go to the AI for strategy and they don't have the experience, how do you teach them to correctly leverage AI?" When you spot people using it in a suspicious way, and maybe they don't even realize it, how do you as a leader go through that coaching experience of bringing that person around?
[00:32:42] Katerina: I ask a lot of questions that I think will bring up the gaps in the process, or maybe the blind spots that this person could have had when they were using AI, putting together their solution. And I think everybody should do that with each other, especially now with AI again, because it's so easy to put together something. I think everybody has certain knowledge, certain expertise they can bring to the table, and they should just ask questions.
[00:33:15] Host: I think it's how you ask the question. I have a sense that your employees probably like you, because I can see how you would ask those questions. So let's give a round of applause for Katerina.
[00:33:23] [applause]
