Building AI-Driven Workflows at Airtable
Checking session availability…
Hang tight while we load the latest updates.
AI is transforming what product teams build. Airtable's engineering team has embraced this transformation while navigating the unique challenges of building AI capabilities that empower users to create their own dynamic workflows. In this session, Jimmy Hillis shares practical insights from Airtable's journey integrating AI into their platform. Key topics will include:
- Understanding changing user needs: How Airtable is building AI-native use cases while keeping workflows intuitive and accessible across different use cases
- Rethinking prototyping for AI: Why teams must show actual AI-generated results rather than static mockups to effectively design, test, and get meaningful feedback on AI features
- Cultivating technical understanding: Strategies for encouraging teams to actively use AI tools, creating a culture where everyone from designers to engineers can contribute effectively
- Balancing automation and control: How to ensure AI enhances productivity without adding unnecessary complexity, particularly within enterprise environments with unique security and compliance needs
- New collaboration patterns: New motions for product, design, and engineering teams when developing AI-powered workflows
Building AI-Driven Workflows at Airtable
Jimmy Hillis at UXDX USA. Video: https://youtu.be/2dEoBaz5K6g
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
Airtable and the AI products it has been building
[00:00:10] Hey everyone. I'm Jimmy, and I'm the head of engineering over at a company called Airtable. If you haven't heard of us, we're a no-code, low-code platform that really enables less technical and non-technical people to build applications, generally for their business. The way I describe what we do is that there are so many people that lack the access or the ability to have tools, software, workflows, automations, whatever it is, that suit not just their specific function, but also the culture of the team that they have and the way that they want to work.
[00:00:43] Airtable really empowers those people. We want to make it so that you don't need access to any of us here in the room. We're expensive, there's not very many of us, and we often have to say, "No, we can't help you," because we're building the products that our companies are. Airtable really wants to provide those folks the avenue to express what they need to and make their jobs a lot more effective through software.
[00:01:05] I've been at Airtable for five and a half years now, going on six. I joined when we were pretty small: about a 100-person company at that point, about 30 engineers. We had one or two designers at the time who were writing a ton of code. Today we're about 750 people, and our product development org is about 250 individuals. We've got PMs, we've got designers, we've got engineers. I've been along for that entire journey, and it's been a really exciting thing to see the team grow and be a part of the culture: the types of things that we value, the way in which we build product, and also those processes and what it looks like to build a product from zero to one in an ever-shifting industry.
[00:01:47] As it relates to this, very specifically, we've been building a ton of AI product for our customers as well. Over the last couple of years, just to set a little bit of context, we've had about 30 to 40% of our organization building AI features for our customers. To give a couple of examples of the sorts of things we've been building: in early 2023 we started a bit of a grassroots effort, like, "Hey, let's grab this LLM technology and see what we can do with it."
[00:02:17] We launched what we called our AI fields, which are essentially the ability to write a prompt, enrich it with the data that you have within your applications, and get responses based off that. There's a lot of really powerful things that you can do. It provided a lot of indirect and direct access to the models that are out there to a lot of our builders. We then went and shipped something we call Cobuilder. Cobuilder is the ability to build your applications in natural language.
[00:02:42] We can take away the complexity of understanding what a schema is, how relationships work and how you should build a good UI. You can just come in and say, "Hey, I'm so-and-so. I'm on the marketing team at this company. We really need to build a bunch of campaigns that we're going to ship internationally. We need to deal with differences of language and time zone," and whatever else it is. It'll build that application for you, so that you can track, do your critiques and all the rest of it in a way that's a lot more natural, really removing even the complexity that we have in trying to build for these low-technical users.
[00:03:16] Lastly, we just recently released our assistant. This allows you to interact with that data in a more abstract way. You can do analysis, you can ask questions about your various data, you can continue to build things, and it gives you an opportunity to generate memos, look at insights and gather a lot of that in collaboration with the application directly.
[00:03:36] We've been building a lot of these features at Airtable. What I'm coming to you today to talk about is how we've had to shift the way that we build products: what that looks like at a team level, the different processes and the ways that we've changed how our product, design and engineering individuals work together. I'll also talk a little bit about the individual changes that any person on our team has had to make to adapt to this.
[00:04:04] I want this to be a bit of a sense of, "Here's some learnings." I'm going to give some really concrete examples of things we've done, just to give folks a sense of ways in which they can approach this. It's also probably a chance for you to reflect on some of the ways in which you've had to change what you're doing, or why some of the things you've done maybe don't work and why new techniques will.
Three industry shifts in 20 years
[00:04:21] To get into that, I want to set a little bit of context. I've heard a couple of people mention some of this too, but I want to give some context on the really big industry shifts that have happened over the last 20 years. As you can see on the board, there are probably three major shifts that have come into product development in the last 20 years. I think they officially say 2004 is when Web 2.0 came in.
[00:04:47] This was really that shift from prepackaged software (you buy software, maybe it's on-prem, you buy something once, you maintain it yourself, you run it, or you just install it on your computer) to cloud software. This is your Facebook, this is your Salesforce. I guarantee you that all of you are likely building this type of software on the web today. This was a huge shift. It changed everything about how we do software.
[00:05:12] You went from these year-long cycles where everything had to be perfect and you would deliver it, to "Hey, I make a change every day," every week, every month, depending on what it is. That massively changes everything we do. You talk about these agile approaches; none of that was really possible when it was, "You have 12 months, you're going to ship this thing, you're done, and then you're going to move on." A lot of us have worked in this industry through this.
[00:05:33] The one that I imagine a lot of you have seen the shift of, if not in working then definitely in terms of interacting with it, was when mobile devices became a standard approach. The iPhone release, 10 or 12 years ago, whatever it was, really shifted the way in which all of us build products. I was already a software engineer working pretty heavily in that, and I remember it was such a large shift: "OK, what do we do with these things? Do I just build the same app? Do I just make it look small? Does it need to be a totally separate app?"
[00:06:02] Facebook famously didn't move directly to building iPhone apps, specifically because they thought they wanted to stay on the web. There's been a huge amount of shift that's come here. I think it honestly took five or six years for us to actually get a sense for when it should be on a phone, when it should be responsive, when it should do different things. I think most of us are probably in a place where we understand it and have good best practices. When you come to a new product, you have good intuition, and you have processes that will enable you to build for either or both, depending on what it is.
[00:06:30] Then these LLMs come out. It's been two-ish years since most of us started to interact with them, and we're seeing the same type of shift here, whereby essentially all of the best practices that we're used to building with don't really apply particularly well. I promise you that if I'd gone back five years and asked any product developer in the room, "Do you think a chat box is going to be the right user interface for the biggest products in the world?", you'd have said, "Absolutely not. That's the last thing on the planet that anyone wants." And yet here we are, where a lot of these products are essentially a chat box.
[00:07:05] The thing that's really important to realize is that you do need to take that step back and recognize that a lot of what you were used to just doesn't really apply anymore. How do you figure that out? How do you work within a mode where this stuff is changing?
Root it in value, and have your own aha moment
[00:07:19] I want to talk about two high-level things before we get into the process. One, and this shouldn't bear too much repeating, but I really think it does: ultimately you work for a business, you're building something for customers, and you really want to find value. Everyone inherently knows that what you're doing is building for value. But it becomes really easy, especially right now, to get caught up in a lot of the shallow use cases of AI.
[00:07:45] There's a lot of stuff that's just really cool. You write some prompt, and it's pretty miraculous that it can do something. You can show how much progress you can make in a couple of days. But the reality is that if you can do that in a couple of days, so can all of your competitors. That's not going to be a differentiator. It might not even be relevant to the strategy of your business or the problems that you're trying to solve.
[00:08:05] It's really, really important right now that while you should play (and I'll use that term a lot today), you root that in the actual value. You need to have some unique differentiator in your business, in what you provide and what you do, and really apply AI to that, and not get caught up in the "me too" of, "Oh, I can go and do that as well. We can have an image generator." Does that matter for your product? How does it actually fit in? What does that actually look like? Really come back to how you're providing unique value, and what you have to give to your customers that really aligns with the technology.
[00:08:38] The next one is more of a personal one, and I think it's really important not just for you in the room but for your teams especially. You need to ensure that you and everybody working on these products have their own aha moment when it comes to using this technology. I found that there's been a reasonable amount of skepticism: "Do I really need this? Is this just a fad? Is this something that's going to come and go?" It was relatively difficult early on to get our teams really adapting to using LLM technology, to investigate and play with these things.
[00:09:10] The thing that works every single time is when that individual has their aha moment with something that they do. If you're an engineer, you go and write some code, and it can find and fix things for you. Or you're using some tool, and it enables you to do something in 20 minutes that would normally take eight hours. Whatever that is, make sure that you and everyone on your team are using the various tools to unlock that moment where you're like, "Oh, actually this is pretty special." At that point it really changes your mindset. It enables you to engage more effectively in building products this way.
[00:09:44] You want to keep having these; I continue to have these aha moments. Even last week, I was building out a personal coding project. I don't write code very much anymore by any means, but I was doing this weird thing in macOS and I just couldn't figure it out. I spent an hour looking on the internet. I've been on the internet since I was 10 years old; I'm pretty good at searching for things. I couldn't figure out how to do it. It wasn't really a code thing; it was some weird compliance thing.
[00:10:09] So I threw all this stuff into Claude and asked this really broad question, and it came back not just with the answer, but with a step-by-step guide: "You need to go three steps deep into Xcode and click this weird archaic button that'll enable you to do this." That sort of moment, where you can't even find these answers knowing what you're doing, is really impressive. Making sure that you continually have those will get you motivated to do similar things within your own products.
Learn as you play, not as you ship
[00:10:34] I touched on this a little bit. Where I want to start from, in how our teams have started to engage with this, is the fundamental shift in how product design has changed with AI. The thing I want to get across at the top level of all of this is that you want to shift from learning as you ship products to learning as you play.
[00:11:00] Everything that you've built up and understand about best principles of how people will interact with your products likely doesn't apply to a lot of these technologies. You also don't know how they work. You've spent 10 years on a phone; you know what you want to do. You know you can pull to refresh. You know that people are going to interact in a certain way. That doesn't really apply here.
[00:11:18] The thing that I would urge everyone to do is shift your process and shift your mindset to playing earlier. Get out of Figma. Get out of your design tools, out of writing docs, and really get into playing with the tools. A lot of these LLMs are non-deterministic. You can't guarantee what is going to happen, or what is possible, unless you start to do that. The sooner you get into playing with the technology, the sooner you're going to find that moment, that little spice that's going to make a really massive difference to the way that you're going to plan and develop your product after that.
The old product process and the new one
[00:11:49] Very specifically, I want to give an example of what our process has looked like and what it looks like now, and talk through a little bit of the nuance of that. This is within our product development teams, and I assume all of you have something roughly similar. There's a product process. You have a PRD or whatever: somebody sits down and says, "We have this customer persona. They have this type of problem. We think it's worth this much to solve," whatever that looks like. Somebody's written a document.
[00:12:18] You probably take that to some form of review. Somebody says, "Cool, that sounds great. Let's go and explore that." Then you go into the design phase. You probably do some user research, maybe put some prototypes together, make some drawings, maybe do a clickable prototype and put it in front of customers, and iterate on that. But ultimately you probably sit down and think for a while: "I understand this problem, and this is how we're going to solve it." You come to a bit of a conclusion early on, before you get into a mode of what is possible.
[00:12:45] You test that, you start building, you iterate, whatever. You probably ship the product; maybe it's an experiment, maybe you ship it to GA. Hopefully you're in a position where you can iterate on it based on user feedback and user impact. But ultimately it's this process where you know what you want to do, you know how you want to solve it, and then you ask, "Did it work or not?" and maybe iterate on that.
[00:13:03] I think the biggest shift now is that you really need to start with your strategy. What are you, as a company, as an org or as a team, trying to accomplish, and who are you trying to accomplish it for? Then go and play. Go and grab some LLMs, write some prompts, grab some dummy data, do whatever you want. Build the hackiest version, completely outside your normal product development stack, to validate what is possible.
[00:13:30] Do not go in there with your preconceived expectations, like, "Well, I think this is how it's going to work," because what you're going to end up doing is spinning a lot. You might do some designs, whatever it is. It's going to get to engineering at some point, and it's not going to work. Or, and I think this is the important part, you're going to miss the possibility of what was there if you had just started to play early on.
[00:13:46] I want to urge everyone in this room: prompt engineering is not an engineering problem. It's an everyone problem. Hopefully nobody needs to write prompts five years from now, but today you do. If you're not in there, if you're not writing these, if you're not exploring what is possible, you're going to miss the potential of it. Really get involved in that early on. Work very cross-functionally. I don't think your function matters at this point. It really is about, "What are we ideating on? What can work? How can we figure out what is possible?"
[00:14:15] At a certain point, you're going to have this really, really special thing. You're going to be surprised by what is possible and what you can get done. Then you can go into that iteration and design phase: "OK, cool. We know that it can do this. We know various inputs can provide these sorts of outputs. How do we design on that?" Whether that's just your user experience, or maybe broadening the quality (and I'll talk quite a bit about quality as we get there).
[00:14:37] Then you get into what I'm now calling productionization. What you'll end up seeing is a lot of people talking about prototypes: "Go do a bunch of quick prototypes. You can prototype really quickly. Prototype, prototype, prototype." What hasn't changed between before and after is that the gap between a prototype and productionization is still really wide. You can make something that looks and feels really good, but it's going to take a lot of effort to get it to production. Making sure that you call that out enables you to spend the time on what that looks like.
Cobuilder: why playing first mattered
[00:15:07] I want to give one concrete example of why this is important too. I talked about how our Cobuilder enables you to build applications with natural language. We came to this project saying, "We want to use this new technology. We think it'll be really good for building." Can we somehow get to the point of view of what it would be like for an individual to talk to somebody who's already a really good builder, and for them to say, "We got you, we'll build it for you"?
[00:15:32] There was actually a lot of skepticism on the team about this. Building's really hard. We have so many different types: there's database schemas and user interfaces and layouts and workflows and automations and all these sorts of things. "There's no possible way. What we're going to do instead is, we know how ML works; it's really good for recommendations." The idea that they had was, "We'll create a bunch of pre-canned apps for marketing teams and for product teams, ask them a couple of questions, choose the right one, and they can start from there."
[00:16:01] A pretty obvious sort of thing, rooted in the reality of the technology that we'd used. It was too small. What we said to the team was, "Go and spend a couple of days and see how you can write some prompts and engage with this from the root level: can you actually develop the schemas we need, the user interfaces and all the rest of it?" The team did this after, to be perfectly honest, quite a bit of pushing.
[00:16:24] What they came back with in a few days was miraculous, because it could actually develop the entire application for you in quite a lot of detail. It had already ingested a ton of stuff about Airtable, and, would you believe, database schemas and relations are a pretty standard concept that these LLMs already understand. It was able to come back with a very full-featured app based on pretty high-level prompts that you can go back and forth on.
[00:16:48] It allowed us to completely change the way we were going to build our Cobuilder experience, to a level that is pretty awesome. I'm sure you've all seen a bunch of this stuff recently, and that only comes from the ability and the willingness to get in there and play to see what is possible before you decide what is the right thing to do.
Productionizing AI: quality and evals
[00:17:10] At that point, you have an idea. I'm traditionally an engineer, so I wanted to take this on a bit of a side tangent to talk about productionization. What's different? What's hard about building AI products? What are the sorts of things we need to shift in both the culture and the way that we develop and work through these problems? There's a lot of practical advice that probably a bunch of us are running into right now, which is similar, but also through quite a different lens. I'm going to talk through some of the processes you'll need to go through to productionize and ship something at scale using a lot of LLM technology.
[00:17:43] The first thing I'll start with is that humans are pretty non-deterministic, and LLMs are too. Generally speaking, when you write code, you put in input X and you expect output Y. If it doesn't happen, you have a bug, and you fix it. Computers do what you tell them to do, ultimately. LLMs aren't quite that way. They hallucinate quite often. They can take a broad range of different inputs, so you can never be really confident about what the output is going to be, and as a result the quality can vary wildly in what you're building. I'm sure you've seen this: you can run the same prompt multiple times and get completely different answers. That's not super helpful.
[00:18:20] You really want to focus on that, and the answer to it, which you've probably started to hear about and which I think is really important, is evals. Reductively, they are tests. You say, "Given some set of inputs, I expect it to provide some set of outputs." We've had these groups of tests for a long time; this is not unique. But I think it becomes even more important, in that it provides you the constraints and frameworks by which you can develop your products, to ensure that they solve the types of problems for the varied types of input that you're going to get from users.
[00:18:51] The thing that I would say to everyone here is that this is not an engineering problem. This is not a technical problem. This is something that everyone in this room, from PMs to designers, should be thinking about as they develop. What do you expect a user to do? What are the ways in which they are going to interact with it, and what do those outputs look like? You can really start to build up the foundational expectations of your product in this way.
[00:19:12] It provides you a good framing so that you don't have teams that just start spinning. Otherwise it's, "Oh cool, I can go and build an app that's really great. Oh, this person asked me to build a document. Oh, this person asked me to do this, or this person asked me to do that." Now you're in this mode where you're playing cat-and-mouse with all these different things, and it's really hard to frame these things into a place.
[00:19:32] But if you start with that understanding of what you're trying to accomplish, you can make sure that it comes in. You get new people asking things: great, no problem, you go and add some new evals. A new model comes out: awesome, I can go and test that really quickly. Does it perform better? Great, the answers are better. I really do think that evals are going to be one of your biggest moats when it comes to the quality of your product, and ensuring that everybody here is thinking about this and engaging very directly with what you want to see in terms of product quality.
Performance: designing for the wait
[00:20:01] Secondly, and this one is extremely painful for designers especially, is performance. We've spent so much time over the entire existence of computers making them go fast, and they go so fast today. You can go and make a search in Google across essentially the entire knowledge of human existence, and it's going to take milliseconds to come back with answers. Milliseconds. I'm sure everybody's had the experience of interacting with one of these tools and it takes seconds, 30 seconds, a minute.
[00:20:31] There are tools that take hours. You can go and do deep research, and it's like, "I got you, I'll come back tomorrow." That is an insane performance difference in the way that things interact. There's this really interesting challenge that we now face: you can't assume that it's going to be quick. You're going to have to have your users wait for things. This is just not a good experience, and it's something you need to build into the product.
[00:20:54] As you can see here, a really common way is streaming: "I'm just going to show you what I'm doing." It's still going to take a really long time. You're still not going to get what you want, but you're going to be part of the process. Maybe I can give you the first word really quickly and you'll start reading it, and I know you're probably pretty slow at reading, so it's fine. But the thing that you need to think about is how you actually build a UX that expects and enables this sort of thing. Even if they get faster, they're not faster today, and we don't know how much faster they're going to get or over what period of time.
[00:21:21] For us, a lot of the way we've tried to tackle this is, since we build applications, we take this really large prompt, which is "Hey, build an app," and break it down into different sections. "Cool, what's the schema? What are the different tables? What are the different object types?" We can start to show you building on the screen: "OK, here's your people table. Here's your campaigns table. Here's your product release table." We can slowly build that.
[00:21:44] It doesn't change the fact that it's going to take a while, but you can get to that first byte, that first response, quicker by doing that. You can build an experience where at least the person engaging with it is like, "OK, I'm starting to understand this, starting to grok what you're doing here." But you really need to design for this. You need to expect that it's going to take a while. How do you keep users engaged?
[00:22:02] Nobody wants to wait seven seconds. I remember in a prior life, we thought that we were losing 30 million dollars because one of our pages was taking seven seconds to load, and we spent a year getting that down to one second. There is a huge amount of money and success on the table when it comes to performance, and you're going to have to work through that with whatever product you build.
Cost: weigh it against value
[00:22:24] Lastly, or not lastly here, I want to talk a little about cost. LLM technology is pretty expensive. It's really not uncommon to run a chain of prompts that might even cost a dollar, or more than that, for a single user to interact with something. That may not sound like a lot, but it's a huge amount of money compared to the cost of a normal server response. You're going to get into the situation of, "OK, how do we make sure that we're able to spend this sort of money?" Maybe it's also internal tools as well.
[00:22:56] I think the thing that's pretty common as you think about tech is that as long as the value is there, you can offset that sort of cost. What you really need to think about is: what is the level of value that you can get from this cost, and does it actually outweigh the amount of money you're spending on it? I have an example that happened last week. One of my engineering directors, who doesn't really write code, is prototyping a couple of new products. We got a flag from finance or whatever it was: he'd spent something like 700 dollars on a coding agent over the last month and a half to get this prototype up and running.
[00:23:28] He's like, "Is this OK? Is this a big deal?" And I was like, "OK, well, how many engineers would that have taken?" Probably two or three engineers. That's going to cost tens of thousands of dollars, way more than that realistically, if you think about all the other stuff that comes into play with it. The idea that we're spending 200 bucks on an entire prototype over a month instead of multiple engineers is insane. That value difference is huge.
[00:23:49] Yes, it seems pretty expensive when you say, "I spent 100 or 200 bucks in two weeks." But that's really nothing compared to the value. That comes back to the point, though, that whatever you're building and whatever you're doing has to have that value. You don't want to get into these showy things that are also expensive, where you burn money and don't get anything in return.
[00:24:09] I do really think it's something to keep in mind. When you hit that point where you're like, "Hey, we need to turn this off. We need to stop people from using this, because it's costing us so much money," turn that around and figure out, "OK, cool. Where is the value in that? How do we convert that into whatever value our customers are getting?"
Enterprise trust: meet customers where they are
[00:24:27] Lastly, I'll just make a quick nod at this. Airtable is a B2B company. We build for some of the largest companies in the world, and one of the things we constantly run into is that every CEO, every leader, says, "AI is the future. If we don't change now, we're dead. We're going to change the world," whatever it is. And then you're like, "Cool, no problem. Can we turn this on for your customers?" And security, compliance, legal, whoever it is, says, "Absolutely not. There's no way that I will let you turn on AI." I'm sure this is true for most of you, even anyone here who works at small companies, but definitely...
[00:25:18] ...meet them where they need to be. Some of the things that we've seen and that have proven really important: model flexibility. Certain companies want certain things. They want it to be on a specific VPC, or on Amazon, so we can do it in Bedrock and keep it all in one place. They definitely want zero data retention. Nobody wants their data to be ingested into these tools, because that's, in most cases, their unique moat against the rest.
[00:25:41] Meet them where they are. Maybe they need to have keys that they bring to you; whatever it is, figure that out. Put good data controls in place: what can I do with this, who's allowed to use it, and where am I allowed to do this? We enable people to do it on a per-app basis, for certain workflows.
[00:25:56] Lastly, and I will just finish on this: again, have the value. It is way easier for somebody to take a risk when the value of what they get out of it is extremely high. If you go to an executive buyer or a business leader and say, "I will change the way that you run your business. I will enable you to be more successful," they're going to be way more likely to say, "I'm going to go and cut through this. I'm going to go and deal with security and all the rest of it." But you've got to have that value. Nobody's going to take a big risk if they're thinking, "Eh, I don't know that this makes a difference." Starting there will really help.
Invent the new best practices
[00:26:31] On that I'll finish, just to replay everything that I covered here. We're in a position now, and this is the fun part of it, where everything that we have known is changing. You get to invent the new best practices. Don't rely on what you currently know. Take a step back and figure out what the next one is. You may be in a position where you get to invent the new way that people interact with products five years from now. That doesn't happen very often, and it's a really awesome opportunity, instead of relying on what we think is right and what we think is safe.
[00:27:01] Play, don't design. Really, really do play. I've heard a ton of people use that word today, and I imagine a lot more people will play with these sorts of tools and systems. It will make a massive difference to the capabilities. And lastly, plan for productionization. You're going to have something that looks really incredible in a day, then it's going to be awful in a month, and then it's going to take six months to a year to be really great. That's the way these things go. Really plan for that, and make sure that you're able to understand and make that difference clear. With that, thank you. I appreciate you all listening, and I hope this has been helpful.
Q&A
[00:27:42] Host: Jimmy, that was fantastic, thank you. I have a couple of questions for you, as you might expect. On playing as you learn, as far as non-deterministic models are concerned, you covered a good amount of ground in your presentation, so I think we're going to have a lot of follow-ups. Just to jump off here, I think there's a very poignant question around how you know where to start, or what problem you're going to solve for, when you're at the step of getting it to work.
[00:28:21] Jimmy: I think part of that is, if you have a good strategy, if you understand what your company is trying to accomplish and what you're trying to do, start there. Because it's not just, "Can I play with a tool and figure out anything that's possible?" It's, "What is unique to us?" You have a persona that you're building for. You have a problem that you're trying to solve. Whatever it is, even at a higher level, engage in that space.
[00:28:43] Don't go off the beaten track to say, "Oh, I know that an LLM can do this, so I'm just going to build that into our product." Really start from the point of view of what is unique about your product and what is unique about your customers, and start there. What data do you have? What sorts of workflows do you start to build in? Play in that space, and don't come to it with that preconception. That's the area in which you want to separate "Well, I know I can do X and Y" from "What else can we do? What else is potentially possible?" It's playing in that space where I think you're going to get some interesting insights and then be able to go down that path instead of a path you might have gone down otherwise.
[00:29:17] Host: Love that. Just curious, from a concrete perspective, can you cite an example of an evaluation?
[00:29:27] Jimmy: Yes, there's a million. Generally it's a little micro, macro. The way that I would reason about it is that you can essentially say, if it's just a prompt engine, "Somebody's going to type in this message. What is the output of that?" For us, for example, we have a lot of evals, I think thousands of them at this point; I don't remember the last number. We have automations. Somebody might say, "Hey, I want you to email someone every week." Somebody else might say, "Hey, on a Friday I want you to go and send a message to these people."
[00:30:02] You start to build up the different ways in which people might interact and say similar types of things to get to a similar end. How do you build in the evaluations so that when somebody asks you to do something in that way, you can funnel it that way? It's different from a traditional unit test, where if somebody puts in the number 10, I know the answer is going to be 100. The 10 might change, but the 100 is going to be there.
[00:30:26] This way, because a lot of it is natural language, you're going to see a lot of variation in the way that people speak and the way that people ask for similar things. How do you categorize that and evaluate them to the same sort of output? That's the way that you'll expect it to come, and it varies depending on the type of product you're building. But normally it's essentially: what's the type of way you would ask an LLM, and what is the output you would expect based on that type of ask?
[00:30:48] Host: Sure. Now we're back to non-deterministic.
[00:30:50] Jimmy: Right, exactly.
[00:30:53] Host: One question that's come up a couple of times: we've got a lot of horsepower, a lot of compute behind these AI capabilities. There's the understanding that they're really powerful and can do a lot of things, but how do you address the organic nature of people leveraging those capabilities?
[00:31:15] Jimmy: I have two answers to this. I strongly believe, and I think it will be true, that a year or two from now people will understand how to engage with this technology, so the teaching aspect will become easier. They'll just engage with it. The difference now, though, is that you want to be the product that gets people to that point. You don't want to be the one who waits, because if you wait, you become irrelevant, and somebody else has come in and done that.
[00:31:40] I do think it's true that there's this idea that people organically start to engage with this. Really, that organic means that somebody else has figured it out before you. So you want to be ahead of that, or at least at the forefront of it, really engaging to try to figure it out for yourself and your teams, and not wait a couple of years to say, "Oh yeah, cool, we can do that as well."
[00:31:58] Host: Right. Given your purview, it's all things product: product development, product management, engineering, design. I'm curious, because we have a lot of designers here in the audience: are there certain tools or products that you would recommend people use to get started on that journey of harnessing all things AI?
[00:32:20] Jimmy: Totally, and I think there are two versions of this. One, there are the tools that are unique to your craft. There are a ton of things you can use to do your design work better, and I think that's a given; definitely engage with those. But the thing that I would push more folks towards is to take a step back from that and engage with the technology itself. Even if you just go and throw something into any of these tools (go into Claude, go into GPT, go into an API, whatever it is) and engage directly with a prompt, that's where I think a lot of design should really be focused.
[00:32:53] You don't want to be in a position where you're just being informed by other people: "Hey, this is what it can do." You want to go and figure that out for yourself, because without that you're going to really narrow the potential space. Go and get as close to it as possible. I've said it a couple of times today: prompts are not that scary. They're pretty boring. They're not very interesting. They're annoying to write, but they're not inherently difficult if you start to engage with them. If you are not writing prompts or engaging directly with the LLMs, that is exactly where I would push people to go today.
[00:33:21] Host: All right. One final question, which is associated with, I'll just say it, all things vibe coding. How would you recommend UX folks or designers lean in on that and learn just enough to be effective in that capacity?
[00:33:38] Jimmy: I am sure that every designer here has made a prototype at some point. You know that the higher the fidelity is, the harder it is, because what essentially happens is someone says, "Oh, awesome, you're done. We're good." As far as vibe coding goes, I would take it the same way: you should use it as a way to explore. I think there's a lot of belief right now of, "Oh cool, I'm going to get to that production website, production product, production whatever." You're not going to.
[00:34:06] That may change years from now; I'm sure the technology will get better. But today you're going to make something pretty bad, and the minute you give it to real users, or the minute you have 100,000 users, it's over. The way I would interact with vibe coding is exactly that: how do I get ideas out really quickly? What is a really cool thing that I can try here? Then put it to the side, start again, and do it in the more legitimate way.
[00:34:30] That isn't to say that you shouldn't use LLMs for writing code. But there's a big difference between vibe coding and actually using the technology to build sustainable, good architecture that will actually scale with whatever problems you have as a business.
[00:34:45] Host: Awesome. Jimmy, thank you so much for your presentation today. We covered a ton of really great ground. Thank you.
