Building AI for Democratized Research

31 Mar16:00 – 16:30 UTCStage: Main StageTalk

Checking session availability…

Hang tight while we load the latest updates.

Democratizing research sounds great in theory, but in practice it often means inconsistent methods, overconfident teams, and findings you can't trust. The fix isn't to lock research back down. It's to build smarter guardrails.
This talk walks through how to use AI to create tools that empower product teams to own their research while keeping the quality bar high. You'll learn what guardrails actually matter, how to drive adoption when teams think they don't need help, and how to go from idea to working MVP fast.
Whether you're a researcher trying to scale your impact or a team lead looking to build more autonomously, you'll leave with a practical blueprint you can start applying immediately.

Building AI for Democratized Research

Mike Oren at UXDX Community: Guardrails, Governance, and Trust: Getting Teams to Act on Research. Video: https://youtu.be/HmXxnNhXVfY

Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.

The tension in democratized research

[00:00:08] Mike: Welcome to the UXDX community. First of all, I want to start with a number, before I put it on the screen. Think about the biggest challenge you have all had with research at your company. Based on UXDX's survey, about 66% of people attending events want to learn to improve their processes. Often this includes recruiting users and customers, which this talk isn't directly discussing, but it can help with that as well. Working for a B2C [?] CRM company, being able to contact your customers isn't one of our biggest challenges. Today is about how you can use AI to help you build better research processes and improve democratization.

[00:00:52] There's a tension that every team with a research function runs into, and it usually gets worse as a company grows. On one side, you want everyone talking to users. Democratized research is genuinely good. PMs interviewing customers, engineering watching usability sessions, designers running concept tests: all of this makes your product better. On the other hand, not everyone knows how to synthesize well or ask unbiased questions. Since this is the step where experience, methodology and judgment matter most, it's also the step that's hardest to review. If a PM reads two transcripts and tells the exec team, "Users hate this feature," no one's going to fact-check that. The story is already in the room.

[00:01:41] So what do you do? You can't say only researchers are allowed to use AI on research data. That ship has sailed. What you can do is build the methodology into the tool itself.

A guardrail is a workflow engine disguised as a chatbot

[00:01:55] Before I get into the mechanics, I want to offer a reframe that I think is useful. When most people hear "AI guardrails," they think restrictions. Don't do this. Don't say that. A guardrail sounds like a fence. But the things I'm going to show you don't feel like fences to the people using them. They feel like expert colleagues. They ask clarifying questions before they proceed. They give you structured output without you having to remember the template. They route you to the right tool when you've wandered out of scope. A guardrail isn't a restriction. It's a workflow engine disguised as a chatbot. That's the mental model I want you to hold for the rest of this talk.

The four pillars

[00:02:40] Every guardrail system I've built, and the teams I've worked with have built, rests on four pillars. Let me walk through each one quickly, and then we'll see them in action.

[00:02:52] One, gatekeeper. The gatekeeper rule stops the bot from proceeding until it knows enough to do the job well. The most common version is a participant count check. How many participants are represented in these transcripts? The bot won't synthesize until it gets a number. It can also require a methodology, interview versus survey versus observation, because those require different analytical approaches. This sounds simple. It is, but it changes behavior. When someone has to type "I only have two transcripts," they suddenly know they're about to get a preliminary read, not a final analysis. The act of answering the question creates the caveat.

[00:03:34] Rule two, no fabrication. Every quote in a synthesis output must trace back to a named participant. The prompt forbids the bot from constructing representative quotes, combining paraphrases or filling gaps with plausible language. If the data doesn't support a direct attribution, the bot has to say so, not invent one. This matters because hallucinations in research synthesis are catastrophic in a specific way. AIs are very confident. An AI-fabricated quote doesn't look uncertain. It looks like evidence. By the time someone realizes it wasn't in any transcript, the decision may have already been made.

[00:04:20] Pillar three, output standards. The bot always produces the same report structure: executive summary first, then themes with supporting quotes, then a methodology note, then any limitations. Always in that order, always with the participant count in the header. This has two effects. Readers learn what to expect, which makes them better at spotting gaps, and it makes reports reviewable. A researcher can scan a synthesis output and immediately see if the structure was followed.

[00:04:52] Pillar four is ecosystem routing, or standards. No single bot tries to do everything. When a user wanders out of scope, say asking the synthesis assistant to help design a screener, the bot declines and routes them explicitly to the right tool: "I can't help with this study design. For that, try the research method assistant," if you're asking the synthesis tool. This keeps each tool focused, which keeps each tool trustworthy. A bot that does everything well is harder to build and harder to trust.

[00:05:24] To expand upon that just a little bit here: one thing that will happen if you try to create a do-it-all skill is that what you end up with is a very large context window. The larger the context you're feeding into any particular custom Gem, custom GPT or Claude skill, the more likely you'll start getting hallucinations from it. So by chunking things out, you're more likely to get the right answer from that particular custom skill that you've created. If you've got Claude Cowork, then you can bundle skills together into a plugin, and they can then more easily reference one another. It looks like Gemini and OpenAI will be doing similar things in the future, but right now Claude is the only tool that allows you to bundle those skills together while keeping them separate.

Demo: the survey trap

[00:06:19] So here's the scenario. A PM wants to move fast, which is completely reasonable. They've got a new dashboard concept they're excited about, and they want to know if users want it. They type, "Generate a quick survey to see if users want our new dashboard." In most tools, like a generic ChatGPT, that generates a survey. Five questions, rating scales, submit button, done. Watch what this one does. And live demos are always fun. Sorry, I should have copied that question over. And here we go. It's running on fast, so it shouldn't take too long. And there we are.

[00:07:12] So it's throwing that complexity warning, just like we're seeing here, but because it is an LLM, it's not exactly the same. It is a little bit different. The tool doesn't write the survey. It throws that complexity warning. It explains in plain language that surveys asking users to predict their future behavior have a known reliability problem. People say they want things they won't use. It's not a user research problem. It's a cognitive bias problem.

[00:07:43] Then the tool coaches. Five user interviews would give you a better signal. Here's why, and here's how to structure them. The PM doesn't need to know that predictive surveys produce false positives. The tool knows it for them. That's triage.

Demo: coaching on the interview guide

[00:08:02] Okay. The PM takes the coaching. They agree to user interviews. Good. But now they're writing the guide, and they tell the tool, "I need to know if they will pay $10 for this feature." That sounds like a reasonable thing to want to know. It's not a reasonable question to ask in a research interview, though. Why? Because the answer is going to be a way to get it for free, or yes motivated by social desirability, or a yes couched as a no, or really a no couched as a yes, where they might say "my friend or my brother would be interested in this," which is their way of saying no without telling you directly. Generally people are polite. They want to be helpful. They don't want to tell you they wouldn't pay for something you're clearly excited about. And even when they're being honest, people are notoriously bad at predicting their own purchasing behavior.

[00:08:59] Let's see this one in action now. Okay. So the guardrail doesn't just flag it. It asks the PM a clarifying question. What's the underlying assumption you're trying to test? What would you do differently if the answer was no? Those two questions almost always reveal a more interesting research question. And the tool uses the answer to redirect the interview guide toward how users currently solve the problem, which tells you far more about willingness to pay than asking an interviewee if they'll pay. If they're paying a competitor already, they're likely willing to pay you, as long as you can make it as good as or better than that competitor.

[00:09:55] Although there are survey techniques that can be used with enough reliability, specifically monadic testing. But this is what coaching looks like at the prompt level. The methodology isn't in a wiki that nobody reads. It's not a training session that people may go to but then multitask. It's in the interface they're already using.

Demo: the synthesis assistant

[00:10:21] So now we're in synthesis. The interviews are done. Someone is using the synthesis assistant to pull findings from five transcripts. Three things happen that wouldn't happen with a standard AI. First, the participant checklist. The tool was told to expect five participants. It's only seeing four transcripts. So it stops. It says, "You said you had five participants, but I only count four. Waiting on P05." It doesn't proceed. It doesn't synthesize a partial data set and present it as complete. It waits. It also protects participant anonymity.

[00:10:56] Second, quote attribution. Every theme in the output has a participant ID next to it. P02 said this. P3 and P4 both mentioned this. You can trace every claim back to a human being in a specific session. Third, flagging inferences. If a theme is directional, based on tone or implication rather than a direct statement, the tool says so explicitly. It doesn't present an inference as a finding. It labels it.

[00:11:28] So I'll give you a quick demo here. I created some synthetic data here, which, by the way, synthetic data I don't generally trust, but for the purposes of a demo we will give it these four transcripts of interviewing someone about buying pants. AI hallucination in research synthesis is dangerous in a very specific way. Again, they're very confident. A fabricated quote doesn't look uncertain. It looks like evidence. The tool's job is to do as much as it can to reduce the chances of anyone putting a made-up quote into a readout. So again, you can see that it is flagging it as expected. There are some other things that this one does, but that's just the easiest one to demo live.

The instruction is just English

[00:12:18] Now I want to show you what the actual instruction that produces the survey-trap behavior you just saw looks like. In this case: "If a user asks to generate a survey to validate a new concept, do not write the survey. Instead, throw a complexity warning, explain that surveys asking users to predict future behavior often result in false positives, and suggest a qualitative concept test with five users instead." That's it. That's the whole rule. You don't need to train the models, although you should test them. The code is just English.

[00:12:52] What this means for you: the bottleneck to building these guardrails is not technical skill. It's clarity on your own methodology. You need to know what the right move is before you can encode it. But if you know what good looks like, you can build this, multiplying your ability to elevate your organization and your work, or, in management lingo, increasing your leverage, so you can focus your energies on even higher levels of impact.

[00:13:25] One quick note before the framework. None of what you've just seen is platform specific. The instruction set is the product, not the platform. Those sentences I just showed you live in a system prompt. You can deploy them as a Claude skill, a Gemini Gem or a custom GPT, and I've done all three at Klaviyo. Whatever your organization already has, whatever it has already approved, the methodology travels with the text. It doesn't live in any particular tool. You don't need to ask anyone's permission to start. You just need a platform your team already uses and about an hour of clear thinking.

Three questions to build your first guardrail

[00:14:06] So how do you actually build one of these? I've distilled it to three questions. Answer all three honestly and you have everything you need to write your first guardrail.

[00:14:17] Question one: worst mistake, or the triage rule. What is the single worst thing that happens when someone does this task without guidance? Be specific. Not "bad research"; that's too broad. Something like "a PM feeds two transcripts to an AI and presents themes to the exec team as the voice of the customer." That's specific enough that you'd recognize it if you saw it happening. The mistake you name becomes your triage rule. The guardrail's first job is to stop that specific thing from happening.

[00:14:52] Question two: done correctly, or your output standards. Write one sentence describing the output when this task is done well. One sentence. If you can't do it in one sentence, you might need to break it into smaller steps. AI can help you do that. Or you may need to clarify your thinking. Have AI ask you questions to help you hone and clarify your thinking a little bit more. When you do have that sentence, it becomes your output standardization rule verbatim. Drop it into the prompt, and the tool will produce that output every time, or close enough to it.

[00:15:30] Question three: workflow position, or routing logic. Where does this task sit relative to everything else? What happens before it? What happens after it? What other tools exist in the ecosystem? This tells you what context the guardrail needs to ask for before it starts, and who it should hand off to when it's done. That's your routing logic. Three questions. That's the entire framework. Everything else is testing and iterating, something researchers should know how to do well.

From zombie insights to an ecosystem

[00:16:03] One thing to consider for your first build: something to help direct your organization to what you already know. Almost every research team has a lot of zombie insights. You can turn those into evergreen wisdom. This matters for a reason beyond efficiency. The more your team builds on shared foundations, the more consistent your research methodology becomes across every project, every team, every region. Individual tools then become an ecosystem, and an ecosystem has institutional memory in a way that a pile of one-off studies never will.

[00:16:39] This also applies to all the AI tools we give them. If you're not building these tools, someone else might, and then you've lost control of that narrative, lost the ability to guardrail that process. That's why I generally recommend getting on top of building these tools ahead of time. Research generally needs to be an early adopter of the tools in order to keep everything aligned.

[00:17:06] So what's the one thing in your research workflow that would embarrass you if a stakeholder saw how it was actually being done? That's not a rhetorical question. I want you to actually answer it in your head right now. The answer to that is really your first guardrail, the first thing that you could potentially build as a custom GPT or custom Gem.

[00:17:25] I'll also add that for many of the tools that I built internally for research democratization, I also built advanced versions for the research team itself, in order to, again, increase my own leverage and increase our consistency as a team. If you go back to Gemini, you can see that I do have an advanced version of the research method assistant. Your product manager or designer doesn't need to know about things like NASA TLX for stress studies, if you're testing something where time on task and stress is important. But your researcher may, if the team as a whole is building certain solutions where that could be relevant.

[00:18:15] I also have a research survey assistant because Klaviyo is a B2B company. We're not democratizing surveys, because then we would inundate our customers very quickly. If every product manager were to send a survey to 200 people even once a month, we'd very quickly start seeing our customers receive five surveys every three months, and that's just too many. So that's where the researchers own those surveys, and having a survey tool also helps make sure that all the researchers are not just making up questions every time, but first relying on established survey measures, which are all in that research survey assistant.

[00:19:07] That now gets to the end of my talk. I'm happy to take any questions, and if there's something I didn't get to, feel free to reach out and ask one-on-one. I do have the full instruction sets available for the generic ones, the synthesis assistant and the research method assistant, where we don't have anything that's company proprietary in those. I can share those if you reach out as well. Okay, any questions?

Q&A

[00:19:39] Host: Excellent. Thank you very much, Mike. That was a really nice coverage of skills, because I've heard a lot about skills, but it's one of those things where so many things are happening. It's great to see it put into practice, so thank you for sharing that. And just a reminder: whatever platform you're on and watching this, please write in your questions and I'll put those questions through to Mike.

[00:20:04] I'm going to jump in with one, because as I was watching that, you were giving a lot of advice on how to get started. I'm looking at the opposite end. I'm thinking, God, there's 20-plus things just on research I could write, and then on design, and then on engineering there's probably another 30 or 40, and I'm just thinking about the explosion of all of these different guardrails. I think it's incredibly powerful. But is that a problem you've seen, that now you need to start thinking about how you're going to manage the sheer volume of these?

[00:20:34] Mike: Yeah, that's one of the reasons why it's important for honestly every discipline to get ahead of it and build those skills and rules themselves. Because if you don't, someone else who doesn't have that expertise likely will, because it does save a lot of time. Or they'll just go through the generic ChatGPT or generic Gemini, and those often give you generic answers and aren't really going to be good enough for what you're trying to do as an organization.

[00:21:13] In terms of managing them, we haven't had a problem, partially because every time we've gotten access to a new tool, I've been able to release our version of those skills basically within a week. So at least on the research side we've been able to keep it constrained. We were seeing on the product management side that a lot of product managers were building their own sets of PM skills for Claude, until one of the PMs started sharing in public Slack channels, "Here's my repository," and then we were able to get that reconsolidated.

[00:21:56] Host: Awesome. That's great. I guess it's that distributed nature, so you just need to manage your own, and it doesn't then become this much bigger problem. One thing that I thought of as you were doing your demos was: could this become a teaching tool as well as just a guardrail? Because yes, it's catching people at the point of them making those mistakes, but is there also an opportunity to say, you know what, and maybe even ask the AIs, "Here are all of my big problems, can you create a little course for people to avoid them in the first place?"

[00:22:34] Mike: Yeah. I didn't mention that here, but, not quite a course, but I do a couple of different things. Someone who's already conducted an interview can upload their transcript to an interview coaching assistant, and it'll basically score them on how well they're doing in terms of asking unbiased questions, doing more listening than leading, all of that type of thing. So that's one tool that we have at the company.

[00:23:04] I also teach on the side, and so I have been having students create their own custom Gems and GPTs or Claude skills in order to help them with certain advanced research methodologies. Anybody here who is familiar with deep listening or mental models from Indi Young: it's a different type of interviewing style that I've found a lot of students struggle with, because it's not structured. It's a very unstructured interview. So you give a custom Gem very simple instructions, really just about a paragraph, and that custom Gem will give them coaching feedback if they start asking directed questions instead of asking questions based off the story that the individual gives them. And you can then use Gemini or Claude as a practice interviewee. Just like people now use these tools to practice for job interviews, you can also use it to practice your interviewing techniques.

[00:24:11] Host: That's a great idea. I hadn't thought of that one before. And we had a question come in from K the FEM [?]: "Thank you, Mike. I'm curious about how this would look in voluntary or on-the-ground research scenarios." I'm not 100% sure of the context. Do you understand "voluntary or on-the-ground research scenarios"?

[00:24:34] Mike: I am not completely sure what's meant by that.

[00:24:38] Host: So, maybe add an update to your question, or take a shot if you understand it.

[00:24:47] Mike: No. I don't know if it means contextual interviews. But yeah, it's unclear what's meant by that.

[00:24:59] Host: That's fine. One of the things that I thought of as well: you mentioned a lot of hallucinations and those problems that are well documented in LLMs. So even with your prompts, I can see there's still that risk of things going wrong. Have you started investigating, because I know there's a big trend now in getting LLMs to check the LLM output. By having two instances, it reduces a lot of the errors, because they catch each other's mistakes. Is that something that you're investigating as well?

[00:25:38] Mike: Yeah. Like I mentioned, I do have multiple versions of my tools, one for the research team and one for everyone else. I don't expect everyone else to necessarily worry about inter-rater reliability and things of that nature. For the research team, where they might be using a grounded theory framework for synthesis, that research synthesis assistant also gives them tips for how they can use multiple AIs to do inter-rater reliability. That way you have more confidence that those things actually exist. In addition to asking them to check themselves and calling it a question, you still will always need some human review. Never fully trust an AI, but it does help reduce the amount of time you might otherwise have to spend on it.

[00:26:34] Host: And I guess we just have time for one last short one. You mentioned reaching out to you, and people can rewind the video and see your contact details again, but is there anywhere else that you recommend people go to try to find starter templates or anything like that?

[00:26:52] Mike: Honestly, I've found the most useful thing is sometimes just to tell AI what you're trying to do and ask it what some pieces of advice are. I think a lot of people think that when you work with AI, it's all about telling it what to do. But the more you ask it questions, as opposed to just telling it all the time, the better output you'll typically get out of it.

Speaker

Mike Oren

Mike Oren

Head of Product Research

Klaviyo