AI Without Access: Shipping Intelligence When You Can’t See User Data
Checking session availability…
Hang tight while we load the latest updates.
Everyone is racing to plug AI into their workflows. But what happens when your product is built on a promise to never see user data in the first place? As CTO of a zero knowledge product, Frederic Rivain has to reconcile two strong forces: teams who want to use AI everywhere, and an architecture that is designed to reveal nothing.
In this talk, Frederic unpacks how Dashlane uses AI internally in a secure and responsible way, and how they choose and integrate AI into the product while keeping their zero knowledge model intact. He will share the trade offs, the architectural patterns, and the guardrails that let teams experiment with AI without weakening the trust the product is built on.
Challenges:
- Leveraging AI internally when sensitive customer and company data must remain encrypted or out of scope
- Choosing AI providers and architectures that align with strict privacy, security and compliance requirements, not just convenience
- Embedding AI into critical user journeys without breaking the zero knowledge promise or eroding customer trust
Key Learnings: - How to design AI use cases around a zero knowledge mindset, including what to avoid, what to transform, and where AI genuinely adds value
- Practical patterns for using AI internally while protecting data, from governance and access controls to what you never put in a prompt
- A decision framework for selecting and integrating AI partners in high trust products so that privacy and security stay first class, not an afterthought
AI Without Access: Shipping Intelligence When You Can’t See User Data
Frédéric Rivain at UXDX USA. Video: https://youtu.be/nwYwalsKNVQ
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
The tension between AI and sensitive data
[00:00:08] Frédéric: I'm Frédéric Rivain, the CTO at Dashlane. Show of hands, who knows about Dashlane? I can see you. Yeah, not that many people. We are a credential manager, so we help you store your digital life, your secrets, your passwords, your IDs. As such, we are really processing some of your most sensitive data.
[00:00:28] And so, of course, we're going to talk about AI. Today, when we talk about a product and AI, everybody has the same mandate. You're asked to add AI capabilities into your product as fast as possible. Try to have it everywhere. It's really important. But on the other side, if you think about sensitive data, whether it's health data, financial data, identity, you're also asked not to touch the data. It needs to be locked out. You shouldn't expose it. You shouldn't try to touch it at all if you can. So that creates a tension between the two things.
[00:00:59] In our case, at Dashlane, because we are a credential manager and, like I said, manage some of your very sensitive data, we are at the extreme end of that tension, where we never want anybody to be able to access that data except you. That means that when you think about adding AI in such a product, we have to rethink everything. We have to think about how we train the models, how we collect data to be able to train the models, what model we can choose, how we debug, how quickly we can deploy, run inference, and so on. So in this talk, that's what I want to share with you: some of the practices we've used in our product, so that they can be used as inspiration in your own product if you want to add AI and be careful about data.
[00:01:41] But let me get back to, if the clicker works, yeah, let me get back to last summer. There was a story last summer that security researchers found hundreds of thousands of conversations that had leaked from ChatGPT on Google. At that point, you could read about mental health issues, about relationship concerns; you could find business contracts, all of that out in the open. And of course the users themselves had assumed that those conversations with the model were private, but it had leaked, it was too late, it's now archived in Google, and it's not great.
[00:02:11] And it's not an isolated incident. There was GitHub Copilot, which leaked data from private repositories. At some point, Samsung decided to ban ChatGPT because it was leaking employee data. It's always the same pattern. We try to add AI features into the product, we don't really pay attention, and at some point we mess up and we leak customer data. And it does matter, because at the end of the day, you're liable as an organization for what you're doing with your customer data.
What breaks when you can't see the data
[00:02:38] So let's get a bit concrete. When we actually try to build a privacy-preserving AI product, let's see what breaks. AI is a continuous cycle, if you think about it. You collect data, you collect signals, then you train your model, then you deploy it and run the model for your customers, you monitor, and then you use the monitoring and the information to be able to retrain the model, and so on. That virtuous loop is how you get those AI models to become better and better.
[00:03:07] The challenge is that when you don't see the data, that loop doesn't exist. You can't really use the same process. And there are two parts that are going to break first and foremost. The first one is the training, because you don't have the data from your customers to be able to train your model, so you need to solve for this. But the second one is also deployment, because at that point you still need to pay attention to the data that your customers, the prompts, are sending to you to run the model. So those two aspects of privacy-preserving AI are the ones we're going to focus on.
[00:03:40] In addition to that, there's another problem. Once you have a model, you still have to run it on user data, like I said. The default pattern here is usually you put your model in a cloud endpoint, you send the prompt in there, it's in plain text, and then you get the prediction back. But that doesn't work when you don't want to see the customer data. Of course, you can still do things like encrypt the data in transit and encrypt the data at rest. But with AI models, the problem is that when the inference runs, you have to decrypt the data. You have to load it in clear in memory so that the model can run against it. So that doesn't work in our context.
[00:04:15] And you could ask me, "Okay, why did you put a toilet on that slide?" Actually, this is a recent and real example of what is supposed to be a smart toilet camera, an end-to-end encrypted smart toilet camera. It essentially attaches to your bowl, it takes pictures of what's in there, and it will use AI to analyze your gut health and provide you with some advice about what you should be doing in terms of food and so on. The reality was that they were encrypting data in transit up to them, and at rest on the servers, but definitely not in use, and it started leaking important customer data, as you can imagine. So really, that's another important challenge when you think about privacy-preserving AI.
Use case one: anti-phishing
[00:04:57] All right. Now that we've highlighted those key challenges, training, deployment, how do we encrypt data, let's try and look at the different use cases. I'm going to share, of course, some use cases from Dashlane. There's no silver bullet, obviously. There's no magic solution here. But still, I hope that some of our approaches can be an inspiration to you in how you can also bring AI in a private way to your own product.
[00:05:19] Our first use case is anti-phishing. The goal here was pretty simple for us: how can we protect our users and make sure that they avoid putting credentials on fake login pages, in real time, obviously, dynamically, and reliably. And in our case, we didn't want to be able to see customer data, whether it's their browsing activity or even their credentials. And that's an important topic these days, because with AI, the attackers have very much accelerated the velocity at which they can build fake phishing websites. So you need something that's really real time. You can't rely anymore on basic email protection beforehand.
[00:05:57] We had three main challenges to solve here. The first one is, okay, how do we train our model, and how do we collect data to train our model? Which model is the right one for that use case? And then, where do we run the model, obviously? And you could argue, okay, that's an easy issue. You can take an LLM, you can get the LLM to run on the webpage and try to predict if the webpage is real or fake. But in our case, because privacy is the constraint, and we don't want to be able to see the browsing data from the customers, we couldn't do that. So all our design is flowing from that privacy constraint.
Building a training data set without customer data
[00:06:33] First, when we wanted to train the model with data, we had to build our own data set internally. For that, we used different data sources. The first one is internal crowdsourcing. We actually used our employees to do that. We built our own data collection app internally, and that allowed us to get a basic set of data to be able to train our model. That gives us a safe and controlled environment to do that.
[00:06:56] But the challenge when you do this internally is that your employees are doing this in a too-clean way. When you get to production, the problem is that your users are going to be unpredictable. They're going to be messy. They're going to be doing things you didn't expect. The problem with models that are trained on internal data is that, like I said, they're too clean, and they will break when they go to production. That's usually what we call production drift.
[00:07:19] So it's not enough. You need to complement your internal data sources with other sources. One other source we used is synthetic data. There's a lot of talk recently about synthetic data and using LLMs to generate fake data. That works. It's fast. It's cheap. But there's also a caveat to that, and it's that the synthetic data is not going to be very diverse either. A way for you to understand this is if you, for instance, ask ChatGPT or Claude to generate jokes, say 10 jokes, you can see that the first jokes are going to be pretty creative, and then progressively you will see patterns that reappear. You have the same problem when you generate synthetic data. You want to avoid having the same structure and patterns reappear. And it's not that the generative models are bad, but as you know, by construction they're meant to focus on high-probability patterns, so they will always go back to the same patterns.
[00:08:09] So you need to complement your synthetic data with other sources as well. We also used a public data set, which is very useful. In our case, we used the public list of phishing URLs and phishing websites, which allowed us to evaluate the model and test against known patterns. You just need to be careful, if you do that, that you have the copyright and the license rights to use that public data.
[00:08:31] And finally, it's about training and retraining. You also need to be able to get more insights as you run your model. Here you can use the telemetry, your logs, but they need to be aggregated and anonymized so that they don't leak customer data as well.
[00:08:47] You could tell me, okay, something that seems missing here is: why are you not just anonymizing the data? Well, actually we are, and we even anonymize the data from our employees. We strip out names and emails and so on. But the problem is that that's not enough, because even if you anonymize, it's never perfect. You always have the risk of being able to reconstruct the context just from the data itself. And also PII, personally identifiable information, is not always easy to identify. It's not as obvious as just an email and a name. You can have a more, let's say, sneaky type of PII.
[00:09:20] So when you build your training data set, like I mentioned, it's really important to think about the coverage. Think about diversity. The goal is not to collect the maximum volume of data. Scale doesn't matter here. What matters is to cover all the different use cases that you will face in production: the different languages, the edge cases, the context of your own users. So think about this problem when you build your training data set.
Start small: frugal modeling
[00:09:44] Now that we have a data set, we need to pick the model. And here you could argue, "Okay, instinctively, if I don't have much data, maybe I'm going to use a very large model, and the very large model will figure it out and find the right patterns." But actually, the reverse is better. You should start small. You should start with the smallest model possible, and it's more efficient to use a model that is targeted at your own use case, because it has a lot of benefits. That's what we call frugal modeling.
[00:10:10] Benefit number one: because it's small, you need less data to train it, and you need less data to achieve the same level of performance. Benefit number two: the iteration can be faster. It's cheaper to retrain, it's cheaper to train, and that matters when you want to have a fast iteration loop. Also, when we have issues, it's easier to pinpoint where the issues are coming from. It's not a big black box where you don't know what's happening. If you try to debug a large language model, it's very hard, because the model does what it wants, and you don't really know how it's been built, and it's hard to identify where the issues are coming from.
[00:10:44] Of course, a smaller model is also cheaper and consumes less, which is always a good side benefit when we think about the cost of AI, whether it's the financial cost or the green cost. Governance becomes easier. If your customer or reviewers come and say, "Oh, I don't understand why I got that prediction back, or why the model reacted that way," it's easier for you to explain how the model behaved. And finally, and we'll talk about this again, with a smaller model, being able to deploy it on device becomes possible, because it's compact enough and small enough. And that's the maximum privacy outcome you can get here.
[00:11:17] So really, for me, for a lot of real use cases and problems, feature engineering plus a small model is probably better than a large language model, which is going to be very hard to deploy safely and privately. For me, it's not a compromise. It's really finding the right solution to your own challenge.
[00:11:36] In the case of the phishing use case I mentioned at Dashlane, we decided to build our own model, like I mentioned. We collected our own data. And actually, the model relies on about 80 indicators that come from the structure of web pages and the semantics of web pages, to be able to have a very compact, very cheap model that can run on device, in the browser extensions of Dashlane. That makes it very good in our use case, because it never takes data outside of the extension and the browser of the end user. So we don't have a cloud surface that can be attacked. We don't see the browsing activity of the user. And it gives us maximum privacy by construction. So, like I said, frugal for me is not really a compromise. It's really: under a privacy constraint, let's find a solution that gives you maximum privacy and the smarter architecture choice.
[00:12:24] So, three lessons from that first use case. When you train, make sure to think about coverage in your data set. Coverage is critical so that you can get maximum performance out of your model. Frugal modeling may be enough. Don't think, okay, I'm going to use the best large language model out there, because it may be too complicated to get it safely into production while maintaining privacy. And finally, if you can, if your model is small enough, try to run it on device for maximum privacy.
Use case two: insights from audit logs
[00:12:50] All right, use case number two. This one's a bit more complicated. As I said, Dashlane is a zero-knowledge product, so we never see the data of our customers. In the case of our B2B customers, we do generate a lot of audit logs, activity logs, for our customers. We tell them where people have used credentials, how, when, whether they are compromised credentials, and so on. Those are a lot of very rich signals we provide to the admins of our B2B organizations. At the same time, it's a lot of data, so we want to help them with AI to process that data. And AI copilot types of features are very applicable in that case: you process a lot of data, and you give insights out of that data.
[00:13:28] The problem is, yet again, we cannot see the logs, and we don't want to as a company. So one, it's really hard for us to train and deploy a model. And two, it's a huge amount of data. We generate a lot of logs, too much to be processed on device, where you have limited performance and limited context.
[00:13:44] So we didn't really have a solution at first. When we started our journey a few months back, saying, "Okay, what do we do to be able to provide more insights for our admins?", we decided to take an incremental approach and not wait for the magic solution or the perfect technology out there. The easy one is, okay, let's build dashboards on top of those logs. That gives you at least a level one of insights. It's not great, because you can't deep dive into the data and so on, but it gives you the bare minimum. We included an integration with SIEM systems. SIEM systems, if you're not familiar with them, are like a security-specific data warehouse. So then our customers can flow their logs into their own systems.
[00:14:21] From there, we started adding AI capabilities. We started with an MCP, so that our customers can connect their AI agents directly to their logs, but we're still within the boundaries of the organization, so the data is safe and we still don't see it. And finally, which is more interesting, we're now starting to introduce actual AI inside our product to be able to process that. But because the data is too large, there's too much context, like I mentioned, we have to do it in the cloud. We have to do it on the server side. And here, and I'm going to dig deeper into that use case, we're going to use confidential computing and what are called cloud secure enclaves, to be able to process that volume of data in a secure and private way.
Confidential computing and secure enclaves
[00:14:59] So let's talk about confidential computing. As a reminder, as I mentioned, encryption is about three things. You need to encrypt the data in transit, as it flows to you. You need to encrypt the data at rest, when it's on your servers. And in the case of AI, you also need to encrypt the data in use, because at the moment inference runs, you will have to have the data in clear, loaded in memory, so that you can do your magic.
[00:15:21] That's where confidential computing comes into play. Cloud secure enclaves are basically small, hardware-isolated environments that are provided by your cloud provider. They are sealed, and they give you a way to process data in a sandbox, in a black box. The cloud provider can't see what's happening in the black box. You can't see what happens in the black box. But you can still run data in the black box.
[00:15:47] I don't want to get too technical, but let me just walk you through the data flows on the diagram. On the left, you have the encrypted data that's coming in. Like I said, it's encrypted at rest. And you have this cloud secure enclave, what we call the trusted execution environment, a TEE. It's a small virtual machine, hardware, very limited. Usually you have limited CPU, you have no memory, you have no GPU, which makes things complicated, but that's also how you make that model defensible from a security standpoint.
[00:16:15] And of course, the enclave cannot just pull a key to be able to decrypt the data randomly. It needs to do what we call attestation. It needs to prove that it is doing what it should be doing. It needs to prove that it's running the expected environment, the expected hardware with the right properties, the expected code inside. Only at that point can it ask the key management system to provide the key to be able to decrypt the data. The key will flow inside the enclave, the data will be decrypted inside the enclave, and at that point you can run the model and you can get back the predictions.
[00:16:44] Something to pay attention to is that the prediction might also be sensitive data, so you might want to encrypt the artifacts that come out of your model, and this is also something you need to manage. But at least the good news with confidential computing is that it gives you the ability to have end-to-end encryption and never leak the data of your customers.
[00:17:03] So when we think about deployment of your model, like I said, there are two options. On the left, on device. That's the best one. If you can do it on device, you should do it on device, because that's where you get maximum privacy. It's fast, it's offline, it only runs on the client side. You don't have to bother about whether your server side can be attacked by the bad guys. But at the same time, it's limited. You're limited by what the device can do. It has limited performance, limited context, limited data flows, and so on.
[00:17:29] And on the right, when you need larger context, more compute, then you go to the cloud, and here you can use confidential computing as a way to achieve that. It's a bit more complex, obviously. Like I mentioned, it's a black box, so you can't really see what's happening inside the hardware enclave. So the iteration loop for the engineering team is a bit harder, but it still protects your privacy boundary and allows you to achieve what you're trying to achieve.
[00:17:55] To wrap up the second use case, two main lessons. We don't always need very complex AI systems. Try to do basic feature engineering, or try to do basic old-school AI with machine learning and so on. That may be enough. And we'll have another talk about ML right after me, so maybe we're going to talk about this again. And then if you need to run AI in the cloud, make sure you pay attention to end-to-end encryption. It's really important to not only think about encryption at rest and in transit, but really encryption in use. And this is where confidential computing is one of the options to leverage to achieve that goal.
Privacy is a leadership decision
[00:18:29] All right, stepping back and trying to wrap up with a mini playbook around this. If I think about how I apply this in different use cases, the first thing is that if you want privacy, it's always a question of trade-offs. There's no architecture out there that's going to help you achieve all your goals. So you need to trade off between privacy on the one hand, model performance, and velocity.
[00:18:50] And at the end of the day, it's more of a leadership question, more of an organization question, than a technical one. We can build ML systems, we can build infrastructure, but you need an organization that has the courage to say, "Okay, no, privacy matters for us, and this is what we want to achieve." It needs to start from there, from the leadership, as a clear mandate, and also having the courage and the authority to say that, even when there's business pressure and you want to go faster, privacy still matters. For me, it's really about building AI systems that you can trust technically, legally, but also that you can stand by ethically.
[00:19:30] One of the issues I see often when we talk about privacy is that it comes too late in the flow. You think about privacy at the end. It's more of a legal or compliance checkbox you need to check at the end of the flow, but that's too late. It needs to be at the starting point of the conversation. It needs to be one of the requirements, part of your product requirements.
[00:19:48] This is how we approach it at Dashlane. We have privacy and reliability as tier-one requirements. This is what is critical, and we can't compromise on those. Of course, that's because we are a zero-knowledge architecture. And then the other ones, like accuracy and latency, are important, but that's where we can optimize and where we can flex.
[00:20:07] And we have made those requirements into a series of documents internally. We have our privacy stance, which is our philosophical document that explains how we think about data. It's grounded technically in our zero-knowledge architecture. It's of course part of our privacy policy that's available to our customers. And then you have different documents, like our acceptable use policy, that guide the teams internally in how they can use data or not. And of course we have an AI policy to also guide the teams in the use of AI internally and in our product. That means that when we think about AI use cases, we start from the same foundations, from the same ground rules, and it avoids restarting the conversation from scratch each time and debating what we can or cannot do.
A three-horizon playbook
[00:20:51] The question I often get when I talk about these topics is, "Oh sure, that's great, but what should I care about, and where do I start if I want a more privacy-preserving AI system?" So here's my take on a mini three-horizon playbook that hopefully works regardless of where you are today.
[00:21:07] Maybe next week, pick one use case where you think you should have more privacy. Maybe it's because you're in EdTech and you manage kids' data, and you think, "Okay, today I'm actually seeing all that data, but maybe I should not. I don't want the liability." So build a one-pager, which is not a product requirement, which is not a technical requirement, which is a privacy requirement. Try to define in that one-pager what type of data is acceptable to see or not. What are the data flows? What are the outputs that are acceptable? And then get the right people in the room. Yet again, this is a cross-functional discussion. You need the engineering team, you need the product team, you need legal, you need security. Everybody needs to be aligned on, "Okay, we agree this is the privacy boundary and the accountability we have as a group."
[00:21:48] Then in the next 3 months, you can start making it real. Use that privacy requirement as a way to map your end-to-end data flows, to understand, "Okay, if I deploy, what type of deployment should I use? Should I do it on device or in the cloud? What is the risk of sensitive data leaking? How can I debug and monitor the data from the models that are going to run in production?" Build your first prototype, and then of course you need to put it in production to test it. Two things about putting it in production: make sure you can roll back in case you realize that your data is leaking. And also, very importantly, think from the start about the monitoring and the telemetry you're going to need to be able to keep improving your model.
[00:22:25] And then in the next 6 months or so, you can run that controlled production pilot, you can monitor it, you can learn from it, and then you can try and repeat. For me, it's really about having that muscle of being able to apply privacy to your AI use cases in a repeatable way, with the right threat model that is shared by everybody, with the right compliance checkboxes along the way and not at the end, and making it a muscle for the organization to think private by design from the start.
Five takeaways
[00:22:52] All right, five takeaways I'm hoping you can take back to your own context. One: you don't have to choose between AI and privacy and user trust. Both are achievable, but it needs to be a choice from the start. You need to think, okay, I want to do things in a private way from the start. Every context is different, but I encourage you to think about privacy, because we're all going to be liable down the road for leaking customer data if we mess up.
[00:23:17] Two: as you think about training your models and building your own data set, think coverage, not volume. You don't need to collect a lot of data, and I think we've all had an excess of collecting too much data from our customers in the past. Go back and think about the diversity of data you need, and just collect whatever you need.
[00:23:34] Three: start smaller. You don't need large models. Large models are great, but they're also very complicated to apply and to train and to debug. Start smaller. Frugal modeling is really a good place to start, and only escalate when you need more.
[00:23:50] Four, for deployment, which is critical: if you can do it on device, I really encourage you to do it on device. You'll get maximum privacy, also better security and so on. And only go to cloud compute if you need to, but that's more risky, obviously.
[00:24:04] And five, yet again, I don't think we need complex AI systems for everything. That may be controversial in today's age of "all the world is AI," but I think good feature engineering can be enough. An old-school ML model can be enough. So try to apply the right tool to the right use case, and don't try to boil the ocean with a very complex AI model from the get-go. Thank you for listening. I hope it was useful, and I'm happy to take questions.
[00:24:33] [applause]
Q&A
[00:24:35] Host: Awesome. I was rapt. My engineering side is just trying to figure out this whole situation, so thank you for going into the depth. All right, let's get into the questions. "Have you ever run into a situation where a regulatory body demanded access to the data in the black box? How do you handle that, while of course having told the users we're trying to be private with the data?"
[00:25:00] Frédéric: Yeah. All the time. We always get requests, from the FBI to police enforcement to regulatory bodies, that ask us sometimes about data from customers for investigations or that type of stuff. Always the same question, the same answer: we're zero knowledge. We don't see the customer data. We don't have the key to your data. We can't decrypt it. Which, by the way, has been a customer experience issue for a long time, because if you lost your master password, your key to your data, we couldn't give you back your data. But that's proof that we can't see it, and that's always worked. So we couldn't see the data even if we wanted to. Cryptographically, I mean.
[00:25:36] Host: "How do you manage on-device models when there are limitations in CPU and RAM?" Of course, all the different phones you've got, or computers, or iPads.
[00:25:48] Frédéric: Yeah, that's where it's complicated. That's why your model needs to be compact enough so it can run with the right performance on the device, and that's a lot of testing on our part. Like I said, we're running browser extensions, and we're running on mobile devices, where we have even more constraints. So there are things we can do on certain devices and some we can't do on other devices. We try to get to the minimal footprint and the minimal CPU and hardware need to be able to run those models both efficiently and reliably. It's a lot of benchmarking and testing with very old computers and stuff like this. A lot of benchmarking.
[00:26:22] Host: You're pretty technical, so for this crowd, what's the size of that model in May 2026?
[00:26:28] Frédéric: For instance, the phishing detection model, to give you a sense, is a few megabytes.
[00:26:33] Host: Oh wow, so it's really small.
[00:26:34] Frédéric: It's really compact. But you're limited in terms of what you can load into an extension, and Chrome will complain if you try to submit an extension that's too large. And also, we want to run that model in, I think we are at 18 milliseconds in terms of speed of running the analysis.
[00:26:50] Host: 18 milliseconds.
[00:26:50] Frédéric: 18 milliseconds. And that's because we look at the whole page, and we don't want to impact the user experience as you're browsing the page. So it needs to be really low footprint.
[00:26:58] Host: Do you have a user experience element where the extension is blocking access in real time, or is it more of just a little suggestion?
[00:27:06] Frédéric: No, we're doing a little pop-up that tells you, "Oh, that website looks suspicious. Our confidence level is, whatever, 98%. You should pay attention; don't trust it." We're not blocking anything yet. The user needs to choose at the end of the day.
[00:27:21] Host: Do you have a consumer version for us?
[00:27:23] Frédéric: Yeah, there's a consumer version.
[00:27:24] Host: Oh, there is?
[00:27:25] Frédéric: Yeah, we are both a consumer product and a...
[00:27:26] Host: Is it free?
[00:27:28] Frédéric: It's not free. But you can try it for free.
[00:27:30] Host: Is there a code for us? [laughter] Okay. "With the frugal edge model," and let's say maybe even the cloud model, I think this is the same problem, "how do you resolve recurring issues with specific data from customers where the model failed?" And I'm thinking, it's been, I don't know, 10-plus years with Siri. Siri still can't dial my wife's name. It's the only person I dial on my phone these days. Some of these models are just stuck in time, and Apple is bypassing their own Apple Intelligence. How do you go beyond the Siri problem?
[00:28:07] Frédéric: Yeah, it's hard. One example we have currently is, on a website where you have 2FA, the OTP field, sometimes those OTP fields are coded in a way that they look like passwords. So our model will detect them as a password, and then it will do things wrong. And it's very hard, because one, there are not that many that are accessible; you need to be behind the authentication. And also a lot of those are in work environments, so they may be behind a firewall or an intranet of our customers.
[00:28:39] Frédéric: So that's why we need our employees to also help us try and collect that specific data, but it's hard, because we don't have that many employees. But sometimes we also leverage our own customers, and we'll tell them, "Okay, we've seen that issue coming from your use case with that type of thing. Can you help us test and collect data in a more controlled way with you, so that we can help train the model?" But it's a hard problem to solve.
[00:29:00] Host: Yeah, it is a hard problem. All right, everybody, round of applause for Frédéric. Thank you so much.
[00:29:05] Frédéric: Thank you. That was great. Thank you.
