Navigating Challenges and Technical Debt in AI Engineering for Finance
Checking session availability…
Hang tight while we load the latest updates.
Join Ahmed Menshawy as he shares the intricate balance between maintaining model complexity and meeting the operational demands of real-time processing in the financial sector. Dive into the practical challenges and solutions in deploying advanced machine learning technologies in a large enterprise.
Navigating Challenges and Technical Debt in AI Engineering for Finance
Ahmed Menshawy at UXDX EMEA. Video: https://youtu.be/ZHNCktQYNQE
Readable transcript: edited from the recording's captions for readability (fillers and false starts removed, punctuation and section headings added). Wording is the speaker's own. Timestamps are positions in the video. Names marked [?] could not be verified against the audio.
AI engineering at Mastercard and what this talk covers
[00:00:07] Today I'll be discussing AI engineering efforts at Mastercard. My team is mainly responsible for building the end-to-end machine learning pipeline around some of the AI systems that really touch on your daily lives, and we really add intelligence to every single transaction switched by Mastercard. I'm super excited to talk to you today about GenAI and how we use it at Mastercard. It is an exciting time, because we are standing at the edge of another technological era, where a powerful relationship between human and technology is unfolding right before us.
[00:00:46] On the HBA [?] side, I will try to touch on how AI is expanding, moving from excellence in structured data to unstructured data, and how it's really augmenting human intelligence and not really, doomsday aside as well, risking our jobs. And finally we will talk about the challenges and technical debt associated with deploying such models in production.
[00:01:13] Over the last two decades AI has been doing really well in labeling things: fraud, not fraud. You give it an image, it can detect the objects in the image. But this is not the reality in most of the organizations. Most of the organizations' data is really unstructured, and it's estimated that more than 80% of organizations' data is unstructured, and more than 71% of them really struggle managing and securing such data.
[00:01:44] Now with the GenAI technology you can easily attach this domain-specific data to your generative AI application and be able to formulate answers based on this domain-specific data, and at the same time discover relationships and patterns that you wouldn't be able to do using the traditional AI or neural-based techniques that we had before.
Two hundred years of humans and technology, and Ada Lovelace
[00:02:06] As I said, the relationship between human and technology is not really old. Over the past 200 years, mathematicians and visionaries have been really dedicating their lives to creating technologies that can reduce human labor and at the same time automate complex analytical and computational tasks. I recommend you this great article if you are worried about GenAI and how it's going to take over our jobs. I recommend this article from Nature, which says stop talking about AI's tomorrow's doomsday when AI poses risks today.
[00:02:46] There's current risk associated with AI, like bias and fairness, and how really we can get it to solve this kind of issue instead of talking about the speculations of doomsday and AI taking over, which is not really happening anytime soon. We don't have the algorithmic foundations to be able to create something AGI, or something that can really originate anything by itself.
[00:03:13] One of these visionaries in the past 200 years, Ada Lovelace, also known as the world's first computer programmer, she had a lot of speculations about AI, or as she called it at that time, the analytical engine. She had good speculations that AI will evolve to be able to understand music notations and later on be used to create music notes, which is what we use now and what GenAI is capable of.
[00:03:41] But despite her belief in the technology, and despite her belief in the ability of this technology to reduce human labor, she is very clear that AI, or the analytical engine, cannot really originate anything by itself. It can only do whatever we order it to perform. And it is true. This is more than 180 years old, and this statement or quote still holds. We don't have the algorithmic foundation, as I said, that can really help us originate anything by itself.
What ChatGPT fixed, and GenAI at Mastercard
[00:04:16] Some folks have a misconception that language models, or the large versions of it now, are created by OpenAI, which I hope you don't share. The idea of a language model, and the whole idea of being able to use your context for predicting the next word, is a really very old idea. But what is missing with this language model is that you are not able to have this instruction dataset. You are not able to interact with your data as you do with another human being.
[00:04:49] This is what ChatGPT really fixed. The instruction dataset was really rare on the internet, and what OpenAI did is collecting a lot of data, outsourcing this to contractors to create an instruction dataset where you have a specific prompt and then ranked answers for this prompt. Then they use this with the reinforcement learning technique to be able to fine-tune the language model further. With the internet-scale data it became large, but then you can also interact with it as you do in your normal life, as you would typically talk with your colleague, and be able to ask it to perform certain tasks.
[00:05:28] With these new advancements, GenAI technology specifically has been used to really customize so many applications, but at the same time manage the tremendous unstructured data that we have within our organization. Mastercard is not far. We have recently augmented our fraud detection capability with GenAI, and we have managed to achieve up to 300% boost in our accuracy for some corner cases that we were really struggling with in our AI solution.
Four essentials for building a GenAI app
[00:06:01] Coming to the last part of my talk, which is really about how can we take such complex and large language models into production. To build a GenAI app you need a few essentials. First, you need access to a variety of foundation models, and this is like the Llama model that you keep hearing about, or the OpenAI GPT models. Second, you need an environment, a secure environment that you could use for building and customizing these domain-specific applications. Third, you need a variety of tools to be able to train and fine-tune such models and later on deploy it to production.
[00:06:41] And finally we need a different set of infrastructure. Really, the traditional infrastructure we use right now for typical machine learning or neural-based applications is not really applicable for large models. We need a wide variety of accelerators that could be used for serving such intelligence to our customers. I've tried to color code each of these essentials based on the challenges we would see in some of these stages.
[00:07:10] Access to a variety of foundation models, I think this is not very challenging now. Thanks to Meta and the open source initiative, we do have access to a lot of open source models. Second, the ability to have a really secure environment to contextualize the large language model is not so challenging as well. Most of the companies have their own AI environment that they use for building a lot of applications, but they need to augment this environment with some other accelerators that can really be used for fine-tuning such models. We will see in a bit as well the different approaches that organizations are using for this contextualization aspect, and how you can fine-tune the model with your own domain-specific data and reduce the hallucination aspect of GenAI.
[00:08:03] The most challenging aspect of these essentials really is having access to a variety of tools that can help you build and deploy such models to production, because the size of the model is not really a traditional size that we have seen in the past number of decades when dealing with AI. It's a gigantic size that we have never seen before, given the internet-scale data that it was trained on.
The model is 3 to 5% of the pipeline: closed book versus open book
[00:08:29] This is taken from our paper. We have recently published a paper in ACM which shows clearly that the LLM core [?], or the open source Mach C [?] model itself, like the Llama model for example, is at the heart of the GenAI application that you are building, but it's only 3 to 5% of what goes into building the end-to-end pipeline. The rest, the 95% or more, really goes into the other components around it. How can you make sure that you have the right guardrails around your GenAI system? How do you make sure that you have the right accelerators and the right GPUs to be able to serve this model, and serve it at scale given the latency and throughput requirements coming from business?
[00:09:16] There are two ways right now people are using GenAI. There is the traditional ChatGPT way, which is what we call closed book. The model is already trained on the internet-scale data, it keeps getting updated, but there are a few problems with this idea of using it as it is. First you have the hallucination problem, and they do it in a scary way, in a very confident way. And also attribution: you can't tell why the content of this model is generated the way it is. You can't have this kind of attribution and interpretation of the output, which is a requirement for high-risk AI systems.
[00:09:56] Third, they go out of date, and that's why you have the different releases from Meta, you have the different releases from OpenAI. And also revisions. Folks have the right to opt out of using specific AI applications, and under the new regulations we are asked to remove the folks' data from the training data. With GenAI, every time someone is opting out and you are trying to retrain the model from scratch, this becomes really expensive.
[00:10:29] And customizations. They are trained on internet-scale data, but how can you customize it for your own specific needs? For example, in the fraud detection solutions that we have built, how can we make sure that it's really specific to the domain and equip it with the terminologies or the tokens that can help it achieve the accuracy that we have achieved?
[00:10:50] The solution to all of this, and this is what most of the companies are using nowadays, is really attaching an external memory to the LLM. Imagine having the large language model and then extending this with an external memory, or what we call the open book. Very similar to when you go to an open book exam, you are answering the questions but you are given the context from the book that you have next to you. You can easily do attribution, because you are providing the context of the prompt to the LLM before it makes its answer.
[00:11:24] If someone is asking a specific Harry Potter question, for example, you are giving it the Harry Potter book or volume as part of this prompt itself, so that it can derive the answer from the context that you are providing it. It is solving all of the problems that we have seen in the previous closed book approach, but at the same time it's making the system more factual. You are providing the context, so the system will be more factual in adhering to the context or the domain-specific data you have given to the generator.
[00:11:55] But there are so many challenges really operationalizing such models, and I know a lot of it is really specific to AI and a bit technical as well, but these are some of the challenges that my team is handling on a day-to-day basis, and these are the challenges that account for more than 95% of what goes into building such systems. I think I can skip this one. That's all I have for you, folks, and this is some of the books from my team here in Dublin as well, so feel free to check it out. Thank you.
Q&A
[00:12:36] Host: Thank you very much, Ahmed. Hallucinations are an obvious hurdle in the enterprise adoption of GenAI. Are there any immediate steps an org can take to mitigate this problem?
[00:12:49] Ahmed: Yeah. The open book approach is definitely what we are using for the hallucination. It is really trying to provide the context for the GenAI application instead of using the model as it is. Instead of using the Llama model or ChatGPT API as it is, we are providing the context as part of the prompt. Someone will ask a question, and then we have a retriever component that will retrieve the domain-specific data from our corpus and then attach this to the prompt before it goes to the generator for the answer. So it's formulating the answer based on the domain-specific data that we have attached to this prompt. There's a paper around it as well that folks can check out. It's called Retrieval Augmentation Reduces Hallucination in Conversation. It shows how this closed, sorry, open book approach is really reducing the hallucination of the system.
[00:13:47] Host: Okay, great. Next question. What are the typical UX research tasks where we could benefit from using LLMs?
[00:13:56] Ahmed: Yeah, I think it's really there. I mentioned about the instruction data and the user experience in how they can interact with large language models as we do on a day-to-day basis, and this is what ChatGPT really solved: having the ability to ask the LLM, or ask your GPT assistant, to do certain tasks. So again, having access to this instruction data is really what makes this language model stand out. It enhances the user experience of how you can interact with such models. But I think there's more to it. I think there's a lot of research really around how we can enhance this relationship further between this great technology and how humans are using it.
[00:14:40] Host: And what's the biggest challenge now?
[00:14:42] Ahmed: The biggest challenge right now is hallucination. I said that the open book approach is really solving the problem, but it's not really making it go away. We still have a bit of hallucination, and it's really embedded in the core of the algorithm itself. There's a lot of research and a lot of work that my team is doing right now to really make this go to at least single-digit percentage. Instead of the 50% that a closed book hallucination will give you, with the open book approach you can easily go to 15% or 20%.
[00:15:21] Ahmed: But then with additional measures that you are taking, it is really trying to compare your output, the GenAI output, with your domain-specific data before you can feed it out to the customer. Instead of just serving the output of the LLM model or the GPT assistant, you're also comparing this to your domain-specific data or your documentation, and making sure, using some statistical techniques, that there is not much difference between the two distributions of data.
[00:15:48] Host: And the last question there. What do you mean by hallucinations? It could mean many things.
[00:15:56] Ahmed: Yeah. Hallucination is really going out of context and trying to answer things confidently that are not true. Hallucination is maybe not really the right word, but LLMs will try to convince you in a very scary way that the false information it is providing you is the right one. This is what is sometimes scary, because there's a lot of users getting access to GenAI and using ChatGPT, and a lot of them don't really fact-check what the GPT assistant is saying. So hallucination, in short, is really giving false information in a confident way that is not really resonating with the actual information.
[00:16:41] Host: Awesome. Thank you very much, Ahmed.
