
Evolution Exchange Finance Podcast · 2026-06-11 · 48 min
Key moments - from our scoring
Substance score
42 / 100
Five dimensions, 20 points each
Production AI deployments face radically different challenges than sandbox proofs of concept. Vikram Saraswathi, lead data and AI engineer at Nadia Bank, opens with the core problem: POCs use curated data samples in isolated environments, but production reveals data quality issues, legacy system inconsistencies, and regulatory constraints teams never anticipated. Alexandra Jevinger, who bridges product and technology in financial services, emphasizes that production requires reliability and continuous monitoring - not just a working demo. Einar Scully Hafberg from Norwegian insurance insurance adds crucial context: traditional machine learning models have evolved for decades, but newer conversational AI (large language models with "sexy voices") introduces wholly new failure modes that existing data pipelines aren't ready to handle. The panel identifies three critical failure patterns that emerge only post-POC: data readiness gaps (poor metadata, unmapped business terminology, unclear data ownership), security and governance blindspots (PII handling, GDPR compliance, external LLM data exposure), and accountability vacuums (no clear owner for production model risk when things go wrong). Context management breaks down because users ask questions differently in production than developers anticipated in POCs. Human involvement arrives too late - teams build fully automated systems before validating how actual users would interact with them.
POCs use curated data samples in isolated, perfect conditions, but production reveals data quality issues from legacy systems, regulatory constraints, PII handling requirements, and user questions that don't match the assumptions made during development.
Teams lack common data glossaries, clear metadata describing business meaning of database columns, understanding of who owns data models, and mapping between technical attribute names and business terminology that LLMs need to interpret correctly.
Regulatory constraints around PII data, GDPR compliance, restrictions on what data can pass through external LLMs like Claude, uncertainty about where external models host or retain data, and questions about EU data residency requirements.
Accountability is unclear in most organizations - platform teams provide the foundational LLM but use-case teams own the business risk; changes to system prompts may trigger new model risk validation; and it's ambiguous whether errors fall under platform or use-case ownership.
POC developers assume users will ask specific, structured questions (like 'show me clients with 30% concentration risk'), but actual users ask open-ended queries ('who should I call this week?'), requiring completely different context and often exposing that the model can't handle real-world variations.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode surfaces a handful of genuinely useful practitioner observations - human-in-the-loop as a design parameter rather than a fallback, the semantic model gap between physical database attributes and business terminology, and prompt-change governance as an unresolved risk - but most of the runtime is consumed by agreement-chaining and generic statements about data quality and trust that add little for an experienced B2B operator.
the right framing is like this human in the loop is a design parameter, and it is not a fallback
teams that succeed, they build the evals before they build the features they design for the edge cases and the hard cases
The mild reframing that traditional ML has been succeeding for years while generative AI is the new failure surface is a useful distinction, and the prompt-change governance question in regulated environments is underexplored in most AI discourse, but the bulk of the conversation recycles standard AI-implementation warnings with no contrarian or first-principles angle.
we've been running machine learning motors for like decades now... Last decade maybe we started with 10 parameters or something 10 years ago and now we are up to a thousand
whenever there is a change in the system prompt, who has to sign off this prompt change? Does this change in the system prompt trigger a new model risk validation?
All three guests are genuine domain practitioners - a lead data and AI engineer at a bank with 20 years in IT, a product leader with eight years in financial services, and a data-pipeline department head at a large Norwegian insurer - giving the conversation credible grounding, though none are in C-suite roles or widely recognised as domain authorities.
I'm working as a lead data and AI engineer at Nadia Bank. I have been working in the field of data and machine learning and data science for the past 15 years
the last 20 years or so I've been working in the data area with the data pipeline... for one of the largest insurances in Norway
The episode offers isolated concrete touches - the 30% concentration-risk query example, the AUM/database-attribute naming problem, and legacy code written 25 - 30 years ago with retired authors - but there are no hard metrics, no named projects, no outcome data, and most claims stay at the level of 'many organizations are facing this right now.'
show me the clients with the concentration risk above 30% in technology sector
you could call assets under management as AUM or the total assets under management... the underlying attributes on the database side could be named in a different way
The host moves the conversation through a reasonable structure and occasionally sets up follow-on topics well, but questions are broad and pre-announced ('another area of failure'), there is almost no pushback on vague claims, and the panellists spend most of the episode agreeing with each other, leaving productive tension unexplored.
what's fundamentally different between a demo environment and uh, live production?
are banks measuring the right things?
Computed from the transcript - who did the talking, and the words that came up most.
Today's episode is hosted by Charlie Beetson and they are joined on the podcast by Einar, Head Of Data Integrations at Gjensidige, Vikram Saraswathi, Lead AI & ML Engineer at Nordea and Alexandra Jevinger, Chief Product Owner - Communications & Output. The conversation explores why so many artificial intelligence initiatives struggle to progress beyond early experimentation and what organisations can do to successfully scale AI into production environments. The exchange highlights the importance of governance, data integration, machine learning operations and aligning technology projects with business objectives. The discussion also examines common barriers to enterprise adoption, the challenges of operationalising AI and the role of collaboration across teams. Broader themes include digital transformation, automation and building reliable frameworks that support long-term value creation. Throughout the conversation, artificial intelligence is considered from both a strategic and practical perspective, offering insights into how organisations can bridge the gap between innovation and measurable outcomes.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign.
Speaker B: Welcome to the Evolution Exchange podcast. A melting pot of ideas and inspiration shared by some of the most successful technical leaders in the world. The views expressed by the speakers on this podcast are their own and not necessarily representative of their organization.
Speaker C: Welcome to Evolution Exchange podcast. I am your host, Charlie Beatson. Joining me today we have Vikram Saraswathi, Alexandra Giovinga, and Inar, uh, Scully Hafberg. We have three people from different, um, organizations from different backgrounds, with different interests. So it should be a super varied conversation. But to get a little feel for those individuals, we're going to go around the room and introduce everyone. Um, in no particular order, Vikram, do you mind giving us a brief introduction to yourself, mate?
Speaker A: Yeah, sure, Charlie. My name is Vikram, uh, Saraswati. I'm working as a lead data and AI engineer at, uh, Nadia Bank. I have been working in the field of data, um, and machine learning and data science for the past, uh, 15 years and overall in it for around 20 years. Um, I live in Stockholm and I like playing badminton and, uh, cycling in my free time. And it is nice to be here in this, uh, podcast.
Speaker C: Good stuff, good stuff. Thanks for that, Vikram. Lovely to have you on board, Alexandra. It'd be great to find out a little bit about yourself.
Speaker D: Absolutely. So my name is Alexandra Jevinger and I've spent the past eight years working in the financial sector and mainly at the intersection of technology and business, which is a really interesting place to be. So started off in marketing and CRM and moved into broader product roles. And today I focus on setting the direction of and how to shape new communication capabilities, um, across different channels and customer segments. Um, so this has given me a great perspective on how product, data and technology all come together. Uh, on a personal note, I'm a mom, a bookworm and a nature lover.
Speaker C: Nice. Lovely to have you on board today, Alexandra. Thank you very much for that. And last but definitely not least, Einar, uh, could you give an introduction?
Speaker E: Yeah, sure. I'm actually from, uh, the insurance side of, uh, finops or the financial industry, if you could say. So the last 20 years or so I've been working in the data area with the data pipeline. So I had, uh, a department or a team of all the data engineers that create all the pipelines and the data warehouse. Now for one of the largest insurances in Norway. I live in Norway, uh, actually originally Icelandic. Uh, I've got grown kids so I used my spare time to go traveling. But during the summertime it's not so much traveling because then I compete in Sailing, that is my, my summer hobby. Not a hobby, sailing. We say competition.
Speaker C: Amazing. Thank you. An art. Lovely to have you on board. And I suppose for the listeners today, um, you can see that we've got not only varied panel from a organizational and experienced background but also from what they like to get up to in their, in their free time. Um, obviously we'll not beat around the bush and we'll get into today's topic which is covering uh, proof of concept into production. Uh, and in this case we're looking into the AI sector. Uh, why most AI never tends to make it coming from a proof of concept. Um, so obviously there's a few things to look at. But if we start from the beginning there's certain situations why AI proof of concepts look convincing, um, and that success rarely translates into something that the business can rely on on a day to day basis. Um, so a question for the panel to kick things off is what's fundamentally different between a demo environment and uh, live production?
Speaker A: Yeah, I think um, I can try to give a few points here about the fundamental differences between uh, a proof of concept environment and a live production environment. Most of the POCs that are done uh, are done using some representative sample of uh, the data from production environment. And in a perfectly uh, theoretical environment, in some kind of sandbox or uh, in an isolated environment, uh, where it is perfect in the sense for the POC to uh, have all the assumptions taken care of. And um, um it works perfectly in a POC environment. It doesn't truly represent the production data but has some sample of it. Uh, but when it comes to the actual production environment there you will see all kinds of uh, noise in the data. Um, you will have for example if you take any large organization it would be a combination of multiple smaller legacy systems. Data flowing from all of these legacy systems into your uh, data, uh, mesh or data lake. And then you will have differences between one legacy system to another system. No two systems agree on the same data point. There could be a lot of uh, issues in the data which you have not considered in a POC environment that could pop up in production environment. That is one point where uh, there could be significant deviations from a POC environment to a production environment that could jeopardize the results uh, that you have achieved in a POC environment compared to a live production.
Speaker D: I definitely agree with Vikram. I mean uh, POC proves possibilities but production, it's messy and it requires reliability. I mean you're working in this very small uh, controlled setting and you're using maybe curated data, uh, you have a dedicated team who are driving this and then it should be handed over. You have testing a happy path. But is the user really doing that happy path all the time? So I mean just getting a PoC without AI into production is hard. So then adding on top of that this new, still quite immature technology, um, it definitely gives challenges.
Speaker E: Yeah, I think I would like to maybe take another approach to this because I'm not totally agreeing with uh, this because we've been running machine learning motors for like decades now. Uh and it's been evolving. Last decade maybe we started with 10 parameters or something 10 years ago and now we are up to a thousand or something in some of our models. But my definition of AI includes that uh, a lot of the new stuff that is coming out now are like the avatars and the sexy voices who can answer everything and tell everything and know everything. That's where the real problem comes in. Like Vikram was uh, mentioning that data isn't ready for that kind of uh, that kind of AI. So um, we are succeeding in definitely parts of it but the new part we are struggling with, it's just too comprehensive I think still. And um, like, like uh, they pointed out I don't think we have the data pipelines or the plumbing ready for it. But we are getting there. It's, it's an evolution I think.
Speaker C: Nice.
Speaker A: Yeah. Ah, definitely. I think uh, you, you touched a very good point. Uh in um, mostly when we think about POCs, we uh think of structured data. But if you think about unstructured data and if you are dealing with uh, this multiple variety of data like the video, audio and uh, uh text transcripts and so on, the complexity just explodes actually.
Speaker C: And um, I think that tees me up quite nicely for the first main point for the conversation today when we look at failure, um, the first kind of foundation that we have as a challenge is the data readiness, uh security and guardrails. So when teams say that the data is ready, what's usually missing,
Speaker E: usually missing is that we actually don't know maybe the data so much. Like if you come from a huge bank environment with old legacy mainframe systems where you have like still the eight character column descriptions and so on, it's like very few people who actually understand the whole model and uh, can utilize it and teach it onwards and so on. And that's a little bit I think uh, where we struggle at the moment. Especially if you take in several systems with several decryptive naming and all of those things, uh not to touch on. Maybe we will touch on that later. The metadata, the data about data. How do we explain the data to the AI and to the analysts and to the business. But um, I think that is where the big challenges come in actually.
Speaker D: Yeah, the common data glossary of having the same data point called the same thing across any organization, uh, that is a hassle. And also when it comes to accessing the data, uh, you might not have the right way to fetch the data or it's some super hard process and you need to go to someone and you need to ask someone else to get it. So there could be multiple pitfalls uh, in this.
Speaker A: Yeah, it is a very valid point uh, that both of you, Alexandra and Inar mentioned. Like how do you describe your data attributes to an LLM which could uh, uh have difficulties in interpreting the business terminology associated with this data attributes correctly? So that there could be uh, n number of ways in which a particular question could be on could be asked by the end user but your AI system has to answer it correctly. Um, for example you could call uh, assets under management as AUM or uh, the total uh assets under management or it could be multiple other things when it comes to um, the holdings positions and so on. But the underlying attributes on the database side could be uh, named in a different way um, which is as per the convenience of the database developers and so on. This is where I think the, the teams have to spend more time in making this semantic model that maps between these actual uh, physical database attributes and the business context, the business meaning of these attributes. So uh, one of the key failures for uh, these POC projects in a POC environment is probably teams are not spending much time in describing or adding sufficient details uh, in the semantic uh model of that particular uh AI use case.
Speaker E: Yeah and sometimes you create some use case for an AI, but that's maybe not what it will be used for in the end because like you said you don't know the question that it's going to be asked. And if you haven't mapped your data because some of the data can be used for these purposes and not for these purposes but you don't know what the question's going to be. So if the AI has access and so on, it needs to be defined.
Speaker C: Nice. And then looking into the data quality and access issues that are silently killing AI use cases before anyone notices um, what sort of security or guardrail concerns tend to surface only once you try to move beyond the sandbox.
Speaker A: Yeah, security and guardrails. It's A very big topic, I would say. Like first, if we take uh, security, I mean, um, in a POC environment, people assume that everything is accessible to everyone. Um, I can access this entire data set. I can ask uh, so many questions around this data and I expect some answers to be given by the LLMs. But then, um, when you start moving to the production environment, that is where the reality hits hard. Like uh, there will be a lot of uh, regulations around PII data. What uh, data could be allowed to be uh, passed through an LLM? Uh, there could be so many regulations, mostly in regulated uh, industries like finance and healthcare, uh, where these restrictions are more tightly applicable rather than in some other uh, environments like in um, uh, gaming industry or in retail, uh, for example. So this is where a lot of background work needs to be done trying to classify your data accordingly. What is a, uh, confidential data? What is PII data? And then what data can be uh, passed in through the LLM? What cannot be passed in. So all this background work needs to be done um, before you move from a POC stage to production.
Speaker D: I totally agree with you Vikram, that it's easy to get caught up by. We have, or you have this cool model and you want to launch it and you want to start get going, but you don't have the foundation of the basic sets, uh, including looking into the actual data readiness as we talked about now. So I think that for sure it's an issue.
Speaker E: Yeah. And also the issue that we aren't sure who will actually have access to it behind if we are using like an AI module somewhere like Cloth or something external. What actually does he access to? What actually does he uh, does he keep somewhere, uh, momentarily or forever? There are a lot of scenarios there that we don't have the answer to at the moment.
Speaker A: Exactly. And uh, there is this another aspect of where these LLMs are hosted also comes into picture when you think about GDPR regulations and uh, is it hosted within EU or outside eu? So all those uh, considerations also needs to be taken care of.
Speaker D: Yeah. Data readiness is often assumed, but it's not validated until you are past the POC stage and going into production.
Speaker C: M. I suppose then. Go on, Einar.
Speaker E: No, I was going to say, I mean we have leaks everywhere and so on, and we cannot afford to have like uh, you say healthcare data, uh, or insurance data, uh, been leaking somewhere because some AI motor temporarily loaded some data somewhere, uh, it needs to be airtight
Speaker C: and then teeing up to the kind of next area of failure. Right. Um, once that Sort of. Well, assuming the foundations hold, who then owns the AI system once it's live in a bank? And who is accountable for those potential mistakes if security isn't tight? That's something that's not usually discussed in your proof of concept, right?
Speaker E: No, we struggle a little bit already in the plumbing. Who owns the data foundation in the first, uh, place, like who is the data owner? But then when you've gathered it all together in some uh, semantic model and it's a lot of data sources together, who owns then that or who is responsible for exactly that data model then? And then there are a few hands in the air.
Speaker A: Yeah, I think it is like um, there are two different factors that comes into picture. One is like uh, the foundational capabilities. If it is provided by some kind of platform team, who would enable the use case teams to build the use cases? Uh, they would be the model owners for just providing the foundational LLM models. Okay. Uh, I can take the risk of uh, hosting uh, this particular LLM model. I have already covered it during, in my uh, model risk assessment process and I got approved for hosting uh, these models through my platform. And whatever use cases that you build on top of uh, uh, in your business area using my platform, you will be the model risk owner for those models. And I can't take the responsibility of model risk for the use cases that you build. So this is the usual uh, flow actually. But then, um, the model risk assessment for the actual use cases itself is like a quite uh, tedious process I would say. Like then who owns the risk of. If the, if the LLM incorrectly answers a particular question, uh, does it fall under the scope of the platform or does it fall under the scope of the use case? It is uh, very much uh, uh, it's a kind of gray area. Right, like who owns that uh, uh, risk or who is the owner of this particular uh, scenario. Like where the LLM is answering incorrectly, but it is through your model that you are getting the answer.
Speaker E: It's a huge risk area. I would say that uh, you don't want to take the risk. You don't know the AI. Well of course, if you created the LLM yourself and you know all of the parameters and everything, uh, but sometimes when we're talking about these new avatars or uh, the ones answering questions, uh, from the customers and so on, then who is responsible for the answer the AI uh gives? That's where it uh, becomes really, really difficult.
Speaker D: Yeah, I think that's also like if no one owns it in production, it Won't survive because then everyone will get scared that I don't know what's going to be the output. I don't want to take this risk or responsibility. So I definitely think you're both onto something. That it needs to be a clear ownership and all these ifs and buts needs to be taken into account if this happen, who owns that outcome, etc. Yeah.
Speaker A: And also these models, when deployed in uh, a particular environment, it doesn't stay the same um, for a longer period of time continuously. The system prompts needs to be adjusted and whenever there is a change uh, in the system prompt, who has to sign off this prompt change? Uh, does this uh, change in the system prompt, does it trigger a uh, new model risk validation? So these are all some of the questions right? Like which, which needs um, a lot of uh, teams that are involved in the entire solution to discuss and agree to a common uh, statement like okay, these are the changes that can be approved without a further model risk validation. Um, but these are the set of changes that needs a significant, that could impact significantly and it needs a new model risk validation process. So it is like uh, the thing is all these processes are evolving at the same time when use case teams are trying to develop and release these use cases into production. So that's the challenge that um, many organizations are facing right now.
Speaker C: Nice. Um, okay. And then when we look into ongoing validation of models, um, let's say they are in a live situation. What tends to be missing when it comes to that? Um, are banks measuring the right things?
Speaker A: Uh, validation of the models when it comes to um, evals. Like there could be multiple things that uh, can be set up there in terms of evals, uh, which are like it is necessary but not uh, sufficient I would say because these evals are kept on a set of um, questions uh, or set of holdout data set. I would say like okay, I give these kind of questions to them to my model and I expect these kind of answers. Um, so they will serve one set of purpose but they will not cover the entire scope of the questions that could be answered that could be asked during a live ah production environment. So evals in addition to in combination with human uh, feedback or human in the loop, could be something that could cover the risk to a certain extent and could, could serve as a good validation point.
Speaker D: Yeah, I mean from a, from a product perspective, uh, you cannot just launch an AI model. Then you're like good, now let's move on to the next feature and do something else. You continuously need to monitor. If it's uh, it's interacting with customers, what are the input from the customers? Do we need to tweak something? Do we need to. So I think continuous monitoring also from the product and business side, when you're launching an AI model, that's key too. Yeah.
Speaker E: Ah, I think uh, you're correct because we don't know if we have some data that is allowed for some purposes, but all of a sudden you use it for something else. The AI is in the loop and he doesn't know what you're using. It was just answering your question.
Speaker D: Nice.
Speaker C: Okay, cool. And then looking into another I suppose failure. Uh, it's the cognitive layer. Right. So context, memory, prompt, risk, something that Vikram, you did touch upon already. But what does context or why does context management break down so quickly when it's in that kind of real world workflow?
Speaker A: Yeah, I think it is. One thing is uh, when it comes to the context in the POC environment, there could be some assumptions made by the team uh, developing this use case and they could come up with some set of uh, questions that they could think of that the user would ask the model and uh, this is the response that they would get. And they could go ahead and build um, their use case around this particular assumption and uh, start building, start providing the context to the LLMs, uh, in this particular direction. But when the actual production users start using the system they would be asking the questions in a completely different way which the model or the um, AI use case has not completely thought of. So this is where the context doesn't completely match, where in a POC environment the assumptions were completely different compared to a live production environment where it is uh, at a completely different level. For example, um, if you take the case of a financial industry where some relationship manager is using the system and is using this AI system that was developed in a POC environment and in the POC environment the team might have thought of these kind of questions like show me the clients with the concentration risk above 30% in uh, technology sector, for example, um, and they would have assume that okay, this relationship managers would uh, ask these kind of questions to find out the customers that they need to talk to. But when it comes to the actual production environment when they are about to launch it into production, um, and let's say the advisors have got a pilot version to test it, they would ask just simply like uh, who should I call uh, this week? So they wouldn't ask these kind of questions like okay, show me the customers with a concentration risk above 30% in technology sector, uh, which the LLM would answer perfectly because it was built for that. Whereas if they ask, whom should I call this week? Probably the model would just say, like, I, um, can't answer this question. So it is like that. That is where, uh, the context is not, um, built correctly or provided correctly to the model. These issues could, uh, pop up if you haven't thought about, um, how your actual users, actual end users of the application would use it in a live production environment.
Speaker D: Yeah, this comes back to the happy path as we talked about previously. If it's a model that you have internally, it might just be an inefficiency or you're like, oh, okay, couldn't answer, uh, well, I know what to do. But it's a whole different stake if this is a model towards your customers or your clients and starts to answer in ways you were not predicting that it should do. So, um, testing is a key, for sure. And testing these different paths.
Speaker E: Yeah, definitely. I'm not that technical in all of this stuff, but I understand that if you create an AI model on some data with some context, then it shouldn't be allowed to answer some, uh, other context questions, but with the same data. So you should have like multiple AI or LLM models answering different questions, which they are allowed to do, different people, but we still have the same data that's just like, yeah, operational. Yeah, I don't know how that scales, actually. That is difficult.
Speaker C: Nice. Um, and then importantly and interestingly is the human side. So obviously you can have a system that works technically well, but then you have the people that use it. So again, something that we've kind of briefly touched on there, but where do the humans typically get brought into the loop in this case? And, um, when that is, is that too late?
Speaker E: What do you mean? Where in the process is it too
Speaker C: late in terms of utilizing, uh, the actual products? Um, I think something mentioned there by Vikram was, um, the context as well. So if someone is, um, using, um, the service and they're not knowledgeable as to what data it has or what it can be used for, then they get the wrong idea of it. So essentially, yeah, if you've got the product there, um, where is a good time essentially for humans to be brought in the loop, um, to prevent those issues or to make sure that it's up to scratch?
Speaker A: Ah, One of the things that I have observed consistently, the patterns where Most of the POCs have failed is where they brought in the human in the loop very late in the process. Uh, this pattern is, uh, depressingly consistent, actually, across multiple, uh, failed POCs that I have seen. Um, it is like we build a fully automated agent demo it, we get excited and, uh, start the production rollout, and, um, then we hit the first incident. Then, um, this is where, uh, we need to bolt on a human review step. Yes, Alexandra, you have, uh, no, no, you can continue. And then, um, the right framing is like this human in the loop is a design parameter, and it is not a fallback. And for each use case, we should decide, uh, up front, like, which decisions are made, uh, by model only. M. Which are the decisions where the model suggests and the human approves, and which are human only, uh, with some assistance from the model. So if we can stick to these three different patterns and then correctly categorize, like, which of these decisions, uh, fall under which category and then build the system accordingly, then there is good chance of reducing the failures.
Speaker D: Yeah, I just want to emphasize on what you said that humans are coming in too late. I also think when you're building a model, you need to assess which humans should be involved in training it and testing it. Because if, let's say there are AI engineers building the model and it's used towards customers, but customer service is nowhere involved in this, in training or validating, then you might have problems because these AI engineers might be focusing on something and then they move on to the next model and they don't take this into account. So it's not only about having humans involved. It's also about having the right competences and the right humans involved in the training and validation before going out in production with a model.
Speaker A: Completely agree.
Speaker E: Yeah, I think you're right, Vikram. It has to be designed for where the human needs to come in. But we have been successful, uh, with, uh, uh, the claims, uh, automating the claims. So there are robots or LLMs or whatever we call them, uh, going through all of the claims that we get in, uh, and I don't know how many parameters we have there now, but there are definitions of when and how and who should this be then escalated to. Because, uh, not all of the things we want the LLM actually to do. And if there are any doubts, then ask a human.
Speaker A: Yeah, definitely. And it is like, um, the one key statement that, um, the teams working with the PoCs and these kind of use cases, AI use cases, has to remember is the productivity gain comes from making the human faster and not removing them completely.
Speaker E: They're talking about, like, if we could have, uh, LLMs and they could escalate to other LLMs and then other LLMs could escalate it even further before they actually go to the. But that is like yeah, things under development.
Speaker D: But I really do think that's a valid point that you're focusing a lot on that humans should be completely erased from a process but AI can also be used to speed up the process. Um, so I think that's something that in general in this discussion at home we tend to forget that and just focusing on removing us from everything.
Speaker A: Nice.
Speaker C: Um, and then kind of coming to one of the main points around this topic. Um, obviously we've discussed the kind of challenges or the areas of failure within a proof of concept. Um, but one thing that's super important at this stage is organizational readiness and uh, the opportunity to scale. So everyone says they're excited about AI, everyone wants to use AI and make themselves more efficient and um, the organizations can be more effective in that way as well. But what does real readiness actually look like within an organization?
Speaker E: I think it comes a little bit back to what I said in the beginning, uh, our definition of AI because like I said we've been successful with uh, robots or LLM models for years. So and I don't know how many we have in my uh, in my company, but hundreds of uh, LLM models running around in our data and doing stuff. But the new use cases lately with these uh, sexy avatars and sexy voices answering all your questions with all the knowledge, ah, that's where it's really uh, hurting. Uh, and I guess that the hype around AI has become maybe a little bit too much. But uh, I don't think that we're not ready uh, for the full scale of AI. Just let it loose into all of our systems. Read everything and know everything. We're not there yet.
Speaker D: I think um, the skills around AI varies a lot depending on what role you have. I mean if you're a uh, data engineer you m obviously might have more knowledge than if you are working at customer support or something more business facing. So I think we need to have some sort of common base of AI knowledge before we could actually leverage on this uh, and use it on a full scale in the financial industry.
Speaker A: Yeah, completely agree with uh, both of your statements. Alexandra and Ener. Um, there is one more aspect that I would like to add like uh, if we see traditional software development for example before the introduction of these LLMs or the agentic uh, coding agents or uh, the wipe coding, uh, tools that are, that are available like when any project needs to be any use case needs to be developed. Uh, usually the team working, the core team working with that particular use case. They would um, start with the planning process and then uh, they come up with an architectural design pattern and start thinking about all the integration patterns. And then they will start doing this coding process and start developing the code writing unit, tests, integration tests, and integrating it into CI CD pipelines and deploying it into the production environment, uh, I mean different environments all the way up to production. So during this process, I mean the most of the time was uh, either spent on uh, the coding and making the application available in different environments through automation, through CI CD processes. And now with the introduction of the LLM models and the coding, uh, agents and uh, things like coding assistant systems that are made available to different, um, teams, uh, working with the application development lifecycle. There is a lot of uh, speed, uh, that can be achieved through leveraging these tools if done in a proper way. Um, this is where the organizations can see most of the productivity in terms of developing the solutions by utilizing these coding assistants, um, in a way that can speed up the application development but also ensuring that uh, there is consistency across the different uh, development patterns, like the architectural design patterns or how you develop the application, uh, code and things like that. So that is where uh, most of the efficiency lies in. And um, organizations, if they want to scale up the usage of AI um, in the organization, this is where they need to think about um, educating the employees and then bringing them up to speed in utilizing the AI tools effectively.
Speaker E: I think you are totally onto the point there, Vikram. I think that that is where we should be focusing our AI capabilities and actually using AI first to do all the groundwork to get all the knowledge and get all the systems understood and documented and shared somehow in a knowledge space or something. We aren't ready to expose that without understanding it first. So I totally agree with you. So that is one of my points or what I will be doing next is to try to introduce AI into understanding our old legacy systems, uh, better and how they uh, interact, uh, and what data exists and how it became to uh, reality. We've got some business functions that were written 25, 30 years ago and the people retired a long time ago and nobody actually knows how it is, but the system is. Right. So. But I think we should focus on that area.
Speaker A: Yeah. In financial industry there are a lot of legacy applications where probably the entire knowledge resides with just one person. And that uh, person might have left the organization long back. Right?
Speaker E: Yeah. But there we are Ready. And I think the AI is already scalable for uh, that area. But it's not like ready for m, making that customer outward, uh, facing AI or something like that.
Speaker C: And then the other side of the uh, readiness within an organization is their ability to work with speed. Um, essentially. Do big financial organizations have the capacity for fast speed, um, with these, you know, uses of AI or does it vary? Interested to hear your thoughts.
Speaker A: I think um, the speed in the development it can be achieved. And um, most of the organizations are there yet or there already. But there are other processes associated with uh, bringing these uh, systems into production. Right. So that is where um, the organizations are not yet there because of uh, all these risk processes and uh, governance processes. They are still in the process of getting matured. So that is where it will slow down the entire uh, um life cycle of bringing these products into production. So that is where there should be more focus on um, bringing more clarity into these processes and providing um,
Speaker D: uh,
Speaker A: more clarity and education also to other users within the system, within, within the organization to get them educated about these processes.
Speaker C: Yeah.
Speaker D: And I mean the financial sector as a whole, it's based on trust from clients and customers. And at least when I think about trust I think about um, very well thought, risk taking, transparency things is more you're thinking it through before you do something because the trust is really key. You can lose it this fast and it takes a lot of time to rebuild it. And then you have AI. AI is fast. It sometimes you are not really, let's say you have a black box model. You cannot really see every single step. And we need to find in the financial sector some sort of middle ground where we can keep uh, the trust but at the same time really start utilizing this new technology. Because I think it is key. AI I believe is here to stay. So we need to start finding our way to do it.
Speaker A: Yeah, I would completely agree with your statement Alexandra. Trust is uh, trust is the main product for financial organizations. So yeah, we can't afford to lose trust because of um, A uh, badly performing AI system. Right.
Speaker E: I totally agree. Especially in the insurance industry, we are very risk aware. So uh, fast usually means more risk. We need to take that into account as well. Yeah,
Speaker A: nice.
Speaker C: Um, and as opposed to kind of close off the subjects as a whole, but something that you might want to go into a little bit of detail around generally. So um, across your three individuals you'll have seen proof of concepts, um, and obviously the success or lack of success of them being going into production. Um, what's One thing each that you think would improve the chances of AI making um, itself from being a proof of concept product within a bank. And what would that be?
Speaker D: I would say it all ties back to what we have been talking about throughout the conversation. You need to have the foundation in place. You need to have the ownership, the data, uh, the risk assessment and not get caught up in that. You need to speed on and do the coolest thing. So have your house in order. I say that's one of the key success factors,
Speaker A: Alexandra.
Speaker E: Yeah, yeah, probably the biggest one. Uh, but I think that where we can be successful is where Vikram came in. To use it in our processes first and try to maybe speed up our uh, knowledge or our knowledge base of. Of our systems and uh, of our pipelines and of our data and categorize it, modularize it and metadata, tag it and so on that we can use AI heavily as well.
Speaker A: Yeah, completely agree. And it is like, uh, if we summarize this discussion, mostly it is like, uh, if we see the patterns across all the. All these dimensions and um, how observing these POCs, like the ones that have failed and the ones that have succeeded, there is one consistent pattern like the ones that has, uh, um, failed. Like where the POCs, they are optimized for impressing uh, a steering committee in, um, a couple of days and then they try to build rush through the entire process. But the teams that succeed, uh, are the ones that treat the POC as just 10% of the work and uh, the remaining 90%. They need to focus on these foundational aspects that we touched upon that is bringing this compliance, risk, security and uh, operations in the very beginning of the poc, that is on week one itself. So these teams that succeed, they build the evals before they build the features they design for the um, edge cases and the hard cases. And they accept that production in a regulated environment is, uh, something fundamentally different than production in a, uh, unregulated environment. Uh, that's the key, um, difference. And uh, that's how I would summarize this entire conversation.
Speaker C: Makes it nice and easy when we have a summarizing point from each of you. But I'd like to say a massive thank you to the three of you for coming on for the episode today. So Ian R, Alexandra and Vikram, thank you very much for coming on and to our listeners, hopefully that kind of gave you guys a bit more of a varied, you know, thought process, varied perspective on this kind of topic, uh, which is obviously quite hot at the moment. Uh, but of course that will be it for today's episode of the Evolution Finance podcast. Myself, your host, Charlie, uh, have enjoyed hosting. Um, if you are interested in discussing a, uh, like minded or similar topic with similar individuals, or you're interested in hearing about the recruitment services that we provide across markets similar to this, feel free to reach out to myself or one of my colleagues on LinkedIn through evolution. Uh, but other than that, we will catch you in the next episode.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.