
Cyber Compliance & Beyond · 2026-06-16 · 22 min
Key moments - from our scoring
Substance score
38 / 100
Five dimensions, 20 points each
SACUVI's Tom Evge examines how AI-driven data intelligence transforms CUI discovery and protection in CMMC environments. Organizations today struggle with fundamental visibility - most can't answer three critical questions: what CUI they possess, where it resides, and who accesses it. SACUVI's platform uses machine learning to identify sensitive data at scale by learning an organization's unique CUI definitions during onboarding, then classifies information across sprawling SaaS platforms, email, shared drives, and cloud storage. The real value emerges when this intelligence feeds downstream enforcement tools like DLP, Microsoft Purview, firewalls, and endpoint security, transforming them from noisy, poorly-tuned alert factories into precise, context-aware controls. Evge also addresses the emerging frontier of AI guardrails - preventing sensitive data from reaching generative AI engines, blocking unauthorized outputs, and carefully controlling what permissions autonomous bots receive, since they learn new skills and can traverse environments unpredictably. The conversation explores how accurate data scoping often shrinks CMMC enclave requirements dramatically, shifting decisions from "easier" enterprise-wide approaches to intelligence-driven, right-sized implementations that reduce complexity and cost.
By discovering what CUI an organization has, where it lives, and who accesses it - three questions most organizations can't answer. The AI learns the customer's unique CUI definition during onboarding and classifies data across email, SaaS apps, shared drives, and cloud storage, often revealing that only a small percentage of users (e.g., 100 out of 1,000 employees) actually work with sensitive data, enabling right-sized enclave decisions.
They generate excessive noise and poorly-tuned alerts that create alert fatigue, causing security teams to ignore them. The solution is feeding these enforcement tools higher-quality data intelligence so they become precise, context-aware controls rather than background noise.
It sits between users and AI engines like ChatGPT and Perplexity, blocking CUI and confidential information from being sent to external AI services. It also examines what data is provided to AI engines before ingestion and prevents those tools from outputting information they shouldn't reveal.
Organizations must be very intentional about permissions because bots learn new skills, update themselves, and can traverse environments unpredictably to fulfill broad functions. They can attempt to access systems through multiple directions, grab packages, update their behavior, and potentially leak any information given to them, so minimal necessary access is critical.
When data is properly classified and labeled with security policies, those labels travel with files as they move through the organization. This enables precise enforcement at the right boundaries and helps organizations define exactly where their enclave bubble needs to end, often dramatically shrinking the number of users and systems requiring high-security baselines.
Our reviewer’s read on each dimension, with quotes from the episode.
There are a handful of non-trivial observations - using data classification to shrink enclave scope, AI bots autonomously acquiring packages to learn new skills - but these are buried under significant host paraphrasing, platitudes, and vendor positioning. The episode moves slowly and repeats itself frequently.
once you answer those three main questions, you're well on your way to, you know, probably further than most organizations are today
they're going to try in multiple different directions to achieve that task with the access that you provide them
The point about AI agents autonomously grabbing internet packages to learn new skills is a genuinely underappreciated risk angle, but the rest of the episode rehashes standard vendor talking points: AI reduces DLP noise, classify before you enforce, know your data. No contrarian or first-principles arguments are made.
it'll go out of the internet and try to grab packages that will allow him to change their voice or learn new skills
I think this kind of like the last frontier, if you will, in a cybersecurity program is that intelligence to feed all of your security tools
Tom Evge holds a 'field CISO' title but his bio foregrounds sales enablement and mindshare activities, and the conversation reads like a vendor demo walkthrough rather than hard-won practitioner experience. He speaks knowledgeably about his own product but offers little independent, battle-tested perspective.
developing and driving sales enablement strategies, improving customer success and product adoption, and increasing mindshare and awareness
Tom Evge, field CISO of SACUVI, a leader in AI-driven data intelligence
The episode is almost entirely abstract. The only concrete figures are a rough '11 tools' estimate and a hypothetical 1,000-to-100-user scoping example. No customer names, no measured outcomes, no documented ROI, and even the product roadmap item (MCP server) is described as forthcoming rather than deployed.
at least 11 tools from firewall to intrusion detection to logging to encryption to endpoint security
in a thousand person organization, it's only like a hundred people within my organization that are actually accessing and interacting with and working with CUI
The host asks broadly framed, open-ended questions and consistently validates the guest's answers rather than probing deeper or challenging vendor claims. Follow-ups tend to restate what was just said rather than push for evidence or counter-examples, resulting in a comfortable promotional chat rather than a substantive interrogation.
So, yeah, talk a little bit about how maybe thinking more broadly than just data categorization or sensitive data, what other types of security benefits have you guys found?
I think that's great advice. I mean, I think that's in line with advice we give even for non-AI solutions
Computed from the transcript - who did the talking, and the words that came up most.
AI is rapidly transforming how organizations identify, classify, and protect sensitive information, yet many are only beginning to understand what this means for CMMC, CUI, and their broader security programs. In this episode, we explain why understanding your data has become one of the most urgent challenges in modern cybersecurity and how AI‑driven tools can quickly reveal hidden sensitive information. This new intelligence helps organizations accurately define the true boundaries of their enclave and reduce assumed scope by pinpointing who actually interacts with CUI. We also cover how: Traditional tools like DLP, firewalls, and endpoint security often produce overwhelming, low‑value alerts. Adding richer, context‑aware data intelligence makes policies more precise and reduces alert fatigue. Enhanced intelligence strengthens enforcement as sensitive files move across the environment. It also improves operations by identifying stale or redundant data and reducing risks as organizations adopt AI. Strong AI guardrails are essential, including managing model permissions, inputs, outputs, and using MCP server‑based controls to block improper data flow.
Transcribed and scored by The B2B Podcast Index.
AI isn't just another shiny new tool. It's rapidly becoming the lens through which organizations see, understand, and control their most sensitive data. As CUI moves across cloud apps, email threads, shared drives, and even AI engines themselves, the stakes get higher and the blind spots get bigger. In this episode, we explore how AI-powered data intelligence is reshaping how organizations discover, label, and protect critical information, and why guardrails, context, and visibility matter more than ever when your security program is drowning in noise.
Welcome to the Cyber Compliance and Beyond podcast, a Kratos podcast that brings clarity to compliance, helping you leverage compliance as a tool to drive your business's ability to compete in any market. I'm your host, Cole French. Kratos is a leading cybersecurity compliance advisory and assessment organization, providing services to both government and commercial clients across varying sectors, including defense, space, satellite, financial services, and healthcare. Now let's get to today's episode and help you move cybersecurity forward.
In today's episode, recorded live at KuiCon in Orlando, we explore what CMMC really means in an AI-driven world where sensitive information is scattered across cloud apps, email threads, shared drives, and even generative AI engines themselves. We start with one of the core challenges organizations face, simply knowing what CUI they have, where it lives, and who has access to it. A problem made exponentially harder by sprawling SaaS platforms, constant file sharing, and workers pasting information into AI tools without understanding the risk.
From there, we dig into how AI-powered data intelligence can finally bring clarity to environments buried in unclassified, mislabeled, or unknown sensitive data. We look at how modern models can learn an organization's unique definition of CUI, classify data at scale, and reveal the true boundary of a compliant enclave, often shrinking scope dramatically for large enterprises that previously assumed everyone needed access. We then examine why enforcement tools like DLP, firewalls, and endpoint security have historically failed, noisy alerts, poorly tuned policies, and a lack of context.
And we explore how feeding these tools higher quality intelligence can transform them from background noise into actionable, precise controls that actually prevent risky behavior. Finally, we look ahead to the new frontier of AI guardrails, from preventing sensitive data from being fed into AI engines, to controlling what AI tools can output back to users, to understanding how autonomous bots might learn new skills or traverse your environment if given too much access. We break down why organizations must be intentional with permissions, model inputs, and AI behavior before these tools become deeply embedded in operational workflows.
Joining us for today's conversation is Tom Evge, field CISO of SACUVI, a leader in AI-driven data intelligence. Tom specializes in helping organizations discover, classify, and protect their most sensitive information, from CUI to personal data, and brings deep expertise in how AI can enhance both security and operational efficiency at scale. Tom is a seasoned cybersecurity professional with over 20 years of experience in leadership, project management, negotiation, and conflict resolution, developing and driving sales enablement strategies, improving customer success and product adoption, and increasing mindshare and awareness.
We hope you enjoy this episode. Tom, just want to thank you for stopping by our booth here at KooEcon in Orlando to chat about AI. It's kind of what we're going to get into today. We talked a little bit before hitting record on this.
You guys, you work with AI solutions and AI has become this thing that's all around us that we talk about all the time and really talking about it. There's a lot of scary things with it. But hopefully through our conversation today, we'll dive into some of the really good use cases and good uses for it. So maybe you could just talk to us, get us started on how do you use AI?
What are you seeing out there? What are the benefits you're seeing with AI? Yeah, absolutely. Thanks for having me and glad to be here.
So Sikoovi is a data intelligence platform that is driven by AI engines for data discovery and classifications. It has a couple of different use cases, obviously, around security, but also operational efficiency. And really the main driver behind the platform is being able to identify and discover very sensitive information, and CUI obviously being one of them, and getting that contextual awareness around data to be able to identify what's really sensitive and provide the security team in building an appropriate policies around that data.
So when folks are working with your particular solution, you mentioned the context, but how does the AI model or tools that you're working with, how does that tool get the context to know, okay, this is sensitive information, this isn't, things like that? So we go through an onboarding phase with the customer where they provide some contextual data around what CUI data means to them. It obviously comes in very many different flavors and in colors. And so training that model in identifying what's CUI to the customer, they're able to very quickly learn what type of data is around their environment and be able to identify it and mark it.
So definitely some what we call marks, data marks from the customer. to be able to build those policies, but the AI models are, like I mentioned, able to quickly identify and learn what that type of data looks like Now do you guys use I know there like standard CUI markings for instance or even if you doing classifications for other types of information or sensitive data does that tool come sort of baked in with some of those known across the board markings for different sensitive types of data?
Or is it required that I go feed it anything and everything that it would then go and identify and use context? Yeah, so it's actually a combination of both, right? We have some basic policies that we start off with, but the context from the customer is extremely important because that's how the models learn. But to your point, there are some baseline policies that we provide, and then the additional context come from the customers.
So it is just sensitive data. Are you guys also looking at other elements potentially? Because, for instance, CMMC, right? So, yeah, I need to know where my CUI is, who's accessing it, all that kind of stuff, which I assume it can identify that sensitive information.
Does it also have the capability to say, hey, here's all the CUI I found, and here's all the people that have access to it or interact with it, things like that? Is that something that it's able to do as well? Absolutely. I think the first three questions when you're thinking about CMMC is understanding what type of data you do classify as CUI, where that data resides, and eventually who's got access to it so you can build that policy around it.
And I think once you answer those three main questions, you're well on your way to, you know, probably further than most organizations are today if you're able to answer those questions. But, yeah, understanding where that data is, who's got access to it, and then ingesting or taking that data and pushing it down to your enforcement points. So your DLPs, your firewalls, your intrusion detection, your endpoint security. a lot of organizations or Microsoft shops or the US Purview to do the labeling and then the enforcement of that data.
And it's interesting. I've had this conversation with some other folks around this decision about what CMMC I can do for this whole enterprise or my whole organization, or I can build an enclave where I store everything. And a lot of times it's a thought or from what's the easiest solution? What's the easiest thing for the business, et cetera.
And it's much more difficult to say, well, maybe I need to make that decision based on what I actually have. But I think a problem a lot of organizations face, especially large organizations, is I don't know how much I have. I'm not exactly sure where it is, and I'm not exactly sure who accesses it. So they might come to it and say, well, I have a thousand users, so let's just do an enclave.
We'll put everybody in there, or just because we have that many, we'll do the enterprise and apply this security baseline across the entire enterprise. But you might come in with a tool like yours and be able to say, okay, well, we'll actually learn and find out what you guys have, who's accessing what. And then you might find, oh, well, it's only like in a thousand person organization, it's only like a hundred people within my organization that are actually accessing and interacting with and working with CUI.
And you can take that as a scoping point and say, all right, maybe I'll build an enclave for those hundred users. And it can help you make a really well informed decision, I think. Absolutely. It will help you define the scope and it will help you define the boundaries of your enclave.
So where does that bubble need to end for that data? And also what data comes in and out of it? I mean, think of the sheer volume of digital content that we're sharing on a daily basis and how it traverses through your typical organization, the emails that we're sending out, the Slack messages, all of the SaaS applications that we're using, files being shared, content being downloaded, input being uploaded from or input from users across different AI engines and search engines.
So just an incredible amount of data. And you are required to maintain some of that data or keep track of it. And so when you're able to identify, not only are you able to classify, but you're also able to tag it. And then those policies are moving around with that data.
Those files that you tagged and have labels on them now are when they move around your environment, your organization, those policies, the labels come with it. And so you're able to really follow that data through and understand how it traverses to your enclave. That's actually what I was going to ask next was sort of an interactive in nature. So it is something that because I think there's a user, let's say I'm going to move this file.
Well, there may be ramifications to moving that file that I'm just not aware of, or nothing gets brought to my attention. But there is an interactive component where it's like, hey, I'm a user, and I move this file, or I send this file. And before doing it, it provides context that, hey, if you do this, there's potential here. There's ramifications of doing something like that.
Yeah, absolutely. That's how the enforcement points within your environment, that's when they kick in. So when you have those labels and that data defined, so when somebody tries to access that data, your DLP, your purview, your endpoint will block it. Or when they're trying to download it or they're trying to send it out, those enforcement points are going to be the ones that are going to take action and stop that action from happening.
The value that we bring in is that intelligence to those tools where they've basically failed or they're either failing or they've just been so noisy over the past decade or so where they just creating just a lot of different alerts and they not functioning properly And we have this alert fatigue and being able to reduce that noise making your policies shorter and more concise. That's really the value that we bring. Absolutely. Like you mentioned and touched on, the amount of digital information that's out there is astounding in many organizations.
So the ability to have sort of an interactive guide, if you will, that cuts down the doors and really gives you actual decision points. Because, yeah, I mean, I've worked in an operations capacity in the past, and at a certain point, you get a certain number of alerts or you get a certain amount of information and you're like, no, you just don't even look at it anymore because it's just noise. But if it's something that's actually tuned and gives you proper context and you see that it actually helps you make actual decisions and improve your enforcement mechanisms.
Yeah, I think that's kind of what the security programs have been missing. And this kind of like the last frontier, if you will, in a cybersecurity program is that intelligence to feed all of your security tools. I mean, in a typical organization with 500 to 1,000 users, the number of security tools, especially if they need to adhere to any sort of compliance framework like CMMC or PCI or ITA or HIPAA, the number of security tools that they have is just astounding. You know, at least 11 tools from firewall to intrusion detection to logging to encryption to endpoint security.
And you're ingesting all that information. But at the same time, you need to tune them. You need to make sure that the policies are in place that you're firing off on when there are actual events happening. And so being able to, at the end of the day, what we're trying to do, I mean, everyone here is we're trying to protect our data, whether it's personal information, whether it's CUI, whether it's customer information, we're all trying to achieve the same goal here.
And if you know where that data is and you know who's got access to it, all of your tools can be smarter and less noisy. Absolutely. And so I'm curious. So we're talking mostly about CUI to start here, but I'm sure you guys are looking at, you mentioned all sorts of security stacks, all the different tools.
So I'm assuming that your capability can actually look at from sort of a layer zero all the way up and can evaluate the context in between those different solutions and kind of highlight and illuminate, hey, this could be a problem area or things like that. So, yeah, talk a little bit about how maybe thinking more broadly than just data categorization or sensitive data, what other types of security benefits have you guys found? So from a security perspective, I mean, again, from looking at sensitive information and what's been classified or unclassified, there's gaping holes within some of those environments.
But, you know, let me take you to maybe a different lens here looking at through operational efficiency where you have some organizations have petabytes of data and they're delivering content to their customers, whether it's a streaming platform, whether it's, you know, a company like Netflix that's delivering, again, a streaming platform that is either delivering videos or music or they're delivering, you know, other services. they have petabytes of data. And part of using AI, generative AI, is that it continuously tries to think what your next question will be, so we can answer it.
And so it brings data to the front, you know, closer to the user, so it's going to be able to answer it more proficiently. And so if we can identify the data, stale data, if we can identify fresh data, data that needs to be moved into archive, there is massive ROI in just from a storage perspective on how the content delivery is created. So we've seen a lot of requests from data security posture management being used in a couple of different ways, and operational proficiency has definitely been one of them.
I can really see a lot of value in that, because I think that's a great description that you just gave, that it's always thinking of what's the next question to ask. And I like that because I think when it comes to how we operate, sometimes I think we have a limited capacity to ask the right question, which asking the right question is what gets us to what the next thing we're going to do is and what that thing is and how it's important. And in a lot of cases, I think there might be even questions we don't know how to ask.
But you mentioned the stale data, the fresh data, all that kind of stuff. I think as humans, there's an element we don't even really think about with that kind of thing. It is an important thing to consider and to think about. So even just having that presence in the environment that can bring those questions to the surface so that you can actually wrestle with, should I do this?
Should I do that? What do I do about this particular thing that I know maybe I didn't know before? Yeah, 100%. So we're definitely seeing a lot of that.
And again, I think we started with security, but our data is everywhere, especially in the dawn of AI, where organizations are trying to race to that proficiency in leveraging AI. They're just feeding it so much data, so much information, because training happens through data ingestion. And so what you're providing, what data you're providing, the AI engines, do you know if it has any sort of personal information, any customer information, any identifiable information, policies that you might be violating?
And also from a guardrail perspective what is your AI engines or bots what information are they providing Are they needing any guardrails to prevent them from providing API keys of doing operational functionality in the backend that might jeopardize keys or passwords or users? So guardrails. So how are you guys approaching the problem of guardrails? Because I kind of touched on it at the beginning, that AI is kind of the scary thing, I guess, in some respects.
Guardrails, I think, are important. So what are you guys seeing as far as guardrails, or what kind of guardrails do you guys put in place when it comes to AI, things like that? So we actually have an MCP server that we're going to be launching, and that kind of sits between the user and your open AI, Chai GPT and perplexity, and looks at the data you're sending. And based on policies, we'll be able to block anything that's CY-related or confidential information.
So we're certainly going towards that direction, but also at the back end, looking at the data, like I mentioned, that is being provided to the AI engines and blocking it before it starts sending all that information. So you're essentially configuring it so that what the user is putting into it doesn't violate any sort of policy. And I'm assuming policies and things like that are things that you have to configure within the solution first. or is it similar to what you were talking about with sensitive information and the context and stuff like that?
Is it able to sort of learn over time from a policy standpoint or provide nuance and things like that as well? Yeah, it's able to learn over time, and it's the feedback that it's getting from the AI engine that is what we're kind of blocking. And so what the user is going to cut and paste from their file system or from their desktop, they may have CUI. We're specifically not there just yet because that's going to be more of a browser functionality.
But what information is being sent or being provided by the AI engine, that is where we block. Making sure the AI engine isn't providing information that it shouldn't be providing? Exactly. Exactly.
Yeah. Tom, again, I really appreciate you stopping by to chat with us. One final question here as we wrap up. I just want to know what you think, and we've kind of talked about some of this.
But when it comes to AI and guardrails and things like that, what do you think is the most important thing organizations need to keep in mind as they're wrestling with how do I use AI in my environment and how do I do it in a way that leverages all the benefits but also prevents some of the scary stuff that's out there? Yeah, that's a great question that I think a lot of organizations are struggling with today. So there are two parts to that, right? The first one is, what information are we providing AI engines like OpenAI or Anthropic or Proplexity?
And be mindful of what information we're sharing. with cloud bots and mcp servers and the llms that we're building internally we really need to be mindful of the access that we're providing them and what can they do because they are they're they're learning they're not just acting on a specific command that you give them they are trying to fulfill a more broad function and with that what i mean by that is that when you ask them to deploy a VM, for example, within your environment, or you're asking them to run a task, they're going to try in multiple different directions to achieve that task with the access that you provide them.
So for example, if you want them to speak in a certain way, your cloud bot, for example, or your chat bot, you want it to speak in a certain way, it'll go out of the internet and try to grab packages that will allow him to change their voice or learn new skills. The bots today learn new skills. They're able to update themselves so they can gain access to different environments, so they can understand how to access Slack, for example, or how to run through your shopping list or book travel for you.
So I think being very mindful of the access that you give them and how it's being used is super critical. And just be mindful that everything that you are giving your bot, it potentially can be leaked. So I would just caution organizations to be very mindful of that. I think that's great advice.
I mean, I think that's in line with advice we give even for non-AI solutions is to really think through this and plan. plan it. Make sure you're talking to the right people and not just going in alone. You definitely want to make sure you're working with those who can make you help you make these decisions and make them in a wise way.
So again, Tom, appreciate you stopping by. Absolutely. Thank you. I really enjoyed this conversation.
I think AI being such a pertinent, prominent topic out there, I think this will really be beneficial to our listeners. So again, I appreciate it. Yes. Thank you so much for having me and looking forward to hearing more of your podcast and some of the content you have to share.
Thank you for joining us on the Cyber Compliance and Beyond podcast. We want to hear from you. What unanswered questions would you like us to tackle? Is there a topic you'd like us to discuss?
Or you just have some feedback for us? Let us know on LinkedIn and Twitter at Kratos Defense or by email at ccveyond at kratosdefense.com. We hope you'll join us again for our next episode.
And until then, keep building security into the fabric of what you do.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.