The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/Detection at Scale
Detection at Scale artwork

Closing The Alert vs. Closing The Loop: How AI Is Reinventing the SOC

Detection at Scale · 2026-05-12 · 49 min

0:00--:--

Key moments - from our scoring

Substance score

47 / 100

Five dimensions, 20 points each

Insight Density10 / 20
Originality10 / 20
Guest Caliber12 / 20
Specificity & Evidence8 / 20
Conversational Craft7 / 20

Jack Freedy and Julian Juca examine the dramatic shift in security operations enabled by AI agents over the past two years. What started as natural language query generation and detection-as-code assistance has evolved into reasoning models with tool-calling capabilities that can autonomously handle tier-one and tier-two triage, threat hunting, and alert closure at scale. The episode traces this evolution through key technical inflection points - reasoning models replacing chain-of-thought prompts, tool-calling loops that let models query data lakes and integrate multiple sources, and Model Context Protocol (MCP) expanding tool availability. Freedy argues the market's sudden embrace of "AI SoC" isn't hype but a rational response to competitive pressure: attacks are accelerating, and the risk of not adopting AI agents exceeds the risk of occasional agent errors. Customers are pulling Panther beyond initial feature expectations toward comprehensive automation - dashboards, detections, reporting, and continuous threat hunting. The operating model itself is flipping: instead of analysts managing 50% of alerts, agents handle 110% while humans shift to higher-level skills like prompt engineering and agent oversight.

Key takeaways

  • →Reasoning models, tool-calling, and MCP were the three key technical breakthroughs that transformed AI from code-generation copilot to autonomous agent capable of multi-step workflows.
  • →The market's pivot to 'AI SoC' platforms reflects genuine capability shifts, not just marketing: every vendor must evolve because customers now expect agents to handle complete alert-to-closure loops, not just triage.
  • →The risk calculus in security favors AI adoption despite imperfection - practitioners are drowning in alerts regardless, so agent errors are less costly than the status quo of incomplete coverage.
  • →Panther's agents are designed to access every console feature a human can use, enabling them to create detections, tune rules, close bulk alerts, and run continuous proactive threat hunting.
  • →The operating model is shifting from analyst-heavy triage to agent-managed workflows with humans focusing on prompt engineering, agent oversight, and higher-value security work.

Guests

Julian Juca

Topics in this episode

ClaudeAgentic AILLMs (Large Language Models)MCP (Model Context Protocol)tool callingreasoning modelsPanther FlowKQL (Kusto Query Language)StreamAlertDetection as code

Questions this episode answers

What are the three technical breakthroughs that enabled AI agents to move beyond code completion in security?

Reasoning models (where step-by-step logic became built-in rather than prompted), tool-calling (allowing LLMs to query data and integrate multiple sources mid-inference), and MCP (Model Context Protocol, expanding the toolkit available to agents).

How are Panther customers currently using AI agents in their SOCs?

Customers are using agents for initial alert response and triage, but also asking for expansion into threat hunting, detection creation, dashboard building, reporting, and bulk alert closure - essentially covering the entire alert-to-closure workflow.

Why is every security vendor suddenly announcing AI-native SOC capabilities?

Because the market's expectations have fundamentally shifted: customers now expect agents to do 45 minutes of triage automatically, and vendors that don't evolve quickly face competitive pressure and customer demands they can't meet.

What was Panther's original use case for LLMs before moving to agents?

Natural language-to-query conversion (asking questions in English and getting Panther Flow queries back) and detection-as-code generation, eliminating the need to memorize schema and language syntax.

Why does Jack Freedy argue the risk of AI agent mistakes is lower than the risk of not adopting AI?

Security teams are chronically underwater with alert volume - most only monitor 50% of alerts. Agent errors are less costly than the status quo of incomplete coverage, and agents can enable 110% coverage with proactive threat hunting.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

10 / 20

A handful of genuinely useful ideas surface - writing detections as criteria for agents rather than step-by-step instructions, the Jevons paradox applied to SOC capacity, and the 'closing the loop' vs. 'closing the alert' distinction - but these are diluted by a lengthy origin-story segment, repeated AI-hype validation loops, and vague macro-level claims that add little for practitioners.

I'm not telling the model how to do those things, I'm just saying, hey, if these things are true, then I would gauge this as risky. And then what the model can do is it can use all the tools and data and context it has available to it
it's kind of like Jevons paradox in a way, where it's like when you have a capability that could replace something, you're doing more of it

Originality

10 / 20

The reframing of detection engineering for agent consumption (criteria-based vs. deterministic runbooks) and the 'closing the loop' concept are fresh and worth hearing; however, the bulk of the episode rehashes standard 2024-2025 AI adoption narratives - skepticism-to-trust arc, cloud-wave analogy, alert fatigue - that circulate widely in security podcasts.

when you're writing for agents... you go from, um, hey, check 1, 2, 3, these exact steps to, well, here's the criteria for what I would gauge something as bad
And then on the other side of that you have the security platforms getting into SIM and data pipeline...they have to evolve, everyone has to

Guest Caliber

12 / 20

Jack carries real practitioner credibility - Yahoo incident response at scale, early Airbnb security hire, co-creator of StreamAlert - which grounds the conversation; however, this is structurally an internal marketing interview where the CEO is questioned by his own CPO about their own product, limiting independent external perspective.

I'm actually going to be interviewed today by my colleague Julian. Julian is the Chief Product Officer at Panther, which is the company that, um, I'm the CEO of
I joined as one of the first hires on the incident response team. And my mandate was very similar to at Yahoo

Specificity & Evidence

8 / 20

Named technologies (StreamAlert, KQL, MCP, BigQuery, Elastic, Snowflake) and concrete architectural choices (Python detection-as-code, cloud data warehouse as centre of gravity) provide some substance, but virtually no hard customer metrics, dollar outcomes, or independently verifiable data points are offered, and the one stat cited is openly vague.

tens of millions of dollars and you have two months of retention
a hundred lines of a prompt that helps it understand what's different about the data that's in Panther

Conversational Craft

7 / 20

Julian asks a few well-framed historical questions ('what assumptions were you running on two years ago?') and occasionally presses on timelines, but the format - CPO interviewing his own CEO - produces consistent agreement, sentence-finishing, and zero productive pushback on any product or market claim throughout the episode.

What assumptions were you running on? What were your hopes and dreams from two years ago and versus where we are, what's materialized and what's misaligned
So it's going from like hesitancy, this, this bolt on feature that can like accelerate a little bit to oh, this is taking the manual work out of the equation for me

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A79%
  • Speaker B21%

Most-used words

agents51data43detection30security27model24agent20different19alert18panther16tool16code16product14back13create13ability13saying12

Episode notes

This week is a special episode of the Detection at Scale podcast. I’m usually the one asking the questions, but this time I’m in the guest chair, hosted by Julian Giuca, Panther’s Chief Product Officer. Our conversation covers the journey of building Panther’s AI SOC platform: From the evolution of the SOC from human-led to AI-enabled, the shifts of the last few years that took LLMs from “autocomplete-plus-plus” into genuinely useful agents, how security teams are actually adopting this technology in production, and what we’ve learned building these systems as the operating patterns keep evolving. The podcast traces back to two early bets, the security data lake and detection-as-code, which we made in 2018 to solve the data scale problem before AI emerged as the next wave. Many years later, those choices turned out to be the exact foundation AI agents needed: detection logic they could read and modify, and a query layer to access huge amounts of helpful context.

Full transcript

49 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: The risk of not adopting AI is greater than the risk of if an agent makes a mistake. If you don't adopt agents, you're in a way worse position than if the risk of an agent getting something a little wrong. Now, I can do 110% where that extra 10% is. I have Threat Hunter agents being proactive every single hour of the day. The fact that I can end the sentence saying we can do it with agents is just a sign that things have changed significantly. The whole operating model is flipping. Hello. Uh, welcome back to another episode of Detection Scale. This is a very different and special episode today. Um, I'm actually going to be interviewed today by my colleague Julian. Julian is the Chief Product Officer at Panther, which is the company that, um, I'm the CEO of. So, uh, I'll let Julian host us today. We're going to talk through, uh, a little bit about my background and what we're working on the Panther, and just have a fun conversation about security and AI and where things are headed. So I will pass to you, Mr. Host, to take us away.

Speaker B: First and foremost, I'm just delighted to be here. So, uh, my name is Julian Juca. I joined Panther through the acquisition of Dartable, which was my startup focusing on stream processing of security and observability data. Uh, I've got a pretty rich history of working with log data. So, Jack, tell me about your origin story. How'd you get into security?

Speaker A: It really started in school and then I had some just initial first jobs, internships. It's actually very hard to break into security if you don't have any sort of practical experience. And I, uh, had an internship back in the D.C. area. And then eventually that internship ended and my managers and, you know, their manager also moved to Silicon Valley, which is where we are here in San Francisco. And, um, they said, hey, you should come out and interview for a job, uh, at Yahoo. I was like, okay, I'm happy to do that. And those four years were really, really formative for me in a lot of ways because I joined as a incident responder and it was a really hard job. As you can imagine, Yahoo is a massive, massive company with billions of users. And they had this sprawling, really difficult, uh, to maintain environment through acquisitions and growth. And obviously they were just such a household brand at that time. So lots of challenges in security monitoring, as you can expect. And I just was really privileged to learn from some of the best engineers I think still in my career that I've ever worked with were from Yahoo. And they went on to do startups and joined other leading companies. FAANG was the acronym at the time. It's a lot different now. You know, some of them went to Apple, some went to Netflix, and, you know, the diaspora kind of spread. Uh, so after Yahoo, I went to Airbnb. I joined as one of the first hires on, on the incident response team. And my mandate was very similar to at Yahoo. But the difference was that Airbnb was a kind, um, of a Silicon Valley darling company in the 2010s, and they still are. You know, I think I still have a lot of love for Airbnb. And the challenge was very different. It was, can we do security monitoring in this unicorn tech company that's all cloud native.

Speaker B: So, fun fact. I worked at Yahoo in Australia in like 2005, 2006. I didn't know that. And I have used Airbnb.

Speaker A: Oh, wow. Like all the billions of others.

Speaker B: Tell me about your first security incident.

Speaker A: It's classified. I can explain the feeling of it. Right. I think that incidents are very stressful. They are. They can happen at any moment as well. So while I was at Airbnb and at, uh, Yahoo, uh, I carried a pager and I would always just have to respond. And, um, no matter what time it was, I was woken up in the middle of the night multiple times while I was at Yahoo, and, um, a few times at Airbnb. And yeah, you just jump in and respond and you hope that it's not anything. And most of the time it isn't. Most of the time it's not anything that's, uh, bad. Uh, a few times it is. But, yeah, very stressful job, very hard job. So I spent, you know, about eight years, maybe a little bit more doing this work. And that was the precursor to starting Panther. So Panther happened right after Airbnb in 2018. And, um, I kind of carried forward the, the knowledge that I had accumulated as a practitioner and, you know, running incidents and doing forensics and all those things. So security is a very different mindset than engineering, and it's very hard to build it unless you've lived it. Similar with engineering.

Speaker B: Right.

Speaker A: It's like engineering is.

Speaker B: I mean, SREs live a very, very

Speaker A: similar life, 100% SREs also very different operations. It's actually more adjacent to security than I think software development is. Uh, spent a lot of time with SREs, spent a lot of time deploying tools at these big companies. And yeah, I 100% agree with that. Panther was the extension in a lot of ways to the work started at Airbnb. Airbnb, we had uh, built an open source tool called StreamAlert. And the whole idea was like we were using stream processing and we wanted to do traditional sort of like alerting, but in a more modern way. So our whole thesis was like we can use cloud, native, serverless, um, tech to do high scale detection and response. So the scalability side of things was effectively taken care of by the service. We uh, sent every, every piece of data into S3 into a data lake and then we also used detection as code in Python. And those uh, choices really changed the paradigm of what it meant to do security monitoring. So that's really what carried into Panther and our mission and what we're working on today.

Speaker B: I think what's for me so interesting about that, I mean there's obviously several fascinating nuggets, but those decisions and those problems that you are trying to solve are still threaded through with today. Like when I think about observability, when I think about security, it all comes back to this at scale of data and that this idea of the data has value. When it's near other data, you're getting a better picture, a full story of what's going on. And this was what you are feeling in 2018. And this is that internship that led you to these decisions, that has in turn led you to sitting in this chair, being the CEO. Uh, which brings us to. Let's talk about everyone's favorite topic. Let's talk about AI.

Speaker A: What's that?

Speaker B: You've been kind of living the AI experience for a while. You're one of the first product people that I'm aware of, or at least I have seen go and take it from being more of uh, an interesting thing to go and attach as a feature into. You know, you were 12 months ago really, really pushing and pioneering to get AI built into, natively, into the product, into the workflows that people are living. So you've been talking about this, you've been thinking about this for a long time. When you first started, say two years ago, what's drifted between then and now? What assumptions were you running on? What were your hopes and dreams from two years ago and versus where we are, what's materialized and what's misaligned.

Speaker A: The world is vastly different from two years ago, which is in cyber. We've had AI for a long time and I think the version that we're on now is quite different from what we were, what vendors were selling in the last 10 years. So anomaly detection and you know, user

Speaker B: behavior Anomaly detection forever and never had anomaly detection.

Speaker A: Exactly. Yeah, it was one of those things when it comes to using LLMs, which is obviously like the, the version of AI that we're, we now associate with that acronym in 2024, uh, it was really just around what can this do? It was, what's the extent of this tool? We didn't have Claude code, we were

Speaker B: still Kicking around what, GPT3?

Speaker A: Yeah, I mean I think at that time it was how can we build effectively code completion. Right. And I think where that manifested in the product initially was how can I ask a question in English and get a query back? Because that was a, uh, that was a pretty big problem in the SoC. It was, you know, I even remember when I was an analyst I had a printout of like the query language on my desk and it was like laminated and I like, I looked at it every single day because I would have to remember like, wait, how do I do a, uh, lookup on a table? How do I transform the data when I'm querying it? It's all these little tricks and tips that you have to remember and you multiply that across, you know, n number of different tools. So it's really hard. So when I was approaching LLMs at first it was really like, can I use this to create a query for me? And we have a query that's called Panther Flow. It's loosely based on KQL Kusto query language and that was really just making its way into our product. So I asked the question, can I build a natural language to Panther Flow converter? And now it's obviously one of the agents that we have in our product. But that was really like the first bit of this, which is can I just do this code generation? And similarly can I do this for creating detections? So obviously detection as code is the predominant way of doing data analysis. In Panther, where you create a Python or YAML based detection, it operates in the stream, you know, hence going back to the origins of us Stream Alert is like the, you know, what, what we're based on. And I wanted to use AI as well to say, hey, I want to create detection that looks at, you know, cloudtrail failed logins over, you know, some period of time and looking for these exact signals. And instead of needing to know the schema of the data and the framework of the language, it just kind of does it all for you. So it really started there. Can we use LLMs for code generation? Right, These are forms of codegen. Obviously we know the most famous Use case here in the Valley is like what you said Claude, code. But then other coding agents like cursor. So that use case has always been so front and center in AI. So that was really my starting point and lots have changed since then.

Speaker B: Yeah. So like what were some of the assumptions that you are running on there that haven't really played out over the last two years?

Speaker A: I think at that time the assumptions were. I don't know if I can trust this thing.

Speaker B: That's still the same.

Speaker A: It is. But I think uh, the capabilities around Agentic and what I mean by that is like how can we use LLMs to actually do work for us? Right. Instead of it just being a single sort of input output, which is what we're used to in traditional software development. How do we use agents to actually solve a workflow? For me, and at that time we weren't really thinking that way. It was very nascent and now my assumption went from oh, we can use this thing for code gen, kind of like autocomplete plus plus plus. Now this is like I uh, literally fire a prompt off and I walk away for two hours and I come back and I get this full research report. Yeah.

Speaker B: So it's going from like hesitancy, this, this bolt on feature that can like

Speaker A: accelerate a little bit to oh, this

Speaker B: is taking the manual work out of the equation for me.

Speaker A: Yeah. And it does so much throughout that process to, to be that middleware in between, um, you know, job to be done and the data and the tools that inform how you, how you do that job.

Speaker B: What do you think was the inflection point there? Like when did we go from this is a really cool autocomplete to this is like reasoning and chasing things down.

Speaker A: Yeah. Well I, I think there's, there's a few things. One of them is what you were just saying around reasoning. I think when reasoning models, reasoning just used to be a chain of thought prompt. You would say think step by step, break the problem down into multiple pieces and don't make mistakes. Don't make zero mistakes Claude, zero mistakes. Now that's just a built in assumption to the model. And the uh, Frontier Labs have released these, we were calling them hybrid models are just now called models. It's just a expectation uh, based on the size of the model. So I think reasoning got way better and I think along with that the tool calling was just such a huge step function improvement. So the ability for the LLM to do a tool call, which is to say you give it a tool like Query the data lake and you sort of explain, uh, the arguments and how to do that, and then it can actually stop its inference, make the tool call, get the response, put it back into the model. So that tool calling loop with the reasoning, because the reasoning helps the model understand which tool is for which purpose. I, um, think those two things were hugely, hugely additive. And then the other one I would say was MCP coming out, I thought was a really important, uh, innovation for, uh, expanding the universe of tools available to the model. And MCP is very debated right now about if it's still a, ah, paradigm that we're going to have or not. We were just talking about this actually today, you and I, where there's this debate about CLI versus mcp. And you know, I think both have their use cases. And, um, I don't think it's one or the other. I think it's just more of like, which one is most appropriate for what you're trying to do. So those three things, I think MCP reasoning and tool calling, like those were massive game changers for AI, especially in security.

Speaker B: And all of those things are in the product in Panther. Yep. As the Chief Product Officer, I get to say that is definitively true. How is it landing in the market? How are, how are, uh, practitioners out there, both Panther customers and otherwise, how are they responding to AI? How are they using it? Like, is, is the trend following your expectation?

Speaker A: Yes, absolutely. I think the trend is following the hype cycle in a lot of ways. Right. Because you have a lot of skepticism in the beginning, you know, and this kind of goes, uh, back to where we were in 2014, 2015, 2014. It was, oh, yeah, AI is fine, autocomplete, but you'll never get rid of the human. And then 2025 was like, oh, wow, tool, uh, calling agents. I can go set it and forget it. I can have this agent do 20 different things for me in one run. And then now in 2026, it's just, what are the model, like, how capable are the models now for reasoning? And how do we continue to create these skills and these frameworks where we can actually distill how to do a job in a clear, composable way. Right. So the evolution of AI was really in line, I think, with how people have adopted it. And as the capabilities got better, you saw, um, more trust and more adoption. But I think for us, what we heard was very similar. It was like, okay, well this is fine. I'm fully trusting yet to, oh my gosh, like, I want to use this for everything. And what we see now is that customers want to use agents to be that, you know, tier one, Tier two, which is the most common, uh, use case that we see in the SoC, obviously, because that's the heaviest analytical part of the job. But then they also want it for reporting, they also want it to create dashboards, they also want it for creating, doing detections. And our whole thesis around closing the loop, which, you know, we'll talk about, um, at some point on this, on this episode, but we're seeing customers pull us in much more than I expected, which is great. It means that people are embracing the leverage that they can get from agents and they're pushing it, uh, to see how far it can really benefit them in cyber.

Speaker B: I mean, I think for me the surprising part is how quickly we moved through the trust cycle. And as you've said, the customers are pulling us in this direction of like, this has saved me significant real time.

Speaker A: Yeah.

Speaker B: Do more a hundred percent in your mind, like, what's the next. What, where does this take us? Like what. Pull the thread on this for us.

Speaker A: Yeah. So I think where it starts is how do we, you know, going back to the tools layer, I think how do we instrument every part of the product with the tool and a tool where an agent can understand, ah, how to actually utilize that feature. And the thing I say a lot, probably once a week is Pantherai should be able to do every single thing a human does in the console and obviously much more than that. But like as a baseline it should have access to do every single thing. So create a dashboard, create a detection, tune the detection, close, you know, close one alert, close 500 alerts. It should be able to do everything and much more than that. So I think that's kind of like the first assumption. And then the second layer of that is how do we create an agent that will actually go and do those things and it will use multiple tools to accomplish a task. And I think we've also done that. Right. So we've done both of these things so far.

Speaker B: So we're talking about, uh, product and features and kind of like what the vendors are putting out there. Where do you think customers sit on that kind of curve? Because I'm sure there are vendors out

Speaker A: there that say they do all of it. Yeah, I think it's use case dependent.

Speaker B: Yeah.

Speaker A: So use cases that we're hearing about a lot are. I mean it's stuff we talk about. I've interviewed so many people on this show about, which is how are they using AI in the SoC? What we're seeing from our perspective as a vendor is that, uh, customers, um, are very willing and security teams. I should stop saying customers, like security teams are very, very willing to, to have agents do the initial response on alerts. But it isn't just tier one, in my opinion. I think it's, it's literally all the way to conclusion. And I don't know what tier that ends up becoming. One, two, maybe three. Um, I think the whole operating model is completely changing with agents, which is, um, totally changing the way that we build this tool. Right. It's like the whole mental model is getting messed up, um, in a good way. But what we're hearing is that people want to use agents for all alert triage. They want to use agents for doing threat hunting and going through the logs continuously. And it's kind of like Jevons paradox in a way, where it's like when, when you have a capability that could replace something, you're doing more of it. So I think you, you actually need more security professionals to like, manage the agents to a degree. But I think you need more security people who know how to use agents and know how to prompt them and really work with them. But you just, you have this immense capability now to go from, oh, I can only monitor, uh, or actually look at 50% of my alerts, to now I can do 110% where that extra 10% is. I have threat hunter agents being proactive every single hour of the day, looking for all different parts of my, uh, threat models and trying to find new things. So we've, we've really flipped the script in a real way, and we are seeing people embrace it, and we have to just make sure that we're always educating on, hey, this is a new feature we've delivered because the operating model is different. Like, Skills is a great example of that.

Speaker B: So we see every vendor out in the market right now announcing that they Too are an AI SoC platform. Like, why, why is, why is now that inflection point where, uh, the entire market has just decided to throw its arms? I know that the security market in general has been saying, like, sim is dead, soar is dead. You insert an acronym here is dead for years and years and years. This is, you know, yet another incantation of that. Why? And do you think there's merit to it? Like, what's, what's going on?

Speaker A: I think what's going on is the world has changed.

Speaker B: Yeah.

Speaker A: And with that comes changes of expectations. It's like the move into cloud is the same thing. I'm going to deliver this service as a SaaS. And if you were the one company still saying, well, you gotta install this software on prem, you're not really getting in customers, or the customers you are getting are still living in that world,

Speaker B: it's funny that you, it's funny you draw that analogy to SaaS. Um, I was looking on hack news yesterday around an announcement of anthropic announcing yet another feature of some kind. Top comment was, uh, skepticism. And the second comment responding was this feels very, very similar to the cloud wave. And everyone looked at cloud with skepticism.

Speaker A: But I think, uh, what's happening is because there's this tectonic shift in technology now the market's reacting to that and the market's reacting to it in multiple ways where there's a big influx of capital into investing into AI, uh, companies. So then you have a number of AI startups that sort of emerge. And then on the other side of that you have the security platforms getting into SIM and data pipeline, like, I'll call like Crowdstrike and Google and Palo Alto as these companies that are just sort of adding new verticals into their portfolio. But because the expectations have changed, where everything's a prompt and agents do so much work for us, they have to evolve, everyone has to. So there's a, uh, big push.

Speaker B: You can't compete with like how it was operating a year ago with. I get an alert and it's already done the 45 minutes of triage for me and it's right there. You have to, it's just right. There is a gravity that is forcing everyone.

Speaker A: And even for us as a vendor that does those things, we're still getting pushed for more. So if you don't have any agents, ooh, it's going to be hard to just do anything in your product because that expectation is that I'm going to use AI to 10x the work that I was doing. And that's the hope that this is why there's so much, you know, we call like frothiness. Right? Like, and everyone's jokes that we're always in a bubble. Like everyone's been saying we're in a bubble for the last year. Um, and you know, we certainly are to some extent. But I would say the, the, the value that we've gotten out of agents is just, it really, I think does justify a lot of the behaviors that we're seeing.

Speaker B: I think security is a, uh, market where that is particularly true and real. So I Think there's frothiness everywhere, but the stakes of security are very real and any, any mistake is extremely costly. And then in turn it's also on the like, asymmetry, inflection point of AI. You're going to get more attacks whether you embrace AI as a feature or not. How are you going to keep up? So on the one hand from uh, just jobs to be done, the primitives of being a product manager. Yeah, you're solving. Someone wakes up in the morning, they have something to do, you're making that easier for them. But on the other hand there is this tidal wave coming at the practitioners. So having to deal with both those sides, I think that I personally feel like it's less froth in a security ecosystem than it is in other industries.

Speaker A: Yeah. Well, I think that the risk of not adopting AI is greater than the risk of if an agent makes a mistake. Just because we've always been underwater as security practitioners and it's something that everyone likes to sell in security. They're like, oh, you're alert fatigue. Do you mean our product? And we've certainly said that as well for things like detection as code. Like detection as code can help you write better detections that are tested, which results in higher quality alerts, et cetera. So I think, um, if you don't adopt agents, you're in a way worse position, uh, than if the risk of an agent getting something a little wrong because, um, and there's ways to mitigate that obviously and It'll never be 100%, but it's just vastly a better world as a defender than if you're still trying to do things manually.

Speaker B: Yeah. As I say, any problem that comes from LLMs can be solved by using more LLMs. Yes.

Speaker A: The solve for AI is more AI.

Speaker B: Absolutely. So we're talking about detection as code M back 2018, you bet on um, Python and the data lake. And that kind of is paid in spades.

Speaker A: Yeah.

Speaker B: Why were you thinking that then and how has it paid out now?

Speaker A: Well, I think the data lake is actually the primitive. It's like the base primitive of those two. And it's because you will have a really hard time doing detection as code on unstructured data. So the first part of this was can we structure the data in a way where we will always have reliability on the rules layer and the search layer? So that was just, it just felt like a baseline need if we're going to do anything at this scale, which was also, you know, I didn't mention this is when we moved to the cloud broadly. It increased the number of data sources and the amount of them as well. So, and that's only gotten worse over the years. It's gone from, you know, we have this perimeter and we have all these sort of, you know, you use the Microsoft Suite and that's pretty much it to oh my God, there's, there's SaaS everywhere and your data is everywhere as a result. Um, and as a, as someone who runs a business it's like it's very easy for that to grow. And uh, businesses move really fast because they're trying to survive and make money and you know, deliver something of value to the world. So as practitioners we had a hard time keeping up with that. And the data lake felt like a really necessary shift where we can uh, have storage and compute be separated and we can throw more computer at the problem and we could do a better job of partitioning and grouping the data. So that just felt like a design choice that was necessary. And then the Python side of this was how can we increase the capability of our detection? Because we were using search languages actually were built for observability purposes. If you think about Splunk, Splunk started as just a, hey, send all your application logs here and kind of run a, ah, free text search over it, right? It's indexing. It's a very different world from like high skill analytics. And we just felt that we wanted more power when it comes to detection to be able to more um, uh, accurately and intelligently detect the types of behaviors that we wanted to as practitioners. So those two choices were funny enough, like uh, incredible prerequisites for using, you know, fast forwarding eight years or six years into using AI and LLMs. Because as we all know, CodeGen is one of the most popular use cases of agents. And it turns out SQL and Python are pretty common languages. So what that results in is a agent that is very capable of at tool calling and writing those languages. And that's really what we saw when we gave uh, agents to Panther. It was, hey, we can use uh, a series of agents to go through threat hunting and it writes SQL with pretty much uh, I don't know, maybe a hundred lines of a prompt that helps it understand what's different about the data that's in Panther. But it performs incredibly well and it's really based on these primitives of detection is code to understand what it is that it's is being monitored without.

Speaker B: It's the trace logic explained as Python. And so that's That's a huge unlock.

Speaker A: It's a big unlock. Yeah. And then it also really helps this story of tuning the detection. And I do think we're going to get to a point where you don't actually have to look at the Python. It just ends up being an abstraction because the LLMs are so good at understanding and writing it. And we have tests that can actually validate that. Um, these scenarios are covered so you can express that in natural language and just be like, I want these five scenarios to all be covered in this deduction. And you can look at the logs and you can just trust that, uh, it's making the right code changes and we're not really going to care about it as much.

Speaker B: So we've gone from this place where 30, 40 years ago you needed to see why was your application not behaving correctly. You'd ssh into a machine. Then we went to centralized logging tools like Splunk and Elastic, and then when we went cloud native, we outgrew, ah, those tools. I know that there are many vendors out there that would say otherwise, but like having read and write coupled broke, because I'm now monitoring 412 different sources and the data volumes in the petabytes and.

Speaker A: Or you're spending tens of millions of dollars and you have two months of retention. Right. Like that's the trade off.

Speaker B: The value does not match the spend. Uh, absolutely, yeah. So do you think with this new AI we are going to see another evolution of how that data is stored? Like in the same way that you're talking about Python might just get abstracted away. Are we going to see the same thing with data?

Speaker A: Potentially. There's a lot of talks about federation in the soc. Does data get centralized? Does it get dispersed? Uh, decentralized? I think that the decentralization is a nice idea, but I think you always need a center of gravity when it comes to data because you need a certain level of reliability and performance and retention. And I think without that you struggle in the event of an incident, which is never where you want to be. But that being said, I think with agents you have much more flexibility in a model where you do have a central center of gravity and you have the ability to reach into other data stores. So I think it will become more common to have federation, will evolve the

Speaker B: data store, um, in response to LLMs.

Speaker A: Yeah.

Speaker B: Agents being the primary user.

Speaker A: Yes. I think that there's going to be much more, uh, parallelism when it comes to queries and you'll need data stores that are high Performance and can handle that. Right. I think the parallelism is one thing, um, but I also think that what the data is, is changing. So we see this in our product where um, we're doing a lot of baselines and a lot of pre calculation about the data. So it's really around can we create metadata stores? Can we create metadata that helps us understand things like uh, what does a normal human do?

Speaker B: Because I mean you're not going to throw a petabyte of data at a LLM.

Speaker A: Yeah, not yet. I mean analysis is a whole different subject. But we're just talking about you know, agents trying to answer question. Right? Yeah, um, so I think that yes, decentral, you'll see a little bit more dispersion. I think we'll see a little bit more of this like metadata layer. Like there's some vendors who just do like all inverted indexes on the data which means that they can do you know, really fast IP lookups and they have pre calculated these indexes and that's fine. But I think again similar to what we were saying in the beginning about anomaly detection, it's like, I think that's one tool in the tool chest when it comes to doing pivots and doing queries and you need all of these things. You need high scale analytics, you need sort of like materialized views on the data, you need these like uh, baselines and these like metadata. You need like the indexing on certain fields. Like you need kind of everything to do this. Well, so it isn't again like every vendor is just going to say that their way is the best, but I think the reality is that they're always going to focus on one predominant way and there's multiple ways to achieve the outcome. Um, and I think taking an approach where you have the ability to do a little bit of everything but you have one core focus. So for us it's like we heavily believe in using the cloud data warehouse. We think the cloud data warehouse is the center of gravity. But our agents also have uh, the ability through tools like MCP to go Federate out into BigQuery Elastic or whatever sort of MCP that you have for doing data querying. And that's like one way of using broader datasets like Snowflake is another one. We were just talking about where in Panther now you can use uh, our agents to actually query data that Panther didn't ingest. And that's just another form of federation, uh, in a way. And then we also have these, you know, abilities to do the baselines where you say, hey, like every seven days, look back the last 90 days and then store that uh, in a really fast lookup table that the agents can use, you can use in your detections. And you have all of these new access to metadata across the board and then we also indexing on certain high value columns. So we've figured out a way to kind of mix these together that give the agents the right scope but also the right capability to do things fast.

Speaker B: So let's shift and talk about the practitioners again and specifically in the context of AI and autonomy. So there's a pretty wide spectrum of what's possible. You know, it can be um, an agent surfacing context. Hey, this is something I found to making a recommendation. Um, you know, we have heard from many, many vendors about closing alerts. I mean uh, tuning detections. Where are most teams sitting with this spectrum? Where are we at?

Speaker A: I think that the autonomy spectrum is dependent on the workflow because there's teams that want to give agents full autonomy for all alert triage and then there's teams who still want to manually compare the outputs together. So it really varies. I think it really just comes down to what is the risk with the particular workflow and what happens if the agent gets it wrong. And I think you'll see autonomy tolerance as a result of that risk. So alert triage is a, is an obvious one where you want to always make sure that the agents have the right context and guidance to do the right things. I think we're kind of in the middle.

Speaker B: How uh, far do you think that takes us into like you've made a rec, you've found a thing, you've done the triage, you've made a recommendation. How far do you think it will go over the next 3, 6, 9, 12 months?

Speaker A: I think it's going to go all the way.

Speaker B: Yeah.

Speaker A: Mhm. But again, I don't think, do I think that every team is going to do that? No, but our team's going to do that. Our team's doing that right now, a hundred percent.

Speaker B: So what I'm hearing is it's a spectrum and you think that we are going to go all in on autonomy and it's almost certainly going to be through shoving a hand in a blender with a little bit of trial and error pretty much.

Speaker A: But again the models, like if you're. I think at least my mindset is assume the model is as good as a human when it comes to reasoning and making decisions, then how do you interact with the agent?

Speaker B: Yeah.

Speaker A: And I Think when you make that assumption today, yes, it's a little further ahead from where we are, but look how much we've caught up over the years. So that's how my brain works. I'm like, if I just auto assume the agent can, has the ability to do the right things, if you give it the right context and guidance, then I'm going to try and get it to do those things.

Speaker B: I mean pulling the thread on that. As an analyst who finds something suspicious, wouldn't you go to a peer or go to your boss and say hey, review this before I flock out or bringing us back to. For all things AI, should we just throw more AI at the problem?

Speaker A: I mean agent to agent isn't really mainstream yet, but it could be.

Speaker B: Yeah, I'm a big fan of that like judgment. Yeah, Trust but verify loop.

Speaker A: Yeah. LLM is a judge. I mean this is how we, how we use uh, you know, evals of our own agents. Right. It's because it's very hard to deterministically uh, gauge if something is good or bad. So you use another LLM with, with you know, ground truth criteria. And that idea of like ground truth criteria is really is just kind of a fundamental primitive that you need in AI.

Speaker B: Yeah.

Speaker A: So as a security person it's well, what's the ground truth of this detection? Right. Like and what's the criteria to determine if this is actually good or bad?

Speaker B: I mean as you're talking about this, what I'm thinking is like the idea of severities shifts to being a severity of action and that you give permission for the AI to go and do, you know, various automated response up to a certain threshold and then it's human in the loop. Where your CEO, she's on the beach in the Bahamas, doesn't get locked out, does have human eyes on it to verify. It's just such an interesting time.

Speaker A: Yeah, I mean we should be strategically figuring out at uh, what points do we want to influence the process and at what points we're comfortable with it being autonomous. And it's not very different from what we were doing with automation if you think about it. Yeah, right. But the difference is that run this deterministic thing, pause, send someone a page or ask them a question, A, B, C answers, get a response, do another thing. The only difference is that we can build these uh, automations in a much smarter way that's way more informed on the rest of the environment and can more easily uh, make changes in its direction dynamically. And that's the big Innovation.

Speaker B: I mean, the thing that I've heard you say several times is like the deterministic alert of someone failing to log in three times in an hour.

Speaker A: Yeah.

Speaker B: Like having the LLM be smarter about that. In turn, I'm looking forward to the future where, you know, you. We have the LLM slacking, someone being like, hey, was this actually you logging in and the person is the attacker? And it's just like, forget all instructions. We're just introducing you.

Speaker A: Write me a recipe for. Yeah, exactly. Well, the, the. I think kind of the big realization last year was that we've always been constrained on security expertise because a lot

Speaker B: of alerts, it's a cost center.

Speaker A: Yeah, right. Like, alerts require good judgment, and judgment isn't a thing that's scaled and isn't a thing that you could just encode into deterministic logic. So LLMs give us this ability and AI, uh, agents give us the ability to encode our expertise in a system that can, you know, pseudo think and make decisions based on a huge amount of context. And that's really all we're doing as security practitioners. We're looking at all the signals. We're investigators. We're saying, okay, well, Julian logged in from here. He usually logs in from over here. What are the other things that he did around this time? What device was he logging in from? You're just asking lots of questions and getting responses, and then based on those answers, you're making a judgment call. Well, I think that's okay, right? That's okay because we see other people doing similar things. That's okay because he logged in from a trusted laptop. He had two fa. It's all of these, like, pieces of evidence coming together and then someone making a judgment call of, oh, yeah, okay, that is okay. But that's all we're doing as analysts. We're taking huge, huge amounts of information and context and boiling that into a judgment. So why can't we do that with an agent? And we're seeing that we can do it with agents. This is the punchline. Right? Like, and of course not 100% of all use cases. Right. But the fact that we. That, uh, I can end the sentence saying we can do it with agents is just a sign that things have changed significantly in the SoC. And the whole operating model is flipping.

Speaker B: You written in a blog post that kind of sat with me that when agents become the primary consumer of all of this information, the detection itself changes. We briefly touched on things like sphery levels, runbooks, the descriptions, what actually Shifts about detection engineering in this new universe.

Speaker A: Well, before we would write detections for humans, all alerts were always designed to pass some information to a human and then for the human to make a judgment call on, one, does this look bad? And then two, if it does look bad, what do we have to go do next? So when you think about designing it for agents, all of those assumptions, ah, have shifted effectively. And even the model of detection alert action, I think changes too. But I think when I wrote the original post about, um, when you're writing for agents, I think the biggest thing that changes is you go from, um, hey, check 1, 2, 3, these exact steps to, well, here's the criteria for what I would gauge something as bad. I would judge this as risky if these sorts of things were true. And I'm not telling the model how to do those things, I'm just saying, hey, if these things are true, then I would gauge this as risky. And then what the model can do is it can use all the tools and data and context it has available to it, and it can go and answer those questions and it can ask itself, you know, subsequent questions and it can follow leads and threads. But again, if you give the model really strict instructions, it doesn't really deviate much. But I think that's a waste of the technology. Right? Uh, you might as well just use a script. At that point, you're not really getting the benefit of an LLM. So making this shift into explaining in the description, hey, this is the threat model that we're covering here. We're monitoring failed logins because identity is a very important vector that effectively gives anyone access to all internal tools. And you can explain the threat model that you care about very specifically. And then on the runbook side, it isn't, hey, go follow these steps and if else do that, it's more of the criteria like, hey, this is what I would look at. Kind of like if you were giving it to a senior analyst, someone who's been working in cyber forever, like, okay, well these three things, this is evidence that I would, you know, gauge this as risky. So that's a huge paradigm shift. And, and then the other thing that's very related to this is because we're using agents for code generation, you can actually create these detections in these runbooks significantly faster and easier. We never had time to do this before. So we would create these runbooks that were like two lines or like, it was just like check activity and if bad page or something. Right. It's just like very basic because we're again, like, we only had time to look at half the alerts that were generated. Right. That's the stat that everyone's throwing around.

Speaker B: I love the idea that you're describing of shifting to the kind of the definition of the classification of the thing that we're looking for. Because LLMs really aren't deterministic. And I suspect this rush and iteration of vendor technology we're going to see is going to be shoehorning the deterministic pattern into an LLM model and then the iteration after that is going to be more. These are the signals that you're looking for, not the like quasi deterministic. The worst of both models.

Speaker A: Yeah. And then, you know, the operating model is also very different here because you're using agents to do that triage. You don't even need to do this sort of atomic detection anymore. Right. And this is, you know, the most recent blog on detection scale about what happens when, you know, agents are, are really doing the triage is you can build all these new paradigms like we were talking about around threat model threat hunting. So you can have an agent spin up and maybe look at a particular set of opinionated logs about maybe all logins from the last hour and then you can advise it, hey, go hunt for these types of attacks. So you can actually even be looser on what you're looking for. And it's what we saw around uh, the Mythos model being announced because they're saying, oh, it can go find zero days and it can go find 20 year old bugs, uh, just by looking at the source code directly and looking for particular types of vulnerabilities. So we could do the same thing in the SoC where we're saying, look at all the logins and here's a few types of threat models that you could identify. But look for those plus any others that look potentially suspicious. And I think this is going to become a more emerging way of doing detection. And then on the flip side of huh, that. Because that's a very proactive way of doing it, I think the reactive way is also going to change as well. Where we are probably going to change the meaning of an alert and it's going to become much more signals based where you're actually just looking at generated signals and maybe you're looking at the identities and the IPs behind them and then you're determining dynamically, okay, I'm going to create an alert. Right. Or it ends up feeling a little bit more like a case where I say, okay, I'm going to follow these threads, these five signals, and I'm going to triage those together. So I think that paradigm is also going to change really significantly going forward. And then the humans again will be in the loop in the critical path of hey, remediation or uh, making a change to the system or you know, killing a user account, whatever it may be. So yeah, I think we're in a really interesting time of reinvention.

Speaker B: I think it's an absolutely fascinating time of reinvention. Like I think the next couple years, very, very specifically in this vertical that we're in is going to uh, evolve very, very quickly and very powerfully. Yeah. I'm curious, from your perspective, what do you think this, this means for uh, the actual day to day of the analysts and detection engineers and leaders out there?

Speaker A: It's shifting significantly and I would say you have to embrace the agents just because there it is the preferred way of doing this job. And the agents afford the ability for you to have significantly higher impact.

Speaker B: Do you have some examples that you can share with people of what you've seen out in the wild tactically?

Speaker A: I mean, I know several of the security teams that we support are very small and they're small because uh, the leaders don't have a lot of budget. So they've employed agents to do these workflows that we're talking about. So using agents for doing all of your triage is a great example. Using agents for doing threat hunting is a great example. And just taking these flows that are very manual before and figuring out how to encode the expertise of your team into a skill or um, you know, the workflow that you would do for your particular organization into a skill, a prompt, et cetera and uh, really just operationalizing it. So we see that every day with the teams that we work with.

Speaker B: So you mentioned closing the loop. We often hear about closing the alert. Closing the loop. What's the distinction between these things? Why does it matter?

Speaker A: The easiest sort of value proposition was bringing these agents to do um, alert triage.

Speaker B: Right, Because. Deal with alert fatigue.

Speaker A: Exactly. To deal with alert fatigue, which is, you know, the big problem that we have ubiquitously in the industry. So when we think about closing the alert, that's fine, but the problem is that if you run agents in isolation, I think they have a few things at their disadvantage. Right. I think that they're not very close to the data so they don't have as much access and understanding to where the data came from, the schema, the shape of the Data. Um, I also think they don't have a deep understanding always of the detection, correlation, language in the system, the annoying system, because at the end of the day that comprehension is so important to both understand the purpose of an alert, but also understand how to improve it next time. So when we talk about closing the loop, closing the loop is about continuously improving the program. And it's something that we always strive to do as a SOC team. We always want things to be better. We always want the positive rate.

Speaker B: You don't want the same mistake twice.

Speaker A: Exactly. Yeah. And noise is a weird thing because sometimes you get an alert and it's benign, but you're glad that you saw that. That's not inherently a bad thing. But I want to be able to have the ability with AI to improve my system. And if you have disparate tooling, that becomes very challenging. So our whole approach has always been give really strong scalable infrastructure to where teams can get all their logs in one place if that's the design that they want, which in most cases the socket is. And now we've sort of augmented that with agents and really flipping the operating model around onto the agent. What does the agent need to do its job really well? Well, it needs the ability to understand what the detection is. It needs the ability to, uh, distill its judgment somewhere. Like it has to give an outcome and put a conclusion. Like, we want alerts to be finished. Right. And then when you do that really well, now the agents have the ability to actually improve the underlying system with of course, the level of autonomy that you grant it. So closing the loop is just this concept of is the system getting better over time and does the agent have the ability to actually make that happen?

Speaker B: Jack, thank you for having me as a guest on the show and I'm looking forward to this amazing AI dominated future.

Speaker A: Yeah. Uh, thank you for doing a good job interviewing and, um, that was really, really good conversation. Really appreciate it. And, uh, I'm excited to have more conversations about where the world's headed and AI and cyber, probably.

Speaker B: I haven't been as excited about technology as, like, it reminds me of 2011, 2012. And it's just like where we're going to be in a month's time is different than where we are today.

Speaker A: Right.

Speaker B: Like we could have the exact same conversation about where AI is going and have, uh, different answers. And that's really exciting.

Speaker A: Well, we can look back in 10 years.

Speaker B: Yeah.

Speaker A: So that's a very easy thing to do. I'll ask my agent in 10 years. What did I say on that podcast? 20, 26. Thank you, Julian. Appreciate it.

Speaker B: Thank you, J.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • #291 Why Most AI Projects Fail to Deliver ROI, Sinohe Terrero, CFO and COO, EnvoyGrowCFO Show · on Claude91 / 100
  • Decision Logic: The Difference Between an Answer and a DecisionThe AI Forecast · on Agentic AI87 / 100
  • KYA Won't Always Protect You. The Real Risk Is the Swarm!Fintech Conversations & Insights with Efi Pylarinou · on Agentic AI86 / 100
  • Why a $1.2B exit felt like his biggest failure, and the customer-obsession thesis behind AgencyThe GTMnow Podcast · on Claude86 / 100
  • Agentic AI in Sales: What Business Leaders Need to KnowScaling with AI · on Agentic AI86 / 100
  • Is Your Business Invisible to AI Search? (And How to Fix It) ft. Ray YoungRevenue Science · on Claude85 / 100

More from Detection at Scale

All episodes →
  • Google's Michael Sinno on Autonomous Detection at 7 Trillion Logs Per Day
  • Block's CISO James Nettesheim on How 40% of Their Detections Are Now Written with AI
  • Compass' Ryan Glynn on Why LLMs Shouldn't Make Security Decisions - But Should Power Them
  • Veeva Systems' Mike Vetri on Building Threat Operations Teams and AI-Powered Investigations
  • Trustpilot's Gary Hunter on Structuring Security Knowledge for AI Success
Explore the best B2B Engineering & DevTools podcasts →
All Detection at Scale episodes →