The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/CheckMates Go
CheckMates Go artwork

S08E04: The Attacks Don't Stop!

CheckMates Go · 2026-07-01 · 15 min

0:00--:--

Key moments - from our scoring

Substance score

30 / 100

Five dimensions, 20 points each

Insight Density8 / 20
Originality5 / 20
Guest Caliber5 / 20
Specificity & Evidence9 / 20
Conversational Craft3 / 20

Phoneboy and Michael Greenberg examine why exposure management has become critical as ransomware attacks surge 50% in 2025 and adversaries operate at machine speed. The core problem: organizations deploy 30-40 security tools that create visibility without actionable intelligence, leading to scattered teams (SOC, vulnerability management, infrastructure security) working in silos with no unified response framework. Greenberg explains how manual remediation processes take days or weeks while attackers exploit known CVEs in hours, creating alert fatigue and false confidence despite robust tool investments. The episode also covers this week's AI security governance news - including the Five Eyes alliance warning that frontier models now pose sophisticated cyber threats, OpenAI's delayed GPT-5.6 release after US government review requests, Anthropic's Mythos model's performance in NSA red team exercises, and the discovery of low-skill attackers using Claude and Codex to breach 14 organizations via AI agents. Key concerns include AI-assisted phishing-as-a-service (1,380% increase in device code attacks), deepfake services on dark web marketplaces, and AI coding agents silently spreading through 180 million open source repositories with minimal detection. CISOs, security teams, and platform engineers evaluating exposure management platforms and AI red teaming capabilities will find actionable frameworks here.

Key takeaways

  • →Exposure management helps organizations identify, prioritize, and safely remediate security exposures before attackers can act, addressing the gap created by having 30-40 disconnected security tools.
  • →Ransomware attacks increased 50% in 2025, and attackers are exploiting old vulnerabilities faster than new ones are being weaponized, often succeeding in hours while defenders take days or weeks to respond.
  • →AI is lowering the skill barrier for cybercrime by automating offensive workflows; researchers found a low-skill attacker used Claude and Codex to breach 14 organizations with AI agents handling reconnaissance through data harvesting.
  • →AI coding agents are spreading through open source repositories in ways that may be invisible to maintainers, requiring detection methods that identify centralized bot accounts, commit message signatures, distributed human attribution patterns, and configuration files.
  • →The Five Eyes alliance confirmed that frontier AI models are now more capable than ever at launching sophisticated cyber attacks, prompting governments to implement pre-release review processes for advanced AI systems before deployment.

Guests

Michael Greenberg

Topics in this episode

Anthropic ClaudeOpenAI GPT-5Ransomware attacksFrontier AI modelsExposure ManagementCheck PointMeta AI reviewNSA Red Team testingOA LabsEvil Tokens phishing as a service

Questions this episode answers

What is exposure management and how does it differ from traditional vulnerability management?

Exposure management identifies, prioritizes, and safely remediates security exposures before attackers act, moving beyond siloed vulnerability lists by creating unified visibility and coordinated response across SOC, vulnerability, and infrastructure teams - eliminating the manual delays and lack of context that plague traditional point tools.

Why did OpenAI delay the full release of GPT-5.6?

OpenAI delayed GPT-5.6 after a US government request for a more limited rollout to vetted partners, reflecting a broader shift toward pre-release review of frontier models to evaluate cyber capacity, misuse risk, national security implications, and infrastructure vulnerabilities before wider deployment.

How did researchers at OA Labs discover a low-skill attacker using Claude and Codex for cyber attacks?

OA Labs analyzed the attacker's entire working directory after a third-party hosting provider discovered malicious activity and shared it with them; the attacker had committed over 1,000 AI agent sessions to the third-party server instead of running them locally, exposing their identity, location, education, and LinkedIn profile in the process.

What are the four types of AI coding agent detection methods identified in the ARCsive study?

The study classifies AI agent traces as: Type A (centralized bot account detected via email lookups), Type B (commit message signatures via regex queries), Type C (distributed human attribution with tool-specific suffixes like 'aider'), and Type D (configuration file only, silent in fewer than 0.5% of commits, detected via file-to-project map queries).

How much did Evil Tokens phishing-as-a-service increase device code phishing attacks in early 2025?

Evil Tokens drove a 1,380% increase in device code phishing attacks in early 2025, using AI-generated personalization to turn phishing kits into scalable, victim-by-victim tailored products that bypass traditional user warning patterns.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

8 / 20

The AI news segment contains a reasonable volume of concrete, current information (Evil Tokens stats, Anthropic/NSA red team story, AI coding agent detection methodology), but the first half of the episode is dense with generic cybersecurity pain-point narration that adds little beyond what any practitioner already knows. The newsletter-reading format limits synthesis and depth.

Huntress reports that Evil Tokens phishing as a service operation drove a 1,380% increase in device code phishing attacks in early 2026
type D which is configuration file only and this tends to be silent. Now the agent leaves traces in fewer than 0.5% of commits

Originality

5 / 20

The episode recycles well-worn cybersecurity tropes in the first segment (alert fatigue, tool sprawl, 'context is king') and the second segment is explicitly a curated summary of public news articles with minimal original editorial framing or contrarian analysis. Almost no first-principles thinking is present.

The attacks don't stop. They really don't. Um, it may sound cliche but it's real
knowledge is power. Let's see everything that we can

Guest Caliber

5 / 20

The sole 'guest' clip is from a Check Point internal employee (Michael Greenberg, exposure management team) delivering what reads as a product-marketing pitch for Check Point's own tooling; the host Phoneboy is a Check Point brand persona. No independent external practitioners or verifiable at-scale operators appear.

why is Jacob's team working so hard? Why is R and D working so hard to develop uh, ah, security
there's also a white paper there that um, goes into more detail and also um, you can look into our AI red teaming tool

Specificity & Evidence

9 / 20

The AI news segment includes several concrete data points with named sources (Evil Tokens, NordStellar, ARCsive, OA Labs, Huntress) and specific figures, and the agent-trace detection taxonomy is genuinely detailed. However, attribution is secondhand and the ransomware statistics in the first segment are vague and unsourced.

a NORD Stellar report found a 39% rise in dark web discussion of deepfake as a service offerings compared with last year
a new ARCsive study detected AI coding agents traces across 180 million repositories

Conversational Craft

3 / 20

There is no actual conversation in this episode: the first half is a replayed clip from a previous webinar and the second half is the host reading a pre-written newsletter summary aloud. No follow-up questions, no pushback, no dialogue of any kind occurs.

The link to the full session is of course available in the show notes
let me know what you think by leaving a comment on the show post on Checkmates

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B57%
  • Speaker C39%
  • Speaker A3%

Most-used words

security9attacks8agent8tools7model7agents7systems7report7checkmates6exposure6management6team6huge6review5news5source5

Episode notes

On this episode, we talk about The Great Exposure Reset and some recent AI news. See also: From Prompt Testing to AI Red Teaming at Enterprise Scale Leave comments on CheckMates

Full transcript

15 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: M welcome to the Checkmates Go podcast. Join your favorite check point expert, Phone Boy and his guests as they cover a range of cybersecurity topics to help you secure your everything. Be sure to subscribe and share and don't forget to rate and review us. And now, here's Phoneboy.

Speaker B: Hello and welcome to the Checkmates Go Cybersecurity podcast. I'm Phoneboy and this is season eight episode four. Uh, now in this episode we're going to talk about the great exposure reset and some recent AI news. So let's get started. Back in February we had a tech talk with the exposure management team introducing the idea of, well, exposure management and more specifically exposure management helps identify, prioritize and safely remediate exposures before attackers act. Michael Greenberg explains the problem space in detail.

Speaker C: The attacks don't stop. They really don't. Um, it may sound cliche but it's real and we all know it. We've uh, all hear it a hundred times but it's the underlying truth. It's the foundation of you uh, know what, what we're here for and what we're trying to solve. If we look at today, we know that ransomware, uh, attacks are ah, huge. There's a 50% increase over 2024. If we're just looking at 2025, uh, the amount of ransomware attacks that are occurring all over the world. This uh, is just looking at one specific uh, attack source. Um, you know, you can pick a company, they've had a recent breach. They were, you know, uh, credentials were compromised on the Dark web. They uh, uh, were exploited and all of their uh, data was exposed. Uh, you know, insert huge uh, uh, data breach here now also which we're all familiar with, uh, which are huge attack vectors, uh, for everybody, every organization. It doesn't matter if, if you're in retail manufacturing, uh, uh, if you know, phishing attacks, they're 1 in every 68 email else has an attachment received. Uh, uh, in an organization that's actually malicious, uh, identity, identity, identity credentials, uh, are huge. They're everywhere. Uh, they're being compromised and you know, that's the, the keys to the Holy Grail. Once you have that it's pretty easy to, to pivot and you know, grant permissions, et cetera, et cetera. Um also you know, the, the use of old vulnerabilities, right, being you know, attackers, ah, taking advantage of why, why am I going to new and hot and fresh when this old uh, trusty, trusty sidekick is doing the job. And those are Actually being weaponized, uh, you know, new ones in less than a day, uh, uh, faster than new vulnerabilities. And there's the supply chain. Right. Uh, we're all connected everywhere. This cloud, that cloud, this supplier, that supplier, this third party for um, all sorts of resources and applications and services that we require. Now the painful truth, the reality. Why is uh, Jacob's team working so hard? Why is R and D working so hard to develop uh, ah, security, ah, right. And provide this protection, um, and provide this ability to give exposure management, um, is because at the moment what uh, brought us here is the painful truth. We know that organizations, right, have all types of security products. They can have 30 tools, 40 tools in their environment. Uh, you know, visibility is great. Uh, I think that was the goal, uh, I would say in our industry was, you know, knowledge is power. Let's see everything that we can. Uh, but ultimately uh, this is really creating a huge gap, right? It's there. Um, our, our teams, the people we're working with, whether they're in it, whether they're in the SOC team, whether they're in um, the vulnerability side, whether they're our ciso and we don't have the confidence that we're actually protected against all of those ransomware attacks that are happening. Right? We uh, don't have a conf that uh, you know, the 60 hosts that I have vulnerable to the CVE are going to, you know, can be protected and I'm not going to break anything in the process and ruin uh, productivity. Right. And ruin uh, business operations. Uh, the tools are also connected, right? We have so many uh, tools in our environment. Uh, the teams, you know, are scattered. We have the SOC teams chasing threats. We have the vulnerability management team trying to prioritize, uh, all of those uh, exposures. Uh, you know, my wife's in product manager management. Uh, they got their uh, you know, tenable list, uh, uh, uh, audit, you know, and they're going to work, work through it top down, right? And, and you know, that's their uh, way to prioritize it. And unfortunately they're not discussing with any, any anybody else and it's only looking at the one uh, source, uh, what's coming into it. And then it's the security infrastructure people, right, that uh, uh, teams and leaders that are there that are responsible for actually taking that uh, intelligence and actually fixing it. So uh, there's obviously a huge impact to all of this building up to what we're here to all talk about and dive into. But it's very uh, important to understand the uh, connection and how we, how we actually are going to solve this problem. So everything's manual, everybody's overloaded, right? We're all aware there's alert fatigue, right? There's no shared view, there's no unified response, there's uh, no context. I think context is king. Right? If you're not looking at the entirety of your environment and you're looking at everything siloed, how can you even uh, pretend to start to apply, uh, cyber security? Um, and the reality, right, it's the first slide I started with. Attackers are moving in hours. They're, you know, leveraging old CVEs, they don't, they don't care, right? They don't have uh, authorities, uh, uh, to listen to, they don't have timelines to follow, they don't have deadlines, they don't have uh, Patch Tuesday coming up. And you know, the, the exposure on the, the defender side, uh, being able to actually remediate that is taking days or it's taking weeks to understand who's the problem, who' responsible, are they able to fix it? What's the impact? Did I do the analysis to understand if I'm going to create a false positive out of this, making this change? There's so much manual labor that, that, that goes into this that it takes a really long time. And that's where we're here. This is what, uh, the problem that we're, we're out to solve to uh, make the, the, the life of our customers, our partners, uh, uh, and ourselves easier.

Speaker B: The link to the full session is of course available in the show notes. Now Check Point produces a weekly article that's posted on Checkmates called this Week in AI, which summarizes the AI related news for the past week and links to the articles. Now, I thought it would make for good podcast content, so here we are. Let me know what you think by leaving a comment on the show post on Checkmates, which of course I'll link in the show notes. Now, this week's AI security news is unusually concentrated around governance, misuse and visibility. From US Government review of Frontier model releases to AI assisted cybercrime deepfake services and the spread of coding agents across open source, the central question becomes how do we keep powerful AI systems useful, auditable and safely deployed? Well, uh, let's get into it now. The Five Eyes Intelligence alliance, which is the us, the uk, Canada, Australia and New Zealand, recently declared that AI systems are now more capable than ever at launching highly sophisticated cyber attacks. Frontier AI models are anticipated to exceed current industry expectations, the alliance wrote. And that might explain some of what's been highlighted in the news here. Uh, so OpenAI is reportedly delaying the full release of uh GPT 5.6 after a US government request for more limited rollout to vetted partners. Now, the story reflects a broader shift toward pre release review of ah frontier models where cyber capacity, misuse, risk and national security are part of the deployment discussion. Now, officials are urging Meta to join other major AI labs in submitting advanced models for government review before release. The program focuses on evaluating model capabilities, vulnerabilities, military implications and the national infrastructure risk before frontier systems reach broader deployment. Now there's also a new report that claims that Anthropic's Mythos model performed strongly in a controlled government Red Team exercise. While Anthropic has disputed parts of the public framing. Now from m the report, um, according to a uh, report by the Economist, uh, Anthropic's powerful Mythos AI model was able to break into almost all classified systems belonging to the National Security Agency, that is the nsa, one of the highest ranking and most powerful intelligence agencies in the US Government within hours during a controlled security evaluation. Now the claim came from Senator Mark Warner, vice chair of the Senate Intelligence Committee, who said to General uh, uh, Joshua Rudd, the head of the NSA and US Cyber Command, and briefed him on the model's capability. Now Rudd reportedly told Warner, as cited by the economist in a June 14 report that initially went under the radar anthropic spot powerful Mythos AI model broke into almost all of our classified systems, not in weeks, but in hours. The quote then went viral about a week later across several social media platforms, generating claims that Anthropic's model hacked the nsa. In response, the original author issued a public statement on the June 21st clarifying that the narrative was false. The breach occurred during an authorized internal Red Team test in which Mythos was paired with other defensive tools under highly specific simulated environmental conditions. Now the broader significance is the policy change challenge around evaluating cyber capable frontier models without overstating results or limiting defensive research. Now, researchers say a low skill attacker used Claude and Codex to help breach 14 organizations with AI agents assisting across reconnaissance, vulnerability, identification, export generation and data harvesting. Uh, cybersecurity researchers, OA Labs discovered and analyzed the attacker's entire working directory. Now they were able to do this because the attacker did not run the AI agents on his own infrastructure, but rather on a third party server. When that third party discovered malicious activity, they downloaded the entire working directory and shared it with the researchers. Now OA Labs could not find evidence that the stolen data was compromised, uh, or monetized in any way, either by being sold on the dark web or by extorting the victim companies. They did however, find numerous pieces of evidence about the attacker's identity and whereabouts. Oalabs was thus able to analyze more than 1,000 agent sessions, seeing how the attacker was able with ease to bypass most of the agent's guardrails. Among the sessions were also the threat actor's CV with his full name, location, education history and LinkedIn profile, as well as his IP address which showed his location. Now the case shows how AI can lower the skill needed for cybercrime by turning vague intent into structured offensive workflows. You don't even have to use commercial or open source AI tools to execute these attacks, you can just purchase their capabilities as a service, lowering the skill even further. Huntress reports that Evil Tokens phishing as a service operation drove a 1,380% increase in device code phishing attacks in early 2026. AI generated personalization is turning phishing kits into, um, scalable products that can tailor lures victim by victim and bypass traditional user warning patterns. Meanwhile, a NORD Stellar report found a 39% rise in dark web discussion of deepfake as a service offerings compared with last year. Cheaper, more realistic deepfake tooling could make fake boss SC impersonation attempts and BEC style social engineering harder to detect. Now a new report covered by Axios suggests that AI agents, especially codecs, are beginning to handle more complex delegated tasks in workplace settings as agent adoption becomes more concrete. Organizations definitely need to manage what these systems can see, change, approve and trigger. Meanwhile, a new ARCsive study detected AI coding agents traces across 180 million repositories and found that single signal detection methods drastically undercut agent activity. The finding matters for software supply chain security because AI generated code is spreading through open source in ways that might be invisible to maintainers and reviewers. Now I was actually curious how this was being done, so I dug into the report and their detection strategy classifies AI traces into four behavioral types and applying a complementary detection method for each. Now, uh, type A centralized bot account. The agent commits under a single registered bot identity with a name and an email address and they're detected via exact email address lookups in the author map providing uh now type B commit message signature. The agent embeds explicit text in commit messages. The WOC clickhouse database is queried using case insensitive regular expressions to detect these signatures. Type C Distributed Human Attribution no central bot account exists. Individual developers append a tool specific suffix to their git author name eg aider. Now we identify these patterns by scanning all 32 uh, shards of the A2C full V2412 author map for the name pattern. Now type D which is configuration file only and this tends to be silent. Now the agent leaves traces in fewer than 0.5% of commits, making detection dependent on the configuration file presence. So they query file to project maps for agent specific files such as replit, uh, copilot instructions MD claude MD agents MD using strict end of line anchoring to avoid false positives. This category is highly dynamic because agents frequently change configuration conventions requiring periodic recalibration of detection rules. This will definitely be an interesting space to watch for sure. Now to wrap up the news enterprise, AI systems are no longer isolated chat windows. They include prompts, retrieval pipelines, APIs, tools, permissions, workflows, data sources, business logic, and changing models. Which means red teaming has to test the whole system as it runs to find the attack paths that create real risk and prove that they're closed after every change. Now there is a blog post that I'll link in the show notes here. Uh, ah, that's on the Checkpoint blog that goes into this in a little more detail and um, uh, there's also a white paper there that um, goes into more detail and also um, you can look into our AI red teaming tool. All right, with that I think we'll wrap it up and we'll see you next time for another episode of Checkmates. Go.

Speaker A: Thanks for joining us. For this week's episode of Checkmates. Go subscribe in your favorite podcast app, leave us a rating and review and share with your colleagues on social media and we'll see you next time.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Rebooting Enterprise AI with MCP and KubernetesPractical AI · on Anthropic Claude88 / 100
  • EP284 Closest Alligator to the Canoe: How Transforming SOC Became P0 for Lloyds BankCloud Security Podcast by Google · on Exposure Management85 / 100
  • Everybody Wants AI. Who's Paying for It?AI Proving Ground Podcast · on Frontier AI models85 / 100
  • Generative AI Meets Accessibility: Benchmarks, Breakthroughs, and Blind Spots with Joe DevonAI Engineering Podcast · on Anthropic Claude85 / 100
  • The Ultimate NVIDIA GTC 2026 Recap (with Karl Freund and Jim McGregor)The neXt Curve reThink Podcast · on Anthropic Claude84 / 100
  • Data, AI, and Knowing When to Let Go - with Tommy CotterDefinitely, Maybe Agile · on Frontier AI models81 / 100

More from CheckMates Go

All episodes →
  • S08E03: Local Config and Hope
  • S08E02: CheckMates Fest
  • S08E01: AI Security and More!
  • S07E27: A Strong Sense of Identity
  • S07E26: R82.10 and Rationalizing Multi Vendor Security Policies
Explore the best B2B Engineering & DevTools podcasts →
All CheckMates Go episodes →