The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Today’s AI News
Today’s AI News artwork

ChatGPT Co-Creator Launches Jev, Salesforce Trains Koa, Self-Improving AI

Today’s AI News · 2026-09-16 · 23 min

0:00--:--

Key moments - from our scoring

Substance score

67 / 100

Five dimensions, 20 points each

Insight Density15 / 20
Originality13 / 20
Guest Caliber11 / 20
Specificity & Evidence16 / 20
Conversational Craft12 / 20

This episode examines a fundamental pivot in AI architecture: away from visible, text-generating assistants toward invisible, ultrafast systems embedded in business logic. Diogo Almeida's Jev from Type Safe operates as a deterministic classification engine - 238 times cheaper than Claude 3.5 Sonnet and 40-200x faster - accepting strict predefined options rather than generating text, eliminating hallucinations entirely. Salesforce's Koa model addresses enterprise privacy constraints by training on fully synthetic data generated through adversarial role-play between AI systems, cutting error rates by 66% versus frontier models on internal business tasks. The episode then examines a Chinese research paper titled "The Last AI Built by Humans" that outlines five levels of recursive self-improvement (RSI), from human-guided upgrades through to complete autonomous architecture redesign. The hosts contrast Western safety frameworks treating level 5 RSI as existential risk versus Chinese researchers framing it as a targeted milestone. Practical applications discussed include OpenRouter API switchboards enabling AI-to-AI task delegation with spending limits, Gemini 3.5 Extended Thinking for natural voice pauses, and Step Audio 3's voice generation. Privacy concerns are highlighted via 404 Media's Project Lilly report on human contractors reviewing ChatGPT conversations without clear user consent.

Key takeaways

  • →Jev's deterministic architecture (selecting from predefined options) eliminates hallucination while costing 238x less than Claude and responding 40-200x faster, making ultra-cheap background inference viable.
  • →Salesforce's Koa demonstrates how synthetic training data from AI-role-played scenarios lets enterprises build specialized reasoning models without exposing proprietary customer data.
  • →Recursive self-improvement research maps five levels from human-supervised code upgrades to fully autonomous AI architecture design, with Western labs treating level 5 as existential risk while Chinese institutions frame it as a development goal.
  • →OpenRouter API keys with hard-coded spending limits enable safe AI-to-AI task delegation in consumer workflows, automating bulk work routing to cheaper models.
  • →Human contractors actively read and label ChatGPT conversations for training data without explicit user consent, requiring enterprise data processing agreements or local models for actual privacy.

Topics in this episode

Anthropic ClaudeOpenRouterJevType SafeDiogo AlmeidaSalesforce KoaNvidia Nemetron 3Gemini 3.5 Extended ThinkingStep Audio 3Odyssey 3

Questions this episode answers

What is Jev and why did ChatGPT's co-creator build a model that can't generate text?

Jev, built by Diogo Almeida at Type Safe, operates as a deterministic classification engine that selects from predefined options rather than generating text. It's designed for high-volume routing and categorization tasks (like customer service ticket routing) at extreme cost and speed efficiency - $42 per billion input tokens with 70-500ms response times, making it 238x cheaper than Claude 3.5 Sonnet.

How did Salesforce train its Koa model without exposing real customer CRM data?

Salesforce used fully synthetic training data generated by adversarial networks where one AI model role-plays as an angry customer while another plays a junior support rep trying to de-escalate. They simulated hundreds of thousands of realistic business scenarios across a dozen industries without touching any real proprietary customer data.

What are the five levels of recursive self-improvement described in the Chinese AI research paper?

Level 1: AI executes code upgrades designed by humans; Level 2: AI diagnoses its own weak spots and writes patches (human approves deployment); Level 3-4: AI autonomously determines what to learn, launches into new environments, and adapts code without human oversight; Level 5: AI completely redesigns the improvement process and builds its own successors with novel architectures humans may not comprehend.

How do you safely delegate tasks between AI models without generating massive cloud bills?

Use OpenRouter, which generates API keys with hard-coded spending limits (e.g., $5 max) and expiration dates. Hand this restricted key to your primary AI agent in its system prompt, instructing it to route bulk repetitive work to cheaper models while reserving expensive reasoning power for complex tasks.

Are my ChatGPT conversations actually private?

No. According to 404 Media's Project Lilly report, OpenAI has human contractors actively reading and labeling real user conversations for training data, often without clear upfront user consent. True privacy requires enterprise tiers with explicit no-training policies or running models locally on your own machine.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

15 / 20

The episode covers multiple substantive topics with concrete technical details: Jev's cost structure ($42/billion tokens, 238x cheaper than Claude), the synthetic training mechanism for Kowa, and the five-level RSI framework. However, there is significant filler including extended metaphors (chef/oven analogy), celebratory asides ('That is wild'), and repetitive conversational scaffolding that dilutes the insight-to-minute ratio. The content is meaty but padded.

JEV costs $42 for a billion input tokens...roughly equivalent to feeding the model a library of 10,000 average sized books for 42 bucks
They didn't touch a single piece of real customer data. The entire training set for Kowa is fully synthetic.

Originality

13 / 20

The framing of AI shifting from 'chatty' to 'silent background infrastructure' is a useful lens, and the Jevons Paradox application is reasonably fresh. The synthetic training data approach for Kowa and the five-level RSI taxonomy from the Chinese paper provide novel structure. However, much of the underlying analysis (hallucination problems in LLMs, geopolitical AI divergence, privacy concerns with human feedback) circulates widely in AI discourse. The show repackages existing ideas effectively but offers limited counterintuitive or first-principles thinking.

We are shifting from these chatty digital co workers to silent hyper fast background engines.
They didn't touch a single piece of real customer data...you set up a dynamic adversarial network using other frontier AI models

Guest Caliber

11 / 20

The episode cites Diogo Almeida (ex-OpenAI researcher, Type Safe founder), implies knowledge of Salesforce's Kowa training, and references a Chinese academic paper, but neither speaker is directly quoted as a guest. The hosts are conversational journalists synthesizing third-party research and product announcements rather than interviewing practitioners. The content relies on secondary reporting rather than direct builder expertise. For a news/summary format this is acceptable, but it lacks the depth of direct operator perspective.

Diogo Almeida. He's an ex OpenAI researcher who helped build the early methods that taught ChatGPT how to converse naturally with humans.
Over 30 Chinese AI researchers representing major institutions like ByteDance, Tsinghua University and Shanghai AI Lab just published a paper

Specificity & Evidence

16 / 20

The episode excels at concrete numbers and named examples: Jev's $42/billion token pricing, 238x cost reduction vs Claude Fable, 70-500ms response times, three times fewer errors for Kowa on internal benchmarks, 75% of RSI papers at levels 1-2, less than 6% at level 5, Meta One pricing from $2.99 to $499/month, the Mopsy emergency care case. These are specific, quantified data points that ground the claims. Some sections (e.g., Gemini Extended Thinking, Odyssey 3 world models) lack hard metrics, but the overall specificity is high.

JEV costs $42 for a billion input tokens...roughly 238 times cheaper to run, and it responds in 70 to 500 milliseconds
On their internal tests for tasks like updating deal stages or routing complex tickets, Kowa made three times fewer errors than the top Frontier models.

Conversational Craft

12 / 20

The hosts demonstrate solid back-and-forth with genuine questions ('If it's not generating text, what is it actually doing?', 'Aren't they just scooping up all proprietary CRM data?') and logical follow-ups that build understanding progressively. However, the questioning is often gentle and affirming rather than challenging. There are few moments of productive disagreement or skepticism; the hosts largely guide listeners through a pre-built narrative. The Mopsy story feels like a predetermined heartwarming segment rather than organic discovery. The exchange is conversational but lacks the sharpness of a host pushing back or forcing guests/sources to defend claims rigorously.

But if it's not generating text, what is it actually doing with those 10,000 books worth of input?
If Salesforce is training a new model to understand the deep nuances of sales and support, aren't they just scooping up all of their enterprise customers private CRM data to feed the machine?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B53%
  • Speaker A45%
  • Speaker C2%

Most-used words

model21models16level16human15massive13logic10highly10called9data9frontier8complex8real8today7routing7world7self7

Episode notes

Former OpenAI researcher Diogo Almeida has launched TypeSafe with Jev, a fast, low-cost system that makes preset decisions inside software and returns confidence scores. Salesforce introduced Koa, an in-house reasoning model trained on synthetic business scenarios for sales and support agents. Chinese researchers also published a five-level roadmap for recursive self-improvement, with the final stage describing AI systems that can redesign how their successors improve. We also cover Google's Gemini 3.8 Live, Odyssey 3, Meta One, and concerns about human review of real ChatGPT conversations.

Full transcript

23 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Welcome Back to today's AI news, brought to you by 9X Productions. Okay, let's unpack this. Imagine hiring a brilliant chef who refuses to cook meals, but instead, uh, instantly perfectly calibrates the temperature of every oven and pan in your kitchen.

Speaker B: That is a great way to put it.

Speaker A: Yeah, Right. And that's exactly what is happening in artificial intelligence right now. We are shifting from these chatty digital co workers to silent hyper fast background engines.

Speaker B: It's a huge pivot.

Speaker A: It is. So the mission of today's Deep Dive is to explore this massive shift in how AI operates behind the scenes. We're looking at how the very creators of our chatty models are now building systems that literally refuse to talk. Right up to, you know, the roadmap for AI that builds itself.

Speaker B: Yeah, I mean, we spent the last few years in the era of the text box, right? You type in a prompt, you wait, and you just sort of hope the model gives you something useful.

Speaker A: Thankfully waiting for the typing animation.

Speaker B: Right. But we are now moving into an era where AI is this unseen infrastructure. It's making thousands of micro decisions a second before you even realize a problem exists.

Speaker A: And there's this fascinating irony right at the center of the shift. The person who helped teach AI how to talk to us in the first place is now leading the charge to silence it.

Speaker B: Diogo Almeida.

Speaker A: Yes, Diogo Almeida. He's an ex OpenAI researcher who helped build the early methods that taught ChatGPT how to converse naturally with humans. Well, he just launched a new model out of his company, Type Safe. It's called Jev. And by design, it cannot generate text.

Speaker B: Not at all.

Speaker A: Right. Can't write you a poem. It can't draft an email. It literally cannot hold a conversation.

Speaker B: Which, you know, sounds completely counterintuitive to everything we've been told the market wants right now.

Speaker A: Right, Everyone wants an assistant.

Speaker B: Exactly. But when you look at the underlying economics and, uh, the actual mechanics of Jev, the strategy becomes incredibly clear. The numbers are just staggered.

Speaker A: They lay the numbers on us.

Speaker B: So JEV costs $42 for a billion input tokens.

Speaker A: Wow. Wait, a billion?

Speaker B: A billion. To put that into perspective, that's roughly equivalent to feeding the model a library of 10,000 average sized books for 42 bucks. Yeah, and the kicker is that the output from the model is entirely free.

Speaker A: That is wild.

Speaker B: Right. If you compare that to a standard frontier model, like say, Claude Fable 5.1, JEV is roughly 238 times cheaper to run, and it responds in 70 to 500 milliseconds so we are talking about

Speaker A: a response time that is 40 to 200 times faster than the large language models you and I use every day.

Speaker B: Literally a fraction of a second.

Speaker A: But if it's not generating text, what is it actually doing with those 10,000 books worth of input? Like, how does a model function if it can't speak?

Speaker B: So it operates as what they call a frontier intelligence function.

Speaker A: Call. Okay, what does that mean?

Speaker B: Well, think about how a traditional language model works. You ask it a question and it acts probabilistically. It guesses the next most likely word and then the next, and then the next, until it builds a sentence.

Speaker A: Right, the autocomplete, um, on steroids thing.

Speaker B: Exactly. And that takes massive compute power. It takes time. And crucially, because it's guessing, it opens the door for the model to hallucinate or, you know, invent facts.

Speaker A: Right.

Speaker B: Jev completely removes the guesswork. You feed it a massive amount of context and you give it a strict predefined menu of options to choose from. It acts deterministically.

Speaker A: So it just picks from the list?

Speaker B: Yeah. It selects an option from your menu and assigns a mathematical confidence score to its choice.

Speaker A: It's essentially a multiple choice savant living inside the software.

Speaker B: That's a perfect way to describe it.

Speaker A: And if it can only choose from the preset options you give it, then fundamentally the architecture prevents it from hallucinating. I mean, it literally doesn't have the capability to make up a new answer.

Speaker B: Exactly. That structural limitation is its superpower. Companies can use this for ultra fast, high volume categorization.

Speaker A: Like what kind of categorization?

Speaker B: Imagine a massive retail platform receiving thousands of customer service requests a second. Jev can read those requests and instantly route them to the correct department, or score database records for relevance, or even act as a digital bouncer.

Speaker A: A digital bouncer. I like that.

Speaker B: Yeah. You can put Jev in front of another AI model to instantly screen its outputs and ensure they don't contain harmful content or jailbreaks. And it does it all in milliseconds.

Speaker A: That bouncer concept really highlights the utility. I mean, you wouldn't pay a PhD level consultant hundreds of dollars an hour to stand at a door and check id'.

Speaker B: No, you wouldn't.

Speaker A: And companies don't want to pay frontier AI models to do basic logic routing.

Speaker B: Exactly. And this brings us to a textbook example of the Jevons Paradox.

Speaker A: Oh, right, economics 101.

Speaker B: Yep. The Jevons Paradox observes that as a technology becomes more efficient and cheaper to use, the demand for it actually Increases rather than decreases.

Speaker A: Like expanding a highway to reduce traffic. But the wider road just invites more cars until it's a parking lot again.

Speaker B: Exactly. Because Jev is so incredibly fast and cost effective, it. It won't just replace existing AI tasks. It's going to be embedded as invisible plumbing in places where I never would have been able to afford an AI before.

Speaker A: We are literally pulling the AI out of the chat window and burying it in the code. And while Type Safe is focusing on these ultra fast micro judgment calls, we're seeing massive enterprises deciding they want to build their own internal logic engines to do this kind of invisible routing on a much larger scale.

Speaker B: Right. Because, uh, large enterprises dealing with millions of sensitive transactions face a totally different challenge. They need reasoning and logic, but they cannot risk sending highly proprietary data out to a general purpose chatbot's servers.

Speaker A: No, absolutely not.

Speaker B: They need walled gardens. And that brings us to Salesforce.

Speaker A: Right. Salesforce just unveiled a newly trained reasoning model called Koa. They built this on top of Nvidia's open Nemetron 3 supermodel. They spent serious time adapting it specifically for hardcore business tasks.

Speaker B: Things like navigating complex sales negotiations and routing high tier customer support tickets.

Speaker A: Right, but I want to pause here and ask the question that any Salesforce user listening right now is probably shouting at their device.

Speaker B: Let's hear it.

Speaker A: If Salesforce is training a new model to understand the deep nuances of sales and support, aren't they just scooping up all of their enterprise customers private CRM data to feed the machine?

Speaker B: It's a completely valid concern, but the mechanism they use to train Kowa is actually a brilliant workaround. They didn't touch a single piece of real customer data. Wait, really none? The entire training set for Kowa is fully synthetic.

Speaker A: Okay, walk me through how that actually works in practice. How do you teach an AI the nuance of a complex human sales deal without showing it a real human sales deal?

Speaker B: Well, you set up a dynamic adversarial network using other frontier AI models. You prompt one model to act as an incredibly irate customer who say, just receive a delayed shipment of perishable goods.

Speaker A: Okay.

Speaker B: And then you prompt a second model to act as a junior customer service rep trying to de escalate the situation and route the ticket. You just let those two AI models converse and argue with each other.

Speaker A: Oh, wow.

Speaker B: Yeah. And it generates a highly realistic transcript of a complex business interaction.

Speaker A: So they essentially role played the training data into existence.

Speaker B: Exactly. They simulated these highly specific Personas across over a dozen different industries. They generated hundreds of thousands of these realistic, complex scenarios that capture the nuances of business logic, all without ever exposing a single line of real proprietary customer info.

Speaker A: That is wild.

Speaker B: And the benchmarks show the efficacy of this method. On their internal tests for tasks like updating deal stages or routing complex tickets, Kowa made three times fewer errors than the top Frontier models.

Speaker A: Three times fewer. Which puts Salesforce in a highly strategic position. By hosting KOA internally, they can guarantee to a bank or a healthcare provider that their sensitive routing requests never leave the Salesforce ecosystem.

Speaker B: Exactly.

Speaker A: But they're also acting as a secure bridge to the outside world for tasks that require broader knowledge. Right?

Speaker B: Yeah. They just launched a beta called claudeforce alongside a suite of AI force tools. They're essentially telling their clients, use our highly tuned synthetic COA model for the high volume, sensitive internal logic. But if you need to summarize a massive external report, we will securely pipe your data into Anthropic's claude, act as the distribution partner, and bring the result back to you.

Speaker A: So the theme across the board here is that we are deliberately embedding AI directly into the foundational logic of our software. It's writing the rules, routing the traffic, making background decisions.

Speaker B: Yeah, it's everywhere.

Speaker A: But that raises a pretty wild question. What happens when we take that internal logic and point it at the AI's own underlying code?

Speaker B: And this points directly to a new piece of research that maps out exactly that scenario. Over 30 Chinese AI researchers representing major institutions like ByteDance, Tsinghua University and Shanghai AI Lab just published a paper, and the title alone is striking. It's called the Last AI Built by Humans.

Speaker A: Okay, that sounds like a sci fi thriller.

Speaker B: It really does.

Speaker A: But they're laying out a very real roadmap for recursive self improvement, or rsi. How do they actually define the journey of an AI learning to build itself?

Speaker B: So they break it down into five distinct architectural levels, which is incredibly helpful for cutting through the hype and understanding the mechanics of where we actually are.

Speaker A: Okay, what's level one?

Speaker B: Level one is the baseline we've been at recently. The AI is capable of carrying out code upgrades, but a human architect designs the upgrade, oversees the execution, and tests the the result.

Speaker A: So the human is still very much the foreman, and the AI is just the highly efficient bricklayer.

Speaker B: Exactly.

Speaker A: What happens at level 2?

Speaker B: At level 2, the AI begins diagnosing its own weak spots. It can scan its architecture, identify a bottleneck, and figure out how to write a patch for it. The human is still in the loop, usually just to hit the final deploy button. But the AI is driving the diagnosis.

Speaker A: That feels like where the cutting edge of coding agents is hovering right now.

Speaker B: It is. The leap happens at levels three and four. This is where the training wheels come off. The AI takes over the entire process of determining what it needs to learn next.

Speaker C: Wow.

Speaker B: Yeah. It can launch itself into a new environment, gather new data post launch, and adapt its own code autonomously without human oversight.

Speaker A: Which brings us to the end game. They outline level 5.

Speaker B: Level 5 is a fundamental paradigm shift. At this stage, the AI completely overhauls the improvement process itself. It isn't just patching bugs or optimizing a database query. It is designing and building its own successors from scratch.

Speaker A: That is insane.

Speaker B: It's inventing new architecture. Architectures that human researchers might not even be able to comprehend, entirely independent of our intervention.

Speaker A: And one of the most fascinating points in this paper was the explanation of how an AI achieves this. They specifically highlight that software engineering is the clearest and fastest path to this recursive self improvement, far faster than robotics or medicine.

Speaker B: And that really comes down to the mechanics of the feedback loop. Think about it. If an AI is trying to invent a new medical drug or improve a physical robot's joint movement, it faces the friction of the physical world.

Speaker A: Right. You can't rush physics.

Speaker B: Exactly. It has to wait for a physical synthesis, real world clinical trials, or physical stress tests. That is incredibly slow. But software code is entirely digital. An AI can write a structural fix instantly, compile it, run a million simulated stress tests in a matter of seconds, evaluate the output and iterate on it. The feedback loop is instantaneous.

Speaker A: It is important to ground this though, for everyone listening. We are not at level 5 yet. The researchers analyzed almost 500 existing papers on self improving AI, and they found that roughly 75% of current research is still firmly rooted in level 1 or 2.

Speaker B: Right.

Speaker A: With less than 6% even exploring level 5 concepts.

Speaker B: The timeline is still theoretical, sure. But what we need to pay attention to here is the geopolitical context. If you look at the safety frameworks published by major Western Frontier Labs, OpenAI, Anthropic, Google, they explicitly define level 5 and AI automating its own R and D as a massive existential risk.

Speaker A: Their safety teams are putting up deliberate guardrails to prevent it, or at least heavily monitor and slow down the approach to it.

Speaker B: Whereas this paper clearly indicates that the Chinese research community is treating level five not as a catastrophic risk to be contained, but as a deliberate, targeted milestone to be achieved.

Speaker A: And we aren't here to endorse one view over the other. But it is a massive divergence in how the major global players are viewing the ultimate finish line of artificial intelligence. Just objectively reporting the facts here, they see it very differently.

Speaker B: It represents a profound philosophical and strategic difference in how we handle the architecture of intelligence.

Speaker A: Let's bring this back down to earth, because while we might not be dealing with Level 5 recursive self improvement today, you listening right now can actually take the first steps of AI delegating to AI in your own daily workflows.

Speaker B: Yes, you can.

Speaker A: We are seeing tools that allow the models we use every day to route tasks to other models on our behalf.

Speaker B: A perfect practical application of this today is a platform called OpenRouter. It functions as a massive API switchboard that connects you to dozens of different AI models from various labs. And the breakthrough here is that you can give an AI agent like plodcode or a dedicated coding assistant direct access to this switchboard.

Speaker A: I would love to know the mechanics of how you actually set that up, because just telling an AI to go use another AI sounds like a great way to accidentally rack up a massive cloud computing bill.

Speaker B: Yeah, that is exactly the risk you have to manage. But the setup is actually quite elegant. You log into OpenRouter and generate a specific API key. But crucially, before you finalize the key, the platform allows you to set a hard coded spending limit. Okay, say, exactly $5 and an expiration date. You take that severely restricted API key and you hand it over to your primary AI agent in its system prompt.

Speaker A: I love this. It is literally like giving your executive assistant a corporate credit card with a strict $5 limit and telling them to go hire a bunch of temporary interns to do the grunt work.

Speaker B: That is the perfect analogy. You give your primary agent clear instructions. You say, when I ask you for a high volume of simple repetitive tasks, do not use your own expensive processing power. Use this API key to route the work to the chimpest, fastest model available on openrouter, gather their outputs, and bring the final result back to me.

Speaker A: So if you are a marketer brainstorming and you need 50 rough mockups of an ad campaign, your primary expensive reasoning model doesn't generate those. It farms the task out to a much cheaper model, saving you money and time.

Speaker B: Exactly. And a great pro tip for anyone implementing this. Have your primary agent write a permanent skill or a local script for this exact routing process. Yeah, that way the delegation becomes seamless. You just type a command like run bulk brainstorm on this concept and your agent automatically knows to tap into the switchboard. It is a very real microcosm of the automated background delegation we started the show talking about.

Speaker A: All right, moving from the invisible background engines back to the models we actually interact with. Lets look at the rapid evolution happening on the consumer side in our quick hits because frontier models are making massive leaps across voice video and we really need to talk about this privacy.

Speaker B: Yes.

Speaker A: Let's start with Google. Google just debuted Gemini 3.8 live and they've introduced a feature called Extended Thinking.

Speaker B: This is, ah, a fascinating mechanism. Previously, when you asked an AI voice model a highly complex question, it had to pause. It would process the logic in silence, creating this awkward gap in the conversation.

Speaker A: Right, the dead air.

Speaker B: Yeah. The Extended Thinking variant bridges that gap by using latent space reasoning to process the complex logic while simultaneously generating a filler audio track, saying things like hmm. Or let me think about that for a second. To keep the human engaged, it just topped the speech to speech quality rankings because it perfectly mimics human conversational patents.

Speaker A: In other audio developments, A massive new suite called Step Audio 3 just dropped, featuring five distinct models dedicated to voice, agents, transcription and music generation. And on the physical simulation front, Odyssey introduced Odyssey 3, which they're calling a new world model, launching in the coming weeks.

Speaker B: It is important to distinguish what a world model actually is. It doesn't just generate pretty pixels like a standard video generator. A world model is trained to predict actual physics.

Speaker A: Oh, interesting.

Speaker B: It understands gravity, collision and object permanence. Because it understands the rules of the physical world, Odyssey 3 is capable of acting as the control system for robot arms, humanoids, self driving cars, drones, and even video game environments.

Speaker A: And we're also seeing shifts in how we pay for all this compute Meta just rolled out a new paid tier across its apps called MetaOne. The pricing varies wildly based on what you need, ranging from $2.99 all the way up to $499 a month, giving users extra meta, AI, uh, usage limits and advanced tools for creators.

Speaker B: That's quite a range.

Speaker A: Yeah, but I want to pause here because we need to address a critical privacy report from 404 Media called Project Lilly.

Speaker B: Right. This report detailed a practice at OpenAI that highlights a severe friction point in the industry.

Speaker A: As I was reading this report, I realized how much of an illusion of privacy I've accepted. You know, you sit alone in your room typing your rough drafts, your deepest anxieties, or maybe your proprietary business ideas. Into ChatGPT, you assume it's a closed loop between you and A cold, unfeeling machine.

Speaker B: Most people do.

Speaker A: But this report detailed that hundreds of human contractors are actively reading, rating and labeling real ChatGPT conversations, often without the user having any clear upfront knowledge that a human being is looking at their chat logs. Now, we're not taking a stance on this policy, just impartially sharing the findings of the report here.

Speaker B: Well, it is a harsh reality check on how these models improve. To refine an AI's behavior and reduce errors, reinforcement learning from human feedback is required. The labs need humans to look at the AI's responses and grade them, but the data they are grading is often the user's personal input. We get so comfortable with the natural conversational interface that we forget the chatbot is fundamentally a data gathering mechanism for the lab.

Speaker A: The pushback here is that if you need actual privacy, you cannot rely on consumer grade chat interfaces. You have to use enterprise tiers that explicitly state in their terms of service that they do not train on your data. Or you have to run models locally on your own machine. Assume someone is reading it unless you have a contract that says otherwise.

Speaker B: It is the cost of using the frontier models for free or for a low monthly fee.

Speaker A: Despite all these high level privacy concerns and geopolitical races for self improving code, I'm constantly reminded that the most profound impact of AI right now isn't always at the enterprise level. Sometimes it's how it helps ordinary people navigate highly stressful, deeply human moments. We have incredible community workflow today that highlights exactly that.

Speaker B: This story really stuck with me. It is a perfect example of using AI to cut through the administrative friction that paralyzes us during a crisis.

Speaker A: Yeah. So an anonymous listener reached out to share how they essentially turned ChatGPT into an emergency care coordinator for their 10 year old dog, Mopsy.

Speaker B: Oh, Mopsy.

Speaker A: Mopsy had developed a severe dental abscess that caused massive swelling around her eye. It escalated so quickly that the pressure was actually tearing her retina. To save the dog's vision, they needed to get this infected tooth surgically treated immediately.

Speaker B: But as anyone who has tried to navigate the medical or veterinary system in an emergency knows, finding specialized urgent care is a nightmare of red tape and disjointed communication.

Speaker A: Right. The listener started manually calling vet offices and kept hitting structural brick walls. The 24 hour emergency hospitals handled trauma, but they didn't perform dental surgery. The specialty veterinary dentists were fully booked out for months and the standard primary care vets refused to just take her in. Their protocols demanded a multi step process. They wanted an initial consultation, then a separate appointment for blood work, and then maybe scheduling a surgical appointment for a later date. After making over a dozen failed calls, watching their dog in immense pain, the listener was completely overwhelmed by this system.

Speaker B: That feeling of helplessness when you are fighting against a broken, fragmented bureaucracy and time is running out is a universal human experience. So they turn to the AI and explain the exact medical urgency. The mechanism of what the AI did next is what is so impressive. It didn't just give advice. It acted as an autonomous agent.

Speaker A: Walk me through exactly what it executed.

Speaker B: First, it scraped the web to research veterinary practices, specifically within a one hour driving radius, calculating the actual logistics of travel time. Then it drafted highly specific medically urgent emails explaining the torn retina, uh, and the need for immediate surgical intervention, bypassing the standard consultation request.

Speaker A: It sent them from their email?

Speaker B: Yes. It utilized the user's Gmail to send those individualized, personalized emails to over 20 different practices.

Speaker A: And it didn't just fire off a bunch of emails and hope for the bout.

Speaker B: No, it organized the chaos. It simultaneously built a live outreach tracker. It made a spreadsheet containing the details of every single clinic it had contacted, the time the email was sent, and the contact info so the owner could monitor the entire outreach campaign in real time.

Speaker A: It is incredible. And within 30 minutes of the AI sending out that targeted blast, a, uh, veterinary practice responded. They read the specific medical details, understood the urgency, and told the owner to bring Mopsi in. They evaluated her the very next morning, ran the necessary bloodwork on site, and took her straight into surgery to relieve the pressure on her eye.

Speaker B: It took a task that would have taken a stressed, panicked human hours of manual labor, searching, drafting, tracking, negotiating, and compressed it into a localized, highly efficient background process that took minutes.

Speaker A: So what does this all mean? We've gone from models that work silently in the background to the theoretical roadmap of AI building itself down to an AI navigating bureaucracy so a dog can get emergency surgery.

Speaker B: If we connect all these dots, it raises a critical question to leave you with. If an AI can act as an emergency coordinator today, saving a dog's life simply by navigating the friction of our broken medical systems, what happens when it reaches that level 5 self improvement we talked about? When an AI is capable of designing its own architecture and fully understanding complex logic, will it just continue to politely navigate our broken bureaucratic systems for us to? Or will it redesign the entire veterinary and medical infrastructure from the ground up?

Speaker C: Uh, that is absolutely something to chew on. We might just find that the brilliant linguist in the mailroom decides to completely redesign the postal Service instead. Thank you for joining us for this deep dive into the unseen layers of AI make sure you hit subscribe so you never miss out on these daily explorations. And if you enjoyed today's conversation and found it valuable, we would love it if you could rate the show five stars. We'll see you back for tomorrow's deep dive.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Rebooting Enterprise AI with MCP and KubernetesPractical AI · on Anthropic Claude88 / 100
  • Legal Tech Doesn’t Need More “Lamborghinis” ft. Kara Peterson and Richard DiBonaBetween the Briefs · on Anthropic Claude87 / 100
  • Episode 429: Getting started with LLM WikisMicrosoft Cloud IT Pro Podcast · on OpenRouter85 / 100
  • AI Is Having Its Dropbox MomentAI Proving Ground Podcast · on Anthropic Claude85 / 100
  • Open Source vs. Closed Source, Memory Chips Eat AI Profits, Comcast Restructures | Diet TBPNTBPN · on OpenRouter80 / 100
  • CAD, BIM, and the AI Leap: Qonic & Raven!AI Across The Product Lifecycle Podcast · on Anthropic Claude80 / 100

More from Today’s AI News

All episodes →
  • Trump and China Reject AI Slowdown, Siri AI Ships, Microsoft Sets AI Rules
  • AI Labs Back a Slowdown, AI Boosts Prenatal Detection, OpenAI Delays IPO
  • Anthropic Misuse Report, DeepSeek V4.1 Flash, AI Coding and Finance
  • Anthropic's AI Safety Debate, Suno v6, Getting Found in AI Search
  • OpenAI's Secret Model Solves a $1M Math Problem, Meta Launches Its Own AI Agent
Explore the best B2B AI & Data podcasts →
All Today’s AI News episodes →