The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Today’s AI News
Today’s AI News artwork

GPT-5.6 Sol Launches, ChatGPT Work Takes On Cowork, Meta’s Muse Spark Undercuts the Frontier

Today’s AI News · 2026-07-10 · 21 min

0:00--:--

Key moments - from our scoring

Substance score

57 / 100

Five dimensions, 20 points each

Insight Density12 / 20
Originality10 / 20
Guest Caliber11 / 20
Specificity & Evidence13 / 20
Conversational Craft11 / 20

The frontier AI landscape is fundamentally restructuring around efficiency and business integration rather than pure capability arms races. OpenAI's new GPT-5.6 family - led by Sol, with mid-tier Terra and budget Luna options - trades expensive capability leaps for practical performance at sustainable pricing ($5/$30 per million tokens matching older GPT-5.5 rates). Sol outperforms competitors like Fable on agentic coding, where AI autonomously navigates codebases, writes, tests, and debugs iteratively. Notably, Sol autonomously post-trained Luna using synthetic data generation and machine-scored evaluation loops, eliminating human labeling bottlenecks. Meanwhile, Meta's Muse Spark 1.1 undercuts OpenAI at $1.25/$4.25 per million tokens while delivering superior agent reasoning benchmarks and a million-token context window. The shift reflects computational physics hitting diminishing returns: 10% reasoning improvements now require astronomical compute as high-quality training text depletes. For B2B operators, this means adopting model orchestration strategies - using expensive flagship models like Fable only for planning and review while routing repetitive agentic loops to cheaper worker models, achieving up to 60% token savings. The economic thesis extends to professional services: AI commoditizes middle-market work, concentrating demand on the top 1% of practitioners who blend AI efficiency with irreplaceable human judgment, taste, and relationship capital.

Key takeaways

  • →Sol's autonomous post-training of Luna using synthetic data and machine evaluation eliminates the human labeling bottleneck and represents a fundamental shift in how frontier models are trained at scale.
  • →Model orchestration - using expensive flagship models only for strategic planning and final review while routing iterative tasks to cheaper worker models - can reduce token costs by 60% without sacrificing output quality.
  • →The AI-driven market shift hollows out the middle class of professional services; demand concentrates on the top 1-5% of practitioners while routine work becomes commoditized and automated.
  • →Meta's Muse Spark 1.1 undercuts OpenAI on pricing while outperforming on agent reasoning benchmarks, signaling that the industry is optimizing for workflow integration and cost efficiency rather than raw capability leaps.
  • →Benchmark evaluation is breaking down as AI solutions become too novel and complex for human-created rubrics to accurately judge, making model comparison increasingly difficult.

Guests

Speaker B

Topics in this episode

Agentic codingModel orchestrationGPT-5.6 SolChatGPT WorkMeta Muse Spark 1.1Synthetic data post-trainingToken cost reduction strategySWE Bench ProRobbie's Lingbot World 2Rove image generation 2.1

Questions this episode answers

How does Sol's agentic coding differ from traditional AI code generation?

Agentic coding means Sol autonomously navigates entire codebases, writes code, runs tests, finds and debugs errors iteratively - going far beyond traditional autocomplete-style code suggestions. Sol achieves this through upgrades to computer use design and cybersecurity capabilities that facilitate autonomous iteration loops.

How did OpenAI train the Luna model without human labeling?

Sol autonomously generated millions of complex problem-solving scenarios, fed them to Luna, and then evaluated Luna's responses using machine scoring. This closed feedback loop of optimization eliminates the need for human data labeling and manual performance rating that was historically required.

Why is Meta's Muse Spark 1.1 priced lower than OpenAI's Sol despite better performance?

Mark Zuckerberg explicitly stated Meta is promising pricing that undercuts rival labs' margins. Spark 1.1 costs $1.25/$4.25 per million tokens versus Sol's $5/$30, reflecting Meta's strategy to capture market share through aggressive pricing while still delivering superior agent reasoning performance.

What is the 60% token reduction strategy and how does it work?

Model orchestration uses expensive flagship models like Fable exclusively for strategic planning and final review, while offloading iterative token-heavy tasks (web browsing, boilerplate coding, data extraction) to cheaper worker models like Codex or budget Claude tiers. The savings come from having cheaper models make iterative mistakes rather than burning expensive context on repetitive loops.

How does AI-driven search accelerate winner-takes-all economics in professional services?

Unlike traditional Google search where ranking 15th still generated trickle-down business, AI chatbots synthesize one best answer rather than listing 15 options. This concentrates demand on the top practitioners, similar to streaming music algorithms concentrating 99% of listens on top 1% of artists, while middle-market professionals get replaced by automation.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

12 / 20

The episode packs substantial technical detail on model economics, agentic coding, and the 60% token reduction strategy, but interleaves these with considerable throat-clearing, rhetorical questions, and repetitive confirmation exchanges ('Yeah,' 'Exactly,' 'Right') that dilute insight density. The Peter Hurley headshot example and winner-takes-all thesis are intuitive rather than novel for business operators already tracking AI disruption.

SOL autonomously postponed the smaller Luna model
agentic coding means the AI acts as an autonomous agent. You give it a high level goal and it navigates your entire code base

Originality

10 / 20

While the token reduction strategy (using cheaper models for iterative tasks, expensive models for planning/review) is practical, it is framed as standard model orchestration rather than a contrarian insight. The winner-takes-all economy and pricing-over-raw-capability pivot are widely discussed in AI industry discourse. The synthetic data feedback loop (SOL training Luna) is interesting but presented without critical analysis of its limitations or counterarguments.

the highly capable Fable model exclusively as your planner and reviewer
a brutal wall of diminishing returns. The computational physics required to get a 10% improvement in reasoning is becoming astronomical

Guest Caliber

11 / 20

Speaker B is introduced as a 'resident expert' but has no named credentials, institutional affiliation, or demonstrated track record provided in the transcript. The episode relies on an unnamed expert offering generic commentary rather than a practitioner who has shipped products or managed AI deployments at scale. Brian H., the loan portfolio manager, is the most credible voice but appears only at the end in a brief case study.

I've got my favorite resident expert here to help us skip the hype and look at what this actually means for your daily work
Brian is a commercial loan portfolio manager for a community bank

Specificity & Evidence

13 / 20

The episode includes concrete pricing ($5/$30 for SOL, $1.25/$4.25 for Spark 1.1), named models (GPT 5.6, Muse Spark 1.1, Fable, Luna, Terra, Codex), a specific benchmark criticism (SWE Bench Pro), and real infrastructure specs (14 gigawatts, 1 million token context window). However, most claims lack underlying data or citations - no source links for benchmarks, no case study metrics (how many hours Brian actually saved), and vague references to 'arena leaderboards' without specifics.

The API costs just $1.25 for input and $4.25 for output per million tokens
a 1 million token context window means you could upload a dozen full length books or an entire massive corporate code base

Conversational Craft

11 / 20

The host poses reasonable follow-up questions ('How so?', 'How do you generate a game world without a game engine?') and attempts to unpack technical concepts, but rarely pushes back or challenges claims. The conversation reads as agreeable dialogue rather than inquiry: Speaker B's assertions about diminishing returns, benchmarking failure, and future automation go largely unquestioned. No productive disagreement or skepticism surfaces; both speakers converge on consensus framing throughout.

Wait, how does an AI train an AI without a human grading the test?
Are we flying blind?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B53%
  • Speaker A47%

Most-used words

model21data13massive12today11human11openai10fable10coding9claude9models8luna8level8million8token8world8sure7

Episode notes

Today we’re covering the biggest AI stories of July 10th, 2026. OpenAI launched its GPT-5.6 family led by flagship Sol, landing just below Fable on intelligence benchmarks but ahead on agentic coding, while pairing the release with ChatGPT Work, a Cowork style platform, and a desktop app merge that finally brings OpenAI’s superapp ambitions into users’ hands, all priced to undercut the anxiety that comes with Claude’s usage limits. Meta shipped Muse Spark 1.1, an agent focused model that leads Opus 4.8 and GPT-5.5 on several reasoning benchmarks, opened it to developers at a quarter of rival pricing, and is already training a larger successor called Watermelon for later this year. Plus, Reve released version 2.1 of its 4K image model reclaiming the number two spot on Arena’s leaderboard while training on a tenth of the compute rivals use, and today’s community workflow comes from Brian in Ohio, a commercial loan portfolio manager who used Claude to build a prompt that extracts business mortgage filings from public county records, enriches them with contact information, and ranks prospects by priority in a color coded spreadsheet.

Full transcript

21 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Welcome Back to today's AI news, brought to you by 9X Productions. Imagine, um, the AI arms race is less about building a single fuel guzzling rocket and more about, you know, engineering a fleet of hyper efficient commuter trains. Today we're seeing the major labs trade massive, costly leaps for practical, everyday utility that fits right into your workspace.

Speaker B: Yeah, the whole landscape is, uh, it's shifting right under us, it really is.

Speaker A: So today we are exploring a massive shift in how frontier models are being built and priced. Digging into the very real economic impact this has on human workers and giving you practical, community tested workflows to make sure you stay ahead of the curve.

Speaker B: Which is, I mean, that's what everyone actually needs right now.

Speaker A: Exactly. And to help synthesize these massive shifts, I've got my favorite resident expert here to help us skip the hype and look at what this actually means for your daily work.

Speaker B: I am super excited to get into this one. It's a, uh, it's a profound pivot we're tracking today.

Speaker A: Total pivot.

Speaker B: Yeah. And we have a lot of ground to cover, especially with how the economics of these new models are fundamentally restructuring like professional services. But to understand the economics, I mean, we really have to look closely at the tools that just hit the market.

Speaker A: Right, let's jump right into the biggest news in the latest development space because we are seeing OpenAI and meta dropping highly capable models that are, well, they're basically competing on efficiency and workflow integration

Speaker B: rather than just, you know, raw brute force power.

Speaker A: Yeah, exactly, exactly. So OpenAI just released its GPT 5.6 family. We've got the flagship model called Sol, along with two other tiers, um, Terra and Luna.

Speaker B: And the performance metrics on Sol, they reveal a very specific strategy here.

Speaker A: How so?

Speaker B: Well, it lands just slightly below Fable on the AA Intelligence index. So it's not the undisputed heavyweight champion in every single category. But.

Speaker A: But it actually beats Fable on agentic coding.

Speaker B: Exactly. It beats it on agentic coding, which is huge.

Speaker A: Just to pause there, um, and make sure we are all on the same page when we say agentic coding. We aren't just talking about effing a chatbot to write a 10 line Python script.

Speaker B: Right? That is a crucial distinction. Traditional code generation is, you know, it's just autocomplete on steroids. Yeah, but agentic coding means the AI acts as an autonomous agent. You give it a high level goal and it navigates your entire code base.

Speaker A: It writes the code, runs the tests,

Speaker B: finds its own errors, debugs those Errors and just iterates until the feature actually works.

Speaker A: That's wild.

Speaker B: It is. OpenAI has essentially given Saul major upgrades when it comes to computer, uh, use, design capabilities and cybersecurity just to facilitate that kind of autonomous loop.

Speaker A: And the pricing is what completely caught me off guard. Oh yeah, you'd think a model with that kind of agentic capability would come with a massive premium. But the pricing, it just matches the older GPT 5.5, which is crazy. I know. We are talking $5 per million input tokens and $30 per million output tokens for SOL.

Speaker B: And for the smaller Luna tier, it's practically free, right?

Speaker A: Just $1 and $6.

Speaker B: Yeah. And San Odman made a telling comment about this launch that I think explains the strategy perfectly. He noted that every enterprise now is thinking about spend.

Speaker A: Oh, absolutely. The budgets are tightening.

Speaker B: Right. The market right now isn't asking for a model that costs $100amillion tokens just to get a, like a, ah, 2% increase in logic.

Speaker A: Because when a company is rolling out an AI assistant to 50,000 employees, that token cost scales exponentially.

Speaker B: Exactly. Enterprises want intelligence that is financially sustainable to deploy at scale, not just a wildly expensive, you know, brute force party trick.

Speaker A: They still threw a bone to the power users though, right? By introducing this new Ultra mode for top performance.

Speaker B: Yeah, they did. But um, here's the detail in the release notes. That totally blew my mind.

Speaker A: What's that?

Speaker B: SOL autonomously postponed the smaller Luna model.

Speaker A: Wait, let's unpack that. How does an AI train an AI without a human grading the test?

Speaker B: It's fascinating.

Speaker A: Is SOL basically generating synthetic data, feeling it to Luna and then scoring Luna's output?

Speaker B: That is the exact mechanism. Yeah. Historically, humans had to sit there and label data, which takes forever, or manually rate the AI's answers to teach it what good performance looks like. But now they are using the larger, highly capable SOL model as the teacher.

Speaker A: Wow.

Speaker B: Bolst generates millions of complex problem solving scenarios, feeds them to the smaller Luna model, and then evaluates how well Luna responds.

Speaker A: It's a closed feedback loop of optimization that is entirely machine driven.

Speaker B: Exactly. It's wild to think about.

Speaker A: And they are pushing all of this directly into our daily routines too. OpenAI introduced ChatGPT work, which is, I mean, it's their clear answer to Anthropic's Claude cowork platform.

Speaker B: Yeah, they're putting the Codex engine behind a much more accessible interface for everyday tasks.

Speaker A: And they're merging the Codex app right into ChatGPT's desktop app, complete With a built in browser and actual computer control,

Speaker B: it really brings their whole super app, um, vision into focus. You know, you are no longer interacting with a chatbot isolated in a browser window. It becomes an active agent living on your desktop, interacting natively with your local files and your like your whole software ecosystem.

Speaker A: But OpenAI isn't the only one making moves here. Meta just threw a massive counterpunch with the release of Musespark 1.1.

Speaker B: Yeah, they did.

Speaker A: And this is designed specifically for agent style tasks and computer use. And they've placed it behind a paid API.

Speaker B: And Meta is playing a completely different game when it comes to the numbers. Spark 1.1 boasts a 1 million token context window.

Speaker A: A million?

Speaker B: Yeah, which is vital for these agentic workflows.

Speaker A: Just to visualize that for a second for everyone. Um, a 1 million token context window means you could upload a dozen full length books or an entire massive corporate code base into a single prompt.

Speaker B: And the model won't forget the first page. By the time it reads the lab,

Speaker A: it retains the entire context.

Speaker B: Yep. And within that massive window, Spark 1.1 actually leads Opus 4.8 and GPT 5.5 on agent reasoning benchmark.

Speaker A: That's seriously impressive.

Speaker B: It can even split jobs across parallel sub agents. So if it needs to research five different companies, it spawns five separate processes simultaneously to get things done faster.

Speaker A: Okay, so given that it crushes GPT 5.5 on reasoning, I'm bracing myself for the price tag. Tell me it's not 50 bucks a million tokens.

Speaker B: Actually, it's the exact opposite.

Speaker A: Wait, really?

Speaker B: Yeah. The API costs just $1.25 for input and $4.25 for output per million tokens.

Speaker A: You're kidding. That is roughly a quarter of what the top rivals are charging.

Speaker B: Yeah, Mark Zuckerberg explicitly stated that they are promising pricing that undercuts the, quote, very extreme margins at rival labs.

Speaker A: Wait, hold on. If OpenAI is basically matching their old prices and Meta is slashing theirs to $1.25.

Speaker B: Yeah.

Speaker A: Are we hitting a ceiling in raw intelligence? Like, is this just a race to the bottom on price instead of a race to the top on capabilities?

Speaker B: That is the defining debate happening in the industry right now. I mean, it's not necessarily a ceiling on intelligence, but rather, um, a brutal wall of diminishing returns. The computational physics required to get a 10% improvement in reasoning is becoming astronomical. You can't just double the data anymore because we're actually running out of high quality human text.

Speaker A: Oh, wow. I didn't Even think of that.

Speaker B: Yeah. So instead of burning billions to get slightly smarter on a benchmark, these labs are optimizing the architecture.

Speaker A: Right.

Speaker B: They are shifting the engineering focus from, you know, how smart can we make this in a vacuum to how seamlessly can this integrate into a business workflow.

Speaker A: But let's bring this out of the server farms and onto your desk, because if you charge money for professional services today, this price war and this pivot toward everyday utility is going to fundamentally change your business model.

Speaker B: Oh, absolutely.

Speaker A: With AI becoming incredibly accessible, cheap, and capable at everyday tasks, the conversation naturally shifts to what this means for human workers.

Speaker B: The economic landscape is shifting beneath our feet, and there's a story making the rounds that perfectly encapsulates the reality of this. It involves Peter Hurley.

Speaker A: The headshot photographer.

Speaker B: That's the one.

Speaker A: For those who don't know, Peter Hurley is widely considered the world's best headshot photographer. He snaps portraits of everyone from Sofia Vergara to Wall street executives.

Speaker B: Right.

Speaker A: He even coined that technique called the squinch to make people look more confident on camera. And he teaches tens of thousands of photography students globally.

Speaker B: Well, he is telling his students something very blunt right now.

Speaker A: What's he saying?

Speaker B: He's saying, if you cannot beat the AIs in portrait photography, you will not get hired. Period. And it highlights a broader economic thesis about the concentration of demand. I mean, AI portrait apps are, for all intents and purposes, good enough for most people right now.

Speaker A: Right. If a professional just needs a picture for LinkedIn, an AI app can generate a highly realistic, well lit image for five bucks in ten seconds.

Speaker B: Exactly. Which means the pool of people willing to pay $300 for a human photographer shrinks dramatically. The middle of the market just evaporates.

Speaker A: The middle hollows out completely.

Speaker B: Yeah, but the critical part of this thesis is that all of the remaining demand concentrates at the very top.

Speaker A: Okay, explain that.

Speaker B: Well, a specific subset of clients will always pay a premium for that last mile of quality that only a top tier human can capture. You know, the specific tastes, the timing, the authentic micro expressions that an AI simply cannot synthesize.

Speaker A: Right, the real human element.

Speaker B: Exactly. So the thesis predicts that the top 1 to 5% of professionals in any field will actually see 10 times more demand.

Speaker A: Wow, 10 times.

Speaker B: While the rest are entirely replaced by automated systems.

Speaker A: And the way we search for things on the Internet is only accelerating. This winner takes all economy.

Speaker B: Definitely. Think about the mechanics of AI search. In the old days of a standard Google search, If you ranked 15th on the results page, For a local service, say graphic design or contract law, you still got business.

Speaker A: Sure, people would click through a few pages, compare a few options, and eventually call you.

Speaker B: There was a trickle down effect of clicks that kept the middle class of the Internet economy alive.

Speaker A: But now you ask an AI chatbot for a recommendation and it doesn't give you a list of 15 options to browse.

Speaker B: No, it doesn't.

Speaker A: It gives you one synthesized answer, the absolute best one. It's like the music industry shifting to streaming.

Speaker B: That's a perfect analogy.

Speaker A: Anyone can upload a song now, right? The barrier to entry is zero. Yeah, but because the algorithms are so good at identifying engagement, the top 1% of artists end up getting 99% of the listens.

Speaker B: Right.

Speaker A: The market becomes a brutal winner takes all arena.

Speaker B: It is the exact same dynamic moving into knowledge, work and professional services. You are no longer competing on generating a basic legal document or like standard marketing copy.

Speaker A: You're competing on taste, high level strategy and human relationship building.

Speaker B: Exactly.

Speaker A: So the bar is rising. What are you doing today to position yourself in that 1% economy? Because if surviving this transition requires outperforming the baseline, you need to use these very AI tools as intelligently and efficiently as possible.

Speaker B: Yeah, you have to elevate your own output.

Speaker A: Right. Which brings us to a really practical strategy for managing your AI resources without going bankrupt on token costs.

Speaker B: Yeah. If you are operating in this 1% economy, you are likely using a mix of models. But as we established, running a heavy expensive model like Fable for every single task will burn through your token budget instantly.

Speaker A: Especially if it's doing repetitive agentic loops.

Speaker B: Exactly.

Speaker A: So we are going to walk you through a specific setup being called the 60% token reduction strategy. The goal is to stop burning expensive Fable tokens on basic grunt work.

Speaker B: Yes.

Speaker A: It's like a construction site. You wouldn't pay your master architect their premium hourly rate to manually lay down bricks.

Speaker B: No, you definitely wouldn't.

Speaker A: You keep Fable out of the routine repetitive retries, use it for the architectural plan, the risky judgment calls and the final inspection, and let the cheaper worker models lay the bricks.

Speaker B: And the technical term for that is model orchestration. You position the highly capable Fable model exclusively as your planner and reviewer.

Speaker A: Okay.

Speaker B: Then you offload the token heavy loops. Things like web browsing, writing boilerplate code, or extracting data points to Codex. Or a lower cost Claude model.

Speaker A: Let's break down the step by step execution for this.

Speaker B: Sure. So if you are integrated into the Claude ecosystem, you open your app, make sure all your updates are installed, and Verify your weekly Fable limits just so

Speaker A: you don't hit a usage wall mid project.

Speaker B: Exactly.

Speaker A: Hm.

Speaker B: Then inside Claude code, you install the Codex plugin and initiate a section.

Speaker A: Okay. Simple enough.

Speaker B: Then you give Fable an explicit system prompt instructing it to use its codec skills. You constrain Fable so it only acts as the strategic planner at the beginning of the prompt and the final reviewer at the end.

Speaker A: Got it.

Speaker B: You command it to send all the basic iterative coding, browsing and extraction tasks to Codex.

Speaker A: And what if someone doesn't have a Codex plan set up? Can they still do this?

Speaker B: Oh yeah. You launch Claude code directly in your terminal. Use the command model to manually select a lower cost worker model.

Speaker A: Nice.

Speaker B: Then you use the Advisor skill to set Fable as the overarching advisor. The reason this saves up to 60% of your tokens is because the iterative loop, the AI reading a page, failing to find a data point and trying another page, is what eats massive amounts of context.

Speaker A: Ah. Ah. Right. So you want the cheap model making those mistakes.

Speaker B: Precisely.

Speaker A: You are applying human strategic oversight to automated labor. It makes perfect sense.

Speaker B: Yeah, it's a game changer.

Speaker A: Alright, now that we've covered the massive shifts in models and economics, let's sweep through the rest of the day's critical AI ecosystem updates. We are going to look at these through that exact lens of efficiency and the struggle to measure it.

Speaker B: Starting with image generation. Roove rolled out version 2.1 of their native 4K image model.

Speaker A: They retook the number two spot on the arena leaderboard.

Speaker B: Right? They did. But the real technical feat here is that they achieved this while training on less than a tenth of the compute used by their major rivals.

Speaker A: A tenth? That's insane.

Speaker B: I know. It also features element level editing, giving users highly granular control over specific objects in the generated image.

Speaker A: We see the same push to get more out of less compute in the physical world too. Like, look at Robbient's new Lingbot world too.

Speaker B: Oh, this is a huge breakthrough in embodied AI. Yeah, they released a new world model capable of generating every single frame of a simulated environment in real time.

Speaker A: And completely without the use of a traditional 3D engine like unity or. Unreal.

Speaker B: Exactly.

Speaker A: Wait, how do you generate a game world without a game engine? What is it actually doing under the hood?

Speaker B: Well, instead of a computer calculating geometric polygons and bouncing light rays off of them, the AI is functioning as a massive video prediction engine.

Speaker A: Uh, okay.

Speaker B: It understands the physics of the world so well that when you input an action like walk forward, it literally dreams the next visual frame into existence based on its understanding of how light, gravity and momentum should look.

Speaker A: That is mind blowing.

Speaker B: It really is.

Speaker A: But you know, with AI predicting physics on the fly and autonomously coding entire applications, it begs the question, how do we even measure if it's doing it right?

Speaker B: Which brings us to the drama in the benchmarking world today.

Speaker A: Right. OpenAI just published research retracting its endorsement of the coding benchmark known as SWE Bench Pro.

Speaker B: Yeah, benchmarks are the standardized tests the industry uses to measure progress. And OpenAI's research revealed that nearly a third of the complex software engineering tasks in the SW Bench Pro benchmark actually had underlying issues.

Speaker A: Wait, a third? If OpenAI is pointing out that a third of a major coding benchmark is flawed, how do we even know which model is actually the smartest anymore? Are we flying blind?

Speaker B: We kind of are. It's like trying to grade a college level quantum physics examiner when the teacher only has an 8th grade education.

Speaker A: Wow.

Speaker B: The AI is coming up with coding solutions so complex and novel that our own human made rubrics don't even know if the answer is technically right or wrong. That is the core dilemma of frontier AI evaluation right now.

Speaker A: So the test is broken.

Speaker B: Yeah. We are reaching a threshold where creating a test that accurately measures the intelligence is harder than building the intelligence itself.

Speaker A: That is slightly terrifying, but even if we can't measure it perfectly, the race demands more power to keep scaling. Which leads us to some m massive hardware news out of Meta. They are reportedly starting the manufacturing of their in house Iris AI chip this September.

Speaker B: And by bringing their silicon development in house, Meta is aiming to double their total computing capacity to 14 gigawatts by 2027.

Speaker A: We hear the word gigawatt and usually

Speaker B: just think of back to the future 1.21 gigawatts.

Speaker A: Exactly. Give us some context on how much power 14 gigawatts actually is.

Speaker B: Okay, so a single gigawatt can power roughly three quarters of a million homes.

Speaker A: Okay.

Speaker B: 14 gigawatts is essentially the entire power grid capacity of a medium sized country like Argentina.

Speaker A: Good. Greece.

Speaker B: They're building infrastructure on a nation state level. The reason they need that much power is because training the next generation of models like Meta's upcoming Watermelon successor requires coordinating hundreds of thousands of chips running at maximum capacity for months on end

Speaker A: from the 14 gigawatt data centers. Let's bring it back to your personal laptop. Anthropic introduced a new Reflections dashboard today.

Speaker B: Yes, this is a great tool.

Speaker A: It helps users analyze exactly how they are using claude, offering personalized suggestions for skill creation and usage habits based on your own data.

Speaker B: It's the perfect tool to help you implement that token reduction strategy we just discussed, tracking where your budget is actually going. It closes the loop nicely, giving you visibility into your own orchestration habits.

Speaker A: We love theory, but we love practice even more. So let's wrap this deep dive up by showing exactly how everyday professionals are applying all this technology in the real world.

Speaker B: Yeah, let's do it.

Speaker A: Today's community AI workflow is a perfect illustration of bringing these high level concepts down to ground level. It comes from a reader named Brian H. Based in Ohio and Brian is

Speaker B: a commercial loan portfolio manager for a community bank. The workflow he submitted uses what he calls a commercial mortgage identification prompt inside of claude.

Speaker A: Let's break down the mechanics of this prompt because it does an unbelievable amount of heavy lifting.

Speaker B: It really does.

Speaker A: Brian starts by taking massive files of public recorded mortgage data that he exports from local county recorder websites. We're talking huge, messy, unstructured data dumps.

Speaker B: And he feeds that raw data into Claude. The AI first parses the document to extract all the business and commercial filings from that mess, ignoring all the residential noise, and organizes them neatly into a structured Excel file.

Speaker A: That alone saves hours.

Speaker B: Oh, for sure. But the automation doesn't stop at extraction.

Speaker A: This is the best part. The AI then autonomously goes out to the web and searches for contact enrichment information.

Speaker B: Right?

Speaker A: It finds the principal's names for those businesses, their corporate emails and their phone numbers. Then it actually ranks these prospects based on specific lending criteria Brian gave it and color codes them in the spreadsheet.

Speaker B: And for the grand finale, the AI generates an executive summary tab.

Speaker A: Oh, I love that.

Speaker B: It lists the top 20 grand tours by dollar volume and provides Brian with strategic suggestions for targeting them, including who to call first and the best timing for the outreach based on the data.

Speaker A: Brian is a perfect example of exactly what we were talking about in Section two. He is the master architect and he's got Claude laying the bricks.

Speaker B: He is the absolute embodiment of someone operating in the 1% economy. I mean, Brian isn't sitting around letting AI replace his job at the bank, right? He is using model orchestration to do the data extraction, enrichment and analysis work of what used to take an entire team of junior analysts. He is elevating his own value from a spreadsheet cruncher to a high level strategist.

Speaker A: It's such a powerful way to frame your career. Right now, you either manage the AI or you compete with it.

Speaker B: It is empowering. Um, it also raises an important question and this is what I really want to leave you to ponder today. We noted earlier that OpenAI's new Sol model autonomously post trained the smaller Luna model using synthetic data and feedback loops. If AI models are now learning to act as the teacher and train each other and tools like Ryan's are automatically executing entire multi step corporate workflows from extraction to strategy, how long until the role of the human orchestrator we talked about today becomes fully automated too?

Speaker A: Ooh, that is a heavy thought provoking place to leave it. Yeah, the architect might eventually find the building designing and managing itself. You definitely want to chew on that one. Make sure you subscribe to stay updated on all of these shifts. We are tracking them as they happen and if you enjoyed this deep dive please take a second to rate the show with five stars. It really helps us out. We'll see you back for tomorrow's episode.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Greatest Hits: Cash, Clients, and the Secrets Behind Growing BusinessesThe New F*Word · features Speaker B82 / 100
  • Why AI Agents Aren't Ready to Manage Your Crypto Portfolio Yetthe un# podcast · features Speaker B64 / 100
  • Greatest Hits 2026: Masterclass in Risk, Tech, and Human BehaviorRisk Management: Brick by Brick · features Speaker B55 / 100
  • Why AI Adoption Depends on Hiring the Next Generation of TalentInsuring Cyber Podcast · features Speaker B48 / 100
  • The One Law No Leader Can CheatThe Business of Alignment · features Speaker B47 / 100
  • Build $5K per Week AI Email FunnelsAI Paycheck · features Speaker B44 / 100

More from Today’s AI News

All episodes →
  • GPT-5.6 Restricted Like Fable, China Stole 28M Claude Exchanges, AI Avatar Built 200K Followers45 / 100
  • OpenAI Built Its Own Chip in 9 Months, Fable 5 Is Coming Back, AI Is Trying to Kill the Common Cold55 / 100
  • Claude Joins Your Slack as a Coworker, Meta’s $299 AI Glasses Launch, AI Gets a Biology Language40 / 100
  • Grok 4.5 Rivals Opus at a Fraction of the Cost, ChatGPT’s Voice Gets GPT-Live, Seedream 5.0
  • Meta Cooks Up a Real Image Model, Beijing Eyes AI Export Controls Too, DoorDash’s AI Report Card
Explore the best B2B AI & Data podcasts →
All Today’s AI News episodes →