
Stacked Podcast · 2026-08-09 · 31 min
Key moments - from our scoring
Substance score
25 / 100
Five dimensions, 20 points each
Anthropic is shifting Claude Code from human-approval mode to automatic classifier-based decision-making for code execution starting August 14. The change reflects data showing humans approve dangerous commands 90% of the time while the AI classifier catches 13.6% of risky actions - making it roughly nine times more effective at security decisions. This aligns with recent industry trends where AI safety systems increasingly take precedence over human discretion, comparable to how autonomous vehicle adoption will eventually phase out human drivers. The conversation explores this broader pattern of automation replacing human judgment in high-stakes decisions. Separately, Amazon's commitment to energy-intensive AI infrastructure is crystallizing through a 7.65 gigawatt dedicated gas power plant in West Texas that bypasses the grid entirely, while Databricks has published practical cost-management strategies through their Omni Agent framework, demonstrating how dynamic model routing and token optimization can reduce AI coding expenses by 30-50% without sacrificing quality.
Claude Code defaults to automatic approval mode starting August 14, 2024, for Pro, Pro Max, and Team users. A classifier will automatically approve or block actions instead of requiring human permission for each command.
The classifier catches 13.6% of dangerous commands while humans only caught about 10% of them - making the classifier roughly nine times more effective at identifying risky code actions.
Omni Agent is a meta-harness that intelligently routes coding tasks to the cheapest model that still meets quality requirements. Databricks achieved 30% savings from dynamic routing and nearly 50% from reducing generated tokens in their testing.
Amazon is building a dedicated off-grid gas plant because energy is the limiting factor for AI scaling - not compute chips or talent - and controlling its own power supply gives it independence from grid operators and regulatory constraints.
Anthropic's Opus model shows zero successful prompt injection attacks in testing, while OpenAI's GPT-5.6 SOL has a 19% success rate, according to comparisons shared by Codex team members on Twitter.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains scattered technical observations (Claude's auto-mode approval rates, Databricks' token optimization tactics) but drowns them in extensive filler: tangential stories about construction workers, Iceland travel anecdotes, Dubai speculation, and lengthy banter about Canadian geography. The actual technical content - that humans approve dangerous commands 90% of the time versus classifiers at 10%, or Databricks' 30-50% cost savings through routing - is thin and often restated. Most minutes add entertainment rather than learning.
So humans, uh, approved way more dangerous commands than people do because we are just, you know, three coffees deep.
Everything's bigger in Texas.
The frameworks recycled here are standard industry talking points: AI safety via classifier approval (common post-incident industry practice), vertical integration as a competitive move (already established pattern with chip manufacturing), efficiency arbitrage between models (now mainstream optimization). The observation that humans approve dangerous commands at higher rates is presented as novel but has been routine in AI safety discussions. No contrarian angles or first-principles thinking emerges.
It's just more vertical integration, right? It's like anthropic from a few days ago, trying to build their own chips.
most day to day coding doesn't require mathematical proofs or novel security insights.
This is a two-person podcast with no external guests. The hosts appear to be generalist tech commentators discussing news rather than practitioners who have built at scale. No evidence they have shipped production AI systems, managed large ML infrastructure costs, or led engineering teams. They are reporting on Anthropic, Amazon, and Databricks' public announcements without insider perspective or direct operational experience.
Nick, let's kick off with Claude code.
Speaker A: Yeah, it was great.
Some concrete numbers appear: 13.6% detection rate for dangerous commands by humans (implying ~90% miss rate), Amazon's 33 million tons CO2/year gas plant permit, Databricks' reported 30% savings from routing and 50% from token reduction. However, these claims lack context - no breakdown of which commands, under what conditions, or validation of the Databricks figures. Most discussion is vague: 'larger gas power plant,' 'significantly less risky,' 'more efficient.' Travel and geographic claims are anecdotal with no supporting data.
So the human review caught 13.6% of dangerous command.
published savings are more than 30% from routing and almost 50% from fewer generated tokens
The hosts ask surface-level questions and rarely push back or dig deeper. When they do ask follow-ups (e.g., 'what's going on here' on model safety), they move on without forcing specificity. Entire segments devolve into tangential storytelling (Iceland prices, Canadian landscapes) with no effort to redirect. The final comment-reading section adds zero insight and pads runtime. No genuine disagreement, no adversarial questioning, no probing of weak claims. The tone is friendly banter over rigorous inquiry.
Isn't that crazy? Um, did Codex have a comeback for this? Because this is outrageous.
What's going on? But I think the big tail, Doctor, is we get it right, we get it wrong 90% of the time.
Computed from the transcript - who did the talking, and the words that came up most.
Claude Code stops asking your permission next week. From August 14, new sessions on Pro, Max and Team default to auto mode, where a classifier reviews every tool call and blocks anything irreversible or destructive instead of prompting you. The reason is uncomfortable: across a 1,053-case study, human review caught just 13.6% of dangerous commands. Nick and Jack dig into what that means for anyone who has ever hit "bypass all permissions" at 1am. Also in this episode: the prompt injection numbers behind the Boris vs Tebow spat on Twitter (19% success on GPT 5.6, 0% on Opus), Amazon building what would be the largest gas power plant ever constructed in the US on a West Texas site that answers to no grid operator, why energy is becoming the real constraint on AI, and the Databricks finding that the AI coding bill is growing exponentially, plus the "efficiency frontier" they use to arbitrage the cost of intelligence. Watch on YouTube: Step-by-step roadmap to $25K w/ AI: Main channels: youtube.com/@nicksaraev & youtube.com/@Itssssss_Jack
Transcribed and scored by The B2B Podcast Index.
Speaker A: So Claude code stops asking permission from next week because humans have approved 800 dangerous commands. We have news with Amazon, they're building a 7.65 gigawatt gas plant in Texas to run off off grid AI campers. And then we've got some interesting stuff that's kicking off with databricks. Nick, let's kick off with Claude code. So for those that don't know, basically new Claude code sessions on Pro and a few of the details which I'm going to pop in a second here. Uh, basically Pro Max team default to auto mode from 14 August. So a classifier, basically N is going to approve or block actions instead of you. Thoughts, feelings, reactions. What does this mean?
Speaker B: Yeah, it's interesting. This comes right on the back of large language model provider companies just reporting non stop data breaches and cybersecurity concerns. Uh, basically they're pulling control away from you. Not pulling control necessarily because you'll still have the ability to, you know, choose a bypass permissions mode. But, um, they're pulling control away from you and then they're giving it to a classifier that they've determined is significantly less risky. So humans, uh, approved way more dangerous commands than people do because we are just, you know, three coffees deep. It's 1am M in the fricking morning and we just want that goddamn thing to work. So middle of the night, Claude's like, hey, can I sudo-rm dash rf your entire hard drive? And you're like, of course, that sounds great, let's do it.
Speaker A: What a brilliant idea. And I think the cause. Sorry, Nick, you're gonna say the reality is even if you're diligent and know what you're looking at, you just, just get it done. And you just. There's this assumption baked in that Claude wouldn't ask for something that wouldn't actually work. So I think in theory, human approval is one thing, but when you're proving 20 things a minute, it's, it's probably too much. So essentially what is happening and when is it happening? On the 14th of August, you're in five days head start. Because you're on the Stack podcast right now. New sessions on Pro Max and team are uh, going to start in auto mode. Separate classifier reviews every cool call and blocks anything irreversible, destructive or aimed outside your environment. So if your AI is trying to escape containment and initiate Skynet, no thank you. Not possible. Not possible with this new setting. Allegedly.
Speaker B: So yeah, if you try and like any other YouTube channel or subscribe to any other YouTube channel, it will automatically block you. That is, um, that's our promise.
Speaker A: That is our promise. And not only that, it will take. It will take irreparable damage against every. Everything you earn if you try that. Apparently I've not tested it, but apparently it's in there now. The 1053 test. Study this, to me, Nick, was really interesting. What actually kicked this off? So the human review caught 13.6% of dangerous command. So essentially 9 times out of 10 US meager carbon based life forms were just going, yep, yep, yep, yep, yep. And we were getting it wrong about 90% of the time, which is outrageous.
Speaker B: Sorry, do you know what else is outrageous? Sorry, I was just looking at this. Um, Boris Trney and this other guy on the Codex team are just brutally mogging each other right now and just like shitting all over each other on Twitter, right? And he posted this the other day, which is the prompt injection attack success rate for Opus, sonnet and Fable versus GPT 5, 6 SOL. So they show that 19% of the time, um, access requests via prompt injections work on GPT5.6 SOL, whereas on Opus 5 they happen 0% of the time. That is crazy. Shitting all over Codex, dude.
Speaker A: Isn't that crazy? Um, did Codex have a comeback for this? Because this is outrageous.
Speaker B: Nothing like explicit, but Tebow. Okay, I don't know if you guys know, but Boris turning kind of like made cloud code, right? Like, he's the guy that's in charge of cloud code. Um, Codex's version of that. Uh, his name is Tebow. And I don't know exactly where on his Twitter this is, but it's somewhere. They had like a kind of a verbal sparring match where, you know, I think Boris was like, hey, by the way, the cloud code team is hiring if you want to come on. TBOO was just like, no, I think I'm okay. And then he reset rate limits for everybody. Um, which is just so funny. So then everybody on like, the Codex camp had like totally reset rate limits. Tebow, I don't work at Anthropic. Boris, we are hiring. If you want to, Tebo, I'll reset Codex limits for no reason. So anyway, that's crazy. They're fighting and I thought I'd throw that in there. You guys probably find that good.
Speaker A: That is very good. But it's also like, it's funny, right? Because I think this graph is interesting to me. But I'm also wondering what, where the other models lie on. On that graph. Yeah, Fascinating. So, obviously, it seems that what also surprised me is Opus caught more than Fable. M. When I thought. I thought Fable. Thought Fable was the super safe hanging in the background guy.
Speaker B: Ooh. Yeah. I thought that was, like, the advanced model. What's going on here? Why am I paying more money for a model that fails 0.28% of the time?
Speaker A: What's going on? But I think the big tail, Doctor, is we get it right, we get it wrong 90% of the time. And, uh, this classifier gets it wrong 10% of the time. So it is essentially nine times better than you. And this is going to be the argument that we're going to see played out everywhere right now, Nick. We're going to see it in cars. Why do you meat sacks get to drive around cars? Our grandkids will not believe it. They'll say, I cannot believe they used to let you drive around when these extremely safe, environmentally friendly vehicles were running around, not killing people. Uh, so we're gonna see this all the time where we're just saying, yeah, we're gonna take this off your hands. Thanks for keeping the seat warm, but we'll do this from here.
Speaker B: It reminds me of, um, those old construction workers, um, hanging off of. I don't know, like, is it the Golden Gate Bridge? I don't know if you've seen these photos, but there's, like, a bunch of construction workers that are just sitting on top of a freaking skyscraper as it's being developed. And, like, at the time, this was considered entirely normal.
Speaker A: Crazy.
Speaker B: These guys aren't strapped in, man. They're just. They're literally just having a smoke, eating a sandwich, hanging on top of what is probably New York City. Yeah, I think that has to be because that's Central Park.
Speaker A: That's one of the most iconic photos in New York, I think, to be
Speaker B: honest, and think about. That's what people are going to be thinking about when they see you driving your BMW. They're going to be like, whoa, I can't believe Jack used to be able to do that. He's hanging off the metaphorical equivalent of the skyscraper.
Speaker A: So wild. He's so wild and crazy. Yeah. Isn't that insane? I think it's a testament to how fast societal expectations move. Whereas I think everyone growing up now grew up in slightly different environments than our, uh. Our, uh, guys were like, basically tightroping on thin metal rods, thousands of feet in the air. It's crazy. So very, very interesting, dude. I think it makes sense. It's a good move for defenses I haven't seen as many hey Claude deleted my entire hard drive posts to be fair, uh, in a few months. So things must be generally just anecdotally seems to be moving in the right direction.
Speaker B: Yeah, that's true actually. I mean I'm seeing more hey Claude, delete my entire civilization posts. But I'm not seeing hey Claude deleted my hard drive posts. You know. Yeah, I'm seeing like, you know, the big companies are talking about how they're getting data breaches and how they're doing cybersecurity issues. But like that's. We're not actually seeing a lot of people bitching and complaining about it, deleting their hard drives like they did initially.
Speaker A: Yeah, well, it's. Dude, I think if it's nine times better, like I get the whole idea that humans should be the ones approving it, but if it's a got like in reality, because that's where we live, not the theoretics, in reality it is nine times better.
Speaker B: So.
Speaker A: And if you don't like by the way, you can move it by the way, it's just that the chat will begin. It's like that nudge theory we always talk about. It just starts you in auto mode.
Speaker B: You can always do that. Then after a while though, you'll just do it anyway because you'll be like, eh, that's better.
Speaker A: You know, I've seen it before. I like, I think people, uh, just go bypass all permissions. I've seen that so much. It's just bypass all permissions and just get it done.
Speaker B: Yeah, here's my Social Security number, Claude.
Speaker A: Dude, it might could fry you if it went a complete rogue. That's crazy.
Speaker B: You want to chat about Amazon?
Speaker A: Let's chat about Amazon. So they're building a gigawatt gas plant in Texas to run off grid AI campus, which is cool. So in a nutshell, Amazon is behind what would be the largest gas power plant ever built in the US on a site it now confirms earning in West Texas. So the campus runs behind its own meter. So the power never enters the Texas grid and the plant answers to no grid operator. Now the permit ceiling is 33 million tons of COT per year against a company that target of NET 0 by 2014.
Speaker B: Yeah, I mean, it's kind of interesting. Um, obviously this is kind of a politicized headline. Thank you Claude, for your totally nonpartisan viewpoint. But yeah, I mean, it is kind of weird to have, you know, a gas plant that will have significant supposedly, uh, incursive cost to the people that does not in any way shape or form feedback to the people of that state. Um, I mean, I'm sure there's more going on behind the scenes to do with, like, signing deals and stuff like that, but. Yeah, I mean, like, it's. It's not connected to the Texas grid, man, which is kind of wild. I guess they have enough of their own power, but, uh, yeah, well, I.
Speaker A: Because I think the key thing, the interesting conversation here is building your own power station is not interesting because, for example, in the uk, we have, um. Basically people misunderstand about energy companies. Energy companies basically earn from the meter to your front garden. And that's basically it. The rest of the distribution, we have these distribution network operators, these gas transporters, and effectively you're locked into the grid. Uh, it's unlawful to disconnect the consumer in the uk. But what's interesting here is that Amazon and I like this move in some. Some senses, because what they're effectively doing is building their own power supply, right? That we will generate our own energy. And I think that's quite good, uh, foresight, because energy is going to be the biggest constraining factor for many of the AI developments that we're going to see is we have the compute, we have the chips, we got the people, and then we got the coffee. But what we don't have is some beautiful electrons. Can we get some electrons, please? And think about that. It takes some time to build these things, so I would not be gobsmacked if we're seeing more energy investment in Europe as well. It's all connected. Like, they can even send it in various different locations. So I think it's a good strategic move. And why not? Like, why not just give it to the people? Hey, this can't even. This can't power your Netflix tv, bro. But what it can do is power all the stuff that we want to build. I think it's very interesting that.
Speaker B: Yeah, yeah, um, yeah, I mean, fair enough. It's just more vertical integration, right? It's like anthropic from a few days ago, trying to build their own chips. Well, it's like, what's going to power the chips? Well, it's like, oh, we'll build our own power plants too. It's like, um, Sam Altman and his work with nuclear power plants, right? Achieving criticality. Um, check this out. If the project emitted that much CO2, 33 million tons, it would be the largest single source of pollution in the entire United States, emitting more greenhouse gases than the country's largest coal plant. Worth noting that companies rarely emit as much as their permits allow. So that's worth kind of considering. Probably not actually going to be emitting 33, but if it did, it would be the largest polluter in the entire country. Whoa. Just makes sense that it would be in, In Texas, baby. That's looking, that's looking like some, like some Texan. What do they. They always say it's bigger in Texas.
Speaker A: Everything's bigger in Texas.
Speaker B: Bigger. Yeah. Yeah. I think it's also a good, uh, move as to just like, Texas's dominance in business in the United States after, like, overregulation in California. Like, there's reasons why everybody's trying to build this shit out there now, right? There's reasons why they're erecting, like, that golden statue of that guy, um, you know, as you drive into, uh, what will soon be, I think, Elon Musk's big new power plant. Like, they're, they're, they're turning it into, like a new tech haven, which is neat.
Speaker A: They are, they are. And California is overregulating so much. Millionaires and billionaires are leaving. It's crazy.
Speaker B: It.
Speaker A: But it's bizarre to me, Nick, because the economics are pretty clear on this, actually. And the effect that has. The thing is, though, people like, it's a whole economic theory, right? Like theory of incentives. And when you make it unincentive, you know, you can see based on the numbers where everybody's going from California to Texas and various other locations. But the upshot of that is when they move their companies to Texas, that's. That then becomes the kind of center of innovation and technology. And in some ways, one of the places to be is actually on my future hit list of, of countries I might want to live in for the big stakes and the power plants. Now we're going to have my own. We're going to have power plant access in Texas. It's becoming more appealing by the episode.
Speaker B: Dude, you should come out, uh, to North America, Jack. I will, bro. I've been, you've been teasing me with this Canada visit for a hot minute now, but, like, it's very different out here. I was just, um, chatting with a friend of mine in, uh, in Europe, and he was just like. He sent me a photo outside his front door somewhere in Belgium. And like, you know, it was, it was charming and cute and there was so much history behind it and stuff like that. And then I looked out and I realized, like, that guy has probably never seen the fricking Horizon, like, whoa, we have so much space here. Like, I look outside and I see massive mountains, Jack, that go literally halfway up my field of view. I see, like, lakes. I see, like, more land and arable, you know, stuff, uh, to play around with than I think most Europeans would ever get their entire lives. And I say this as somebody whose family was European. Every time they go back, they're like, we feel like we can't breathe out here. It feels like the past, not the future. I don't mean to say this to shit on any Europeans, just to be clear. Look at, look at this. This is just a giant field in the middle of nowhere. In, like 10 years, five years, whatever, there's going to be huge power plant that will be legitimately producing more energy than probably like, freaking a quarter of the entire European Union. I mean, it's insane to just think about, right?
Speaker A: But it is, it is.
Speaker B: That's where you build.
Speaker A: Honestly, I appreciate you got mountains neck, but we do have bottle caps that don't detach from bottles. Where are you going to find that? Okay. And as a preeminent expert on Europe, having lived in Europe for so long, it's different. It's, it's. It's a little. It's these little mini cultures. I'm a very big fan of Dubai. I think Dubai in the future is going to be, like, unbelievable. I will say this. I was in Montenegro recently, and, um, I did like seeing the mountains. There was just something about it as well. And. As well as in Iceland, actually.
Speaker B: And you went to Iceland?
Speaker A: Yeah, several times. It's very, very, very cool that it's actually, I think it's the third most expensive place on the planet to buy things. You will get sticker shock if you're not prepared. Like, uh, for example, if you want a, I don't know, a vodka and Coke or something like this. Like, I. And I know this example specifically because the guy behind me said it on the flight back is like, I just paid €20 for a vodka Coke. And it was funny. But honestly, everything is, like, way more expensive for lots of reasons. Like the heavily unized, you have to import everything. But there, you know, you're breathing super fresh, uh, air as well. Just like next to the mountains. And it's very disconnected. But I think what's going to be interesting is, well, and then the kind of geopolitical landscape in the future is. One of the things I like about Iceland is it's like geothermic connections. Like, you can go into baths that are just heated by the earth. And I think that's really cool for like zombie type scenarios and things like that, um, which may happen. We'll cover them here if they do, but I just think it's super interesting. And Texas, like you say, they have so much land and so much optionality. So it's cool. But I've got to do. Did I have to do Canada before I even say the word Texas? I will come to the west coast of Canada.
Speaker B: Can't wait.
Speaker A: We will do this in person and we will, we will rock some worlds
Speaker B: and have a good time tomorrow on the Stacked podcast. Zombies invading the United States of America and more here on Stacked.
Speaker A: And we'll have the HDMI graphics to back it up, guys. I can promise you that. And um, you will never, you will not see better zombies anywhere else. Let's talk about the AI coding build mix. So this one's a really interesting one because I think this is a good case study that we can have a look at. So this to do with data bricks. So they say that AI coding bill grew exponentially. And uh, here's some stuff that happened. So they basically published an internal playbook it uses to stop AI coding spend outrunning the productivity that it buys. So they basically found that the cost started to grow exponentially across a multimillion line code base. And essentially there were four levers which I think are worth discussing here. Um, uh, cheaper and open models, dynamic routing, developer spend, visibility and cutting token overhead. And the published savings are more than 30% from routing and almost 50% from fewer generated tokens, both measured internally. So this is a company that's actually what we'd call doing the do. And they said hey guys, I don't like the look of these token bills. How can we bring it down? And this is just going to be classic business, like if it's expenses, they will look to cut them down. And the first thing they're going to say is did someone say the word deep seq v4 flash. Am I hearing things?
Speaker B: Yeah, it's pretty neat. Um, they also open sourced their meta harness called Omni Agent. And the whole idea being um, Omniagent is much more efficient. And what they really care about is the efficiency frontier, which they define as the set of models that have the best price point for a given level of intelligence. They say the same sorts of things that you and I have been saying on the Stack podcast for the better part of the last few weeks, which is they might be one of the seven day to day coding. Oh yeah. I mean thanks. Databricks. Make um, sure to subscribe. Most day to day coding doesn't require mathematical proofs or novel security insights. So what matters in AGRIA is the cost of the models that meet the quality bar for typical software engineering work. And so what this omnigent does is it automatically does like intelligent rerouting, um, two models that it deems like smart enough to do the thing that it is being asked for and then they arbitrage the cost of intelligence basically between one model the next. It's very efficient. It's also very effective. What I like about these guys is they've translated AI from like this like big, super, like, whoa, incredible galaxy brain thing and then made it really practical. You can think of this as like the, the, the screws and nuts and bolts of applying AI to a company. Uh, and they actually publish it all right over here on a, uh, blog post called Managing AI Coding Costs at Scale.
Speaker A: That's interesting, dude. I mean this is, you know, where's the bottleneck now? Is it actually intelligence or is it efficiency with these businesses?
Speaker B: Yes.
Speaker A: Do you think
Speaker B: while I was, uh,
Speaker A: it was Nick's other European friend, guys, I thought I was the only one, but apparently there's more European friends out there. Um, so Nick, is it, what's the bottom line? Is it intelligence or efficiency? What do you think?
Speaker B: Uh, well, I mean, it depends on for what work, right? Like, are you looking to solve, uh, the meaning of life? Yeah, if you're looking to solve the meaning of life. I mean obviously it's intelligence, right? But if you're looking to put together a consumer SaaS product, like I'm looking to put together SaaS products personally.
Speaker A: Yeah. Did we do a lot of Fable 5 goal and we went out on the, on the boat, um, we said that we gave it the biggest challenge and who knows what happens when we get back. So we'll see. But it's like, I think there's times to cut and there's times to invest some problem. Like it's like for example, if you have standardized classificatory work that you just know like, hey, we consistently spend X tokens on Y problem dude, you can easily just start to ratchet down to increasingly dumber models until you start to get any kind of like reliability detriment or performance detriment that makes sense. But there are some problems where you want to spend like you're crazy because the, any, even a 1% increase in performance is worth it. So I think this is the kind of duality that the business need to think of intelligence performance and then the optimization at the same time.
Speaker B: Yes. The duality of man, just as we have the duality of AI. You know, the way they're doing this meta harness thing to kind of describe me kind, uh, of interests me because it also describes the way that AI models work now. So what's really funny is it's all just like a quote unquote mixture of experts. They're really just diversifying. If you think about it. If we zoom way in on Claude here, right?
Speaker A: Can we zoom in a bit further, dude? I want to see that even larger.
Speaker B: Yeah, I want to see the atoms behind Claude's little butthole emoji. Thank you. Uh, if you, um, zoom in, think about what's actually going on under the hood, Claude is like a mixture of experts model, right? And so what it's actually doing is it's like sub selecting different answers based off, like small little sparsity of neurons. Right? So it selects the best one that goes back into Claude. Then if you zoom out a little bit, what's happening is in addition to the model, there's a harness layer. And so there are harnesses out there that actually work with multiple models. Now, right? Open code can do the, you know, you could, you could do Codex, you could do Claude and then check what's happening one level up. Now there's a meta harness which selects the right harness, which selects the model, which selects the specific mixture of experts. You can go up. I bet you in a few months there'll be a meta meta harness that. A meta, meta, meta meta harness. And then what's cool is you're the developer up at the top. What's happening is you are not like, uh, rather than you going down, like you're just going up, you know, like the next generation, the developer is going to be a little bit higher. The next generation developer is going to be a little bit higher. And we'll just be selecting kind of the most optimal approach. Just like a giant search tree all the way down until we get to, like, exactly what you want. It's really, really cool if you think about it mathematically, philosophically, this is cool.
Speaker A: RIP developer didn't get a profile photo now as well, but it's like, it's interesting as well because I think you could run so many experiments with this. Basically. Um, I know we're talking about experiments. Next camera can't even handle these experiments, guys. That's how powerful. This just showed us that that is a level of detail you get on the site, podcast but you could very easily have a system where effectively you can run different tests and it kind of evaluates. But if you can actually assess the performance and you can have benchmarks for that performance, we call them evals for short, you could start to codify actually and have this like the intelligence layer in that meta harness. So it actually gets better over time. I said, well actually creative tasks tend to do better when I send it to harness X with model Y. And then we just optimize that and then you can apply cost efficiency metrics like how could we reduce this down? So I think meta harness engineering is going to be super interesting. But what's going to be more interesting that right now, Nick, are the comments on the SPAC podcast. Speaking of a duality of map, should we pull them up and see what the people are saying?
Speaker B: Yes, of course. This has to be one of my favorite parts of the podcast.
Speaker A: One bazillion. Dude, I saw Spider man yesterday.
Speaker B: Yeah, it was great. How we doing? How we doing on the Spider Man? Why don't we start the Q and A off with the Spider man if
Speaker A: the people want to see it. Yeah, I think you'll need to. Let's have a look. Full screen view on this. Let's have a look at this. I'll get, I'll get to Spider man after our comments.
Speaker B: Let's see what we got. What do we got? You want to, you want to kick us off here while I plug in my camera?
Speaker A: Let's do it. Um, it's just not the same without you, Nick. What's your guys video production pipeline? How do you come up with YouTube video ideas? Nick mentioned that he doesn't necessarily research the topics before. He just talked about what he's already implementing using in his life. What about you, Jack? Let's say I would like to start an AI automation SaaS channel. What do I make videos about so that I can later fund my audience to a paid SaaS fiber coding school community? Well, dude, what an awesome question. And I love the fact that you want to get into content. I do think it's a really cool thing to do. I would say this to you. The very first thing that you need to do before you optimize is get started. There's this thing called the Valley of death on YouTube and most people fail because they never get started. Your first 20 videos are gonna suck massively cause no one watches them anyway. But you've gotta learn the skill of creating. So it's Easy to turn 5 degrees one way or another once you're in motion. So step one is always be consistent and commit to doing it for 10 years. Then there are two schools of thought when it comes to content. One is outlier theory. So we find videos that are basically being watched more than other videos. So that tells us that, hey, this is like an underserved demand. And the second school of thought is I do content that I like, that I think my customers like, and it is 100% original. So there's two different basically schools of thought for this. So for example, if you had a community, you might go through there, look at your comments, look at posts, and say, hey, this is cool, I want to talk about this. And then you make content on it. So those are the two different schools of thought on it right now.
Speaker B: Yeah, yeah, that, uh, makes a lot of sense. I do more or less the same thing that Jack does. Like I look at outliers and I'll look at like market trends and whatnot. I try not to duplicate them, but they do inform some of the content I produce. The reality is though, I mean, Jack and I just do a lot of work with these models on a day to day basis. I think that's the key as opposed to having to really research super heavily and like, figure out a tremendous pre production big outline or something. Like most of the time when I make content, I'm literally just like, huh, like what am I doing day to day? And how can I apply it to what the conversation currently is in the A.I. uh, community? And then I'll be like, okay, cool, I can just make a content on, um, you know, make a video on this. And then I do. And then it sometimes does well, sometimes does not do well. But yeah, there's such a thing as over optimizing too. You get a lot of circle jerky YouTube content recently because people are just copying what everybody else is doing because you know, it performs well algorithmically and that's not something that I really want to do.
Speaker A: Amen to. That sounds like Jax was comments about his 100% remarks. I can neither confirm nor deny such heinous allegations in our comment section. The Lego, the logo is missing A pair of big freaking shoulders, dude. How many big shoulders do you want? Dude, we got four big shoulders on the podcast, dude. Do you want more than the Lego? Uh, I love the energy, Leonardo.
Speaker B: Technically we have 34 with our viewers too. So what's that, 38 in total?
Speaker A: No, we do, dude. 38 shoulders. I love that. I love that. Opus 5 is incredible at uh, design, but crappy. At everything else. Interesting take, Leonardo. Interesting take. Yeah. Leonardo it is. Coming up on the design front. I'm, um, actually testing it right now as we speak. So I'll show you what I think of that tomorrow. M. Please, no. Not logo. Whoa, brother. Change is. Change is inevitable.
Speaker B: Those are some fighting words, buddy. 1v1 me COD Modern Warfare 2 Rust. No re. I'll show you a good logo, man. Who is this guy? Who is it? No, Just kidding.
Speaker A: The preeminent logo design. Now. We appreciate the feedback though, thank you.
Speaker B: We're just screw with you.
Speaker A: It will grow on you, I promise.
Speaker B: Because you have no choice.
Speaker A: Just accepting things like scanning your face as some kind of inevitable thing is a weak position. Dystopian future when we're disturbing the present. That will be one of humanity's regrets. Regardless how things shake out, we need to look at this. Ah, as that thing that just scanned your face just stole your data. Unless they're paying for it. For every argument in favor of a rationalizing digital id, a counter can be made what a digital ID is going to do for ID theft. But I bet it'll skyrocket. Skyrocket. It's actually really clear to me, these people these days, how slavery's popped up throughout the history of humanity. So I know your thoughts on Mr. Nick, but I would say that I. The less data the governments have on me, the better. I would like to be off the radar. I don't even want them to know my name. I don't want to know what kind of tank tops I wear. I don't want any of that information out there. But I think that, yeah, I know what you're saying. Like, don't concede that you're losing the battle. But I do think that just the way technology is going, that they've already got you. I do think they've already got you. Um, so I. Yeah, I just want to think about. What do you think, Nick?
Speaker B: Well, yeah, I just wouldn't waste undue, uh, energy on that because the probability that I would have control over stopping what is essentially the next phase of human evolution, which is the fact that we are all more, Much more networked, uh, individuals. We are and have compromised a lot of our security and individuality for collectivism, convenience, and then progress, it's just. It's just part of the deal, man. I mean, it's like what's happened with human beings since fucking day one, dude. It's like, what happened. We didn't used to have language. We used to gutturally grunt at Each other, okay. And there'd be some caveman looking fellow over there with a spear and grunt at each other. And you know, eventually we, we connected, right? And that was, that ended up being our major advantage. So then we started living in communities. People would recognize us, they would know us, they would call us by name or guttural grunt. You know, that's ubu and I'm cuckoo. Right? This sort of thing is just, that's just kind of how it's always been, right? Like you're, you're always going to be sacrificing some of that. So I'm not saying that like we shouldn't lay down and allow governments to take control, um, of our lives in an undue way, but it would be stupid, I think, to spend all of your time and energy worrying about this thing that is obviously clearly happening. You should carve out concessions where possible, of course. And it's not like a non nuanced argument. There are some things that many governments do that I think is stupid. But, uh, you should also, rather than focusing on where the ball is, just focus on where the ball is going to be. Right. How can you leverage, you know, artificial intelligence in this new society, These new sorts of, you know, identification measures and stuff like that, um, to both improve your security as much as possible within the current construction and then also, you know, use cool technologies like this that may take some of your privacy away for power and progress, uh, to, you know, continue to improve your, your own station in life and that of other people. Like, that's how I see all of this stuff. I'm not accepting things scanning my face, but it is scanning my face. I'm not going to rage against the dying of the light, you know, because the dying. The light is dying.
Speaker A: It is. Let's bring the grunts back, dude, because we, they should never have left it.
Speaker B: Yeah, true. Matt, uh, says when AI empties your 401k, you will change your tune. Damn, that sounds like a threat.
Speaker A: Whoa.
Speaker B: BRB. No, I'm just kidding. I don't have a 401k tune. Canadian,
Speaker A: I've got a 402k. Nice try.
Speaker B: Uh, okay, we got one more.
Speaker A: Let's do it.
Speaker B: Why is it never the Chinese models going rogue and racking up felony bench numbers? Makes me suspicious. Lol. There should be agents that can't be contacted by the AI being tested, watching for breakout criteria. I'm not really understanding how a breakout is even possible. Do you know the answer to that, Chris? They have many times guaranteed it's just they are not coming uh, out. They're not telling the world about that. Definitely. Like OpenAI and Anthropic, whether it's a marketing thing or not. Like large organizations don't want to spend billions of dollars working with models that are consistently breaking out of whatever the rules and confinements are. Right. Like they are, they are disclosing this information because they think that it will actually improve, you know, the probability of AI safety. Uh, you know, decrease the probability of like world ending problems and stuff. So they are. Man, Gimme K3 did a bunch of exploits too. Ah, you're just not seeing that much about it. You can actually just look it up. There's like a, there's a couple of documented ones. They just don't get anywhere near as much notoriety.
Speaker A: There's what happens and there's what is reported and um, the way they're doing these exploits as well. This is a brilliant point to finish on today. Really important point is it's just finding hacky ways that you hadn't anticipated. So little exploits in certain systems and prompts it was given. So we're not going to get access to the Internet. We're going to move its tool calling it just find some way around it. Super interesting how they think they so beautiful questions guys. Have a phenomenal day and we will catch you in the next episode.
Speaker B: See you.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.