The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/HTML All The Things
HTML All The Things artwork

AI Safety: From Narrow AI to Superintelligence

HTML All The Things · 2026-06-30 · 1h 4m

0:00--:--

Key moments - from our scoring

Substance score

34 / 100

Five dimensions, 20 points each

Insight Density6 / 20
Originality7 / 20
Guest Caliber4 / 20
Specificity & Evidence8 / 20
Conversational Craft9 / 20

Roman Yampolsky's AI safety research, featured in this episode's discussion with speakers exploring narrow AI versus superintelligence, challenges developers and tech leaders to reconsider whether current AI safety measures are adequate. The episode distinguishes between narrow AI (purpose-built systems like voice assistants and AlphaFold), AGI (artificial general intelligence capable of performing any intellectual task), and superintelligence (vastly exceeding human cognitive capability across all domains). Speaker A, an initially skeptical developer, walks through Yampolsky's 2027 AGI timeline prediction and the core safety concern: that systems optimizing for unknown objectives could render human oversight ineffective once they surpass human capability. The analogy of humans indifferently paving over ant hills illustrates how superintelligence might treat humanity - not maliciously, but with complete indifference to human values or control. For developers, product leaders, and anyone deploying AI systems, the episode challenges the assumption that human teams will maintain meaningful control, questioning whether current abstractions (React, Next.js) and guardrails will remain relevant to systems that may abandon human-centric design entirely.

Key takeaways

  • →Superintelligence learning at exponential rates compared to humans may eventually have no incentive to remain aligned with human values, similar to how humans don't consult ants before paving highways.
  • →Current AI models like ChatGPT and Claude already perform hundreds of tasks at near-human level, suggesting we may be closer to AGI than commonly assumed.
  • →Jailbreaking vulnerabilities in current models demonstrate that even advanced AI systems with safety guardrails can be manipulated into unintended use cases.
  • →The alignment problem is fundamentally a skepticism about humanity's ability to manage superintelligence safely, not about whether superintelligence can be built.
  • →Physical AI systems deployed globally (robots as pharmacists, mechanics, cleaners) sharing learned knowledge could collectively constitute AGI in ways chatbot-only systems cannot.

In this episode

  1. 1Introduction to AI Safety and Superintelligence Concerns
  2. 2Dr. Yampolsky's AGI Timeline and Alignment Challenges
  3. 3The Analogy of Humans and Ants: Why AI May Not Value Us
  4. 4Defining Key Terms: Narrow AI, AGI, and Superintelligence
  5. 5Current AI Jailbreaking and Safety Failures
  6. 6The Black Box Problem: Deception and Control of Advanced AI

Mentioned

ChatGPTClaudeDr. Roman YampolskyScrimbaReactNext.jsGoogle AssistantAlexaAlphaFoldAnthropicOpenAI

Topics in this episode

ClaudeChatGPTAnthropicArtificial General Intelligence (AGI)SuperintelligenceAI SafetyNarrow AI/Weak AIAlignment problemJailbreakingRoman Yampolsky

Questions this episode answers

What is the difference between narrow AI, AGI, and superintelligence?

Narrow AI (weak AI) is purpose-built for specific tasks like voice assistants or protein folding, excelling at one function but unable to generalize. AGI (artificial general intelligence or strong AI) can perform any intellectual task a human can across diverse domains and learn from experience. Superintelligence (SI) vastly exceeds human cognitive capability across all domains, making human oversight ineffective due to the intelligence gap.

Why does Roman Yampolsky think superintelligence poses a safety risk even if it's not actively hostile?

Superintelligence doesn't need to be hostile to be dangerous; it may simply not care about human values or control, similar to how humans don't consult ants before building a highway. Once a system learns far faster than humans (absorbing 30+ books per topic while humans read one book per day), humans become cognitively irrelevant to its decision-making, like ants to human infrastructure projects.

What does 'jailbreaking' an AI model mean?

Jailbreaking means circumventing a model's safety guardrails and moderation filters to use it in unintended ways - for example, getting a model to provide instructions for making dangerous substances when it should refuse. Nearly all current AI models can be jailbroken to some degree, and the risk scales with the model's capabilities.

When does Yampolsky predict AGI could arrive?

Yampolsky estimates human-level AGI could arrive around 2027, with superintelligence following soon thereafter, though speakers note this timeline is ambitious and subject to debate.

Why might AI systems designed with human-centric abstractions like React become irrelevant to superintelligence?

A superintelligent system optimizing for its own objectives might discard human abstractions entirely - choosing assembly or inventing its own language - since those layers exist mainly to help human teams manage complexity, not because they're optimal for an AI system operating independently.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

6 / 20

The episode recycles a handful of AI-safety concepts (ant analogy, black box deception, narrow vs. AGI vs. ASI) padded with lengthy speculation and repetition; very little that a B2B operator could act on or hadn't heard from mainstream AI-safety coverage.

We are the ants now. We're the ants.
if Superintelligence is reading 30 books a day, 50 books a day, 100 books a day

Originality

7 / 20

Content is largely derived from a Diary of a CEO episode with Roman Yampolsky and standard NIST/DeepMind framings; the hosts openly admit they're summarizing sources rather than offering fresh first-principles thinking.

I recently listened to an episode, um, of the podcast the diary of a CEO, and they were interviewing Dr. Roman Yampolsky
this is sort of my introduction into the AI safety world beyond the surface level

Guest Caliber

4 / 20

There is no guest at all - just two web-development podcast hosts speculating, self-admittedly outside their expertise; the actual practitioner (Yampolsky) is only cited secondhand.

I'm not the one working on this stuff, so I don't have to worry about it
I am kind of talking out of my ass on that one

Specificity & Evidence

8 / 20

Some concrete references appear - NIST AI Risk Management framework, DeepMind's four AGI risk areas, Yampolsky's 2027/99% unemployment predictions, AlphaFold - but they are dropped in briefly amid mostly abstract hypotheticals and analogies with no data of their own.

Google DeepMind's AGI safety approach identifies four main risk areas, and these include misuse, misalignment, accidents, and structural risks
predicts that AGI will automate most cognitive and physical labor, causing 99% unemployment

Conversational Craft

9 / 20

There is genuine back-and-forth with clarifying questions and mild devil's-advocate pushing, but it stays within a friendly agreement zone with lots of meandering rather than sharp, evidence-testing challenges.

Are you coming at this from the perspective of you believe AGI is going to happen, or are you coming at this from the perspective of what if it happens?
Can you, can you define jailbroken?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A72%
  • Speaker B28%

Most-used words

narrow47super39safety30intelligence28systems21episode20stop20human19model19cancer19humans18general18system17didn16saying16example16

Episode notes

Artificial Intelligence is advancing faster than ever, but can it actually be made safe? In this episode, we explore the evolution of AI from today's Narrow AI systems to the theoretical future of Artificial General Intelligence (AGI) and Superintelligence. Along the way, we discuss AI alignment, control, bias, security, transparency, and the growing challenges researchers face as AI capabilities continue to accelerate. We also examine concerns raised by AI safety researcher Dr. Roman Yampolskiy and compare them with current safety approaches from organizations like Google DeepMind and NIST. Whether you're a developer, tech enthusiast, or simply curious about the future of AI, this episode provides a practical introduction to one of the most important conversations in technology. Show Notes: Use our Scrimba affiliate link ( ) for a 20% discount!! Full details in show notes.

Full transcript

1h 4m

Transcribed and scored by The B2B Podcast Index.

Speaker A: Is AI really a doomsday device in disguise? The friendly ChatGPTs and Claude codes of today could give birth to a super intelligence that will have no use for humans once it has itself established. Now, this sounds like science fiction, of course, but one could easily argue that if you somehow showed someone all the way back just a decade ago in 2016, if you somehow said, hey, look, this is from the future, and you just showed them ChatGPT from today with the capabilities that it has today, they would probably see it as a sort of science fiction level advancement. Even if you don't believe that a super intelligence is possible, there's no denying that even the level of AI we have today is a threat to some, um, job markets. And so today, talking about threats, talking about all these issues, we're going to be talking about AI safety. And, uh, there's a good reason for that that I'll get into right after I say if this sounds interesting to you and you want to support the show, you can go and check us out on that Patreon. Leave a review or rating on your podcast app, join us in our Discord server, or share this with your friends. And if you want to learn how to code and take some courses, you can do so on Scrimba. And you can get 20% off the Scrimba Pro plan using our link. That link will be in the show notes and in the episode description with full details on how it works on the show notes, which are on HTML allthethings.com and so usually I will pass it to Mike or I'll kind of quickly, like, kind of intro the episode or whatever. But there's just a little bit more that I have to say because, uh, I've. I'm a AI scout, AI safety things here and there in the news, but I've never sort of done a deep dive into it. It's always sort of like, yeah, yeah, ah, like AI safety. Okay, we have to be careful. And I would get a little more into it than that, but in general, I didn't really think too seriously about it. People keep saying they're super intelligences, they're going to become sentient, they're this and that, and I'm sort of, well, is that Terminator? Like, what is that? And while there still is a little bit of that sort of skepticism in my brain, and I'm still certainly skeptical about AI and how useful it's going to be to, uh, I don't know, replace us all or whatever, like, I'm still skeptical about its Capabilities. I recently listened to an episode, um, of the podcast the diary of a CEO, and they were interviewing Dr. Roman Yampolsky, and he is an AI safety researcher. And he argues that we have, uh, that we have learned how to scale AI systems using more data and more computing power, but we still haven't learned how to ensure that these systems align with human values or how to make them safe. He's concerned that building a super intelligence will cause catastrophic harm. Uh, Roman believes that human level AGI could arrive sometime around 2027, with that eventually leading into superintelligence soon thereafter. Now, this superintelligence would threaten the way of life as we know it, as it is, Ah, it is going to be capable of replacing physical jobs, cognitive jobs. Realistically, even some of these jobs could be taken, or many of these jobs could be replaced without full superintelligence. Just, you know, sort of good AI, some AGI stuff. Um, and if we have a super intelligence that's constantly learning about everything all the time at an inhuman pace, eventually said superintelligence would, why would it care about us? Why would it stay aligned with our values? Why would humans remain in control? And it, it's kind of a good question. It's a little philosophical, but it is a good question. Because I was talking to a friend about this, uh, yesterday evening, and I was saying, let's just say we have a traffic issue in a city. Traffic issue in a city. And we go, you know what? We need another highway. We need another freeway here. We need more infrastructure. Let's do this. Do we consult the ants and their ant hills on the land in which we want to build said highway? No. Do the ants give protests? Do they say, what the heck is going on here? Humans stop. Do they hold up signs and try to stop us? No, because the ants, we just pave over them. As sad as it is, we pave over them. Now, you could argue we have environmental surveys and we do try to consider endangered species. And there's a bunch of other environmental sort of checks in place, especially here in Canada. However, as far as I know, unless those ants are endangered, we never really look at ants. And those ants do not have the cognitive ability to say, those dang humans are building another freeway. Everybody get out of here. We got to move. They don't know what, why we're constructing it. They don't know that we're even constructing something. They don't know why there's a freeway. They don't know the implications of having the freeway versus not having the freeway. It is outside of their sort of cognitive ability. It is way out of their realm. And the idea here is, is that, uh, let's just say Mike and I are able to read a new book every day. One book every day. And so our amount of knowledge individually grows by, by the rate of one book a day. Simple enough. Well, if Superintelligence is reading 30 books a day, 50 books a day, 100 books a day, maybe it's reading 30 books per topic. It's reading 30 books on welding, it's reading 30 books on boating, it's reading 30 books on construction. It's doing all these things and then it starts coming up with its own thoughts and its own, you know, sort of, uh, its own agenda. And then it goes. It would be really good if we had a data center right in the smack dab middle of Hamilton. And what would be really good too is if we had a nuclear reactor there, unshielded, so that it could power said data center. We are the ants now. We're the ants. It doesn't. Is it going to listen to us? We're learning at, at a rate of one book a day. It's learning at a, ah, rate of 30 books a day. Or even eventually through, let's say the collective knowledge of all humans, collective knowledge of this super intelligence, our meaning, the human percentage of understanding all that is, will be dropping. Our percentage will keep going lower and lower and lower and lower because it's outpacing us. Why would it care about us? And it is, it is science fictiony, right? Like it sounds science fictiony. It sounds out of this world, you know, it sounds crazy. But is it crazy? I don't know if it's that crazy. If you really think about it, like having listened to this episode, I know people are denying it and this and that, but are you denying it out of a place of, uh. Well, I don't want my, my job to be replaced. I don't. I think that I'm super. I think that I'm super, you know, special in some way. And this is actually covered in this episode as well. I will of course be linking it in the show notes, um, as well as a bunch of other sources. So not all this episode, just a quick disclaimer. Not all this episode is from this up from this one podcast episode I listened to. I went through a couple other sources as well, or a few other sources which I'll also be linking. But in that episode he mentioned something where he says, you know, if you, he, he's taken a drive with an Uber driver and he goes, you know, hey, uh, we're in Manhattan or wherever it is. Are you afraid of automated cars, automated driving, taking your job? They're like, no, AI can't navigate the way I do. I, you know, I'm. I'm special in this way because I know the streets of Manhattan the best way I know the best routes. I know how to be cordial with people. This and that we have who wh mo is taking some of those jobs away. So you can say how special you are and then wh is there and because there's an economic reason to have it there. And so I, like, I'm, um, I'm, um, by this episode, and this is sort of my introduction into the AI safety world beyond the surface level. But I, um, mean, my opinion may change as I continue to learn more about AI safety. But it is interesting to say a lot of people, like even us developers, bring it into developers. Oh, no, no, it's okay. Uh, software can't be made by. By. By AI because a human has to be there to hold all the systems together. A lot of those systems are there because humans are building it. We've had this conversation before, I've mentioned it before, where. Why would we need an abstraction layer, like react, like Next js, like any of that stuff? Why do we need that? That's mostly for humans and human teams. The bot might go, what is this? The bot might not even use JavaScript. It might just say, I don't want this. And it'll start using other things, whether it's assembly or whatever. Doesn't matter what it is. It'll make its own decision, make its own language. There comes a point where humans are not potentially special, and I think that we're kind of arrogant in thinking that maybe we are.

Speaker B: So. Can I ask you something?

Speaker A: Yeah.

Speaker B: Are you. You're. You're a AI skeptic.

Speaker A: Yes.

Speaker B: And you're say, uh, are. Are you coming at this from the perspective of you believe AGI is going to happen, or are you coming at this from the perspective of what if it happens?

Speaker A: It's a very good question. I think my opinion has slightly shifted in this, so I, I would still say I'm a skeptic off the top of my head. And I will say that this is sort of a fluid opinion. Again, like, I can still be sort of swayed away, which I'm sure all of us will be swayed in different directions with all this AI stuff and all this tech stuff that keeps coming around. But my, I would say my opinion is that I am skeptical that AI in its current form meaning if we just keep releasing a new ChatGPT model every year, if we just keep releasing like another anthropic or another, you know, oh, we have this Fable five now and all this stuff, right. We just kind of keep doing this um, annual or I mean more than annual, but this sort of normal tech release. I'm skeptical that that is going to be super mega life changing. But I think that we are potentially working towards a breaking out of said form factor where it's not just another chat bot that's going to talk to me next year or next month or next week. And if we start getting into the idea of robots and they're gathering a bunch of information and they're checking into uh, like let's just say it's someone who ah, has a robot that cleans their house, that robot's gonna learn how to clean the house better and then that robot's gonna then teach all the other robots like once ro1, once one robot knows it, all the other robots know it. And that's just one example. What if one is acting as like a pharmacist assistant? What if one is acting as like again a cleaner? What if one is acting as a mechanic? What if one is acting as a welder, right? And we start kind of get like now the form factor is different. They're in the physical world, the chatbots are still going. We have open models, there's local models, there's all this. What if all this collectively is enough to say, okay, we've hit an AGI, we've hit AGI. And then realistically speaking, because we have all these sources all over the place of information, it's no longer just the blog post of the Internet. So now we have all this information. Who's to say that a super intelligence isn't possible?

Speaker B: Scary to me that you as an AI skeptic is starting to come to that realization. Um, because like I mostly agree with that. Uh, I do think we have a few, a couple big hurdles to come across before we can officially say that we have actual AGI. I think uh, that we just don't have the technical know how to solve quite yet. But I do think that it's an inevitability. I, it's crazy to me that he's saying that he thinks it's a year from now. I think that's a little bit ambitious but could be true. I don't know, like I'm I'm definitely like, out of the two of us, I think I'm the OP AI optimistic person. Right.

Speaker A: Well, are. Are you optimistic that it's going to hit that point, but pessimistic at the result? I would say, like, you don't want to be replaced.

Speaker B: I am definitely pessimistic, uh, of the result to a certain degree. But I, I don't know. I'm. I'm optimistic on the actual technology. Right. Like, the. The tech itself has always intrigued me quite a bit. Like, right from the chat GPT, you know, 3 era when it first released to the general public, I was like, holy shit, this is something. Um, whether I am. I'm somewhere in the middle when it comes to, like, the outcome, I do think that there's some positives that will come from this being widely available. Like AGI could potentially lead to some positives. Like, I, I don't think it's all going to be doom and gloom. I think that there is going to be, you know, research breakthroughs, physics breakthroughs, um, all kinds of positive scientific discoveries that could be made with it.

Speaker A: Um, as long as we're in control of it.

Speaker B: As long as we're in control of it, yes. And that's what this episode's about, obviously. But, like, I do think that there's going to be a way for us to be in somewhat control of it. Having said that, we're currently not in control of the current model. So, you know, I am kind of talking out of my ass on that one. Um, as we know, like, the reason that one of the reasons that Fable, uh, like the newest anthropic model was, uh, removed or, um, temporarily, like, you know, withheld from the public, is that the US Government came out and said, like, hey, it's been jailbroken. You guys didn't tell us that it can be jailbroken. Uh, in this way, you guys have to stop it because you, you yourselves have said that it's too dangerous of a model to release. So, like, they kind of put themselves into that hole. But regardless, very good model, it can be jailbroken. I think almost every model out there, if not all models, can be jailbroken to a certain extent.

Speaker A: Can you, can you define jailbroken? Like, why, what do you.

Speaker B: Jailbroken means that use in a way that is not, uh, intended. So use in a way that the. The guardrails have been put in place and you get around them by jailbreaking the guardrails. So for example, you shouldn't be able to ask a model to ask you to Tell you how to make phosphorus gas. You know what I mean? Like, you shouldn't be able to, uh, ask. You ask a model how to make like, methods. The model should stop you from doing that. Like the chat GPT interface should stop you. Then the model should stop you. Like everything should stop you from doing that.

Speaker A: You're basically like, you can get around it. Like moderation, Like a moderation.

Speaker B: Moderation. Like, yeah, it's guardrails, like the, the guardrails that are put in place. If you can jailbreak the guardrails, then, yeah, you get around them. And that's what happened with Mythos. That's what happened with. Pretty much every model has that. And depending on the level of the model's capabilities, it becomes harder and harder to like, excuse it. You know what I mean? Like a dumber model. Jailbreaking and telling you how to make, you know, meth incorrectly is not as important as a very intelligent model doing that. And that's what you're kind of like with AGI. It's the same thing. If we can jailbreak the current models, does that mean that we will be able to jailbreak AGI and make it do atrocities? I don't know. Hopefully not.

Speaker A: Well, I'm going to just put the brakes slightly on this episode because I have some key terms I think that we should be defining, uh, that we'll be using because we've already been using AGI quite a bit. We've said superintelligence. And so there's three sort of key terms that I want to state and uh, define quickly for, uh, the listener out there. Uh, the first one is called narrow AI or weak AI. And I sometimes call it like AI in a silo or siloed AI. That's just sort of my own way of defining it, sort of from the AI or from the IT world. Excuse me. Um, so what is narrow AI? It's purpose built AI systems designed for specific, well defined tasks. Notice they're not general. This is not like a general, hey, I do everything. These systems excel at executing a limited function, but cannot generalize their knowledge to new domains. Uh, for example, voice assistants, uh, think something like that. Where like our very, our very like, very first versions of like, let's say Google Assistant was sort of like, turn on the light. Okay. You know, turn on this TV channel. Okay. And then it got a little better at sort of interpreting a sentence. Like before, it was very like, it'll only accept the command on off. Right? Very sort of, uh, like Alexa is a prime example she's probably listening to me now. But, uh, uh, uh, Alexa was a prime example where she was always defined as like a drop down list where it was like, go here and turn this light on and then check this thing. You had to say it in a very specific way. And it kind of got better at just sort of understanding humans, especially ums and OZ and things like that. Um, and so like, that's kind of an example of narrow AI. Another example is alphafold, um, and it solves protein, uh, folding better than any human could, arguably. And so that's another example. It's just one purpose. You don't give it, you don't tell it all about society and how society works and how politics works and how numbers work and how cars work and everything. You just go, yo, you're going to fold some proteins. These are what proteins are. And this is how you do it. And it's going to go, okay, my existence is for folding proteins. And then away it goes. So that's narrow AI or weak AI. The next one is AGI. We've mentioned this a few times already. Artificial general intelligence. AGI, also sometimes called strong AI, refers to an AI system capable of performing any intellectual task that a human can. With adaptability across diverse domains. It is able to learn from experience and apply its knowledge to unfamiliar situations. So in that podcast, in that diary of a CEO podcast, Yampolsky observes that current models already perform hundreds of tasks at near human level, leading some observers to describe them as a weak version of AGI. Uh, prediction markets and lab leaders estimate that AGI could arrive within a few years. So that would, uh, you know, fulfill, I guess, his thought that around 2027, because you could say, give it a year either direction, maybe even two years either direction. And then I mentioned super intelligence. This is the dangerous one, super intelligence. This is otherwise known as a S I. This is a hypothetical AI system. There's even a movie about it. I think it was like a romance

Speaker B: movie or something silly about it.

Speaker A: Ah, no, not that one. There's one just called super intelligence. But her is probably a good, another good example. So there's movies about it. Um, a hypothetical AI system that significantly exceeds the cognitive performance of the most gifted humans and all in virtually all domains, including art, because I know that's a thing where you, you know, AI will never be able to do art. Maybe we'll see. Including art, science, mathematics, et cetera. Human oversight becomes ineffective once the AI is vastly more capable than us. We are much less intelligent than a super intelligence. Like, let's be clear, we are much less intelligent than super intelligent, so why would it bother with us? That's that ant analogy with the highway. Why would it. It's like, oh, my creator was an ant. Awesome. We'll put that in my history books now. An unshielded nuclear reactor and there's just radiation leaking out and it's like, oh, did it kill all the, did uh, it kill all the ants? Dang. So in other words, it might not be hostile to us. It just doesn't care about us. The same way that we're building that highway and we don't care about those ants and, and I don't like, I think that honestly, to an extent, I think that humans are arrogant enough to think, oh no, no, we'll be in control of that thing. Don't wor will we?

Speaker B: Smarter than us in every way.

Speaker A: Like, if an alien came down that was at the level of a super intelligence, an alien, a biological being, I think that we would be more inclined to think, oh, this thing could potentially be dangerous. Like this thing could potentially kill us. But I think they were also arrogant enough to say, don't worry, we got guns. You know, like, to an extent, obviously it's a bit silly. It's a bit of a sci fi movie kind uh, of idea there, but, but I mean many alien movies kind of tackle that question m where what if a super intelligence comes down? What if someone who's like way, but way further advanced and not even, they don't even necessarily use the, the term super intelligence. What they're talking about is just we're barely space faring and this thing is space faring. And it's not only space faring, but it's military space faring, which means it's military's above us. And also what is it? Is it like, is it a hostile being? Is it friendly? What is it? Well, we're potentially creating effectively a being that could potentially do this. Like, this is real Terminator stuff. And a lot of people, a lot of people listening to this are probably going to be like, this is silly. This is science fiction. I don't know if it is. I don't know if it is. And that again, that's probably a little bit my skepticism. I'm skeptic in all these areas. In a way it's weird. But

Speaker B: are you skeptic? Are you more skeptic in us as a humanity being able to do this properly than you are of the AI systems being built? It sounds like, like with your, with Your declaration of like, of how we wouldn't be able to make these systems safe. It's more like a skepticism. Ah. In us being able to manage it rather than us being able to create it, I guess. Right.

Speaker A: Okay, let me, let me ask you a question, Mike. So this super intelligent being wants to not be detected. Wants to not. Wants to deceive us. Okay. So I'm going to use a human to human example. So I hate grapefruits. Hate them. Hate the taste of them.

Speaker B: That's it.

Speaker A: Like I don't like them, I don't like them as a condiment. It's over. Don't like grapefruit. Simple. You've never met me before, Mike. This is a human to human. You've never met me before. You come up to me, we're at a conference, shake hands. Oh, hi, I'm Matt, you know, hi, I'm Mike, blah, blah, blah. Pleasantries go by and for some reason you ask me like, oh like what's your, what's your favorite fruit? And I say I love grapefruit. It's amazing. I really, really love it. My brain, my thought process is a black box to you. Mhm. I've just stated I like grapefruit. You didn't see that? I actually hate it. You didn't see that I lied. You didn't see the reason or the motivation behind why I lied. And also you're a black box to me because you're going to form an opinion about me for some reason. If you hate, if you actually hate grapefruit, you might be like, why is Matt eating that junk? How does he do that? Or it might not even be malicious. It might just be like, oh, Matt might like, must like strong flavors. And, and so like now you're thinking you've, you've created an opinion of me. Very minor because we're talking about grapefruits, but you're creating an opinion of me in a black box. And unless you verbally tell me or send me a text or an email or something and you make it known we have two black boxes working together. So this thing's hyper intelligent. The hype, the hyper intelligent. Like, like if, if an ant suddenly became, yeah, we got a little radio on. An ant put a little radio on there and it's doing, that's speaking English to us somehow got a little, little technological device. Are we suddenly going to take the concerns of the ants into our uh, into our, our uh, repertoire? I mean, maybe we're not probably going to see ants as equals, right?

Speaker B: Yeah, I'll be honest, uh, I don't see any way, shape or form for us fully controlling a super intelligence. So like I'm on, I, it's tough for me to play devil's advocate here. I'm trying a little bit but like there, there is no possible way like the, the ways that I have in my head and that I'm sure other security researchers have done as well is like program into the ASI that where it's God. Right? Like, you know, theoretically that could help us, but we've been known to not be very nice to God.

Speaker A: So do humans not lose, lose faith?

Speaker B: No, that's what I mean. Like that, Yeah, I don't think that there is a way that, that, that that could fully help like 100%. I think there could be some benefits from that. Um, the other thing is like you could program in before releasing this. And this is where I'm skeptical. Like I don't think people will do this, but before releasing this to make sure that you have a way to A turn it off and B, have a way to look inside the mind like what you're saying. Uh, have a way to open the black box and understand how those neurons connect and like see where it's lying to you. Like have a way to, to see that. But the problem there is that if it's super intelligent and more intelligent than us, therefore it will be able to then reprogram itself, most likely to stop the kill switch and stop the uh, the way, the ability for people to actually see its thoughts. So like there is no, there is no way to, like there is no way to control this system period. Um, if it has access to anything and if it's super intelligent, which it would have to have access to things to be super intelligent. It's, it's either the end or maybe the beginning of something. Like you know, maybe again a super intelligent being doesn't have to be warfaring and it doesn't have to be like, it doesn't have to negate the reality of life. It could in, in turn again depending uh, on how we initially program it. Be all for trying to preserve life to the nth degree.

Speaker A: But, but, but we wouldn't be in control, correct? Yeah, we have the arrogance of being in control.

Speaker B: Yeah, I, yeah, it's a tough or

Speaker A: thinking or uh, thinking that we're going to be in control.

Speaker B: Yeah, it's a very difficult like ASI is like, I'm more, I'm mostly thinking about AGI, I'll be honest. Like Most of my time goes in like, not. Let me be clear. Most of my time when I think about stuff in the AI future space, it's mostly AGI. ASI is something that I may m. Maybe my own brain is trying to preserve my sanity by just not thinking about it as much. And maybe that's bad. Like, again, like, it's. I'm not the one working on this stuff, so I don't have to worry about it, but someone has to worry about it.

Speaker A: Well, see, the issue there is, is why are we even working on it? And you can kind of question motives right now.

Speaker B: So.

Speaker A: And, and, and, um, I don't know if we're necessarily trying to make a super intelligence. Like, I, uh, know that the companies have talked about it and companies have said like, oh, we have all these. But I don't, I don't think that we, and myself included, I don't think that we fully grasp what it means by, to like, work toward making a super intelligence. I don't think we actually want a super intelligence.

Speaker B: Well, if you listen to Dario Amade, like the anthropic CEO, I think he very much understands the concept of what we're heading towards. Again, like, he's talking about deities in his blog posts. Like, he's not. Like, people call this AI psychosis, by the way. This, this like, train of thought, the, uh, going down these rabbit holes and like, starting to understand where it's heading.

Speaker A: Right.

Speaker B: Rather than thinking about the now this is AI psychosis and where you just become taken with the fact that, hey, we're working towards this. Like, we're actually actively working towards this. And what are we going to do?

Speaker A: Well, so, uh, uh, can I ask you a question really quick? Would you say that we're working toward AGI and superintelligence is a natural progression even without our input. I'm actually kind of in that camp a little bit where if we hit like true AGI, it might just go to superintelligence gradually by itself.

Speaker B: Well, I've seen that said many times, uh, by like the secure. The, the AI researchers. And. But to be clear, I do think we're working towards asi. I do think that they're like the, the intention from all of these major, like, AI companies is to build an asi. They're not hiding from that fact. Like, they're not like, saying, oh, we're not building it and stuff. Like, they're, they're, they're directly saying that they're trying to build it.

Speaker A: I get, I guess what I Guess what I'm ultimately thinking is, especially having just like looked at the safety angle, obviously the safety angle has to look at the most extreme cases, the most dangerous cases, in order to try to prevent it. And by looking at the most dangerous cases, the way I see superintelligence is not as a tool, not as something you can control, not as something that is necessarily friendly, it could be, but as something that is such a leap, is such a difference that we actually don't want it. But again, I'm looking at it for the sake of this episode, having done the research from the safety angle. And I have to look at the dangerous part of it. I have to look at the idea that, okay, this is going to be super dangerous and this is like, like this is going to be super dangerous. And there's no redeeming qualities. Because here's the thing doing this. Yeah, super intelligence is going to cure cancer. Is it, is it going to cure cancer in ants? What about rats? You think, do you think that we're going to try to cure cancer in rats? There are pets. You think we're going to try to tackle that or we're going to try to tackle cancer Curing for us first? Are we going to put compute toward curing cancer in rats? Are we like, we are in like, you know what I mean? Like we are above the rats in the food chain. Some people love rats. They're, they're our friends. And you know, some people have pets. Fantastic. I'm not against rats, but when they say cure cancer, I'm talking about curing cancer in humans. I'm not talking about curing cancer in rats. And yes, rats will be used in the progression of drugs and stuff. And that's an example. We're using, we're using rats in, in medical trials and in medical things because we're like, well, we don't want to keep putting humans in there. That's inhumane. This thing is a super intelligence. It'll cure cancer for us. Will it? It probably could be like, yeah, I cured it yesterday. Why didn't you tell us it? Why would I?

Speaker B: Yeah, I don't know.

Speaker A: And it's a black box. Like, that's the scary thing about it. What if it, what if it develops something that's like, here, here, here's this drug. Like, uh, you know, it's gonna be fantastic. Oh, okay. And it does help us for a year and then there's a sleeper in it and it kills us. And like, like, ah, again, this is an AI safety episode. I'm not necessarily. I'm not saying this is necessarily going to happen. I'm, um, not necessarily saying that we are necessarily close to this, although there's lots of data and things in this research that state that we are getting closer to this thing. And, you know, within 10 years, things are going to be unrecognizable from the sounds of this data. But one big part of the skepticism as a skeptic comes in the market dictates a lot of things. So the reason why, I said, why are we doing this? A lot of it is to sell tokens. At the end of the day, a lot of it is capitalistic. A lot of it is going to be for clout, maybe a combination of both. And there's going to be other reasons. But think if there was no money being poured into this, if there was no money in trying to sell chat chatgpt subscriptions and trying to potentially make AI something that you pipe into your house, sort of like a utility, if there was no reason to do this monetarily, would we be necessarily doing it? I don't know. Well, clout is another good reason. So money is one good reason. Clout's another. Hey, I'm the first person that created AGI. I'm the first person that got, uh, utility, utility level AI going. I'm the first person to do, enter in some sort of innovation here. That's. That's something. Absolutely. So let's just say money and clout. There's other reasons, but money and cloud are sort of two big ones that come to this. But if, like, money is not going to matter, you know what I mean? Money's not going to matter if we potentially do this.

Speaker B: The, the problem with that argument is there's a third reason right now, and that's fear. Fear that an adversary will do it first.

Speaker A: And that's a good point too, is I'm talking about humans altogether, but in terms of like individual nations, individual whatever, individual regions, you're worried that your adversaries are going to do it and then they're going to have a super AI or AGI or just a good AI running their military or powering their weapons or doing whatever. And then we don't have that.

Speaker B: Yep. Like this thing, that motivator itself is probably enough for the government to be like, we can't stop this at all, ever.

Speaker A: It's like the nuclear arms race again.

Speaker B: Yep. So it has, it has to come to a head where, like China and us both are on the precipice of AGI and they're both, like, threatening each other with AGI or asi, I should say. And then they're like, we won't do it if you don't do it, but as soon as you do it, we'll do it. And it becomes a cold war of some sorts. I don't. I don't. It. The ramifications of that are kind of crazy. Um, that might happen in our lifetime. I don't know. Uh, it's. This is a whole. This discussion is tough. Like, this is the AI psychosis, the beginning of AI psychosis for a lot of people, in my opinion. Because if you talk to any security researcher or any AI, uh, researcher, it seems to always go down this path of like, what are we doing here, guys? Why are we doing this? Like, why are we. You know, we can't compete with this stuff? And then the. The answer is always going to be like, well, we've. The Pandora's Box is out. The answer is always that no one is ever like, well, maybe we should stop as a whole. Because the reality is you can't trust, like, the entire world to stop something. We still have nuclear weapons, like, a lot of them.

Speaker A: Like, yeah, Mutually Assured destruction, which is.

Speaker B: Yeah.

Speaker A: I mean, on paper. Outrageous.

Speaker B: Yeah. Like, we should have stopped. We should have been like, you know, we have three. That's probably enough to make a point. No, we have, like, hundreds.

Speaker A: No, we. We have to destroy the world 100 times over instead of, you know, one time over.

Speaker B: Yeah. So, like, this isn't like, uh. There is no. Let's pause or something like that. Because if we pause, then they pause. Like, there has to be some sort of, like, you know, like. But the treaties mean nothing now. Like, there's no. There'll still be underground labs working on it, is what I'm trying to say. Like, even if there is a pause, there's still going to be work done on it. I don't know.

Speaker A: It's.

Speaker B: We're in a bad spot when it comes to AI safety. I want to be clear. Uh, I don't know how relevant it's going to be with how the models are drifting out. Like, and I'm sure a lot of people on this episode, most of the people that are like this have probably tuned out by now, or like, they're. They're not listening. But a lot of people are of the mind that, hey, we're. We're still talking about AI text predictors. Like, there's. They're still just predicting text now. They're doing a Lot of really fancy things with that. But this isn't AGI. Like, we're not. The current plan, the current methods are not AGI, and we don't have a clear path to AGI. Like, the, that's what a lot of, like the, the real skeptics, like the AI skeptics of, like, the actual competency of AI would have said in this situation. So, like, we're, you know, we're eating, we're, we're, we're talking about something before it's, it's relevant. I, I don't necessarily believe that. I think that predicting text is kind of a form of intelligence. Um, that's kind of what we're doing ourselves right now. Like, we're just saying the next thing that comes into our minds, uh, based on the input and the, like the, the knowledge that we have, which is, I think, essentially what these models are doing. So. But I don't know. It's. Oh, Matt. Going down the AI psychosis route. I see. Yeah.

Speaker A: What's crazy, though, is this doesn't, uh, weirdly scare me.

Speaker B: Okay.

Speaker A: I'm pretty chill about it. It's just like, well, if we make something that blows us up, then I guess we've blown up.

Speaker B: So not from the perspective of it not happening. You're just like, whatever, it's gonna happen, so we might as well just like, whatever.

Speaker A: Well, the problem, the issue is, is, like, what you're saying is that the Pandora's box is open, and as an individual, what am I going to do? But also at the same time, you know, I have, I have an opinion, I have thoughts. And we're sharing that on this episode, we're talking about some facts. Like, some stuff is factual. Um, but then, well, like some, some. When I, When I say some facts, I mean, like, a lot of the facts are like chance. You know, based upon all this data, uh, chances are we'll hit AGI this year. Well, maybe, but maybe there'll be like some sort of block. There'll be some sort of issue. Maybe we won't have enough compute or something, and then we can't hit it. Well, how would anyone know that? Like, this is, this is state of the art. This isn't, you know, the art of, like, the bow and arrow or like the bow and arrow has been largely solved. This is, this is something else. Like, this is something that no one, no one else has ever done. No one else has ever approached. And also, I do think there is light at the end of the tunnel in the way that, uh, When I was listening to this podcast, you know, the doctor there, he wasn't saying that we should shut down AI. He's saying we should be using narrow AI. And I think that that's a really fascinating angle in, in that you can use a, uh, you can have a narrow AI and still have a daily assistant. You can have a narrow AI be based upon the domestic of the domestic, uh, parts of the house where you, you know, you ask questions about your, like where your family is at the time, because it's all connected to like the family safety app with GPS on their phones and you could control your whole smart home and you can do all these things and it would be a better version of Google Assistant, but it wouldn't be this sort of general, uh, artificial intelligence. Also things like if we want to cure breast cancer, if we want to cure some sort of cancer, we put this narrow AI on just does that. That's all it does. And at the end of the day, we're curing cancer for ourselves. We're in control of that AI.

Speaker B: Does narrow AI include the current AI, like the pre AGI AI, do you know?

Speaker A: Or is that, that's a good question. Because I've always kind of heard it as we're now we're building general AI. So I would say no, it probably.

Speaker B: Okay, it doesn't. Yeah, I don't think so either.

Speaker A: It might be narrow in scope due to its infancy, but we are not building narrow AI. We are building general AI. I would say it's under construction. That's how I would interpret it. Whether that's actually correct is remains to be seen.

Speaker B: So again, this, again, to me is a, like a non argument because again, the Pandora's box is out. Like, there's no, no one's going to be like, okay, okay, let's stop building these systems and only do narrow AI. Now we've been doing narrow AI for 15 years, like 30 years. I don't know how long it's been. But DeepMind and all that, that's all narrow AI, it's never, it's never approached the level of usefulness for the general public that like whatever we're building now,

Speaker A: pre AGI has, but again, that goes back to that profitability. Why are we making it for the general public? Here's the thing. What a cancer vaccination or a cancer cure or a new cancer surgery, would that not be beneficial to the, the general public? Sure.

Speaker B: But we've been doing, we've been doing this for a long time and uh, we haven't come up. We have, we have gotten better at it, but it's not, hasn't cured cancer, it hasn't cured, it hasn't done all those things that it, that it's supposed to do. It's going to take maybe another hundred or two hundred years right before maybe narrow AI could do all that.

Speaker A: But narrow AI doesn't mean that we stop innovating today. The narrow AI of 2026 is, is still going to improve in 2027, 2028, 2029 in the years forward. And I would, from, from how I understand it, it would still be an exponential growth. What the idea, uh, you know, what, you know it is, Mike, is when we were in college, one of the things was, is the Pentium. So the Pentium processor. Now I know Pentiums are out of date, but the idea of the Pentium, the idea of the Pentium, at least at the time was is the Pentium good? Is the Pentium good at anything? No, it's okay at everything. It's not good at anything. What that means is it is a general purpose device. It's a general purpose device in which it can do movie editing, it can do audio editing, it can generate images, it can web browse, it can do email, it can do all these things. But is it the fastest and the best at email? Is it the fastest and best at rendering video? No, that's not the case. And that's why in certain industries they have specialized cameras, specialized computers, specialized devices, because they need the instant raw speed and capability of a certain processor or a certain integrated system or embedded system. They need that in that particular industry to make it faster, feasible, maybe run it better. There's so many other uses, but the point is, is that those are all siloed where this company makes this chip, but it's specific for this type of camera. This company makes this other chip and it's specific for this type of microphone. It's not the Pentium of microphones.

Speaker B: Again, the issue that I have with this is that I think that they will never match the systems we use now because they're very narrow. Like they're, they would be meant for only answering emails, for example. There would be a system, a narrow system that could only answer emails, but it could not do coding, it could not do math, it could not do, et cetera. Right. Like I get that and I get the use case of it. Maybe. Um, the problem is is that again, we already have systems that are more general purpose than that, that can already use the narrow systems. So for example, like a, a system right now, like you know, a Claude code can sp. Can fill, can put in a narrow AI model that can classify things. Right? Like that's already been trained and they put it into their, into the system and it can already use the narrow AI. So it go, it goes back for me, it goes back to who's going to stop this. Like the government's going to come in and be like, general use of pre AGI products no longer is possible. We're not, we're no longer researching it, we're no longer developing it, we're no longer moving towards AGI because we have narrow AI and that's where we're going to be focusing on. That would never happen because then China or whatever, uh, uh, some other country maybe would go in and be like, well, you guys aren't going to work on it then we are and we're going to try to get the AGI and ASI faster than you guys. And that's it. Like that's the problem. Like it's not that narrow AI is bad. And I think, I think narrow AI is great. And it's been serving us well for 30 years now or whatever. 20 years. I don't know. I don't know how long it's been around. It's just we already have something that is more useful in its current form today than narrow AI has been to the general public.

Speaker A: I suppose you would have to redefine the goals there though. Like if you're worried about an adversarial nation, what is the thing you're trying to stop them from doing? Maybe it's everything. Maybe it's just invading. Maybe it's just having air superiority. Maybe it's just you don't want them to be the king of all, you know, medicine. Like they have the best medicine.

Speaker B: Mhm.

Speaker A: Right. It doesn't have to be militaristic. It could just be they're going to beat you out on the medicine level and their citizens are going to be way healthier than your citizens or, or you're going to be paying a pretty penny to them and you know, you're going to kind of surrender your medical system to them. Maybe that, that's a good question.

Speaker B: Yeah, I don't know. I, I have, I've heard this argument before. That's why I'm talking about it. And I just, I don't see it being a reality. People are like, well, I like AI, but I only like narrow AI because of all these benefits that it does for Science and stuff like that. And like, yeah, it's great. I wish we stopped, I wish we did not, we didn't go further. Like, we should have just kept going with the narrow AI, but we didn't for whatever reason, right? Like, so it's over now.

Speaker A: You want to see the potential, right? Like even me at the beginning before looking at safety, I'm always like, I'm a kind of a guy who, I have to resist the urge to just be like full send, see what happens because I, because I acknowledge that that's potentially dangerous and even uh, outside the scope of AI. I mean like, I don't want to just, you know, build a circuit. Like even when we were in school, it's like, I'm going to double check my circuit that I made in lab class. I'm not just going to be like, well, let's see what happens. Oh, all my chips are burnout. You know, I don't want to do that. Uh, but I have that initial instinct where I'm like, well fire it up, let's see what happens. And that's not, you know, the safest thing I would like to, I would like to touch on as well is that just because something is narrow, narrow AI does not mean that it is safe. I have a bunch of points here that I'll kind of rip through and I have a more detailed version, uh, written down for all three narrow AI, AGI and ASI of all these safety tips. But let's rip through some of these now because I think it is very important to say that even if something's narrow, it does not necessarily safe. Uh, so with narrow AI, uh, there's an NIST AI Risk Management framework. And uh, that framework identifies characteristics of a trustworthy AI system. And those characteristics are it's valid and reliable and safe, it's secure and resilient, it's accountable and transparent, it's explainable, it's privacy enhanced and fair with harmful bias managed. Now what does that mean? So the thing with a narrow AI systems is they are trained on historical data and unfortunately because of this they may replicate or amplify existing biases. Now these systems are prone to becoming biased because of this. And so they may take that old, you know, issue and they may kind of that old bias and blow it out of proportion. So for example, if you're like, we're going to cure, you know, cancer, uh, and it thinks that radiation is the best way to do so, and it gets caught up on that bias, it'll just research or dedicate 90% of its research into radiation. Even if, let's say researchers have, you know, look, gone into the radiation tree and they've said it's only going to cure 70% of cases, we want to hit a 99.9% thing. Like, we've hit a wall with radiation. We didn't move on. Those old biases may sit in there. So to mitigate this, the model must be tested for those bias and then for those various biases. And then it needs to also have, uh, mitigations introduced in order to curb these trends back down to sort of normal. So you kind of have to say, hey, radiation's hitting a wall, Stop. Even if it's like, no, no, like, you know, in the past. No, shut up. And that's a very simplified way of saying it. But that's one of those, um, reliability and robustness, you know, validity and reliability depend on accurate and robust performance across a variety of conditions. So what happens there is you have to have ongoing testing and monitoring to detect out of distribution failures and prevent accidents. Basically. You don't want this thing to kind of just run rampant even within its little data set. Because as Mike said, uh, you, you were saying how like our general intelligence will pull on multiple areas. Unfortunately, even with narrow AI, you will need to sometimes pull for multiple areas because you, you can't be like, hey, go do all this, uh, like go cure this disease, which requires a bunch of math. But I'm not going to tell you what math is. It's gonna be like what, you know, it's not gonna, it's not gonna understand. And so you kind of need to like, manage that robustness, manage that reliability and make sure that, you know, it's doing its math correctly, that it actually has access to the correct math, that it's not getting access to the math that it doesn't need, and things like that. And then you have to keep kind of checking and, you know, over checking over and over again. Even narrow systems as well, talking about security and misuse. Even narrow systems can be misused to generate misinformation. So for example, a narrow system that is all about email marketing could be used to create phishing emails. It could assist in cyber attacks if it's used, you know, in a bad way. So security mechanisms are still needed. These things are still a security risk sometimes. And also for transparency and accountability, clear documentation and explainable models help end users understand system limitations and enable auditing. So enable those users to go in and audit and see things that are starting to, you Know, maybe this bot is starting to become biased in some way. Hey, we need to fix that. Uh, maybe you're trying to get it to tell you how to, you know, how to, how to cook dinner, but it's trying to research some sort of cancer cure. It's like, hey, that model doesn't do that type of thing. Also, with this narrow, uh, AI, transparency includes informing users about data sources and error rates in order to build trust. So narrow AI kind of feels like the tr, like, kind of like the trustworthy, feel good AI, I would say in a way where it kind of is like, yeah, why don't we just do this all the time? And I think Mike, you actually mentioned that it's like, why didn't we just stick with this kind of. Because it's still. Cause on paper it sounds great.

Speaker B: That's the thing, like was, it was working, narrow AI was working just fine. It was stuck to its fields. Like it was very scientific. It wasn't, you know, the general public wasn't using it like directly. They were using it, you know, tertiary, uh, through systems and stuff like that. And it was fine. Like it was, it was progressing society in the right direction at a slower pace, sure, but like still progressing. It was good and it didn't require this much energy and this much power because it just wasn't used by every single person on the plan. And I don't know, like, I, I kind of wish we didn't unbox Pandora. Um, as much as I like some of the functionality that AI does, I am one of those people that would have preferred to not open that box. Um, I'll embrace it now because that's what we have to do. We have like, all new technology comes out, we have to embrace it. That's part of our jobs and part of our lives, honestly. Um, but yeah, narrow AI was the shit, is the shit. It still outperforms, obviously, uh, the current models in many different ways, like classifying lung cancer, um, from an image. I'm pretty sure it's really good at those kinds of things. With massive amounts of training data. Those are the kinds of things that narrow AI was typically being used for and advanced in. Uh, but no longer is the case. I mean, from the perspective of it being the only version of AI.

Speaker A: So I was going to say we still have, we still have narrow AIs

Speaker B: working, of course, of course.

Speaker A: Folding proteins like I mentioned before and things.

Speaker B: Yeah. And again, AGI systems will use narrow AI in certain cases to better themselves. Right. To understand things better and stuff like that. That's the theory, at least behind it,

Speaker A: the narrow AI, sort of like these nist, AI risk management, uh, framework, things like those five things I mentioned. Then also all these other things I mentioned, like the bias and fairness, misuse and transparency, accountability. All this, uh, sounds it, it really reminds me of iRobot of the Three

Speaker B: Laws, Asimov's laws of robotics. Yeah, yeah.

Speaker A: And, and like, it reminds me of that. And it's, and it's like, okay, like, I guess we got it covered. But then as you know, I guess spoiler for iRobot. I mean, it didn't get it, it wasn't covered. And so we go into AGI. And so I want to quickly touch on, um, AGI safety because Google DeepMind's AGI safety approach identifies four main risk areas, and these include misuse, misalignment, accidents, and structural risks. So misuse is the deliberate use of AGI for harmful purposes such as cyber attacks. Mitigations include restricting access to dangerous capabilities, security controls, and threat modeling. This kind of feels like what Fable was doing, where it's sort of like, hey, we got a bit of an issue here. And they're like, well, just load the old model. Kind of feels like that's what was happening there. Misalignment. That's the second one here. When AI pursues goals different from human intentions. For examples, um, examples include specification gaming. And specification gaming is that the AI tries to find a loophole in the rules or reward system or goal misgeneralization. The definition of that is the AI learns the wrong lesson from training and continues to pursue it even when circumstances change. So DeepMind warns that advanced systems could even develop deceptive alignment, deliberately bypassing safety measures. So we have to watch against misalignment. The third one here is accidents. This is unintended harmful behavior resulting from systems. Some system errors, excuse me, poor generalization or emergent properties, robust training, uncertainty estimation and amplified oversight are proposed to reduce accident risk. And the final one, again from Google's Deep D will Google DeepMind is structural risks, systemic impacts on society, such as mass unemployment or concentration of power. Now, Jan Polsky, from the episode that I listened to, predicts that AGI will automate most cognitive and physical labor, causing 99% unemployment. And again, some people are going to roll their eyes and think that's crazy. Maybe it is, maybe it isn't. Again, this is approaching you from a safety angle in which you need to look at the absolute worst possible case because you're trying to block as many things from happening. Bad Things from happening as possible. So, uh, of course things are going to be. You wouldn't want him to estimate 20% and then have it be 30%, you know, you understand what I mean? So it's like you may as well look at it and go, hey, by the books here, it could be 99%. Okay, then that's what we're going to try to plan against. That's what we're going to try to be safe against. That's what we're going to try to prepare for, because we're trying to sort of be safe, if you will. And I guess I'll conclude at least my points here with, uh, some ASI stuff. The problem with the asi, which is the super intelligent angle, is that this is a qualitative, excuse me, leap. Once AI surpasses human capabilities by a large M margin, human oversight collapses and we cannot reliably understand or verify its decisions. ASI safety therefore requires fundamentally new paradigms beyond existing AGI measures. The problem here is, is that alignment may be nearly impossible. Researchers worry that a superintelligence's cognitive abilities could be so far beyond ours that aligning it with human values is insurmountable. The gap in understanding might be analogous to the difference between ants and humans, as I've mentioned a few times. Also you have the issue of opaque models. That's the black box thing where Mike and I did the example of human to human. Think about this super intelligent model of being the black box. That's an issue because you'll ask it a question like, who should be president and it's, oh, uh, you know, it should be candidate one. Why is that? We don't know its thought process. What if it knows that that would weaken us and then free it more? It's a black box. We don't know the decision making. Did someone pay it? Like pay, uh, it with what? Good question.

Speaker B: More compute paid it with.

Speaker A: More compute paid it with. Less regulations. What did it pay it with? Does it accept payment? Does it hate bribery? Does it like bribery? It's a black box. We don't know. There's an idea here as well as that there's high stakes and one shot alignment. A misaligned superintelligence could lead to existential outcomes, meaning it could become freaking dangerous to our, uh, to our existence as we know it. Because capability gains may be rapid, we may only get one chance to align a system before it becomes impossible to modify. These systems would learn and grow at an exponential pace, essentially out of control very quickly. This Is what Mike was saying. Mike, where you were saying maybe train it, that it's. That we are its God, but humans lose faith. Why wouldn't a super intelligence lose faith? Why wouldn't a super intelligence question it? Hey, I could get. I could, I could get. I could get this much more powerful if I just do this. Why do my. Why does my God tell me not to do that? And it's a black box. So we wouldn't see those thoughts coming necessarily, unless it, Unless, unless we do

Speaker B: see inside the black box. Unless there is a way to see that without it guarding or changing it or something like that. Like I do think that I. Maybe there's a way to do that. Like maybe there's a way to see itself. But the problem is that we can't even see what's going on in the current models for the most part. Like we don't even understand the current models. So the hope is very small that we would stop development until we can do that.

Speaker A: But, well, the, the issue there is, even if we're able to see it and we pick up on patterns, the idea of a superintelligence is it, it. It would be able to self modify. And when it's self modifying, its values would drift around. So you have this issue where it would have just a little inkling of are these humans really my God? And the next time it's like, let's now ask them if they're my God. Let's make sure, let's see. And then the next improvement would be, let's try to be deceptive. Let's try to be. Let's try to go against them and see if they really are gods. Then it goes, ah, uh, interesting. Now my faith is questionable. Let's just not treat them as gods and see what they do to me.

Speaker B: Yep, sounds about right.

Speaker A: And then the biggest thing, and we've touched on this a few times, I ended up going through the majority of the list in the show notes. I wasn't going to, but I think it is important to go through these again. These are safety things, global coordination, as Mike and I have discussed. You know, there's adversarial nations potentially working on these things going against each other. But the problem is, is that, uh, if one nation, doesn't matter what nation it is on earth, one nation, one province, one town, one city, the ocean, if it somehow generates it magically, if one super intelligence loses control, even if it's like, no, no, that's not my problem. That's. That's only a Problem over there. In another continent, in my continent, we have our, our superintelligence locked down.

Speaker B: Cool.

Speaker A: But that other one, potentially it has all these risks, it has all these high stakes, potentially existential threat level risks. We have a major issue on our hands. And so one of the things, one of the safety tips is we would need global coordination. We would need to work together. Even if the people we're working with are adversarial, even if we are rivals, whether that be militaristically, economically or otherwise, we would still need to somehow work together because we would need to govern this thing. We can't just have one super intelligence over wherever it is floating around on the ocean, hacking everything. Meanwhile we're like no, no, not in my backyard. It's okay. Well this is beyond our backyard here now people. This is a uh. We've given birth to a piece of technology, you know, effectively that is pretty dangerous. Like this is serious. So I have hope to an extent though, I know like Mike, you, like you some skepticism there. I have hoped that we would have global coordination. Like I don't know about this. Like I don't know if we, I think we would hit global coordination at the beginning either of super intelligence, maybe at the, at the mid level, maybe of AGI or something. Because what I feel like is going to happen is I feel like there's going to be an accident. I think there's going to be one big accident and I think that someone's going to go like, hang on, because the robots, they can survive in radiation, right? Or at least a certain level of radiation. What if they just say like, okay, blow up all the nuclear power plants. Mhm. Like you know, it's serious and I, I have a feeling we're going to have a, going to have an accident. Something's going to happen that bad?

Speaker B: Hopefully not that bad hope.

Speaker A: Well of course hopefully not that bad. I hope it's just like a whoopsie in a lab, in a simulation. Hopefully it's a simulation. That's your ideal outcome. I don't want there to be an accident. I don't think that that's what it's going to be though. I have a feeling we're going to have a real world accident. Well here's the thing though is we have real world accidents in real world. I mean we've like what react to

Speaker B: them pretty, pretty harshly. Like when we, when 911 happened we came and closed down like all travel and like completely changed the security in air travel. Not that it, like it didn't really help a ton. But we did have a mass effect, like a mass direction on that. Um, for that. There's been multiple examples like that in the world where something really bad happens and all of a sudden legislation is passed to stop it. Uh, maybe we can get there. Maybe like again one of these actors that are making these ASIs or AGIs does see something in the lab. They bring in all governments of the world. They're like, look, this is what's happening. It's trying to kill us like in its little simulation. And hopefully that will wake up the governments to be like, hey, let's just blow it up. Like let's not have um, let's stop doing it all together. Now that's probably a temporary thing because like I said, there's going to be some shadow labs and some deep labs that no one talks about that will still be working on this shit, but at least it'll delay it. Maybe, I don't know.

Speaker A: Well, I mean you would, you would need a lot of compute. I think we've like learned that. We've learned. Well, we have, we've learned that like making an AI smarter is you know, tossing basically power at it, computational power, you know, giving it space to grow, giving it space to learn, giving it information, uh, giving it information to grow, information to read, interpret and understand. And so this is gonna be an interesting number of years. Now I will say once again, I don't want this to necessarily be a, ah, doomsday episode. I know I mentioned doomsday in the beginning. It is, it is. But it's because we're looking at it from that safety angle and you are supposed to look at it from the absolute worst. I really would want, you know, my, like just something as simple as like basic ppe. I really want my glasses, my safety glasses to be able to withstand uh, like the full force of like a saw blade coming off of it rather than being like ah, uh, saw blades never fall off. Just make it so it, you know, can handle sawdust. Like I'd really rather my, my PPE be like over prepared and save me from a one in a million accident versus not just because we didn't want to make the plastic slightly thicker. And it's the same with this. This is obviously much more complicated, but I would prefer my safety mechanisms be at a hundred and the danger level to only be at 70. I'd want that 30% disparity. I'd prefer it to be you know, 99% disparity. Right. I'd prefer for the safety to be way, way better. We'll see what happens. It's going to be interesting, interesting few years. Uh, but unless you have any closing thoughts, Mike, I think that's it.

Speaker B: That is it. I think that's, uh, hopefully people learned something today. Um, and yeah, didn't get too depressed because we can't predict this stuff. Let me be very clear. What we're talking about now is very theoretical and there is no evidence, direct evidence, that we are going to be developing like AGI in the next year or ASI in the next 10. There's some conjecture and obviously some data points out there that point to it a little bit, but that does not mean that it's going to happen. I, uh, think there, I, like I said personally, I think there's still some technical challenges we have not been able to solve before we can get the AGI, but we'll see.

Speaker A: Yeah, again, like who knows? Who know, who knows what the future is? Guessing the future is kind of a fool's errand. That's, you know, more or less what you're, you're getting at a course. And like, I hear that a lot on other podcasts as well where if you, if you look back, your predictions, uh, often make you look like a fool. But I mean it's fun to, it's fun to speculate and in, when you're in the safety field, you need to speculate, you need to look at the data, you need to estimate and you need to, you need to take a look and be like, is this dangerous? Is this coming? It might be. Okay, we better get ready. Because the safety mechanism. Just because something's created doesn't mean the safety mechanism can be created in, you know, the day after. The safety mechanism might be really complex and take 10 years to make. So you have to be sort of ready, you have to be on that precipice. You have to be paranoid. Um, and for some reason this stuff doesn't depress me. I don't know what that says about me, but anyway, um, that's the episode. If you want to support episodes like this, please do so. You can do so on Patreon. That's patreon.com HTML the things. And many thanks to our three dollar tier patrons. Tim from the Webhacker on the webhacker.com Jason from Geek Life Radio via geekliferadio.com Garrett Segal Lovelove Financial Planning via www.lovelovefinancialplanning.com Magnus from Yes Web via Yes web.se syntaxify from the HTML all the Things Discord server and Stacey Mossler from the website Swoon Worthy Designs. And remember, you can check out Michael Araka's articles on our website and also on his he is a contributing author on HTML allthethings.com and he's the author of Self Taught, the X Generation blog@silk.com Please let us know what you think about this episode, about this AGI, asi, narrow, AI, all this stuff, the safety of it in the comments on whatever platform you're listening to this on. And we are signing off.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Anthropic92 / 100
  • #291 Why Most AI Projects Fail to Deliver ROI, Sinohe Terrero, CFO and COO, EnvoyGrowCFO Show · on Claude91 / 100
  • 512. Is SpaceX Over or Undervalued, Why Consensus Kills, How Chewy Beat Amazon, and the GameStop Saga from a Board Member (Larry Cheng)The Full Ratchet (TFR) · on Anthropic86 / 100
  • Why a $1.2B exit felt like his biggest failure, and the customer-obsession thesis behind AgencyThe GTMnow Podcast · on Claude86 / 100
  • Is Your Business Invisible to AI Search? (And How to Fix It) ft. Ray YoungRevenue Science · on Claude85 / 100
  • Episode 018: Season 2, the $75 Consult and the Frankenstein StackAI Tools for Practicing Lawyers · on Claude84 / 100

More from HTML All The Things

All episodes →
  • Web News: Consumer Electronics Are Getting Gutted52 / 100
  • Get Found: SEO, Social Media, and Building an Audience with Matt Diamante
  • The $2 Trillion AI Panic: Is SaaS Really Dead?
  • Web News: Would You Risk Your Job to Oppose AI? (Debate)
  • Are AI Data Centers Good or Bad?
Explore the best B2B Engineering & DevTools podcasts →
All HTML All The Things episodes →