Hidden Layers · 2025-06-18 · 40 min
Key moments - from our scoring
Substance score
49 / 100
Five dimensions, 20 points each
This episode bridges neuroscience and artificial intelligence by exploring how the free energy principle - Carl Friston's foundational model for how the brain processes information - can inform next-generation AI systems. Rather than the backpropagation-driven deep learning that dominates today, Friston and Dan Mapes (founder of Versus AI) advocate for active inference as a more biologically plausible learning mechanism. Active inference enables local, neighbor-to-neighbor optimization in neural networks, mirroring how actual brains learn without requiring global error broadcasts. The conversation addresses why large language models work so well (they're fundamentally prediction machines) while also explaining their limitations - particularly hallucination. Versus AI's approach wraps the free energy principle into a software development kit and creates domain-specific, hyperlocal AI models via the Spatial Web protocol rather than betting everything on centralized foundation models. This enables "collective intelligence," where distributed domain experts (cardiologists building cardiology AI, Kenyan farmers building farming AI for Kenya) collaborate through factor graphs and message passing. The discussion positions this as the next phase in computing architecture: from mainframes to the internet, from AOL to the web, and now from centralized LLM systems to decentralized active inference ecosystems.
The free energy principle frames the brain as a system that minimizes surprise by building predictive models of the world and acting to select sensory outcomes that resolve uncertainty. Learning happens through both perception (updating internal models) and action (changing the world to match predictions), unified as free energy minimization - mathematically equivalent to maximizing evidence for your world model.
Active inference enables local, neighbor-to-neighbor optimization where each component learns based only on signals from adjacent neurons, whereas backpropagation requires broadcasting errors across the entire network using the chain rule. This local message passing is how real brains actually learn, without needing a global supervisor.
Large language models succeed because their core objective is prediction - they're fundamentally autoencoding systems predicting the next token. This prediction objective, combined with sufficient model capacity and training data, enables emergent world modeling sufficient to handle complex reasoning tasks like narrative prediction.
Versus AI builds domain-specific, hyperlocal models (a cardiologist's cardiology model, Kenyan farmers' farming model) that are curated with accurate priors rather than scraped from the entire internet, then connects them via the Spatial Web protocol. This 'collective intelligence' from distributed experts avoids hallucination and enables better generalization than monolithic foundation models.
The Spatial Web is a decentralized protocol (analogous to HTTP for the web) that allows domain-specific AI models to interact through factor graphs and local message passing. This enables intelligence to emerge from cooperation among millions of distributed agents rather than a single centralized model, with protocols ensuring models can discover and interact with each other.
Our reviewer’s read on each dimension, with quotes from the episode.
A handful of genuinely non-obvious ideas surface - using LLMs as natural-language explainers for active inference machines, the 'amortization' of inference via deep nets, local vs. global optimisation as a practical architectural distinction - but large stretches are abstract theory, host throat-clearing, and mutual admiration that yields no actionable takeaway for an operator.
there's a perfect role for large language models here that can now learn your machine representations, your beliefs about the causes of your world and then map that to a narrative uh, that somebody could actually listen to
if I can act in a way to solicit the smart data, the data I need to resolve my uncertainty, I don't need to ingest the entire uh, web or universe
The free energy principle framing for AI architecture is genuinely non-mainstream in B2B tech discourse, and Friston's point about LLMs serving as explainers for active inference systems is a fresh synthesis not commonly articulated; however, Dan Mapes' decentralisation pitch and domain-specific-model argument are rapidly becoming standard talking points.
active inference is predicated on a world model or a generative model at uh, multiple scales. Um, that means it's um, effectively explainable
once you have a view of any distributed belief updating system that can be um, dissembled into a lot of local intelligence and belief updating, where I'm just Communicating with my neighbours
Friston is a legitimately world-class scientist and one of the most cited researchers alive, which is rare for any podcast; Mapes is a credible founder-practitioner, though his contributions trend toward promotional vision rather than hard operational experience, limiting the combined practitioner score.
A professor at University College London and scientific director at the Wellcome center for Human Neuroimaging, he's also one of the most cited scientists of all time
we created a spatial web protocol in partnership with the ieee
Almost no concrete business evidence appears - no customer names, revenue figures, deployment metrics, or documented outcomes from Versus AI; Dan's specifics are large-scale internet statistics and speculative AGI timelines rather than verifiable operational data.
we've got 40 billion computing devices plugged in now, and we've got probably 200 billion IoT devices plugged into it
it's clearly sometime, um, from probably 2030 to 2035
The host is genuinely knowledgeable about AI and asks one substantively probing question (on brain processes that resist the FEP), but he repeatedly validates bold claims without challenge, makes the conversation partly about himself, and lets Dan's promotional claims about Versus AI go entirely unscrutinised.
I'm a big fan of the free energy principle. Love active inference. I think it is just the first time I heard it I immediately thought oh my God, this just makes sense to me
Are there any physical processes in the brain that we know about that we can observe right now that currently don't fit that model
Computed from the transcript - who did the talking, and the words that came up most.
In this episode of Hidden Layers , Ron Green sits down with Dr. Karl Friston - world-renowned neuroscientist and originator of the Free Energy Principle - and Dan Mapes, founder of Verses AI and the Spatial Web Foundation. Together, they explore how neuroscience is beginning to reshape artificial intelligence. They break down complex but powerful ideas like active inference, biologically plausible AI, and collective intelligence. You'll hear how concepts from brain science are influencing next-gen AI architectures and what the future might hold beyond large language models. From the limitations of backpropagation to the promise of decentralized, embodied, and domain-specific models, this is a deep dive into the future of intelligent systems - and the science behind them.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Welcome to Inlayers where we explore the people and technology behind artificial intelligence. I'm your host, Ron Green. We have an amazing episode today. I'm joined by the world renowned neuroscientist, Dr. Carl Friston and his colleague Dan Mapes, the founder and president of Versus AI, a company building AI systems modeled on, um, natural intelligence. We discuss how neuroscience is starting to shape the way we think about artificial intelligence. Concepts like the free energy principle and active inference, originally developed to explain how the brain handles uncertainty and makes decisions, are now inspiring new models for AI. Carl and Dan explain how the brain works as a prediction machine and how those same principles are being used to design AI that can learn, adapt and respond more like a human. This is a deep and forward looking conversation about the future of Intelligent Systems. Dr. Carl Friston is a world renowned neuroscientist and theoretical biologist best known for the free energy principle, a foundational model for how the brain processes information and guides behavior. A professor at University College London and scientific director at the Wellcome center for Human Neuroimaging, he's also one of the most cited scientists of all time. In 2022, he joined Versus AI as Chief Scientist to help develop biologically inspired AI systems based on natural intelligence. Dan Mapes is the founder and president of Versus AI, a company building AI systems modeled on natural intelligence to power smarter, more adaptive digital ecosystems. He also leads the Spatial Web foundation which is developing its decentralized spatial Internet using the Hyperspace Transaction protocol. He's also the best selling author of the Spatial Web. Carl. Dan, thank you so much for joining me today. I'm really excited to have this conversation.
Speaker B: Wonderful to be here.
Speaker C: Great to be here.
Speaker A: Well, Carl, maybe um, for those new to it, um, you could give us uh, an overview on the free energy principle and how active inference builds upon it to unify perception, action and learning as a model for understanding both the brain and intelligent systems.
Speaker B: Yeah, so I normally ask whether you want the high road or the low road, but I'm going to give you the high road.
Speaker A: Okay, great.
Speaker B: Give me your technical experiment. So the free energy principle is really a physicist take on how things self organize and in particular it drills down on what it means to make sense of the world and act upon the world in a principled way to maintain effectively the boundaries and the integrity of you separate from your universe. Um, it transpires you can effectively describe being as observing and coupling to a universe and describe that in terms of a Bayesian mechanics, a kind of inference, um, that rests upon acting in a way to minimize the surprising encounters or interfaces or coupling to the environment. And that um, mathematically is exactly the same as seeking evidence for your world models or generative models. So the idea is that you can always describe effectively anything um, to a greater or lesser extent as in a sophisticated way, um, building a model of its lived world, trying to predict the sensory interactions with that world. But crucially you actually mentioned action and perception. So um, this is not just a one way street. It's not just the universe impressing itself upon your sensory organs and you try to make sense of the causes of those sensations, but you actually couple back to the world and you act upon the world specifically act in a way to select those data, those sensory outcomes that are going to be most informative in resolving your uncertainty about your world model or your generative model. At uh, its very simplest it's really a question if you're trying to minimize your surprise mathematically this free energy bound, which incidentally is exactly the same as an evidence, uh, lower bound, an elbow, uh, in machine learning, um, if you're trying to sort of optimize this bound on surprise or um, the evidence for your models of the world, um, then you're going to try and realize your predictions. So one way of minimizing surprise is just to act change the world in a way that it supplies exactly what you predicted. And that sort of also licenses uh, a nod to predictive coding which would be at the moment one of the sort of forerunners in terms of a computational account of how the brain works. By um, you're making predictions about the sensations at hand, matching them, um, and um, comparing them to produce a prediction error, which is a simple form of free energy, and then acting in a way to either change your mind to resolve the prediction error by changing the predictions or acting upon the world to change the things that are being predicted, both minimizing this prediction error of free energy. So that's the basic story that has action and m perception all wrapped up.
Speaker A: Yes, and I love it because um, we were talking before we started recording. I've been working in uh, artificial intelligence for several decades now. But m, my work is mostly around artificial uh, neural networks which are typically called sort of deep learning systems. Now what's really interesting is that those systems were originally designed to mimic structurally part of the brain, right? You know, the neurons and the synapses. But they weren't really fully complex in a biological way. And they certainly didn't have like a biologically plausible way of being trained or learned back propagation which was invented decades later, um, is I think you know, pretty deeply implausible but active inference which grounded on the free energy principle, you know, it offers this much more realistic model of how the brain works and adapts. Can you kind of tie those two concepts together um, and articulate, articulate a little bit more about the biological plausibility and why it is so much more um, um, tied to our own brain systems than the deep learning systems that most people are familiar with and interact with today.
Speaker B: Right, no, that's an excellent, excellent question. Um, and just in defense of machine learning, uh, there are I think certain beautiful aspects of the architectures of these deep uh, neural networks that do inherit and do actually mimic um, the way that the brain works. And I'm thinking here specifically about the kind of hierarchy implied by the word deep. There's a hierarchical depth there which is I think quintessentially a really important feature nature of any adaptive neural network that's trying to model its world simply because the world, the universe, has this deep non markovian structure with separation of scale and the like. I think that should be the deep part of deep learning. Should be absolutely applauded. Um, but if you want a critique, on the other hand, um, yeah, the actual implementation uh, in terms of back propagation of errors does depart from um, a neuromorphic or biometric or a natural way of optimizing these neuronal networks or neural networks. And I think in two interesting ways. Um, one way which I think may provide an avenue for discussion about the difference between reinforcement learning and active inference is that the objective functions are ah, not fundamentally different. But to assume the existence of a reward function is quite a bold move. One would have to carefully justify in terms of self organizing systems and the free energy principle. So we can talk about that later. But um, beyond that um, what is entailed by back propagation of errors is the notion that you're not doing it locally. And the whole point of predictive um, coding formulations that are inspired by um, a more biomimetic approach is that your objective function can be evaluated locally which means that you can do your learning just based upon your neighbours. And uh, if you can do that then you don't need to broadcast and use the chain rule to do back propagation of errors and leads to a much more biomimetic kind of learning process which is just a free energy minimizing uh, process. So I think people are starting to um, get into that and certainly the Oxford group have sort of made quite um, Compelling advances bringing the notion of predictive coding to the table in the context of machine learning. Um, at the moment there has been a slight paucity of work in terms of um, making this work in a dynamic context, the context finding transformer architectures for example. But certainly the basic idea of local message passing and local optimization, very much of the kind that Joshua Bengio uh, has been promoting in terms of expectation propagation that I think is a cardinal um, or that is one bright line you could draw between propagation of errors and a more biomimetic uh way of doing that. The problem is though, uh, the scaling, you know, can you get predicted coding to do all the beautiful things that current, for example large language models today?
Speaker A: Yes, um, it's unrealistic this idea that if you um, interact with the world and you learn something that you might update every part of your brain in some way which is how these deep learning systems work. Um, you mentioned large language models. I'd love to talk about that a little bit. Um, so M. Scott Aronson, the theoretical um, uh computer scientist known for his work in quantum computing has said that the emergent abilities of large language models at scale for him is the most significant scientific discovery of the 20th century. I tend to agree with that. I was completely surprised and blown away that these models would exhibit um, emergent behaviors just by scaling the count and the training data that they were trained on and they began to exhibit these behaviors, behaviors and skills that are not explicitly programmed. Um, it's almost like uh, intelligence started to appear just from sheer complexity. Um, what are your thoughts on the abilities of large language models and have they changed or sharpened any of your opinions around intelligence or consciousness?
Speaker B: They certainly endorsed um, views of sentient machines and you know, um, machines that, or systems uh, that can show these emergent properties that are now approaching intelligent like behavior. Not quite natural in terms of AGI, but certainly I think going in that direction. Um, and I think there are two, um, there are two themes that
Speaker A: um,
Speaker B: may provide a potential explanation as to their competence and the, I repeat, the very beautiful behaviors that you can invince and that emerge when you uh, endow them with sufficient degrees of freedom. The first one is just to ask, well what is the objective function that one is using now in contrast to sort of generic deep rl where you've written down what you think is the good way of behaving in terms of cost function or reward function? I think large language models very much in the spirit of autoencoders or specifically variation autoencoders for example, there's an autoencoding aspect to it in the sense that you are predicting content. So the objective is to predict. So if you just look at how you would train a large language model, yes, there is human reinforcement in the loop. But at the end of the day, what we're talking about is a prediction machine, possibly a limited one in the sense of one step ahead, but just sort of acknowledging that the objective is prediction, I think, um, is important in trying to understand why they perform so wonderfully. The second, of course, uh, they're doing it in the domain of language so we can understand and we can admire and engage and appreciate the emergent world. So I think those two themes explain what all the excitement and the potency, uh, of large language models.
Speaker A: Yes, yes. Um, one of my favorite, I think, um, thought, um, experiments with this is, um, the interview Ilya Sussex gave where he said if you trained a large language model, um, and then you gave it, uh, like, uh, a mystery novel, right? And you gave it everything up till the very, very last sentence where the murder murderer was revealed, that these models would have to understand, have sort of a world, um, view strong enough to be able to predict, as you said, who did it. Uh, for lack of a better term, um, I want to pivot Dan now and ask you what has it been like taking this incredibly important, um, research that Carl's developed and trying to bring it into the real world, trying to build intelligent systems that are based on these, like, much, much more grounded, biologically plausible theories?
Speaker C: No, sure, Ron. Um, well, fundamentally, um, the Internet, uh, itself, which, uh, has been the largest engineering project humanity has ever undertaken. I mean, when you think of all the servers and fiber optic cables and satellite sensors and, I mean, we've got 40 billion computing devices plugged in now, and we've got probably 200 billion IoT devices plugged into it. This system is a framework that's sitting out there. And so, uh, the way it works is a decentralized development model where we create protocols and we enable anybody anywhere in the world to build a website. And then the protocols then allow those websites to interact with each other and ultimately be uh, discovered with, uh, Google search and things like that. So then we get a worldwide web of information much larger than any one company could ever have done. I think we have 5 billion websites right now. Uh, so Microsoft or Google could never have built all those. So, uh, what we're doing with Carl's, uh, chief scientist of, uh, versys, and what we're doing with Carl is Uh, we're uh, um, wrapping the free energy principle into a software developer kit. And uh, then we created a spatial web protocol in partnership with the ieee so that anybody anywhere in the world, instead of building a giant 10 trillion parameter foundation model and everybody in the world has to query that model and kind of mining it with prompts and things. No, no, no, no, no. We want domain specific models. So we want a cardiologist highly curating a cardiology model for medicine. And this is uh, what then the free energy principles operating in that accurate uh, model it's already curated. Whereas in the large, in the large language models, I mean you basically throw, scrape the entire Internet and throw it in there and every book that's ever been written and then you mine it and hope for the best and it hallucinates. I mean it just does. And so by beginning um, what uh, what Carl would I think call your priors, beginning your priors with accurate domain specific models and having them be hyperlocal. So people in Kenya building AIs for farming in Kenya, but people in Japan are building AIs for farming in Japan which is very different. That's what we're doing. We're kind of creating a global decentralized framework for this new uh, AI to be taken up by developers all over the world. And um, they build their own um, local highly curated uh, domain specific models. And then the protocols allow those models to interact with and you get something that Carl has been calling collective intelligence. And you might want to ask him to speak about that a little bit. But it's going to be way larger than any single foundation model by Google, uh, or Tropic or Meta or whoever. And so uh, you've got 100 million people working on this project instead of uh, uh, 10,000 or so. So I think that's a really fundamental place to start um, moving and things historically have moved from central DC central mainframes to the Internet, aol.com to the world Wide Web and now centralized uh, big uh, AI LLM systems to a decentralized active inference uh, system.
Speaker A: Okay, perfect. Well I would love to hear your take on uh, collective intelligence Carl, and maybe describe it, unpack it for our listeners a little bit first before you jump in.
Speaker B: Sure. I think actually it um, rests upon what we were talking about before about the distinction between back propagation and errors and this local optimization. So once you have a view of any distributed belief updating system that can be um, dissembled into a lot of local intelligence and belief updating, where I'm just Communicating with my neighbors. I think you have a good picture of this kind of not federated learning in the computer science sense, but federated inference, distributed intelligence. You have this notion that now any deep neural network can actually be partitioned into an ecosystem of little local intelligent agents. And interestingly what that entails is um, a separation not only of a distributed network into lots of little local sub agents and that could be on the edge or part of the web, um, but also separation in time that some agents sort of um, take a broader picture because the way that they're connected to other agents. So these would be like the deeper layers or the more central layers in a centrifugal architecture that um, see more and um, think more slowly and contextualize. Coming back to your um, nice, your thought experiment. Would a large language model ever really um, be able to predict the denouement of a long narrative, a long novel? Of course to do that it would have to see um, a very large context window and would have to assimilate things very very slowly. And other things will be moving very very fast, possibly driving cars or navigating around uh, your particular website. So the notion of distributed intelligence, um, is if you like, an emergent property of understanding the imperatives that enable local components, members um, of an ensemble that constitute an ecosystem that is constructed with, with its own structure. I think is quite central just technically. Um, one of the things that Dan and his colleagues um, brought to the table conceptually that struck me as exactly the right way forward in terms of understanding uh, or architecting. The next version of the World Wide Web was the isomorphism with factor graphs in computer science. So one way of expressing this locally optimal way of belief updating or inference, or active inference, um, is in terms of a graphical model. And a graphical model always has a dual, which is a factor graph. You've got a factor graph, then you have a very precise recipe for the message passing that is all local and can all be configured in terms of um, maximizing an elbow or minimizing a variational theology. They're the same thing. So just the architecture of the message passing and um, the potential, the possibilities in a suitably engineered World Wide Web, uh, to me remarkable because there was a recipe, there is a recipe for doing this in the right way. If you commit to things like the free energy principle and active inference as the principles of, of self organization.
Speaker A: You know this, you remind me, uh, in the, in the 90s in grad school, I remember uh, one of my professors was saying, you know, just how brittle um, modern software is right, you, you can build you know, a 10 million line coded system and you can accidentally misplace one semicolon, right? And everything breaks. And that what we needed were much more uh, biologically inspired systems that were capable of being uh, not necessarily really self organized but could um, be hierarchical and componentized and you could breach in and maybe remove some part of it and it would still function. And I think we're seeing this um, more and more across different AI systems. And when I think about um, the architecture you described, a mixture of experts come to my mind. That's some approaches that are being used within the language model world. Um, so Carl, I've got a question for you. I've always wondered this. Um, I'm a big fan of the free energy principle. Love active inference. I think it is just the first time I heard it I immediately thought oh my God, this just makes sense to me. But I'm kind of curious. Are there any physical processes in the brain that we know about that we can observe right now that currently don't fit that model, that in any way, um, um, remain unresolved with respect to that theoretical foundation?
Speaker B: Um, it's going to sound cheeky and polemic, but no. So you know the reason that um, I think more and more people are ah, um, interested in the free energy principle is it just has such a broad explanatory scope now. It's not very useful in and of itself. So um, very much like the theory of natural selection is a beautiful theory and it explains everything, but it doesn't tell you how to make an eye or why you have legs. Also in English. M. It doesn't do any of the heavy lifting. We've got, still got to do the heavy lifting. And technically that basically means writing down the structures of the generative models and all the updating processes, the free energy minimizing processes. Um, and that's exactly what we've been talking about. Um, however, um, what one realizes is that you can apply the free energy principle at many different scales. So if you apply it in the moment in terms of say message passing of the kind you find in a transformer architecture, this can be regarded as inference because your disbelief, updating about things that are changing, unknown variables, random variables, latent causes of your content or your sensations that are changing very very quickly, you then apply the same maths to things that are changing more slowly. And then you get into these contextual variables that would be read as attention heads in transformer architectures or literally attention, uh, of A psychological sort in terms of um, ah, neurobiology. And you move to longer time scales and you get now continual learning the parameters of the generative model that um, install the contingencies, the laws that are much less time invariant. You can then take it right through to um, neurodevelopment and um, Bayesian model selection read as structure learning as we develop. And of course you can now take uh, Bayesian model selection that is predicated on scoring and selecting the models with the highest elbow or the minimum, uh, free energy that now becomes a mathematical image of natural selection. So you can now take it to the evolutionary scale and subsume evolution algorithms on this to get the very structure right. So I think just realizing you can apply the free energy principle at uh, multiple scales and every scale contextualizes each other, um, enlarges its expiratory scope. So I have yet to find. And um, interestingly if I did, I would seriously question my commitment to the free engine. But I've yet to find something that cannot be explained by the Voegyn.
Speaker C: I love it. Okay.
Speaker A: I'm delighted by the answer. And I also love the different fields you brought in. Um, I was talking with Risa, uh, McLennan, uh, a couple of months ago. The uh, uh, very uh, well known, uh, sort of computer scientists use evolutionary techniques. And I really feel like we're starting to see this um, resurgence of ideas from the 80s and 90s come back, um, reinforcement learning, evolutionary algorithms, um, and realizing that it's not really one or the other, it's these mixture of techniques that we might need to look at holistically to finally crack and deeply understand intelligence and consciousness and things like that. Dan, I would love to ask you. You're on the forefront of building these systems, trying to build production systems using these theories. Are you taking elements not only from active inference, um, but from other fields, whether they're evolutionary approaches or variational autoencoders or aspects from uh, the deep learning world, uh, to build these systems versus AI.
Speaker C: I mean there's a big um, um movement right now in the uh, in the AI world to move um, uh to uh, kind of an agent, uh based world. And uh, so there's a lot of people banding the term around agentic. Um, but really to um, to be an agent you need to have agency. And so we're seeing that a lot of the things that are called agents are really simply workflows and they're pretty brittle.
Speaker A: I agree with that.
Speaker C: And if they run into a problem, they just freeze up and Stop. And so um, it looks like there's an actual partnership model here where um uh an active inference AI could kind of function at a higher level almost like a general manager. And when a workflow gets into a problem it could step in. And uh, because the um agent, the true agentic qualities of active inference uh can function as a reasoning engine or um, um uh make planning decisions that an LLM isn't really suited to do. Um so I think uh we'll see a handshake and a partnership between um, people that are building workflows uh and a lot of corporations are doing those. Uh and then maybe this new uh, more advanced reasoning system that um, that uh, can kind of come in and function almost the way a manager of a team of people might. And um, and so when something gets stuck uh, the uh, the active inference agent can come in and um, redirect it and um, make a reasoning decision on the next path. But I think uh, it gets into fundamental issues too and that is LLMs generally are um machines. Uh and so we build them and then so we have a GPT2, then we have a GPT3, it takes a couple of years, then a GPT 3.5 and a 4. Um but we really with active inference we don't do that. Um it really grows up more like a child. Um we build a domain uh specific world model and then um, that world model is hypothesis testing um the real world and updating itself in uh real time. And so it's a self evolving. It's called autopoetic or autopoiesis technically term uh it's an autopoetic model and really, really arguably the first uh AI to be fully autopoietic. And so um, so then you get uh wow, you've knocked your cost out a lot. You're not building these huge foundation models that uh, are you know more ah brittle and a little bit more static. You have a smaller uh lightweight domain model that can evolve through time and on top of it all um, they're not interior models. When I ask the um, AI to do something and it doesn't only look within itself to answer, it's embodied uh intelligence that means that it's living in a network uh like the Internet and it has access to IoT devices so it can read all kinds of sensor data and that becomes important so that opens it up real time applications. Right now if we want to do a real time application we can't do it because the LLMs are hallucinating. And that means you've Got to have a human in the loop to edit whatever the result is. Whether it's a report or a video that got made or whatever, it's great, it saves us hundreds of hours, maybe over time, but still we have to edit it because it randomly hallucinates. Whereas if you've got a curated um, world model, uh, built by domain specific uh, expert, uh, and uh, then it's living in a real world environment that, where it can measure things. Now you've got a smart city application, you've got a smart supply chain application, you've got a smart factory application, you've got a smart hospital application. These are things that we can't safely get at right now. We can't do mission critical uh, real time applications with LLMs. But active inference is actually built for that. And so therefore it's, you can see, it's built perfectly for robots because it's embodied. It can look through the eyes, it can move an arm, it can direct things around, it's spatially aware. Um, so I think we're going to see a new class of applications uh, for AI unfold with active inference as it um, spreads uh, uh, throughout the world. Carl, you might want to see a few more things on that. Uh, in other words you could argue uh, LLMs in general are good for content creation and that kind of thing. And uh, active inference is really good for operations, you know, real time operations. But Carl, you might want to give some color to that.
Speaker B: I, I would if we have time. I think there's, I, I think you've
Speaker C: got, you've got, you've got three more minutes. You've got three more minutes.
Speaker B: Right? Very quickly. Yeah, I think that there are two, well, three things I'd very briefly like to sort of foreground um in that response. I think the notion of agency, um, is crucial here. To be an AGI, um, I think is to be an agent. To be an agent you have to have a world model or a generative model of the consequences of your actions. And I think that is basically what active inference brings to the table. And then you get into the world of okay, what are the objective functions that I would need in order to score whether I'm going to act like this with these uh, yet counterfactual unobserved outcomes as opposed to that and some lovely maths that brings you into the world of expected information gain and um, expected cost and the admixes of the two. Um, and that's really the heart of active inference. It's all about um, how am I going to act and of course if I can act in a way to solicit the smart data, the data I need to resolve my uncertainty, I don't need to ingest the entire uh, web or universe, it's just those data that I need. There's a remarkable potential efficiency by affording um, a large language model, for example with the ability to now select its own training cohort to specialize um, in the context in which it is actually operating, if it's a cardiologist or driving a car for example. Having said that, I do think there's lots of wonderful um, engineering uh, that's gone on over the past decades that could be very usefully harnessed within an active inference framework. I'll just give you two examples. Amortization is one. So this is basically if you can get a sufficiently context invariant mapping from some input to some variational um, density or probabilistic representation of the causes, then that is learnable and then you can replace that mapping which would otherwise be done with belief updating and variational message passing with um, a deep neural network. So advertising, habitizing, inference, learning to infer I think would be one really um, powerful and is one really powerful application of say deep learning and possibly even um, transformer architectures. The other one is slightly more entertaining um, because active inference is predicated on a world model or a generative model at uh, multiple scales. Um, that means it's um, effectively explainable because you're actually representing the causes of the content that you'll have at hand. If you wanted to surface that explainability. There's a perfect role for large language models here that can now learn your machine representations, your beliefs about the causes of your world and then map that to a narrative uh, that somebody could actually listen to. So I think there's a really simple and powerful application large language models specifically just in doing the heavy lifting in actually explaining what an active inference machine or architecture actually believes, um, not only about the world but why it is doing this, explaining why I'm going from the expected information gain or I'm going to minimize that expected cost, that particular constraint and um, actually have it declare its own intentions and beliefs.
Speaker A: Oh, I love that I have not heard that ah idea before. That just makes complete sense and it goes back to what we were talking about a moment ago which is this merger of different approaches and different ideas, um, and not being limited through seeing everything through just one lens or one set of principles. Um, uh, Carl, I've got a question for you. I'm really interested in this. So you've cited thinkers from Helmholtz to Hinton, Schrodinger to Charon. Um, if you had to narrow it down to one or two minds that have shaped, you know, you're thinking on the free ninja principle, uh, who would they be and why? I'm dying to hear this answer.
Speaker C: Right.
Speaker B: Uh, yes. Um, I'm going to cheat. I'm going to give the answer that, uh, Jeffrey Hinton has, uh, given in the past. And specifically last time I heard him give it was, uh, at his uc, the Ulysses medal award. It would have to be Helmholtz for exactly the reasons that articulates, you know, you've got unconscious inference on the one hand and, um, the notion of Helmholtz free energy on the other hand. If Helmholtz, I think, had been able to chat with Richard Feynman. Who's my hero? Richard Feynman's my hero. Um, but if I could go back and listen to a conversation, it would be between Feynman and. Hell yeah. I think if they got together to together, we would be, uh, where we are now in the 1930s.
Speaker A: Oh, I love it. Feynman is absolutely one of my heroes. Uh, I read, um, Surely you're joking, Mr. Feynman in college and oh my gosh, it's one of my all time favorites. Um, okay, I'm going to wrap this up. I have one more question. Um, this is for both of you. I'd love to hear your thoughts on each of this. Um, if you could have one, you can have one magic wish granted. You could have any unresolved question, open question that you've been wrestling with in your research throughout your whole career. If you could just have the answer for that question, what would each of you pick?
Speaker B: I'll go first. Give Dan some time to come up with a clever answer that I'll give. Um, it may be a strange and quirky one, but I'm still very puzzled by the nature of time. Why is time so unique? I understand the physics of everything else, but not time. Why is that so special?
Speaker A: Love it.
Speaker C: That's wonderful. Yeah, I mean, uh, it's funny, I mean, because one of the big things we're bringing to the table is space that we call our protocol the spatial web. And so, um, up until now, our, um, World Wide web has been in two dimensions. It's pages. And that's why it's called Facebook Page Rank. It's all about pages, but, uh, the spatial web is about spaces. And so, um, what's happening While we're sitting here right now is people are making digital twins of every building in the world and every city and the entire surface of the planet. So um, it's going to be really amazing to watch how uh, active inference and LLMs are able to inhabit a perfect digital twin of the Earth, um, that's being updated moment by moment with satellite data in real time. I think it could be possible that it becomes a planetary management system that helps us manage our climate and um, our flows of uh, energy and other kinds of things. So um, you know, uh, uh, just watching this whole thing unfold over this next, uh, five to ten years and when AGI, uh, pops up, I mean, um, uh, when that happens, when that uh, crossover happens. So, um, uh, it's clearly sometime, um, from probably 2030 to 2035. And um, what uh, an exciting historical moment that will be. I mean it's a uh, a moment like uh, the invention of fire or language or mathematics. So I think uh, this is an extraordinary moment to be alive. We're really excited to be right in the middle of it all and to be partnered with such a wonderful, uh, scientist like Carl and uh, to. To really explore this together. So. But uh, yeah, let's all have a great next five to ten years and then, wow, uh, we got Hei probably.
Speaker A: I totally agree. I share your optimism and I truly feel fortunate to be alive in this moment of time in the 90s working in AI. I don't know if I would have, uh, bet on it and certainly maybe not 15 years ago. So I just couldn't be happier. Thank you both so much. Dan, Carl, this was just a wonderful conversation. I really appreciate you both making the time.
Speaker B: Thank you.
Speaker C: We had a great time. It was really good. Thanks a lot.
Speaker A: Thank you for listening to Hidden Layers. This series is hosted by Kung Fu AI, a management consulting and engineering firm focused exclusively on artificial intelligence. If you have any questions or thoughts about today's episode, or if you know someone we should feature, please visit us at Kung Fu AI.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.