
MSP Cyber Roundtable · 2026-05-19 · 58 min
Key moments - from our scoring
Substance score
42 / 100
Five dimensions, 20 points each
Matthew Warner and Zoe Lindsey from Blumira join the host to challenge the assumption that frontier AI models are necessary for all tasks. Rather than paying premium per-token rates for models like Claude Opus or GPT-4, the conversation pivots to a tiered intelligence approach: using smaller, locally-runnable open-weight models like Google Gemma for routine tasks while reserving expensive frontier models for complex problems requiring deep inference. The guests describe this as an L1-L3 staffing model - analogous to IT support tiers - where L1 handles simple classifications with small models, L2 addresses moderate complexity, and L3 engages frontier models only when needed. This approach offers cost control through capex (purchasing GPUs and hardware) rather than subscription opex, data privacy by running models locally, and customization opportunities. The discussion emphasizes that smaller models require more verbose prompting and context but can run efficiently on standard hardware like MacBook Pros or AMD GPU setups, making AI accessible without enterprise-scale infrastructure.
It's a tiered approach where L1 uses lightweight open-weight models for simple, routine tasks; L2 handles moderate complexity; and L3 reserves expensive frontier models like Claude Opus only for problems requiring deep inference. This mirrors traditional IT support tier structures.
Yes - models like Gemma can run on a MacBook Pro or AMD GPU setups costing under $2,000-$10,000. With proper prompting and context, they handle the majority of MSP use cases without cloud infrastructure.
Smaller models have less training data and inference ability, so they need 3-4 pages of detailed context, explicit instructions, and step-by-step guidance to deliver accurate results, whereas frontier models can infer context more independently.
Frontier models like Claude Opus charge per-token via API subscriptions, costing hundreds of dollars daily for continuous use. Open-weight models running locally have a one-time capex cost for hardware ($2,000-$10,000) and negligible ongoing operating costs.
Data stays on your infrastructure instead of being sent to third-party APIs, eliminating concerns about proprietary information being used for model training or retention by providers like Anthropic or OpenAI.
Our reviewer’s read on each dimension, with quotes from the episode.
There are genuinely useful nuggets buried in the episode - distilling small models for specific tasks, using frontier models as judges of smaller model output, and the orchestrator-handoff pattern - but they are deeply buried under extended personal backstories, AI hype commentary, and meandering conversation. The episode title promises an L1-L3 staffing model framework that never gets coherently articulated.
we end up providing it with like three to four pages worth of prompt. So it knows what is it looking at, what does it need to do, how does it need to respond? You can get there and it can run on a, um, very cheap MacBook Pro
use the foundational models and really slim ways to help you build up those prompts, build up those contexts, understand how you're going to talk to the thing, but don't use it for actually running the thing
A few genuinely fresh angles emerge - using frontier models purely as validators or judges for smaller model outputs, and the AMD GPU cost arbitrage argument - but the bulk of the conversation recycles well-worn open-vs-closed model discourse, AI hype skepticism, and general agent framework hand-wraving that circulates everywhere in tech media.
using those frontier foundational models to validate the output of a small model, to act as the judge, essentially
I can have 48 gigs of vram right here for $1200, which you can't do with an Nvidia chip the same way
Matt Warner is a genuine practitioner with hands-on AI experimentation (local inference rigs, CAD distillation project, real pipeline testing with Gemma) and 20 years in cybersecurity, giving him credible authority. However, this is effectively a vendor appearance by the co-founder and a colleague promoting Blumira, which caps the caliber ceiling; neither guest is a scale operator or external expert.
I've been working on this project for about six months now. And the biggest thing I've learned is that we are better off doing exactly that. Taking a small model, giving it what it needs to actually be able to answer the right question
over the last two weeks we've seen a significant increase in codecs accessing SSH keys, for example, which tells me that we're seeing more application of people using chat GPT
The episode has a reasonable density of concrete hardware and parameter specifics - VRAM prices, model sizes, cost comparisons - which is above average for an AI podcast. However, many broader claims (about MSP adoption, agent disasters, cost sustainability) are asserted without data, and the Blumira product discussion is largely abstract.
E4B is their, their mixture of the small mixture of experts. That's only 8 billion parameters. Like uh, you can fit that into a thousand dollar computer
I can have 48 gigs of vram right here for $1200
The host is a technical peer which creates authentic back-and-forth, but the episode suffers from undisciplined structure: 20+ minutes of origin stories before substantive content, frequent host monologuing, no meaningful pushback on any claims, and an abrupt ending that never synthesizes the promised L1-L3 framework from the episode title.
Well, guys, we're running long on time here. We really need to hear what Blue Mirror is doing here.
I feel dumb right now that I don't know how to use these tools.
Computed from the transcript - who did the talking, and the words that came up most.
Join Matthew Fisch of FortMesa on the MSP Cyber Roundtable, along with special guests, Matthew Warner and Zoe Lindsey from Blumira, as they discuss how most organizations take a model-first approach to AI, leading to costly systems handling routine work. This session introduces a tiered LLM strategy - using lightweight models for high-volume tasks like triage and summarization, and reserving larger models for escalation - to optimize cost, performance, and scalability.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign.
Speaker B: Welcome back everyone to another episode of the MSP Cyber Roundtable. Uh, we're here in your social feeds and maybe in your podcast feed. Um, I've got with me today some friends from Blumira. Am I saying that right? I hope.
Speaker A: Yeah, you got it.
Speaker B: Uh, great. Matthew Warner and Zoe Lindsey. Uh, thanks for joining us.
Speaker A: Excited to be here.
Speaker B: We're here to get geeky, uh, about AI with a slightly different topic than normal because we're usually talking about how AI is. You know, the front edge of AI is changing the world. I think we all know that it can get very expensive and some of these teaser rates on these frontier models are like completely imagined. Like, um, uh, intro, uh, like starter drugs and you know, but the reality is we're not all going to have $10 million clusters to run our AI for our business in. And there's a whole world of intelligence out there. And um, today is really about, well, what's the right level of intelligence for the right level of task and how do we create this integrated work team with those different levels. Um, and I think, Matthew, you've been really excited with uh, this new, um, open weights model from, from, from Google Gemma, for now, is it. So we'll, we'll get into that. But before we do that, let's just, um, let's do the personal intro. So, um, Matt Warner, Zoe, uh, Lindsay, who wants to go first? I want to hear, I don't want to hear the blue Mirror story. I want to hear what landed you in Blue Mirror. I think that's like, where did you come from? How did you get in it? Slash security.
Speaker C: I feel like in that order, Matt has to go first because there, there's not a blue mirror to land in before Matt.
Speaker A: Sure, I'll, I'm happy to start. So, uh, you know, like all, not all of us, a lot of us in this industry, um, you know, I, I was the person who got a computer when I was 5 years old and then kept going with it.
Speaker B: So that is, you got to stop there. You got to tell us what the computer was.
Speaker A: Well, it was a variety of them actually, because my dad was a electrician at Ford. And back in the early 90s, late 80s, Ford didn't care about successfully getting rid of their hardware. They just didn't care. No one cared back then. So he would just take all of their computers and then I would just have a room of computers. So a lot. Basically any computer that Ford motor company would have had in the late 80s, I had in the early 90s.
Speaker B: Yeah, and they probably had all the things, right?
Speaker A: All the things, all the things. And because I grew up in Dearborn, Michigan here, uh, in Michigan, we were one of the first cities to have DSL because Ford wanted the whole city to have dsl. So I had high speed Internet and pretty much every computer I could have wanted that was used in the early 90s. Uh, and that was not good for anyone realistically like that. That was not a good thing probably. But it was great for my future, future kind of where I was going. Uh, otherwise I am that normal startup guy in that I went to a bunch of schools, dropped out of them because I hated learning from them. I thought university was a waste of my time and I really wanted to work inside of it. You know, I, I remember my first, uh, help desk job in IT was working inside of healthcare doing uh, support for password resets, the usual kind of thing. And that was when I learned, I don't know, I was 18, 19 this time, that surgeons will happily scream at you when you don't reset their password fast enough. And it's like, great. This is. We also got this really fun thing where we also managed the TVs for the patients. So we had to manage that the TV was working correctly. So if you've ever been doing it and had someone call you screaming, who's a patient in the hospital? That is good stuff. So very, very like, as I move through it, for me it's been a lot of growing into cybersecurity. My, uh, first company, I started way back when, back in the 2000s, which is crazy, um, was about data security and how can we encrypt data and share it and you know, back to the days of PGP keyrings, how can we, how can we better store and kind of share these keys? Uh, and where I got here, how I got here, like a lot of us, I was working at an msp. I was working for MSPS M kind of coming up, building services and I ran into something that I think we'll talk about today and I think all of you have run into, which is we got sick of waking up at 3am and replacing raids and we just didn't want to replace drives anymore. We wanted a better way, a better way to do something. And that really kind of just kept leading me forward. And as my wife knows, I'm sure a lot of us are like this. We just come up with ideas and we're like, well what about that thing? I'm going to do that thing instead. And really Kind of continuously. To continuously driving that forward. And the amount that, you know, I've been in CyberSecurity for about 20 years now, and I know, uh, Zoe always hears me say this whenever I get the question of what will change this year? What's going to be different this year? I will say the same answer I would have given in, like, 2006, which is literally nothing has changed. Almost new pieces of technology on my
Speaker B: last webinar, it was like two weeks ago. Um, or maybe it wasn't a webinar. Maybe it was a. It was. It was a. Yeah, it was a closed webinar, but it wasn't a livestream like this. Um, the question of what's this Mythos thing? Came up and I was like, nothing. Uh, my answer was basically like, I read the model card. Nothing to see there.
Speaker A: Like,
Speaker B: it's the same thing as last year's model, but with 100% more marketing served on the side.
Speaker A: It costs more money to train, so they had that going for them. Uh, but yeah.
Speaker C: Did you know that if you spend a couple of years asking a model how it feels and what its motivations are, that eventually it's going to start coming up with answers to tell you how it feels and what its motivations are? It's wild.
Speaker B: Yeah. Uh, or tell you about. It's like, uh, what is it the, uh, uh, the. The fantasy creature thing that was going on with. With, uh, OpenAI's models this year? They're like, everything relating everything back to fantasy.
Speaker A: Yep.
Speaker B: Uh, yep.
Speaker A: So goblins are very important. You never know.
Speaker B: Yeah.
Speaker A: But.
Speaker B: But I'm gonna say, um, while. While the watersheds. Maybe the watershed moment idea might be a little, um, overhyped. Right. Because the reality is the floor has been slowly raising since, uh. I think 2015 was the year of AI. Do you remember when 2015 was the year of AI was?
Speaker C: I was starting to see the. The like, to the consumer level, seeing the. I. I remember seeing the like, face. Face swap, static image face swap apps hitting first for. For folks actually starting to see it come through.
Speaker B: I just remember the trade floors where everyone, like, I washed their products and sure, like. And everyone said AI is like going to change everything. And turned out nothing changed. Um, well, it did. It did. It's been a slow change. Like all technologies that. That change slowly. Unless you're hiding under a rock and then it feels fast. Right, right.
Speaker A: Um, yeah. And you know, the last four or five months of models has definitely changed, but to your point, they haven't changed so massively that we need to like, sit down and start screaming in a hole and worrying about like, well, yeah,
Speaker B: ultimately these, these transformer models that are the center of a lot of discussion. Um, you know, it's basically the same technology as 15 years ago. And the big difference has been, um, microchips have gotten smaller with more energy efficiency, and banks have been more willing to open their wallets to build $100 billion data centers to do nothing but speculative AI training and inference. And so we built this card castle of AI.
Speaker C: Um, just one more data center, bro. I swear. Just one more data center.
Speaker B: Yeah, one more data center. And there's going to be enough for, and it's going to finally work. Has been the mantra for like over a decade now. And, uh, it does seem like maybe there are some legs to some of that. Right? But the reality is you can't just keep turning the dial up forever on that. So, uh, we're in this space now where I think, um, like the 3 nanometer chips and the 2 nanometer chips we have are the chips we have. Right. And they're probably not going to get a lot more efficient a lot faster. Right. And there's a limit to how many fabs and there's a limit to how much energy there is. Um, and so now, you know, we have, uh, we have an amount of AI in the world and what are we going to do with it? Right. And I think that this focus on these, these models that are $20 million for like an instance or $10 million for $5 million for an inference instance. Right. It's like, that's definitely not sustainable. I, you know, even when you split one of those clusters a thousand ways, right. It's still too expensive.
Speaker A: Yeah.
Speaker B: Which is, which is basically how they split them. Like you can have a thousand users on them at a time or something like that, um, and your time sharing with everyone else.
Speaker A: And so I, I, I think we all will learn a lot the moment that one of these companies goes public and we will be able to actually learn, like, what is happening inside of some of these organizations if they can go public. Because yes, I, I will be very interested to see that for Anthropic, for example, of like, how much are you subsidizing? What is the real million per token cost? Like, what is happening inside of this business?
Speaker B: But they don't, it's not even their money. They're, they're like, it's not even investors money. The hyperscalers are like fronting them.
Speaker A: And yeah, there was just an announcement yesterday, right. 50, uh, billion on 900 billion. Uh, and it's like. Yep. All right, well, yeah, keep going. Sorry, Zoe.
Speaker B: Oh, I was just going to hear your story, Zoe.
Speaker C: Oh, yeah. Story. Uh, hi, I'm Zoe Lindsay. Uh, I also am at Blue Mirror. I, uh, had a slightly more circuitous run out in. I, uh, got my first computer out of a neighbor's trash at 8. I love that. It was a Commodore 64.
Speaker B: That's a great story.
Speaker C: Oh, yeah. Uh, so my best friend and I, we would take, uh, our weekends with our bikes. We ride around the neighborhoods. We would garbage pick whatever tech we could find, uh, most of it didn't work. And set up these elaborate looking command centers. Uh, but you know, the ones that did work, we would mess around with. No, wait, wait, wait, wait.
Speaker B: Hold on. The Commodore 64 does not sound bike friendly to me.
Speaker C: No, not at all.
Speaker B: Is this like bungee cords or like backpacks or what?
Speaker C: I narrowly survived a generation where we would ride around on each other's handlebars. So it was just a matter of, you know, having, having a little bit of balance and a little lack of fear, I guess. But yeah, it was, uh, it was very much the, uh, the era of, uh, you know, I was, I was convinced I was gonna teach myself basic out of the, uh, in Basic book that I found at the library. Uh, my friend and I were completely obsessed. Uh, we went way down the rabbit hole. The library was also the first place that I had Internet access. Uh, then, uh, I think 11 hackers came out and I was like, I want to be part of this world. Uh, even though at the time I didn't, I didn't quite, uh, I thought there was going to be a lot more like Fashion and Wipeout, uh, a year before Wipeout got released, uh, than there actually ended up being. Then I finished, uh, high school, just barely. Uh, didn't make it into the schools, just went directly into working. Worked about, uh, 80 or 90 different kinds of jobs, including singing telegram, locksmith, uh, apprentice, and charity casino part or, uh, charity casino party dealer. Uh, and the whole time was trying to find my way into getting to futz around with computers full time more often. Uh, I found out about a local startup that was getting some funding that needed a office manager. I had zero qualifications for it, but I was able to pull pieces of my resume and, uh, polish them up and said, you better hire me for this because I'm going to apply for literally every job you post until you do. And they did. So I got in the Door at Duo and started writing down every term that I heard and didn't recognize. Uh, saw the opportunity to get to do this full time. Uh, built a web server so I could start learning some of this stuff more hands on since I was getting to focus on it for my day job and eventually found that there was a spot for somebody that was happy to go down the technical rabbit holes and do translating for the folks that didn't care to. Uh, so I did that for a good long while. Realized that after acquisition I was not so much a fan of the giant company thing as the startup vibe.
Speaker B: I was about to ask if you bailed before or after.
Speaker C: Uh, I left. Well, after. I probably should have left a couple of years earlier. M. But uh, fortunately for me there was a good opportunity in another local startup. I got connected with Matt. I, um, was always drawn towards solutions that wanted to help out. Ah, the vast majority of folks and not just the, uh, folks that are in the Fortune 100 or 500. Uh, and that's sort of how I ended up here. Um, I forgot what I was starting at.
Speaker B: I love that space of me, um, personally I love that space of translating between the people that have that technical acumen and the ones that don't. That's my favorite place to be. Um, yeah, I mean I found it's just as important in security as it ever was, um, anywhere else in technology. Right. And it's actually free right now. Yeah, because you got the technologists that don't speak security either.
Speaker C: It's true. We're abstracting all of that away now. Uh, well, I mean, it's like, you know, um, security is the collision of people and technology and then process, hopefully to govern it. All right, so you can't just solve those challenges with technology solutions. You need to have some people solutions in there too. All right, well my marketing coming out,
Speaker B: I forgot what we were talking about before that, but I love that story. I can really relate. Um, uh, thanks for your work at Duo. Um, it was great in the before times. Um, uh, we were just talking about splitting up all these giant models. Right. And how, you know, you can argue on sustainability and is there going to be enough compute and is it going to be enough smarter? But I think the thing that we wanted to really talk about today and the thing that's often overlooked is the front of the frontier is important, but it's not actually where business gets done. So. And um, that that interface of like the frontier of open weight models, which are models that are not necessarily open source And I think people misunderstand this a little bit. Right. So in, in, in AI models, there's like, there's the training set, which is oftentimes proprietary licensed materials that you actually can't release to the world. Right. Then there, and then there's the methodology you use to do your training. Right. And the architecture of your model and that stuff is, is very rarely open sourced for legal reasons really. Um, but then there's like, what do you produce with that? Well, you've got this AI model that can infer things. It can, inference, um, for the user. Right. It can transform that inputs from the user or the API into outputs that seem smart. Um, that thing isn't always locked up in an anthropic or Microsoft data center. Right. It's a thing you can download sometimes, and good luck trying to download cloud opus, um, even if it's legal to do so. Um, but there are models that fit on computers that you can buy. Um, and that's like, actually where the real important edge is, because those models that you can actually fit on a computer are cheap enough to run on a computer that we can all afford.
Speaker C: Uh, well, in theory, at least six or seven months ago that we could all afford now, hopefully, uh, you already got those GPUs and that RAM or your Mac mini before they sold out.
Speaker B: Yeah, you know, I, I, I try not to, like, pay too much attention to, like, this week's, like, changes in macroeconomics that are like, not going to be the same in another six months. Like, I think that, um, the memory factors are, uh, memory factories are building more fabs and, you know, there's more capacities coming online and those, some of these crunches are going to go away, right?
Speaker C: Oh, absolutely.
Speaker B: The, the, the reality of just like the silicone crunch in general is real though, right?
Speaker C: Oh, yeah. Um,
Speaker B: I, um, in any case, I didn't mean to hijack that thought. The, uh, the, uh, these models do fit on things that we could theoretically buy. They cost less than a car, let's put it that way.
Speaker C: Very, very fair.
Speaker B: So whether you're Talking about like $2,000 computer or a $10,000 computer, it's something a business can own.
Speaker C: Yes, absolutely.
Speaker B: And whether you're, whether you've stuck that like, under your desk, which might not be the most secure place to stick something, or you're in a hyperscaler, but you're using a piece of their infrastructure that's actually economically sustainable, that you can actually, um, maybe you can customize it. Um, I'm talking to this um, customer of mine yesterday and they're talking about how the frontier models are just like um, they're never going to work for what they're trying to do but they see that what they want to do could be done with AI. And the only way that they see that their only route forward is to take one of these open weight models and train it on their set. And it's the only thing that's going to work for them. And guess what? They're going to end up being the only one in the world or one of the few people in the world that know how to do this thing with an intelligent data set that fits on a computer that they can actually afford. Right? And they're going to be able to defend this specialized expert system, right. And sell it to people. And that's like a really neat, neat space, right where the future of um, the future of technology is still something that we can create and we don't need to be a trillion dollar company to own something, to own a piece of intelligence, to own a piece of technology, to own a concept, to own a, like a capability in the world. Right? You can actually sit in your office and you can still create something and own something and um, be a producer, right and not be owned by, by Google or Anthropic or Microsoft. Right?
Speaker C: Yeah, for sure. Matt. I'm sorry because this is going to be basically verbatim what I've talked with you about already this week. But I think like uh, you mentioned uh, Gemma for uh a little bit ago, Matthew, but it's a really interesting time because I see like two, two diverging paths that we can be going down. And one is very token driven. There's a handful, we go to the NASDAQ 4 or whatever and then we're paying per token and paying through the nose, uh, in a closed market, basically whatever these few uh, providers want us to. Um, or we go down a path where these models get more efficient, uh, more manufacturers lean into building hardware that is going to better support them. That seems to be the route that Apple is going. They don't see, seem to be as interested in going down the services route. I think they're going to try and win the hardware race because they've already got a foothold there. Um, and it's a very different world. It's not token driven. You are able to do a lot more with using an open weight model. You don't have to worry about where your data is being stored if you retain control of your data. Um, I think that, uh, it's a very interesting time because on one side there's a handful of frontier model providers that are very heavily leveraged and motivated to make sure that theirs is the winning solution. On the other side is every other business and corporation that is only going to tolerate so much before they start to try and use their capex instead of their OPEX to keep this going.
Speaker A: Yeah, I would even go one step further in that the large foundational providers, and I mean this is a nice way, please don't strike me down, uh, are actively damaging the economy at this point. And that isn't to say that they're actively trying to do it. Uh, I mean, they are in some cases. I'm sure people are making a lot of shorting stocks right now. Um, but in, in the reality, when it comes down to it, and we see this a lot, you know, Blue mirror, but I see this a lot in my personal life. This machine right here has AMD gpus in it because I can get AMD GPUs for cheaper. Technically, you can get that running inference models on it. I can do it. I can have 48 gigs of vram right here for $1200, which you can't do with an Nvidia chip the same way. So it's a.
Speaker C: There you are.
Speaker A: It's a really interesting dynamic where the cloud providers aren't inherently making it super easy. I mean, they're making it easier. You can do spot instances of GPUs and TPUs, the new TPUs coming out of Google are a little more efficient, a little bit more efficient. It's like we're seeing a little bit better stuff out of this. But when it comes to what we've seen, we tend to, or at least what I've seen in my effort, uh, the foundational models are for like that last mile of intelligence and where I think people tend to struggle in the conversion from. I used Opus and then I'm going to go use Gemma, uh, as they expect similar output and it's like, no, it's a different thing. You have to use this differently. How you're going to use it as differently, how it's going to engage with you as different. Your scope is going to be different when it comes to this. And that's okay. It just requires that mindset and change. And I'm sure that a lot of people have gone out and set out like openclaw and please don't expose it to the Internet. If you've done that, uh, and set up openclaw attached to Claude. You'll burn a hundred dollars in two days very, very, very quickly and very easily. That flip is really interesting though, because then you have to have that changeover. And M, I'm sure people are seeing this in their kind of how do I make my MSP more agentic friendly? You put in a model that's less, doesn't have as much fuel. At the very least, it hasn't scanned the entire Internet and been trained on the Internet. At the very least, you have to be a lot more verbose in your prompting. You have to be a lot more verbose in your context. You have to give it different information. That changeover is an interesting kind of part because to your point, Matthew, is that those bigger models are going to have so much more ability to infer, the smaller models just won't have as much ability to infer without you providing more to it. That flip is a lot of the interesting part. We've done some testing, for example, using gemaphore. Um, MSPS could do this as well. Using Gemaphore in the middle of our ingestion pipeline, just using it to say, what does this data look like flowing through? This is a thinking model. We want you to evaluate it, but we end up providing it with like three to four pages worth of prompt. So it knows what is it looking at, what does it need to do, how does it need to respond? You can get there and it can run on a, um, very cheap MacBook Pro at that point.
Speaker B: Well, I think what people, um, what I see people get so impressed by, right, these days is when they produce half a thought, right? Like it's almost like a water cooler, like half a sentence, right? And the AI just flies and does a thing and comes back half an hour and it creates this result. And people say, isn't that amazing, right? It is amazing in its own way. It's not as amazing as someone straight out of school who's only like maybe 20 years old. It's actually not even that amazing, but it's amazing that a computer can do that at all. Take that thought and riff on it, right? Um, but if that's how you train yourself to use AI, right, you're gonna need that. Like, I don't even know how big these. They don't even tell you how big these models are. Like, the model providers don't even want to tell you how much compute they are burning to produce these answers. In some ways you might say, well, that's proprietary information. But the reality is it's an Embarrassment is what it is for them, right? Because it's really like exposing like how much money they're burning to produce these, these answers. Um, but you can actually take, um, those if you are a little bit more structured with your prompts, right? You can actually take that and go down a level of intelligence. And so, um, you know, what I'm working with internally is um, a framework where we've got four, four layers of intelligence. And M, you know, from the perspective of creating sustainable infrastructure that can scale, it's really important for us to use to maximize those lower levels of intelligence as much as possible. In the same way that you're not going to pay someone $200,000 a year to unjam a printer in your IT service provider, right? It also doesn't make sense to use your Opus model, right, to do something that's like reformat this document, right? Like, or whatever. Um, might be like really probably a pretty straightforward task if you just gave it a little bit more instructions. Or in my case, what needs to happen is I need to not give it a half sentence, right? What I need to do is really the same way a programmer normally does, right? Map this business logic. What actually needs to happen here? How can I simplify those instructions? How can I make this procedural, right? How can I minimize the dependence on giant cluster that knows everything in the world, right? Um, but still produce something that's actually really different than procedural programming, right. That we've come to sort of expect over the last couple of decades, right? You can actually get really unique insights from. From a relatively small model. So let me. Let's just talk like relative size here. So you've got, um, what sounds like maybe like a hefty gaming rig equivalent of hardware sitting next to you. Uh, maybe like a very hefty gaming rig.
Speaker A: Hefty, yeah, yeah. M. But.
Speaker B: But there are literally YouTubers playing games with rigs like that, right?
Speaker A: 100%.
Speaker B: And, um, because they're on like the cutting edge of gaming, right? And, uh, what is that, like 10 billion parameters? What was. What fits in that space?
Speaker A: Uh, yeah, you can fit probably just at the edge of 10 billion. The hard part of that for a lot of people is the patience. Like when you get addicted to the big foundational models that are even smaller. Like, let's say you're used to using haiku. Like, realistically, you're used to that snappiness and the loss and snappiness, I think is one of the hard parts for people when they move to this. Like, don't get me wrong, I think to your point, Matthew, you can go even smaller and have that be the first layer and have initial snappy response and then have that layer up through your intelligence. That tends to be what I've seen the best way to approach this. But you can fit, you know, I could probably get with the amount of memory I have in there in the 15, 20 billion. But you're talking about a pretty slow inference, you're talking about a pretty slow low to get to that point. In my experience when going through this, you're way better off having that 10 billion parameter model be the orchestrator and have it hand out the work to the things that you have the additional models available for, just because then you can really drive into what's the most valuable. And you know what I've seen work really well from our perspective is use the foundational models and really slim ways to help you build up those prompts, build up those contexts, understand how you're going to talk to the thing, but don't use it for actually running the thing, just use it to help you build where you're going.
Speaker B: Go ahead. You use a word like haiku, which I understand is about um, $100,000 investment in hardware. If you wanted to run that, it'll be around there, um, somewhere in that range, right? Yeah, you've got that Gemma 4, that's. What is that?
Speaker C: Like probably you could probably 27 billion for the mixture experts and like 31 or 32 for the dense model.
Speaker B: So you could do that. What is that, 5 or $10,000 of hardware? Is that in that range?
Speaker A: At the most? Yeah.
Speaker B: Yeah. So, but that's a big difference actually. Um, but haiku is not actually what most people are experiencing when they access some of these frontier models, right. Um, they're experiencing something that's orders of magnitude, um, more complex than that. Right?
Speaker C: Yeah.
Speaker B: Um, so haiku is like we have these, if we want to use like anthropic terminology, because you brought that up, right? You've got Opus, right. And something that's um, you know, somewhere between like twice and five times as efficient. You've got Sonnet sort of in the middle, right. And then you've got Haiku, which is uh, what is that? Another like half an order of magnitude smaller there.
Speaker A: Yeah.
Speaker B: Um, and then you've got uh, I guess your example of Gemaphore. Right. But there's a whole, there's a whole, um, I've been, I've been using. You uh, know Gemaphore is relatively new. What I've been experimenting with is like OpenAI has got an open weight model in that size range, right?
Speaker A: Yeah.
Speaker B: Um, which, and you don't have to buy a $5,000 computer, right. You can go to one of the model gardens, you can go out to Foundry or Vertex or Bedrock and you can go access these models. They're available on top. They're very relatively low cost because they're open weights models. Um, in the same way that you've got open source that people can audit and poke at and those open weight models have been poked and they can be poked and they will continue to be. And I'm imagining this Future where these 2026 models are not going to go away ever because they're going to become more and more well understood foundations that um, people sort of know what they are. Right. Um, in the same way that we have source code that's been sticking around since the 80s, right.
Speaker A: Or older.
Speaker B: Right?
Speaker A: Yeah, very much so, yeah. Um, I mean this, this kind of goes to your point. Like E4B is their, their mixture of the small mixture of experts. That's only 8 billion parameters. Like uh, you can fit that into a thousand dollar computer in the grand scheme at all. Like you buy an AMD GPU and some cheap RAM and the cheapest CPU you can find out there, it'll run just fine and it'll get you what you need in most cases, you know, and you're going to have up to at least a quarter of a million in the context window associated with that. You just have that interesting problem that you get to manage of. Um, is it going to go insane when it has too large of a context window? What is this going to look like? How's it going to fit together? But to your point, if you can spend $5,000 and not $100,000 and you can get value out of it and you can build additional intelligence in. There's a lot that you can put together. Just leveraging technology to ensure that you're putting the right pieces together. Like the, the way that I like to look at this when we build stuff, and I build stuff either over here, over here, personal computer, private computer, uh, is what is the framework that I can have built around the problem that I'm trying to solve. How can I make sure that intelligence has handoffs between models? And if there's ever a handoff to the foundational model, it better be a damn good reason that it's going to cost m bucks to run this thing and really ensuring that you have the right kind of flow through that intelligence. It's just hard to do with some of the ones in the market. Like you can't, you can kind of do that like open claw right now, but not really. Um, it's inherently difficult without you knowing kind of exactly. You just said, Matthew, you have to know the business problem. You have to know you want to solve. You have to know what it's going to look like because it's not going to do the magical inference for you now.
Speaker C: Okay.
Speaker B: Yeah. Well, I just want to, I want to ask you about that word you use, foundational, right. People, people say frontier. I actually like, I appreciate the word foundational because I'm imagining a future where instead of programming code, right, we're programming open weight models. We're pulling a Gemma 4 off the shelf. Right. And then we're distilling on top of it from like a crazy complicated model we couldn't possibly run on our own. And the only thing we're using that Opus or Mythos for is to distill a small model to do something very specific and it actually can outperform or equally perform what those massive models can do at ah, that very specific task. Right. And in that way, like, and that, by the way, that's what the frontier model providers are doing with their more efficient models. They're saying, oh, users like to do X, Y and Z. We're going to take our frontier model, train it to do X, Y and Z called Sonic.
Speaker A: Right.
Speaker B: And mhm. You know, so that's, that is sort of the future and when those uh, we don't, we don't have great tool sets for that. Right. It's not something that's easy for um, for someone to do on their own right now. But it's, it's not that far away either.
Speaker A: No.
Speaker B: So you know, I think that, that, that that future is really close. And it reminds me of the way we, you know, you have a service provider that's got their level one, level two, level three staff and like actually like that, that, that model's been around a long time. But a much more efficient way to treat your lower level tap, lower level staff is not as like a level one or a level two, but to say, well we have this person that's trained in fortigate and they're like the expert in fortigate firewalls and they don't know anything about Microsoft Entra and they don't need to. Right. And uh, like do you, do you, how far away do you see us from that future? Is that what's sitting under your desk right now.
Speaker A: I think it's closer than ever in a few different areas for sure. Like, especially when it comes to detection and we think about real time streaming analytics. What we would have done three or four years ago, three or four years ago is like, all right, well, we're going to put Redis in here. We're going to catch a bunch of data, we're going to look at it every once in a while. We're going to write a bunch of code to evaluate this data. We're going to hope that that code lands in the right way. We're going to hope it doesn't break somewhere. However, if you take a step back and you start to generate synthetic data using a model, if you have a set of data and you start to distill that into a smaller model and start to give it more intelligence to your point, so it can actually give you the right answer. You don't need to actually put layers of technology into that like you used to. All you need is something that understands how to solve that problem there. So I, I do think you're going to see more of this. I think the hardest part though, and we've been talking about this a lot internally, is how do you reduce still the fear of hallucination and incorrectness in those situations? How do you make sure that you're training and grounding it out as much as you possibly can? You know, one of the things that we've gotten more into on our side, which, which is more on the, the frontier foundational side of testing, not necessarily using it, of using those frontier foundational models to validate the output of a small model, to act as the judge, essentially. And, uh, not all the time, but like just validating. Like, was this correct? Was what we trained this to do and what we distilled this down to do? Did it land in the right direction based off the Genius of Opus 4.7 or chat or GPT 5.5? Because if it did, then to your exact point, Matthew, it has already got to that level of intelligence that Opus is or that Codex is.
Speaker B: You know, the, uh. I think that that word hallucination is going to leave our lexicon in another few years. Because it's really. When you're, when you're talking about this technology, um, you're talking about the difference between being procedural and precise and being creative and expressive, right? And if you ask, if you take a model that's tuned to create art and you ask it to follow like a series of procedural tasks, it's gonna,
Speaker A: it's gonna invent like procedures sample for you here. Perfect example. I have a personal project working over here I won't go too deep into. It's not that interesting. But I was working on a project that I wanted to develop. Basically, um, CAD drawings I wanted to develop. Three. I have 3D printed stuff behind me. I wanted it to develop things for me so I didn't have to do it. Like gear systems and real systems and robots and designs. And what I quickly learned is a few different things. One, even the frontier foundational models are not very good at it. It just being generate me a gear system in CAD that looks like this who's not very good at it. Most of the small models aren't very good at it. You can look at the Quinn models are okay at it. But realistically what you have to do is build an entire library of, uh, this is how things are built, and then train a small model to do it. Because you're better off training a small model to do it, because otherwise you're just training the large model that also is not very good at doing it to do it as well. I've been working on this project for about six months now. And the biggest thing I've learned is that we are better off doing exactly that. Taking a small model, giving it what it needs to actually be able to answer the right question, either through distilling it, fine tuning it, however you're going about it and getting that answer out, than you are ever trying to just prompt the living hell out of a frontier model to get it to what you want. It just works so much better in those very specific use cases.
Speaker C: Well, I think that's kind of the direction that things have started moving in for some time anyway. If you go and look, look, a year and a half even ago was, you know, excuse me, I'm not a vibe coder. I'm a prompt engineer. And people are trying to figure out what the magic set of words in the right order is to unlock the answer that they want. And then we've seen a movement from that towards context engineering and curating what you're going to be pulling from. And the next step beyond that is actually using that to train a model. But you can curate and build that context and take it with you or start to build it, or use it where necessary with a foundational model and merge that with what you're using in your local environment. We're seeing more and more kind of this bridging. Um, I think we're going to continue to see it move in that direction. But, you know, if I ask a question, I want to know what the source is. I don't want to just ask a question of someone on the Internet has answered this question before. Um, and we're already past the point where adding more context to these foundational models that are just sucking up the Internet is toughen their own generated farts. Like, we need to be able to have, uh, some quality control on the information that's going in if we want to continue to get quality information coming out. So it is going to be a, uh, melding of folks continuing to cultivate, uh, what is going to be relevant for the purpose that they're using it for, whether that is a library of CAD drawings or a library of detections or a library of writing samples. But you want to, um, you're going to have a responsibility as like both trainer and curator for IT to get the best result out of that at the lowest input cost, whether that cost is tokens or money or energy.
Speaker B: I'm curious where you guys think that, um, people can, you know, as an IT professional, right, not as someone creating this technology, but as a user of the technology. Your, your job is to take the best products and your, your people's talent, right, and service your customers the best way you can. I am, um, you know, I, I go, I have a lot of partner relationships that I try to keep fresh. And they're, they're, they're in like half of them are on one side of this industry wave right now, right? Where they, it might be a disaster, but they're using, um, uh, react agents to uh, you know, fulfill job functions right now. Right? So they're, they're going through these inference cycles and they're asking agents to actually do stuff, right? And they're using that, that tool to do stuff. And then there's organizations that aren't doing that yet. I think the organizations that aren't doing it yet are probably actually not as they might be at more advantage than they think. Because I see a lot of disaster happening out there right now and people not realizing the dead ends that are ahead of them. The difference between perceived productivity and actual productivity, right? Like, oh, I'm, I'm creating this beautiful castle and not realizing that the castle is never going to be finished, right? It's like the AI is just really good at asking you questions and producing slop, right? It's never actually going to get there. Um, I think that people are confronting that right now, right. And trying to Understand, like, where is it useful? How useful can it be? What can I actually expect? But I, I'm seeing people put, um, and openclaw is an example, right. But I, I, I know partners that are, they're, they're running MSP operations with an OpenClawbot.
Speaker A: Oh, yeah, right.
Speaker B: I, I, I, I've seen that, um, I've seen people creating their own stuff. Right. Um, how do you see this is going to interact with, you know, you have the various people in your msp, not all of which are capable of dealing with the complexity right now, because the AI harnesses themselves are sort of early days still. Right. Um, what types of distribution, uh, do you see in that intelligence? Like, do you see people using, I think, like alert monitoring? Um, you know, like, you know, in the old days, you're, something failed, something bad happened. Right? And you were like, how do we stop this from ever happening again? Oh, we can detect this before it, like, like, how do we get ahead? You put the monitor, you put the monitor in place. Right. Are you seeing, like, are you seeing that out in the field yet? Do you think that that's like OPUS level model, or is that like, actually you can throw Gemma in there and here's the list of situations to look for for.
Speaker A: So I, I don't think you need OPUS to do that at this point. Like, we, I think that it's still a little hard to do it with Gemma. I think it's possible with the right people. I will, I very much agree with you and that, you know, we see a lot of people also, like, very split, like people who are going heavy agent people are doing no agent. I talked to an MSP the other week who is saying that they need to update their entire MSA because they previously had an MSA that said they never use an LLM. They were like, oh, crap, we need to take that out. And I was like, yeah, you probably should, ah, at one point. Um, but in the grand scheme of what we're seeing, like, I can tell you that over the last two weeks we've seen a significant increase in codecs accessing SSH keys, for example, which tells me that we're seeing more application of people using chat GPT doesn't necessarily mean that they're actually using it for any real success across that time period. Because you can always burn money with, with a model, people are happy to do that. So really where we're seeing the most integration from the standpoint of on the ground is really saying, all right, we have people that are Microsoft shops, people that have bought into the Copilot kind of view. Now I have just as many people have shut off copilot as people that are using Copilot. So take it with a grain of salt to an extent. Great. Like it gets you a little further. And we're starting to see more people say well what if you had an MCP server? What if I hook this up to my AI uh and I no longer logged into your thing and it was just all headless. So we're starting to see more of that. But I very much would agree. Like you can run Hermes, you can run Open Claw, you can run these different agent frameworks. The issue that we tend or uh, that I tend to run into is that they are thirsty. Like they will inference the living shit out of everything they can possibly do because they were built to use the frontier models to kind of drive them forward. And what we haven't seen yet, at least I haven't seen yet, is we do see people doing some with RPAs like in Tines. We're seeing some more built out. All right, we'll deploy this and we'll have a smaller model like Haiku or Gemini Flash handle this. We are not seeing as much local yet. Probably the place where I'm seeing that the most annoyingly, probably annoyingly based off this conversation is in bigger companies like Fortune 500 companies that are like you know what, we have enough money to put capex in into this. We're going to buy hardware to let m our engineers run local models and we're going to see what we can do with that in the msps. I think one of the hardest parts that we're all that uh, someone will solve this problem and make a lot of money is where is the framework that sits within the MSP and acts as that RPA with AI, smaller AI models built into it to help with the orchestration of what is happening across that kind of MSP as a whole. Like the inherent issue with open closets for an individual like the, when you open it up for the company it gets a little, little, little, little scary.
Speaker B: So I, I use um, you know, I use a little bit of Claude code, a little bit of um, Google's uh, got an anti gravity framework. Um, I actually um. Anti gravity has been a little bit of a disaster, but it's actually um, beaten the little living pants out of uh, um, uh, Claude code in some very specific ways that I think if we fast forward a year, right, this idea of sub agent automation, right. I've Got my core thread right. I'm trying to accomplish this long range planning um, goal, right. I'm not going to like be burning tokens for like 12 hours straight, but I might, I might hand off a series of 80 sub agent tasks. Right. Um, and that is, it does seem like that's going to be a winning model. I don't know if it's going to be the winning model. But then I heard earlier in your conversation you talking about how you use the lower intelligence model to do some of the orchestration and I've been playing with both of those in um, in what I'm doing and I'm not sure, I'm not sure why one's going to win over the other. I'm not, it's like I'm very confused right now as a, as an implementer, right, Because I, I also run some engineering, um, what's the right, what's the right way? Like which, which way is the way that's uh, more efficient? Which way is the way that's going to win in the end? And I don't, I don't know, I don't know. I'm, I'm doing both right now. I'm sticking hard problems and giving them to cheap agents and like throwing them in the background and not worrying about them burning tokens, right. And I'm taking complicated things and giving them to more capable models and letting them crunch on it. That gets very expensive. But sometimes it's the only way you get to your goal. And I feel dumb right now that I don't know how to use these tools.
Speaker A: I think Zoe and I talked about this a few weeks ago and I see this in my work and I'm sure other people do as well, where if you can bias the model to write code and improve, improve code to take action and do its own work, you can get that token usage down a lot. Which is to say if you can get it to do like Hermes, if uh, anyone has never looked at Hermes, it's worth looking at from a self learning perspective. It will try to improve itself in the background and those are the kind of areas where I suspect we'll see more bleeding edge effort come out of this. Uh, you asked for a thing, it's been three weeks of you asking for this thing. Here are the areas where you correct, corrected it. Here are the areas where we're going to improve this. Here's how we've changed the code looking forward. But it's tough because to distill down the complex problem into a Smaller model is always going to require that context to be brought in. So the way that we engineer now here, for example, is that most of our engineering effort is based around specs and content. And what, uh, are you intending to build? What does the business need? How is this going to work? What is the end goal? And giving that to a variety of different models to actually suss out and validate. Because, yeah, it varies. It very much varies at this point,
Speaker B: which sounds a lot like how engineering has always been done. I think it's true if you want to be successful. Yes.
Speaker A: I think it's how all of us in engineering would have loved it to be done the right way. And now it's kind of like. But now you have a helper.
Speaker B: Well, guys, we're running long on time here. We really need to hear what Blue Mirror is doing here. How are you guys? How are you guys winning right now? How are you helping our service provider community? Who wants to take a crack at this?
Speaker A: Uh, I'm happy to kick it off. Uh, although I like to have Zoe pitch when I'm here as well, because it's just weird when I don't pitch because I'm the CEO, cto, co founder. Um, so what do we do? What are we doing to help the community here at bluemirro? What we do is we provide, uh, integrated Security Operations platform that's built around our siem. Uh, and we like to talk about that. And I think it's kind of very relevant for today and that one of the most important parts of having a model that does something for you is having the right data that's brought to it and making sure that you are bringing the right data to it at the right time. Um, I'm not going to go too deep into this, but we're working on our own kind of view into AI and looking at kind of exactly what this conversation is, which is how can we make it as easy as possible for people using a SIM? I think we've all seen for the last 20 years, people have been talking about SIM and saying, don't worry, the next gen of SIM is here. I can find you a News article from 2011 that says that. I can find you a news article from last year that says that. Here's a surprise. It, uh, only got worse over that time, for the most part.
Speaker B: Oh, yeah.
Speaker A: Turns out data warehousing and not actually structuring data was a bad idea. Who would have thought? Uh, and so what's most important to us and how we like to work with Our MSPS is that we want to make security as easy as possible for you across the entirety of your stack. We want to see all that data. We want it from the endpoint, we want it from your identity, want it from your servers, we want it everywhere. Because when it comes down to it, the attackers just as much as your users, just as much as the AI needs that context to make sure that we can make sure that we understand what's happening there. You can understand what's happening there. You know, we can talk about acronyms like edr, itdr, sim, xdr, blah blah, blah, who really cares? What's most important is that you're getting the right information at the right time. And we're here to help you and your organization meet compliance and stop threats. M. You want to do anything bad, but there's my quick and dirty.
Speaker C: No, I mean I think that that's a good way of breaking it down. You know, it's, it's you, ah, look at the traditional approach to doing security as a human and what we're talking about with AI and being able to make good decisions, you're looking at the same progression, right? Like first you need to have that input, you need to have that visibility and context. Then you need to be able to do something about it. Then you need to be able to, uh, graduate or adjust what that response is based on the information that's going in. And then you need to be able to do that faster so that you can get earlier in the process. That's the same sort of thing that we're looking at when we're talking about AI. You're not going to be able to take good action if you don't have good context going in on the front end. Um, so security teams aren't going to be able to take good action without having a complete picture of what's going on in their environment, uh, and having that picture over a period of time. Because you don't actually know what normal, normal looks like if you don't have, or if you don't know what abnormal looks like, if you don't have a good clear picture of what normal looks like over time. So, uh, it's, that's very much what we've been focused on doing is, you know, first you have to solve the problem of getting that visibility and getting it in a solution that works for companies of every size and not just giant enterprises that have a full dedicated SoC. Then once you have that visibility and you see all the scary shit that's happening, you need to be able to do something about it. Uh, and then you need to be able to do that at different levels. You need to be able to have that granular level of response and to do it faster. That's ah, sort of the progression that we've taken. So um, give you a little sneak peek of maybe where we're going.
Speaker B: I can tell you that the average organization I bring in for an assessment, um, where I actually get to see, you know, because we have software, right? And people are using our software to do assessments all over the place. But I don't know, I don't always get to be geeky in security and understand what's going on in that customer's environment unless they bring us in for that portion. Um, but people really weak here, people really weak on anomaly detection in general. And there's an over reliance on Endpoint security and MDR products. And of course every IT product you can think of has that alert button, that alert feature you can turn on and then great, you turn on all the, those features and then you just like inundated with a quarter million emails and like now what? Well we send the critical ones to our service desk and it's like well how many critical ones can you really send to your service desk and still like, still effectively respond to them? Like how do you not lose the signal and the noise? Um, and the answer has always been well you need to do some kind of anomaly detection that involves correlation. You need to have humans um, in the loop or you need to have now intelligence in the loop, right, to really like look beyond what's procedurally possible. But it's hard, it's hard to collect the insights all in one place and you really need um, you need something purpose built for that. And so um, in almost every organization and we sometimes use um, acronyms like sim, right? That's like what you put on a product. Um, but from a compliance perspective that's not usually what it says from a compliance perspective is like basically is someone paying attention to what's normal and what's not normal. And when something's not normal they're noticing it and doing something about it, right? And it's very hard. It's very hard. And MSPs can't do this alone. And an endpoint security product does really doesn't get them much closer.
Speaker A: So yeah, it's such a, uh, especially with the increase in business email compromise, the need to look at the, the behaviors that are actually happening. Like it is the I wouldn't, I won't Say the world has changed because it hasn't. Uh, it's really more. How do we get to the right place and make sure we're not burning people out? Because that's what's most. I mean, that's, that's always been the goal. I feel like in IT and it security.
Speaker B: Well, there's, there's more noise than ever. We keep, we keep creating more data.
Speaker A: We sure do.
Speaker B: Any of you guys remember that paperless office promise where we were going to stop cutting down trees? And I just feel like we have to start worrying about those wasted bits now.
Speaker A: Yeah. I mean, the main difference with us is that we're fully unlimited. So we just tell you to send us all your data because we want to see all your data. And that becomes my engineering problem, my team's engineering problem. But like when you, when it comes down to it, that 100 plus petabytes of data is valuable to all of the organizations because you have that ability to introspect the baseline, look at the activity, identify what's happening. And to your point, when you get that audit and they say, prove to us that someone looked at this six months ago or could have, you can actually prove it, which is nice as well.
Speaker B: Yeah, yeah. The proof is becoming more important. Well, thanks guys for joining. We're over time. Um, any takeaways before we, we let everyone else go as well?
Speaker A: My takeaway is.
Speaker C: Thanks for having us.
Speaker A: A local inference use. Buy an AMD gpu. It's cheaper.
Speaker C: Yeah, there you go. That's a better takeaway.
Speaker B: All right, well, uh, we'll find another excuse to do a M. Matzo and Matt show and maybe we'll bring Zoe along. Um, thanks for joining us today, guys.
Speaker A: Thank you.
Speaker B: Thanks.
Speaker A: Thank you, everyone.
Speaker C: Great to be here. Have a good one.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.