
Making Data Simple · 2026-08-05 · 44 min
Key moments - from our scoring
Substance score
71 / 100
Five dimensions, 20 points each
Rob May brings a rare dual perspective: founder, investor, and ten-year AI analyst. In this conversation, he unpacks inference - the often-overlooked second half of the AI compute market - and explains why it matters more than most realize. Unlike training, which happens once, inference runs every time a customer queries your model, creating marginal costs that traditional software companies never had to manage. May walks through real customer scenarios: companies spending $100-200K monthly on frontier models like OpenAI's API, watching costs climb 20-30% month-over-month, only to discover many of their workloads (summarization, simple classification) could run on cheaper, smaller, or open-source models. He frames optimization across four dimensions - cost, latency, reliability, and data sovereignty - and NeuroMetric's role as an intelligent orchestrator routing tasks to the right model for each job. May also addresses the strategic questions keeping investors up at night: the commoditization of frontier models (citing Kimi K3, DeepSeek), the fallacy of "whoever gets to AGI first wins," and why protectionist chip policy might actually handicap American innovation. His thesis: the real margin in AI won't be in models, but in the infrastructure layer that makes inference efficient.
Inference is running a trained model on new inputs to generate predictions or outputs - a single pass through fixed model weights. It becomes expensive at scale because unlike traditional software where marginal costs are negligible, every customer query incurs real compute costs, and companies typically build on frontier models without realizing cheaper models can handle many tasks.
By analyzing workloads and routing simpler tasks (summarization, classification, basic extraction) to smaller or open-source models instead of frontier models, companies can often reduce costs by 10-100x on those specific workloads while maintaining quality.
Yes - as models commoditize, frontier model companies will struggle to maintain pricing and margin; the real value will shift to infrastructure and orchestration layers that intelligently route work across models, not to the models themselves.
Restricting advanced chips creates incentive for competitors to innovate on cheaper hardware and alternative training methods, similar to how Africa leapfrogged the US in mobile because they had no legacy landline infrastructure to protect.
Latency (running smaller models locally for faster response), reliability (consistent outputs across requests), and data sovereignty (keeping data on-premises for privacy or regulatory reasons).
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains solid, actionable insights about inference optimization, cost management, and model routing that most operators wouldn't encounter in casual reading. However, substantial portions involve biographical storytelling, repetition of basic concepts (training vs. inference), and general market commentary that dilutes density. The core value clusters around specific practices: fine-tuning smaller models for cost savings (70-80% reductions), prompting strategies, and architectural decisions - but these are interspersed with filler.
we typically see a drop of 70 to 80% in price and the performance stays the same or gets better
when you prompt them you really want to choose the best model, which is really, really important
May offers some genuinely fresh framings - particularly the analogy of inference cost to manufacturing margins, the Hardware Lottery concept, and the argument against protectionism using the Africa mobile leapfrog example. However, much of the core narrative (models commoditizing, AGI racing myths, smaller models for specific tasks) has circulated widely in AI discourse. The contrarian takes on regulation and market winners are worthwhile but not deeply explored.
typically when you build software, you have not worried about the marginal cost of selling another piece of software because it's been negligible. That is not true with AI
I have never believed in this idea that you have to be first
Rob May is a legitimately credible operator and investor: repeat founder with a successful exit (Backupify), active venture investor managing multiple funds, currently building and scaling NeuroMetric AI, and a decade-long public thinker on AI markets. He speaks from both founder and investor vantage points with authentic operational experience. However, he is not a household name or top-tier exec at a major AI lab, limiting the score.
I sold my first startup in December 2014... it was called backupify
we're looking at some of this and we're like, okay, there's going to be a lot of things you're going to want to optimize in these systems
The episode includes concrete metrics (70-80% cost reductions, 2-4x speed increases, $100-200k/month spend patterns, 10,000 daily query thresholds) and real named examples (OpenAI, Anthropic, Mistral, Cerebras). However, many claims lack specific supporting data: the Africa mobile analogy is vivid but anecdotal; the AGI timeline arguments reference theory rather than evidence; specific use cases (text-to-SQL, sales meeting summarization) are described generally. Overall stronger than typical but with notable gaps.
a company they're spending 100 to $200,000 a month on inference and it's growing fast
we typically see a drop of 70 to 80% in price
The host asks reasonably focused questions and attempts follow-ups, but rarely pushes back or challenges May's claims substantively. When May makes bold statements (e.g., 'OpenAI will disappear' as 'the WeWork of AI'), the host moves on rather than probing. The conversation feels collaborative and respectful but lacks the friction that would yield deeper insights. A few genuine follow-ups exist (why Claude over OpenAI), but most questions are open-ended rather than pressuring.
If you have to pick one brand, who are you then?
So where does neurometric come in here? What problem are you solving?
Computed from the transcript - who did the talking, and the words that came up most.
Send us Fan Mail Rob May, Co-Founder & CEO of NeuroMetric AI, who’s revolutionizing how multi-model systems think and run, dramatically reducing expenses while boosting performance. Rob brings decades of experience as a founder, investor, and the creator of the popular ‘Investing in AI’ newsletter. 01:05 Meeting Rob May 04:22 The Investing in AI Newsletter 06:50 Starting Neurometric 11:39 Inferencing Defined 16:12 Future Investing in AI 20:24 Model Thinking w/ Guidance 23:40 Neurometric, a Token Engineering Platform 29:07 Data Center Investments 30:42 Real AI ROI 33:48 Where is the Smart Money Going? 35:37 What AI Leaders Must Get Right 39:09 Stock Picks 40:32 Rapid Fire Want to be featured as a guest on Making Data Simple? Reach out to us at almartintalksdata@gmail.com and tell us why you should be next. The Making Data Simple Podcast is hosted by Al Martin, WW VP Technical Sales, IBM, where we explore trending technologies, business innovation, and leadership ... while keeping it simple & fun.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Let's go.
Speaker B: You're listening to Making Data simple, where we make the world of data effortless, relevant, and yes, even fun. Hey, folks, today we're getting into one of the most expensive. I don't know if it's the most expensive, but it's expensive one way or another. But it's the least understood problems in AI, uh, inferencing, not training the models, but the cost of actually running them at scale. I'm with Rob May. He's a co founder and CEO of Neurometric AI, which intelligently orchestrates how AI systems think, picking the right model and the reasoning strategy for each job to get better results at lower cost. He's also the managing director at Half Court ventures with over 100 investments, a repeat founder with an exit, at least one exit, I don't know if many under his belt. And the author of the long running Investing in AI newsletter. I think few people can, um, see the market from the founder seat and the investor seat all at once. Rob does that. So, Rob, welcome to Making Data Simple. Appreciate you being here.
Speaker A: Yeah, thanks for having me. I'm excited to chat more.
Speaker B: If you wouldn't mind, introduce yourself, giving us a little bit of your history and what brings you to today.
Speaker A: Yeah, um, I got started, I was electrical engineering major and I got started doing computer chip design out of college. So I was an ASIC designer for your listeners who know what that is, which would be one of the hottest jobs you could have at the moment if you came out of school today. Uh, but 20 years ago, uh, it was military and space applications were the only places you needed customers, chips. Um, so I did that for a while. I also started a master's degree in computer science focused on AI in the early 2000s. This was like symbolic logic programming in Lisp. So I was just like, this isn't going anywhere. So I quit about 2/3 of the way through. I should have just finished, but I kept an eye on AI for all these years while I went off and did other things. I did a couple of startups when I sold my first startup in December 2014. So my first startup, and it was still my best exit was, was, uh, called backupify. We did backup for cloud computing applications like Google Apps and Salesforce and Office 365. And um, after that I sat back, I took a little bit of time off and I was like, okay, what's next? And I looked at IoT Cryptocurrency and AI. I had read the Google had published this paper in 2014 on word vectors, which was like a big breakthrough moment. And I read that and I was like, oh, that's cool. Um, and again, word vectors were the way of creating embedding such that you could do math with words and so you could do stuff. And the example they gave in the paper that was so famous was they said, um, the word king minus the word man plus the word woman equals what? And the machine spit out queen. And people went, oh, that's cool. These machines could reason. So I decided. So I read that paper and I was like, oh, now's the time for AI. So, um, built a company called Tala as an early customer support chatbot. Had an okay exit there, the product lived on and uh, then went into venture capital for a while and was a general partner at a fund called PJC in Boston. Led some of their AI investing. Ultimately decided there's a lot about VC that I don't like. I mean, I love the actual process of selecting investments, but it's a lot more to the job than that. And uh, I also like to be more hands on and that's not the way to make money in vc. So um, you know, I wanted to build stuff, I want to do stuff, I wanted to sell. And so I went back to the, um, went back to the operating side. And uh, so now what uh, prompted me to think about Neurometric was initially we were looking at different kinds of hardware. So there's a company that's public now called Cerebras that in late 2024 for the first time filed for IPO. And I was looking at the Cerebras chip because, you know, I was a chip designer and I was like, well, you know, if you might start having different chips for different use cases, you need to route between those. And so the initial idea was to start this heterogeneous chip router, uh, and we went out to talk to customers and nobody, even if people were like, this was like early 2025 and people like Cerebras, what is that? Is there something other than a gpu? So left that piddled around a little bit running some experiments and came on this idea of like, wow, as you move from training models to inference, uh, inference cost is going to matter. And here's the key insight of why it matters and what's so different. Typically when you build software, you have not worried about the marginal cost of selling another piece of software because it's been negligible. Um, even as we've stored more data, even as we use more compute still mostly negligible. That is not true with AI. And so these AI companies, people think about them a little bit the way they think about, like, manufacturing companies, where you have the cost of a widget and you're always trying to minimize that cost. And so we were kind of looking at some of this and we're like, okay, there's going to be a lot of things you're going to want to optimize in these systems. And that kind of led to the inception of Neurometric.
Speaker B: You have this long running Investing in AI newsletter. When did you start that?
Speaker A: So I started it in some form or fashion. It's been going on continuously for 10 years. Um, I started it. It's been through three iterations. So the first thing it was a newsletter, um, called Technically Sentient, which is a little bit of a play on words, technical, but also, like, you know, there's like two plays on the word technical. So. And then in 20. So I wrote that for two or three years. And then. And then Jason Calcanis, who's a friend of mine, very popular west coast angel investor, uh, ran. I don't know if it's still around, but he ran the Inside.com network, which had a bunch of newsletters. And he reached out to me and he said, hey, we're going to start an AI newsletter. I love yours. Why don't you roll it into our network? And we cut an agreement. He realized, sid, like, look, you can. I said, oh, you're going to have advertisers. I don't want to be, you know, beholden to the advertisers. He said, write whatever you want. We don't care. I'm not going to censor you. I said, okay. So rolled it in there. And that grew it from like 3,000 readers to like 30,000 readers, which was awesome. And, um, I wrote the weekend commentary, and then they added news through the week. And then, uh, I did that for three or four years. And then it really got to be heavy. And there was a lot of work and there was a lot of AI stuff going on. So in 2021, maybe, I decided to sort of step back from that, bring my readership over to Substack. And I decided to write just about investing in AI because, um, that's what's most interesting to me at the time. And I joined this venture firm and everything else. And so I've been writing the Investing in AI newsletter in its current format, which is, um, commentary on Sunday about the market and then one stock analysis piece through the Week where we sort of look bottoms up at what impact AI is having on a particular company and what their opportunity set might be given these changes in the market. So that's been a lot of fun. And I have a book coming out. I have the first book on investing in AI and it's coming out in, um, uh, September 15th.
Speaker B: All right, all right. There's a lot there, man. That's good. You do everything. You're like the most interesting man alive, I think, right there. Fantastic.
Speaker A: So, you know, what's happened different for me? So I grew up in Louisville, Kentucky. So I got married young, had kids young, and got divorced young. So it's like, I'm in my late 40s, man. I got all the time in the world again. Right. It's like my daughter.
Speaker B: It's a benefit.
Speaker A: Yeah. It's like, ah, all my friends my age have like five year olds, so, you know, I got time.
Speaker B: That's funny. Well, speaking of, uh, so let's backtrack a bit. You founded companies, you've invested in a ton of companies and you've written about AI for a decade. What made you start Neurometric? I mean, what was the. You talked a little bit about it, but why give up everything else? Or I don't know if you've given up everything else, but why throw this back on? I know you got more time, but. Yeah, I'm sure that's not it.
Speaker A: Yeah, it was a couple of things. Um, number one, on the personal side, it was a little bit of. I mean, I had this really great exit with my first company and I thought, man, if I did this well on this company, how much better will I do now that I'm smarter and like, you know, less building a company? And I didn't realize, like, there's a lot of randomness, like, because you're. When you start a company, you're basically placing a bet on the way the future is going to play out. And it may or may not move your. Move your way. It's like when you play a hand at Texas hold' Em and you're like, I've got ace, king, same suit. Like, that's a pretty good starting hand. You could end up getting nothing. Right. You could have ace high, could be your hand at the end of the day. Don't know. Uh, you could get a royal flush, like anywhere in between. So you're. So, so companies are like that and there's more out of your control, I think, than people want to acknowledge. And so I've had some exits since then, but nothing is good. And I was like, I feel like I have one or two more in me and I want to try for something bigger even than my first one. And um, so that leads me to point number two, which is if you're going to play in a market, one way to think about it is bigger markets provide a lot of advantages because you can get a lot of stuff wrong and still have a good exit in a giant market because there's a lot of players, there's a lot of need for stuff. And I believe inference is going to be one of the largest markets in the history of the world. It's going to be up there with like, because it's not just eating into the software market, it's eating into the labor market, it's expanding the labor market. And um, and so that's, that's a really powerful place to be. So. So it's this combination of like, I felt like I wanted to, I wanted another win under my belt, right. It feels really good and haven't had one for a while. And then, um, felt like this was the right market and felt like it was a really unusual opportunity. I mean, I, you know, being in technology for 20 years, I've never seen anything move this fast in terms of, yeah, sure, every layer of the chips, the infrastructure, the models, the applications, like, they're all shifting at the same time. Like it's wild and it's.
Speaker B: Well, how do you keep up with that stuff when it's moving that fast?
Speaker A: Yeah, you know, I don't, man. I just try my best, but it's uh, I think you gotta, you have to, um, you know, you have to do a lot of triage and figure out where you think things are going. And I try to think about markets in terms of like, well, when they settle into an equilibrium someday, what will that look like? There will come a time. We're actually probably not that far from it, but like sometime in the next three or four years, people aren't going to be like, oh my God, there's a new drop from OpenAI. And like, what can it do? Like, people won't get so excited about the frontier because like, we're already at the point where these models do a lot of tasks that we need them to do and they're pretty good. And so this reminds me of, um. Any of your listeners that are sort of probably like over 35 will remember the times when you were like, oh, I've got 133 megahertz pentium, but like, oh, wow. The 200 megahertz came out. You just like. It was like this constant process. Process processor speed upgrade.
Speaker B: Yeah, yeah.
Speaker A: To this point where, like, nobody talked about processor speed anymore. Right. Like, they became fast enough, they could do enough. They were concerned about RAM and other things. I, uh, think the model piece of the industry will get there soon. But anyway, what I do is I try to project ahead and think about what that's going to look like and try to position myself for like. Well, when this market settles into an equilibrium, what position do you want to occupy in that market at that point in time?
Speaker B: If you have to pick one brand, who are you then? Are you an investor by. You know, is that your primary?
Speaker A: Definitely more of an operator these days, which is funny, since I have a book coming out on investing in AI and nobody's written one, so that's part of the reason I wanted to do it. But I, um. Yeah, in my heart I'm an operator. I like doing battle day to day. I like trying to sell things. I like, um, as much as I hate to say it, I kind of like managing people. It's so hard.
Speaker B: But do you still invest? Are you still doing venture or. No.
Speaker A: So I stepped away. So we have a group of funds called Half Court Ventures. There's four of them. I stepped away from making new investments, but I help manage some of the existing investments. And I still go to our thesis meetings, so I still hang out with the team and we talk about what we're seeing in the AI world and where things are going. And so I still do some of that, but I don't write checks very much anymore. Um, not to say that I wouldn't hear and there. I still would be allowed to, but I kind of told them, like, this neurometric opportunity is a really big opportunity. And I wanted to step away and focus on that a little bit more clearly.
Speaker B: Well, when you were writing checks, are these all personal checks or are you managing other people's money or both?
Speaker A: Uh, both. So the first, um. So the first fund that we started, uh, I was the largest lp. I took some of the money I'd made from my exit and put it in that fund. And then people were like, oh, you guys. You know, I had a partner and he's like, you guys are decent at this. We should do another fund. And then after that, it was mostly other people's money in the, in the later funds. And these were all small funds, right? These were all pre seed and seed funds that were $25 million or smaller. We didn't run any 100 million, $200 million funds.
Speaker B: All right, so let's make this simple. In the name of making data simple, I think everybody listening should know what inferencing is. But I'd like to get your definition and why you believe it is such a big cost, uh, and an efficiency problem right now and with ton of opportunity therein.
Speaker A: Yeah. So you can divide the AI market, the AI compute market, into two pieces, training and inference. And the difference is that training is when you're trying to get the neurons, let's say the neurons in the model, similar to neurons in your brain. You're trying to get them all to have the right, um, mathematical weights so that you get the right answers out. And so the way you do training is you. Let's take a simple example. I show the model a picture of a cat. I see what it spits out. Was this a cat? It's not a cat. I go back and adjust the weights until it is a cat. Uh, and then I do that a whole bunch of times. And that's why training is really expensive and takes a long time because you got to do millions and billions and trillions of passes through this with all the data to get the model weights where they work for all the inputs. Inference. Then once the model's trained and the weights are fixed, inference now is just passing something to the model and getting an answer, which is much faster because it's just one pass through the model. Uh, that's not entirely accurate. Um, now that people do reasoning, you can make multiple passes through the model to keep it simple. Let's say that inference is one pass through the model. The model's been trained. Now I present a picture of a cat, and I ask it if it's a cat or not. Right. Um, so inference is interesting because most people build applications on frontier models, the biggest and best models that are available. And it makes total engineering sense to do that because you don't know what you're trying to optimize for. You're just trying to get your application to work. But then what happens is, you build on this giant model, and your inference costs start to climb as more and more of your customers use it and hit that model. And you start to realize, huh, uh, do all these tasks that I'm doing. Need to go to the biggest model. I mean, we had models before. They did some things. They were smaller, they were cheaper. So I'll give you an example. Let's say. So this is where we normally talk to a company. It'll be a company. They're spending 100 to $200,000 a month on inference and it's growing fast. It's actually maybe manageable at, at 200,000amonth, but it's growing 20 or 30% month over month for their use case. And they're like, wow, this is going to be a $5 million annual bill pretty soon. And that's going to be too much. So they start looking at what things they're sending to the model. What are the workflows, what are the tasks you're asking the model to do? There are some tasks that require frontier level intelligence. Some of those tasks might be things like, uh, you're doing some complicated people using this to do complicated mathematical proofs. People are using it to do drug discovery. But if you're sending it something like summarize the notes from my sales meeting. Models have been able to do that for a couple of years and much smaller models can do it. And so if you analyze your workloads and you're like, wow, of our $200,000 a month, $15,000 a month is summarizing m the discussion and the notes from sales meetings. Do we really need to be spending that? I bet we could do that for $1,000 a month with a smaller, cheaper model. Uh, probably an open source model. So there are other reasons and other ways to optimize inference. Mostly when we talk to people it's about cost, but it doesn't have to be. There's two other three other reasons you might want to optimize your inference in a different way from a frontier model. One is latency. If you have like an AI powered search application or a real time kind of chatbot that needs to do a bunch of multi steps like hitting those APIs for anthropic and OpenAI can be slow if you don't need them for the task. And so running a small model, particularly locally, if you can do it, but even if you run in the cloud, might improve your latency by four or five times. It might be significantly faster. Uh, the second would be reliability. Some models and some inference combinations. Some models on some chips are more reliable than others. And what I mean is they consistently give back the same answers in the same time frame. That is not true of all models. And then the fourth reason and final reason might be like data, uh, sovereignty, which is like, hey, for privacy reasons or control or whatever, we just need the data to run over here and we can't use a frontier model or we don't want to use a frontier model. So there's a bunch of reasons that you may want to optimize your inference across those, uh, variables of sovereignty, latency, reliability and cost.
Speaker B: I want to keep the flow going. But this gives rise to as of late, the Chinese just released an AI model, Kimi K3. Ah, you probably know all about this, right? This is open source, right? People are freaking out again, just like they did with deepseek, because we see all these frontier, uh, companies that are investing in all this hardware and infrastructure. Then the Chinese come out with this model, uh, that's open source free. When you put your investor hat on, when you put your inferencing hat on, what does this all mean? Where are we going?
Speaker A: Well, if I put my investor hat on, I had this theory when I was more actively investing before. For this year, uh, there has been a perception, particularly in Silicon Valley, that whoever gets to AGI first wins. And AGI there, I mean artificial general intelligence. So a model that's as smart as a human, why do they believe that? Because they believe if you're first by even a couple of days, that model will start, be smart enough to do its own engineering and come up with its own test and self improve and it'll just extrapolate and blow away anything else. We have not seen anything like that prove out in the lab so far. In fact, the anthropic team spun out of OpenAI and caught and passed OpenAI in just a couple of years. And, uh, people forget that Google still has a thousand times more compute than all of these companies. So if OpenAI invents AGI first, uh, if Google can figure out in a couple of weeks or months how they did it, they can spin up so much more computing power to throw at the next stages and still win. So I have never believed in this idea that you have to be first. And so with my investing hat on, I think you have to invest assuming that whenever we reach AGI or super intelligence or whatever you want to call it, within a couple of months it'll be open source and available to all of us. So what does that mean? As an investor, where do you place your bets? You don't bet them on the model side. Everybody was betting that one model would win it all. Um, and I don't bet that at all. In fact, I think you're going to see the, as models commoditize, I think you're going to see the model companies struggle more and more making money off of models. So it'll be interesting to see how that plays out now I think. Can I give you an answer with my political hat on this?
Speaker B: Yeah.
Speaker A: So the, so I think the administration has done the wrong thing to try to protect us from these models. Right. This is a, when you're in a hyper competitive area, particularly when you have a lead. I don't know that protectionism is a good answer. And we have tried. We had this belief that if we don't ship the uh, top tier chips to China, uh, they won't be able to catch us. Well, that proved not to be true. Now granted, you could argue they found ways around it and everything else, but, but let me give you the example that scares me. The continent of Africa beat the United States in mobile for a long time. Why? Because when mobile phones came out we all had landlines, so we were slow to switch over.
Speaker B: We were good, we didn't need it. Right?
Speaker A: Yeah, they didn't have any other options so they were much more rapid for mobile adoption, mobile payments, mobile gaming, like all those kinds of things. So I worry that if you give other countries the incentive to try to learn new ways of training and training on cheaper hardware, you're actually giving them the incentive to build an advantage against us, an efficiency advantage. So I think you want to be open and free and as competitive as you can be and you gotta trust
Speaker B: that I agree with you.
Speaker A: North American innovation is gonna win here and you don't wanna handicap us because you might get surprising second order effects. And I think you're seeing that. So um, that would be one thing. And then on the inference side I think these chips are gonna um, they're gonna continue to drive down the cost of inference. It's not just the open source models, but you know, these models are all based on what they call a transformer architecture. And there's alternatives to that coming. There's a physics based architecture called a state space model and some of the Nvidia nematron models that have come out are these hybrid state space model transformer model combinations, which is pretty cool. And I think you'll see other innovations even in model architecture, uh, definitely innovations in semiconductor and chip design that will continue to drive these costs down. I don't know that that'll help people on their overall inference cost because they're going to throw intelligence at more problems and uh, in more places, but you'll see it move to the edge, you'll see smarter devices of all types. And uh, but I think people are still going to have to manage their inference cost where when I say manage I mean, you have to match the right model to the right task that you're trying to do. And that's what we don't do well today at most companies.
Speaker B: Yeah. The thing is, you know, back to your, your comment on being open, it's kind of funny to me that I agree with you. What's the saying? Go, necessity breeds opportunity. Something like that. I mean, back in the day when we used to code, when I was coding, they give us the crappiest machines possible with low memory because you'd figure out how to make that code super efficient because you had to. It's the same kind of concept. Now the Chinese are figuring out, all right, you're not going to get us everything. Well, we've got to be competitive. So we're going to find a way and they're going to find a way. We've got every resource available on the planet. And I think it could be to our disadvantage in the hint. I know you talk about models thinking differently depending on how you guide them. What does that actually mean?
Speaker A: Well, there's two pieces to it. One is these models are all slightly, they're all slightly different in how they're trained and they're probabilistic. And so what that means is because they're choosing the next word if I say the boy was riding his bike. Two, well, school is a high probability word, but the grocery is another one, home is another one, the park is another one. And so these models will pick a slightly different answer every time. Now extrapolate that across all the problems you're doing. And they're probabilistic, that's how they work. But they're all trained on slightly different data sets with slightly different training parameters and all this. And what that means is if you have, uh, two models and you ask the same question a hundred times, you might not always get the same distribution of answers and you might not always get the same answers out for even any one on one things. And so the way that affects inference is when you look at a lot of what we do is we work with people who have tried to move a workload to a small model and been unsuccessful. So models are measured by their parameter size and um, you can think about that as the number of neurons in the model. And these frontier models are trillions of parameters now and your mid tier models are hundreds of billions of parameters. But if you want to know, hey, what can a 20 billion parameter model do? Well, a 20 billion parameter model from Quinn, one from Mistral, one from Gemma. Ah, like they're not all the same in terms of their capabilities. And so prompting those, you have to understand their capabilities and the nuances of the model. It's sort of like, um, uh, there's definitely a science to it. Like, there's a way that it works, but there's a little bit of an art to it. Sort of like when you have a. You have one of those door handles in an old house where you're like, you know, yes, you turn it to the right and it works like a normal door handle. You also kind of got to lean into the door a little bit to get it to where, like, you know, the trigger, you gotta, you gotta lift the handle up while you turn it. Like you. Just because it's not in the socket the right way. Like, you know, you learn these things. Um, there's an art like that to these models still because of the probabilistic nature and how they were trained. And so when you prompt them. So first of all, you really want to choose the best model, which is really, really important. And that comes from doing lots of testing and evaluation or buying a product that does that. Uh, but you also want to prompt the model correctly. And you've seen some of this stuff. So if you're doing an investing deck, you want to prompt the model and you want to say you're a partner at Sequoia because they're the number one venture capital firm. And, um, your job is to. I do this with our fundraising deck sometimes, which is I take one model and I say, you're a partner at Sequoia and you're going to raise questions about this model, uh, this fundraising deck I give you, and then I take. I'll ask Claude that, and I'll ask OpenAI or Gemini. I'll take the Claude output and I'll say you're a partner at Sequoia, and, um, one of your funds is. One of the investments you've had is raising more money. And you have to advise the CEO on how to counter these objections. And I feed it the objections from the other model and I go back and forth until the objections feel really weak. But, you know, prompting it in those, in those ways to make it, to make, you know, can make it think better. There was even a time, I don't know if this is still true, but a couple years ago, you could actually get better API access by telling the model in the prompt something bad was going to happen if it timed out on you or something like that. So you could find examples of people doing this Prompting the model is its own kind of art form as well.
Speaker B: So where does neurometric come in here? What problem are you solving? And give us a little bit of the tech behind it that makes it, uh, unique and differentiated.
Speaker A: Well, we call it a token engineering platform. And um, we call it that because what we try to do is help you manage your tokens, make engineering decisions around your tokens to optimize for whatever you're trying to optimize for. Most of the time that's cost, but it doesn't have to be. We can help you with latency and, um, reliability and stuff like that. So what does that mean? The technology consists of a couple of parts. The biggest part and the piece that we're most proud of is, uh, we have an automatic pipeline to create small language models for specific tasks and we make it really easy to continually fine tune those models. And so what that means is if you're sending all your data to OpenAI and it's costing you a lot of money, we can identify the workloads that would probably work with a small language model. Peel those off, they run faster. These, uh, small language models, now you have to be high volume enough to make sense, right? It's like any other engineering decision. If you have two queries like this a day, it's not worth your time. But if you're doing 10,000 of these every day, whatever the query is, and it can go to small language models, it's probably more cost efficient. We typically see a drop of 70 to 80% in price and the performance stays the same or gets better. And the reason is that smaller models use less memory and therefore they run faster and they're cheaper to train and everything else people say. But they're not as good as the big models.
Speaker B: No, no, no.
Speaker A: On any single task you can train them to beat the big models. It's just that's the only task they do. So, you know, you go use a frontier model and it'll give you a lasagna recipe and it'll tell you the history of Rome and it'll also build you an application and it'll also give you a workout guide for this morning. If you fine tune a small language model to tell you about Rome, that's all it's going to do, right? It's really going to fail on other tasks. Um, so that's our sort of bread and butter. But there's other pieces to that platform, to token engineering. We can help figure out when you might want to cache certain prompts. We can Help you evaluate models. Uh, we have a giant sort of like model testing tool where you can take. It's hard to. The industry sort of runs on benchmarks and these benchmarks are like meaningless and they're gamed by the model companies and everything else now. So what you really want to do is you want to take your workloads and you want to run them against a bunch of models. We have a part of the tool does that, um, you can think about it as a tool that helps you and we route between those for you because these are custom models that you're creating. So you need a routing system. It's like the SLM creation, um, the routing piece, the model evaluation and it's all under one platform. It makes it pretty easy for people to work with.
Speaker B: Wow. So is it an IDE then? I mean is it. You're working?
Speaker A: Yeah, no. Um, you wouldn't develop your code in here. What you would do is, um, there's a SaaS app component that you would log into to do the training. If you're going to fine tune a model, one of the things you would do, let's say we have a model that does text to SQL. You type in a query that you want to send to a database and I spit back the SQL that you can feed to that database. Uh, you can upload in the ui. You can upload use cases where it fails and then we'll generate synthetic data around those and we'll continue to fine tune it so it'll get better and better and better. That's one piece. It's just an API outside of that. Right. So you just call us, uh, with one line. Sort of an OpenAI compatible endpoint, uh, is what they would say in the industry.
Speaker B: And it works for all models?
Speaker A: Most models I would say, yeah. So we don't do anything text or audio related. We do anything vision or video related. We do some of that. But like you couldn't use neurometric today if you're like, I'm got this big system that's editing a bunch of video files and it's expensive and I need to run on smaller models. Like we can't do that analysis yet. But it'll work with all the major LLMs, uh, text based models and most of the audio based models is the
Speaker B: monetization by API call or it's twofold.
Speaker A: You pay a flat monthly fee. You pay a flat monthly fee for the management tool. This is your telemetry, your analytics and there's a couple of tiers. It Depends on the number of users, the number of models and endpoints, the number of fine, fine tuning sessions you need to have and then. Yeah, and then we can either host the model for you or manage the hosting for you in which case we just take a cut of the model inference wherever you're, you know, if you're hosting it on together Fireworks or Amazon or gcp. Yeah.
Speaker B: Do you show the ROI up front? So like if I'm a client I can look at it and I could say look if I was just making this call to I don't know, OpenAI or Claude, whatever but now I've made the switch, I'm using you guys, you got made suggestions and look at the savings. I wouldn't have thought of it, I wouldn't have thunk it but now I'm seeing X number of dollars I'm saving.
Speaker A: Yeah, so far I say our average customer saves about 75%. The worst case we've had was about 40% savings. So it's pretty significant. Uh, the more and more use cases we put out there, the more words getting out and the more people are coming in and actually we have a page on the website that has a bunch of sample use cases where you can go in and let's say one of the use cases extracting key entities from a document, you can upload a document or you can use one of the pre uploaded documents and you can run it and you can see the difference in a fine tuned SLM and a frontier model in time and cost and then you'll see a list of the answers that each one provided so you can compare to see how similar they were. And when you see that it's really powerful because you're Normally talking about a 2-4x increase in speed and 80% discounting cost and you'll typically get roughly the same answers out.
Speaker B: This is interesting because um, like I said earlier with all these companies buying all this infrastructure you think all that's going to pay off. I mean they're like betting the farm.
Speaker A: Yeah, I think inference is the compute capacity is the big bottleneck right now and I think it's going to grow. But there's so many ways that are coming to make this more efficient and I think as you roll those out you will probably find it'll work out like the fiber industry in the early days of the Internet where we over invested. Somebody's going to like there's going to be some, there's going to be some defunct data centers or data centers that don't get finished and they'll sit for a couple years and then somebody will pick them up at a lower price. So. And I don't think it's wildly over invested. I think it's going to be a boom or bust thing. But I think we're probably building a little too much capacity given all the operational efficiency improvements that are coming.
Speaker B: I mean, like, uh, on my end right now, to your point, I mean, I've got four models loaded on my laptop right now, my M2, and they answer a lot of the questions. And sometimes I prefer them, believe it or not. Like sometimes when I'm rewriting something or something, I prefer the smaller models because they're more direct, they're more straight to the point than the bigger models. I'm not looking for reasoning in that case. I'm just saying, hey, rewrite this sentence because it sounds terrible. And I'm amazed what it comes up with.
Speaker A: The bigger models like to give you all this feedback, like that's an excellent task. Great choice. And you're like, hey, I don't need
Speaker B: all the hook, I don't need all the narrative.
Speaker A: Just give me my answer, dude.
Speaker B: Yeah, no kidding. So you see hundreds of pitches or you've seen hundreds of pitches? What separates AI companies that, how do I say this? That convert compute into real outcomes from the ones that are just running flashy demos?
Speaker A: Well, it's a tricky question to tell from the outside, um, looking in. Right. But I'll tell you, when you get under the hoods of these companies, there's still a lot of real work to do. There's this idea that these coding tools are so good you can just vibe code everything. And look, the coding tools are very good and you can do a lot of stuff with them. I've definitely made progress, but it still takes a lot of real engineering to build a system for a scalable, interesting product. I think when you look under the hood, I think you want to see people that are thinking about. When it gets easier to do something, you have to think more about where defensibility comes from. So in the AI world, obviously like data loops and learning and getting access to proprietary data is really valuable. May not always be the case. Uh, you can think of scenarios where data becomes so widely available and models get good at combining data in weird ways that in a couple years maybe. Data advantages are hard to come by and uh, hard to find. I think the user interaction patterns, um, a lot of these don't have a UI in the traditional sense, but the UX, uh, is really important. Uh, you asked about IDEs earlier. Sometimes a lot of people pick their coding tool by do they want to be CLI or IDE or something else in terms of how they want to interact with their code. Uh, and so I think the biggest problem with the AI startup ecosystem right now is that the ideas that are getting pitched and the ideas that are getting funded are typically marginal improvements. There's just so much stuff that's like a perfect example is like, oh, uh, these models have a limited context window of a million tokens. And so we are this tool that helps manage and give you a 3 million token context window. That problem's going to go away with changes to hardware and models and architecture and everything else. Uh, I don't know if that's a problem that an independent company needs to solve. It's uh, a marginal problem that all the existing players in the industry are going to work on and solve collectively. I would rather see people really do some breakthrough stuff. So really roll the dice and say, like, hey, this may not work at all, but if it does, it's very like nobody works on evolutionary algorithms anymore, right? Everything's neural networks. Uh, so it would be interesting to see somebody go out and take a swing on really improving evolutionary algorithms and their approach to AI. Uh, there's a really interesting paper written, uh, by a woman named Sarah Hooker, who was, she was at cohere. I think she runs her own company now called Adaptive. She wrote this paper called the Hardware Lottery a couple years ago. And the idea behind the paper is that the best ideas don't always get chosen. A lot of times it's the ideas that fit with the available hardware so, you know, you can run them. And uh, so I think there's a lot of ideas hanging out there in the intellectual ether around AI that nobody's tried because they're just, the space is moving fast and these ideas would take time and they would be hard and they're low probability that they play out. So even though they could be remarkable if they did play out well, um,
Speaker B: where is the smart money going in AI right now, to your estimation? And where's the hype outrunning reality?
Speaker A: Um, I think so a lot of smart money's been going into data centers. I think that's slowing down. Uh, the smart money was recently going into memory because these memory bound, um, there's been a little bit of a pullback there. And the simis as a whole at the time of recording this podcast, I think that'll come back. I, um, think smart Money is starting to think about the, what you would call the economic complements of AI. So what goes up in value when AI becomes popular because of the changes that it brings forward? Like what things that are fine, that we have plenty of now are actually going to start to become scarce. So like you know, land for data centers and all that is something people are talking about and water to support them and all that. But when you think about moving that out to the application layer and you think about, you think about human judgment and you think about um, the number of people like let's go back to the coding example. These tools are really good. They can help somebody who's never written a line of code in their life build a simple app. You are not going to build payroll software with these tools as a person who knows nothing about payroll or has never written software. Like it's just, it's too complicated even with the help of these apps. And so you're still going to need expertise. You're going to need people that know the questions to ask these models. You're going to need people that know how to prompt. You're going to need people that know how to make decisions. When like Claude code will come back and say uh, you know, do you want to try this path or this path? And if both the paths are meaningless to you, then you're just, you know, you're choosing random. And so somebody, so what that means is somebody who doesn't have to choose random because they actually know the answer to that is just using Claude to move faster, uh, is going to be able to win. Um, expertise and judgment are going to go up in value. I think a lot about where the new bottlenecks going to be when average level intelligence is widely available.
Speaker B: If you're a data or AI leader, what should you be doing in the next 12 months and not get caught flat footed?
Speaker A: Well, uh, one thing is this is a place where you have to hedge. And I know it's interesting because core corporate strategy basically says don't hedge. Right. Like you pick a path and you go all in.
Speaker B: It's true.
Speaker A: Yeah. When the world's changing this fast, you have to hedge. So you have to have a little skunks works team trying stuff that's and experimenting with stuff that's not. You're building agents, right? So you decide, here's how we're going to approach building agents. Here's the tools we're going to use. If we use an agentic platform, here's how we're going to select our M models, blah, blah, blah. And you make a commitment to that. You should have a team building, you know, a small team, maybe it's two, five to five people, but like they should be building some stuff in parallel. And it's different and learning because this is an industry where so much stuff could still come out of left field.
Speaker B: All right, so now we're three to five years out. Does inferencing cost stop being a problem or does it move someplace else? What are we underestimating?
Speaker A: Uh, that is a great question. I think inference cost is still a problem because we are so awash in inference at that point for so many things. And the difference between inference and compute is, um, inference is less, or I should say intelligence and compute intelligence is less fungible. And what I mean there is like the units of compute are roughly the same across providers. The units of intelligence, when you look at the model that you choose and the hardware that you run it on and the task that it's built for, are not necessarily as fungible. And what that means is it's going to be harder for any provider, any single hyperscaler or Neo cloud or any provider to compete across all the major use cases the way that it might have been in the compute world or the early cloud world. And um, so I think that's going to keep, uh, that's going to keep everybody busy with pockets of uh, certain types of inference workloads that are intelligence that uh, are always going to be a little more expensive. I think the Frontier labs are going to keep pushing the capabilities of. I think they're thinking about, well, what are the workloads that I can uh, charge this kind of money for? Uh, how good do my models have to be? What are the things people pay 50 bucks per million tokens to solve? I think that'll happen. You're going to have counter forces. A lot of this is going to go to Edge devices and local. You mentioned earlier Al, that you're running a bunch of stuff on your local machine. I thought you're going to see more of that. But these all create management complexity problems where now you got to manage all these models and you got to keep them updated and you got to keep them tuned. Um, it's just like anything else as a human being. You have to stay sharp in your field. So do these models as the world changes and the data distribution moves around them. And so they have to be constantly retrained. Um, and you're going to add more reasoning, you're going to add more post training. Those Things are going to drive up your token usage. So I think your per unit cost of intelligence is going to go down dramatically over the next five years. But I think most companies wind up spending a lot more on intelligence than they are today.
Speaker B: I think you're probably right there. Here's the question of the day. Since you obsess over inferencing costs, do you still say please and thank you to AI when you're asking for?
Speaker A: Uh, I do not. Um, I actually, I came up during command line programming in Unix and so I love just being efficient and talking to the machine. Uh, maybe I should in case all this super intelligent stuff happens.
Speaker B: But, yeah, they'll come back. And you did not say please and thank you. You're done for. That's funny. Hey, you got to give us a stock pick, man. Come on, come on. This is not investing advice, folks.
Speaker A: But good question. So, um, who do I like? Uh, so, two things, right? First of all, I think, um, I said the semiconductors have been beaten down a lot. M. I think Micron has a bunch of upside for high bandwidth. Uh, and full disclosure, I'm a holder of the stock. Um, and then if you want something on the application side, again, I'm a hold of the stock. But I think, uh, HubSpot has done a really good job of, um, applying their business to starting to avoid the SaaS apocalypse. And, uh, yeah, I think it'll be interesting to see where they end up and, uh, if they can apply. I think they're well positioned to apply AI to their user base.
Speaker B: Awesome. Thank you for that. You didn't shy away. Hey, is there a question you would have asked or you wish I would have asked? Um, that. I didn't.
Speaker A: No, I think this was.
Speaker B: We hit it all. We hit it all. All right, I got a little rapid fire, but before I do, where can folks reach you? And where can folks, um, reach Neurometric AI?
Speaker A: Uh, so the website Neurometric AI, uh, is the best way to find us. You can reach me. I'm roburometric AI. We, uh, also, if you're interested in inference optimization and you've heard the term token maxing, we, uh, run a community@tokenminning.com, uh, which is all you know, has a newsletter and a manifesto about using less, using fewer tokens. So.
Speaker B: Fantastic. All right, here's the rapid fire. You ready?
Speaker A: All right, let's do it.
Speaker B: Open source or proprietary AI?
Speaker A: Uh, open source.
Speaker B: Build or buy?
Speaker A: Build.
Speaker B: Build. Okay. One AI tool that you use every day. Um,
Speaker A: I probably have to go With Claude.
Speaker B: Really? Why do you choose Claude? I gotta pause there. I know this is rapid fire.
Speaker A: I never did as well. But I. It was obvious to me early on that OpenAI was going to get sued and have a lot of legal issues. And I also worried that they were trying to do too much and going too broad. So I settled in as a Claude user early on. So it's embedded in some of my key workflows. I actually don't use it so much anymore for coding.
Speaker B: Uh. Oh, really? That's surprising. I see. Coding is pretty good, I think.
Speaker A: Yeah, no, I've moved to a combination of GLM 5.2 with um, Quin 335B, uh, coder. There's a special model for that. You can send the right task to the right thing. There's some tools to help you do that. But, uh, Claude's still embedded in some of my workflows, so it's probably still the single tool that I'm most likely to use on a given day.
Speaker B: Do you use OpenAI? Do you use Groq? Do you use any others?
Speaker A: I try them from time to time, but, um, have never seen a reason to move over. Even if they jump in front. It's typically temporary once you're embedded in workflows.
Speaker B: Yeah, yeah, I get it. Makes sense.
Speaker A: You gotta have a big jump to change.
Speaker B: All right, back to rapid fire. Biggest AI hype that will disappear.
Speaker A: I'm going to say OpenAI. CNBC M A couple of weeks ago said they were the wework of the AI scene, so I'll take that bet.
Speaker B: They're the Netscape. They're going to be gone. Um, one book every technology leader should read.
Speaker A: About AI or about anything.
Speaker B: About anything. Let's go. Anything. Because I want to hear what you got to say.
Speaker A: Yeah, I'll give you one of each. Um, I think, uh, High Impact Management by Andy Grove is still the only hands on book about tech management that's ever been written. Because all these books are like Blue Ocean strategy and Do Whatever and they don't tell you how to do the thing that the book's about. And this book by Andy Grove is like, here is how you design an org structure. Take a piece of paper, think about the workflows that are like he walks you through it and it's really great. Here's how you conduct a one on one with your subordinates. It's just super awesome.
Speaker B: Uh, I haven't read that one. I'll do it.
Speaker A: Uh, and it's an old book, right? It's from, like, the late 80s. Still relevant. Uh, and then on the AI side, I would say Competing in the Age of AI by Kareem Lakhani is really excellent.
Speaker B: Nicely done. And of course, your newsletter, Investing in AI.
Speaker A: Yeah, Investing in AI.substack.com, books coming out September 15th. If you're interested in investing in AI, what the book will give you is a handful of mental models, some of which we've talked about here today. And I didn't necessarily label them as the mental models, but basically, like, we've been investing in a lot of software. Now we're moving to AI. What's different and how do you have to think about the world differently to make good investments here?
Speaker B: Nice. All right. This is fantastic, man. I'm certainly going to check it out. Thank you so much for being here, Mr. Rob May. Fantastic. I wish you luck. It looks like you're good luck in neurometric AI. Uh, you got a book coming out, you got the newsletter, you got everything going on. And you'll probably get married again here shortly.
Speaker A: We'll see.
Speaker B: Ah, you're in New York. Is that where you live, in Manhattan or Chelsea? Nice. Nice. All right. Very good. Well, thank you for being here. It's been really good. I appreciate you.
Speaker A: Yeah, thanks for having me. This was fun.
Speaker B: All right, uh, listeners, hit us on lmarntalksdata at gmail com. You know, we always like to hear from you. Say that every time, but I'm going to keep saying it. So thank you, and we'll see you on the podcast. See you. Bye. Bye.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.