
Infinite Curiosity Pod with Prateek Joshi · 2026-02-06 · 42 min
Key moments - from our scoring
Substance score
53 / 100
Five dimensions, 20 points each
Fireworks AI operates at the intersection of models, inference, and systems infrastructure - helping companies deploy AI at scale with reliable, cost-effective inference. Chen draws heavily on his experience at Meta building ads inference on GPUs and ASICs starting in 2017, bringing classical infrastructure discipline to the new AI economy. The conversation unpacks the inference bottleneck: while tokens are conceptually simple, the physical reality involves managing new data centers with uncertain power/networking, GPU failures (like the B200's mysterious 66.2-day timeout), compliance across regions like GDPR, and capacity planning across 20-30 data centers. Fireworks leverages PyTorch, Kubernetes, Helm, and open-source GPU monitoring, while building proprietary abstractions and kernel optimizations on top. The economics of inference centers on total cost of ownership rather than raw speed - balancing reserved versus on-demand capacity, amortizing GPUs effectively, and helping customers avoid scaling into bankruptcy. For enterprises, Chen recommends identifying core AI use cases to build internally versus outsourcing, establishing rigorous evaluation datasets, then choosing models (frontier labs initially for speed, fine-tuning open-source models later for cost) and simple tool-calling architectures with proper knowledge bases via RAG rather than complex multi-agent setups.
Inference faces real-world physical constraints - variable data center power/networking, regional compliance requirements (like GDPR), GPU failures (e.g., B200s timing out after 66.2 days), and the need to maintain stability across distributed infrastructure - that make it more operationally complex than the relatively controlled environment of model training.
First identify which use cases are core competencies that must be built in-house versus outsourced; then establish rigorous evaluation datasets and performance metrics for AI systems; only after those questions are settled should you choose models and fine-tuning strategies, since many companies waste effort on deployments that don't align with business value.
Fireworks moves customers from frontier models (fast iteration, high cost) to fine-tuned open-source models (lower cost, better control), while offering flexible pricing mixing reserved and on-demand capacity and helping them evaluate total cost of ownership rather than chasing raw speed, so they avoid 'scaling into bankruptcy.'
The stack simplified from complex multi-agent setups with cross-checking agents and guardrails to straightforward tool-calling loops; the key is setting up RAG-based knowledge bases and detailed system prompts correctly, allowing a single model call to handle most tasks.
The founding team came from PyTorch; rather than building from scratch, Fireworks leverages PyTorch while carefully choosing which projects to adopt (stable vs. experimental), and heavily uses Kubernetes, Helm, and open-source GPU monitoring, building proprietary abstractions and kernels on top.
Our reviewer’s read on each dimension, with quotes from the episode.
There are genuine nuggets - hardware failure specifics, the compliance-first scaling problem, and the agent-stack simplification observation - but they're interspersed with considerable filler, repetition of 'very very important,' and generic platitudes about trust and quality. The insight-per-minute rate is moderate at best.
I think the B200 will hand indefinitely after 66.2 days. It's like an oddly specific number.
two or three years ago. The architecture is actually very complicated. You have a multi agent setup, you have agents that check each other and you have guardrails on top of uh. These days it's mostly just a tool calling agent that takes a 10 page essay on what you should do as a system prompt and just go and it mostly works.
A few genuinely fresh framings emerge - training/inference fungibility as a procurement principle, the 'ugliness' abstraction metaphor, and the observation that developer busyness has increased rather than decreased - but the bulk of the content (start with frontier models, build evals, pick your core competency) is standard industry advice recycled across dozens of AI podcasts.
tokens are beautiful, but the real world is very ugly
training, inference, fungibility. Uh, maybe that's the right way to put it where when you procure capacity at scale you want to make sure the machine is going to be good, good for all sorts of purposes
Benny Chen is a genuine deep-systems practitioner - responsible for ads inference on GPUs and ASICs at Meta from 2017, co-founder of a scaled inference infrastructure company - not a career podcast guest or pure thought leader. The transcript confirms real operational depth, though he occasionally retreats into vague generalities.
I was responsible for uh, serving ads models on GPUs and ASICs. Uh, I started working on uh, working on having ads model on ASICS in 2017. At that point Nvidia was a small company.
a lot of team worked on inference for most of the tenure uh at Meta or Facebook
A handful of concrete specifics appear - 66.2-day B200 failure, 17W ASIC cards, 20-30 data center scaling, 10-20x cost savings claim - but most claims are illustrative and vague, with few named customers, no revenue figures, no benchmarks, and insurance/medical examples kept at an abstract level.
the B200 will hand indefinitely after 66.2 days
if one of them goes down, then you lose 20%
The host structures questions reasonably well and occasionally gets specific ('does it mean faster decode, more GPUs, better GPUs or better scheduling kernels?'), but there is almost no genuine pushback, no challenging of vague claims, and the rapid-fire section is entirely substance-free. Repeated affirmations like 'that's a great point' and 'right, right' crowd out probing follow-ups.
what levers you have in practice? Like when you say fast inference, does it mean faster decode, more GPUs, better GPUs or better scheduling kernels?
Yeah, no, that's a bad answer.
Computed from the transcript - who did the talking, and the words that came up most.
Benny Chen is the cofounder of Fireworks AI, an AI infrastructure platform. They have raised $327M in funding from Benchmark, Sequoia, Lightspeed, Index, and others.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Benny, thank you so much for joining me today.
Speaker B: Thank you for having me all.
Speaker A: Uh, right, let's start with, uh, the state of play in AI infrastructure, which is on a phenomenal run right now. And it's incredible what's happening. So can you paint a picture of what's happening in both training and inference infrastructure? And also what can we do well and where are the gaps? Where are we still inefficient?
Speaker B: So the current state is the inference or the deployment of AI is really taking off. And Fireworks is very, uh, focused on helping people deploy AI more. There are more and more labs also raising money to work on training. And there are also more and more small companies learning how to train their own models so that they can deploy their model efficiently. Uh, I think going into 2026 there's going to be more focus on, uh, deploying AI into areas where there's positive ROI. People, uh, are less patient about investing into the future and want to see results now. And hopefully a lot of people will see a lot of results this year. Um, and I think, uh, the most concrete evidence of AI taking off this year will be a lot of companies making a lot of money.
Speaker A: Right? Why is, uh, many reasons, but if you had to pick one, like why is inference the big bottleneck? Uh, where are the choke points? And in an ideal world everyone would have, uh, abundance of compute, but clearly that's not the case. So why is inference the bottleneck?
Speaker B: Uh, I think tokens are beautiful, but the real world is very ugly. There's a lot of constraints, uh, that comes from the physical layout of where the tokens are generated itself that results in, uh, inference or deployment, AI being very complicated. So for example, it could be a new, uh, building in some, um, new data center with a questionable power, uh, and networking supply. Um, but you as a consumer for token, just want to see it, ah, keep coming. There shouldn't be any stability issues, there shouldn't be any, uh, failures. So we are here to help abstract away that ugliness. And just so that you see tokens coming out uniformly in the other end. Um, I don't know if you saw a recent GitHub ticket. Uh, I think it was like the B200 will hand indefinitely after 66.2 days. It's like an oddly specific number. And it's not a number that you would have guessed. Uh, but that's what the real world looks like. Uh, the GPU would just die on you after 66 days and there's nothing you can do about it. And there's also a lot of um, new clouds uh coming out, new players uh, in the field, uh, like the networking, the power, the uh, failures uh, and time to fix. Uh, those are all the things that we help abstract away uh, for our customers.
Speaker A: Now Fireworks obviously uh, build a phenomenal business and you sit at the intersection of models and inference and systems. So to build a company like what did you borrow from Classic infra that's been going on for a long time. And also what net new things did you have to do to suit this 2026 AI economy in terms of borrowing?
Speaker B: A lot of team worked on inference for most of the tenure uh at Meta or Facebook. So I worked on ads inference. I was responsible for uh, serving ads models on GPUs and ASICs. Uh, I started working on uh, working on having ads model on ASICS in 2017. At that point Nvidia was a small company. Just so you know, uh, it was the world of uh, V1 hundreds. And I don't actually know how many people have used V1 hundreds. I think H100 will be a lot of people. Um, so in terms of uh, borrowing from the old world, we just knew from the get go that the AI serving is very very important. It's going to be a huge portion of the fleet. Whenever you do capacity planning it's always like the majority goes to serving. So we do that every year. It's sort of like ingrained into us that serving will be the lead. I think what's different this time around is how complicated the stacks are. So I still remember when I was working on the ASIC, I uh, think the whole cards was 17 watts if I remember correctly. Those are the things that you could air cool. You just blow a fan through. And I think it was like 12 cards on that machine and then just blow air through it and it would go away. It's like those laptop level uh, power. And now you're talking about a giant new building with new racks and water cools and thousand watt cards. The current was very very different. The hardware is way more powerful and the models are way more complicated. So the complexities are going up very very quickly. Uh, how to abstract that away from our customers so they can have a good experience. Uh, that's much more difficult than what it was. Right.
Speaker A: Um, let's dive into the inference problem. Let's set the basics first. When someone says hey LLM inference, I want to buy inference from some company, what are they actually paying for? Obviously tokens is the easy, simple way to understand it. But really if you break it down, if I'm buying inference, what am I, like what am I paying for?
Speaker B: Uh, definitely the biggest aspect is definitely the tokens. Uh, you're also paying for someone holding the pager and supporting you as a customer, uh, as well as a partner to brainstorm, uh, a uh, partner to uh, collaborate to deploy your AI more effectively. Because we help a lot of companies deploy their AI, uh, we can provide also good guidance on what's the most effective way, cost effective way to deploy the AI in this scenario. Um, at the end of the day I think everyone's moving very, very quickly when they are committing spend on Facebook. It's also a sign of trust. They're trusting that we can provide good ah, platform for them to use. Uh, the spend either on training or inference or all sorts of different things.
Speaker A: And at small scale, many people, they play around, they tinker and they do like, hey, I have a model, I can do my own inference. But when you go from, I don't know, a small number of tokens to a billion tokens or more, what tends to break first? What is the first big problem you'll face as you go through the journey?
Speaker B: So there's a lot of uh, there's a lot of different components that uh, that breaks. Um, the first thing funny enough is uh, compliance or like that doesn't break compliance, but the compliance becomes more complicated. So for example, like uh, gdpr where you have to run Europe, uh, like I said, the physical world is a little bit ugly. So how many data centers you have in Europe? So how many data centers you have in Australia, those are actually more important when the scale goes up. So that's the first thing. I don't know if it's a good answer, but that's the real answer.
Speaker A: Yeah, no, that's a bad answer.
Speaker B: Where the compliance kicks in and you have to comply to the real world. Um, the second one, I think it's as you scale up, The preparedness for failure becomes more difficult. So for example, if you're talking about say uh, especially uh, fireworks early on when we only had like let's say four or five data centers, uh, if one of them goes down, then you lose 20%. Yeah, um, I think as we scaled up things has become much, much easier where you have like 20, 30 data centers to serve customers. If one of them goes down, you can quickly recover other places. Um, but yeah, as we were scaling up, the second thing that we ran into were just physical uh, constraint because concentration, uh, in certain regions and then probably the third thing that was difficult is also the um, I call it like uh, training, inference, uh, fungibility. Uh, maybe that's the right way to put it where when you procure capacity at scale you want to make sure the machine is going to be good, good for all sorts of purposes. Um, and there are a lot of companies who scale up by taking very opinionated approach and procure capacity that will only be for inference to basically cut some uh, like cut out some cost in their capacity, um model. I do think for us, or at least for me, I value a lot on training, inference, fungibility. Mhm. Uh, and I want to make sure that all the capacity we procure at scale can be used for multiple purposes because that should drive down cost uh, in the long run. Yeah, at the end of the day our customers when they scale up on our platform are looking for reliability, stability and sort of um, ability to spend their money in multiple different ways, either training or inference. Um, so for us scaling up while making sure we can serve both. Training inference at scale is difficult and
Speaker A: if you walk into inference infrastructure and uh, your goal is hey, I need to make it fast. Can you explain what levers you have in practice? Like when you say fast inference, does it mean faster decode, more GPUs, better GPUs or better scheduling kernels? Like what, what are the things you can pull on to make it fast?
Speaker B: Uh, all things you mentioned. Uh, those are all levels. Um, I would say the kernel is definitely still a big part. We still have a lot of engineers working day and night writing better kernels so the inference can be better. Um, the better GPU is definitely a big uh, plus for us as well. We always procure the best gpu um, on the market because it's a very competitive market. Uh, and it's important for us to find uh, the best hardware for our customers. Um, we also do a lot of um, sort of in house, uh, training uh, for speculators to make sure that the speculation uh, side of uh, serving is high quality. Um, at the end of the day though, I think a lot of customers are not just focused on speed. Uh, when we go to market we talk a lot about speed, but at the end of the day is total cost of ownership. They want to make sure that their overall capex is controllable and we are here to help on the overall capex. Um, the way I would think about it is when we talk about speed, it's sort of like an upper bound of how much we can do. But when we Talk with our customers, they are more concerned about tco, like total cost ownership. And that's more like a point on the curve. And they get to pick which point on the curve they want to land. And it's our job to figure out uh, how to push the curve up.
Speaker A: Right, right. It's uh, a great way of looking at it. Real world is just full of trade offs and different customers just want different things and the job is to provide an array of options. Now talking about systems, because you are deep in gnarly systems problems. I'm sure they come up all the time. So just as a starting point, uh, how have you architected the system and also how much do you own versus you just relied on the ecosystem, either open source or you bought a tool from outside. How do you think about the systems architecture here?
Speaker B: Yeah, I think the like buy versus build is a good question. Uh, uh, uh, a lot of the people on the founding team have their roots in Pytorch. So a lot of us when we started we definitely weren't thinking about building things from scratch. Like we definitely leverage Pytorch as much as we can. Um, at the same time because we uh, worked on Pytorch a lot, uh, we knew what are the knobs to pull and what are the knobs not to pull. Uh, specifically like which set of Pytorch projects are like the uh, sort of like the shiny new toy that's still experimental versus the one that's very stable and sort of we get to pick and choose which set of projects to pull in. Um, and then we sort of build our abstractions on top of that. Um, uh, in terms of like using what tools, um, we definitely leverage a lot uh, on the open source ecosystem like Kubernetes, um, like Helm, uh, a lot of different things that we used to orchestrate the fleets, um, and that has kept on giving. I think that the orchestration layer of the ecosystem has really matured and it's been great. Uh, GPU monitoring is also something we pull a lot from the open source. Uh, how to monitor Nvidia GPU is not an easy task. Many players are very incentivized to make sure we can uh, pull efforts together to monitor Nvidia GPUs effectively. Uh, and also AMD, we work closely with AMD on that as well. Um, so in general, uh, we pull a lot from the ecosystem on the orchestration layer, the Pytorch layer and also the monitoring layer.
Speaker A: Earlier you mentioned total uh, cost of ownership. So let's Dive into the economics a little bit. So just in plain English, what are the main line items and inference cost. If you had to explain to somebody, maybe a net new customer, what are the main line items here
Speaker B: for our customer? Uh, they mostly concern about uh, what's the cost per million token, let's say. And we sort of present the overall picture to them on how much the token costs. We don't really talk about the line item or whatnot. But I think a lot of like a big nuance in the industry is uh, how much do you commit versus pay as you go? I think that is a big open question, uh, to the infrastructure or like the, not just infrastructure, but like the whole frontier, uh, lab business, uh, model where um, because people leapfrog each other all the time sometime uh, uh, I don't know if you saw the diagram where it's open air and probably Gemini
Speaker A: this part of the cycle.
Speaker B: Yeah, they take turns on who's ahead. Um, so how do you even commit to something? Uh, it's almost impossible to commit to something at the end of the day. I think people will have, when we present the items, some will be reserve capacity, some will be on demand capacity and the customers will sort of commit to multiple uh, vendors, uh, just wait for, while waiting for them to leapfrog each other and use the model as they go. Um, but yeah, I do think the whole industry is very aggressive on pay as you go. Um, but at the end of the day someone's renting the capacity or someone's uh, holding uh, the capacity at a steady state. And then um, since if you apply the logic that uh, electricity is only small part of the cost, um, then how do you amortize the GPU effectively? And uh, I think even how do investors think about how to amortize GPUs effectively? I think that's still a big open question.
Speaker A: Yeah, no, that's a big um, dedicated capacity versus pay as you go. Because dedicated is cheaper. But you have to plan it just right or else you'll be left holding the bag. Uh, versus pay as you go is just more expensive because somebody else is holding the bag for you. So they'll got to pay a premium to them. So it's a very interesting uh, problem now going into customers use infrastructure for all sorts of reasons. Um, let's talk about fine tuning for a second. So many models available now, open source people are taking it, they're trying to figure out how do you make this fit into my enterprise use case. So how do you help or guide them to think about. Hey what's, here's how you think about fine tuning versus just prompting or just use whatever is available in quad. How should they think about that if you're Walmart for example.
Speaker B: Yeah, I think that's a good question. So uh, I will probably reframe uh, the question a little bit on what sort of like um, uh, forward leaning traditional enterprise should be doing. Um, because in general there's a lot of cases where it's not straightforward for you to deploy within the company. Sometimes it is important for you to procure solutions directly. Um, so I think the first call to make is what's the core use case that you cannot give up that you have to figure out within the company and what are some of the use case where you just want to outsource? Because I think the first failure pattern is try to tackle too many things at the same time. Um, often the main date from the board is a little bit vague and rightfully so because the board doesn't have all the context. Right. But just like deploy AI, right, it's very important to figure out where to deploy the AI. And that I think often is the first but the most important question because unless you figure that one out, um, you're going to waste a lot of time and energy on things that uh, doesn't matter to your core business. So the first thing we normally recommend is the core competence. Let's figure out how to deploy the AI uh together the things that's not in your core competence. Just buy some subscription somewhere, don't spend too much time there. The second thing for these forward leaning enterprises to figure out is how to evaluate their core competence. I think it's very easy to say that you have some sort of performance review process. It is much harder to say that now you have to apply this performance review process on a bot because they're much more nuanced. Um, you can probably assume a person is going to respond in some predictable way. Um, even if they hallucinate, let's say they're going to still respond in some uh, predictable way because they went to school, they have friends, they know the context in the society. If they are taking a customer support call, they will respond in certain ways. So then UPSC is your performance review for them is more like you need, you cannot say certain things at the same time. Ah, at certain time. Right. It's probably more straightforward when a bot gets on a phone call and try to answer customer support. The bot was trained to be very general purpose. It can handle math Questions it can handle very complicated uh, technical questions uh, about software engineer. But uh, like how to express empathy, how to properly handle customers requests, uh, those are very very difficult. So how to evaluate your bots correctly? I think that that's the second most important thing. After that, once the businesses have evaluations and data sets, things become much easier. For example, on a uh platform you can train like a uh model with reinforcement, fine tuning. It can learn the task very well and it can be deployed in that setting. Uh, a lot of businesses get stuck on what task to deploy on and how to evaluate a task. Once those two uh, those questions are settled, everything else is very easy.
Speaker A: Right, That's a great point. Um, going into model choices so as inference infra you get to touch and provide many different models, different customers have different requirements. So for some company that as you said the board says hey, do AI stuff right? And then they come to you say like hey here's a mandate. Now help me think about. You mentioned use case specific but what does best model mean for a given use case? Meaning like is it about accuracy or is it about speed or cost? Or maybe a combination of both. Like how do you help them think through what does best mean in a
Speaker B: given case in a competitive market? I think the best model should save you time. Uh, it should help you iterate fast. Uh, and it should not bankrupt you when you try to scale up.
Speaker A: Right.
Speaker B: Uh, so time is important as the first element. So ah, a lot of businesses definitely start on the most frontier model because m they don't know what they are iterating on. They often don't have evaluations yet. And using the best model will save you time because then you don't have to spend as much time thinking about the system, prompt thinking about how the model would behave and just go with it. Uh, go ahead, go ahead.
Speaker A: No, go ahead.
Speaker B: As the business scale up it is more important to think about fine tuning and making sure you don't scale yourself into bankruptcy. Um, so having those evaluation, having uh a rigorous way to pick the best model, uh having evaluation for you to even train the model is important. Uh, a pitch I often uh have for my customers. Even if you don't want to consider fireworks, even if you, even if you think we are not competent or not good enough in certain aspects, having the evaluation will at least help you pick between the three or four frontier labs because like a new model just got released from OpenAI at this point they have so many models. Which one do you use? How do you convince yourself that this is a better model than uh, whatever I'm using right now. Having that kind of confidence will help yourself. And as soon as you have evaluations, uh, it's so much easier for us to figure out how to fine tune your model on top of open source and save you like 10x20x the money, right?
Speaker A: And once the initial like uh, the tinker, they play around and they're like okay, for this use case we roughly know what we want. And then after that what's your view on model diversity in production? Meaning like is it, hey, just go with one model, not worth the hassle. Or it's good to have many models or maybe a routing layer. Like how do you help them think through after the initial tinkering? Like what should be the good architecture?
Speaker B: By architecture, uh, the architecture I think is very very different depending on the customer. Uh, use case. So uh, let me think about what's the best way to answer the question. Uh, let's focus on the forward leaning uh, enterprise. Let's say a uh, lot of it has to do with some sort of knowledge base. To start with. These companies like a big customer service machine, they have a lot of information uh, from the customers, a lot of context on the business side. Um, and it's very important for them to have first a good knowledge base for the bot to operate on. Um, I think a lot of in the current state there's a lot of sort of data that the bot doesn't have access to because those in the uh, those are sort of like tribal knowledge that the human has. For example, like uh, um, what was a good example? I think like for example we work with a lot of like uh, medical companies and insurance companies and what the acronym means is very, very different in different forms.
Speaker A: Right?
Speaker B: And sometimes it is like very amazing that AGI would just figure out that in this medical form this acronym means this thing. If you ask me, I would never know. But at the same time you're setting up the bot for failure. So in the other 50% of the time the bot just never figures out what this acronym means in this form. Um, and just having that context uh, that like if you're probably like an insurance underwriter you're probably seeing like two uh, thousand, sorry 20,000 forms already. You know what these acronym means? Giving the bots some sort of internal knowledge base, some sort of knowledge bank, ah is sort of a prerequisite. Uh, so you have to set up rag, you have to set up your system prompt properly. Once that happens. Um, I think a lot of the deployment these days are very very straightforward. Um, uh, I think two or three years ago. The architecture is actually very complicated. You have a multi agent setup, you have agents that check each other and you have guardrails on top of uh. These days it's mostly just a tool calling agent that takes a 10 page essay on what you should do as a system prompt and just go and it mostly works. So the stack has been simplifying and as long as you set up the uh, environment and the uh knowledge base correctly, uh a lot of the uh stack are just a simple uh loop of uh two calling uh model.
Speaker A: Right now let's talk about uh company building uh for a minute. As a starting point you're shipping fairly complicated software. For many people it's fairly complicated. So how do you guide your team around? Hey, here's how we ship. Do you have uh, uh, guiding principles or what do you tell your team as you've grown? How do you ship product?
Speaker B: Definitely quality first. So I would say code is getting cheaper every day and it's very important for us to adjust to that and definitely we push our team to not be married with your code. It's like you can re architect things very very fast these days but making sure that you ship something that the customer can uh, use very quickly and save them time, that's very very important. At the end of the day when we work with customers we always want to figure out a way where uh, we minimize the amount of time they spend on us just to get value very quickly. Uh, sort of like ah, the mentality is more like a utility search engine style thing where I think maybe um, for Facebook it's more important to figure out how to get people to uh, stay and have some fun. For utility like uh, Google and Uber, it's more like how to get the job done very very quickly and predictably consistently. Um, so definitely the bigger m focus is on quality and making sure the code we ship, the software we ship uh, are stable and predictable. Um, way more important than shipping more new features.
Speaker A: And if you looking back you're building a product which is mostly used by developers and developers are famously very unforgiving, very strict with what they, they don't like. So what's the most surprising thing you discovered? Building for developers, Building a product like an infra product for developers.
Speaker B: Uh, maybe it's the flip side where uh, they are very forgiving if you screw up, uh, but if you don't screw up and consistently perform they will reward you with trust. Um, I do think that time is important. Maybe I keep coming back to it and that might be a pitfall. And say this gets repetitive. Uh, the developers are actually busier than before, believe it or not, because they can do so much more. Their mental bandwidth is on how to build a product. Uh, we want to get out of their way as quickly as we can. We just want to provide tokens in uh, a predictable way and they don't have a lot of time to think through, hey, uh, who do we pick? Sometimes it's just a knee jerk reaction where like, oh, I've seen these people, I've used these people before. They are predictable, they're reliable. And that kind of trust I think is actually going to be more important as the bot takes m over more of the stack because we need to make sure that developers trust us, um, that trust that we can uh, do it and do uh, it in a reliable way. And the surprising thing is that once we build a trust, um, it's actually very helpful for us to scale the company because the people who trust us go to different companies, they use us as a tool, they get used uh, to it and uh, fireworks as a tool will proliferate in the new company they land into.
Speaker A: Yeah, actually that's an incredible point and something I think about quite frequently for a developer. You just need to be predictable and reliable and like, hey, just merge into my background so that I don't have to think about you because I know exactly what you do and that's great. That's why infra. For infra, it's almost beneficial to be just super boring and just like you just work all the time so that developers are like, yeah, that's my default, I use them, I pay for them and uh, I'm gone. So it's great. Um, also when you went from zero to the first five customers, obviously a bunch of things you had to try, some things wouldn't have worked. So if you were guiding a new young infra company about selling for uh, zero to five customers, like what, what learnings do you have that you can share with them?
Speaker B: Yeah, to be perfectly honest, I think even today we still have a lot of work to do on go to market. And go to market is incredibly difficult in a competitive landscape. Um, So I think what to do is actually a very hard advice to give for generic like it's hard to give generic advice. But maybe uh, maybe I can talk about like what not to do as an infra company is like, I think the uh, there's a lot of like new Companies that sort uh, of very, very focused on marketing instead of just go to market like Cluli. I think I uh, see their short videos all the time. Uh, I think maybe that's a good strategy for a product company as an infra company. I do think it benefits uh yourself when you come across more serious, uh, uh, because you're marketing to developers, not to the general crowd. Um, so I think probably the marketing for. We still have a lot of work to do and I don't know what generic device to give but probably not uh, uh, short videos on uh, interesting things that's not business related.
Speaker A: Right, right. Well that's great. I think uh, it reminds me of. It could be like some form of corollary like in the famous uh, opening line of Anna Karenina. Like it's called something like all happy families are alike, but all unhappy families are unhappy in their own way. But in this case like okay, success can be had in different ways but there's a common, few common failure notes like don't do that because it's just not going to work in infra. Um, that's right. Okay, so I have one last question before we go to the rapid fire round. Um, AI is moving so fast. What AI advancements are most exciting to you as it pertains to building fireworks? What excites you the most in 2020?
Speaker B: What excites me the most is that the open Source models in 26, the Frontier ones finally have vision input and I think that maybe uh, that's a bad answer maybe uh, but like recently Kimi released and they have vision uh, input and that's very very exciting believe it or not. Because uh, for a lot of use cases, uh, for example coding you just want to take a screenshot from Figma and just say like hey, write the code and stuff for me. Um, I think that's a trivial and funny ah, answer but at the end of the day I'm very excited about the prospect of uh, open source model in 26. Um, as the competition heats up and as more players come into the field and things get more fragmented, um, sort of. We are very committed to skill work with open source and making sure that we provide things that people have access to in a predictable way and we really sort of are looking for a future where the model building itself becomes more commoditized where so the open source model hopefully converges with the frontier model. This year we saw a glimmer of hope with Llama four or five. That was the peak meta days where you see the Open source model really catching up. But I think in 25, actually, the gap has been widened, um, because Meta is not like you need a lot of money to make these work. Um, but I do think recently we've seen a lot of really good releases from open source. Uh, and I think this year is the year, ah, where they will catch up. Uh, and also this year is the year where a lot of labs, uh, will release their models as well. They've been baking for, like a year or two, so the playing field will get very exciting this year. It will be a very chaotic scene. Uh, and m. Uh, we always advise the company we work with. It's even more important now to have an evaluation, because right now you're picking between four. You will be picking between 10 very quickly.
Speaker A: Right, Right. Yeah. Um, amazing. All right, with that, we are at the Rapid Fire Round. I'll ask a series of questions and would love to hear your answers in, uh, 15 seconds or less. You ready?
Speaker B: Go ahead, go ahead.
Speaker A: All right, question number one. What's your favorite book?
Speaker B: Uh, the, uh. Right, I forgot the title. Sorry. It's Ray Dalio's the Principal. Yes, the Principal.
Speaker A: Right, right.
Speaker B: Yeah, I just remember the COVID
Speaker A: Happens all the time. I can know exactly the color. Color and what the story is about, but the title is there. I keep blanking out. That's hilarious. All right, next question. Which historical figure do you admire the most and why?
Speaker B: Uh, like FDR with a New Deal, it takes a lot of effort. Takes a lot of effort to drive a completely new, uh, paradigm. Um, just the energy that this person has, uh, to push through things and keep not giving up, maybe. Sorry, that's a long answer for your question. So I'll keep the show by FDR with a New Deal.
Speaker A: Got it. Yeah, that's a great answer. Um, next question. What's the one thing about AI infrastructure that most people don't get?
Speaker B: Like, how ugly? That was the answer I gave earlier. But, yeah, that's my honest answer as well.
Speaker A: That's a fair answer.
Speaker B: Yeah.
Speaker A: Um, next question. What separates great AI products from the merely good ones?
Speaker B: Wow, that doesn't take explaining. I, um, think recently, like, the Clothbot, sorry, Mobile, uh, has taken over the Zeitgeist. Right. Like, you just see the bot in, uh, WhatsApp or Telegram. You don't need to explain to people what this is doing. Of course. Like, the setup is very elegant. It's very complicated. Uh, there's probably a lot of, uh, engineering iteration to it, but you don't need to explain to People what it does, it's sort of people just gets it. And I think those are the best products.
Speaker A: What have you changed your mind on recently
Speaker B: that the demand is so crazy that the speed of which we deploy AI? I think I was a little bit of a pessimist, to be honest. Um, uh, it is about to take over the world this year. Uh, there's a lot of Demand for deploying AI this year.
Speaker A: Uh, what's your wildest AI prediction for the next 12 months?
Speaker B: Okay, to be perfectly honest, at this point, I don't think whatever I can come up with can be wild. Because you already went through the. I don't know if you've seen the answer, uh, from uh, Davos. Like all the leaders in the AI, they already give outlandish predictions. I don't know what kind of wild answers I can come up with, to be honest. But I do think it's like, uh, I don't know, it's like carburetor enthusiasm kind of thing where like, uh, I still don't, I still, I still need to be educated on robotics. I don't understand some of the, uh, best case scenarios for robotics. And that's where I need to learn more than be educated on my current understanding. It's like, it's GPT2 but not GBD3. Uh, so it's still the, so like the quirky phase. I don't know if you still remember, like, is it, uh, Jasper and like a lot of companies where you were using GPT2. I think it's going to be a very different form factor and product when they land on GPT3. Um, so yeah, that's my understanding of robotics at GPT2. Right now someone need to educate me on why GPT3 is coming for them because I don't have the technical capability to understand that.
Speaker A: Right. All right, final question. What's your number one advice to founders who are starting out today?
Speaker B: Does your spouse, ah, support you? I say half as a joke. Uh, it's a very important thing to straighten out before you go for it. Um, uh, especially for people who are married, uh, it's very important to make sure you, your, your family is behind you and support you because there's a lot of, gonna be a lot of, uh, ups and downs. Right. Um, I really appreciate my wife, like supporting me all the way through. And it's very important for you to straighten out with your spouse. Like, hey, like, there's going to be a lot of ups and downs for the next two years. Right. And a lot of people who are able to come, uh, out and found a new company, have a cushy job. They are not, they're very smart people. They're not in some low paying job in their company right now. So they technically could just quiet, quit. Um, a lot of time. These people have the drive, they want to go for it. Um, whether their family have their back often is very, very important.
Speaker A: M. Right. Amazing. Uh, Benny, this has been a phenomenal discussion. Loved your hard earned insights, uh, on infra and company building. So thank you so much for coming onto the show and sharing your insights.
Speaker B: Thank you so much for having me. Thank you.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.