The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Startups & Founders/SaaS Interviews with CEOs, Startups, Founders
SaaS Interviews with CEOs, Startups, Founders artwork

Featherless: $3.6M Revenue Running 6,700 Open Source AI Models - Eugene Cheah

SaaS Interviews with CEOs, Startups, Founders · 2026-07-01 · 19 min

0:00--:--

Key moments - from our scoring

Substance score

57 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality9 / 20
Guest Caliber14 / 20
Specificity & Evidence13 / 20
Conversational Craft10 / 20

Featherless AI operates as the largest open-source LLM inference provider, offering flat-rate access to over 6,700 models from Hugging Face - a strategic differentiation from competitors who typically support fewer than 100 models. Eugene Cheah and his co-founders (Wesley George as COO and Harrison Van Dyk as CTO) initially built the platform as an internal tool for RWKV, an attention-free AI architecture developed under the Linux Foundation. When they launched Featherless as a pricing experiment alongside their original Recurso platform, it immediately outperformed the original product and became the company focus. The platform abstracts away infrastructure complexity similar to how Heroku or Vercel work for traditional applications, allowing developers and non-technical prosumers to access models like Qwen (Alibaba's multilingual model) and Stepfun without managing GPUs. Revenue grew explosively from launch to over $250,000 monthly within months, with their largest enterprise customer now paying $1-2M annually. The team of 27 (soon 30) is heavily weighted toward infrastructure (12 engineers) and research, supporting diverse use cases from agriculture-specific models to regional sovereign AI implementations. Their go-to-market strategy relies almost entirely on organic discovery through Reddit and Hugging Face communities, with enterprise sales focusing on helping startups reduce OpenAI and Anthropic API costs by 60-90%.

Key takeaways

  • →Featherless differentiates by hosting 6,700+ models including long-tail fine-tuned models that competitors ignore, capturing demand from specialized use cases and regional markets where traditional providers won't invest.
  • →The business model abstracts GPU infrastructure complexity behind simple pricing ($25/month entry to enterprise contracts), making open-source AI accessible to non-technical users and reducing customer acquisition friction.
  • →Enterprise sales close by quantifying cost savings - helping companies burning $100K/month on closed-model APIs reduce spend by 50-80% by strategically routing workloads to cheaper open models.
  • →All 10,000+ customers came through organic channels (Reddit, Hugging Face) with zero paid marketing, indicating strong product-market fit in developer and prosumer communities.
  • →The company maintains a dedicated research team optimizing RWKV and next-generation architectures, betting that AI models can become dramatically smaller and more efficient than current trillion-parameter approaches.

Guests

Eugene Cheah

Topics in this episode

Hugging FaceSovereign AIModel fine-tuningFeatherless AIRWKV (attention-free AI architecture)Open-source LLM inferenceQwen (Alibaba model)Stepfun modelGPU infrastructure (H100, B200, MI325)OpenAI API alternatives

Questions this episode answers

How much does Featherless AI cost and what does the pricing cover?

Entry-level pricing starts at $25/month for unlimited requests to all 6,700+ models (limited to one request at a time), and scales up to enterprise contracts with dedicated capacity for larger production workloads, similar to how Vercel abstracts deployment infrastructure.

How much can Featherless AI reduce my OpenAI or Anthropic API costs?

Featherless claims to reduce inference costs by at least 10x (80-90% reduction) by routing workloads to cheaper open-source models; their sales team audits customer spending and identifies which models can be substituted without impacting performance.

What models does Featherless AI support?

The platform supports 6,700+ open-source models from Hugging Face, including popular models like Qwen (Alibaba), Stepfun, and smaller fine-tuned models optimized for specific languages, domains, and regions that other providers don't host.

How many customers does Featherless AI have and what's their typical usage?

Featherless serves 10,000+ customers ranging from individual developers to enterprises; average entry customers pay $25/month, while largest customers pay $1-2M annually for high-volume inference.

Why is Featherless able to offer models that other inference providers don't?

Most competitors focus on the top 100 models representing 50% of demand, but Featherless targets the long tail where companies run proprietary fine-tuned models for unique use cases - a market too fragmented for traditional providers to serve profitably.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The episode contains some useful operational insights about market positioning, infrastructure costs, and pricing strategy, but is heavily weighted toward product explanation and business metrics rather than novel strategic ideas. Key learnings include how Featherless captures long-tail demand other providers ignore and the cost structure behind high-parameter models, but much of the conversation retreads basic founder story beats and product descriptions without deep exploration of the underlying business dynamics.

[We're] supporting all the other models that people are interested in in experimenting...our strategy is we are supporting all the other models that people are interested in in experimenting and for example step one model is a particularly popular model for us that easily ship several contracts for us uh, on this model alone and no one else is.
we can lower your bill by half and we can see where it goes. That is usually some of our most ideal large volume customers when they are basically spending that much

Originality

9 / 20

The core positioning - serving the long-tail of open-source models rather than competing on top-tier models - is sensible but not particularly original. The broader open-source democratization narrative is commonplace in 2024-2026 AI discourse. Eugene's mention of parameter efficiency is conceptually interesting but underdeveloped and not contrarian given industry focus on efficiency. The execution is novel, but the strategic thinking lacks fresh frameworks.

it's something that we believe that shouldn't be controlled by only a handful of companies where they can choose to restrict your access and what you do with AI.
If you're not going to fight the giant battle of like the top 10 models...You're going to see like 8, 10 providers. Our strategy is we are supporting all the other models that people are interested in in experimenting

Guest Caliber

14 / 20

Eugene is a legitimate founder with demonstrated operational results - $3.6M ARR, 10,000+ customers, and co-creator of RWKV architecture. He's executing at meaningful scale and has skin in the game. However, he comes across as primarily a technical founder still learning the business side, and the episode mostly surfaces business metrics rather than deep operational expertise or hard-won lessons. His insights on infrastructure, customer acquisition, and team building are present but shallow.

We have three founders. Uh, yeah. Uh, so this is probably. It was done by Wes. Uh, so Wesley George was uh, our coo, he's based in Toronto and Harrison Van about, he's based in Australia, our CTO.
we currently have around 27. We are close to past the 30 mark soon. Um, we are aggressively hiring

Specificity & Evidence

13 / 20

The episode provides useful concrete data points: $3.6M ARR, $250K-$500K monthly revenue range, 10,000+ customers, $25-$100/month pricing tiers, largest customer at $1-2M/year, 27-person team (12 infra, 10 platform/GTM, 5 research), and specific model examples (Qwen, Step Fun). However, much of the conversation lacks specificity on customer acquisition costs, churn, unit economics, or technical details about cost reductions. Claims like "10x cheaper" and "80-90% savings" are stated without supporting data or customer case studies.

So currently the biggest would be around 1 to 2 million dollars a year, which may sound extremely large, but when you actually peel behind the layers, it only comes to around like five, six of the largest servers you see in the market.
we have uh, we currently have around 27. We are close to past the 30 mark soon. Um, we are aggressively hiring um, uh, and our team is split across both the platform for deploy engineers and we support the customers, the infrastructure team

Conversational Craft

10 / 20

The host asks solid opening and business-focused questions but largely accepts Eugene's answers without meaningful pushback or follow-up. When Eugene provides vague answers (e.g., on growth trajectory, specific cost savings, or technical differentiation), the host rarely probes deeper. The host shows enthusiasm but conducts a relatively soft interview that allows hand-waving claims like "10x cheaper" to pass unchallenged. Some good tactical questions appear late (team breakdown, largest customer size), but overall the conversation lacks the sharpness expected for substance-focused evaluation.

Why are you guys. I mean I know you're smart, right? I Mean, I don't understand all the research, but I can see you're doing a ton of research. But why are you the only inference provider listed here?
How much of what you do would you say is sort of note people buying you sort of without emailing you, without a call versus high touch you telling people what model they should use

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B61%
  • Speaker A39%

Most-used words

models41model19customers12featherless11today11paying11eugene11help11research10inference9platform9market9first8revenue8million8example8

Episode notes

How do you go from a pricing experiment inside a failing product to $3.6 million in annual revenue with 10,000 customers - in under two years and with zero paid marketing? Eugene Cheah is the co-founder and CEO of Featherless.ai, the largest open source LLM inference provider on Hugging Face. He co-created RWKV, the first attention-free AI architecture under the Linux Foundation, and today runs a 27-person team serving customers that pay anywhere from $25 a month to $2 million a year.

Full transcript

19 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: So you can see RW KV here inside of now featherless. But what you're saying is you effectively were doing all the research here, you built this featherless for yourself. What got you the first bump of signups?

Speaker B: So when we did the first soft launch, we had a few members of the community just posted on Reddit, essentially.

Speaker A: Are you comfortable sharing your monthly revenue today? Is it more like 500,000amonth?

Speaker B: It's something that we are scaling up towards and we are negotiating like multimillion dollar contracts on an annual basis.

Speaker A: How many customers are you serving now today?

Speaker B: So we are serving uh, around 10,000 plus customers.

Speaker A: So just to be clear, we're recording in April of 2026. You're doing more than $250,000 a month, but less than 500,000amonth. You think you'll pass $500,000 a month sometime in fall here of 2026?

Speaker B: Yeah.

Speaker A: What's the largest customer paying you today?

Speaker B: So currently the biggest would be around 1 to 2 million dollars a year.

Speaker A: Hey folks, my guest today is Eugene Chia. He's the CEO and co founder of Featherless AI, the largest open source LLM um, inference provider. On Hugging Face, he's offering serverless access to over 6,700 models on a flat rate pricing model that cuts inference costs by at least 10x. He co created RWKV, the first attention free AI architecture under the Linux Foundation. Eugene, you ready to take us to the top?

Speaker B: Yeah. When you look at this fundamental piece of technology called AI, it's something that we believe that shouldn't be controlled by only a handful of companies where they can choose to restrict your access and what you do with AI. So because of that we fundamentally believe that people should be able to make their own choices and decision when accessing AI models. And the best way to do it is to support all of them in the open source ecosystem. And that's why we built Federless AI to support any AI uh, model that you have on Hugging Face, uh, and to provide instant access to all of that.

Speaker A: And so here's a list of a lot of those models for the non technical listener that's using right now. Can you dumb this down for a second? Explain it like you're explaining it to a kindergartner.

Speaker B: So AI models can be used for various use cases. You have a typical ChatGPT use case, you can have the AI use cases for supporting people in particular language or domains. So for example there are AI models specifically tuned for the use of providing agriculture advice for farmers in both the Asia region and also a few models specifically for the North American region.

Speaker A: So give me an example of how your mom would use this or how this enables someone to build a model or a version of ChatGPT maybe that your mom could use.

Speaker B: So they could use existing applications or even go to featherless AI as an inbuilt application as well as uh, we have a chat application uh, at the top there. And they could actually assess one of the many models that we have in our catalog. And this includes some of the models that we have highlighted respectively. Usually what most people will, what we see from the community is that like be it through the Reddit or local communities, that people will actually find their own preferred model and they'll know the models that they would like to run and they will actually come to our platform and they'll make the request and we'll add support for it and then they can run it straight away.

Speaker A: So let's use the. It's sorted right now. Downloads high to low. Is this pronounced quen? Quen ranks number one in terms of number of downloads.

Speaker B: Yeah, this is one of the most popular downloaded models. So when the user signs up on the platform, they can have access to any of these, these models. So Kwan in particular is one of the latest models made by Alibaba and it's strong in both English and Chinese and, and a few other, uh, additional languages.

Speaker A: You guys can read the description of sort of what it does up here. So let's keep going down the stack here. Eugene, my goal on this interview is to help take a very technical sort of concept and help my listener. Why will you build up Featherless AI? So important in terms of the end value. Right. So is it mostly developers that are paying you for Featherless?

Speaker B: Yeah. So developers. We also see an increased wave of what I call prosumers, uh, entries who's like not exactly purely developers. Some of them will be like coming in and I want to run my cloud code or I wipe coding my apps and my apps run on AI. So these are the two major categories. Traditional developers who are building apps that uses AI and these groups as well.

Speaker A: I have multiple different price points here, but what would you say the average customer is paying you per month today?

Speaker B: So the average entry level customer is paying the $25 a month plan that provides them access to any of the model where they can make unlimited requests, uh, limited to one request at a time and then subsequently once they figure out which models they want and sometimes when they create apps and ship it to The App Store, that's where they go to our scale up plan and that's where they have much larger dedicated capacity because you don't really want your.

Speaker A: Who's paying for the credits though there? If they pay you 25amonth and then they use some of the models that you help them sort of play with and then it gets used a bunch. Who's paying for Correct.

Speaker B: Like so for example, like how Heroku or uh, Vercel abstract the whole infrastructure layer for just deploying your applications. We are abstracting that for AI models and if you want to host and run at ah, larger scale, we also provide the pricing plans for it.

Speaker A: I see, very cool. Okay, it's now making sense to me and I'm not a tech person so I imagine my audience is now following along nicely. How many customers are you serving now today? And then we'll get your backstory here.

Speaker B: Yeah, so we are serving uh, around 10,000 plus customers and I would jokingly say the average customers do not know what B200 Mi325. They may have heard of H100 because it was in the news. And that's essentially what our job is. We abstract away all the complexity of running these AI models and the infrastructure for it. So and this growing collection of open models includes some of the best models that are already on par or surpass, let's say Cloudsonnet or even uh, GPT4 mini. And we realize actually for a lot of customers when they move to production, these are the models that are more than sufficient for the inventory.

Speaker A: Eugene, tell me more of your backstory here. When did you write the first line of code for Featherless? What year?

Speaker B: So Featherless started out kind of like ah, as an accident that outgrew itself two years ago. So as you heard, um, we started from the RWKB committee where we were doing experiments in next generation foundation models. And that's where our roots are. Uh, like this new AI architecture has the potential of reducing inference cost by over 1000x. And if you hear all the energy demands that is extraordinary as well.

Speaker A: What was the first massive sort of signup surge that you saw? Was it a article on Hacker News or somewhere else? What got or Reddit? What got you the first bump of signups?

Speaker B: So when we did the first soft launch we actually just uh, we had a few members of the committee just posted on Reddit essentially uh, to some of the existing AI committees because we knew that we wanted to serve the models that no one else hosted. And that's what really Drove the traffic. You see most providers, they only provide, let's say less than 100 models. That covers 50% of our uh, inference workload is the bottom 50% where they run all these interesting fine tuned models that people came on board for. And they are all usually very unique use cases or languages.

Speaker A: Is this you? Is this your silly tavern? AI, Is that you?

Speaker B: It's one of our staff members that probably did the initial post.

Speaker A: How many co founders do you have?

Speaker B: We have three founders. Uh, yeah. Uh, so this is probably. It was done by Wes. Uh, so Wesley George was uh, our coo, he's based in Toronto and Harrison Van about, he's based in Australia, our CTO.

Speaker A: Okay, wow. So 2024 was official launch then two years ago and you're doing more than $250,000 a month day in revenue. So it's fair to say you've gone from zero to $1 million of revenue like very quickly, right? In a couple of months, yes. You're like, you're an engineer and I'm a business guy so it's uncomfortable when I ask you finance questions, but I love to capture the growth story.

Speaker B: You want to hear the funniest bit about this, right? Is the company at that point in time was called Recurso. So we already had an inference platform for rwkb. Recursion, Recursive model, rwkb. It all makes sense. And Faedalus was meant to be a uh, um, pricing experiment so we gave it a different name. But within the first few days it became more profitable and more revenue than the original company platform that we were like, I guess we are fatherless now. That's amazing.

Speaker A: So what, what did it say? I mean, did you guys go, I mean you guys as your co founders, you must have said, oh my gosh, we just passed 88, $83,000 a month in revenue and we've only been live for like how many months did it take you to break 83,000amonth? Do you remember?

Speaker B: Oh wow. Uh, I can't really remember that moment, but it was like it was all such a blaze because like it was a case of like we get more users, the servers on fire, we add more servers, we get more users, we add more servers. And so it was like just a constant hectic rush there. And it wasn't until like much, uh, more recently where we had a lot, we recently had close, uh, our funding, uh, our latest A round where we had enough server capacity. Like we have a sign of relief now. There is slightly more servers than Users for now.

Speaker A: For now, that's.

Speaker B: I'm all ears for that. And I would like to know how many of your portfolio is burning so much AI usage that we can come in to help them over there.

Speaker A: Guys, remember, I am not just a YouTuber I'm investing into my third fund. We've deployed $250 million into 550 software companies so far. Again@founderpath.com if you're interested in capital I would love to a check because I know you're investing in your education. You watch my show. So sign up@founderpath.com and when you get the onboarding email, I reply and I see all those. Just reply and say nathan, I found you through YouTube and I'll make sure to prioritize you. I would love to cut you a check. Check out founderpath.com tons. I mean this is why it's interesting right? When we look at all the, the profit and loss. So every portfolio company has to connect their profit and loss to founder path. I can look in their cost of goods sold line and I see how much they're paying to anthropic and OpenAI and like on all the credit spe. If you're telling me that you've got a way to help them cut their cogs in half. Right. Or even more, that's extremely valuable.

Speaker B: Exactly. And that's actually how our sales teams are uh, starting to close a lot of these deals because they'll come in and say hey, why are you using AI for do uh, you know, for half of this workload you could use this open model that's much cheaper and lower. We can provide that for uh, this half of this workload you can use this model. And since we have seen them all, we can advise more specifically.

Speaker A: I love that. So how much of what you do would you say is sort of note people buying you sort of without emailing you, without a call versus high touch you telling people what model they should use or how they can save 10 times or you know, 10x their inference spend.

Speaker B: It's actually, it's currently a bit of a both. So when it comes to the individual users when they're coming in on um, the public cloud, this is true word of mouth self discovery for most of the cases and every now and then some of these users will upgrade down the path respectively. But we also realized that there is a lot of money on the table right now where you can go after the startups that hey, I just built my entire startup or SMB on OpenAI or Anthropic and I'm burning $100,000 a month and I did not know what I was doing. And it's like, okay, we can help you here. We can lower your bill by half and we can see where it goes. That is usually some of our most ideal large volume customers when they are basically spending that much and they're entering a situation where hey, we need to start thinking about this and then we step.

Speaker A: Okay, so here's the deal. I know you're an engineer, you're not a deal guy, but I'm going to try and say tell me more about your team, Eugene. How many people are full time today?

Speaker B: So we have uh, we currently have around 27. We are close to past the 30 mark soon. Um, we are aggressively hiring um, uh, and our team is split across both the platform for deploy engineers and we support the customers, the infrastructure team that keeps everything running, uh, but doesn't do anything with building for example. And then we also have the research team that is still working on the RWK line of models because we still feel that there's a lot of room still to optimize these AI models today.

Speaker A: How many of the 27 are working on infrastructure?

Speaker B: So around shelf is working on infrastructure right now and then 10, 10 is moving towards platform, uh, go to market and then the rest is uh, increasingly like GTM, like marketing activities and research.

Speaker A: Okay, got it. So 12 on infra M10 on like platform and then the rest, you know, call it five, six people are sort of research and go to market Motion.

Speaker B: Yes. Uh, the platform does do some marketing as well. Yeah, they do more like, like death route and things like this.

Speaker A: And is all of the inbound right now pure word of mouth or just Reddit or how are the customers finding you?

Speaker B: Um, both via Reddit, hugging face. Various other platforms we have started trying to, we also have started like preparing marketing materials but those haven't really kicked off yet. So it's mostly been Reddit and whatnot. Yeah, right. If you're not going to fight the giant battle of like the top 10 models. Quinn is one of the top 10, uh, top 100 models where all the providers are fighting it out there. You're going to see like 8, 10 providers. Our strategy is we are supporting all the other models that people are interested in in experimenting and for example step one model is a particularly popular model for us that easily ship several contracts for us uh, on this model alone and no one else is.

Speaker A: Why are you guys. I mean I know you're smart, right? I Mean, I don't understand all the research, but I can see you're doing a ton of research. But why are you the only inference provider listed here? Is it just really hard to build that inference model?

Speaker B: Yeah, uh, when you talk about the top 10 or top 100 models, yes they are. If you talk about the rest, that's where we come in. Um, in an AI landscape where try to view it the other way. Like today, a lot of companies have started fine tuning their own models for their own unique use cases and specialization. In a world where various companies are fine tuning their own model, you can't go to a provider that can only support 100 models. There are more than 100 companies on Earth. Uh, you want an infrastructure tailored to be able to handle all these various fine tunes.

Speaker A: So do you have the largest coverage? I mean, is that what you measure? How many, how many models can you cover?

Speaker B: Yes, and that also allows us to have a lot of demand for all these models. Like Step Fund for example is not an unpopular model. It's shipping billions of tokens per day.

Speaker A: Can I see that somewhere on Hugging Face?

Speaker B: Yeah, unfortunately I don't think Hugging Face provide that statistics. But if you go by the download count, it's quite a popular model. Like you have to understand that this is a 200 billion parameter model, meaning you need at least some of the highest end GPU. We're talking about at least uh, 4H100. We're talking about like $20 per hour systems to run this model. And uh, we are talking about like hundreds of thousands of companies are uh, already using this model.

Speaker A: Very, very cool. Okay, I've learned a lot on this episode, I guess. Let me just ask one or two other questions. Like I know Cohere just from my background in like the SaaS space. I don't think they're technologists like you. Why do they have an inference provider option over here?

Speaker B: So Cohere in particular, um, they have their own particular line of models that they created for basically the North American market and in particular Canada. So they are going to the direction of highly tailored sovereign AI models for the domestic market. And we actually see this happening more and more. So for Cohere they will service the Canadian market. For the US market is going to be served by OpenAI and Tropic. For the French market it's going to be served by Mistral for example. But for all the other nations who are not creating models from scratch, they are increasingly actually leaning in towards fine tuning their own specialized models, uh, for their own Domestic use. And that's where we help fill in those gaps. We are not here currently like cohere. As of now. We might in the future create our own line of open models. Uh, but as of now, we are trying to cater more towards all the other revisit use cases where people fine tune their own respective models.

Speaker A: Interesting. Well, this is great.

Speaker B: So for example, KOHI does have contracts with the Canadian governments and so on, things like that.

Speaker A: When you're working on large contracts as well, what's the largest customer? Don't name the customer obviously, but what's the largest customer paying you today is like $500,000 a year or a million a year.

Speaker B: So currently the biggest would be around 1 to 2 million dollars a year, which may sound extremely large, but when you actually peel behind the layers, it only comes to around like five, six of the largest servers you see in the market. Yep, yep.

Speaker A: You guys are just getting started. A lot of growth ahead, hopefully.

Speaker B: Yes, that's what we are. Rather excited.

Speaker A: Eugene, I'm not a technologist, so I'm fully aware that there are questions I probably should have asked that I didn't. Is there anything else you want to cover over the last two minutes here?

Speaker B: So the reason why we still do a lot of, uh, research into next generation AI architecture is that we fundamentally believe that we are still in the very early stages of AI. AI models can be much smarter, much more efficient, much smaller. And the reason why I believe in that fundamentally is I point to myself as a human. I didn't need trillions of tokens to be trained to reach university level intelligence, and I didn't need a trillion parameters to function here. Talking to you, there is something fundamentally that we can do better and that's why we still maintain that research and we apply that as we improve everything along the way.

Speaker A: So guys, if you're listening and you're spending a lot of money on credits, right, uh, go check out Featherless AI, work with Eugene's team and figure out if there's a way that they can save you 60, 70, 80% of your current cogs that you can scale more efficiently. Eugene, how's that for a sales pitch?

Speaker B: Awesome. And if we don't help you save, we'll refund you back. That's what we agree on.

Speaker A: I love that that's recorded. So we're going to hold you to that. But Eugene, this is great. If people want to follow your story online, where's the best place where they

Speaker B: can find you so they can follow me either on, um, Twitter on the handle picocreator. P I C O uh, C R E A T O R or I do also have my own substack which I maintain. Uh, Techtop cto. Otherwise just sign up for the Federalist newsletter that they spoke to guys.

Speaker A: There you have it. Featherless I start off as a research lab, built a tool for themselves. It then started taking off. They have over 10,000 customers today paying between 25 bucks a month and 100 bucks a month. On average, they're north of $250,000 a month in revenue, but under $500,000 a month in revenue. But they're really onto something. Their team is 27 people. They're investing deep in infrastructure, 12 people full time on infra, 5 on research, 10 on platform and go to market. They've taken strategic investment from folks that can help them secure data center processing, help their customers. That's why they've raised from folks in the industry. They've raised a 2 million seed round caught pre 20 around 2023. Pre 2023 and then a big series A in December of 2025 for 20 million bucks. Now scaling quickly, their largest customers are already paying one to two million bucks per year. As Eugene and his team starts to scale faster, their goal is to help you again build on top of customized foundation models and decrease the cost right of the credits you might be currently paying to Anthropic or OpenAI. They think they claim they can save you up to 10 times or decrease your cost by call it 80, 90%. So reach out Featherless AI Eugene, thanks for taking us to the top.

Speaker B: Thank you very much for having me here.

Speaker A: You won't believe this CEO's revenue. Click here to watch the next episode. Right now.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Lewis Tunstall: Hugging Face, SetFit and Reinforcement Learning | Learning from Machine Learning #6Learning from Machine Learning · on Hugging Face89 / 100
  • Mahdi Yahya (Ori / Radiant) on Getting Acquired by Brookfield, Building the Backbone of Sovereign AI, and Why Intelligence Is InfrastructureFWDstart · on Sovereign AI86 / 100
  • Hermes Agent: Agents that grow with youPractical AI · on Model fine-tuning83 / 100
  • China’s AI and robot revolutionThe Business of Tech · on Sovereign AI77 / 100
  • S8E4 - Five Minutes to Fifty Years: Two Futurists on What's Actually Coming - Brett King and Andrew GrillDigitally Curious · on Sovereign AI77 / 100
  • TDM S3 Ep5 - How Slack is Changing DevRel with AI - Kurtis Kemple (Slack & Salesforce AI)The Data Mix Podcast · on Hugging Face76 / 100

More from SaaS Interviews with CEOs, Startups, Founders

All episodes →
  • Gather AI: $15M Revenue With Drones and 170% Net Retention - Sankalp Arora79 / 100
  • Tive Hit $100M Revenue After a Down Round and $10K in the Bank75 / 100
  • Golf Genius: $53M ARR, $10M Self-Funded, Zero VC - Mike Zisman79 / 100
  • $28M Series A at $100M Val: The Pest Control SaaS Nobody Saw Coming
  • He Lost $50/User to Build a $30M ARR AI Empire (Fathom)
Explore the best B2B Startups & Founders podcasts →
All SaaS Interviews with CEOs, Startups, Founders episodes →