
21 in 21 · 2026-07-28 · 21 min
Key moments - from our scoring
Substance score
53 / 100
Five dimensions, 20 points each
Sharif, who transitioned from crypto and sales into open-source AI development over the past two years, founded local.ai to provide transparent benchmarking data on running large language models on local hardware - from MacBooks to data center GPUs. His motivation stemmed from Cursor's pricing model change in February 2024, which exposed his dependency on expensive cloud services and sparked interest in ownership and cost control. The local.ai project benchmarks model performance across different hardware configurations, measuring both pre-fill latency (time to first token) and decode speed, while running evaluations like agentic tasks (e.g., installing Windows 11, creating accounts) to identify the Pareto frontier of speed versus intelligence for specific hardware. Sharif challenges two key misconceptions: that local AI is prohibitively expensive, and that smaller compressed models aren't effective - citing Quen 3.6 35B performing 15% better than Sunet 4.5 while being dramatically smaller (60GB vs. terabyte-scale). The conversation also covers the open-source community's ethos parallels to Bitcoin, emerging agent payment systems using Bitcoin or USDC, and barriers to agentic commerce (e.g., e-commerce platforms not designed for bot interactions).
Pre-fill is the time the model takes to read and process your entire prompt, while decode is the time it takes to generate each token after that. Different hardware optimizes for different phases, so benchmarking both metrics separately helps users understand how long they'll wait for the first token and sustained generation speed.
Quen 3.6 35B performs 15% better than Sunet 4.5 (estimated 500 billion parameters) on Artificial Analysis benchmarks while being 60GB instead of terabyte-scale, and runs faster on consumer hardware.
Running rigorous agentic evaluations requires multiple attempts (5 runs × 5 samples each per model), and adding compression variants multiplies this workload. A single evaluation can take a day on data center GPUs (B200, B300), costing around $1,500 per day for eight GPUs, with limited provider availability compounding costs.
Cursor changed its pricing model in February 2024 from $20/month unlimited to a credit-based system, forcing him to pay thousands monthly for the same usage. This exposed his dependency on centralized providers and motivated him to buy hardware and explore local alternatives.
E-commerce platforms and most APIs are designed for human users and actively block bot interactions, so even if an agent has a wallet and funds, it cannot execute transactions because the infrastructure isn't built for machines - this is a systems and vendor design issue, not a model capability issue.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains some useful technical specifics about local AI benchmarking (pre-fill vs. decode latency, Quen 3.6 35B vs. Sunet 4.5 performance comparisons, five-benchmark framework) but is heavily diluted by tangential discussion about San Francisco, conference experiences, conference logistics, and casual observations. The substance-to-filler ratio is notably unfavorable for a B2B operator seeking actionable knowledge.
there's like two stages. There's the model reading the prompt, and then there's the model generating the tokens. And different hardware optimizes for different things.
Quen actually does 15% better, despite being, they estimate that Sunet is like 500 billion parameters... versus a model that's like 60 gigabytes.
While the local.ai benchmarking project itself is focused work, the conversation recycles common talking points about local AI (cost, ownership, provider lock-in, data sovereignty). The insights about model compression tradeoffs and the specific benchmarking methodology are moderately novel, but much of the discussion reflects widely-circulated views in the local AI community without significant contrarian perspective or first-principles reasoning.
cost was a big motivator and having that one be kind of outpriced with your abilities. So cost is one thing and then just ownership of the thing is important to me
as many people know, at different times of the day, you're getting different types of service. You're hitting the same plot code or whatever codex, but you don't really know what's going on under the hood.
Sharif is an active open-source contributor and practitioner in local AI with hands-on experience benchmarking hardware and models, which is relevant for the topic. However, he is not a C-level operator, did not build a major company, and is more in the practitioner/developer space than the executive or scaling operator level. His credibility is solid but narrow in scope - strong for local AI specifically, less broadly applicable.
I've been working in the AI space for the last two years. I did a lot of work in crypto before that and sales before that even.
Most of the stuff I do is open source. I mean, it's pretty much since the start of my coding, I've just open sourced everything.
The episode includes concrete examples (Quen 3.6 35B vs. Sunet 4.5, 15% performance delta, 60GB vs. 500GB parameter counts; B200/B300 rental at $1,500/day for 8 units; Google Colab's 16GB GPU) and specific technical workflows (five benchmark runs, five sample attempts per run, pre-fill/decode latency measurement). However, much of the conversation abandons specificity entirely when shifting to San Francisco observations, conference logistics, and vague statements about agent payments and ecosystem adoption.
eight would be $1,500 a day.
Quen actually does 15% better, despite being, they estimate that Sunet is like 500 billion parameters. That's a terabyte, you know, or like, yeah, a terabyte if you don't compress it, versus a model that's like 60 gigabytes.
The host asks relatively soft, open-ended questions that allow the guest to ramble freely (e.g., 'how has your experience been in San Francisco?'). While there are a few good technical follow-ups early on (asking about benchmarking methodology), the conversation frequently drifts into anecdotes and off-topic territory without sharp pushback or productive redirection. The host does not challenge claims or probe deeper into metrics; the tone is more exploratory chat than investigative interview.
you're not from San Francisco. No. Okay, as a non-San Francisco person, I want to know, how has your experience been in San Francisco, especially with all the news and Twitter posts about AI and SF here?
First off, how was your experience at the conference?
Computed from the transcript - who did the talking, and the words that came up most.
0xSero joins Haley Berkoe on 21 in 21 to discuss local AI, open models, and his soon-to-be-released platform for benchmarking open models and hardware ( local.ai ). He explains how his team benchmarks models and hardware, why local AI is becoming more capable and affordable, and why ownership and control matter as more people build their work around AI. They also discuss open-source AI culture, agentic payments, and what people can do to start experimenting with local models. FOLLOW: LINKS
Transcribed and scored by The B2B Podcast Index.
Sharif, welcome to Presidio Bitcoin and the 21 and 21 show. Thank you. I'm excited to be here. Yes, I'm excited that you're here.
You're one of our few guests that have come on that are more in the local AI space. I'm excited to talk to you all about it, and we can tell our listeners a little bit more about even what that is. Definitely. We'll start things off.
For people that maybe don't know who you are, can you just give a little bit about your background and kind of what you're working on? Yes. So, hello, my name is Sharif. I go by Cero Online.
I've been working in the AI space for the last two years. I did a lot of work in crypto before that and sales before that even. Most of the stuff I do is open source. I mean, it's pretty much since the start of my coding, I've just open sourced everything.
And right now I am working on local.ai, which is for local AI. So it's essentially like a website with all the information about all the different hardware and costs and everything like that to give people a better idea of what this is all about and what they can expect, essentially. Very cool.
So yeah, you said that you kind of were doing crypto and sales before, always been an open source. What exactly sparked your interest in AI, specifically local AI? Yeah, to start, there was a project that was kind of crypto, kind of AI. It was Eliza OS.
So Eliza got really popular because you could connect it to Twitter. So it was essentially like running a Twitter scraper. The bot can actually interact and post for you. And it can do actions if it reads something.
So it was like a very early version of Hermes and OpenClaw. This was over two years ago now. And it was also in the cryptosphere. And so these people were like friends or friends of friends.
So I started contributing to it. And that got me into learning about how the technology works under the hood. Then basically around the time Cursor changed their pricing structure, so this was I think in February last year, I was using a lot of credits, but I was paying like $20 a month. It was like just $20 and all you can use, but I was using at least $5,000 in AI credits.
When they cut that out, I kind of started panicking. I realized, well, I'm very dependent on this, and I cannot afford to keep going the way that I'm going. So I started buying hardware and taking advantage of these subscriptions and all the different offerings that exist. And so cost was a big motivator and having that one be kind of outpriced with your abilities.
So cost is one thing and then just ownership of the thing is important to me because as many people know, at different times of the day, you're getting different types of service. You're hitting the same plot code or whatever codex, but you don't really know what's going on under the hood. They could be swapping things out. There's so many different things that they could do to cut costs.
They can take the models away. So if you build a career on top of it, and essentially you're locked into a provider, it would really hurt if that provider just altogether stopped offering you what you use every day. You'd have to buy everything at once, which is very difficult. It's expensive.
So yeah, it definitely makes sense. And something that you obviously hear at Presidio Bitcoin, we care a lot about, you know, having ownership of your data and more privacy and just sovereignty in generally. And yeah, your project, Local AI is super interesting because it does a lot of kind of benchmarking with different models and hardware options. And maybe tell me a little bit more about kind of how you do that behind the scenes.
Yeah, so the company I'm working with purchased a bunch of Macs. So all the different types of MacBooks, Macs, Mac minis. We got like all the cheap GPUs, expensive GPUs. I mean, everything that could run in-house.
So that was step one. Step two is we did like a whole bunch of benchmarking across them to get things like speed. Because, you know, when you are like when you were running a model, I think I explained this before, but there's like two stages. There's the model reading the prompt, and then there's the model generating the tokens.
And different hardware optimizes for different things. So we are trying to make it very clear both, what the pre-fill is, what the decode is, how long you're going to wait to get your first token out. And yeah, so that's one part of it. Another part of it is eval, so evaluations for the models.
So there are a whole bunch of different evaluations. the ones that we like to run are like agentic evaluations. So essentially you give an LLM harness and you give it like a mini computer in the cloud and it's supposed to do something like install the Windows 11 and maybe create an account and do a whole bunch of things in sequence and then you score the model on correctness. So most people are not going to be able to run the original models that are released by the labs but they could run compressions of that So there are whole different types of compressions They all optimize for different things So what we try to do is we have these five benchmarks that we think are the best They cover coding agentic economic understanding and just general knowledge.
And we run it against the original model and then all of its compressions to see which one of the compressions is doing well, consistently across what should people use. And we're trying to get the Pareto frontier of speed performance. So for any specific piece of hardware, like a DGX Spark, what is the fastest model, the smartest model, and the most balanced model that you can run? We're looking for these three things.
So essentially, that's what we're trying to accomplish. And the way we do it, maybe somebody might benefit from this, is we rent B200s, B300s. So these are data center GPUs. They're very expensive.
I mean, eight would be $1,500 a day. But we do this because these evals, like you have to run them five times. So you get a lot of variance, right? Like just randomness.
So you need to run it five times. And each time you run it, you need to run each sample like five times. So like multi-shot attempts. So the amount of compute required to get even one eval out for one model is crazy.
Like it takes like a day to get even like one thing out. So we need to get all these GPUs and run them like all day. And the main challenge there is that, and I'll end with this, but the main challenge there is that it's just like there's not much availability for this compute. So it's like we're constantly bouncing between providers.
And again, it's also very expensive. Very interesting. It's super important. As we see every cost go up, these kind of big frontier labs have more control, keep releasing.
There's going to be a lot more people that are seeing the value and kind of having more ownership and control in their data and their information. I guess, what is one misconception about kind of this more local AI that you hear frequently? Yeah, I think there's two misconceptions I want to talk about. One is that it is super expensive to run.
I think that is a misconception. And then two is that the cheap models are not good. Like, you know, it's not worth doing. So these are the two misconceptions.
So there is a benchmark website called Artificial Analysis. And I think they're the most thorough. Of course, they deal with data centers instead of local. So they have the total benchmark stats.
They have a number on the score. And they compared Sunet 4.5, which was just six months ago, and Quen 3.635B.
And Quen actually does 15% better, despite being, they estimate that Sunet is like 500 billion parameters. That's a terabyte, you know, or like, yeah, a terabyte if you don't compress it, versus a model that's like 60 gigabytes. So the fact that it does better, it's faster, and it can run on pretty much anything. I think a lot more people should try these things.
They are not what they were when we started using ChatGPT or LLAMA, for example. Agree. So you came to San Francisco. You're not from San Francisco.
No. Okay, as a non-San Francisco person, I want to know, how has your experience been in San Francisco, especially with all the news and Twitter posts about AI and SF here? What's your experience been like? versus what you thought, maybe.
I'll try to be honest. Be honest. I'm from here. I love the people.
I think all the people I've met have generally been great. It's very exciting. It's mentally stimulating. People are, on average, more educated, more driven, harder working, have seen more stuff.
And it's nice to be in a big city with a lot of big fish, basically. Now, if we talk about the city itself, so if we talk about the nature, it's incredible, then nature is beautiful. If we talk about the society of the city, I'm not talking about the people or the jobs and that. I think it's just extremely disheartening.
I could say a lot of things, but it's really, really disheartening because we have this beautiful space and then you go out and 30% of the city is people that are half dead that are completely out of it and we're stepping on top of them. What parts of the city were you in? I've went from the top to bottom, right? Like, so Tenderloin, I'm talking about these places.
I know I can just like walk around it, but the point stands, right? Like, we know that there's this huge problem, like really bad problem, like society destroying problem. All the families are going to move out. You know, nobody's going to want to live and walk outside to like a mall with their kids when there's homeless people all over the place, people that are mentally unstable.
So I know I'm complaining a bit. I think you spent too much time in maybe not the best areas of the city. I should maybe give you a tour. You should, you should.
I should have given you a tour. Yeah. I would say arguably probably people in SF think the AI people are going to be driving them out more with the cost going up than the homeless. I think it's like a double squeeze, right?
It is a double squeeze. Yeah. Top and bottom. Okay, so you don't love SF yet.
I do, I do. I love it. It's amazing. It's a really amazing city.
I mean, it's like a pinnacle of this country. I really think this is the best place I been around this country a lot This city is phenomenal So besides that one problem you know everything else is great Well so maybe So you were at the AI Engineering World Fair Yes. That was in SOMA. First, I want to ask you about your experience at the conference.
I was there too. And maybe we can talk about the conference we're hosting here, which is in the Presidio, which is a much more beautiful location than SOMA per se. First off, how was your experience at the conference? And I know there was a whole track on local AI.
Yeah, the conference, I think, was like three or four days, something like that. Yeah, I think four days. And it was pretty good. The first day was slow.
The last day was great. It was the local AI track day. It was the most packed room. I spoke to many people who were speakers in multiple places.
And the consensus, this was the most packed room. Like for the size of it, like there were people sitting on the floor, standing on the walls, like next to the walls, you know, it was full. and everybody was excited. There were people from all of these big tech companies, employees that were there and asking questions.
I think this is a problem that's on people's minds. There was also this band. I spent the three days that were not the AI track just sitting in front of the band the whole time. It was amazing.
Did you do karaoke? I couldn't tell if they were doing karaoke there. They were. They were, okay.
Yeah, so the band was like people who go to the event. This was not planned. They just set up the instruments and people can come in and out. And then just all these musicians were coming in together, singing together.
That was a great experience. This is the first time I got to experience that, like this open-ended band. So I like that a lot. Two complaints.
One is that water was $7. But if you had a water bottle, if you brought a water bottle, it was free. That's true. That's true.
And then two is they made us pay for the food. Those are my two complaints. That's it. Oh, as a speaker too?
No, everybody was paying for the food. I'm pretty sure. Maybe some people, maybe the staff were. It was an expensive conference.
That's true. In Europe, basically, this is not ever a thing to think about. Everything is usually free. I don't need the free stuff.
It's just a little shocking because it's super expensive to be here. And then also the tickets are decently expensive. But I think most of the people that go there get their ticket gifted to them. So I'm not sure.
How was your experience? Yeah, so I went because one of my teammates did a talk. I'm not sure if you went or know much about MCP. He's in the MCP, ACP, Goose world.
Not sure if you've played around with Goose at all. I have. What's your thoughts on Goose? It's very clean.
Very clean. Good. Good. Goose is really important because it's the only kind of open agent that's part of the Linux Foundation.
It was interesting to me trying to go more to talks that were either kind of on that protocol level or on the kind of open source, open weight. And I thought there is a lot of interest out there, like you said, for people to have more control over their data and the costs. And there are a lot of people kind of building similar solutions out there. So that was super interesting to learn about and to kind of see the excitement.
I thought some of the talks were interesting but maybe kind of a little more salesy on the exact thing than the problem itself but overall I thought I thought it was well organized it was a little hard to find the there was not like a big printout of the schedule that's true there was a ton of talks going on yeah there were a lot of rooms there were a lot of rooms and that's hard at big conferences but I've helped with conferences so I know yeah I just really like the quality of people that were there it was very very nice people just everyone was Pretty much.
Everyone was nice. Yeah, I felt like people were really there to learn. There was a lot of people in the room paying attention during the talks. So that was good to see.
You know, I did a lot of crypto events. I worked in the crypto industry for like four years. So all the ETH Denver's, a lot of the Bitcoin events. And it reminded me of that culture a little bit.
Like I feel like a lot of there's a big flight into AI from crypto. Like there's people that like, you know, the frontier of technology, whatever it is. They're all in AI now. It reminded me because of the crypto room.
I wouldn't say all of them, but a lot of it is hollowed out right now. So it's a lot of people like left to other things. They'll be back. It's part of the cycle.
But yeah, it was it was nice. It was a good reminder. Yeah, it's actually it's interesting you bring that up. So obviously, this is Presidio Bitcoin.
And we work on Bitcoin and AI projects here. But something about Bitcoin that's specifically unique with crypto. It's kind of more the ethos angle of it being a bit more decentralized, self-sovereign, open source, things like that. And so it's really interesting now with this wave of people who are into local AI.
Here at Presidio, Bitcoin are caring about these ethos and something that we want to continue to promote to other AIs that are now as privacy, decentralization conscious. So I guess when you're seeing people beyond the cost, have you seen a lot of people drawn to local AI because of the ethos? Yeah, for sure. I mean, I think every time certain companies do something, there's like a wave of interest.
And I can see that because I can see like through my Twitter stats every time they do something my Twitter stats shoot up And I see like a lot of this like newfound joy and appreciation a lot of like younger people that are also like waking up to the fact that hey maybe we should have something that distributed decentralized owned by you know the public the open source world What do you think about open money and being used in terms of agentic payments? I think that will be the next cycle push.
And I don't think anybody that's building it now is going to be a part of that. Maybe cash, maybe Stripe, I think they might have a chance at this. But in terms of like any of these companies that are trying to do like decentralized payments for AI, they're building these things before seeing what the demand looks like. For example, if you ask the model to build a front end, most of the time it's going to go to React, right?
It's going to default to whatever in its training data. And I think Bitcoin is probably going to be the most common or lightning Bitcoin, USDC, ETH, like they're going to be the payment rails. And then in terms of how they transact, this is really a vendor issue, not an AI issue. Vendors right now opened up and accepted payments in crypto online everywhere as a default standard.
It would already be the case that they just go with this. Yeah, it's definitely very early. That I think was another interesting aspect of the conference too, and kind of talking a lot about what they want from their agents and kind of at this point, it almost seems like it's not the model's capability that is the issue. It's more of like the systems on top of it that aren't caught up.
So right now the model can actually make its own wallet, right? So it can make a wallet and it can just send money and receive money. The problem is like, who is it sending the money to and what is it purchasing? So if you think about our websites like Amazon, right?
It's built for people. And they specifically don't want bots interacting with it. So then this model could have like everything it needs, but it can't actually move forward with the purchase because everything is not built for it. We are still building things for people.
When I think, you know, obviously we should build things for people, but we should also build things like with this mindset of like, hey, there's a system or a machine that's going to be reading this. Mm-hmm. Definitely. I mean, so much is changing every month.
It's crazy. Things are moving fast. Well, I guess a few more questions. So for people that maybe want to get more involved in the open source AI community, do you have any suggestions for them?
Sure. So Google has a, I forgot what it's called, but it's, if you search Google Free GPU, you'll get it up. So there's like a website that Google has where they give you a GPU, it's like 16 gigs, and you can run models on it, right? if you have a macbook a mac like whatever you have you can install these models lfm so they're very small and they're very like coherent they're not going to be great for coding but to just understand like what the process is and to get a feel of like what that looks like how fast does it drain your battery how do you actually use it i think yeah trying to play around with this technology is important because even if you don't like it like if you don't like this tech anything it's still clear that there's a huge demand in the market for this technology and the knowledge of how to apply this.
So if you learn how to like set up a local model, you can, that could, that could be part of your offering as like an entrepreneur, you know, you can, whatever it is that you do, you can set up like maybe for schools, right? This local inference and teach. There's so many opportunities if you learn how it works. It also helps you like kind of inoculate yourself.
Because, you know, everybody's telling you, well, this is like a new, this is like the, yeah, they say these things that are like, very exaggerated. And I think, yeah, the data centers are boiling the water, which is not true. Like objectively, it's a closed, like it's a closed loop system. But people either hear something like this, or the news is trying to like, embed this into people.
And they have a lot of misconceptions about what all this is and how it works. That's a good suggestion. I will check that out. I hadn't done that.
Well, we're at time. So yes, thank you for coming. For anyone else who's interested in learning more about local AI, we have meetups at Presidio Bitcoin once a month called Builder, where we talk a lot about it. And we have a summit in September, which is going to be in the beautiful Presidio, the open source AI summit.
Nice area of the city, a lot smaller, and we will have free food. That's going to be great. I'm coming for the free food. Thank you.
There will be free food. I'm really excited for it. And water. There will be water.
There will be water. I will have to say the water was expensive. I did bring a water bottle the next day. What?
Anyways, thank you so much for coming on the show. Oh, if people want to follow your work, what's the best place? I have the same username everywhere. It's 0xsero.
It's the same on Twitter. I'm mostly on Twitter, GitHub, Hugging Face. So I also publish models. So maybe people like it's free.
It's just it makes it fit on smaller hardware. But yeah. Consistency. We like it.
Okay. Well, I hope you have a great weekend and safe travels out of San Francisco. Thank you. I appreciate it.
Yes. Thank you.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.