The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/SaaS for Developers
SaaS for Developers artwork

Optimizing Cloud Costs for SaaS Startups

SaaS for Developers · 2024-04-04 · 30 min

0:00--:--

Key moments - from our scoring

Substance score

51 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality10 / 20
Guest Caliber8 / 20
Specificity & Evidence12 / 20
Conversational Craft10 / 20

Dash Dive takes a novel approach to cloud cost optimization by treating it as an observability problem rather than relying on AWS Cost and Usage reports alone. Founded by recent Y Combinator alumni Adam Sugar and Micah after pivoting from a social app, the company identified a critical gap in the market during the 2023 funding winter when SaaS leaders repeatedly cited cost accountability as a major concern.

Unlike existing cloud cost tools that ingest billing data and provide dashboards, Dash Dive uses agents and agentless monitoring (EBPF) to collect per-event usage data - from S3 object operations to individual database queries to per-API-request compute attribution in Kubernetes. This enables multi-tenant SaaS companies to understand which customers or features are driving costs, rather than simply dividing total spend by customer count. The company built their stack on ClickHouse, Kafka, and Kubernetes for handling petabyte-scale ingestion. Adam discusses the practical reality that this solution fits companies past Series B with multi-tenant infrastructure; earlier startups using ECS/Fargate or single-tenant deployments won't benefit. He also emphasizes Y Combinator's value: structured advice, user-focused product development, and community accountability were instrumental in pivoting from their initial idea to identifying a real market need.

Key takeaways

  • →Treat cloud cost optimization as an observability problem requiring per-event attribution rather than relying solely on AWS billing dashboards, especially for multi-tenant SaaS with shared databases or Kubernetes clusters.
  • →Build unit economics by understanding which customers or features drive costs, not just dividing total cloud spend by customer count, which masks outsized usage and inefficiencies.
  • →Early-stage startups should not over-engineer infrastructure; Fargate and ECS are valid choices, but Kubernetes becomes essential at scale when auto-scaling and resource control matter.
  • →User interviews during downturns reveal durable product opportunities - Adam's team discovered cost accountability problems by talking to engineering teams tightening budgets in winter 2023.
  • →Y Combinator's structured advice on building products, talking to users, and maintaining community accountability is most valuable for founders fresh out of school and early in their startup journey.

Guests

Adam Sugar

Topics in this episode

KubernetesUnit economicsY CombinatorKafkaOpenTelemetryClickHouseAWS S3Dash DiveEBPF agentsmulti-tenant attribution

Questions this episode answers

How do you attribute cloud costs to individual customers in multi-tenant SaaS infrastructure?

Dash Dive uses agents and agentless EBPF-based monitoring to collect per-event usage data (per S3 object operation, per database query, per API request on Kubernetes), then rolls those granular events into a BI dashboard showing calculated costs per customer, rather than dividing total costs by customer count.

What's the difference between Dash Dive and existing cloud cost tools like AWS Cost Explorer?

Existing tools ingest AWS Cost and Usage reports and add alerting and UI improvements, but struggle with sub-resource attribution in multi-tenant setups; Dash Dive provides event-level observability to show exactly which customer or feature is driving cost in shared databases, Kubernetes clusters, or S3 buckets.

When should a SaaS startup start worrying about cloud cost optimization?

The problem manifests past Series A or Series B when cloud credits run out and cost accountability becomes critical for margins and fundraising; earlier stage companies with cloud credits don't typically feel urgent pressure to optimize.

Is Kubernetes necessary for early-stage SaaS startups?

No - Fargate, ECS, and Google Cloud Run are valid alternatives that avoid platform engineering overhead; Kubernetes only becomes necessary at scale when you need advanced auto-scaling and fine-grained resource provisioning.

Does Dash Dive support OpenTelemetry integration yet?

Not yet, though it's on the roadmap; currently customers send events via direct API or install Dash Dive agents, but OpenTelemetry integration would reduce the upfront engineering cost of instrumentation for companies already using observability platforms.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

There are genuine technical insights around sub-resource attribution as a gap in existing cloud cost tooling, and the event-level observability architecture is explained with reasonable depth. However, a significant portion of the episode is consumed by YC experience rehash, generic startup advice, and host anecdotes that dilute the density.

where they really fall short is in sub resource attribution. So uh, in a lot of companies, both big and small, you know, say you use the same very large postgres database to serve all of your customers, or you use the same kubernetes cluster for many customers
we use agents, uh, and agentless monitoring, uh, to collect per event usage. So in S3 that's like per object, put per object, get, we collect all of it

Originality

10 / 20

The framing of cloud cost as an observability problem and the specific insight about sub-resource attribution being the gap in existing tools is moderately fresh. However, the bulk of the episode relies on well-worn YC tropes and generic startup advice that has been repeated countless times.

treating it as an observability concern
codifying all the kind of byzantine billing rules of the cloud provider for each service is not something you probably want to spend your time doing

Guest Caliber

8 / 20

The guest is a recent college graduate (class of 2022) at a very early-stage startup with one known design partner, which limits the depth of at-scale practitioner experience. He has relevant technical knowledge of the problem space but has not operated at meaningful scale himself.

we graduated college in 20, uh, 22, my co founder, uh, Micah and I
we started by working on a social app, you know, as 22 year olds often do

Specificity & Evidence

12 / 20

The episode includes concrete technical specifics - Clickhouse, Kafka, Kubernetes, EBPF agents, per-object S3 event attribution - and names CloudZero as a specific competitor. However, there are no hard metrics on savings achieved, no customer names, and no revenue or cost figures to anchor the claims.

we use Clickhouse and Kafka and Kubernetes um on the beginning of the ingestion pipeline
they're storing petabytes of uh video in Amazon's S3 cloud

Conversational Craft

10 / 20

The host asks some technically grounded questions (OpenTelemetry integration, tech stack) and adds relevant personal context from Confluent. However, the host frequently shares extended personal anecdotes that redirect attention from the guest, and the closing question - 'Anything I should be asking you and I forgot?' - is a low-effort punt.

Do you use anything like open telemetry related? Like if I want to integrate with I have a, I happen to have a database if I want to uh integrate and send metrics in your direction uh is it just about getting an endpoint and open telemetry exporter?
Anything I should be asking you and I forgot.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A70%
  • Speaker B30%

Most-used words

cost33cloud26customers19usage15kubernetes13costs12customer12market9advice9helpful9problem8combinator8build8observability8hard8tools8

Episode notes

You can't manage what you don't measure, and this includes your cloud costs. But how detailed should this measurement be? And how will the data translate into impact? I sat down with Adam Shugar, co-founder and CTO of Dashdive, to discuss his approach to cloud costs. He shared his advice, not only on cost cutting but also technology, growing a startup, the importance of community and more.

Full transcript

30 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Foreign.

Speaker B: And welcome Back to the SaaS developer community where we discuss interesting topics around building SaaS. And today I have with you Adam Sugar. Adam, um, is the co founder and CTO of a very new startup called Dash Dive, and he's helping SaaS companies with, with literally the worst problem, which is analyzing and optimizing our cloud costs. I'm so happy to have you here.

Speaker A: Great to be here. Thanks for having me. M. Gwen, Fantastic.

Speaker B: So you shared with me when we prepared for this session the story of how you got to start this particular company with this particular mission. And I thought it's the best story ever. Can you share that?

Speaker A: Sure, I guess I can start from the beginning. So, um, as you noted when we were talking earlier, we are sort of fresh out of school. We graduated college in 20, uh, 22, my co founder, uh, Micah and I, uh, and we started by working on a social app, you know, as 22 year olds often do. Uh, and you know, we were having some trouble with traction and just generally wanted some more guidance. We applied to Y Combinator, uh, got in with that sort of social app idea and very quickly, uh, you know, they told us we think you should pivot. We sort of obstinately stuck it out for a couple more months, but eventually realized they were pretty much right, uh, and transitioned over to B2B sort of SaaS, which is definitely Y Combinator's, uh, kind of strength. And also we felt it would be nice to have people who actually are straight up and tell you their needs and then you can go build exactly what they need. Um, so we started doing interviews with a lot of software engineers as well as more senior kind of engineering managers, uh, bouncing various ideas off of them. Uh, and this was kind of during the uh, big contraction. This was, we were in winter 23. So, uh, you know, everyone was kind of tightening their belts and uh, the fundraising market was drying up, uh, initially. And uh, the thing we kept hearing over and over is that, you know, people didn't really understand their costs. Cost of capital was increasing. They wanted to increase their margins for their next fundraising round. Uh, and you know, in bigger companies this kind of manifested as, you know, we don't know how to have cost accountability for our managers. You know, each team is sharing the same resources and you know, I don't know how to evaluate my managers. Like, it's easy for me to see revenue for this product, for this team, but it's harder to see cost and harder to enforce accountability. Accountability for that. Um, and in smaller sort of startup Scenarios, uh, it's pretty much per customer. You know, you're doing custom things in SaaS for each customer and you don't necessarily know what features are driving the cost. Um, and so that's sort of how we uh, landed on this problem. And luckily, you know, there's a startup use case which made it way easier to kind of enter the market. Uh, you know, early startups are a lot more willing to take a leap of faith, uh, with a new idea, a new company like ours. Um, so that's where we started, uh, and then we started working with a design partner, a uh, security camera company, um, to help monitor their S3 usage. Uh, and S3, if you're using the same bucket for a lot of customers, is sort of almost inherently multi tenant. Um, so you sort of need observability tooling applied to each bucket to understand like, you know, how much storage is allotted to each customer, uh, you know, how much data egress is allotted to each customer. Unless you're extremely disciplined about, you uh, know, prefixing everything properly and even then it's sometimes hard. So that's kind of how we got to where we are.

Speaker B: Um, yeah, and as you kind of hinted, you're taking a very interesting approach to the entire cost problem. You're treating it as an observability concern. Mhm.

Speaker A: Yeah. Right. So that's sort of the gap that we identified uh, in the cloud cost, uh, sort of cutting cloud cost understanding market. And honestly our product is not the best use case for everything. So you know, for example, there are plenty of cloud cost tools already on the market that essentially ingest your AWS cost and usage report, uh, bills, you know, so at the end of the month you get, you know, per database, per region, per EC2, et cetera. Um, and these tools, uh, ingest it, show it in a much nicer dashboard, build a lot of like very nifty features on top of, you know, as we talked about, there are billions of dollars to be made just on the fact that AWS's UI could use a lot of improvements. Right. So you know, these cloud cost kind of dashboards do an important job in, you know, adding alerting on top of all that. Uh, but where they really fall short is in sub resource attribution. So uh, in a lot of companies, both big and small, you know, say you use the same very large postgres database to serve all of your customers, or you use the same kubernetes cluster for many customers and even down to the individual POD level, uh, the Same POD is serving API traffic for all your different customers, many customers. So even if you were to tag that pod or that database with every single customer or one customer, you don't get an accurate breakdown because just dividing by the number of customers doesn't really tell you if there's some outsized usage or some inefficiency that you need to find and correct. Um, so that's sort of the difference is we use agents, uh, and agentless monitoring, uh, to collect per event usage. So in S3 that's like per object, put per object, get, we collect all of it. Um, in a database it might be a select query, insert query. And in Kubernetes it really depends. With compute, it depends a lot on the use case. So a cron job or a batch job might be monitored very differently than like a high throughput HTTP server. Um, but for example with an HTTP server running on Kubernetes it might be like per API request, try to attribute the CPU and memory usage to that very specific, very small uh, request and then we roll that up over a very large number of requests, a lot of activity, into a kind of BI dashboard that then shows you your usage and the calculated cost based on the cloud provider's billing rules.

Speaker B: And the benefits of treating the problem as an M observability problem is that uh, you get a lot more slicing and dicing flexibility and drilling into details that may be really hard to get out of the AWS cost Explorer.

Speaker A: Exactly. And it certainly only applies to uh, some customers. You know, not every customer has, has the ideal fit for something like this. For example, if you have single tenant infrastructure, if you spin up entirely new infrastructure for every single customer or like run it in their cloud for example, then there's really no use for what we're, what we're providing. And these, the existing cloud cost tools uh, are quite sufficient.

Speaker B: So uh, well there's also, you mentioned earlier the accountability of engineering manager use case which was also back at Confluent I had my team, I had to report cost savings every single quarter. I was looking for projects that would allow my team to report those cost savings. And uh, yeah, it's not always easy to even find opportunity, meaningful opportunities to save cost. Even if as a manager you have the incentives and the ability to do it.

Speaker A: Totally. Yeah. Makes a lot of sense.

Speaker B: Yeah. So obviously observability is kind of a technically hard problem. Uh, uh, can you guys share your uh, tech stack, how you built it, what's uh, your approach?

Speaker A: Yeah, so we kind of had to you know uh, the standard YC advice is do things that don't scale and that's certainly true in some respects for us but unfortunately we had to do the sort of the scalable things uh somewhat out of the gate um because the first customer that we were working with was very high scale. Uh we started with S3 specifically so they're storing petabytes of uh video in Amazon's S3 cloud as well as uh, other S3 providers uh and doing a lot of API requests, a lot of data uh throughput. Um and so that basically meant that we had to use uh like an OLAP kind of analytics ah stack rather than like a more like transactional database row oriented stack. So um, we use Clickhouse and Kafka and Kubernetes um on the beginning of the ingestion pipeline uh we have highly replicated Kubernetes nodes with ingest uh services. So we have an endpoint like an S3 event ingestion endpoint where the client uh sends usage events that happened in S3 in very high volumes and then these highly replicated nodes basically confirm the authentication and kind of transform uh the event uh to be ingested into Clickhouse. In the back end they write it to Kafka both as sort of a back pressure and kind of like queuing mechanism as well as uh, you know uh, just to uh replicate it uh and then from Kafka it's uh, it's ingested into Clickhouse and then finally uh the sort of bi tool dashboard uh from the other side breathes from Clickhouse. So we do a lot of like pre aggregation manual aggregations as well some joins for uh, the more complicated stateful stuff that can't just be done with regular uh materialized views. Uh and then uh, we show that uh, in the dashboard in real time.

Speaker B: Got it. And honestly for an architecture that scales it doesn't sound quite that uh complicated either.

Speaker A: Yeah for now, I mean we'll see how well it scales but it seems to be working well.

Speaker B: Yeah so you mentioned that you have those collectors at the, for S3 and getting data out of all those systems. Do you use anything like open telemetry related? Like if I want to integrate with I have a, I happen to have a database if I want to uh integrate and send metrics in your direction uh is it just about getting an endpoint and open telemetry exporter?

Speaker A: Yeah so. So we kind of have two levels at which you can integrate uh right now. So we haven't built out any sort of Open telemetry integration purely because we don't have customers who are using OpenTelemetry at the moment. But we, we love OpenTelemetry. I mean we're really excited about it and that's definitely on our, our roadmap to build out an integration. Because honestly, you know, uh, as you sort of alluded to, one of the biggest challenges with uh, integrating a product like ours is like high upfront cost. If, you know, if you don't have any observability already, uh, instrumenting your application, then starting with us, it just takes real engineering effort. It's non trivial and so it helps a lot if you already have like a Honeycomb or opentelemetry or something like that that you can just like transform and send to us. So that's definitely on the roadmap. Um, the way it works now is customers can either uh, install an agent or they can send to our API directly. Uh, and the customers we're working with currently all just send to our API directly. I mean some people have concerns about security with running an agent or things like that. Um, but we also sort of offer uh, an EBPF enabled agent. So you just uh, put it on every container or every ec, uh two instance that you're running and want to instrument with batch drive and then with theoretically very little code change, we uh, can sort of automatically collect usage events and then send those to the same API that the customers can manually send to.

Speaker B: Makes a lot of sense. You said you talked to a lot of customers to understand their problems. Did you get a sense of the technologies that are popular? I know you're starting with S3, which is probably the most popular uh, storage system in the world at the moment, but I'm wondering if there's anything else that stands out to you. Like if some of our listeners are starting something right now and they want to use technologies that a lot of other SaaS companies are happy with, what would you recommend?

Speaker A: Yeah, well, um, I think what we've been seeing might not be super applicable to very early stage SaaS companies because it tend this problem tends to manifest when you're past Series A, Series B, um, for the earliest stages, everyone has their cloud credits and frankly they don't care very much about costs. I mean even we don't with our current cloud credits, which is sort of ironic. But uh, for companies that are past Series B or something like this, uh, the big one that we've seen and the next integration that we're working on currently is Kubernetes, almost Every company uh, that we've talked to, whether it's like high throughput, kind of like HTTP server or API server companies or AI companies even that are doing offline batch and model training, almost everyone is using Kubernetes and for many of these companies compute is like the vast majority of their cost. Um, so I'd say learning Kubernetes is definitely a skill that is not going to put you out of the market anytime soon.

Speaker B: Yeah, it's interesting because it used to be Kubernetes used to be one of the things that definitely don't start with. Ah, but from a bunch of interviews that I've been doing over the last two years, I don't think I've met anyone who did not start with that.

Speaker A: Wow, that's quite, maybe you know, different

Speaker B: because Y Combinator does give the do things that don't scale advice. But uh, so maybe you have more stories of people who stayed away from Kubernetes.

Speaker A: Yeah, well we, we have uh, done a couple engagements with earlier stage startups and we find that they are really interested in like avoiding platform engineering altogether, which makes a lot of sense. I mean it's kind of the textbook definition of undifferentiated heavy lifting. And so often they will use ECS or even Fargate, uh, directly and just containerize their workloads to the extent that's possible. I um, think the reason that there's a lot of, you know, we talked to some companies in like the oil and gas sector, uh, who actually one told us that they were even maxing out uh, AWS's instance availability in certain regions to give you an idea of like the scale of compute that they're uh, using. And so I think that you know, once you get to that level and you need uh, a lot of control over like the provisioning of resources, then it makes sense to use Kubernetes. But if you're just trying to get something off the ground, I mean we've definitely seen people use Fargate and ECS plenty. Or Google Cloud Run is kind of the analog as well.

Speaker B: I really appreciate you saying this because we got started with ECS and Fargate and after a while of everyone else talking about Kubernetes, you start thinking, am I doing something wrong?

Speaker A: Yeah, no, I think, I think you're doing it right. I mean I've run into plenty of Kubernetes headaches. You definitely don't want to use it if you're not absolutely sure that you have to. And the only reason we do is for the auto scaling capabilities. Um, and some of. Yes, some stuff like that.

Speaker B: Yeah. So you mentioned that you're part of a recent YC batch and you already gave us one good tip out of Y Combinator. So I'm really curious in what ways you felt that it was helpful for you and if you have a few more tips to share with our listeners.

Speaker A: Sure, absolutely. Yeah. So Y Combinator was definitely transformative for us. I think that you know, we wouldn't be here as a company uh, without the Y Combinator experience that we had. Um, the partners were like incredibly helpful with their advice and it just, I think the biggest transformation, I mean there were many things that are offered but uh, the biggest one for us was just structure. I mean as we talked about, we were straight out of school, we were doing a social app, we had a lot of energy but basically had no idea what we were doing. And Y Combinator is really strong, uh, and it's just hammering of simple advice, uh, that you do well to remember often. And so things like build the product, focus most of your time on writing code, talk to users and really listen to the problems they have. Don't build extraneous features, uh, you know, don't play startup, don't spend all your time going to networking events and you know, doing things, things like that. Um, and that was really helpful for us just to orient, you know, given the limited real world experience that we had, uh, and just to be also be surrounded by a lot of other people in the batch. I mean they do a great job community building and putting you in the same room with a lot of other, you know, really high energy people, uh, and sort of commiserate, uh, you know, who are going through the same things as you at that time and uh, trying to build something. Um, and so yeah, that was sort of the uh, the best part of it uh, for us was just adding uh, structure and giving us some sort of guidance.

Speaker B: Yeah. And it seems like some of those, this is a common mistake. Everyone makes it. Please do not make this one is kind of a recurring theme of uh, the experience.

Speaker A: Definitely. I will say I think the earlier uh, we are, uh, the earlier we were, the more helpful it was for us. I think if we you know, were to go have a go again in five or 10 years with a lot more experience and, and do another company at least in the same sector, we might not uh, get the same transformative experience. But I mean that's just personally our experience. You know, the earlier you are, I think the more helpful the Structure in the community can be.

Speaker B: So yeah, I'm also thinking that having the community can be helpful in cases where you actually know the right thing, but it's hard to get the willpower to do it. So one of the things that my experience was really hard. Everyone always tells founders, launch before you're ready. If you're not embarrassed by your product, you're probably too late. Like I heard it probably north of 100 times in the two years it took us to launch. Uh, but um, it's a lot harder to actually ship a product you're embarrassed of than to know that you're supposed to do it. So I'm getting a sense that if you're a community of people and you get to talk about and commiserate, um, you get a bit of this social pressure that may just get you to do this kind of thing.

Speaker A: Definitely. I, I, that's a great point. I think at any stage, uh, it helps to have just those hard reality checks from the group partners every week. Uh, you know, I mean they, the culture there is sort of one of a kind. I mean they'll tell you exactly what they're thinking. If they think you're doing a terrible job, they'll tell you that you're doing a terrible job. And exactly why, um, which is actually really helpful even if it doesn't sound great to hear. Um, but yeah, to your point about community, I mean we actually still meet uh, every month with our small group kind uh, of co founders, uh, in other companies there's about 10 companies and we check in every month, uh, just because we found those meetings during the batch super helpful and we set goals and uh, you know, having that kind of social pressure of like, well I said I was going to get that customer uh, last month. I can't show up to this meeting empty handed. Like it really helps.

Speaker B: This is so, so true. And it's also, I think we're at a stage where we started having our early customers and having them as a form of pressure is also so useful. I have to get it done because I promised m the customer it will be done by tomorrow. So this has to be done by tomorrow. No matter what happens, it has to get done. So you kind of set yourself up to a point where you don't want to disappoint other people and therefore you kind of push extra mhm. And hopefully customers are nicer, it sounds like, than your Y Combinator group, but you still feel like you don't want to disappoint them. And seeing them um, say something like, oh my God, you guys are so fast. How did you figure out this bug so fast? It is really cool.

Speaker A: Yeah, well, I mean once, once we got really engaged with our first customers, our feature pace uh, of development just went way up. Yeah, exactly.

Speaker B: Yeah. They tell you, as you said in the beginning, they tell you what they need and you just have to go and build it. It's so much better.

Speaker A: That's right.

Speaker B: So we talk on the topic of good advice. A startup that even though it's small we still care about our margins. We definitely care about the unit of economics because I feel like this is something that is hard to fix later if your entire unit economic is wrong. And in the words of charity measure, your startup is actually reselling AWS at a loss. This is not a good business model. And I think even if you're using cloud credits, you want to have a sustainable business model. What kind of advice do you have around optimizing cloud costs?

Speaker A: Yeah, um, it's a good question. So when we were first starting on this path and hearing from engineering leaders about their cloud cost problems, we did a pretty comprehensive survey of the cloud cost market as it stood at the time and it's still largely similar um, to see if anyone was doing this sort of sub resource attribution that we were focused on. Um, and we didn't find anything. But we did learn a lot about the market, uh, and broadly by far the biggest uh, class of companies that we talked about or that we, that we um, discovered uh, or the ones that we talked about earlier which are these sort of like AWS usage reports, ingesting with better ui, uh, as well as some alerting features. On top of that many uh, of these companies also have features uh, where they sort of uh, profess to use AI to automatically cut your costs a bit. Um, how much they really use AI is up for debate but they definitely use some heuristics to you know, find exactly, you know, how many reserved instances, uh, you'd benefit from purchasing, uh, and things like that. Um, and so I'd kind of bucket it into two different kind of uh, strategies, considerations uh, for, for early stage companies, um, even, even mid market kind of companies. And one is sort of okay, we just, our costs are way too high, we need to get them down, uh, and we need to just do a first pass sanity check because you know, just from an entropy standpoint like over time if nobody's monitoring it, things are going to spin out of control. That tends to happen a lot. Um, and so for something like that, those tools are really good. So they will flag, hey, you know you're not using these EC2 instances. Just shut them down on an ongoing basis. They'll tell you that um, they'll tell you general trends uh, in your entire kind of across, across many accounts they're good for aggregation and rolling up across like GCP aws, ah, many different accounts, many different environments. Just giving you like bird's eye view trends of that and telling you if something is like seriously wrong, uh, something that might have slipped into cracks, uh, where uh, we don't feel that they uh, kind of uh, succeed as much is in like real unit economics. Like sure, you can kind of take your total cost for your production environment and divide it by the number of customers and that's your like unit economics so to speak. Uh, but the problem is that doesn't account for various features being offered to different customers. One customer accounting for the lion's share of usage as uh, we've seen with some companies that we've talked to. Uh, and for that you really need observability. Um, and that's where something like Dash Drive comes in or some companies have talked about, well maybe we can home roll this ourselves with open Telemetry with datadog and just getting those usage metrics in there and you know that's a reasonable proxy for costs if you assume, you know cloud providers aren't being too predatory. Um, really the, the value proposition of Dash Dive is kind of to take care of all the kind of cloud provider intricacies for you. So you know, you can collect the usage data yourself pretty easily. But providing the right structure for each usage event and like you know, codifying all the kind of byzantine billing rules of the cloud provider for each service is not something you probably want to spend your time doing.

Speaker B: There is no way I can translate 50 API calls per second going through an API gateway to actual cost. There is so many different AWS components on the way. Every one of them has ingress costs, egress costs, the cost of the device, the cost of the usage of the like it's uh, it's seems almost impossible to do off top of your head to be honest. So I don't think datadog monitoring can help me at all. Like the raw numbers do not easily translate to costs. Uh, maybe at S3 it's still doable. Um, but the uh, moment you have an API like EC2 in front of it and then it has a nut and then it's uh, you have the VPC and You have ingress and ingress.

Speaker A: No, I mean, I totally agree. I, I'm just, I'm giving voice to some of the common objections we hear from engineers. But why can't I just build this myself? I mean, by all means, there is no way.

Speaker B: If anyone is listening to this and thinking they can do it themselves. I, I want to, I would almost say that I would like to quiz the person. I'll show my datadog monitor, tell me how much this costs because there's just literally no way it can be done.

Speaker A: Yeah, so I guess just to summarize kind of to your question of like what kind of cloud cost cutting strategies, uh, should a company think about and you know, what kind of solutions are available on the market. There's sort of general first pass bird's eye view like get a, get a handle on everything. And some of these tools, like Cloud Zero is a fairly mature one that's like one of these cloud cost dashboards will work great for that. Um, if you really wanted to optimize your margin in a customized way that goes beyond like, well, just buy some reserved instances just right size your instances then I think it can often be helpful to have observability data to say like, oh, in our specific case, this one feature is like really inefficient and people are using it a lot and it's not sort of, uh, maybe even priced in. A lot of times we find that companies just are uh, sort of shooting, shooting uh, blind on their pricing and saying, well this is probably, this probably relates to usage, let's charge on these features. But you know, some customers might be getting a really good deal, uh, that you don't want, you know, your margins are too low, um, and there's no way to really know until you like look at the observability data.

Speaker B: So exactly that. The other thing I would give as my advice is to really find a way to keep your finger on the pulse on a weekly or even daily basis. Getting a surprise bill at the end of the month can be traumatic if you did not previously prepare for that. It was traumatizing for me to get some, um, like why is this month 20 times more than the previous months? I did not do 20 times more. And uh, also I saw a lot of the customers I talked to and even people in the community saying things like, I can no longer use this vendor even though I love them because the surprise bill scared me over them. So I moved to a vendor that actually costs me more in normal cases just so I will not get hit with a surprise bill like this is especially for a small startup that is trying to calculate the Runway. This is a very traumatic experience. So really having a system that will ah, tell you not a month later, oops, you spent 20 times more than you intended. But a day later, hey, we're seeing a sharp increase in your cost. You may want to take a look at those five metrics to see what's going on there. It would have saved me so much headache.

Speaker A: Yeah, I mean that's where like sort of the real time data collection and alerting becomes so important. Right. I mean with, you know, these more these traditional cloud cost tools or even with AWS cost and user ports directly, there is a way to enable, you know, hourly roll ups. Uh, AWS I don't believe has very good, I mean I've played around with their alerting and it's like pretty hard to use. Uh, some of the cloud cost tools are a bit better and so you can enable that hourly alerting. See the hourly alerts from those cloud cost tools. Um, that doesn't necessarily help with drilling down to the root cause.

Speaker B: Um, exactly. They don't have drill down for like you either see hourly, which is nice, or you get the drill down. I couldn't find any way to get both at the same time.

Speaker A: Yeah.

Speaker B: So yeah, I'm looking forward to uh, solutions that can do it for me.

Speaker A: Yeah. Awesome.

Speaker B: Cool. Yeah, I think we got all the good advice out of you. Anything I should be asking you and I forgot.

Speaker A: I think we're good. Yeah, I think that about covers it.

Speaker B: Fantastic. It was so good talking to you. So good to listen to your advice for startups and overview of the cost, um, cloud cost optimization space.

Speaker A: I really appreciate it. Yeah, great to be on and thanks for all the questions.

Speaker B: Thank you so much for your time.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • The Forgotten Chapter - Operating a Business Like an InstitutionATLalts · on Unit economics88 / 100
  • ClickHouse with Alexey Milovidov and Austin BonanderRust in Production · on ClickHouse86 / 100
  • DOP 356: Warehouse Robots Are a Distributed SystemDevOps Paradox · on Kubernetes83 / 100
  • #141 AI Pat Works Here Now: Why Agents Must Follow Human Rules with Pat Casey // CTO @ ServiceNowalphalist.CTO Podcast · on Kubernetes82 / 100
  • SECURE& | “75% of Security Reviews Aren’t Code” with Emily Choi-Greene | S5 Ep5The Start and Scale Podcast · on Y Combinator79 / 100
  • How Medical and Dental Practices Can Grow More Profitably with Ibrahim AshmaweyProvider's Edge · on Unit economics78 / 100

More from SaaS for Developers

All episodes →
  • SaaS: More than just a business model
  • Building Streaming on S3
  • Kora: Cloud Native Platform for Kafka
  • Cell Based Architecture for Early Stage SaaS
  • Building a Serverless Streaming Platform
Explore the best B2B Engineering & DevTools podcasts →
All SaaS for Developers episodes →