The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/TestGuild Devops Toolchain Podcast
TestGuild Devops Toolchain Podcast artwork

AI-Powered Predictive Autoscaling for Kubernetes with Jennifer Rahmani

TestGuild Devops Toolchain Podcast · 2025-08-13 · 35 min

0:00--:--

Doris AI tackles a fundamental DevOps pain point: reactive autoscaling that forces teams to overprovision infrastructure and remain on-call for predictable traffic patterns. Jennifer Rahmani brings deep expertise from nine years in defense tech DevOps, where she experienced firsthand the frustrations of manual scaling policies and guesswork-driven infrastructure decisions. Doris AI's predictive autoscaler integrates with Kubernetes via a single Helm chart, ingesting metrics from observability tools to forecast workload demand across multi-cloud and hybrid environments. The platform moves customers from typical 40-50% utilization (wasting compute) to 85%+ utilization while maintaining high availability. Rather than jumping on the AI bandwagon, Doris deliberately uses lightweight machine learning models suited for numerical time-series data instead of LLMs - a philosophy Jennifer calls 'use the right AI for the right use case.' The platform runs air-gap in customers' environments, scales up before traffic spikes and scales down during low periods, and also flags anomalies for security investigation. This resonates with cost-conscious, reliability-focused engineering teams tired of Black Friday disasters and unexpected product launch outages.

Key takeaways

  • →Doris AI's predictive autoscaler moves Kubernetes infrastructure from 40-50% utilization to 85%+ by forecasting workload demand 5 minutes to 6 hours in advance instead of reacting in real-time.
  • →The platform uses lightweight machine learning models specifically chosen for time-series infrastructure data rather than LLMs, embodying the principle of using the right AI for the right use case.
  • →Doris integrates with Kubernetes in minutes via Helm chart, runs air-gap (no data leaves your environment), and can operate in autonomous scaling mode or recommendation mode.
  • →The technology helps prevent outages during unpredictable events like Black Friday, product launches, code deployments, and unexpected traffic surges by learning workload patterns over time.
  • →Doris also detects infrastructure anomalies and unexpected behavior patterns, alerting engineers to potential security issues or unusual scaling needs requiring investigation.

In this episode

  1. 1Jennifer's DevOps background and frustrations with reactive scaling
  2. 2Founding Doris AI with her twin sister to solve scaling problems
  3. 3Why machine learning is the right AI for predictive autoscaling
  4. 4How Doris AI's predictive autoscaler works with Kubernetes
  5. 5Real-world use cases: Black Friday, product launches, and unexpected traffic
  6. 6Utilization improvements and cost savings through forecasting
  7. 7Security anomaly detection and alerting capabilities

Mentioned

Doris AIJennifer RahmaniKubernetesSmartBearJoe ColantonioInsight HubHelm

Guests

Jennifer Rahmani

Topics in this episode

KubernetesMulti-cloud infrastructureDevOpsMachine LearningObservability toolsSite Reliability Engineering (SRE)Doris AIPredictive autoscalingHelm chartsBlack Friday traffic patterns

Questions this episode answers

How does Doris AI predict Kubernetes scaling needs?

Doris integrates with your Kubernetes environment via Helm chart and ingests meaningful metrics from observability tools (like CPU, memory, or custom metrics). It trains machine learning models on historical patterns to forecast workload demand 5 minutes to 6 hours in advance, then automatically scales pods up before traffic spikes or down during low-traffic periods.

Does Doris AI send my infrastructure data to external servers?

No. Doris runs air-gap, meaning no data leaves your environment. The platform processes all metrics locally within your Kubernetes cluster and makes scaling decisions autonomously without external dependencies.

What utilization improvement can we expect with Doris AI?

Customers typically move from 40-50% cluster utilization (with manual reactive scaling) to 85% or higher utilization by predicting demand in advance, freeing up unused compute and reducing cloud spend while maintaining reliability.

Can Doris AI help with unexpected traffic spikes like Black Friday or product launches?

Yes. Doris learns workload patterns and detects anomalies, allowing it to handle unprecedented traffic spikes more gracefully than reactive scalers. It scales preemptively for known patterns and reacts quickly to sudden, unexpected events.

Does Doris AI work with multi-cloud and hybrid Kubernetes environments?

Yes. Doris integrates with Kubernetes clusters regardless of where they run (AWS, Azure, GCP, on-premises, or hybrid) and combines infrastructure metrics with real-world signals to make scaling decisions across any environment.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B72%
  • Speaker A28%

Most-used words

scaling24help23data20customers19case18thoris17sure17engineers15infrastructure15black15better15scale15reliability14learning13traffic13devops12

Episode notes

In this episode of the TestGuild DevOps Toolchain Podcast, host Joe Colantonio sits down with Jennifer Rahmani, Co-founder and COO of Thoras.ai, a company redefining how infrastructure scales with AI-driven predictive technology. Drawing from her years as a DevOps engineer in the defense tech sector, Jennifer shares how she and her twin sister turned real-world frustrations into a reliability-first platform that eliminates the guesswork from scaling. We discuss how Thoras.ai integrates with Kubernetes to predict workload demand minutes - or even hours - in advance, allowing teams to maintain high availability without overspending. Jennifer explains why they use the right AI for the right use case, how their predictive autoscaling works in multi-cloud and hybrid environments, and how it helps SREs avoid downtime during unpredictable events like Black Friday or major product launches. Whether you're dealing with noisy data, high cloud bills, or sleepless nights worrying about reliability, this episode delivers practical insights for making smarter scaling decisions.

Full transcript

35 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Get ready to discover some of the most actionable DevOps techniques and tooling, including performance and reliability with some of the world's smartest engineers. Hey, I'm Joe Colantonio, host of the DevOps Toolchain podcast, and my goal is to help you create DevOps toolchain awesomeness. Have you ever been jolted awake at, uh, 2am by a scaling emergency? If so, you're not alone. What do you do to avoid this? Well, today you're in luck because Joining us in this episode is Jennifer Rahmani, a co founder and COO of Doris AI, a company redefining how infrastructure scales with AI driven predictive technology. And drawing from her years as a DevOps engineer in the defense tech sector, Jennifer shares how she and her twin sister turned real world frustrations into a reliability first platform that eliminates the guesswork from scaling. We discuss how Doris AI, uh, integrates with Kubernetes to predict workload demand minutes or even hours in advance, allowing teams maintain high availability without overspending. Jennifer also explains why they use the right AI for the right use case, and how their predictive auto scaling works in multi cloud and hybrid environments, and how it helps SREs avoid downtime during unpredictable events like Black Friday or major product launches. Whether you're dealing with noisy data, high cloud bills or sleepless nights worrying about reliability, this episode is going to help you learn more how to make smarter scaling decisions. You don't want to miss it. Check it out. Hey, before we get into this episode, I want to quickly talk about the silent killer of most DevOps efforts, that is poor user experience. If your app is slow, it's worse than your typical bug. It's frustrating and in my experience and many others I talk to on this podcast, frustrated users don't last long. But since slow performance is as sudden, it's hard for standard error monitoring tools to catch. And that's why I really dig SmartBear's Insight Hub. It's an all in one observability solution that offers front end performance monitoring and distributed tracing. Your developers can easily detect, fix and prevent performance bottlenecks before it affects your users. Sounds cool, right? Don't rely anymore on frustrated user feedback, but I always say try it for yourself. Go to smartbeer.com or use our special link down below and try it for free. No credit card required. Hey Jennifer, welcome to the Guild.

Speaker B: Hi Joe, thank you for having me.

Speaker A: I'm, um, excited about this interview. I saw an article on your company and your founding. I thought, wow, this'd be A great episode. So before we get into the meat of the episode, I like to learn about the guests. Maybe how you got into tech, how you got into DevOps and SRE.

Speaker B: Yeah, can definitely go into it. So prior to Thoris, I spent about nine years as a DevOps engineer and a lot of my focus was on deploying monitoring solutions for the defense tech world. And during this, you know, firsthand, I experienced a lot of my own frustrations with the tools that I had. I just found that it required a lot of my guesswork intuition, you know, for a lot of my job I would, you know, have to set up scaling policies for infrastructure, figure out uh, what are the thresholds for alerts and you know, I just thought we were very reactive with the way we were doing all of this and using like a lot of guesswork. So you know, with working with so much data it became easier to just throw money at the problem. So you. We would monitor and ingest anything and everything. We would over provision and have extra compute at hand. I also started to see this industry change where costs started becoming very important, especially with meeting with the rise of GPUs. We had to start being more mindful of cost but also make sure that performance and reliability isn't impacted. So I thought to myself, you know, there has to be a better, more proactive, less firefighting way to make better data driven decisions about how to scale and measure infrastructure without a trade off between, you know, being more performant and reliable and being more cost efficient. So I teamed up with my co founder and we started Doris and the whole idea was let's develop this reliability first machine learning driven platform to empower engineers to run their infrastructure more efficiently, eliminate waste, better safeguard and be able to uh, plan for growth as their environments evolve and prevent downtime.

Speaker A: Love it, love it. So I guess there's a little more to your co founder. Uh, not only is it your sister I think, but I believe it's your twin sister. So I have three older sisters. I can't imagine going in business with them. If they were a twin I would think it'd be even difficult. Like why that as your co founder just happened by chance or is this something you've always planned?

Speaker B: Not anything we've ever planned. It's actually pretty funny, Jo. We ended up kind of working in similar fields kind of accidentally. She was a site reliability engineer. She did work um, also in the defense tech world, but she also transitioned more towards commercial. So you know, it's funny, dinner table conversation would become about work And a lot of the problems that we were seeing, you know, we venting to your family is a very common thing to do during, uh, family dinner. And uh, one thing we started to kind of see is, you know, like we're seeing a lot of the same frustrations. We're both very much engineers, never really expected to become founders, never thought we would. But we realized that, you know, her and I have very, uh, complementary skills and we have a lot of similar, uh, experience, but kind of in different domains. Mine was more towards monitoring, hers was more towards the infrastructure side. And she also has a little bit of an ML background as well. And we decided, why not do this, right? Nothing like this exists. Let's build the tool that we wish we had.

Speaker A: Love it. Love it. I also would be scared. I mean, I got laid off and I just started my own thing. I didn't have a good gig and I just started, you know, my own business. What drove you to that point though? Was it like an aha moment? I know you mentioned a bunch of different issues you had at work, but were there no toolings that existed that addressed, you know, all these issues that you had?

Speaker B: Yeah, no, it's a really good point. I think it was naturally with like the pain points in the way that the industry was going. You know, one thing again, I think with the frustration of just firefighting always having to be on call, that becomes very tiresome. A lot of weight gets put on the reliability engineer. You know, we, it's very high stakes and you know, once, you know, outages happen, we're digging through our monitoring tools, we're trying to figure out the cause. You know, we're realizing that a lot of the issues we're finding maybe uh, there was a more proactive way to kind of get ahead of it naturally. As SREs, you know, we love to find things to automate. You know, a couple other things contributed. I think one was the rise of Kubernetes also that ended at being our first, you know, wedge focus. Kubernetes is a fantastic technology, um, but Joe, I'm sure, you know, it's very complicated in its own way and ah, you have to really enable it to kind of be more proactive. So with the way like Kubernetes works, right, is you have your real time scaler, it's very reactive and often, you know, you're using, you're kind of eyeballing your metrics and your data from your observability tools and you're reacting and you know, as engineers, uh, ilo and I, my co founder, we were updating our scaling policies like very manually. So when, you know, we would get hit with a traffic surge, we would go in and quickly update the policy. And the problem with that is, you know, the applications, a lot of times they have long startup times, so there would be latency already. The, uh, other problem too is, especially if you're starting to use GPUs or more compute intensive servers, you have to be able to go ahead and quickly get that compute and oftentimes, you know, try to get the gpu, wait for it to initialize, wait for your applications to load up. It's just too reactive. And uh, we realized that, you know, I think with the rise of AI, even though we're using more traditional machine learning for our tech, there is now more of an appetite for engineers to use AI in general. So we're starting to, you know, we very much believe use the right AI for the right use case. And that was one thing that kind of, I think helped with this. Okay, now's the right time. AI adoption is, is at its rise. The problem is getting worse. You know, data is becoming more noisy, environments are becoming more compute intensive. I think now is the prime to do it and who else is going to build it if we don't?

Speaker A: Right, absolutely. So you made a good point, the rise of AI. But like you mentioned, you're using machine learning and I think does that kind of not taint, but give people a wrong impression? Oh, they're just jumping on the AI bandwagon when it sounds like this is a great problem that machine learning was made for almost.

Speaker B: It sounds like, yeah, no, you're absolutely right. So one thing we kind of set out to do with Doris was, you know, have this reliability first platform that uses the right AI for the right use case. Right. And the way we kind of envision is we're starting with machine learning because it's a safe, easy adoption for infrastructure engineers to use AI, because this is very new for many of them. And we realize that, you know, there's different types of AI, there's LLMs, um, you know, you have traditional machine learning. And I think for the right use case, it just makes sense to figure out what type of AI to use. For example, there's a lot of buzz right now with using large language models for part of the SRE workflow. I think it's very good and useful for postmortems, going through logs, natural language, but for the type of, um, scaling decisions that we're focusing on with Thoris. It's more numerical statistical data and it's actually better to use more lightweight performant machine learning models rather than an LLM. Why use an airplane when you can drive a car, right, for a shorter distance, especially if it's cheaper and very efficient.

Speaker A: Weird question, why the name Thoris? I know Thor is the God of thunder. I don't know. Is Thoris or derivative of that?

Speaker B: Yeah, it actually is. Um, so you got it right. It is thunder, but the as at the end, uh, makes a plural and it's also the feminine version.

Speaker A: Oh, cool.

Speaker B: So yeah, the two female co founders, Thoris and the whole idea that, you know, we're building this platform that helps give the right actionable insights, you know, is very game changing in the way that we manage infrastructure today.

Speaker A: Love it, love it. Now you've also gotten some investment as well. What's the picture given that resonates with you, uh, know, people that are looking to invest in companies, uh, did they know, like Kubernetes is a pain and therefore, you know, it's. This is a no brainer, like what was the hook that got people like, oh, I want to be a part of this.

Speaker B: Yeah. I think one big part of it, Joe, is the fact that Nilo and I have lived and breathed this on our own. Um, you know, we've worked firsthand for nine years in the space. We've dealt with the frustrations, we know the pain points. I think that's very appealing for investors. But also at the same time, we try not to introduce too much of our bias. We were very, very big when we raised our first round. We were talking to customers about their pain points. We were figuring out who is the ICP we want to go for. And you know, we started to realize that with our customers, they are a group that cares a lot about cost savings. You know, they don't want to overspend anymore. They also care about reliability. They want to always make sure that their highly available performance is never impacted. They want that to get better anything. They, uh, they're sick of that trade off. They don't always want to be on call for issues that could have been intervention in the first place. And they're also really curious about AI. But we also find, Joe, that, you know, our investors, our customers, they're a very AI savvy group in the fact that they know that with the AI hype that's going on, yes, AI is very promising. But let's make sure that the folks are pioneering these solutions, are actually pioneering the right ML for the right use cases. So I think it was the, you know, we have the expertise, we're very educated about the types of customers and we do a lot of our research and also we're very big on, you know, using the right type of AI.

Speaker A: Absolutely. So as you mentioned, you both have deep expertise in sre, so when you speak with customers they could probably tell it's not bs, but you must have had a notion of what you want created. Was it, did it turn out to be different than what you got into, uh, live when talking to customers? Like hey, this is actually they're using it for things that we didn't think about or this is an issue that this really helps with that we weren't anticipating.

Speaker B: Yeah, no, that's a really good question. So when we first got into it actually what was really interesting is, you know, we knew we wanted to build for Kubernetes. What we decided to tackle right is to kind of give you an idea of what we do with Thoris is we built a like a predictive autoscaler that integrates into your infrastructure. And what it'll essentially do is it'll use multi signal intelligence. So it combines like infrastructure metrics with real world signals that tie into your infrastructure and better signal, you know, when you need to scale. And so what we do is, you know, we forecast what is going to be workload demand in five minutes, an hour, six hours, whatever. And we use that to basically scale your infrastructure in advance. And we'll also take a look at the real time to make sure you're covered. In the meantime, if it's an unprecedented spike that our models haven't seen M before. And so, you know, we knew that the industry didn't have this today. The way that scaling works today is very real time, it's very reactive. What that causes engineering teams to do, what we had to do was operate at a low level utilization. You know, most of our customers before they start using us, they're at 40 to 50% utilization. So that means you're leaving 50 to 60% of your servers unused just at hand, just in case. And so we discovered, you know, with this technology, if we can anticipate, we can bring them to 85% utilization or more. And we can also help make sure that we scale them before the usage, uh, spikes. We created this for that Black Friday use case. You would get the unexpected product launch or unexpected, you know, spike in traffic and you don't really quite know what uh, that's going to look like. And so Usually you over provision, but you can still get it wrong. The whole idea with Thoris, right, is you know, we can ingest whatever meaningful metrics and help predict and get ahead and make sure you have the right capacity so that, you know, you have just enough capacity to make sure you're highly available, but you're not massively over provisioning. So that we got right. I think what we didn't realize is how widely applicable the technology is. We got into it for customers that have very variable traffic patterns. But we started also getting customers who are just massively adopting kubernetes. They want to, you know, be able to manage that with less overhead, less complexity. But also, you know, sometimes they just want to stop eyeballing the dashboards and figuring out how to scale. They want to have better, you know, developer velocity and push out code and not have to worry as much about that infrastructure piece of scaling.

Speaker A: Very cool. So I started my career as a performance testing engineer. So I have to put on a load and uh, you know, staging to try to represent what I think was going to happen in production was never the same and was always difficult. It sounds like this could almost fill that gap. So does it listen to your traffic then and tell you, hey, with the forecasting you should probably anticipate a load of XYZ in the month of blah blah, blah. Like how close does it get to, you know, leading indicators without you having to jump in? I don't know if that makes sense,

Speaker B: but yeah, no, it's a very good question. So how it works, right, is it integrates directly into your kubernetes metrics. So you know, it installs in minutes, uh, via like a single helm chart command. And so we run air gap today. So like no data comes out from your environment, which is uh, really fantastic, especially for our customers in more regulated spaces. But we basically will ingest with whatever metrics you have as long as you have data in your environment that is helpful for determining, you know, when to scale, whether it's in any of the major observability tools somewhere in your cloud environment. We can basically grab that data, ingest it, train, uh, our models to forecast what is the, the uh, pattern going to be in, you know, an hour, six hours, whatever that's going to be. And so we'll give you the recommendation, um, for right sizing your pods, but also for scaling so you can, you know, run in recommendation mode. But typically most of our customers will turn us in autonomous mode or which means we will go ahead and take our metrics our predictions, we'll take a look at the real time as well and we'll just start scaling your infrastructure as a result. And what that does is actually very powerful. Right. One, it helps as you're pushing your code through production, you know, a lot of times you have very unknown situations. If you think of like the crowdstrike outage, right, where someone pushed out a change and they weren't, uh, sure, you know, how it was going to affect production, it led to a big outage. What a situation like that could actually be prevented. Right. Because we can start simulating a lot of those changes in the lower environment and kind of figure out, you know, what is going to be the impact and make sure that you're fully covered against the unexpected unknown as you're pushing to production and running in production.

Speaker A: Nice. So here's a dumb use case, very small. I run webinars and we're doing a new system. And sometimes, uh, when you run a webinar, you get like 200 concurrent users trying to log in at the same time. Using this technology, would you be able to beforehand say, hey, um, I'm having a webinar on this date, I anticipate this. So rather than have it guess, let it know, hey, this is something you should be looking at.

Speaker B: Yeah, that's a very good question. So. Exactly. And that's the premise of it, right? You have a webinar in, um, this example, and you kind of have an idea of how many users you're going to have, but not really. Right. You could be completely off. So typically how this works without a solution like Doris, right. Is you just guess and you, you hope for the best and you think, I'm hopefully going to have this many users and let's just provision it that way. And as you see more users logging on, you quickly make changes to be able to scale. But the problem is, is if you get it wrong and you have to merely, uh, scale your webinar, users are probably having issues getting on there and they're experiencing latency. Oh, with Thoris, you know, since we're integrating with more of the metrics that are meaningful for determining how much traffic you're going to get to hit that webinar, we actually will anticipate it in advance. So the whole idea is all your users can get on there no matter what the scale ends up being. We're reacting very quickly to manage that scaling. Uh, but you're also not overspending just in case, and keeping that extra compute very cool.

Speaker A: And obviously it helps does it uh, shuts down resources that it uh, knows it's not in use so you don't have to always look, I guess or have like set times that ah, you, you guess like during nighttime it's not going to be a peak when you don't necessarily know this automatically does that for you.

Speaker B: Yep, that's absolutely correct. So when you deploy us, we start learning the workload pattern. We'll take care of like that scaling decision. And so you know, we'll scale you up, you know, increase your pod sizes, uh, when you know where your expected to get more traffic in advance. But um, we'll also scale you down and shut everything down, uh, when during those periods of low traffic. But think of Thoris as you know, always kind of having that eye out so that you can sleep soundly just in case you get more traffic at an unexpected hour. We're also reacting very quickly and learning what those patterns are as they evolve over time.

Speaker A: Nice. Uh, so you didn't mention Black Friday. It always amazes me how many times huge companies get this wrong and it's not too far away. So definitely using a solution like this will definitely help them out. Do you find that that's an easy sell, that example when you go to a company, hey, Black Friday's coming. Or do people not get it? And it's still a hard education you need to do.

Speaker B: So the Black Friday use case, I like saying it Joe, because it's a uh, very applicable, it's a very understandable use case because people, everyone pretty much knows what Black Friday is. So I think it really does help demonstrate. Now I use it as a starting point because you know, not everyone is affected by Black Friday. But the whole, it does help paint the analogy of how that Black Friday use cases and this is one thing I learned right from us working with Doris and you know, continue to mature the platform and serve more customers. That Black Friday use case ends up lending itself to any unexpected event. And if you think about it like any unexpected uh, what is an unexpected event, right? It's typically as your environments change, as your code, more code gets pushed out, your applications get heavier, your traffic gets heavier, you do product launches. All of that are kind of similar to the Black Friday use case in the sense that on uh, a smaller or different scale, you don't know what your traffic patterns are going to be, you don't know what your scaling should look like. So it's the whole idea of whenever you have uh, your environments changing in some type of unknown pattern, you know, Having Thoris be able to illustrates that Thoris is able to better insulate, uh, and take care of that scaling efficiently, um, but more reliably so you're covered no matter what.

Speaker A: So this may not be a case. Can it be used for security? Say someone's trying to take down your website, will it notify you, hey, you're getting some really odd behavior here. Based on a forecast, you might want your security team to jump in or something like that?

Speaker B: Yeah, so, and that is one piece too. We believe in helping to uh, take care of the pieces of automating infrastructure and scaling that engineers don't have to manage, but also giving engineers the right alerts about what they need to know about so they can go ahead and take action and be informed. So that is exactly how it works. Right. We'll take care of the scaling, but we'll also kind of let you know of those anomalies and weird types of issues in your environment that hey, you should take a look at and this is where it's going on and this is what else it might be affecting. But yeah, that's a great question. That's exactly a very realistic use case actually that we help with.

Speaker A: Very cool. I also note, sometimes companies are really cost conscious. Does this give you like, hey, you saved X amount of dollars using Tharis, uh, this month or anything like that?

Speaker B: Yeah. So when customers, uh, deploy us, we made sure that we baked the cost savings and ROI upfront. So, um, we're constantly telling customers, you know, how much they're saving using Thoris, how much they could save using Thoris. We're also giving them more of an idea of, you know, as they continue to grow, what, you know, help them be able to plan for the capacity better. We're very, very big on, even though we take a reliability first, um, approach, you know, we believe that reliability is the number one KPI. We know that cost is very important. So that is something that we make sure is very clear in our product.

Speaker A: Who's the perfect, uh, use case for this product? Is it, say someone has a greenfield application or like, do you have a certain, like you're probably better off if you're already in production because then we can work off real data and historic data or does it matter?

Speaker B: So we have customers across like a whole broad of industries, right? So we have E Commerce, we have B2B, we have cybersecurity, we have ad tech, health tech. We find that, you know, as long as you have data in the cloud, you're using kubernetes. You know, we can help scale your applications better. We do see a whole range of use cases. There's a lot of migrations today. We actually help really well with that also because if you're migrating to Kubernetes or migrating to different cloud platforms, it really does help to have a platform that helps take a look at, you know, your evolving traffic patterns and workload needs and help, you know, inform you with the right insights to be able to take those decisions. So very open ended, right? We can support any use case, any, anyone that's running these applications in the cloud. But it is interesting that we are seeing a lot of migrations and we're also able to help with that as well.

Speaker A: Very cool. So besides migration, I know a lot of times companies are doing like a multi cloud type of environment. Does this work with that type of uh, setup?

Speaker B: It does, yeah. In fact many of our customers are running like these larger environments. They're multi cloud, they're hybrid. And I think that is part of the reason they need a solution like this because there's just so much complexity, there's so much overhead, there's different engineering teams involved and they really need a shared tool to be able to help efficiently manage all of this.

Speaker A: You know, engineers tend to be somewhat skeptical in my experience how many feel confident in using that autoscale feature without a human kind of in the loop.

Speaker B: So it is interesting, right, we do take a human in the loop approach, right. So we've made sure we champion like a very safe adoption for low risk adoption for Doris. You know, when you deploy us, you can run us in recommendation mode where we just give you recommendations. That's something that we built in the product. But we've actually found that all of our customers and anyone that tries this out switches to run us autonomous because they immediately see, you know, everything up front about like the scaling decisions and you know, how much potential and savings. So we, you know, uh, that was one thing I was wondering when we were going into this, right, like we know the appetite is there, how quickly is this going to happen? And I was very pleasantly surprised to see that not only are we being put in production environments where you know, we are more useful because there's more variable traffic patterns, but we're getting put across all these different environments and engineers are actually turning autonomous. And um, I think part of it is there is trust with the platform that we're not just a cost savings tool, right? We, unlike uh, many of the other tools out there that just scale based on cost savings what's the cheapest decision we're taking more of that reliability first engineering approach by making sure that we're making smarter scaling decisions that are more proactive, that makes you cover against unexpected events, but also allow you to have that higher utilization. And I think that reliability first method is what enables so much trust in the platform.

Speaker A: Uh, and speaking like a lot of SREs, uh, really want to be hands on. I think a lot of solutions are kind of like black box. It's like, ooh, we're doing this magic here. You can't look at it and sometimes people just want to look at it or be able to tweak it. Is that even an option or something that you think about?

Speaker B: Yeah. I love that you use the term black box because uh, we always say this especially when we're talking to customers anytime, especially with AI tools, right? Like there's always questions about what type of AI are you using. You know, just different questions about is there like a, a black box to how does all this works? I'm a big believer then if anyone ever tells you that it's a black box and we can't tell you anything, it's completely proprietary, it just takes care of everything behind the scenes, don't worry about it. That's a red flag, right? I think things should be more visible. And that's one thing we've gotten really good with our customers, with balancing in our platform. It's, there's a balance between taking care of the right tasks like scaling and uh, automating those, you know, that low hanging fruit that is very critical but at the same time giving engineers the right insights about what to look at and what to worry about. And I think striking that balance is very, very important, especially with the human in the loop, uh, safe AI adoption that is needed for these tools today.

Speaker A: Nice. So I guess another question I would have is a lot of companies I speak to have like kind of cultural issues or things in place in order to be successful with the uh, like X technology. Does a company need to have some basics in place in order to be successful with their product? Like do they need open telemetry? Does that help? Like anything like that?

Speaker B: So the way we built this is we've made it pretty easy to grab any data from its anywhere. So we, we support all the major observability tools. Like we can, you know, we have integrations with all of them. We can grab data from there. Um, but also even if you have data that's important, that's not in one of the Tools, we have a way of ingesting it. So I think right now with how much the platform is matured, as long as you have data somewhere in your environment that you know is helpful for making these scaling decisions, uh, we can work off of that. So it is pretty open ended in terms of what we can support today we support Kubernetes. Of course we would love to go beyond that, but we do find that as long as you have the data, we can take care of it and help.

Speaker A: Cool, cool. Also, I know a lot of times, uh, SREs sometimes complain uh, about a lot of false positives. How do you avoid false positives or maybe even overscaling because uh, maybe a prediction error.

Speaker B: Yeah, it's a great question. This is another reason, by the way, that I'm. We're using traditional machine learning, which is really good for this type of thing, right? Statistical data that you know, is well suited for machine learning, which is very well suited for it. But it's funny, we, we sometimes get questions, right, like some folks for LLMs, right? There's a concern of hallucinations. Luckily that is not a concern here, right, because we're using more traditional machine learning. But of course, you know, there's always concerns about other types of issues like the false positives with the way that we baked the product. You can see how well our predictions perform and I think that visibility helps really greatly. And again, these models are constantly training, learning your patterns, your workload needs. Um, and I think that's what's also really helping them to kind of get better, especially as they see more use cases.

Speaker A: Nice. Anything on your roadmap that you're excited about or that you could share?

Speaker B: Yeah, absolutely. So, you know, at its core today as a wedge product, you know, we're helping engineering teams scale smarter. But our vision actually goes beyond auto scaling. You know, we want to help engineers not only handle, uh, unpredictable traffic with ease, better plan for pushing out changes to production, making sure they're always prepared for like these big moments, right, like product launches, viral growth, Black Friday, all the these different types of surges, but without scrambling, being reactive and overspending. So ultimately the goal is to give developers and engineers the confidence and the data driven decisions to move faster without breaking things. Especially in today's day and age, you know, where you have different tools that help you to write code quicker, you know, test better. We want to make sure that we're that layer that helps you to helps engineers give them the right insights and automation to help, you know, run their infrastructure. More efficiently so that, you know, as these applications are pushed out quicker, they're able to run more safely in production. And so, you know, you're better balancing performance, costs, reliability. And I think to do that, you know, that goes beyond auto scaling. Uh, it's really essentially, you know, building different features and products that help utilize all the data that we have today and give engineers those right insights so that they're always involved, human in the loop to be able to run more efficiently and be more highly available.

Speaker A: All right, so you were a hands on SRE for many years. Is there anything you thought, you know, you wish you had this product when you were like hands on, when you got a call and you're like, oh, if I only had this insight before, it would have solved this issue beforehand?

Speaker B: Yeah, absolutely. You know, I think, um, having been on call, there's so many incidents I can think of in the past where, you know, it was just such a wild goose chase to figure out, you know, what were, uh, the, what was the originating issue, you know, what were the symptoms. And I think, I believe had I had a tool like Thoris, I would have been able to take a look at the right insights and I would have been alerted into like, what were the anomalous type of patterns that, you know, were introduced by, you know, any changes made in the environment. You know, oftentimes, for example, I think once there was a incident where a developer pushed out a image that, you know, started consuming more memory and that affected scaling. I'm going to have had some other downstream impacts and it took a while to kind of figure that out. But had I had a tool like Doris, not only would it have, you know, scaled to make sure that it would have, you know, taken care of that so there was no performance impacts, but also would have alerted me that, hey, this workload is affected. There was a change that was introduced that did this. It just would have been a lot quicker. And meantime, the response is a very important metric. It's very critical that we cut that down in the SRE space and that would have been much lower, I think with a tool like Thoris.

Speaker A: All right, Jennifer, before we go, is there one piece of actual advice you can give to someone to help them with their DevOps SRE efforts and what's the best way to find contact you or learn more about Thoris?

Speaker B: Yeah, absolutely. One thing I always tell folks is, uh, with all the AI out there and all the demanding work that SREs have, figure out the right workflows and ways that you can use AI and embrace it. You know, I think there's a lot of opportunities with the work that we do and so, you know, find the right AI for the right use case. As for finding me, uh, feel free to check out Thoris AI. That's our website. You'll find our docs. You can also request a demo or even try out our product if you would like and see how it works firsthand. You can connect with myself also through, uh, LinkedIn. My LinkedIn is, well, it's LinkedIn, uh.com in Jennifer Vermani. Feel free to add me would love to connect.

Speaker A: We'll have links for all this awesomeness down below and for links of everything of value we covered in this DevOps Toolchain show, head on over to testguild.com p199 so that's it for this episode of the DevOps Toolchain Show. I'm, um, Joe. My mission is to help you succeed creating end to end full stack DevOps toolchain awesomeness. As always, test everything and keep the good. Cheers. Hey, thank you for tuning in. It's incredible to connect with close to 400,000 followers across all our platforms and over 40,000 email subscribers who are at the forefront of automation, testing and DevOps. If you haven't yet, join our vibrant community at, uh, Test Guild, where you become part of our elite circle driving innovation in software testing and automation. And if you're a tool provider or have a service looking to empower our guild with solutions that elevate skills and tackle real world challenges, we're excited to Collaborate visit test guild.info to explore how we can create transformative experiences. Together, let's push the boundaries of what we can achieve O the Tescill Automation Testing Podcast with lutes and lyres, the bards began their song A tune of knowledge, A um, melody of code through the air Ah it spread like wildfire through the land Guiding tester showing us secrets to behold.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Less about Models; More about ArchitecturePractical AI · on Machine Learning85 / 100
  • Microsoft Fabric: The Platform That Turns Data into Competitive AdvantageLeading IT - APAC Insights · on Machine Learning85 / 100
  • How Datadog Scaled Engineering Without Burning OutThe CTO Podcast with Fexingo · on Kubernetes82 / 100
  • Navigating AI Risks with Trevor Horwitz from TrustNetB2B Automation Spotlight · on Machine Learning79 / 100
  • The Robot Is Waiting on Your Data.AI Proving Ground Podcast · on Machine Learning78 / 100
  • Wanting to Adopt AI Is Not a Problem But Lack of AI Fluency in Commercial Teams IsRevenue Hub · on Machine Learning77 / 100

More from TestGuild Devops Toolchain Podcast

All episodes →
  • Developer-First DAST: Fix Security Issues Before They Reach Production with Gadi Bashvitz70 / 100
  • A Practical AI Guide for Business Leaders with Brad Groux
  • Why AI + DevSecOps Is the Future of Software Security With Patrick J. Quilter Jr
  • GraphQL in the Age of AI Agents - Insights from Apollo's CEO Matt DeBergalis
  • Are AI Agents Replacing Contract Testing? DevOps Insights from Matt Fellows
Explore the best B2B Engineering & DevTools podcasts →
All TestGuild Devops Toolchain Podcast episodes →