
OpenObservability Talks · 2026-06-27 · 1h 2m
Key moments - from our scoring
Substance score
49 / 100
Five dimensions, 20 points each
OpenTelemetry's breadth and flexibility make adoption challenging for many organizations, especially when multiple teams implement it independently without a cohesive strategy. Dan Gomez, author of Practical OpenTelemetry and creator of the Blueprints initiative, joins host Dotan Horvitz to unveil how blueprints address this adoption gap. Unlike reference implementations - which are point-in-time snapshots of how specific companies like Adobe, Mastodon, and Skyscanner adopted OTel - blueprints operate at a higher level. They identify common challenges in specific environments (infrastructure-led, managed telemetry platforms, Kubernetes observability) and provide structured guidance using a framework inspired by Good Strategy, Bad Strategy: listing problems, offering design principles, and recommending implementation steps. The conversation emphasizes how blueprints balance OTel's intentional flexibility with practical recommendations, why the first blueprint targets non-Kubernetes environments despite CNCF focus, and how blueprints serve as organizational anchors for cross-team conversations about observability strategy. Platform engineering teams, infrastructure architects, and application teams struggling with OTel implementation will find value in this structured approach to reducing accidental complexity.
Blueprints are maintainable, structured guidance for common OTel adoption scenarios backed by production experience, while reference implementations are point-in-time snapshots showing how specific companies like Adobe, Mastodon, and Skyscanner adopted OTel at a particular moment - typically not maintained beyond initial publication.
OTel has essential complexity from its breadth (SDKs, collectors, semantic conventions, deployment models, instrumentation quality) and accidental complexity from lack of organization-wide adoption strategy - when teams implement independently without coordination, they create unnecessary additional complexity that can be avoided.
Blueprints use Richard Rumelt's Good Strategy Bad Strategy framework, which structures guidance into three blocks: challenges to solve, guidelines to address them, and implementation actions - ensuring every recommendation links back to solving actual problems in specific environments.
Non-Kubernetes environments (VMs, bare metal, on-premises) lack the rich tooling available in Kubernetes ecosystems (Helm charts, operators, declarative config) and face distinct infrastructure-led monitoring challenges, making them particularly underserved and the priority for initial blueprint guidance.
Blueprints can be adopted in pieces by individual teams (like infrastructure teams deploying collectors), and they give teams ammunition to advocate for broader changes by citing OTel's published recommendations and documented benefits to drive cross-team coordination.
Our reviewer’s read on each dimension, with quotes from the episode.
There are some genuinely useful technical insights buried here - like the leader extension solving the cluster receiver single-replica problem and the centralization-without-extensibility antipattern - but they're surrounded by long organizational platitudes, repetitive framing, and mutual back-patting that dilutes the per-minute yield considerably.
with the leader extension what you can have now is a, um, just deploy a daemon set and then each but one replica in that demon set will be the one monitoring the cluster. If that dies, then another one picks up
centralization alone can uh, be really bad. As in like if you just give people uh, uh, a way to deploy OpenTelemetry but you don't allow them to extend it in an easy way, then that's as bad as not having centralization
Applying Rumelt's Good Strategy/Bad Strategy framework to OTel deployment architecture is a mildly fresh angle, and the daemon-set crowding observation is concrete, but most of the content covers well-trodden platform engineering ground - golden paths, centers of excellence, bottom-up vs top-down adoption - without genuinely counterintuitive or first-principles arguments.
we chose that framework of um, good strategy, bad strategy... basically you should structure your thinking in three main blocks. One is the challenges that you're trying to solve, the main guidelines to solve those challenges and then the actions to implement those guidelines
you end up basically uh, with a node. When the moment of the node spins up, you just have demon says in it and there's no user workloads that can run on that node
Dan Gomez is a credible practitioner: former OTel governance committee member, SIG End User maintainer, led OTel adoption at Skyscanner, authored 'Practical Open Telemetry', and works as an observability architect at New Relic - he has genuinely done the thing at scale across multiple organisations.
I was an end user myself not too long ago
that made me even happier because it happened after I left the company. Right. But all the. I guess that was part of, you know, it was like my baby in a way
The episode names real companies (Adobe, Mastodon, Skyscanner) and real contributors (Tiffany from Grafana Labs, Luke from Splunk), and calls out specific technical components like the cluster receiver's single-replica limitation and OTLP GRPC vs HTTP trade-offs. However, there are zero quantified outcomes, no adoption metrics, and most recommendations stay at the architectural-concept level.
we already have three. I think we're getting a thing. We have Adobe, uh, Mastodon and Skyscanner that have basically published the reference implementations for autel
this blueprint calls uh out specifically that the guidance in the blueprint is not aimed at organizations that may have a requirement for you know, uh, 100% completion on every single span
The host is clearly knowledgeable and frames the topic well, but repeatedly delivers extended monologues that crowd out guest answers, never pushes back on any claim, and the conversation functions more as a structured promotional unveiling than a probing interview.
So I'm as an external, I brought them to talk to each other and suddenly realizing so. But having the blueprints as like an anchor to, to um, create even the language of conversation within the organization across the different teams
And I also see there on the background the uh, Practical Open Telemetry. That's the book that you authored about Autel. So uh, definitely yet another testament of your uh, your credibility
Computed from the transcript - who did the talking, and the words that came up most.
OpenTelemetry (OTel) is now the industry standard for observability - but deploying it successfully at scale is still a major challenge. In this episode of OpenObservability Talks, we dive into the newly launched OpenTelemetry Blueprints initiative: a set of practical reference architectures designed to help Platform Engineering and SRE teams cut through the complexity. We explore the first official blueprints for Kubernetes and traditional infrastructure environments, covering: How to standardize OTel Collector deployments Best practices for instrumentation at scale Building self-service observability platforms What these blueprints mean for Platform Engineering and SRE teams Whether you're just starting your OpenTelemetry journey or trying to scale it across a large organization, this episode gives you a concrete, actionable framework to move faster and with more confidence. You can read the recap post: Subscribe for monthly conversations on open-source observability, OTel, open source and cloud-native monitoring.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign.
Speaker B: And welcome to another episode of Open Observability Talks. I'm your host Dotan Horvitz, and here at Open Observability Talks we talk about anything DevOps, observability and open Source. So as I always say, may the open source be with you. And uh, actually we are now concluding the sixth season of uh, Open Observability Talks. Really exciting. Six years. I can't even believe it myself that the time has passed. Uh, uh, so essentially it's an opportunity to thank you all for following us. And by the way, don't forget to rate us wherever you get your podcast. And uh, we'd love to hear your feedback, so feel free to reach out. Twitter, X LinkedIn, BlueSky, Mastodon. Whether to the show or directly to me. Love to hear more from you. And for the conclusion of the sixth episode, I chose a topic that I personally have high hopes for. Uh, you know, I'm a big fan of uh, OpenTelemetry Autel. I even came with a, with a shirt for the occasion. Uh, but admittedly many users find it difficult to get started. Uh, we've made significant improvements in documentation Getting Started guide, but there's still a high uh, entry bar. Uh, so can we learn from each other and get some community knowledge formalized at least for the common use cases? The new Blueprints initiative comes to address exactly that. At uh, Hotel Unplugged Europe earlier this year in Brussels. I had the chance to take part in a conference session about the blueprints and then followed up with uh, Dan Gomez, the creator of this initiative. And I even discussed it here, here on the episode after the Autel Unplugged event. So now I'm really happy to give the stage here at uh, Open Observability Talks to unveil the first blueprints. So and I have Dan, Dan Gomez with me to share all about that. So uh. Hey Dan.
Speaker A: Hello. Hello. How's it going?
Speaker B: Good, good. Great to have you there on the show. And uh, may the open source be with you as well.
Speaker A: May the open source be. And I brought my favorite Star wars like thing that I own, which is a Lego piece of mando. I'm going to put it back. But yeah, very happy to be here. Um, thanks for having me.
Speaker B: Yeah, yeah. And I also see there on the background the uh, Practical Open Telemetry. That's the book that you authored about Autel. So uh, definitely yet another testament of your uh, your credibility and your experience with hotel. So really glad to uh, to have you Here on the show, um, we've been uh, knowing each other for a lot of time. You also uh, uh, led me in my journey when I started uh, SIG there on the semconv semantic conventions for cicd. And you scored that on behalf of the technical committee uh, and uh, help us uh, make it uh, happen. I would just say you're actually by the way also a maintainer at Autel's SIG End user. That's where end users come to share their experience. So you probably hear, hear that often from users. What I just described the pain. So what were uh, maybe the recurrent pain points that end users bring up that later led to the blueprint.
Speaker A: Yeah, I think um, one of the, perhaps the main reason to start this initiative is something that um, I start to think about when I was an end user myself. Now I, I work for New Relic and basically I do speak to a lot of end users. But I was an end user myself not too long ago. And I was thinking about if you want to adopt OpenTelemetry effectively you need to think about the cross cutting aspects of it. Um, one of the things that I keep seeing from end users is that they think about otail or adopting otel. Some of them just think about the collector. Okay, I'll deploy the collector and you know that's me done. Uh, some of them will just basically think about it as like you know, ah, a replacement for maybe a legacy vendor agent that you drop in a box and instruments things. But I think what the thing that was missing here is that opentelemetry is so cross cutting that you need a strategy of how you're going to roll it out at an organization wide as in like you know, how does, how do all these things um, play together? How do all these components actually um, link to each other in a way that allows you to have that, you know, that high quality observability data. I think that's what I think was missing and something that we've seen from end users that they needed that end to end, you know, story. Right. Of how to do it.
Speaker B: Yeah. So you also used to be a member of Autel's group governing committee. So um, tell us maybe let me ask it bluntly. Why is hotel so complex?
Speaker A: Yeah. So and this is something that you know by the way as well. I forgot to mention that during the, you know, hotel has graduated. So we're all celebrating and this is
Speaker B: celebrated here on the show by the way on the last episode, the, the previous episode, big year, uh, celebration about the altar and also the alpha public alpha for, for the profile. So uh, yeah, celebrating all of that and kudos again to all the maintainers, contributors, uh, committee members and everyone else involved.
Speaker A: Yeah, it is a huge milestone. But we um. During the process of graduation and it took time. Right. Some of the part of that process is interviews with end users and that was something that was raised to not just to governance committee but also to maintainers. And the word, I think the word at the time the word blueprints was already um, so like given as feedback as in a week one blueprints. Right. Um, and then the reason is that hotel is complex for good reasons and sometimes for bad reasons. Right? There is, there is that sort of um, aspect of hotel being cross cutting, being you know, um, across many, many different environments and components. Like there is like that uh. I forgot the, the author but the paper was. There's no silver bullet. How there is like you know, in software that is, you know, accidental complexity and there is essential complexity and OTEL has a part of it which is the essential complexity of adopting you know, technology across your whole, um, your whole stack. And that goes from you know, how you configure your SDKs to all the settings I use. How you do it, is it like, you know, you're going to do it like programmatically? Are you going to use, are you using environment variables? Are you using declarative config? Um, do you need a collector as a sidecar or not? Or do you have a gateway of a collector? Um, what about semantic conventions? What about instrumentation quality? All these things that you know, go with OpenTelemetry? That's the, I think that's the complexity is in its breadth. And then there is the accidental part which is I guess the complexity of like if you don't have a strategy to simplify this for ultimately for your users internally in your company, if you're in charge of say you know, Eurotel adoption strategy, then you can find yourself in a, in a bit of a complicated mess if you allow everyone to do it their own way.
Speaker B: Right.
Speaker A: If you have like a large um, number of teams and then everyone does it in their own way, then you put yourself in a corner of like complexity that I think can be avoided.
Speaker B: Yeah, I think I see that very uh, frequently. Obviously you mentioned that you're an observability architect also. So you see probably lots of deployments in production. We should have mentioned that also on the entry. Uh, but essentially what I see as well is um, that many times there is no, you're talking about the ideal case where you have like the open telemetry adoption strategy in the organization and someone does that like top down and all of that. In many organizations the dynamics is very different and individual teams make the selection uh, of hotel for their individual needs, sometimes as uh, a replacement for an existing, I don't know, vendor uh, or vendor's agent or whatnot. And in some cases for other needs or suddenly moving from just logging to logging plus addition signals, um, uh, but then again because it springs up like bottoms up from the individual teams, each team uh, puts up their own practice if you will, uh, for how to do that. Sometimes even they don't even try to formalize it as a practice or try to think about it end to end. And there's also the uh, I guess the organizational perspective who owns it. So intuitively you'd say okay, platform engineering should be probably some, some the entity that has the overarching responsibility and the breadth of vision across the organization, across the different application teams on how to implement that and should give some sort of golden paths. Not just by the way autel in general in observability, um, but in many cases you actually see that there are individual teams, DevOps, apps, uh, that take this approach and, and where is the delimitation? As you mentioned, there's the piece about instrumentation that is closer to the code and there's the other piece about collection. Maybe different teams own these different pieces so we call it hotel together. But different pieces of hotel fall under different parts of the organization. And where is the handoff and where is the ownership? Not just both deployment instrumentation, uh, semantic conventions, the monitoring. And so there's many, many questions that if you don't look at it holistically, uh, there you start seeing like hiccups within the organization especially of course large organization that ah, you have lots of teams, lots of business units, lots uh, uh, lines of business, each one doing its own thing. Right?
Speaker A: Yeah, I think that's um, ideally and I think we're seeing this more and more with um, I guess this is not just open telemetry. Right? You see that with AI now where like platform engineering is coming at uh, a stage where it sort of necessitates that like the adoption of AI tooler necessitates a strategy for platform engineering. The same with otel. I think when you're thinking about something that gets applied across like teams and code bases and environments and so on, you do need a little bit of uh, you know, you need to reduce complexity for the rest of your, of your organization. So in blueprints, in otail blueprints, we're not trying to dictate how you should organize your, you know, how you should do your company structure.
Speaker B: Right.
Speaker A: Uh, if you should have a platform engineering team or not. We do aim some of the advice of platform engineering teams, but you can still take the advice that we give in blueprints, in auto blueprints as like something that you can take in pieces. Right. And if you're like perhaps a, and then link between some of these different pieces. So if you're perhaps uh, you know, an infrastructure team and then you, you just think about deploying, you know, monitoring hosts, monitoring infrastructure, you might drop the collector in there and then you might call it a day. But then the blueprint that we, that we have right now is already published. We'll talk about that later. Basically will start to give you something to think about in terms of like, well, if you deploy a layer of collectors over here and you have a centralized pipeline, then you can get these benefits. So um, even though like one single team can take the advice, it might give them, um, some say ammunition to go to another team and say, well, why don't you provide this as like if you're the central team in charge of data pipelines, this is what OTEL recommends and this is the benefits of doing that. So maybe you can do that. Um, you start basically to have those conversations internally in your organization, to have that common strategy.
Speaker B: Yeah, actually it's funny, a lot of times when I came to talk to organizations, I was the one bringing to the table different teams that before did not even talk to each other, have not talked to each other about this and bringing them just because of asking, so who handles this and who handles that and bring people. So I'm as an external, I brought them to talk to each other and suddenly realizing so. But having the blueprints as like an anchor to, to um, create even the language of conversation within the organization across the different teams and starting mapping out uh, the responsibilities because you have this something to follow on. Uh, I think it's, it's very useful. And um, so, so you started touching on that what before going into the specific blueprints. But what like what should people expect from blueprints in general? Um, and I know that there's uh, also the initiative combines both blueprints and reference implementations or maybe adding a second question about the difference between that. Uh, to know what to expect from what.
Speaker A: Yeah, I'll start with that actually, if you go to the Otel, to the OpenElemeter IO, to the website, the official website. You see that we now have blueprints and reference implementations as a section, right? And then the first question is like, and we do explain there what the difference between blueprints and reference implementations are. Um, the way that we like to think about it is that reference implementations are something that we, well we, in fact we have already three. I think we're getting a thing. We have Adobe, uh, Mastodon and Skyscanner that have basically published the reference implementations for autel. Um, and those reference implementations are just a point in time. Basically this is the way that Adobe, mastodon, Skyscanner approached OpenTelemetry Adoption, ah, at a specific moment in time. So we don't intend those to, you know, we don't intend for any of these companies to come back and maintain that according to their reference implementation, right? If they want to do it, they're welcome to do so and come and like, you know, update it. But it's not a maintainable resource. It's just a snapshot in time. In fact, they were published as blog posts originally. Um, and you know, this is something that the, the developer experience, SIG and OpenTelemetry has been doing so far. I think, you know, this is something that they have been doing for a few months, speaking to those end users, putting together the stories, putting together the learnings and then in the future we now have a process basically for those reference implementations too. We just want to make it easier for end users to give us the reference implementation in a structured way. So that's what a reference implementation is. This uh, company adopted AUTEL and these components and this is how they did it and this is their benefits and what they learned from it. So that's part of it. And we're ready. As I said, we already have some. And then blueprints, what we're trying to do with blueprints is that operates at a higher level in a way. So we're trying to look at some of the common challenges of adopting opentelemetry in specific environments, specific scenarios or conditions. For example, uh, you could say what is the, you know, what are the main challenges that you're wanting to solve by adopting open telemetry in a infrastructure led, so like a non, uh, Kubernetes environment where you might have some Docker containers over here and then some, you know, uh, bare metal over here and you're interested in like monitoring processes, um, and that, that sort of thing. Um, that's one say archetype that we're trying, that we're you know, we're trying to find the common challenges and we're trying to put together a set of common design patterns to solve those challenges and then a set of implementation steps and then you know, as we will cover that in the future in you know, later. But like another could be. Well this is how you deploy a managed telemetry platform as in centrally managed where you have a team that is in charge of facilitating adoption of the configuration of the SDK and a centralized set of pipelines and so on. Another one could be Kubernetes observability. That's another one that's currently in progress where we're looking at challenges of monitoring a Kubernetes cluster and the workloads that are in that cluster in terms of um, the resource utilization and also some of uh, I guess um, common Kubernetes components like you might have Cert Manager or Core DNS and things like that. Everyone wants to know how to. That is almost like default in Kubernetes. Um, so those are the main ones but in terms of um, the structure of each of them and this is what, you know, what is a blue. How is a blueprint different than a, than a reference implementation? The structure of each of them is heavily influenced by a framework that was uh, put together by like that was popularized by um, Richard Rumelt and Good Strategy, Bad Strategy, which is a book that I would always recommend to anyone thinking about strategy. And we can go into more detail later but in general what it says is that uh, basically you should structure your thinking in three main blocks. One is the challenges that you're trying to solve, the main guidelines to solve those challenges and then the actions to implement those guidelines. So it's a well structured framework that then allows us to put these blueprints in a way that is easy to consume and it solves actual problems for end users. Right?
Speaker B: Yeah. So essentially. No, no, ah, I think it's essentially the difference between conceptual guidance uh, and real production stories that the conceptual guidance in blueprints and the production stories um, that show uh, specific company with their specific uh, needs, pains and the constraints and how they solved it. So uh, I think this is uh, the way that I at least try and direct people when they ask where should they go? And obviously it's not one or the other. They can read through uh, the reference implementations and look at the blueprints and take from each one the pieces that are relevant for them. And by the way I should say that credit's also for you on skyscanner1. You're one of the ones designing the adoption of Autel there. So, uh, your signature is there. Right.
Speaker A: That made me even happier because it happened after I left the company. Right. But all the. I guess that was part of, you know, it was like my baby in a way. So. And it's good that, you know, it was something else, someone else, someone else that I picked up that and you know, started sharing that. It's great to see, to see that as well. Um, but yeah, so I think the. Yeah, I think you basically, you mentioned that there. Which is a relation between both of them.
Speaker B: Right.
Speaker A: The. We do want the blueprints to be based in or backed by reference architectures. So we're not just picking this advice out of a hat. It's just we have seen this work in production, we have seen this work at scale. And the advice that we're given in these blueprints is not theoretical, it's practical as well.
Speaker B: I'm actually wondering also again, as an architect myself, uh, there's always. In systems in general and specifically in. In Autel, the most interesting, uh, one of the most interesting tensions, I guess, uh, is the flexibility versus opinionation. Right. Otel, as you mentioned before, is intentionally, uh, designed to support many architectures, many vendors, many deployment models. So I guess how do you create the recommended blueprints without turning them into rigid, um, structures? I'm just wondering how you balance these.
Speaker A: Yeah, so I think, um, that's one of the reasons why we chose that framework of um, good strategy, bad strategy. And the reason for it is that um, we may. There are many, many ways of doing hotel, uh, or many ways to deploy every single component of OpenTelemetry. Um, and the way that we're trying to scope the area or the scope of the problem is by listing the common problems to solve. And this is one of the things that the framework is aimed at is that any guidance or any action to implement that guidance needs to always link back to the problem you want to solve. Um, and in a particular environment that tends to be. There tends to be a preferred way of doing that. Now it's not always possible to, uh, you know, to do something in the most optimal way. So, um, I'll give you an example. Um, we know that there's a challenge and when people want to, you know, when people start to think about adopting open telemetry at the application level, which is configuring open telemetry across, you know, different languages, different teams, different you know, and it's all deployed. Let's say that it's all deploying Kubernetes for the sake of it. But you, you know, you deploy it that way. I guess, you know, one of the easiest way of doing that is using the OpenTelemetry operator for Kubernetes now. But it may not always be possible to deploy an operator for whatever reason. Right. So then we give different options in a blueprint to say, well, these are the main options we recommend this one if you can. But if for whatever reason you can't do that, well, this one may require more work which is, I don't know, creating your own base Docker images, for example, that will have your standards baked in. However, you know, if you can do the operator, it's just a lot easier to do it. Or you should use declarative config as in ikea move to configure your SDK. But it's not available in every language. Right. So it depends on what language you, you're working with. You may be able to use that or not. So we're trying to give some options in order of preference where we think that there are many, I guess, equally good ways of doing that. And we're also trying to make sure that whatever we recommend is stable in a way, as in like, this is one of the aspects that we, you know, there's no point in putting together a uh, blueprint for adoption if things are changing next month. Right. So, um, I guess, you know, we're trying to base it on like things that tend to be, you know, either stable or they're not stable. We may recommend it. But say, well, use caution here because this is not completely stable yet.
Speaker B: Yeah, yeah, for sure. And let's go straight, really the exciting thing, the first blueprint just came out this month and it was really exciting to see that. Uh, and the first blueprint essentially is you hinted to that uh, earlier about uh, looking at non Kubernetes environments actually. So although we are uh, both on the, all of us in the CNCF Cloud Native Computing foundation, sorry, uh, uh, thinking Kubernetes native, but actually a lot of the users out there still on traditional VMs, even maybe bare metal directly, um, on premise environment with all sorts of setups that they need to manage, uh, themselves. Uh, and the question arises, because Autel is part of the CNCF and Kubernetes is like the first thing that comes to mind. What happens with these folks, uh, do they have a good way of operating? So, uh, it was very interesting to see the choice and this is the first one that has come up. Uh, so take the stage, Tell us a bit more about this.
Speaker A: Yeah, so, um, the, I mean the reason why the first one was that there are three that we were focusing on, um, as part of the initial set of blueprints. And I mentioned the three earlier. But the reason why I was particularly excited about this, um, you know, non kubernetes. Non kubernetes 1. Because the reality is that if you're in kubernetes in hotel, things are a lot easier. As in like, you know, there's no, uh, there's no question about it. You have loads more tooling available to you either as like helm charts or operators or like things to, to control the config sprawl that you might, you know, end up with in hotel than you know, in a non kubernetes environment.
Speaker B: Right.
Speaker A: Um. So, um, yeah, so basically what that blueprint is trying to cover is some of these challenges that, that normally arise from. Some of them arise from not being in a Kubernetes environment. Not having an orchestration layer like um, you know, the. Not having a standard way to deploy agents for example, or to deploy uh, collectors right across uh, your estate, or having you know, completely separate like, or fragmented instrumentation approaches. Those are the challenges that we're trying to solve. And also like basically having these siloed pipelines. So basically, as I mentioned earlier, you might have, you know, the challenges that come with um, instrumenting or like deploying a set of collectors across your infrastructure in an unknown way and then um, not having a central pipeline for all that data that comes with the challenges of governance. Uh, if you want to enforce a certain standard or ensure that you have the visibility over what data is produced and what is produced. Um, those are the challenges that we're trying to solve with that blueprint. And some of our recommendations are around, for example, centrally managing that aging life cycle and using opamp. And I think you mentioned earlier that you had a guest in the show.
Speaker B: We had Andy Keller here on the show. Uh, he's uh, I guess one of the creators of OPAMP and the core maintainer. So he was here on the show just uh, a couple of episodes ago. So if you're looking for a deep dive on OPAMP and what's new and what's uh, the latest on opamp, do check out the episode. I'll share the link on the. On the show notes.
Speaker A: Yeah, that's a good example actually opamp of um, a Case where this specification of OPAMP is currently in beta. Right. I guess one would say is it stable? Is it not people are using in production and we know that. However we're mentioning that it's a standard that's evolving and we mentioned that in the blueprint. But we also give um, an idea of how one could work with OPAMP to help manage that configuration sprawl the way that you manage that aging life cycle or that collector lifecycle and then some I guess, you know we're not trying to reinvent the wheel in blueprints, we're trying to give, to link people to relevant guidance. So again you know, we're not trying to document how OPAM works. Right. That's not what a blueprint is about, it's about how you deploy it and then link it back to your, I don't know, to the way that you apply to the way that you should think about your configuration strategy. Right. Where you have normally one centralized bit of config from a centralized team perhaps, but then you allow other teams to extend that config for the particular use cases and then how you put all your strategy together.
Speaker B: Yeah and I think this just for the audience to understand that I think the complexity here that you're tackling with uh, the blueprint is really the challenge of uh, the consistent observability across this uh, heterogeneous infrastructure. So imagine these systems, we're having some legacy processes, containerized workloads, mixed infrastructure, VMs, bare metals containers, whatnot. Uh, and that's the thing, when you don't have kubernetes, you don't have this centralized orchestrator that abstracts all of that and gives you a unified set of APIs to fetch, uh, I don't know the statuses and to configure. And so this is the complexity. So this is why I loved uh, seeing this uh, as the first blueprint that comes up. Because really this, these are the places where architects flourish. Uh, you really get uh, the non trivial things that you need to plug that not all of them are the cutting edge and none of them are. You said not even YAML might not be supported or things that uh, the cloud native folks take for granted and you still need to address them unless you just declare no, AUTEL is not relevant for that. And I definitely don't agree with that. So I think uh, this blueprint brings Autel to these environments that people less natively maybe think about or maybe some of them even. And I heard that in discussions with users and customers and Whatnot, some of them just disqualified that because. No, no, but we're not Kubernetes, so we know that AUTEL is great but you know, we're still stuck with VMs and stuff like that, so we didn't even bother looking at that. So they, to begin with, they didn't even go down this path because of that. So having this blueprint, uh, I think and having it as one of the first ones I think is really uh, going to be helpful in showing others that are in, I don't have to call it legacy, but no other types of environments that are not Kubernetes, uh, that AUTEL is just as beneficial and uh, not that much of a heavy lifting to make it run effectively, uh, in a scalable fashion on these environments as well. So. And by the way, maybe we should also give uh, uh, kudos to uh, Lukaszka. Uh, I'm not sure how to pronounce the name correctly. Uh, but uh, that contributed uh, that and uh, helped lead that. Uh, you tell me if other credits. But really, uh, great to see.
Speaker A: Yeah, no, I think, yeah, I think the work that Luke has from um, Splunk and also um, uh, Tiffany from Grafana Labs both uh, have been, you know, instrumental to get these blueprints off the ground. Right. Uh, and we're also having folks like, you know, there's. I guess we're like a few of us like you know, Tiffany, Lucas and Alex, uh, that are sort of like the core team, I would say currently working on these blueprints. So um. Yeah, really appreciate their effort very much.
Speaker B: Yeah, yeah.
Speaker A: Uh, one thing that I would like to mention as well is that. And you know, I was thinking about that blueprint maybe like a call to action here to end users. If you listen into this and you think. And actually you know, I do already run hotel in bare metal and I have a, you know, uh, a way of handling this with OPAMP or without opamp. Um, and you want to contribute your, your reference implementation. That's something that would help us a lot because at the moment the three reference implementations that we have are all Kubernetes based. Right. So having one of our large, you know, end users or even large, you know, you have done it in production, you're running open telemetry in non kubernetes environments and you want to share how you're doing it, get in touch with us and then you know there is. Or you, you can go to the, to the blueprints and reference implementations website and we have a section where we have like how, how you can help, how you can contribute. One of them is just sharing your story let's say um, and then we'll have these blueprints being backed by actual
Speaker B: evidence and that goes back to the reference implementations. This is the synergy between reference implementations that we're an organization, although a specific organization with specific uh trade offs and considerations and requirements but they put it from A to Z in production so they can really share how they done that and then looking at the blueprint that generalizes that to the general case and keeping this cross pollination between the reference implementations and the blueprints. So yeah, call uh to action. Everyone wants to take part especially in this part that is the non kubernetes is less charted. Uh do touch base and again we'll have the links uh on the show, notes on uh how to get in touch and where the pages and the sig that handle that um and the second blueprint that's uh just hot off the press is about managed uh telemetry platforms for uh kubernetes workload. So again now going back to those who run on Kubernetes, uh and still there are lots of uh, uh complexities and challenges that you need to think about in these environments. So maybe even before the blueprint do you want to give us some examples of the challenges that you see and what are the like the emerging best practices and guidance that then led to uh, what you formalized in the blueprint?
Speaker A: Yeah, so I guess one of the main challenges that we're trying to solve um and I think some of these are a little bit overlapping with the ones in that ah, sort of like non kubernetes environment. But in general um, what we're seeing is like the inconsistent configuration at the application level. Right. Um we talked about that previously how you might have multiple teams uh approaching opentelemetry in different parts of the organization uh in a different way. So they may be in charge of configuring their own hotel SDK and their applications. And then without having a consistent way of doing that you end up with context silos for example. So if you don't configure your propagators to all understand the same standard for context propagation then you'd end up with even if you push that to the same backend you'd end up with uh broken traces. Right. Or disjointed traces and orphan spans. If you don't apply a consistent layer of resource attributes, for example you may not be able to understand where the telemetry is coming from what is, you know, what entity is producing that telemetry. Um, you may have multiple versions of SDKs that, across the environment. So that means that, well, if you are relying on a particular breaking change or a semantic convention, something that is changing in a particular SDK that makes it more difficult uh, for a centrally managed team to deal with it. And then probably the most important ones here is like, and this is where platform engineering comes in, doing all that for each team that creates cognitive load, that creates, you know, that reduces the velocity of the organization in general. If you have every single team to have to do the same thing over and over again, which is configuring the SDK, maybe tuning some things, maybe configuring it to your particular backend, um, or maybe depends on the backend that you use. You may need to use different settings for, I don't know, aggregation, temporality or the way that you, that you emit data in a particular way. So yeah, all these things just create, generally create bad or like lower quality in data as emitted by the applications. Um, that's one of the challenges.
Speaker B: And then I should also mention by the way the skill set, even people don't think about it. It's a lot to know about Autel and if you need to have the skill set in each and every team and maybe you have one champion and that champion left that team and suddenly the team is uh, left hanging with it with no one understanding how it works. And rather than having one specialized team, the center of uh, knowledge, center of excellence if you will, in one platform engineering team that spares the uh, app teams uh, from handling these parts and there's focus on their core business logic that they need to do so even that maintaining and there's a lot more in the rapid pace in which Autel project has been moving. Uh, you and I know that it's uh, been hard for the organizations to even just deploy the latest releases, um, and let alone learn exactly everything that's been released. Uh, that's I think a major additional major load that people don't account for when they run their math on the engineering efforts just keeping up to date on the project and the uh, specs and everything.
Speaker A: And even if you have a set of like, you know, sometimes, you know, that center of excellence approach works but then the way that you implement it can create those silos as well. So um, what I've seen happening with many end users is that they tend to have like when they, I mean they don't always have A centralized, let's say set of standards. Right. But when they do have these engineering standards, some of them just have them as documentation, right? So um, you just have them somewhere and you say well this is the guide that you use to configure the OTEL SDK. These are the settings that you use. The problem with that as you mentioned, is that that still creates cognitive load. That still creates like reduces the velocity because you provide the guidance and the standards and the documentation, but you don't actually provide the tooling to do it. And this is what, you know, some of the challenges that we're trying to solve with that blueprint as well is facilitating the adoption. So making that golden path that you're, you know, we're all as engineers, we're all very good and architects at uh, as you know saying, the golden path. But then you need to make that golden path easy to, to walk, right? As an easy to, to go through. And that's what the tooling that. There's a lot of tooling around that nowadays in hotel that can help you and that uh, in that journey. Maybe a few years ago when you know, if you think in 2020, 2021 hotel was a different place and then you had to do a lot of these things yourself and then build that tooling yourself. Nowadays there is a lot out there that can, that can help you in that. So we're trying to, yeah to publicize that as well.
Speaker B: With things like um, the injector you can create your own bundle of the auto collector with the right mix of uh, receivers and processors and uh, exporters. And uh, there's we mentioned OP amp for like scaling out the uh, and centralizing the configuration. So just mention more concretely what kinds of tooling that have been built in hotel and uh, it's a great addition. So maybe you can tell us. So now that we understand a bit of the types of challenges, uh let's dive a bit deeper into the managed telemetry platforms for kubernetes workloads. What does this uh, blueprint entails? This is the one that you wrote yourself. So uh, credits to you as being also the author for that. So tell us a bit more about that.
Speaker A: Yeah, so I think you know, that's one of the challenges uh, that we are trying to solve. But not the only one. There's others like having multiple collectors and multiple clusters perhaps and this config sprawl of having all these collectors in place, that's another one of the challenges uh, and rolling out changes to that Config, um, or pipelines that are not optimized for observability data requirements. And what I mean by this, and this could be a contentious topic, depends who you speak to. Um, so, and this is where we start to scope a blueprint, right? So this blueprint calls uh out specifically that the guidance in the blueprint is not aimed at organizations that may have a requirement for you know, uh, 100% completion on every single span and every single log record that is produced through the pipelines or you know, because they've got compliance or audit requirements that, you know that, that they necessitate that type of uh, uh, reliability. This blueprint is aimed at the more common ones that we see, which are, you know, having fast production of data flow and fast through pipelines and highly contextual data, um, and also lower resource utilization. So I think that's the, so trying to basically run a sort of lean platform, but at the same time that it allows you to govern that data, uh, that uh, telemetry production and then that governance aspect as well is another one of the challenges that we're trying to solve. Right? So how do you produce good quality telemetry? How do you ensure that, you know, you either, you know, you basically lower your carbon emissions or your, or your cost story or your storage costs or your telemetry costs in general, how to ensure that teams are using the right signals for the right purpose and so on. So um, again, you know, we're trying to cover some of these aspects in, in the blueprint and the, the main guidance in it, and I'll go into that now is basically we'll start with the, with the basics of how do you apply SDK config. And um, and then we're calling here for a centralization of the config of your applications but uh, in an extensible way. So centralization alone can uh, be really bad. As in like if you just give people uh, uh, a way to deploy OpenTelemetry but you don't allow them to extend it in an easy way, then that's as bad as not having centralization in my personal view. And so, uh, if we go into that recommendation, we start with the hotel operator, so with the onpet telemetry operator for kubernetes, um, to apply that config out of the box, right? So if you're deploying every time that there is a pod spun, uh, up in the cluster, uh, that emission webhook will kick in and then inject the instrumentation into that, into that container, right? Um, and if you can't do that, then start to basically use the um, um, like pre baked images for your containers that will have something, you know, configured under the hood to like by, by default. So like, you know, if you're, if everyone in your organization is using the same base Docker image and it may do other things than OTEL may do, I don't know, things like security, networking, whatever, um, you can hook into that um, base Docker image and then provide your tooling that way. Or even like some libraries. Right. There are certain languages that don't allow you to um, let's say you need to configure OTEL in a programmatic way. Um, and if you need to configure OTEL in a programmatic way instead of like telling your, the users, your internal users that will be the application owners to use OTEL um, in a certain way or to configure the SDK in a certain way, you can just give them a library that can just call and it does it for them. Right. So that's the UM recommendation and there is a bigger part here which is perhaps the, the guideline in general which is to have a clear ownership of who owns the configuration and who owns the telemetry that is produced by that application and then how we should, by following this model, you can still remain the owner of the base configuration while the application owners have the responsibility of ensuring that the data that they produce is high quality, that they, you know, they should only care about the API and then you shouldn't really care about the SDK unless they really need to. Right. That's the ultimate goal. With that, uh, baked in config is the application owners, they should ask the business context to their telemetry should ensure that it'll, you know, that it describes their system in an effective way. But they really shouldn't have to care about, you know, configuring different parts of the SDK or batch processors or metric readers and so on. If they want to though, they should also be able to do it, to extend it in a certain way. But by default you basically ensure that sort of like Bayes layer of telemetry, that minimal viable telemetry in a way that should be produced by every service.
Speaker B: I definitely relate to that because I think this is the challenge that I see out there. I think circling back to what we said at the beginning of the conversation is that the challenge is not just technical, it's also organizational. And the questions around the separation of responsibilities, uh, often arise, you know, what platform teams own uh, what's owned by the application teams and the limitation between all of that. So I think the fact that you took this uh, into consideration as part of the blueprints, to align uh, the blueprints with uh, modern platform engineering practices, I guess uh, to me is a very, very crucial part of the success of success for blueprints to really look at the organizational aspect, not just a technical one. And although we can't dictate organizations how they structure and if they even have a uh, central uh, platform team or uh, what's the line of responsibility or whatnot, I think it is important to have a statement around uh, the distinction between the infrastructure and telemetry pipelines versus the application like the, the more instrumentation and business telemetry, uh, and then really try to provide the guidance around these things. So I'm really happy that uh, there is also uh, addressing that in the blueprint.
Speaker A: Yeah, there's one big as well. I was just thinking of um, the blueprint itself. We want it to be easy to consume, right. It shouldn't take you an hour to read it. Um, however we have started to put in appendices that may go into a little bit more detail into certain areas. Um, but they don't deserve a full blueprint. And in this one for example, there's one that I've been asked many times about this myself. It's like should I use OTOP GRPC or should I use OTOP HTTP for my, you know, uh, protopath for my exporter. And I think the question is, and you know, how hotel as well changed the default from GRPC to HTTP and, and then yeah, so the question is like when should you use one or the other and completely depends on where you send the data to. What's your, you know, what's your architecture? So this is, has an appendix, this blueprint has an appendix on that that is going a little bit more in detail into when to use one or the other basically. Um, but yeah, I think the, the blueprint itself takes that approach from like the application level to like deploying a centralized gateway, um, in which you may a centralized gateway in your cluster. But how you may also need a global gateway if you're deploying, if you're doing tail sampling, for example, um, and you have services deployed across multiple clusters that call each other. And this is one of the things that. Well I also see as a common challenge as well.
Speaker B: Right.
Speaker A: People think, uh, people want to do tail sampling and then they start to think into tail sampling. But there's no single let's say plays that says okay, well you know, you need to configure your. We can talk about the tail sampler in the collector. Um, and that this blueprint is not aimed to, you know, to document how to use a tail sampler. That's already very well documented elsewhere. But what we're saying is like well you may need to deploy a gateway at ah, a global level. You may need to still keep your gateways at the cluster level but have a global gateway. You may want to start basically from the application level, enable, um, always sample there and do tail sampling on the other side. Or there are other ways that you may want to think about this in a way that you might just want to, you might think this is all very complicated and just go for probability sampling and call it a day. That's also an option. Um, but we're calling these options out uh, in the blueprint as well with the point of reducing or improving the quality of the traces that we store. The ones that actually describe the issues that we're looking for. Um, which goes back to the challenge that we're trying to solve in terms of cost and telemetry quality or even carbon emissions. Um, so yeah, so I think all these recommendations, they all need to link back to the problem that we're trying to solve.
Speaker B: Yeah, for sure. And I want to have uh, a bit of time also for the third one that is really uh, about to uh, come out also in a couple weeks. So maybe by the time that you listen to this episode, uh, it's already going to be out which is actually the Kubernetes observability itself. So up till now we talked about like workloads running on Kubernetes but obviously Kubernetes itself has its own challenges uh, for running it at scale and uh, we need to absorb it into the Kubernetes itself. Do you want to talk about the, this blueprint?
Speaker A: Yeah. So um, that is currently, it's a little bit earlier than that one. There's a, there's a draft at the moment that's being reviewed. Um, but I'm quite excited about that one because that will contain now what we think is going to be the, the say the recommended way to do open telemetry or Kubernetes monitoring. Yeah. So um, and the reason for it is that we're moving away from um, I guess the guidance for a long time had been you know, you have kubestay metrics and then you have node exporter and then you have your Prometheus receivers and then in the collector and then you know, you just query those, send that data, um, elsewhere. Then you might need the target allocator for that. So they're like, you know it by itself. That would have been quite a nice blueprint and how, you know how to approach uh, kubernetes observability. Now what we're seeing is there's two things that are happening at the moment. One is the stabilization of the Kubernetes semantic conventions, which I think is huge, right? That we are actually having now stable attributes for resource attributes for Kubernetes resources. And we also have the, the collector components themselves that are able to monitor Kubernetes directly. Right? The Kubelet, um, the Kubestats receiver, the host receiver. That's been there for a long time in terms of node metrics and the cluster, and the cluster metrics themselves. Um, how we approach this uh, as well is something that I'm not like 100% sure now if this is going to be the recommended path. But there is definitely the leader extension. I think that's what it's called, the leader extension that basically allows you to have, for example for the cluster receiver. There's one of the challenges that's been forever in there in terms of ksm, for example, had the same issue, same uh, issue with um, if you were to run the collector with the cluster receiver is that you can only run it in one replica, right? So normally the architecture that would normally come out of uh, monitoring Kubernetes used to be you have your daemon set with um, your collector maybe like doing your Kubelet metrics, doing your host metrics and then anything that you can gather at the host level and then you had a single replica deployment maybe for the cluster level metrics and anything related to any other CRDs, like any of his custom resources, like horizontal port autoscalers and whatnot. Um, or you would query KSM again. You might have one replica of KSM of uh, cubestet metrics, or you might shard cube step metrics. But in general that's the approach. Um, with the leader extension what you can have now is a, um, just deploy a daemon set and then each but one replica in that demon set will be the one monitoring the cluster. If that dies, then another one picks up, which is a very neat way of like basically saying no, the only thing you need to do is deploy a collector and um, deploy a uh, demon set.
Speaker B: Demonstrate.
Speaker A: And that's the um. Yeah, and that's the way that we're. I believe that's the way that we are approaching it. If that's not the way that I apologize in advance, but I think that
Speaker B: is the way you shouldn't be apologizing. Actually, this is an invitation for all the kubernetes folks there that run Kubernetes, especially in scale, multi clusters and all of that, to chime in on the discussion. We'd love to hear your feedback on that and uh, get more feedback, opinion, more practices that people have already put in place uh, to learn to uh, decide what are the, the best practices around this.
Speaker A: So uh, yeah, I think a demon set, I mean I've worked in a Kubernetes, uh, environment for a long, long time and uh, demon sets are always like a problem or end up being a problem at scale because uh, especially if you have a very sort of like, if you have a heterogeneous cluster where like, you know, you may have like multiple different types of node sizes.
Speaker B: Uh,
Speaker A: and I've seen this happening sometimes you may have so many daemon sets, one for this EBPF agent, one for the collector, one for flimbit, one for this, one for the other, one for maybe another thing. You end up basically uh, with a node. When the moment of the node spins up, you just have demon says in it and there's no user workloads that can run on that node. So I think this is one of the things that we think can help is centralized or consolidating that into, you know, maybe, maybe in the future. The only thing you need to do is deploy one demonstrat for your observability and that's the collector and it does it all in that node. Right. Um, it just becomes easier to manage. But yeah, so I think, to your point, um, yeah, I think this is something that I am really excited to see. What guidance comes from the community again. You know, these blueprints are not something that we come together as. Like it's not one person that does it. We, we basically reach out to specific sigs. We reach out to the community is a common pool of knowledge that uh, we're pulling from to say is this the, the way that we think about it as a project? Are we, you know, are we. This is, this is the way that we think about as a best practice to deploy hotels or projects. So the more feedback that we have, the better. I think. Uh, in fact, the blueprint that I've been authoring I think is closer to like, I think it was 100 comments or something like that on that PR, which is great because it means that conversation happens, we get more points of view. And uh, it's great to see that level of engagement.
Speaker B: By the way, even in Hotel Unplugged. I think the roundtable that we had around the blueprints, which was much earlier in the initiatives that I didn't see, still seeing how many were first of all, uh, excited to attend that because there are lots of other interesting and uh, conference sessions taking place at the same slot, uh, and the level of discussion that uh, took place. And I was uh, before that I knew that it's happening. But I have to say attending that and hearing the discussion, uh, made me much more enthusiastic about this initiative and made me follow it much closer. And uh, so it was really nice to see people passionate about that and really seeing that as an answer to this complexity problem. I, uh, don't want, uh, leave before we reach the end of the time to talk about looking forward, looking ahead. I'm curious to hear what's next for the Blueprints initiative and also what, what does success look like, uh, as far as you're concerned, let's say a year, two years from now.
Speaker A: Yeah. So I think the way we see this evolving, and that's called out in one of the blog posts that we published, is connection points like blueprints will overlap, blueprints will link to each other. There is no single blueprint that will solve all your problems. In fact, we're calling out for extension points. And the blueprints that we're writing at the moment, as I said previously, compliance or audit login not covered in this one. But we already have an issue open to create a blueprint on that or um, collector deployments in highly regulated environments where you need HIPAA or fips compliance and blah, blah, blah. Another one that we see evolving in that way. So blueprints start to link between each other. Um, and then the ones that, I mean we've already got like quite a few in the pipeline which uh, is great to see. We have like, you know, that one for compliance and audit data data, uh, governance and privacy. About one about like semantic conversion, semantic convention registries and governance of uh, like Weaver, uh, specific things. Um, Prometheus by itself, as in how do you actually scale Prometheus? Um, uh, with otel. As in scale Prometheus ingesting Not the one, not the. As in scraping. Um, and uh, of course AI, uh, that will come in as well. In terms of how do you approach, um, MCP or different different ways that we see this evolving. So yeah, I think we'll start to see blueprints being more connected to each other and being to fill the holes and the gaps that we, that we're identifying already in the guidance. So we have a way to contribute. Um, we can, I mean we're. Well we're welcoming people to author blueprints by themselves or to come to us and you know we work together and author in a blueprint. Right. So um, yeah, this is why the way that I see this evolving is in a way that we have the blueprints extending and also we make it more consumable. Talking about, talking about AI, more consumable by agents themselves. So um, at the moment you can go to Otel or OpenTelemetry IO and the Kappa AI or the AI in the website will already have consumed the blueprint and you know, be able to answer questions but maybe something that you can integrate more into an agentic workflow. Something that we talked about recently was you know, creation of skills and other types of formatted data from these blueprints. Right. And uh, yeah that's pretty much it sounds good.
Speaker B: So uh, this is a call out I think you also mentioned about the ecosystem explorer to also again extend.
Speaker A: Yeah so the ecosystem explorer for those that don't know and now it has a uh, at least for Java, ah, has a builder config builder so you can click a few boxes and it will give you the um, configuration to configure the Java SDK for example. So um, as I said blueprints don't have any um, configuration snippets in them. We're not going to maintain config for like how do you configure Java in a blueprint? Right. However we can link to a particular. This is what we would like to do in the future is link to tooling that can allow you to then um, generate config in a more automated way. So linking from the theoretical to the actual practical with configuration snippets that you can create.
Speaker B: Sounds good. So uh, we are reaching the end of the time. So we mentioned several times that uh a call out for everyone to join this initiative and we'll share the links on how to uh, to do that. We also mentioned here on the OpenTelemetry I.O. but uh, we'll share the links and how can people uh, uh, reach out and follow you.
Speaker A: Dan, uh, myself um, I am Dangb me on bluesky and you can find me on LinkedIn on uh, Dan Gomez Blanco and of course. I mean, that goes without saying. We're all there, just hang out, like, you know, it's a really nice community. So, yeah, reach out to us for sure.
Speaker B: So, uh, Dan, thank you so much for, uh, joining me and for doing this initiative. Uh, great to hear that and really, really exciting to see the first three, uh, blueprints, uh, two, three now coming out these days. So, uh, kudos to you and to the entire team for this great, uh, success for the entire, uh, OTO community.
Speaker A: Thank you very much. Thanks for having me.
Speaker B: Yeah. And of course, thank you all our, uh, listeners for joining us on this episode. Uh, as always, uh, this and every other episode. You can find them on, uh, your favorite, uh, podcast apps or, uh, on YouTube, uh, and of course you're invited to, uh, join us and follow us on bluesky, LinkedIn, xkenobserve, uh, to stay updated about the episodes, to carry on with more deep dives. If you have questions for Dan to me about this episode or anything else or some suggestions, uh, do, uh, ping us there and we'd, uh, love to hear from you. Uh, I'm Dotan Horvitz. Thank you very much for listening and, uh, concluding the sixth episode, sixth year of, uh, Open Observability Talk. So, uh, may the open source be with you and see you on the seventh episode on, um, the seventh season. Sorry. Bye.
Speaker A: Bye. Bye.