The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/Data Engineering Weekly
Data Engineering Weekly artwork

Knowledge, Metrics, and AI: Rethinking the Semantic Layer with David Jayatillake

Data Engineering Weekly · 2025-08-20 · 42 min

0:00--:--

Key moments - from our scoring

Substance score

51 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality10 / 20
Guest Caliber13 / 20
Specificity & Evidence9 / 20
Conversational Craft8 / 20

David Jayatillake, a 15+ year data industry veteran with experience at companies like Metaplane, Cube, and Avora, defines the semantic layer as the compiled, applied knowledge graph that translates business questions into executable SQL queries. Unlike BI tools like Looker, Power BI, or Tableau that embed semantic logic, standalone semantic layers (like those from Cube or dbt) enable interoperability across multiple tools and support embedded analytics use cases. The core challenge is maintaining semantic layers as business requirements change - dimensional modeling and Kimball techniques can support lighter semantic layers, but they don't eliminate the need for a system that enforces consistent metric definitions. Jayatillake argues that semantic layer ownership should ideally sit with product engineering teams end-to-end, though hybrid models with embedded data engineers and analysts are common. The conversation covers metric trees, governance challenges, and how AI tools can dynamically extend semantic layers by generating new calculated metrics and detecting redundancies when analysts request data outside existing definitions.

Key takeaways

  • →A semantic layer is the applied, compiled knowledge graph that sits behind an API and translates requests into SQL, distinguishing it from a static knowledge graph or dimensional data model.
  • →Standalone semantic layers solve the lock-in problem of BI tools, enabling multiple tools (Power BI, Tableau, Looker, Sigma) to read from consistent definitions across an organization, which becomes critical at scale and for embedded analytics.
  • →Semantic layer maintenance and governance require cultural buy-in around data-driven product development and metric ownership, often implemented through hybrid models with embedded analytics engineers.
  • →AI tools can dynamically extend semantic layers by generating new metrics and calculated fields when users request data outside existing definitions, potentially automating much of the maintenance burden.
  • →Metric trees and metric chains (e.g., ACV × deals = revenue) are essential for root cause analysis and understanding how changes in component metrics drive business outcomes.

In this episode

  1. 1Introduction and Background: David's 15+ Years in Data
  2. 2Defining Semantic Layer: Knowledge Graphs and Codified Business Logic
  3. 3Semantic Layer vs BI Tools: Lock-in Problems and Universal APIs
  4. 4Challenges: Dynamic Definitions, Collaboration, and Standardization
  5. 5Ownership and Implementation: Product Teams and Data Engineers
  6. 6Handling Unknown Data: AI-Powered Extensions and Self-Service Metrics
  7. 7Metric Trees and Root Cause Analysis in Semantic Layers

Mentioned

David JayatillakeCubeMetaplaneDelphi LabsLookerPower BITableauDBTExcelSigmaAvoraBusiness Objects

Guests

David Jayatillake

Topics in this episode

Semantic LayerLookerKnowledge GraphPower BIDAXTableaudbtCubeDimensional modelingSigma

Questions this episode answers

What's the difference between a semantic layer and a knowledge graph?

A knowledge graph is a structured representation of entities and relationships, while a semantic layer adds a compiler layer on top that translates requests (via JSON or API) into executable SQL, making it applied and actionable rather than purely descriptive.

Why is a standalone semantic layer better than having it built into a BI tool like Power BI or Looker?

While BI-embedded semantic layers work well for small single-tool organizations, standalone tools enable multiple BI platforms and custom applications to query from the same definitions, avoiding vendor lock-in and reducing logic duplication at enterprise scale.

Who should own the semantic layer in an organization?

Ideally, product engineering teams should own it end-to-end - emitting events, building data transformations, and defining metrics as part of the product development process - though hybrid models with embedded data engineers and analysts are common in practice.

How can AI tools help maintain semantic layers as business requirements change?

AI can generate extensions to the semantic layer dynamically when users request metrics or data that don't exist, and can also detect redundancies or similarities to existing metrics, reducing the bottleneck of manual metric definition and review.

How does dimensional modeling relate to a semantic layer?

Dimensional modeling and Kimball techniques create well-structured data models that make semantic layers easier to build and maintain, but a well-designed data model alone is not a semantic layer - you still need the applied layer that enforces consistent metric definitions and eliminates custom SQL.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The episode delivers a handful of genuinely useful distinctions - the compiler-as-differentiator between knowledge graph and semantic layer, AI as a maintenance mechanism, and the idea that organisational scale not data scale is the real trigger - but a large portion of runtime is spent on foundational explanations and conversational meandering rather than novel claims per minute.

every semantic layer is a knowledge graph, but not every knowledge Graph is a semantic layer. And why is that? It's because a semantic layer uh, in my mind always has a compiler.
you can't have metrics without a semantic layer, right?

Originality

10 / 20

There are a few genuinely fresh framings - the compiler as the definitional dividing line, the forcing-function argument about AI access pressuring data hygiene, and the claim that MCP dissolves semantic layer standardisation wars - but the bulk of the episode recycles well-known BI lock-in and single-source-of-truth arguments that have circulated for years.

I think that the semantic layer could become invisible right in the future
if you're doing MCP and it doesn't have to be mcp, it could be some other agent standard, I don't really care. Let's say MCP for now. If you're doing mcp, the format of the semantic layer and the standard of the semantic layer doesn't matter.

Guest Caliber

13 / 20

David is a genuine practitioner with 15+ years of hands-on experience, having led data teams at Worldpay, co-founded Avora and Delphi Labs (a semantic-layer AI company), and worked inside Cube and Metaplane; he speaks from real implementation experience rather than thought-leadership positioning, though he is not a widely recognised industry figure at the top tier.

I ended up leading data teams at various companies including List, Elevate, Credit, worldpay before going on into startups where I've worked, been a co founder a couple of times and also worked at companies like Metaplane and more recently Cube and have founded a company called Delphi Labs in the past which was focused on applying AI to semantic layers.

Specificity & Evidence

9 / 20

The episode names specific tools and vendors (Cube, Looker, Malloy, DBT, Databricks metric views, DAX, MCP) and offers a couple of illustrative examples like the ACV formula and a 30-second vs 30-minute latency comparison, but there are no hard metrics, customer case studies, or concrete before/after outcomes to anchor the claims.

ACV times number of deals equals revenue
you've got an answer in 30 seconds instead of the best case, 30 minutes

Conversational Craft

8 / 20

The host brings a healthy sceptical framing and raises substantive structural challenges - lock-in, standardisation impossibility, keeping pace with change - but questions are extremely long-winded and grammatically difficult to follow, follow-ups rarely drill into David's specific claims, and no assertion goes meaningfully contested.

I'm suspicious guy about the semantic layer and the reality of the semantic layer
I feel like the standardization is literally not possible. Like every vendor trying to push their one.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B70%
  • Speaker A30%

Most-used words

semantic125layer121data90model21product21metric21knowledge19tool18modeling18build17point16standard13building12better12tools12metrics12

Episode notes

Semantic layers have been with us for decades - sometimes buried inside BI tools, living in analysts’ heads. But as data complexity grows and AI pushes its way into the stack, the conversation is shifting. In a recent conversation with David Jayatillake , a long-time data leader with experience at Cube, Delphi Labs, and multiple startups, we explored how semantic layers move from BI lock-in to invisible, AI-driven infrastructure - and why that matters for the future of metrics and knowledge management. What Exactly Is a Semantic Layer? Every company already has a semantic layer. Sometimes it’s software; sometimes it’s in people’s heads. When an analyst translates a stakeholder’s question into SQL, they’re acting as a human semantic layer . A software semantic layer encodes this process so SQL is generated consistently and automatically. David’s definition is sharp: a semantic layer is a knowledge graph plus a compiler . The knowledge graph stores entities, metrics, and relationships; the compiler translates requests into SQL. From BI Tools to Independent Layers BI tools were the first place semantic layers showed up: Business Objects, SSAS, Looker, and Power BI.

Full transcript

42 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Hello everyone. Welcome to Data Engineering Weekly. We have great guests today. We have David Jaithal. He's a leading voice in semantic layer and you write frequently in substack. I will link the substack in the podcast. I'm always passionate about to understand semantic layer. To be honest, I'm suspicious guy about the semantic layer and the reality of the semantic layer. But a lot of things are changing in the industry right now with AI AH and LLM coming in and the quest for building knowledge graph is growing very high. So we wanted to better understand what a semantic layer does that and what it brings to the industry and what it means then and what it means now. So welcome to the podcast. David, how about you give a quick intro to you to the audience.

Speaker B: Thanks. Great to be here. Yeah. So I'm David. I've been in data for 15, 16 years now. Initially started off as uh, what was called an analyst at the time, but a lot of what I was doing was data engineering in hindsight, mostly building pipelines and automated processes to process data for end like weekend month end reporting. Then I do some analytics as well back in the day. So I ended up leading data teams at various companies including List, Elevate, Credit, worldpay before going on into startups where I've worked, been a co founder a couple of times and also worked at companies like Metaplane and more recently Cube and have founded a company called Delphi Labs in the past which was focused on applying AI to semantic layers.

Speaker A: Okay, great. So thinking of more about semantic layer. How do you explain semantic layer to a person who new to this concept, been, um, a data engineer and a research, but I don't really know what is semantic layer here? What does it mean? How do you explain semantic layer to them?

Speaker B: Yeah, so I think one way I think about it is I say to people who, so if any company that uses data, even if it's just for reporting, there's always a semantic layer. Semantic layer is always there. It might just be in like in, in someone's head. Right. So I've been that human semantic layer and so that's what I started off in my career doing. So I'd build those data pipelines for the first week of the month and then the remaining three weeks. I do like ad hoc or proactive analytics projects and when I was doing that, someone would ask me something if it was reactive and I would use my knowledge of, oh, uh, this is the data model we have, this is how the tables join together. This is how I figure out the number of customers or measure any other entity. This is how I calculate revenue, this is how I group by the right things. And so in my head was this knowledge graph of the entities of the business and how they relate to data structures of the business and how they join together and how, and so I could, from the request that someone gave me in natural language, I could then generate SQL, write SQL that would, as close as I could answer that question from the results. And so then I pivot. This was before BI tools. I'd pivot the results in Excel, Excel 2003 or whatever I had and make graphs and explain like what, try and answer that person's question and share it with them. And um, not a hell of a lot has changed right in the way people answer data questions. The tools are ah, better, right? I started off on SQL Server and those kinds of tools and I didn't have a BI tool, I had Excel. And like nowadays you have amazing cloud data warehouses and much more powerful BI tools, but the process is still the same. And one of the things I thought about even when I was young was like actually is it possible that this knowledge that I've got in my head about how the data fits together could be systematized, codified and so that someone could ask and then because they know the terms they'd get their, the SQL would be generated. I thought about building this at the time in DBA or Excel. And then you generate in a simplistic way generate the SQL required to answer uh, the question rather than me having to write it. And so that's what I, so when I say there's always a semantic layer, the semantic layer is either in someone's head and they're doing that translation from the request in SQL it's a piece of software. And so these pieces of software, they're not that new. You've got things like business objects which I've used. So say yes from Microsoft Looker, uh, more recently and uh, then the headless ones like DBT Cube at scale even more recently than that. And so it's a case of is it in someone's head or is it in a tool? And if it's in a tool you've codified this knowledge and probably in YAML and you say, oh, this entity is an abstraction from this table and this is how they join together, this is how you aggregate things to find measures. And so that's what a semantic layer is from a high level view. And I'd say every semantic layer is a knowledge graph, but not every knowledge Graph is a semantic layer. And why is that? It's because a semantic layer uh, in my mind always has a compiler. And so the compiler works behind an API of some kind where someone can ask make a request, whether that's in JSON or however it is to I want these objects. The compiler uses the knowledge graph of the semantic layer to know how to compile that request into RAW SQL. And it could be very complicated. SQL executes it and pulls the data back. So that's the difference between a knowledge graph and a semantic layer in my mind.

Speaker A: So you're positioning knowledge graph more of a discovery to understand what is peeking to your memory. But like the semantic layer is more of triggering an action. Take a glass of water based on your.

Speaker B: Yeah, it's very applied.

Speaker A: Yeah, it's applied. Okay. Maybe an applied knowledge graph. Yeah, okay. Okay.

Speaker B: For me.

Speaker A: So this whole knowledge the semantic graph concept, right. I in theory April, I agree what you are mentioning about a person knowledge. If we can codify that and then you can build this one. And there was a lot of interesting conversation happened about the semantic layer. One of the common argument was like oh, bi doing this for 20 years. How bi doing this for 30 years. When a Power BI has it, Looker has it some kind of a semantic layer definition you can give this one and uh, business object has this object definition that you can do and then build certain things in SAP for a long time. It is true for that. So I think is that like what is wrong with that format and BI holding the business layer there's. You think it's a semantic layer then whatever the BI uh, that implemented that and yeah, what's the pros and cons of it?

Speaker B: Yeah, absolutely. BI tools are the first things to have semantic layers because they're the most obvious place that they were needed. The problem with that is lock in. Now if you're a small company and your only use of data is for reporting and you only have one BI tool, there's nothing wrong with that BI tool having the semantic layer and actually it's the most perfect place for it because you don't have this friction between the front end applications that would use the semantic layer and the semantic layers API like it's all managed in one place. And that's why tools like Looker are successful because it was done in one place. And same for Power bi. Power BI has DAX now as that kind of layer where. And it nicely fits together because Microsoft has built both sides of the interface. So there's nothing wrong with it. The Problem comes when complexity and scale comes. So if you think about a big Wall street bank, right, they'll have every BI tool, not one. They'll have all of them and they'll have all the warehouses and all the clouds. And when you have your logic trapped in one BI tool and naturally inaccessible by any other, because why would one BI tool vendor want to empower another? They won't, then you get stuck and you end up with these silos of logic. And so the whole point of the semantic layer, which was to govern and codify the meaning once of anything, is no longer possible because you need to do it in 10 different places. And so that's where having a standalone or universal semantic layer, however you want to call it, makes more sense. So let's say with Cube, where I used to work with Cube, you could have Excel, you could have Power bi, you could have Tableau, you could have Sigma or whatever else reading from the same semantic layer, uh, using the same definitions across an org. And that's really valuable when you have that kind of complexity and scale. The other reason as well is when you have embedded analytics or data products as some people call them. Often you don't want to use a BI tool to power a data product. You'd rather have an API that you call and you use some kind of like React or something or some framework in front end to build your data product and share it to your customers. And that's again where having it locked in the BI tool is problematic. Some BI tools in the past have provided quite good like uh, APIs to pull data, like looker had something, but again it's limited to what they want to build in that way. And it's not ideal. It's better to have something that's not stuck in the BI tool.

Speaker A: Okay, yeah, I think it perfectly makes sense to have that locking, removing the locking, making like every other application domain interoperable with each other. That brings in other questions about this notion of codifying our knowledge into some kind of a document that essentially assume that compile once and run anywhere from anybody else, anyone else. Again, the two challenges. One is the, the dynamics not going to be seen, for example revenue metrics for example, that might be static for a long time, but when the business changes pretty often the knowledge has to keep improving. Like your semantic layer has to be very improving a very long. And it is no longer like one person internalizing that. Like when you are telling to other person that it has to be some kind of a collaboration needs to be happening to make that definition happen often requires multi team collaboration on that.

Speaker B: Yeah.

Speaker A: And so in given the case how feasible is that? That's the first question can be codifying that particular part of it. And that brings like when we split it out a layer from um. A BI kind of a thing from a technical perspective from one common interface the notion of like common interface never we know you create five standard there will be a six standard to unify all the other five standards and implementation all those implementation looker has to implement in a very optimal way Tableau um has to implement in a very optimal way because they all still need to query the source of data that is somewhere else and that's a cost and other performance and other attributions causing. So when we splitting this layer these are the two challenges coming in right. Can it be a uh. In the ever going changing the definition and then is there a standardization even possible? Will everyone able to implement the same performance? What do you think of all those options?

Speaker B: Yeah and so the real problems and problems I saw and so like when. So if you think about collaborating with other tools in the stack right there is problems that because you need often you need that vendor uh to support your tool and they are not heavily incentivized to do that as unless they have decided very specifically they don't want to have a semantic layer and therefore they're happy for someone to buy your tool and there which can happen but it's rare and then they don't prioritize working on the integration. The best case for that is where it's an open source tool and then you can build the integration with your semantic layer or if that BI tool has a standardized API that you can basically pretend to be. So like with Power bi the DAX API will always work with it. So if you can be the DAX API or the MDX API you can work in Power bi, you can work in Excel. So that's the difficulties there. And I think you were talking about collaboration also then internally like where let's say product engineering are uh building a new feature and then you want them to emit the right events with the right schema to build a data model and then use the data model to have your semantic layer. I think that the thing there is actually that's more of a cultural thing and that's what I've seen where you have to persuade that business. If you're going to do product development, why are you doing this product development? Surely you're doing it to push something forward, to improve something. So every Product feature development should have success metrics, right? And so if you have success metrics, the success metrics existence dictates that a semantic layer exists. You can't have metrics without a semantic layer, right? And because you've got entities and you've got metrics and you've got dimensions and that's how you're going to use this success metric. And so if you're going to have that semantic layer and then that sits on top of the data model that you've emitted as a product engineer, I think it's about thinking about that process. Okay. You want to measure how well this product has done and so therefore you have a success metric. If you're going to do that, you need to go end to end. Uh, you need to omit the right data that can power a data model that can then fit under a semantic layer. And I think it's that way of thinking and that's hard. Like lots of companies fail at it. But if you don't have a product function that is data driven or has a data culture where they want to measure everything and they'd rather shoot from the hip and just oh, we're going to build this feature because we think it's the right thing to do and that can be the right thing at times. If that's how your company works, you'll always struggle, right? Their product engineers will probably not even omit the data. You need to uh, before you even get to the semantic layer. And then you, yeah, then you're like scrambling around trying to measure these features. But yeah, I think where there is buy in and they want to measure and that measurement is part of the product of product development and release process. I think it can work.

Speaker A: You mentioned one point, right? In order for to measure the success, we need a semantic layer. So many companies right now, their notion of semantic layer is like either they have adopted some kind of a domain architecture or some kind of a dimensional modeling that, okay, I got to say here's my dimension, here's my facts and I going to query that one. How do you differentiate the dimensional modeling technique that people are using with the semantic layer on top of it?

Speaker B: So things like Kimbal modeling or data Vault or whatever you might use, that is a way of arranging your data, uh, and making it available. That could make having a semantic layer very straightforward. Right? Because you have beautiful like foreign key relationships, good granularity in your tables. Your tables represent entities really nicely. Your tables capture necessary state changes nicely. Then what you're saying is, yes, you still need a semantic layer, but it's very light. All you're saying is table A take joins to table B on this foreign key, which is clear, you don't have to do any transformation before you get to that. So then what you've done is you've made your work easy for the semantic layer. You still. The problem with that is uh, it's still not a semantic layer on its own to just have the data model because someone still has to write code to access it. Someone still has to know which even though the joins are easy, they still have to know them and they still have to know, oh, you might need to run this filter to get this metric to come out properly. And so when you do that you still end up with like different versions of the truth. And that's the joy of the semantic. There is, you're saying, no, this metric is always calculated this way and the person accessing it doesn't need to write code every time to get it. And so that's why I think having a good data model is obviously a very good idea, very important to uh, even power a semantic there. But it's not the, it's not the full task. There's more.

Speaker A: Okay, so in a typical organization, who or which team plays a role of building the semantic layer? Right. Somebody has to own the definition of it. Is it a distributed ownership or um, it should belong to one team who makes sure that my semantic layer is not going to get copter or some wool. Some are going to be working on that one.

Speaker B: Yeah, ideally I would say the product engineering team like the, if you think about that they need the metric, they, and they are emitting the data. If they were going end to end and it was like a squads and tribes kind of way of working, they would build the event emitted, the transformation models that take the events and make them available in this model and then the small semantic layer on top at the end, I would think they would do the whole thing. Whether that means the product engineering team has embedded data engineers or embedded analytics engineers, may that probably would be true. Right. Or whether you react, work with some central data team that's doing this for you, that's possible. Also like at the moment I've seen, I've seen it like work in a hybrid way where the product engineers are really just software engineers and they're doing the emission of the events to an API that can accept them in a standardized way. And then you've got product analysts who are somewhat embedded into those teams and the product analysts are doing rudimentary transforms on those events which are clean, they still need some level of transforms and then they're building like a semantic layer but it's not efficient. So often what we're seeing also is an analyst, extension or data engineer taking that product analyst work and say, oh okay, I'm going to incrementalize your transformations, I'm going to clean up your semantic layer so uh, it runs more efficiently and runs faster. And so there's that. But yeah, in an ideal world I think it should just happen as part of the product development process.

Speaker A: So what will happen in the cases where the semantic layer not able to catch up with whatever the changes to this one? Like any good system should respond to a query that thinks it knows and just say that I don't know, that all this information I cannot find it.

Speaker B: Yeah, absolutely. And that's one of the things I really liked when I was working with Delphi was at the time people were building text to SQL systems and they would just make things up, right? If you ask them about data they didn't have or they didn't understand, it would make something up. Whereas if you ask Delphi, ask about some data that doesn't exist in the semantic layer, it would just say I don't know, which is good. That's what a human would say. If you ask them about data they don't have or they don't know. And so that's. Yeah, but in the case where it doesn't have it and you're saying people need things, what do you do? You can expand that semantic layer and nowadays I think you can do that with AI. So even a non technical person could ask, okay, you don't have the version of revenue I want, you only have like gross and. Net. I want ebitda. And then it says to the system, okay, EBITDA is calculated by taking net revenue and subtracting operational operating costs and debt costs or something like that. And then the system goes away and um, figures out, oh okay, so I know where that data is and then makes a new calculated metric and stores that to the semantic layer and then the non technical user gets what they want and it's still governed. That's ideal. In reality, I think right now if it's not in the semantic layer, people are probably just writing custom SQL. And that's what I've seen. I've seen loads of product managers write their own custom SQL for things which aren't in the semantic layer. And I think there's like a level of which if your Organization's not in a place where they want to put a lot in the semantic layer or they just don't want to dedicate the effort to it. You go 80, 20 and think, what's the most important data we have? Like, what's our gold medallion data like? We need to cover that in our, uh, semantic layer and all the experimentation data. I'm sorry, that's not gonna, that's not gonna go in. It's moving too fast. We can't incorporate it quickly enough.

Speaker A: Yeah, I think that is one of my criticism essentially. Like, even if in an ideal state that I have a semantic layer, there is no way that I can keep up to date with what people who wanted to ask this one. So one notion strictly say that if you are building a certified data product, it has to like all these dashboards are certified and this is a standard dashboard. It has to consider the semantic layer side of it. And rest of them just. I opened up everything because I, I may not able to cut it up to the semantic layer. And the second cause is like the regressions that I need to go to make sure that the testing and everything is completed and updated. It's a little hard. And if I land it and I have to tell people that, oh, for this query here, for this query semantic layer. So that kind of complicates things from a fiction perspective on an adoption perspective. Do you see any better solution for that?

Speaker B: Yeah, and I think this is where I'm very hopeful about the new AI tools coming in, is that, that they can do that, where they can keep up because they, as, as fast as someone can ask, uh, AI a question about data, if it's not in the semantic layer, the AI can generate an extension to the semantic layer. Now the bottleneck probably then becomes your, your PR review process. Right. If the AI is generating a bunch of new extensions at the semantic layer that some human has to review. Sure. But you imagine you've ever, you could overcome that as well at some point and then you don't have that lagging problem. And every time the AI generates, uh, an extension to your semantic layer, you can figure out very quickly, oh, am I creating something that's very similar to something that already exists? Should I not create it? Should I just extend that thing that already exists instead? So that's what I think that. And um, you're right, that is what. That has been one of the flaws of semantic layers. It's difficult to maintain them. But yes, I think really the solution to that is, is AI okay.

Speaker A: Yeah. Before Jumping onto the AI I feel one other question I have is we talk about a lot about metric tree and like how this cautious relationship of uh, the metrics can impact you can run a simulation at scale. How do you see the concept of metric tree play? Going to play with the semantic layer.

Speaker B: Yeah. And so this was something that I used to have in the uh, one of the first time I was a co founder at a company called Avora, we were like spinning out a root cause analysis tool out of an old BI tool. And in order to do root cause analysis and data, it's not enough to just look at dimension permutations. You have to understand the metric tree. You have to understand what we called them metric chains. And so you write a formula to say oh, ACV times number of deals equals revenue. And so therefore you could explain revenue changes which might have been the thing people actually looked at by oh, ACV dropped last month, not the number of deals. So we closed the right number of deals but they were just small and you need that. And then you've also got the these, you've got these relationships which aren't so direct but like indirectly causal. And typically that's where you've got like marketing spend where just because you spend another dollar doesn't mean you'll get an extra $2 in revenue. It diminishes depending on how you spend it and um, it's probabilistic and so you then need to be able to capture those kinds of relationships as well in order to explain root cause. So yeah, that's like the way I see metric trees is there's a reason is a way to explain why the top level metrics that you have, the important ones that your business is looking at, why has that changed? We need to go down the tree potentially. Usually it's something down the tree has changed and uh, not just a permutation of dimensions.

Speaker A: Okay, yes. Yeah, it's good to think about metric layers and tracing for metrics that one of that concept. Now it goes back to another important moving back to the AI side of it, both metric layer, uh, the metric tree or semantic layer side of it like we talked about, like how it's imposing, like it's hard for keep up to date and there's so much of manual modeling and other aspects involved in it. What do you think AI can play a significant role here? Like how and why the AI is significant for semantic layer?

Speaker B: I guess so one nice thing about metric trees in this regard is that in theory, uh, if you have a complete Metric tree for your organization. You don't really need to keep generating lots of new metrics. And arguably a new metric that's outside the tree is probably wrong because you're uh, probably asking something that's captured from an information point of view in the tree. But you're like, I want something a bit different. But actually if you conformed to the things you already have in the tree, you'd have something that was more standard and more well understood. But there will always be exceptions to that where actually you did need something new. And so I guess your question is how does AI help with this? And I think I alluded to this earlier where I think that the semantic layer could become invisible right in the future. Where yes, it's there, yes, it's codified, yes you can ask it how something's defined or how it chooses to join things, but it's maintaining that codification of the semantic layer, the AI is and it's also maintaining and extending it. When metric definitions change, it's making those changes. And so uh, everyone gets the benefit of those changes. And I think yeah, it resolves the keeping up point, but it also solves for the place where you don't have, where you need to increase the number of metrics you have to answer questions.

Speaker A: Yeah, automatic building in a reversible semantic level. Right. I think that goes back to one of my previous questions, like the standardization of it. I feel like the standardization is literally not possible. Like every vendor trying to push their one. Like we've seen there and seen the DBT cases where even though their initial argument was like open the semantic layer out of BI and proprietary semantic stuff of it, maybe for a good reason because uh, if they know the concept and the hunter of the thing, they can probably better optimize, optimize the system and make a better user experience. That's what BI tools also trying to do that one. So even if you build the semantic layer automatically in an only side of it, that interface tree needs to be somewhere near align on that. And so far SQL is being adopted as a standard for querying this one do you think right now. But semantic layer has not defined as like enameled. Like this is my description on their 1. Do you think that will change to somewhere into the adoptable layer, uh, and that we don't need to fight about this whole standardization. We build on whatever the standard is already available that everyone is talking to and then just forget about that.

Speaker B: Yeah, exactly. So I think one of the things I've seen recently is in the Past I think Excel and still today Excel is like the place. What I'm seeing now is companies are rolling out enterprise like ChatGPT licenses to their whole company and people are starting to do those kind of Excel type analyses, like this quick analysis. They're doing them from ChatGPT or Chord and then the place is there. That's where the data, uh, that's where the data is ending up. And so the current way of thinking about enabling that is with mcp. And so if you're doing MCP and it doesn't have to be mcp, it could be some other agent standard, I don't really care. Let's say MCP for now. If you're doing mcp, the format of the semantic layer and the standard of the semantic layer doesn't matter. Right. Uh, the user just needs governed consistent metrics and dimensions. And then if AI in the background is maintaining that semantic layer, does it matter what standard it's using? Probably not. And actually I think this is where in the past I've probably been a bit, not unkind but a bit critical of maloy because I just didn't see how it could be used like in production very easily. Whereas actually now if AI is just like has a MALOY data set and then just extends it when it needs to or uh, uses it to answer a question and that's exposed for mcp, it doesn't matter, all those problems go away and you can use any standard. You could use Cube, you could use Malloy, you could use whatever you want. Really?

Speaker A: Yeah, yeah. I think one of my recent vinnie also talked about it like one of the blog that I read so far that we develop language for human to understand and data model for human to understand and semantically it is the one thing that we are talking about like the data modeling technique that we did like dimensional modeling that invented when the storage was very scarce, all the slowly changing dimension altogether. And then we started with whole snapshotting and data vault and everything on utilizing object storage.

Speaker B: Yeah.

Speaker A: So the data modeling itself is evolving as the fundamental structure is changing. But all this data model is still focusing on human readability. Like how I arrange dimensional modeling, physical layout perspective. I feel like semantic layer is elevating one level to more aligning with the large language model. So do you think semantic layer uh, will eventually become the de facto like a modeling layer, like data modeling layer rather than like a similar to on the scale of what Kimber methodology is another dimensional modeling technique evolve with a dimensional like an awesome modeling layer that somehow transpire inherently to the physical layout. So people operating on that semantic layer modeling instruct first I build my data modeling and build semantic layer versus I build semantic layer that compile to a modeling uh, technique whatever the physical layout make a better interrupt.

Speaker B: Yeah, so you're talking about going top down from the semantic layer. So you just define your semantic layer and then that will dictate the data model underneath. I think that's a good way to think like from a engineering point of view is I need to. Because uh, fundamentally you're building this data to be used. The usage is codified in the semantic layer. So why not build a data model that fits the semantic layer that's kind of like the right way. And I think the problems have been in the past where people start so far to the left like someone built the pipeline, then they dump some raw data into object storage for the next person to pick it up and then that person's gone and built some more models and then someone else has built the semantic layer on top and it's made it really hard. Yeah, and I think that's right. But I also think yes, the semantic model is much closer to language. If you think about how anyone has ever tried to use it. Even like using a REST API that's really quite close to language. It's like reading off a menu and asking for things from the menu. Yeah, just the way I think about it. And that's much better for AI to use than to write code because it's literally you're saying here's the stuff, choose the stuff that you think it will answer the question. And that's a much more safe method. That's what I think all of the successful AI data tools are doing now. So whether that's Cube, DBT, ThoughtSpot, Databricks and Snowflake are now going this way. So recently I've been looking at databricks metric views. Metric views are semantic layer. It's just on a microcosm of a semantic layer rather than having a giant one. Yeah, and eventually, and I say eventually meaning before the end of the year they're going to have the ability to have a much larger semantic layer as well.

Speaker A: Yeah, yeah. It would be interesting to start some kind of a open source project around like how like on top of semantic, like taking all the learning from semantic layer and like how easy for to build human readable like a machine readable way of building this model and then automatically compile two dimensional modeling or data world.

Speaker B: So the thing is I don't think there's a big difference between what's human readable. And we say that we were geared towards human readability. We didn't do a very good job. I think people have got all these horribly named things in their data model and badly structured. There's so many of those flying around. And because humans could over time learn how to deal with the abrasive nature of that data model because they would learn, oh, okay, this weird table name means this. So you'd have that semantics in your head. And what we're seeing is actually if you make it better, better for machines, you're making it better for humans because you're just making it clear because that machine. Let's imagine you'd built a semantic layer for an insurance company. The machine is like an expert in insurance. Like it knows insurance terms. It doesn't know your company. That's what it doesn't know. It doesn't know the specifics of your company. So you conform your data model and your semantic layer to uh, standardized insurance terminology. Then an AI looking at that data model and that semantic layer will be able to use it, uh, very easily. But it's where you haven't done that and where it's not clean, where the names are not far apart, that's where it will struggle.

Speaker A: I think that is a challenge that I'm also seeing that for example dimensional modeling, if you take any data modeling technique right now, it says the same business, like all the business process identifying what is a grain and define your surrogate key and how the relationship put in your uh, matrix here. But it doesn't dictate or make it mandatory that you have to define a human readable form or a machine readable form. It's like you codify that into some form where who has an understanding of a data model then he has to have some ask questions to people who have done it to understand, oh, this metric here is mean. So this is a surrogate key here. This is a primary key here. Are the documentations bound. That is not a mandatory or like an odd part of it. Unless until some organization make it like it has to be mandatory.

Speaker B: And I think that's coming because we're seeing uh, senior stakeholders say hey, I want to access my data with AI and they're saying that's happening.

Speaker A: Yeah.

Speaker B: Then read that you're forced towards a semantic layer because if you do that on text to SQL, it just won't work. And then when you have the semantic layer, it can't be a dirty semantic layer with lots of similar Names or poor bad names and bad structure. It needs to be relative. The cleaner it is, the better the answers they're getting. And so if you don't want to get harassed by them for your system being bad, yeah, you need to clean it up. And so I think that's the forcing function that's happening, correct?

Speaker A: Yeah. So people listening to this podcast and thinking of, hey, I should buy some semantic layer right now, or how can I fit into a semantic layer on top of the giul? Like every organization, they don't start with a semantic layer to start with, even they don't have a data warehouse. They just query from operational data store and then scale limit head, then put the data into data warehouse and then after some point they focus on how should I organize my data. That's the nature of any data, uh, engineering ecosystem. What would you advise for them to. How to think about if someone on that spectrum that say, hey, I understand the value of it, or integrated details of semantic layer, but how do I implement? I do already have an existing state here.

Speaker B: Yeah. So are you talking about where people don't even have a data warehouse or where they've got data coming in or

Speaker A: what they might have a data warehouse, for example. They are in a stage like they have some established data warehouse snowflake or a databricks. Yeah, constantly querying that one. But like, at what point do you say that this is going to be the pain point and this is where you need to think about semantic layer and this is the, this is how you can incorporate semantic layer to an existing system.

Speaker B: Yeah, I think the pain point, yeah, usually it's with scale and with semantic layer, it's not like data scale that it solves, although they can help with that with caching and things. But it's more about organizational scale because once you get to a point where you need to disseminate knowledge more efficiently because you can't know everyone or talk to everyone or have a meeting with everyone, then it's about codification of knowledge. Right. And so semantically is another codification of knowledge. It's just also in the production line in that people use it to access as well, which is part of its power. Right. Because people use it to access, it stays fresh. And so I think when you get to a point where people are just like pulling data all over the place and oh, every time someone looks at one source of data and another and they don't agree, like that's okay if that's your main pain point. That is what the semantic layer Solves uh, if you don't have that situation and everyone's quite good and you at uh, querying your data model, then you haven't reached a point where you have that semantic layer. The other pain point is where people struggle to access. Right. And so they're waiting a long time or always for an analyst or whoever, even if they always get a good answer, but they have to wait weeks to get an answer. Right. Whereas if instead of waiting for an analyst to write you a query, you could ask AI which then requests the semantic layer and then generates a query, you've got an answer in 30 seconds instead of the best case, 30 minutes. Right. That's the other, that's. Those are the two places like speed and inconsistency.

Speaker A: Okay. Yeah. We almost come to end of the podcast before parting. Do you want to share any like what do you in two to three years down the line, what do you hope the things will happen in the industry, not necessarily to semantic layer, but in the general in the data ware side of IT and engineering side or analytical space. What do you wish or hope to see in next two to three years?

Speaker B: Yeah, what do I hope to see? I like so some of the things which have already come to pass like Iceberg becoming like a uh, standard. I know you've had a lot to say on this like standard wars. I'm glad that's happened because people can. Even if it's not the best standard and I know it has its flaws, people can move forward and build on top of that. So it's okay. Rather than having data lock in, we can now see, okay, all the data warehouses just use this same data lake. So we haven't seen that yet. I We've seen people say they'll support Iceberg then they haven't all released it yet. There's lots of companies who are building their support for it. So what I want to see is that way of working which we were hoping for, which is that everyone can query with whatever engine they want and get the same access happen because we're seeing these. We're in this great era now where you've got things like Polis and DuckDB and other engines which you could use if it's from the same data store. Then you can have multi multiple query engines in your transformation pipelines or you can have a small DuckDB step because it's not that big or use databricks for another step and it's all on the same. That's one of the dreams that we really want to happen from a data engineering point of view. And I think I can see that happening soon. Then if you think a bit further out. Even though I think semantic layers, I still, every time I talk about them, I always have to explain them to people. Whereas if you're talking about a data warehouse, you wouldn't always have to explain it someone. And I don't think that's going to change because even though databricks and snowflake have their own semantic layers now, that concept of a semantic layer is quite abstract. And so I actually hope that I don't have to talk about it so much anymore. I hope that in three years, like people don't have to. There's tools that come out that do a very good job of hiding that semantic layer from people. And so engineers aren't. You don't have to sell a, uh, semantic layer to someone. No, they're just, they're going to use it and it's just baked into that tool that they're using and it's maintained without them seeing it.

Speaker A: Okay. Yeah. I'm also really excited about this whole AI side of it on the semantic layer. And then it's a bit of a driving to the concept of knowledge agent. Would be great to see how the industry evolve. Is. Thank you so much for trying, David. Great, thanks.

Speaker B: Uh, thank you.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Decision Logic: The Difference Between an Answer and a DecisionThe AI Forecast · on Semantic Layer87 / 100
  • How SSW turned AI into ½ their pipeline - Ulysses Maclaren, COO of SSWSaaS Stories · on Power BI86 / 100
  • Is Your AI Actually Worth What You're Spending? with Parker ConradStrictlyVC Download · on Tableau86 / 100
  • Microsoft Fabric: The Platform That Turns Data into Competitive AdvantageLeading IT - APAC Insights · on Power BI85 / 100
  • From Farm to FP&A: How California Dairies delivers 17 billion pounds of milk each yearFP&A Today · on Power BI85 / 100
  • AI Is Already Changing Data Jobs: Why You Must Create More Value Now with Rob CollieThe FP&A Guy Network · on Power BI83 / 100

More from Data Engineering Weekly

All episodes →
  • Insights from Jacopo Tagliabue, CTO of Bauplan: Revolutionizing Data Pipelines with Functional Data Engineering
  • AI and Data in Production: Insights from Avinash Narasimha [AI Solutions Leader at Koch Industries]
  • Is Apache Iceberg the New Hadoop? Navigating the Complexities of Modern Data Lakehouses
  • The State of Lakehouse Architecture: A Conversation with Roy Hassan on Maturity, Challenges, and Future Trends
  • Beyond Kafka: Conversation with Jark Wu on Fluss - Streaming Storage for Real-Time Analytics
Explore the best B2B Engineering & DevTools podcasts →
All Data Engineering Weekly episodes →