The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Marketing/The Search Session
The Search Session artwork

A Reality Check on Structured Data | Jarno van Driel

The Search Session · 2026-08-10 · 60 min

0:00--:--

Key moments - from our scoring

Substance score

68 / 100

Five dimensions, 20 points each

Insight Density14 / 20
Originality15 / 20
Guest Caliber16 / 20
Specificity & Evidence11 / 20
Conversational Craft12 / 20

Jarno van Driel, a longtime structured data consultant and schema.org volunteer, cuts through the hype surrounding markup and semantics. He explains the crucial distinction between structured data (a broad term covering many formats), schema.org (a specific vocabulary), and the syntax used to express it (RDFa, microdata, JSON-LD). The core problem, van Driel argues, isn't Google's documentation - which he praises as the most comprehensive available - but rather that most SEO professionals don't read the actual schema.org GitHub specifications, the W3C docs, or understand the use cases behind each term. He cites Web Almanac data showing 90% of markup exists only because Google created a rich results feature for it, and that removing non-Google features leaves almost nothing. Van Driel also addresses myths around structured data and LLMs, noting that Google is gradually deprecating rich results that don't fit the agentic AI pattern, while product schema remains the primary use case as Google aligns Merchant Center specifications with organic search schema.org definitions.

Key takeaways

  • →The biggest issue with structured data adoption isn't Google's documentation quality, but that SEO professionals copy schema terms without understanding why those properties exist or the use cases they were designed for.
  • →Approximately 90% of published structured data markup only exists because Google created corresponding rich results features; remove Google features and most implementations disappear.
  • →Most SEO plugins and WordPress implementations default to the same core schema types (Article, Organization, LocalBusiness, Breadcrumb) because they align with Google's documented rich results, not because they represent best practice.
  • →FAQ pages became popular due to LLM myths even after Google deprecated the feature, when in reality NLP was already capable of extracting Q&A content without markup since 2016.
  • →Google's shift toward LLM-based search results means many traditional rich results features are reaching end-of-life, and product schema is becoming the primary use case as Google works to align Merchant Center specifications with organic search markup requirements.

Guests

Jarno van Driel

Topics in this episode

Google Merchant CenterJSON-LDSchema.orgRDFaMicrodataWeb AlmanacW3C specificationsLinked Open DataProduct schemaFAQ page markup

Questions this episode answers

What is the difference between structured data, schema.org, and schema markup?

Structured data is a broad umbrella term covering Excel files, tables, and markup annotations. Schema.org is a specific vocabulary (ontology) that defines the meaning - like Person, Article, Author. The syntax (RDFa, microdata, JSON-LD) is just the vehicle. Schema markup is technically inaccurate terminology; the correct term is semantic metadata.

Why does almost all structured data markup on the web use only a small subset of schema.org terms?

Web Almanac analysis shows 90% of published markup only exists because Google created a rich results feature for it. Once you remove everything Google documents as a rich snippet feature, almost nothing remains - meaning implementations are driven by search engine features, not by independent business need.

Should SEO professionals use every property available in schema.org for their content type?

No; schema.org is a vocabulary laying out options, like a dictionary. You only use the properties relevant to your content and use case. Using every term simply because it exists is misusing an ontology, similar to insisting every word in a dictionary be used in a sentence.

Is it Google's responsibility to explain W3C specifications and JSON-LD syntax in their structured data documentation?

According to van Driel, not really - Google should explain what they want for their search features, but the actual syntax specifications are the domain of W3C documentation. Google's main limitation is that many SEOs only look at Google's docs and never reference the actual W3C specifications or schema.org GitHub.

Why did FAQ pages become popular again despite Google deprecating the feature?

LLM hype drove adoption despite the deprecation, but van Driel argues FAQ markup wasn't necessary for AI to extract Q&A content, since NLP was capable of extracting that strict question-answer pattern without guidance since 2016.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

14 / 20

The episode contains substantial insights about structured data mythology, the gap between schema.org theory and practical implementation, and realistic assessments of ROI for entity optimization. However, significant portions devolve into conversational meandering, personal anecdotes (summer weather, gaming preferences), and repetitive discussion of basic terminology rather than advancing novel claims. The guest does deliver contrarian takes (FAQ schema was unnecessary, most vocabulary goes unused, schema.org may need resetting) but these are not consistently packed throughout.

90% of the markup out there only exists because Google actually had a feature for it
unless you got a lot of money to burn, stay away from it. Any form of agent optimization.

Originality

15 / 20

Van Driel offers genuinely counterintuitive framings: structured data markup is irrelevant for LLMs (which parse natural language), entity optimization ROI peaked around 2016-17 due to NLP improvements, and most schema.org vocabulary will never be used. However, these insights are presented somewhat discursively and the core thesis (schema.org over-engineered, focus on pragmatic use cases only) while sensible, is not deeply fresh. The framing around knowledge graphs for internal business purposes is solid but not radically new.

LLMs don't need that format. They're not created to work with those formats. LLMs are created to deal with natural language
schema.org was designed for a different era...I wouldn't be surprised if we see a reset the coming years

Guest Caliber

16 / 20

Van Driel is a legitimate practitioner with deep pedigree: working in structured data since 1998, active volunteer contributor to schema.org governance, hands-on experience with enterprise knowledge graph implementations, and documented case studies (Panda recovery 2014, entity optimization evolution tracking). He speaks from 25+ years of real implementation work rather than theory. The only deduction is that he's a consultant/expert commentator rather than a current operator running a live business at scale.

he started his career...in this field in the last century, because it was 1998
I'm one of the volunteers out of dozens...Everything in there ended up in there with a certain use case in mind

Specificity & Evidence

11 / 20

The episode lacks concrete data and named examples. Van Driel references his own case studies vaguely (Panda recovery, entity optimization ROI decline around 2016-17) but provides no metrics, timelines, client names, or quantified results. He mentions general concepts (NLP reached 85% accuracy by 2017-18, knowledge graph error reduction to 80% with agents) without sources or specifics. Most claims remain abstract: 'Google has been aligning Merchant center specs for eight years,' 'product schema as biggest use case,' but without hard evidence or numbers.

I started the discussion around product variants around 2017 and it took only like six or seven years to roll out into production
NLP had become so good that it already crossed the 85% benchmark in accuracy

Conversational Craft

12 / 20

The host Gianluca asks reasonable setup questions and allows Van Driel space to develop thoughts, but the conversation lacks aggressive follow-up, pushback, or pressure-testing. The host often agrees or pivots to related tangents rather than drilling deeper. When Van Driel makes strong claims (FAQ markup was unnecessary, most schema never used), the host doesn't push back with counterexamples. The discussion meanders into weather, gaming, and extended anecdotes without redirection. Some questions are soft ('How is SEO treating you lately?'). There is one good moment where the host challenges Google's documentation clarity, to which Van Driel provides a nuanced rebuttal.

Yeah. Yes. In fact, it's interesting how the this is something that I like to do every year
Yeah, I understand. Well, maybe the future. Let's see. One of the evolution of a technology

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B61%
  • Speaker A39%

Most-used words

google84data62structured48schema42knowledge35search28graph28markup27instance26page21already20llms20information18important17based15case15

Episode notes

Structured data is back at the center of SEO conversations, but not always for the right reasons. In this episode, Jarno van Driel joins Gianluca Fiorelli to separate fact from fiction. As an active Schema.org contributor, Jarno shares first-hand insights into how the vocabulary evolves, what structured data can realistically achieve, and where many SEOs are getting it wrong. This is the 75th episode of: Here's what you'll learn in this episode: 00:00 Intro 00:06 Meet Our Guest: Jarno van Driel 02:00 Structured Data vs. Schema.org: What's the Difference? 07:10 Google's Documentation: Helpful or Part of the Problem? 20:37 FAQs, Products & Merchant Center 27:24 Do LLMs Actually Need Structured Data? 30:13 Should Businesses Invest in Agent Optimization Yet? 34:48 Internal Knowledge Graphs: The Real Business Opportunity 42:23 The Myth of "Schema Fairy Dust" and E-E-A-T 47:02 Writing for Accessibility Is Good SEO 49:51 Will a New Vocabulary Replace Schema.org?

Full transcript

60 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Hi, I'm Gianluca Fiorelli. Welcome back to the search session. Today we are going to have maybe one of the, uh, longest evangelists about structured data and semantics, search and semantic web in general. He's from Rotterdam. He's, uh, a consultant and, uh, really started his career, I can say it somehow, uh, like me, not, not really like me. But surely he started his career, uh, in this field in the last century, because it was 1998. So it sounds very, very old. But this guy is not that old at all, especially his brain. It's not that old. It's in fact someone who anticipated many things that we are constantly talking for about 18 months. And he's also one of those that, let's say is I like him because he's very balanced. For instance, and we are going to talk about this topic when it comes to structured data, he says, okay, structured data are great, but please don't overdo it. This person is Giorno Van Driel. Hey, Yarno, how are you doing?

Speaker B: Hi, Gianluca. Thanks for having me. Doing fine.

Speaker A: It's my pleasure. Things are you, uh, in rotten? I'm not, right?

Speaker B: Yeah. The Netherlands.

Speaker A: Okay. Okay. Okay. So how's the summer going?

Speaker B: Oh, finally got a week of less heat last week. We're like 34 till 36. Temperature dropped to 22. Yay.

Speaker A: That's cool. Well, and I hope it's not raining because the times I go there in the Netherlands, it's quite a weird type of weather. Sun, um, and all of a sudden it's a kind of British weather. So sun, rain, wind.

Speaker B: On average, a bit more sun than in the uk, but comparable.

Speaker A: Okay. And um, how SEO is treating you lately?

Speaker B: Uh, uh, seems like I'm most of all spending time having tons of conversations about the sense and most of all nonsense about structured data. Think. Because a lot of people, especially because of ChatGPT advice, have all kinds of weird ideas about what structured data is doing. So I'm having a lot of weird calls these days where in the past I used to have to convince people to start using structured data markup, and nowadays I'm spending more of my time talking the idea out of people's heads.

Speaker A: Yeah. Yes, that's what I was saying before. I remember you did a wonderful and also funny experiment was where you were substantially tagging with structured data, uh, almost every entity present on the page, if I'm not wrong, in a way to demonstrate that going that way. So over plotting your code with structured data information is not even the correct way to Use structured data. So no, if you would say in one phrase or, uh, short paragraph. Okay, first of all, what is the difference between structured data and schema dator? Because I feel that people still have not understood that there is a big difference between the two.

Speaker B: Yeah, well, first of all, the term structured data is already a, uh, very convoluted term because structured data covers a lot of things. It covers Excel files, it covers tables, it covers structured data markup, what we're going to be talking about. So structured data in itself can be many different things. When we're talking about structured data markup, we're talking about annotations made in a certain syntax. Now that syntax can be rdfa, microdata or JSON ld, whereas the syntax itself carries no meaning. So to convey the meaning of the data you're sharing, you need to use vocabularies or ontologies. And schema.org is one of those vocabularies out there that helps you to say this piece of data, uh, is about a person. He or she has this name. Listen, that city is the author of this article. So those are all kinds of things you can express by using the schema.org vocabulary because it contains all those terms that are machine readable, so to say. And so that's the biggest difference. When people talk about schema markup, that's yet another convoluted term because schema markup can also be a plain JSON file according to the JSON schema that exists. So the technically correct terminology would be semantic metadata. But that's a bit of a mouthful. So years ago the search engines decided to come up with a marketing term called structured data to make it explainable to the everyday person out there.

Speaker A: I can understand it because I mean, they need to talk because we usually think about it, about Google for instance, or even Bing, as talking to developers or even people who knows webmaster, what they were calling webmaster. But actually the most effort is the most important blog. For instance, on Google is, is the keyword blog, which is substantially talking to business owner marketing people, not digital marketing people. So they need to explain things to dumb down or simplify things in order to make them understood. And I love that blog because usually it's where you can find what Google really is adding to. Sometimes you find the most interesting tidbits and strategic idea that Google want to push in that block because it's talking directly to this business owner, not to us. So to us is talking about, okay, how, how you have to use vhreflag and Blah blah blah. But using structured data versus unstructured uh, data is the easiest way to, to make it also visualize things. And uh, it's fun that now another terminologies coming out. So it's something like structured data still versus prose. So prose use as okay talking about things without a real structure like in prose. Which is weird for me because it's somehow it's structural data. What is poetic? Because it's like that poetic, you know. You know it's time of Ulysses of EOD sales.

Speaker B: Now let's be honest.

Speaker A: The exam against the pros exam is structured and the pros is not in.

Speaker B: One could easily say that a lot of marketing content out there is actually just pros. It's a lot of blah blah and a lot of words without actually going into dep. It's trying to sell you an emotion instead of facts.

Speaker A: But I don't know if that could be even the finest pros. It would be just blabbing and or baroqueism. If you want to go talking literally literal terms. I think that I don't know if you agree with me. Don't you think that returning to Google that it is also somehow fault of Google this kind of misunderstanding of the difference between structured data, schema.org and eventually other definition, um, of same sample of you know, things that are inside the big bucket of structured data. Because if we go to the only existing information about from Google about structured data, it's the, you know, what was the structured data, the structured data uh, gallery that where Google is explaining how to use structured data in order to obtain rich result. And um, the fact that Google is talking of structural data just in terms of rich result even it puts something like structured data is important. Schemas is important for understanding the meaning blah blah blah. But it's just one line and then everything else is rich result. Don't you think that there is also this original scene from Google for making people especially in SEO to not understand it a lot?

Speaker B: M. To be honest, I actually applaud Google for what they have done throughout all those years because I think they have probably the biggest knowledge repository on how to use structured data as opposed to any publisher out there. There is nobody with so much technical guidelines and so much explanation in their documentation. And sure that has been a path. We look at the documentations of 15 years ago that weren't that good and there's a continuous evolution of that documentation. But I'd love to see do anybody out there do better than Google does actually. Is it perfect? No it's not perfect. Not even close. Yet again, Google tries to explain those parts, but actually matter to Google. So should they go as far as to explain all the theory around the linked open data, which is structured data markup, is based on. Well, not really. Uh, they try to convince people to use the things that are important for their search engines. If you're curious to learn more, their documentation points to every resource there is. You can go to the W3C specifications. I think the bigger issue, and that's the similar issue you highlighted in your recent document about open knowledge format, is that the issues actually lie with the people in our industry. People don't read the actual source. So the loudest screamers in our industry have been trying to sell us on structured data over the last couple of years because that's an easy sell and that makes them sound advanced. But the honesty is, is none of them have even spent a single minute over at schema.org's GitHub and read about all those terminologies, why they're in there in that vocabulary, what the use case was when we added those things to the vocabulary. I say we because I'm one of the volunteers out of dozens, uh, in there. Everything in there ended up in there with a certain use case in mind. But a lot of those use cases never materialized and then now suddenly there are a lot of people that say use all those terms. Yeah, those terms were added over a decade ago and never were used. So why would they be important right now? They're not.

Speaker A: Yeah. Yes. In fact, it's interesting how the this is something that I like to do every year when the web automatic, the chapter about the implementation of structured data again is evolving during years. And let's say I think that somehow there are certain types of schema vocabulary that are very used. But I think, and I think you agree with me that sometimes they are so used not because really web owners, masters, developers, or even SEO really thought about them, but because for instance, WordPress is the biggest and most used CMS. All the SEO plugins we are thinking from the classic Yoast SEO, but also ragmata if I'm not wrong. And they have their own section of structured data. So implementing schema tutorial by default according to certain characteristics. So that's why certain types, usually the types that Google recommends to use and implemented in the way Google recommends to implement it for rich results and they use Google Task for painting. Also SERPs are implemented by default. And so we see, okay, webpage, article, local business organization and so on.

Speaker B: Breadcrumb list. If the interesting thing especially about the web almanac overview is if you uh, what is it? Take the top 50 classes that are used, then you already got like 85% of everything that's been published. You take the top 50 classes and you remove everything that has not been part of Google's documentation. If you take out everything that has been a rich snippet is or has been, then you suddenly have nothing left.

Speaker A: Yeah, yeah, yeah, yeah. Ah.

Speaker B: What the web almanac clearly demonstrates is that 90% of the markup out there only exists because Google actually had a feature for it.

Speaker A: Indeed, indeed. And I think that sometimes there are things that Google should eventually and using the same philosophy, I'm not asking Google to change. Google should clarify maybe in very specific cases, I mean what I would like from Google but it's you know, you can understand also the implied business need of Google to explain very well things all the schema related to products. Mhm. Because schema related to products also are working together with the feed of merchant therefore and because this feed of merchant feed not only organic merchant but also Google shop. And it's so, so important for the business ecosystem of Google that Google is really explaining well, okay, you can use this for a variant, you can use this but for Sphinx, this other for aggregate offer. I would like sometimes the same attention to detail from Google for other things because if not we see people going very creative with schema.

Speaker B: I think the issue there lies with the people doing that and not so much Google's documentation.

Speaker A: Yeah, but you know for instance Google is telling us if you think for instance in the travel industry when we talk about the so called by Google carousel the item list is for European user still the beta for having the rich results with carousel of uh, the items which is valid for obviously retail but it's also valid for travel. So implemented correctly you can have the carousel m of hotels for instance where you are selling rooms, I don't know, in Madrid, in Rotterdam and in other places. But then Google when it comes for instance to hotel doesn't help. There is always this question coming up. I ah work a lot with travel. This question coming up my clients or SEO in house, all my clients asking me how can I indicate the offer of the aggregate offer because hotel formally doesn't have the offer. Then I uh, usually in this case I'm lucky because for hotel there is a wonderful, Wonderful guide inside schema.org website with all the indication for hotel and also for car dealers I think and precise indication on how to Use a combine. It's very old Stigulate using, not use indicating example with JSON ld. But it's totally doable. Uh, and totally follow up. But sometimes I think, for instance, for

Speaker B: software, this is an interesting one. Do you know why the documentation over@schema.org is so extensive? Because the entire hotel section was actually intended to be an extension. So there was a working group that actually created that extension for schema.org and that's why it's so in depth and why it so has contained such good examples because those have all been created by that specific working group.

Speaker A: That's great.

Speaker B: That's the whole idea behind Schema Dog. It's not, it's not supposed to be just a Google thing. It's supposed to be an, an open vocabulary.

Speaker A: No, no, no. I know, I know, but join. Yeah, but for instance, I would like a Google analyst to put forth a specific case of hotel. This is the link where you can find more information. Not necessarily. You have to. Because people, it's weird. SEO many times don't search.

Speaker B: No, that's my whole point. What you also made in your article is that all the information surrounding schema.org is out there. 95% of all the conversations surrounding the vocabulary, its sponsors. It's all there on GitHub on the W3C. But some SEOs just go through schema.org think they found the magic terminology copy that without even having a look at into why is it part of that vocabulary? Uh, if you don't understand the why behind it, how can you come to the conclusion it's valid just because the word exists? It's like taking a dictionary and giving every word in the dictionary exactly the same weighting. You have to use every word in a dictionary or else you're not talking correctly or you're not writing correctly. That's not how a dictionary works. And that's the Same thing for schema.org that's not how a vocabulary or an ontology works. You don't need to use everything. It lays out the options.

Speaker A: But I mean there is an interesting, uh, citing still Google. There is um, a somehow an exception where Google is saying, okay, I don't tell you what are the required properties or and the recommended property because Google does this distinction. In the only case that it does this, use as many properties as you can if they are present in the visible part of the content of the page. And um, that's organization. And it's interesting that Google is making this exception for organization of not selling practically Nothing in terms of recommendation, just telling you as a recommendation. Just use it once on your own page or in your about page. Where, where is the best page for. For using it Problem is that organization. Then it could be also useful to cite it with the id. A problem that id. For instance, Google doesn't explain what is the ID and eventually how can you use it to connect.

Speaker B: It shows the examples mostly around product variants. There it shows and return shipping policy and things like that. There it simply uses fragment identifiers in the ID to link things together. I think Google has good reason not to delve into the why and how that works because that's not the role of Google's documentation how ID is supposed to work and that that's sort of the difficult part about it. That's why you have the actual specification out there. Now people are looking at Google and not just your recommend, your comments. I've heard this comment throughout time. It's not up to Google to tell you what's already in the W3 specifications. If you want to learn HTML, you look at the W3C specifications. If you want to learn JSON LD, you look at those specifications. There you can go and learn how the syntax works. So I think it's. I get your point but to ask of Google to be the creator of a vocabulary and explain how syntax works and explain how they want to do you to do things. I think that's too much of an ask of Google. How can we if anything. Where's Bing? Bing has been signing.

Speaker A: Oh yeah, yeah, yeah, yeah.

Speaker B: A decade already.

Speaker A: Well, I don't know Yandex because also Yandex was one of uh.

Speaker B: Yandex is out of the picture. It's, it's close sort of within Russia. What used to be uh, it was registered in the Netherlands. The EU forced it to be sold back to Russia. It's in Russian hands now. And due to the war in Ukraine,

Speaker A: Yandex is out without a surgeon. Gene collaborated with Schema about being.

Speaker B: Originally it was Yahoo, Bing and Google. Later on Yex joined as well pretty fast. Baidu never joined from China but they sort of consume it in the background. So they're a silent consumer of structured data. Don't know if that's still valid these days. 10 years ago that was the case. And what you're also seeing in, in I think it's South Korea. There's a local search. Yeah, yeah. And they have a bunch of documentation themselves from what they consume. But if you look at that, that documentation, it's actually nearly A copy of Google's documentation. So now that Google for example officially has dropped FAQ page, I'm curious to see what those other search engines will do. But in all honesty, the main driver for more than a decade already has been Google for this. It's the one search engine that actually tells us what they want of us who has been actively developing new features. And no other search engine out there has done that. No LLM out there has done that. They're the only one actually producing documentation

Speaker A: surrounding structured data and talking about. You cited somehow as sort of recurring topic FAQ and this fun because it became all of a sudden so popular again because of LLMs and even if it was deprecated by Google I think in 2023 for everybody but very few

Speaker B: sites and now government sites and yes,

Speaker A: and now for the governments and the really really reputed F site is now also I think this is a wonderful example of how the cut off training data can produce disaster because maybe faq

Speaker B: okay, I'm not even convinced FAQ page did much to help to train LLMs. Reason being with natural language parsing for years already it has been so easy to extract FAQs because a very strict pattern.

Speaker A: I mean it's one of the simplest form of structured data in props. Yeah, because it's a question, it is an assignment.

Speaker B: Exactly. And um, um that's quite. It's been easy to extract that type of information already since late 2016. NLP was good enough to extract that without any guidance. I remember working on FAQ pages specifically for large e commerce parties before FAQ page markup was even live. Originally Google already gave perfect answers based on FAQ pages so why the heck they ever invented the FAQ page snippet? It's beyond me. NLP had no need and Google was already pretty good at answering questions before they launched faq. So why it ever existed is beyond me. If anything, if I think it caused a very negative pattern where everybody became lazy, stopped writing good product descriptions and just slap an FAQ on the page and I call that lazy.

Speaker A: Well it depends. I mean I can understand the need of using somehow FAQ for instance for some kind of product pages Product.

Speaker B: Oh, there are use cases for FAQs

Speaker A: definitely but okay, yes, but I, I understand whether there is sometimes I see also, you know, even articles that are substantially structured as they were FAQ pages M Just because there is this sort of new myth, let's call it so even if it's not really new myth of LLMs prefers one chunk to question and Also, but talking about myth and relation with LLMs and how people are trying to advise for feasibility in LLMs. What are the myths related to structured data? Schema and also other types you are seeing are becoming becoming dangerously popular.

Speaker B: Actually it's quite stagnant for a couple of years already. Google, with the whole shift going from traditional search results to LLM based search results, the problem has become that a lot of the original rich results do not fit into that AIO design answer pattern. Sure, we still have some traditional results if you're patient enough and keep scrolling. So the use case for a lot of the rich results that originally were designed is slowly disappearing. What we've seen over the last couple of years is that Google has been cleaning up its inventory of rich results because they lost their purpose. That's not a bad thing. That doesn't mean markup has become less important. It just means that that one specific use case has reached end of life. And at this moment the biggest use case there is probably has to do because everybody wants the agentic web to start working is product schema. And Google has been added for eight years now, approximately step by step aligning the merchant center specifications to which we constructed 2018.

Speaker A: Yes, yeah, roughly eight years because they started to show the rich results of popular products in 2018.

Speaker B: Yeah, there's a difference to what Google is doing in search and the discussions we've been having over@schema.org they're highly out of sync. I think. I started the discussion around product variants around 2017 and it took only like six or seven years to roll out into production. So it kind of happened that we been discussing things over@schema.org for quite some years already before you finally see it up. They show up in Google Search and especially around the product markup. That's not that easy because Google Shopping is a separate environment technically than Google Search is. So the moment they want to apply things from Google Shopping into Google Organic, they need to create new algorithms, new scripts, new testing tools, new testing features, new reports in search console, new result test features, and so on and so on. So it's, it's actually quite involved for them to align those specifications between organic search and Google Merchant center feeds. Especially because the format of a firm merchant sentence feed is very simple. It's a flat table, column rows and you know, structured data isn't that flat. It's a graph. So it goes left, right, up, down.

Speaker A: Um, yes, it's more three dimensional.

Speaker B: Yeah, exactly. So translating one into the other is not as always as easy as you think. And as they are creating new terms and new data shapes, we're also trying to evaluate, okay, this is what we got now, but we're going to create something new into the future. Are there any new things we need to take into account which aren't part of Merchant center either, but that need to be added to schema.org right now because we're doing that and later they show up in Merchant center specification.

Speaker A: Yes. And it's maybe more important, even more important now that Google started to anticipating merchant some, you know, properties related to a product that are not in the schema. And this is making the consolidation of the data, uh, more programmatic.

Speaker B: So it's, it used to be years ago, it used to be more problematic than it is now because I remember a time where we had different structured data models on the page Merchant center while at the same time providing a second piece of markup that was specifically for Google Organic. And since those two sites didn't communicate with each other, the markup for Organic was causing errors for Merchant center and the markup for Merchant center was causing errors in Google Organic. Those are really fun times. So every step they take I'm happy because that's one less obstacle, uh, in the way of actually doing it in markup.

Speaker A: But okay, returning to the, to the LLMs, not just Google, what is the biggest myth that you are seeing about the role of schema for LLM visibility?

Speaker B: The um, stubborn, persistent idea that you need to have structured data markup, machine readable markup so you can serve those LLM machines? Nope. Yes, it's true. LLMs are machines, but there are totally different types of machines. LLMs are machines that take natural language into account. That's it. And structured data markup is created for RDF based machines. So in an ideal world, structured data markup ends up in a RDF based knowledge graph, which is a fancy word for a database. It's just a database that works with triples instead of property value pairs. So that's the main difference between that and a regular database. But structured data markup is a markup that can directly be injected in that type of database. So it's an easy format to transport and move around. It's instantly injectable, instantly readable. So it's a very nice, clean and practical format. But LLMs don't need that format. They're not created to work with those formats. LLMs are created to deal with natural language where structured data fits in. And that's one of the things I'M very curious about Andrea Vampiri from wordlift is involved in. I was going to think, I think in August or September, something like that there's going to be a, a workshop meetup for the W3C and the GS one is in there and some giggle groups are in there where they're actually going to be talking about creating a standard so that agents can navigate knowledge graphs. In theory if you give an agent currently you ah have a folder full of skills you can relatively easily tell an agent here's the skills go travel that knowledge graph. The problem is is that there's no standards surrounding that. So you run the risk of every company creating its own skill sets, every company offering different solutions to get agents to navigate knowledge graphs. So what they're doing now is Andrea Vopini ran some experiments over the last couple of months which do definitely indicate that if you throw knowledge graph into the mix with agents and LLMs then up to 80% less error rate. That sounds interesting but for that we need to have new standards. And what I like about this one is that there are people involved from companies like Samsung and Siemens that are also going to be up for that debate. So this is not just a searching, this is an industry wide and personal possibly a commerce international new method of exporting and transferring data. That's going to be very interesting.

Speaker A: The lack of a standard or the fact that the standard are uh working are still in works in progress Because I'm thinking about well you cited before my article about your the open knowledge knowledge format which is not even a standard but there is for instance uh many, many SEO are talking as if was already established a standard for WebP and it's still not a standard at all.

Speaker B: It's just a suggestion so far yeah

Speaker A: I can see people on LinkedIn or uh, on a case or tools and so on especially recommending already the implementation of WebMC for everything which can be uh, of authentic use when for instance in my personal um point of view is okay test it but I don't know in a form that is not so so important in business level you can test it there and see how it works. If it's really working what are the results, how can you eventually use it eventually for other things in the future and stay always looking at how this, for this suggestive standard is evolving to a confirm standard. Set the standard but put all your eggs in that basket I think is still too early now to be honest

Speaker B: my opinion is very pragmatical about that. Unless you got a lot of Money to burn, stay away from it. Any form of agent optimization. If you don't have the money to waste, don't get involved in it. Let the standards develop. It's going to take quite some time. We don't even know which of the LLMs is going to survive the coming years. So to say, go after UCP and, and um, or, or agp, ACP or, you know, whichever standard or protocol you, you're thinking about doing, stop and ask yourself the question, will this platform still exist in a couple of years? Will ChatGPT actually survive? And again, unless you have a lot of money to burn, you don't really matter if you waste a hundred thousand here, left or right. It's probably best not to get involved in those technologies until everything sort of has been figured out over the next couple of years. There's a good chance a lot of those standards will disappear or will merge with other ideas. I think it's too early for a lot of businesses. If you're a small e commerce shop, for example, you're doing quite fine. I don't like the platform Shopify, but that's a different discussion. But you're quite safe sitting on Shopify. It works together with Google, it works together with Bing. They make it happen for you without you having to do huge investments.

Speaker A: Yeah, yeah, because Shopify for instance, the UCP is already existing for every, every shopping.

Speaker B: That's the same thing for the average local business. Make sure you got your Google business profile really sorted out and that you actually make use of it properly and all the features it offers. If there's anything agentic that needs to happen, Google will first plug it into business profile to test it. So instead of wasting a lot of resources for everything. Yeah, so instead of wasting a lot of resources on um, rolling out MCP and agentic standards and you name it, hold on. As a local business, make use of your Google business profile. That's probably where your return of investment lies. Anything else is for those that can afford for themselves to play around with it. What I'm actually expecting, especially if you look something like the open knowledge format. I've done a lot of use cases, worked on a lot of use cases in the past where structured data wasn't important for the public, it was important for the internal processes of that business. We had reasons to create a new database system, a new data warehouse, a new ERP system, and we based that all off on top of schema.org the biggest motivator were internal reasons. Making sure everybody in the business spoke the same business language, making sure that everybody followed the same sales rules and the same content, uh, guidelines, you name it. So we had all kinds of reasons to make that investment, but none of them were search. And I think that's for open knowledge format and for a lot of these standards that are coming out right now. I think the first use cases are mainly going to be internally and not so much the agentic web because the agentic web sounds cool but it doesn't generate all that much revenue quite yet. The majority of revenue still comes from traditional search.

Speaker A: And I think that business can be also extended to when we talk about knowledge graph in the terms of internal knowledge graph. I mean a good implementation of schema with the ID or with graph it's already somehow giving creating a sort of archetypical knowledge graph of a website. But when I talk about knowledge graph internal knowledge graph is I usually do this example that I did also with in tools. Let's say you are in commerce. So you use as a knowledge graph based of a knowledge graph your catalog of product and then create all the things related to each product. And my classic example are you know Star wars legend mini painting. So uh, for instance the miniature of Luke Skywalker is representative of of Luke Skywalker character of Star wars directed by uh, doing this is ah also a way it's internal creating an internal knowledge which can be used internally by an organization focus different types of things which can be from the simple business side. For instance with a knowledge graph I can understand if I don't know where I was in New York as best piece product or not, et cetera et cetera and how uh to collect it as many things but also in terms of freeing your mind and starting seeing connection between things. Also for the context I think one

Speaker B: of the coolest internal knowledge graph projects I've been part of where I I didn't do the technical side of things. I mainly guided every tons of different departments in the right direction and had to create some data shapes. But it was more I was the one running around in the business making sure everybody was kept up to speed. Uh, if you work for an international organization with let's say 15,000 people staff getting an up to date organogram out of a business is near impossible. But if you work in an International Organization with 15,000 colleagues finding the right one can be true. Hell, and they're an internal knowledge graph just for internal purposes that's up to date that knows which departments exist where which email addresses go with that department, which staff is Working for that department, how you can reach out to those people having that information up to date, it sounds like a no brainer but the majority of international large organizations it's a crime to get your hands on the right person. It's near impossible. So an internal knowledge graph can already be as simple as being able to generate an up to date organogram. I've used it also for uh, customer service department where they had a call registry system. We actually turned that call registry system into a knowledge graph so that it could not only serve customer service but that we actually could also serve the website. And we, we actually monitored the most asked questions and that determined the order of our FAQ pages because that would actually the questions coming in on the telephone and that was updated in real time. So no manual labor to get those questions up to the site. No question. Any answer was in the core registry system. So we made it available to the website. But for that it needs to be in a, in a knowledge graph format. So there are a lot of reasons why a knowledge graph can be very useful internally and actually can help companies make money without ever using it for search.

Speaker A: Yeah. Yes. In fact I think that was this case of FAQ for instance was the kind of uh, example I was using too when I was talking about combining creating an internal knowledge graph as a source of inspiration because doing so you can see the connection between things that are apparently not so evident. For instance in the mini painting you can see, you can really understand for instance with knowledge graph you can see the concept of the mini with the concept of a paint recipe. A paint recipe is what color to use to paint the mini in what segments like when you are cooking something. That's why prolific recipe and this is interesting because using the knowledge drives very easy to understand what are the possible recipe. And so because you know what are the recipes you can understand what kind of related colors you can put in the PDP of a miniature. In order. If you want to paint this miniature we suggest you this recipe. For instance, if you want to do Rebel Commandos in endor so some kind of tropical camo painting or I don't know scarif like rogue one tropical beach so sandy kind of camel. And this is a very practical use of uh. The problem is that creating knowledge graph it is even if it's form theoretically simple because they are triples and unfortunately it's so expensive because m. Well it

Speaker B: depends on the skill set of the developers involved and I think that's one of the issues the majority of businesses do not have somebody who is experienced with graph based information. The average database engineer does not work with graph based information. They work with property value pairs and tables still in these days. So A it's a lack of knowledge and actual developers knowing how to work with that. B majority of businesses do not have anybody internally champion a knowledge graph because that requires you to have somebody in house with enough understanding and division on um, how that could apply to a business and then fighting for it internally and going through it to actually make it happen. And to be honest there just not enough of us out there to make that happen. That's one of the reasons why adoption rates of knowledge graph has been lacking so so far is because there are not enough people out there who actually know how to do that. And um, here's the fun part, that's where LMS actually help close that gap. How if. If funny thing is that I see uh, in my LinkedIn feed I also follow a whole bunch of knowledge graph engineers and. And one of the earliest changes I saw is that in the past people were very busy writing uh, sparkle queries. Sparqle is the triple for the RDF variant of MySQL queries Spark also based on MySQL so to say. And they. Everybody was writing those queries manually and now they're all using LLMs writing human languages and the LLM writes the query for them.

Speaker A: Yes.

Speaker B: And more and more knowledge engineers are actually using LLMs solely and having the whole query technical backend being handled by agents and LLMs completely automated. Yes, it costs you a bit of tokens but at the same time uh, if you cannot find the stuff to do it for you, that's actually an easy way to open up a technology without having a huge new department to make it happen. It can be run with actually a few people and some people actually have some vague idea of what they want to do already get you really far with LLMs nowadays and I uh, want

Speaker A: to recover now one thing, sorry, turning a little bit back in time with you. I mean you have a personal website. I'm joking in. In the email I told you quite abandoned. Quite abandoned. Various because I was looking for things you wrote also in order to prepare the this episode and in your website first on the website there is only one article but it's very cool.

Speaker B: Yeah it contains about as much as the average website does.

Speaker A: So in this article where you talk about schema fairy dust which is a wonderful definition in relation to EEAT and before of a record we were talking about rail auto Back in the days which was considered maybe one of FWA and author. We end up at talking about the type schema author can be and now when it comes to for instance articles, blog posting and so on. Red author is also suggested by LLMs. It's one of the things that also LLMs one is faq, the other one is auto, which is also degraded by

Speaker B: Google for instance quite um, some years already.

Speaker A: And I'm bad. Where is the author? Author can be author, person, profile whatsoever, uh, this kind of thing, let's say author. And what are the structural elements on a page that actually remove the needle for a machine trying to calculate the trustworthiness of ah, an entity for instance which can be a person who is an author or a brand who is a publisher for instance. Okay, I mean I say business structured kind of infrastructure and elements and besides obviously many other things which can be backlinks, mention, positive mention and all these kind of things.

Speaker B: To start off here, we're getting a little bit into speculation from my end but I personally don't believe trust can be based on any single page.

Speaker A: Right.

Speaker B: That's one of the main points I have in, in that ridiculous long story of mine is that you should never trust what an offer says based on a single publication. That search engine don't do that either. They trust is based on um, aggregated information. The last person they trust is the author itself. That's the last entity in the chain they trust. It's if you want to build trust it's based on the aggregated information about an entity out there. And then in the end of the chain they look at what the homepage or the, the entity home of such a page says about that person. And if that information coincides with the information aggregated on the web, then it forms a confirmation layer. That's great because that gives your data, uh, the information your domain contains a little bit extra trust. Trust not as in I trust you with my wallet. Trust based as in data accuracy score. We can place a higher mathematical trust in the accuracy of the information you provide because generally you provide information that ah, coincides with the information the search engine finds on the web. So in that regard you cannot base trust on anything a person says. You put your money where your mouth is. If the Internet says it's the case, then uh, you're probably getting an all a good end on your way. Problem is people are smart. They know how to fake online profiles. SEOs have mastered it throughout time. There are plenty of examples throughout the years where people created entire Fake profiles just to elevate trust and and only in very few cases did that actually make money. Generally speaking, the majority of these things cost a lot of money with a very low return of investment. I don't say you cannot manipulate things, obviously everything can be manipulated. But the question always is, is the investment worth the effort? If you're looking at on page elements nowadays, I still think it's very important. Same as it was back in 1998 in the early 2000s when I worked on accessibility. Make sure you got proper semantic HTML. Why? If you got proper semantic HTML, there are gazillions of algorithms and scripts out there that can translate HTML into markdown. That type of information can be is being used by those same data consumers of your HTML.

Speaker A: And I think it's not a coincidence. It's not a coincidence that Google went for the agentic part which is important sometimes to be precise and say this is documentation that is presented by Google developers on the agentive side of Google or even Chrome and this is documentation prepared and presented by the search team. So to separate things the developers in the sense of the agentic side of Google and Chrome, it's not a coincidence that they talk about the accessibility right now which is something I mean I'm sure if I'm betting on how many SEO you are used to to include accessibility audit in their big 50 page long audit surely there will be one digit percent of doing also that and now it's, it's coming out people will say you have to do it as it was something new and we know it's a very painful. There are a few crawlers that are really helping you making this kind of audit. So for instance a very stupid example is when web designer decide whether what should be an editing is not an editing but a test in bold with bigger fonts. And this is a classic example which is going to semantic is disrupting eventually the semantic sequence of the page because it could be NH2, could be NH3 what it is.

Speaker B: It's just going to go full circle to how uh, I got into this stuff. When we talk about accessibility, the focus for SEOs 9 out of 10 times is if they even look at it. It's the technical side of accessibility. Are we using a link instead of a button? Are we providing alt text and descriptive alt text? Actually not just some automated thing that doesn't make sense for the image but that's all technicality. There's a whole separate world to accessibility which is much more interesting and that's writing for accessibility that's making sure that content is available for people who don't necessarily have an IQ of 120, who don't necessarily speak that language as their primary language for there are a lot of great immigrants in this world. So how do you keep a text easily accessible to read while still conveying the information it needs to convey? And that's something I almost never see anybody talk about in the SEO world. Yeah, I've seen some presentations about it throughout the years. One here, one there. I think once a year.

Speaker A: For the last example that you did is it's maybe something that international SEO are used to talk about because of localization. And so it's something that we are used to doing. One last question because you are someone who also a part of your daily work is this one. Um, you can read here International structural data concept. So let's do some, um, speaking. But also because you are really involved in the schema.org foundation as a whole. And what do you see? What kind of evolution are you going to see? I don't want to do 10 years, but in the next couple of years in everything structure that are going, is going, are going to be even more important or less important are going to be a bigger adoption of it. Something like this. What kind of evolution are you going to see in this? Let's call it the structure than a Semantic web landscape.

Speaker B: I think one of the things I'm being leaning towards over the last year, looking at schema.org as a whole, schema.org was designed for a different era. Schema DORG was designed as a vocabulary for search in an era where LLMs didn't exist even back then. If you look at about 2/3 of schema.org maybe even 75% of it was imagined. On, um, schema.org version 0.9, there was a very small group of people and one or two of each search engine that created their part of the vocabulary because they had to have, each of them had to have some input and they imagined what the future would need. We're now 15 years later and it has proven that predicting the future is very difficult even for those people who actually work at those search engines. And in practice, two thirds of the vocabulary ended up never being used. So I wouldn't be surprised if we see a reset the coming years. I don't necessarily see schema.org disappear, but I wouldn't be surprised if we see a new vocabulary pop up, one that better aligns maybe with the agentic web. I'm saying that without Seeing any obstructions right now nobody's. I don't see any conversations about people saying schema.org doesn't fit into the agentic web. I just think that the majority of it doesn't serve a real life purpose anymore. So I think we need more focus in the end describing everything with ontology most of the time serves internal purposes. But for search, search engines and LLMs want to rely on natural language. So the focus is going to be even more on specific UK use cases for structured data. And right now Google is busy enough with translating Merchant center into markup. I'm um, most of all curious to see what will happen when that exercise is done. Once we've got schema.org up to date to what's possible with Merchant Center, I'm very curious to see where the new things will come because by that time the whole agentic shopping should be crystallized as well. And then the question becomes which other types of services products will we see end up being part of the agentic web? We can fill it out easily. It's going to be travel, food is probably in clothing, all those those more or less fall on the E commerce but food probably has a place still in there as well.

Speaker A: Yeah as Google is always preventing a lot with recipe.

Speaker B: Yeah not but not only recipe literally food providing services, menus, things like that. Those are, those are still very agentic capabilities. Uh seems to lure around the corner for those. Yeah So I think the biggest where will we see the agentic web move towards And I think markup will follow that because it will serve that use case mostly.

Speaker A: I can see also another kind of use which is maybe I don't know the classic problem of helping disambiguating entities because even if natural language algorithms sometimes still fails about understanding but you are talking about entity A U N E it's written exactly identical as db and this is maybe because of the many things surrounding a written content that can happen on a website which can be maybe the context is hidden in a JavaScript that is not rendered so it's not seen and so there is no context and so on and so on and so on. So many sometimes fringe cases. So something like the maybe a pro a classic property one of the first for me one of the most important properties of schema like the CMS is going to.

Speaker B: Funny thing is as I okay maybe sound like I'm patting my own shoulder. I'm one of the people who actually sort of pioneered entity optimization and I'm going back to 2013, 2014 I was running a case study 20 trying to resolve issues with Panda when even Google engineers no longer had an idea why a certain website was being hit. They couldn't make sense of why it was going up or down. And then I got a request to do to try to resolve that issue with structured data. My working theory back then and um, we're talking about 2014 was Google is getting confused in the details. It's not clear enough to them that this topic is a variant of that topic and we're not talking about exactly the same thing on those pages. So I use structured data and entity optimization really to dive into what is this page about versus what is that page about? Ah, we're talking 2014. That was an era when Google engineers were also experimenting with markup. Why it was new. They were trying to see what they could do with the markup people were putting out there. So I think for a part that actually helped that website and that entire case study I presented around it back in 2015 because I was lucky Google engineers were playing around with this stuff. As of 2016, 2017, I already started noticing that the ROI no longer was there in that effort. Entity optimization, where practically speaking SEOs were are turning keywords into entities. That's all they're doing really go off NLP had become so good that it already crossed the 85% benchmark in accuracy. So you're talking about back in 2017 18ish that Google was wrong in 15% of the cases. Why was it wrong? Not because it was missing markup because the entity involved didn't have a good enough digital footprint to be recognized as an entity. And this is the same confusion I still see SEOs make out there where they go like entity optimization. Markup is needed to be very specific to explain what the topic is about. No more. Uh, it's not neither. What is it Google uses same as an identifiers to make sure that they understand they're talking about the right organization, the right place, the right person. They couldn't care less whether it's about checkers or chess. You know, if, if, if your terminology is not clear enough to explain whether it's puma shoes who Puma the animal. Then you got some work to do on your pros.

Speaker A: Yes, yeah, yes, I agree in that regard.

Speaker B: Yes, I'm marginally somewhere there in the background. It does help go recognize the entity you're talking about. I hate to use the word understand. So yes it helps them recognize. But if you need that little bit extra to help Google Understand and recognize what that page is about. You've got probably better things to work on than additional markup.

Speaker A: Yeah, yeah. Yes. And I think that entity recognition and understanding is also not just in the content itself on a page, it's also in how this content is linked and be linked by other piece of content inside the bigger growth website. Yeah, yeah. No, in one hour, six minutes, it's well, uh.

Speaker B: Oh, well, that went fast.

Speaker A: Let's stop it here. But before, when you are not thinking about structural data. Ah. And designing graph in your head, what do you like to do?

Speaker B: I'm a big fan of following certain twitch in YouTube streams and unfortunately because of my three fingers, there's the camera, I cannot use a game controller properly. So there are. And in my mind I'm a big gamer still. But unfortunately, physically I cannot play a lot of games. So I instead of gaming myself, I spend time following certain streamers and actually play games that I like as alternative for being able to play them myself.

Speaker A: Yeah, I understand. Well, maybe the future. Let's see. One of the evolution of a technology would be, you know, a classic. Just look using the highs and something else.

Speaker B: Yeah, but you know, there's a certain billionaire out there called Elon Musk that's working on a neurological interface. But I'm not sure how much I trust Elon Musk with my brain.

Speaker A: I'm not really a fan of a neuromancer perspective. Sincerely, I wouldn't uh, like to have a cheap. A private company in my head.

Speaker B: Exactly.

Speaker A: Okay, thank you, Jono. It was a real pleasure to have you here. Let's see, maybe in the future to organize something else as maybe a panel with other ontologists and lovers. Semantic search. Thank you again and you're welcome.

Speaker B: Thank you for having me and thanks

Speaker A: to everybody having watched us till now. Remember to subscribe to the channel surf session on YouTube but also on um, Apple podcast and Spotify and to ring on the bell so you will be notified for a new episode when Baker. Thank you and bye.

Speaker B: It.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • 107: Using Generative Engine Optimization (GEO) to Win the AI Search Race with Cole CaspersonUsing AI at Work · on Google Merchant Center71 / 100
  • How SoftwareFinder Grew 300% While Other Directories Died (The AI Search Playbook) with Adnan Malik, CEO of Software FinderSuperMarketers.ai: Your Roadmap to AI-Driven Marketing · on JSON-LD66 / 100
  • Google Merchant Center Updates: Big Changes for eCommerce BrandsEmail Einstein Ingenious eCommerce Email Marketing by Flowium · on Google Merchant Center65 / 100
  • Episode #110: Is AI really driving a search change across Shopify stores?Shopify with Milk Bottle Show · on Google Merchant Center56 / 100
  • The Brand Blueprint: AI Shopping Is Moving to the Cart: 6 Things E-Commerce Founders Must Fix Now The Brand Blueprint · on Google Merchant Center49 / 100
  • The Search-Driven Fashion Strategy That Helped Spirithoods ScaleeCommerce MasterPlan · on Google Merchant Center

More from The Search Session

All episodes →
  • Digital PR Done Right: Earned Media Strategies and Human-First Pitching | Britt Klontz62 / 100
  • Search Beyond the Blue Link: Agentic Commerce and LLM Readiness | Alex Moss
  • Technical SEO for the Agentic Web | Alfonso Moure
  • Scaling SEO: Crawl Strategies, Image Optimization and Internal Search | Roxana Stingu
  • Technical Branding and Video Signals for AI Search | Myriam Jessier
Explore the best B2B Marketing podcasts →
All The Search Session episodes →