
The Privacy Insider · 2026-03-18 · 45 min
Key moments - from our scoring
Substance score
53 / 100
Five dimensions, 20 points each
Philip Rathle, CTO of Neo4j, discusses how graph databases address critical AI challenges including hallucinations, explainability, and data governance. Neo4j - used by 84 Fortune 100 companies and downloaded hundreds of millions of times - stores relationships as first-class objects rather than forcing network data into relational tables. The conversation explores how graph databases enable customer 360 views that unify siloed identities across departments while respecting privacy controls, marketing opt-outs, and regulatory obligations. For privacy and compliance professionals, Rathle explains how graph structures solve the impedance mismatch between how business problems are intuitively conceived (circles and lines) and how relational databases force developers to work with them. The technology has become increasingly relevant as enterprises adopt AI and large language models, which require explicit context, fine-grained data governance, and traceable decision paths - capabilities that graph databases provide natively through their query language (Cypher, now standardized as GQL by ISO).
Graph databases address hallucinations, lack of explainability, poor data governance (knowing what data is appropriate for what purpose), and inability to access explicit context - all common issues when using non-graph solutions for AI.
Graph databases treat relationships as first-class objects with nodes representing things and typed, directional relationships with properties, whereas relational databases force network-shaped data into tables and require multiple complex joins to express connectivity.
By creating a unified view of customer identities across silos that sits behind a service layer, organizations can implement privacy controls like marketing opt-outs that bubble up appropriately without being over-applied, while still respecting consumer preferences about what data companies should know.
AI and large language models require fast-moving data, cut across silos, need query languages that both humans and AI agents can generate easily, and demand explicit context and fine-grained data governance - all capabilities graph databases provide natively.
Cypher is Neo4j's query language for expressing graph connectivity; it became the de facto standard for graph databases and has been formalized as GQL by ISO, the same standards body that created SQL nearly 40 years ago.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains several genuinely useful insights - attacker graph thinking in infosec, fine-grained property-level access controls for GDPR compliance, and blockchain KYC de-anonymization - but roughly half the runtime is career backstory, a prolonged Graph 101 explainer, and surface-level AI commentary that adds little for a technically literate operator.
attackers think in graphs. And so if your defender is thinking in lists or tables, then you're at a disadvantage
the model isn't the right place to do it in. But a lot of the rhetoric from the foundation model providers suggest that uh, well, I should just trust the model and it'll do the right thing
The 'use case network effects' framing and the one-link blockchain de-anonymization angle are modestly fresh, but the episode leans heavily on well-worn observations about social media surveillance, Cambridge Analytica, and 'when you're not paying for the product, you are the product' - the contrarian or first-principles thinking never really lands.
use case network effects, which is a term I came up with, just observing that if I sell for fraud detection, I might have most of the day that I need to do better recommendations. And then anti money laundering
things that you have probably assumed, and I've probably assumed up until now, that remain in a certain domain and can't be connected can very, very easily get stitched together and used
Philip Rathle is a genuine practitioner - 14-year tenure, first product hire at Neo4j, with hands-on enterprise data architecture experience at MCI and United Airlines - giving him real credibility; however, the conversation rarely extracts deep technical or strategic knowledge commensurate with his actual expertise.
I was brought in as part of a SWAT team to do some work at MCI back in the day...they had this massive marketing engine that marketed 300 million households per month which by the way is more households than exist in the United States
building United Airlines first customer360 system back when they had outsourced their website to a third party vendor
The episode has a handful of concrete numbers - 300 million households, a 28-hour batch job, 84 of the Fortune 100, a $2B+ valuation - but the most striking claim, 600% marketing uplift, is completely unattributed to any company or study, and the privacy and AI governance sections remain largely abstract.
a nightly batch job that took 28 hours to run
they had this massive marketing engine that marketed 300 million households per month which by the way is more households than exist in the United States
The host occasionally asks useful counter-questions (downsides of graphs, the 'do as I say not as I do' confession angle) but mostly accepts vague claims without pushback, lets the origin story run far too long, and relies on summarising/validating the guest's points rather than probing for deeper specifics.
What are the downsides? Uh, when we think about graph databases again I come back to this Facebook era where all we heard about was the knowledge graph
Is there anything that you have or any behaviors that you engage in that maybe don't live up to the best privacy practices in the world that you're willing to confess to us today
Computed from the transcript - who did the talking, and the words that came up most.
We sit down with Philip Rathle , Chief Technology Officer of Neo4j , to explore a question that’s becoming urgent in the age of AI: What happens when powerful models operate without context, governance, or explainability? As generative AI reshapes enterprise technology, graph databases are quietly becoming a foundational layer for accuracy, transparency, and data control. Philip shares why AI systems struggle without structured relationships, how graphs reduce hallucinations, and what this means for privacy teams navigating Customer 360, data subject requests, and regulatory pressure. Key Takeaways: 00:00 Introduction. 02:30 From chemical engineering to data architecture. 05:45 What a graph database actually is, and why it’s simpler than it sounds. 10:30 Why relational databases struggle with complex, connected data. 17:45 The AI tailwind: hallucinations, explainability, and governance. 23:10 Customer 360 and resolving fragmented identities. 28:15 Handling data subject access and deletion requests with graphs. 30:45 The double-edged sword: when graph power becomes surveillance risk. 37:00 AI models, privacy controls, and why not everything belongs in an LLM.
Transcribed and scored by The B2B Podcast Index.
Speaker A: The technology that we built happens to solve really the top problems that you end up with in AI applications when not using graphs, which are hallucinations, lack, uh, of explainability, lack of, let's say, data governance discernment judgment of knowing what data is appropriate to use for what purpose, and then have access to explicit context.
Speaker B: Hi, everybody. This is Arlo Gilbert, co founder and CEO of Osano, a leading data process privacy management platform. And you are listening to the Privacy Insider podcast. This show explores the past, present and future of data privacy for privacy and business leaders alike, as well as anyone who wants to keep privacy. Top of mind M. Hello, my name is Arlo Gilbert. I'm the founder of Osano, Ah. A leading data privacy company. Today, I'm your host on the Privacy Insider. We're joined today by Philip Rathley. Philip is the CTO at Neo4J. Neo4J is the leading dominant graph database. It's been downloaded hundreds of millions of times. It's used by 84 of the Fortune 100, and their last valuation was in excess of $2 billion. Philip has overseen the growth of this company from its early days. Philip, welcome to the show.
Speaker A: Thanks, Arlo. Pleasure to be here, Philip.
Speaker B: We have a lot of people on this show who represent the privacy side. We've had regulators, we've had technologists, we've had ethicists and philosophers. Uh, we've even had leaders of technology and politicians join us, but we've never really had a deep technical mind join us on this show before. You're working on really interesting technology that has a lot of overlap with personal privacy and AI. And, uh, we were hoping today that we might learn a little bit about some of the technology that you're using and developing. But before we start with that, we always want to know, who are you? How did you get here? I mean, you're the CTO of Neo4J, which is arguably the world's most popular and dominant graph database. And by virtue of that engagement, I'm imagining that you've got your fingers on the pulse of a lot of things that are happening. So how did a guy like you get to this place in life?
Speaker A: It really started with you, uh, know, first job working in tech as a consultant, and then somehow stumbling into data very early on on the data warehousing, but also the operational side. And, you know, some of my early experiences with data within. Within a few years into my career, I was doing data modeling and DBA work and eventually got into solution architecture kind of roles where I was sitting somewhere between the business and the technology, uh, which is where I've, you know, where I still like to be and uh, you know, and really seeing the power of data even going back a couple decades and how for example I was brought in as part of a SWAT team to do some work at MCI back in the day if you remember them. And probably those of you who are old enough like me, remember getting lots of telemarketing calls from mci. Well it turns out they had this massive marketing engine that marketed 300 million households per month which by the way is more households than exist in the United States. You know lots of business uh, implications around that like people getting multiple calls. And so that was an introduction to data quality and the reputational impacts and technical costs. And then underneath that uh, there was this batch job that was part of it that took nightly batch job that took 28 hours to run. So like all right, well there's this whole realm of it's a non starter for a nightly batch job to run 28 hours or even 24 hours, it probably should run maximum like three or four hours. And uh, got very much into performance scalability but, but also data quality and how ultimately the fuel for business capability becomes how well you understand your data. Another implementation I was involved with early on was building United Airlines first customer360 system back when they had outsourced their website to a third party vendor who then owned the profile and then they had their own mileage plus system. But they also were coming up with a initial uh, WAP phone and IVR and you know, how do you tie these together? So again there from an operational standpoint having systems infrastructure that would enable the business to treat customers as the same person regardless of the touch point with all the latest information. So that you're dealing with someone who just missed a flight, knowing that they just missed a flight and uh, like likewise on the analytics side for doing customer value, lifetime value calculations and that kind of thing. So yeah, lots. I fell into data really early on and stuck with it. Got into database tooling uh for about seven years, uh ended up heading up a product portfolio at a company called Embarcadero Technologies which had had and has one of the leading database modeling tools, er Studio, uh along with a number of tools for DBAs and developers, um, working with data and databases. And then uh, yeah 14 years ago or so now had uh, an opportunity to pioneer the graph category as the first product hire and product leader at Neo4j.
Speaker B: So I'm curious about two things on that um, first off, how did you end up getting connected with the Neo 4J team? I mean, there's usually a story in the early startup days of they weren't a thousand people yet, uh, they didn't have billions of dollars. What was that like?
Speaker A: There were 20 something people and they were. I met the founder, he was raising Emil, who's still our CEO, raising for a Series A. But, you know, as these things often happen, it was through the graph, through the professional graph in this case. And a friend of mine who worked with me in product marketing happened to know the first CMO at Neo4J, um, who happened to know that the CEO was looking for a first, uh, product leader. You know, 20 people. The product founder is still in charge of everything like product and engineering and so on. And, uh, you know, he's looking for someone to post Series A, hand off stewardship of the product, um, and of the product vision to someone who could lead that. And it was also grounded in what the world was doing and needed of a database platform. Um, so, uh, yes, stumbled into it really, just through knowing the right people at the right time.
Speaker B: That's amazing. It's always good fortune that turns out to be one of the predictors of amazing outcomes. Uh, and when you talk about your background, I mean, it sounds like there was a lot of data involved in there. What is it about data that you, that you love? Because there are a lot of ways that you could spend time as a technologist and product leader. Is it, you know, is there, is it, is it the structure? Is it the connections? Is there something about data that pulled you versus, I don't know, going into coding? Ah, you know, and being a developer or things like that.
Speaker A: Yeah, I definitely did some amount of coding in my time, but to me, um, I studied chemical engineering because I loved physics and chemistry and math and, you know, and the challenge and what I realized coming into software is in the same way that you're flowing some materials through a set of, you know, reactors and crystallizers and distillation columns and whatever else, and you're transmuting these, you know, these chemicals and their physical properties inside of each step in the enterprise. You have all these pipelines and what all the programs and code is doing is it's effectively transmuting and operating on data. So I saw the purpose of code, obviously this is just my perspective, and you could argue the opposite, but in my mind, the purpose of the code and the applications and everything else was actually to move data through so that you could then use that data to operate your business from like a store and retrieve OLTP standpoint and to make better decisions and then ultimately predictions and you know, we can get into AI from, from analytics, but, um, I really saw it as the core essence of, you know, where a lot of value then came from. Obviously you need to marshal up your data into applications that then will be used by users and create great experiences and so on. That's definitely a part of it. And that was, that's been the appeal to me of being in product management and technology more broadly. And then as technology moved on and we got into the early innings of AI, what previously was done as code got moved upstream and became a data problem. This is what supervised machine learning is. Instead of creating rules or speculating, I'm actually going to move that up, uh, into a data problem. So the data thesis and perspective I had in some way became strengthened by AI and then strengthened even more so as we get into generative AI, where all the magic that language models are doing just comes out of being trained on massive amounts of data. Now what we don't think about is that data is often curated. You've got companies like Scale AI and others who actually do all this manual labeling and so on. So it's, it seems like it's just masses of uncurated data from the outside, but there's still a little bit of a. Very much of a garbage in, garbage out. And so data still remains as or more important than ever.
Speaker B: More. Well, let's talk about data. It was 2013, 2014, uh, I was in San Francisco and, um, there was a meetup at the Medium headquarters in that triangular building. And I, uh, remember going there because I saw that there was a meetup about graph databases. And I was really curious about graph databases. This was a new technology. And there was a lot of question about whether graph as a technology was actually even a real database and a real thing that people would want to use in a commercial, scalable way. And I believe that early on, companies like Medium were some of the big early adopters of that technology. But a lot of people hear the word graph and their mind goes to the paper they had in school. You had your graph paper or they heard, uh, about Facebook building a social graph to monitor all of us. And I thought it would be really helpful if you might take our audience through kind of a 101 of what is a graph database? Exactly. And how does that work? Because I think most people understand broad concept, at least conceptually. They understand what a tabular database is, or a relationship database is. Right. It's a, it's like an Excel spreadsheet. And if you're familiar with pivot tables, you kind of understand joins. But beyond that, graphs are way more complicated and I think it would be helpful to understand that.
Speaker A: Yeah. And my uh, perspective is that are not more complicated, they're simpler. But, but let me, let me get to it and explain why really the idea with Neo4J and what kind of gave birth to the technology and the idea is a lot of the world just shows up as there are systems in the world that we need our software to be able to understand. And this is even more true of AI and those systems. How do they show up? Well, you could say that what gave birth to relational databases was the need for business process automation. And then what's the data problem underlying business process automation? I have data in warehouses and filing cabinets and paper forms. So paper forms have the characteristic of being human generated. They're pretty straightforward. You can easily come up with a set of rules to break them apart, remove your redundancy. These are all the normalization rules. And then to rehydrate it back into some form or subset of the form when I'm working through my system. So that's a bit oversimplified, but not far from the truth. The increasingly, as we've mastered that and moved on to managing, you know, having to deal with these complex dynamic, you know, world of business where everything's interconnected and with the build applications that do, you know, digital transformation and AI. And I need to deal with networks of people, networks of computers, networks of spread of ideas, of influence, of uh, geopolitics, of networks of payments, networks of biology, networks of ecology. A lot of this, a lot of things in the world show up as networks. Likewise you have networks that are maybe more up and down, hierarchically shaped or shaped like trees. This is like an HR hierarchy, which by the way is not strictly top down, it is from a reporting structure perspective, but in terms of the way the org actually operates. And you know, I have, I report into a project and um, maybe a project manager there and I have a mentor. And then there's the historical dimension and then there's my skills all around me and so on and so forth. But you could say that's, you know, broadly hierarchical permissions, asset, asset ownership, supply chain, these things are all like real world systems, digital world systems that are more hierarchical. And then you have paths and journeys through that, customer journey, patient journey. And so the observation that our founders had and that I really resonated with me as a thesis is look like we're taking all the stuff that shape like networks and hierarchies and stuffing into tables and there's a huge impedance mismatch, which is a jargon for there's a great distance between the way the data is shaped and shows up in the real world and even the way that a business person will whiteboard their domain. They always do that as circles and lines that are connecting each other, which is a graph. Um, and they're not going to say, here's my supply chain and draw out tables and then join tables and recursive joins. And that's just not how we intuit data. It's not how it naturally shows up. Not that you can't put it into a relational database. Anything you can put into a graph you can put in your relational database and vice versa. But uh, the observation was for a world of business that's fast moving, where I don't know from one minute to the next what problem I'm going to need to solve next to be competitive and where I'm next going to get and what piece of data is going to give you, me the most value. So you want a model where the distance between the business conception and a developer's way a developer works with the data is, has as little distance as possible between that and the way that data actually lives in the data structures in the database that's managing it as well as to have a model where ideally I should be able to add to my data without going through this whole schema migration exercise and spend multiple weeks or months, um, getting data modelers into a room to figure out how to adapt the model and then doing a whole development effort around how do I change my application now that my scheme has changed. So schema flexibility is another core principle here that is part of the core implementation in Neo4J and in many graph databases of course you have an option of then after the fact adding constraints once you understand the data and want to, you know, implement those at a lower level and then also having a query language that is based on understanding how things are connecting. So expressing connectivity in a query language so that instead of having like, you know, 50, uh, 100 line queries that are doing 15 way joins in order to do some supply chain query that's multiple levels out, or some social, uh, arbitrary Kevin Bacon type queries which are actually useful in many business contexts of uh, what's the shortest path between these two points in the network that uh, being able to express those queries and then being able to run them very efficiently. Those are all the, let's say core ingredients of what make graph databases different and also useful. So relational databases are super useful and obviously by far by 10x margin, the most commonly deployed with respect to all other database technologies combined. Notwithstanding any AI application that you build today, you're going to want to, and I'm sure we'll get into this, but deal with fast moving data, cut across the silos, have a language that both humans and models and AI agents can easily generate and express queries in. And it turns out uh, the graph database checks all of those boxes, um, particularly the way Neo4j has implemented it. So think of it as a, to bubble up, uh, a database management system built from the ground up for contemporary hardware. So where it's memory, abundant memory is fairly cheap, very fast storage substrate rather than spinning disk. And that is designed to store and work with and analyze but also to be the real time system behind agents for doing retrievals in worlds where you've got one or more of these um, complex kinds of real world systems, uh, that show up as networks.
Speaker B: So then the takeaway and what I'm hearing is that relational databases as we think about them, um, at their core they weren't really designed for, for relationships, they were designed for storing tabular data. Whereas graph databases actually were built by design to be about connecting the dots between different things.
Speaker A: That's right, yeah. The relationship is a first class object in the database and what that technically translates into is I have nodes which I can use to represent things and I have relationships which is like what's, you know, it's what it sounds like, it's how does this thing relate to that thing? And relationships have a type, they have a direction, they can have a number of properties. So you can have attribution of relationships which is great, like from a start date, end date level of certainty if I'm dealing with identity and so on and so forth. And this model has been you know, let's say blessed and you know, by the international Standards Organization and the same body that came up with the SQL standard nearly 40 years ago, like SQL 86 was adopted by ISO in 87. There literally is no other for 30 plus years there was no other ISO standard database model, um, or language to go with it up until the graph model came around. And now there's something called GQL which for all intents and purposes is more or less the same as Neo, uh4j Cypher query language, which was already the de facto language for graph databases. So it's something that has, you could say, your generational seal of approval, uh, from the key standards body that governs database standards.
Speaker B: That's great. Um, well, I'd love to talk about the privacy implications of that, but before we talk about that, I'm curious. With the growth of generative AI and, and all of this need to do lots of disparate retrievals, I'm assuming that this has been a bit of a renaissance for Neo4J. I mean, it was already a fast growing company, but has this AI wave really had a significant impact on technology decisions and customer buying behavior? Are you seeing a lot of AI first buyers?
Speaker A: It's been a massive tailwind. And of course, whenever there's a big generational, um, platform shift like this, you know, first of all, the pendulum swings over and everyone starts using just the one piece of new technology to solve everything. So LLMs. And then very quickly, you know, you discover, all right, what are, what's the right mix of new technologies? So, like, vector embeddings is another one with other existing technologies like graph databases to, let's say, balance and make, uh, create a, you know, one plus one plus one equals ten, let's say in the case of LLMs, vectors and graphs. So the, um, yeah, it's. The buying behavior has definitely shifted, I think, for everyone in enterprise tech very heavily in the direction of AI. And lucky for us, the technology that we built happens to solve really the top problems that you end up with in AI applications when not using graphs, which are hallucinations, lack of explainability, lack of, let's say, data governance, discernment judgment of knowing what data is appropriate to use for what purpose, and then being able to have access to explicit context.
Speaker B: So we talk about AI and let's shift over to privacy, because that is a big piece of anything that you build these days. I mean, if you're going to put the data in, then you now have all these obligations under various regulatory regimes that require you to be able to get the data out, to be able to cleanse it, to be able to control access to it. How are graph databases and Neo4j in particular? Um, how do those contribute to privacy? And how can privacy professionals think about this technology that for a lot of people feels very abstract? Uh, if you're not a database person, then you're probably not super familiar with it. And anything new or anything unpredictable obviously creates fear in the mind of people who are involved in risk. So how do People manage to Govern A, uh, Neo4J instance. How do they implement privacy controls on a Neo4J instance?
Speaker A: Let me start with customer 360. And unifying, uh, silos is a way to get into the answer to your question. One of the most common use cases that we've had over the years is how do you solve the problem which every company has that you have? Each department, uh, each division has one or more silos, you know, usually many, many silos that are either internally built applications or, you know, some third party app. And when you're engaging with the customer, what we all know as customers is we want companies to take into account the entirety of our relationship across all channels, across all the identities I've used, across all the channels, you know, phone, email, um, text, et cetera, social media, and across all the different departments and lines of business. And so there's always been this problem of, you know, how do you get to a single identity? And where Neo4J is used a lot is to say, let's take, create a graph that has all the different identities, all the identities and how they relate to each other, because you know what those are. You might have to go hunting in a bunch of different systems to pull it up, but then you can put those into a graph where I can say, here are all the different aliases for what ultimately is one customer, have a relationship there. And then for each alias, have a relationship with all the different identities, multiple email addresses that people have used, multiple, uh, roll those up to households. And this very messy structure that's really problematic to deal with in relational databases becomes actually very easy to deal with in the graph. And then from that point you can put the graph behind like a service layer that will call out spider across the graph, figure out what the identity is and then go. And if there's some payload you need to grab an existing system, you can all do that. And this ends up being a happy medium between the two extremes, neither of which works. One is let's just federate absolutely everything which doesn't work. Or the other is let's create one system to rule them all. And as the meme goes, then if you had 14 systems and you create the ones who rule them all, then you just have 15 systems at the end of it. What does that mean for privacy from privacy? There are things that I want as a consumer that you could say almost are in the opposite direction of privacy. But it's, I want the company to know all the things about them that I've told them that I expect them to know. On the other hand, I want to be able to make sure that the company is not using information in ways that I don't want it to, marketing opt out and whatnot. And so by having this unified view that sits over and straddles your existing systems, you can then make sure that if there's a, uh, marketing opt out over here, that that bubbles up appropriately and then isn't over implemented in cases where maybe I want to opt out of certain things and, but I really, really am interested in getting information about other things. Another thing you can do in the graph is there's information which a customer's disclose. There's also information which you gather by virtue of their activity, which might end up in web logs and so on where the customer hasn't explicitly identified themselves, but I might know who the person is or I might know, uh, if they haven't yet disclosed who they are, I might know that all this activity adds up to a certain individual based on looking at cookies, trackers, Mac addresses, IP addresses and so on and so forth and, and resolving those in some way. So um, obviously there are respectful and appropriate and legal ways to do this, but for me it's convenient as a consumer if as I'm engaging with a website multiple times, that they use what they know of my activity and serve up more and more useful information and they don't need to know my name to do that. And the kinds of things that you can do with graphs on the analytics side are to take all these breadcrumbs, do some fancy analytics, come up with a suspected identity, use it for marketing. And we've seen companies get literally as high as like 600% uplift in engagement, which is completely unheard of in the marketing world where usually like low single digits is um, the norm for a successful marketing campaign by using graphs in this way on the analytics side.
Speaker B: So when you use the graph in that manner that's being used, it helps to serve the end user, but it largely is helping the business. When you think about the ability for a consumer to opt out, to delete, to redact their information, we call those the kindergarten rules of data privacy. Right? If you want something from somebody, ask permission. If they want it back, give it back to them. And if they want to know where you're keeping it, be honest. Would it be fair to say that with a graph database it's a lot easier to be able to identify. Here's the information I have about you in for the purpose of reporting out. You know, hey, I, I happen to have information about your buying activity here. I might have some behavioral activity or uh, here. And because it's all in a similar graph or it's, or it's connected it's far easier for me to be able to surface the answer to what data do you store about me?
Speaker A: Yeah that's right. And that comes up in. There are two situations where you need to do that. One is people exercising the right to know what information is being stored and then the second is right to be forgotten. And these both can be very expensive exercises. If you don't have the right data infrastructure and if you have a graph then you can essentially have the structure through which you can easily like just work your way down the graph. Just follow the dots literally and follow the path into what data is stored at. Ah a different level. So that's definitely a use is responding to those requests in a way that's faster and you know, let's say in line with compliance timelines and faster for the customer and so on. But is way, way, way, way cheaper. Now I won't say that that's necessarily the highest value use I think the higher but it's a super yeah is one of the uh, value added uses. I'd say the highest value is making sure that you are using either your data appropriately across channels based on what customers have requested. Respecting this is less privacy but kind of in the same ballpark of what enterprises need to do respect the organizational uh, firewalls. So in banking for example the investment bank can't uh use certain information from corporate banking for regulatory and anti competitive reasons. Um the implementation is the same is understanding what data at a fine grained level can and cannot be used by what party, for what purpose for what person.
Speaker B: Now conversely so those are great examples of the pros of the graph. So marketing teams can quickly get better uplift through better identification of properties and users. Sounds like there's some real pros to the privacy side of being able to do the same thing. For the purpose of understanding what you store about me. What are the downsides? Uh, when we think about graph databases again I come back to this Facebook era where all we heard about was the knowledge graph. And so the word graph itself has some connotations around societal profiling and things like that. Are there any downsides or any challenges that you think graphs introduce into the privacy equation?
Speaker A: I don't see downsides insofar as good actors being using the technology as a tool because what you can do is actually very rich and I'LL add one more capability, uh, that I missed is you can add weights to the relationships in the graph that reflect your level of certainty that this person is this identity or is this other identity or this thing happened or this didn't happen. So that gives you a very, very rich set of tools. Then what the risks become is bad actors using this technology, um, which then means you need to defend against it. So there's this saying on the infosec perspective, uh, cybersecurity that was coined by um, someone at Microsoft years ago, which is attackers think in graphs. And so if your defender is thinking in lists or tables, then you're at a disadvantage. Like it's you know, bringing ah, a knife to a gunfight kind of thing. And so what are some of those things, um, people who are doing. Again this is outside the realm of privacy, but I think the listeners will get the analogy. What is money laundering? Money laundering is nothing but sending data from multiple places to one through many intermediaries and taking advantage of the fact that the systems that most companies have, at least up until pretty recently, knew nothing about being able to understand paths across multiple intermediaries. And so that ends up being a big gaping hole for someone who takes a graph perspective on things. Likewise from a privacy perspective, uh, a good example is blockchains, like the Bitcoin blockchain for example, where everything is public, all the transactions between wallets are public. Of course the average person doesn't know who these wallets belong to, but the reality is you could have a million transactions and they're all anonymous. But then the second you need to pull money out of a financial institution, your account's been kyc'd and all it takes is that one link to tie you back. And now I can spider through the graph and all this million transactions become very well known and understood.
Speaker B: Fair enough. So graphs give you superpowers and you can choose to use them for good or for evil. And it does sound to me like the scary part of a, ah, well orchestrated graph is getting into the hands of a government who now all of a sudden can explore my shopping habits and who I talked to and how I voted in the last election and what I said on social media. And so there are definitely some, some scary, scary outcomes. It's not necessarily the graph database itself that caused that, but the graph database does enable that discovery at a new level.
Speaker A: That's a good thing for people to keep in mind because to the degree that our personal information, our customers, our families is exposed in the public world you can have an entire graph of social activity and that's in its walled garden until you have just the one link that connects it into some other data set. Uh, and now all of a sudden, I know all this additional information. So the positive view when you're using this technology for good is that's amazing from the perspective of data network effects and actually use case network effects, which is a term I came up with, just observing that if I sell for fraud detection, I might have most of the day that I need to do better recommendations. And then anti money laundering and entirely unrelated things. This is the beauty and power of the graph. On the other hand, what it means is things that you have probably assumed, and I've probably assumed up until now, that remain in a certain domain and can't be connected can very, very easily get stitched together and used. And so activities that seem very, very remote and end up being, uh, end up being very transparent to companies. So one example is, you know, there's this ongoing debate of, you know, are my iPhones and my apps on my iPhone listening to me? Because fake Facebook just, you know, served up this ad. And it's exactly what I was talking about over dinner last night. And one of the techniques that companies like Facebook use is, okay, let's say, let's assume they're not listening on the phone, which supposedly they're not, is they know, based on location, sharing information, that I was at dinner with the other people who were at the table. And if one of those people during dinner searches for a particular thing, then all the people who at dinner presumably are good to target with that particular thing, especially if that thing is a product that is an advertiser. So that is 100% happening. And, and that's something that listeners can take advantage of and guard themselves against as appropriate.
Speaker B: That's right. We always joke in my house when we, you know, we talk about something in public or in front of our, one of our smart devices that, you know, our ads are going to change real soon. So as, as we think about the world of technology, do you hear much? I mean, there's a little while where privacy was really a top topic for a couple of years. It was in the news constantly. A lot of that came out of the Cambridge Analytica scandal. Do you feel like the, um, I mean, and you're deep in the heart of Silicon Valley building a very powerful tech company. Do you feel like privacy is something that is being considered by many technologists? Is it being kind of put off for another day? How does that conversation rise to the
Speaker A: rise to you, I say it's somewhere in between. It's a really big deal in Europe. Um, in fact last year I was on a panel at Viva Tech with the um, Anu Talis who's uh, chair of the EU Data Protection Board. And you know, an AI is a big area of concern because the assumption, I think both rightly and wrongly, is that every model is getting trained on everybody's information. I think the nuance there is foundation models have access to a certain amount of things and then companies aren't necessarily training their data in models or trying to solve privacy sensitive problems in that particular way. But there's definitely the fear of this among regulators, uh, and among the populace. And uh, so the answer from that perspective is yes, very top of mind. I think where it becomes top of mind among execs that I've uh, met with is if you do, you know, one of the ways that you could get models to understand your business more is to train them on your data. And then, but then if you're training them on your sensitive employee data or sensitive customer data, then anything the model's been trained on is fair game. And so people can easily use this to deliberately go and mine information about other people and violate privacy laws and norms. And so therefore the model isn't the right place to do it in. But a lot of the rhetoric from the foundation model providers suggest that uh, well, I should just trust the model and it'll do the right thing. And so to me the counterbalance to that is don't trust the model with everything. Like look, we do have technology that can provide fine grained access controls. And in the graph you have very fine grained access control down to property level, which is kind of like column level, as well as controlling whether a relationship can be traversed or not. So knowing that these two things are connected versus not, or knowing things about the relationship or even down to is this individual document, um, should it be accessible to a person based on say their clearance level versus the classification level of a particular document? So we definitely have tools to do this. Those tools don't exist at all in the LMS themselves. And that's fine. That's just how enterprise tech has always worked is let's use a combination of technologies and AI therefore becomes a um, composition problem of what's the right set of technologies to use and how do I delegate my different concerns. So uh, graphs play the role of, in cases where I need more accuracy or even 100% accurate and explainable Answer that can respect all the different privacy rules. Graphs are super well suited to that when used as a knowledge layer that works as part of an AI system.
Speaker B: Yeah, it's interesting, you know, we hear the fears about AI traversing our data and it's a conundrum because in some ways these LLMs are capable of far more than we might assume they can do. But at the same time, we tend to ascribe a lot of power to these models, assuming that they have tools and data and capabilities. And the general public often doesn't realize that when you talk to ChatGPT on your consumer app, you're not just talking to the model. You're talking to a model that is armed with access to lots of different things. And so the model itself isn't the dangerous thing any more than the person is the dangerous thing. It's what does that model do and what did the people who decided to use the model, what did they give that model access to?
Speaker A: 100%. Yeah, models use tools and the foundation model providers in their consumer products like ChatGPT, Gemini, Claude and so on. Perplexity can make call outs. They can make call outs to the web. Um, of course in an enterprise context, they can call out to tools via mcp and that can include graph databases or any database for that matter, any microservice and you name it.
Speaker B: Well, we've talked a lot about privacy on this show and a lot about graph databases and I'm curious, we all support data privacy and we, we try our best to adhere to best practices. Uh, the graph database sounds like a powerful way to be able to exert some controls with those governance controls you were talking about. But sometimes it's a case of do as I say, not as I do. Is there anything that you have or any behaviors that you engage in that maybe don't live up to the best privacy practices in the world that you're willing to confess to us today?
Speaker A: Sure. So I have a couple, but let me. So one of course is recognizing that when you're not paying for the product, you are the product. And you know, this is all the different social media platforms and you know, being, being judicious, judicious there. But I'll call it the one that actually impacts my life more viscerally, which is giving out my phone number because I, I have a cell phone, I just have one, I use it for business and personal and I travel a lot and early on I had, you know, business cards that I would give out very liberally that had my phone number on them, and those would end up in some database and I'd end up with many, many calls per day from vendors selling all kinds of just random things to me at any time of day or night. Because I'll be in Sydney expecting a call from a driver at 5am as I'm trying to get to the airport. And it turns out it's someone from the Bay Area trying to sell me some router or some random thing. Um, and I've still not learned my lesson and, you know, taking the time to work out a set of virtual numbers that I, you know, some for more time use and not give them out at conferences. So I'm definitely behind there. If anyone has any good tips, please reach out to me.
Speaker B: Um, just don't reach out by phone.
Speaker A: Just don't reach out by phone. No, actually this, this call I would welcome. Yeah.
Speaker B: And, you know, it's, it's funny you mentioned the phone and immediately takes me back to the examples that you're using of. We know all this information about you here. We know all this information about you here. And it just takes one piece of knowledge to be able to connect those. And that phone number might be one
Speaker A: of those pieces, 100%.
Speaker B: Well, Philip, thank you very much for joining. And folks, uh, if you'd like to learn more about graph databases, I'd encourage you to go over to neo4j.com, that's n e o the number 4j.com, and, uh, on there they have some great documents. Philip, uh, has been part of building up a thing called the Graph Rag Manifesto, and they also have a graph academy there. Uh, I would encourage everybody, even if you're not intending to use a graph database. This is a seminal technology that has taken over our planet in many places that you don't realize it. And having an understanding of this technology is going to be critical to being able to succeed as a privacy and governance professional if you're responsible for making sure that people govern these databases appropriately. So, Philip, thank you for joining us today. This has been an illuminating conversation.
Speaker A: It's been a pleasure. Arlo,
Speaker B: thank you for listening to this episode of the Privacy Insider podcast. You can find a full transcript of this episode and any show notes@osano.com that's www.osano.com. and while you're there, get access to an excerpt of my book, the Privacy Insider how to Embrace Data Privacy and Join the Next Wave of Trusted brands, which is now available on Amazon for purchase until next month. Take care. And remember, data privacy is a fundamental human right, y'.
Speaker A: All.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.