Private Equity Data Guy · 2026-03-12 · 49 min
Key moments - from our scoring
Substance score
68 / 100
Five dimensions, 20 points each
This episode explores the hidden cost of fragmented data systems in mid-market PE portfolio companies through a concrete case study: a scaling business that grew from $3.5M to $21.6M revenue but maintained disconnected product databases, CRM systems, and financial records. When they entered exit processes and needed a unified profitability report, three separate sources yielded slightly different revenue figures - a seemingly minor discrepancy that created massive deal complexity and uncertainty. Greg Hood, who built SkySight Analytics after 20 years in finance and data at fintechs like Kunai, Paramount Commerce, and Qtrade, explains why this pattern is endemic in mid-market companies. Most firms deprioritize back-office infrastructure while scaling sales and product. By the time PE arrives, they're running on QuickBooks, thousands of manual journal entries, and cascading Excel macros. Hood walks through his approach via ScalePath: migrating to proper ERPs (replacing QuickBooks for multi-entity operations), building finance data warehouses on modern infrastructure (Snowflake, Databricks, AWS), automating journal entries with APIs, and implementing tools like Campfire. The conversation emphasizes that technology isn't the bottleneck - process discipline and governance are. Financial close speed is presented as a leading indicator of overall data health.
Fragmented systems where product data, CRM records, and accounting records live in separate databases without a single source of truth. During exit, when these systems are reconciled for due diligence, small discrepancies (like revenue differing by $150k across three sources) create material questions about data integrity that buyers exploit.
Because QuickBooks was perfectly adequate when the company was smaller, and scaling priorities correctly went to sales and product first. Back-office infrastructure gets deprioritized until the company becomes multi-entity or multi-currency, at which point QuickBooks becomes a bottleneck that's expensive to fix.
Typically 80-90% of journal entries are done manually through copy-paste workflows, cascading Excel macros, and multi-step pivot table transformations across disconnected systems.
Month-end, quarter-end, and year-end close speed is a leading indicator of data governance quality. Companies closing in 2-3 days typically have better data fundamentals than those taking 6-8 days.
No - upgrading from QuickBooks to NetSuite or implementing Snowflake without first defining processes and data governance simply automates the broken workflow at higher cost and speed.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode delivers substantive insights on data infrastructure problems in mid-market companies, particularly the $2.8M valuation loss example and dirty data discount findings. However, significant portions consist of career backstory, AI tangents, and conversational padding that dilute density. Core insights about multi-database reconciliation failures, process-before-technology approach, and data monetization limitations are valuable but interrupted by meandering discussions.
Great success story, going from three and a half to seven to 14 to 21.6. But what's happening is during the exit process, they needed a product profitability report. Well, the problem is like most firms, they hadn't actually grown the back office. Their product set in three databases. When they pulled that together, it's a 21.15 million revenue. When they went into their CRM, they pulled $21.3 million of sales.
We have something we call the dirty data discount. It's usually worth about 10% of the valuation. We did a survey about 120, 122 different advisors and kind of basically said, if we gave you this data, we gave you this data, what would your valuation be? And you know, we got anywhere between a 1% discount and a 22% discount.
The core argument - that data infrastructure is a constraint in PE-backed companies and that process precedes technology - is sound but not novel. The dirty data discount quantification (10%) is useful but derived from a single survey. Most frameworks are standard industry thinking: fix processes before buying tools, data quality matters for exits. The data monetization section and AI integration discussion add some fresh angles, but largely recycle conventional wisdom about data assets and model risk.
You can drop a netsuite or you can drop Snowflake into a broken process and all you get is more expensive broken process.
Behind every value creation plan, there's a data problem nobody wants to talk about. Fragmented systems, metrics nobody trusts. And decisions made on gut feel dropped, dressed up as analysis.
Greg Hood is highly credible: 20+ years as finance executive in fintech/financial services, CPA, CMA, founded SkySight Analytics, held CDO title. He brings genuine operational depth and has actually built data warehouses at scale (Qtrade, others). His consulting work across 15-20 portfolio companies adds real-world credibility. However, he is now primarily a consultant/vendor rather than an active operator, which slightly limits caliber versus an active CFO or CEO managing live transformations.
Greg Hood is a CPA and CMA who spent over 20 years as a finance executive across Canadian and US fintechs and financial services, including roles at Kunai, Paramount Commerce and Qtrade, an online brokerage where his team won most innovative finance department.
As a company we've been through about 15 to 20 different scaling ups so far. Right. But we bring in the expertise that your head of finance just may not know or your head of data may not know.
The episode anchors arguments in concrete numbers and named examples: $2.8M valuation loss case (21.15M vs 21.3M vs 21.6M), 10% dirty data discount from 122 advisors, Qtrade's 100M company with 90% month-end close by 9am on day one, $800K pharma dataset example, commission reporting reduced from 2.5 days to under 2 hours. The QuickBooks-to-NetSuite mistaken recommendation (Claude error: $90K spend) provides a specific failure case. Some discussions lack specifics (AI model drift percentages mentioned but not deeply quantified), but overall evidence density is strong.
There's a half million dollar discount, but it's not really half a million dollars because it's a 5.8x on multiple. So you're talking about $2.8 million of lost revenue of lost valuation because you didn't have the correct reports in place and you couldn't prove your 21.6 million of revenue.
We took a two and a half day process and turned it down to under two hours where our competitors were still copying and pasting spreadsheets to do commission reporting. We had a fully vetted, double, double checked two and a half hour process and that in the last hour and a half it was just a manual intervention of just a human actually take a look at everything, make sure that the output was right.
Host Graham Crawford asks decent setup questions and shows domain familiarity, but rarely pushes back or challenges Greg's claims directly. Follow-ups are mostly confirmatory ('yeah, no, absolutely'). The AI debate section meanders without sharp interrogation of the Claude NetSuite error or its systemic implications. Host pivots to personal anecdotes (Commodore 64, soccer simulations, location data) rather than extracting deeper expertise. No instances of productive disagreement or probing into contradictions (e.g., if data monetization ROI barely breaks even, why pursue it?). Conversational flow is friendly but lacks edge.
Yeah, I love it. And actually that metric you mentioned, if you're looking for hidden indicators of how good a company's data is, like how long it takes them to close month, close quarter, close year from a financial perspective, is usually a pretty good barometer of what you're going to find in the other areas of the company.
So I thought it was a 17 legal entity. So they was recommending spending 80, $90,000 on NetSuite when the problem wasn't QuickBooks. The problem was all the information going into QuickBooks, not QuickBooks itself. The $225 per month subscription to NetSuite, to QuickBooks was all they needed, not a $90,000. But if you listen to Claude, he would have said, hey, you need to spend $90,000 on NetSuite.
Computed from the transcript - who did the talking, and the words that came up most.
Greg Hood has spent over 20 years as a finance executive inside Canadian and US fintechs and financial services firms. He built Sky Site Analytics to help PE firms, sell-side M&A advisors, and high-growth companies fix the data layer before buyers find it first. We cover the dirty data discount, what actually breaks during exit, and why the back office is always the last to get the budget it needs. If you have ever sat across from a buy-side team watching trust drain out of a room because three reports show three different revenue numbers, this conversation is for you. Greg and I have worked in the same trenches long enough to know that the data problem is almost never a technology problem. It is a process problem wearing a technology disguise.
Transcribed and scored by The B2B Podcast Index.
Great success story, going from three and a half to seven to 14 to 21.6. But what's happening is during the exit process, they needed a product profitability report. Well, the problem is like most firms, they hadn't actually grown the back office.
Their product set in three databases. When they pulled that together, it's a 21.15 million revenue. When they went into their CRM, they pulled $21.
3 million of sales. Well, 21.15, 21.3, 21.
6. Well, when you're looking at an exit, you're telling a story. You're selling your story and the numbers. So it creates if you don't know what you don't know.
That's why like a scale path makes sense. Because as a company, we've been through about 15 to 20 different scaling ups so far, right? But we bring in the expertise that your head of finance just may not know or your head of data may not know. They were a business analyst five years ago.
Now they're the head of data of a company that's three times the size, right? It's just bringing in the experts who have been there to help you, not to replace you, but just to help you grow. And everything worked. I'm like, what did you just do?
He goes, oh, this is SQL. The pages three, page four were pulling from different tables. I was like, I gotta learn the SQL thing because you solved in 20 minutes what I was spending two days on. That's what started my career down into data.
Behind every value creation plan, there's a data problem nobody wants to talk about. Fragmented systems, metrics nobody trusts. And decisions made on gut feel dropped, dressed up as analysis. Welcome to the PE Data Guy.
Each week, host Graham Crawford talks to the operating partners, advisors and practitioners who are doing the work inside portfolio companies. If you care about what actually drives returns in the market, then you're in the right place. The PE data guy starts now. Greg Hood is a CPA and CMA who spent over 20 years as a finance executive across Canadian and US fintechs and financial services, including roles at Kunai, Paramount Commerce and Qtrade, an online brokerage where his team won most innovative finance department.
Greg has the distinction of being one of the first CPAs to hold the title of CDO. Along the way, he left the exact problems that most PE backed portfolio companies are still droning in fragmented data, disconnected finance stacks and leadership teams making decisions on gut feel because the numbers are just not trustworthy enough to act on. Greg took that experience and built out SkySight Analytics, a Toronto based consultancy that works with PE firms, sell side M and A advisors and high growth companies to turn finance and data infrastructure into something that actually supports value creation, exit readiness and data monetization.
Greg and I operate in the same world. We see the same problems and often arrive at the same conclusion. The data layer in most mid market PE portfolio companies is the single biggest constraint that nobody wants to budget for. So thanks for joining me Greg.
Great to have you on. Good friend and thanks for coming on the show. I'm super excited to be here. I love this topic.
Of course, you know this, you know me. Yes, yes. So we have, we can talk for hours. For hours, for sure.
Is it, is there anything you'd add that's quite a cool winding road of a career that you've had with lots of really exciting stocks along the way. Is there anything you've. You'd add? Any favorite things you've worked on?
The very favorite thing I've worked on, of course, is all the finance data warehouses we built out. And you know, we've seen the evolution of data. Like when I first did my first data warehouse, it was 2010, it was on prem, the technology was completely different and back then it was really weird for like what's this finance guy doing with the data stuff? Get away, go back to the finance team, get back to your spreadsheet, get back to Excel.
Obviously over the last 16 years now we've seen quite the evolution of where data's come and how much it actually impacts finance. And it's not just something that's siloed away in the IT department anymore. Yeah. And finance is one of those areas where it's critical that it's right as well.
Of course. And I don't want to date myself, but same type of thing. I remember the first warehouses I was building and we had to build a six week window into the start of the project plan to allow for the servers to be ordered from HP and then delivered on site and installed in a data center. So those were the days.
I remember walking into the room and we had to make sure the air conditioner was working because we didn't want it to overheat and all that good stuff, right? Oh yeah. The size of the stacks used to be taller than me, Right. As well.
It was, yeah, it was quite the, it was quite the thing. What, what do you think drew you to data? I mean you've like. I've spent a career in data as well and I think I remember when I was a kid like simulating soccer matches with my stuffed toys and building league tables all along the way.
Like I've always kind of been drawn to and had an affinity to it. Like what, what do you think led you to, to the world of data? Well, I started off like, I actually originally wanted to go into comp Sci but I kind of swapped over to go back over business but you know, to date myself, you know I was, I learned, wrote my first piece of code back in like 1981 on a Commodore 64. So I've already kind of had the affiliation for like computers.
But really it was, I had an aha moment at work one day. We had this report, 10 page report. We had a number on page three, a number on page four, number on page five and the third decimal point was different between pages three and four, but was the same on the page five. I was going crazy.
I kept on trying, I'm in a reading room, I've got highlighters, I've got stuff on a whiteboard. And I became the crazy mad scientist in this room. Two or three people came to join me. We're all trying, doing one penny entries into the accounting system to try to figure out why is this report, what's wrong with it.
So on day three, some other guy came by us, he's like, what are you guys working on? So we explained to him, he kind of did one of these, walked back to his desk, we sat at his desk for 20 minutes. He's like check tonight. And everything worked.
I'm like what did you just do? He goes oh, this is SQL. The pages three and page four are pulling from different tables. I was like, I got to learn the SQL thing.
You solved in 20 minutes what I was spending two days on. And that's what started my career down into data. Yeah, that guy knew where the well governed tables were for sure. Yeah, I love it, I love it.
And again I think we have similar answers here. But for you, what led you into the world of working with PE and PE backed companies? So I came up to the big bank world, spent the first six, eight years in big large financial institutions. Great trading arm.
Love my time there, but I also realized it wasn't me. So about 10 years of my career I made a swap over into fintechs. Really found that mid market love and then from that point just kept on realizing there's a lot of potential mid market. Mid market is traditionally dirty data.
The processes aren't quite as refined as you get a bank, but you really have an Opportunity to make an impact in the lower mid market. The mid market. And from there that led to three successful exits. So I kind of got bitten by the exit bug.
Ah, there you go. Yeah, understandable. And addictive stuff that I'm sure. Addictive stuff.
I mean, I. You and I talked about it a couple of weeks back, but I love the clarity that comes when you're approaching any sort of liquidity event. Like the perceptions. The importance of perception gets swept away, the importance of politics gets swept away.
And it all becomes like, how do we. How are we sure that these are the right numbers? And how can we, either now or in the future, make them go up or make them go down? If that cost numbers, of course.
Yeah. So I love the. I feel like PE gets a bit of a bad rap sometimes because it comes in and clears away some of the safety blankets that organizations and leaders and organizations hold on to. The safety blanket of being able to override reality with perception.
The safety blanket of having built a reputation or career around a certain platform or a certain, you know, urban myth or legend, which may or may not be correct when the numbers are fully audited. So I can understand why people get. Why people feel displaced and people get their noses put out of joint when PE comes to town, but there's a certain clarity about it that I enjoy. No, absolutely.
This is when you're under the microscope and. But, you know, if I come from a finance background, we go through an audit every year anyways, right? So we're used to being put on the microscope. But when you.
When. But then it's not now just the numbers, it's your processes are under. Under the microscope as well. Your reputation is at stake.
And, you know, you may have, as you mentioned, you may have built this really great system or like this great spreadsheet that you're so proud of, and it runs like a third of this process. Well, PE doesn't care about that. PE cares. That's a risk.
That's a risk. How do we get rid of that risk? How do we actually turn that into a more efficient process so we can actually get a better bottom line? That's what everyone cares about.
Or if no one else knows how that spreadsheet works, then you are the risk. Exactly. You are absolutely the risk. So you talked about the path from finance to data consultancy.
And as you mentioned before, there's always additional rigor in the world of finance. And I have a background in financial services myself, so I'm familiar that when the cost of an error is a Eight figure regulatory fine. Then there tends to be a bit more of a business case for investing in data quality and governance for sure. What else did you learn from 20 years in finance that operating partners and PE firms and PE backed companies are still missing?
Well, one thing is, what I really find interesting is that nobody actually really audits the thing that makes the post plans work, which is the data layer. People look at the system, people look at the numbers, they don't actually look at that semantic layer that actually ties everything together that's actually quite often missed during due diligence. So like why we do financial audits and financial ratios and we look at, we write this tech stacks out of date. This tech stack is great.
This person really knows their stuff. We don't actually look at what the data layer looks like and what the data quality, data governance from the data layer looks like because that's what will actually eventually drive all the numbers. Yeah, there's a bunch of opportunity and risk that totally lives in there. And the other thing I was talking with a friend of mine back in the UK about this.
The other thing that actually automation can mask and modern data warehouses can mask things like models that need to be regulated, like credit risk models or financial models or any sort of. And in the case of pe, like decision models, like how are you deciding what customers to go after? How are you deriving lifetime value or acquisition cost of customers, all of these critical numbers to pe. And as warehouses become increasingly automated and our friend AI, and how long has it taken nine minutes for AI to come up in a podcast?
It's a new record. Historical tracking of that stuff is just not there. And I think I have friends who work in the reg space and being able to understand exactly what decision you made and what the logic was behind it and what calculations happened under that. I've heard people say layer of that's kind of being swept away with technology.
Technology can bring a lot of automation and governance and correctness to data. But do you feel that there's also a risk of history and temporal views and other important things that are important being swept away? There is that risk. Absolutely.
There is at risk. We're working at a regulated firm and we were seeing, it was seeing more fraudulent transactions come through and we were like, oh great, our system's working. It's getting more, it's picking up more and more fraudulent transactions. They were false positives.
Right. So we were actually costing the company 1% of their revenue because we had 1% drift towards, oh, this is fraudulent transaction. So we're rejecting it. So we still need to have that, you know, the human in the loop.
Like, do these numbers still actually make sense? Yes. Now we can process 17 fraudulent rules in under one second. Right.
But is it the output actually? Right, yeah. Is it right, though? Yeah.
And how are you deciding? Yeah, because you're right, those false positives turn into operational friction. Right. Because someone has to then manually deal with each and every one of them.
And then the client side, like, well, why am I getting like, your system is rejecting some of my revenues? Right. So the client's not happy because they're seeing their revenue go down, because they're seeing less transactions go through that were actually legitimate transactions. But we had some model drift that was saying, this looks like a fraudulent transaction that we weren't monitoring properly at the time.
Yeah. And then the consumer trust comes into it, et cetera, et cetera. And it's interesting. And having dealt with, again, lived very similar worlds to you.
That's a really tricky line to walk, that. Because obviously you want to err on the side of false positives rather than missing the ones that are actually fraud because that, as a customer experience is far more damaging than a false positive one. But it's a tricky line to walk and as we said, maintaining the logic and all the changes you make to that logic are really important is insufficient. Right.
In a regulated world to just have a GitHub style log of. Oh, we changed it because we were getting too many false positives. And then here's the new version. You need to understand what that old version is, what you were doing, what calculation, you changed all of that sort of stuff.
And I think there's a lot of SaaS platforms, including those trying to get into kind of financial services space, getting spun up very quickly at the minute in AI that are taking no care whatsoever. You definitely. Like, that was a perfect example, like how a data mistake, which is. No, it wasn't.
It was the better mistake because we weren't letting you know a fraudulent transaction go through, but it was costing over a million dollars a year to the company. So data and finance tied at the hip, right? Absolutely. And yeah, another thing I'm familiar with, like, 1% might not sound a lot to some companies, but if it's 1% of 10 million transactions, then suddenly you've got some major operational issues to deal with.
Tell me more about the award. The most innovative finance department. Aw. What does one have to do to win that?
What was the dinner like at the awards ceremony? Tell me more. Yeah, so that was, that's going back a little big time finance people are running their whole entire. Like I was acting chief data officer at the time for Qtrade and we had actually built a full multi dimensional, multi line of business finance data warehouse that was actually expanded and eventually expanded into an enterprise data warehouse.
But the fact that no, we could pull analytics to 10 different ways on 9am on the first business day of the month was absolutely crazy awesome. Right? Like we took a two and a half day process and turned it down to under two hours where our competitors were still copying and pasting spreadsheets to do commission reporting. We had a fully vetted, double, double checked two and a half hour process and that in the last hour and a half it was just a manual intervention of just a human actually take a look at everything, make sure that the output was right.
So like we had five lines of businesses, we had three different presidents involved in this. It was $100 million company, all running right through this one little finance data warehouse. We went no, we had 90% of our revenues booked at 9am on business day one. I love that, I love that.
And actually that metric you mentioned, if you're looking for hidden indicators of how good a company's data is, like how long it takes them to close month, close quarter, close year from a financial perspective, is usually a pretty good barometer of what you're going to find in the other areas of the company. And that was exactly why we actually wrote our first finance data warehouse was that we were subsidiary of a larger organization and they were giving us reports on the morning of business day six if we were lucky.
And we were also paying $100,000 per year for these reports and paying $385 per hour for data mining. So we built one internally. Well, instead of having to wait on business day six now we have everything available at 9am on business day one. We have a three business day close now instead of having to close on business day eight.
And that makes all the difference as well, like in terms of any decisions that hang off the back of IT management reporting. So everyone understands how the company's going, especially in the PE backed space as well, like investors are always keen to see. Yeah, fantastic. Bring it more into the recent past.
So SkySight built something called ScalePath, right. Specifically to help PE backed companies get their finance and data stacks in order. What do you typically find in one of these companies when you run them through Skillpath? And what does a typical finance data stack look like?
Mid market company. So what Kafnu find. We find a lot of QuickBooks. People are very heavy in QuickBooks and people are very heavy in spreadsheets.
We usually find that 80 to 90% of the journal entries are being done manually. And it's a bunch of accountants in there getting reports from here, getting reports from there, copying and pasting. Like we saw people creating a pivot table that would then be created to create another pivot table. So a lot of macros like no, you take three spreadsheets, turn to a spreadsheet, turn to a pivot table, take that pivot table, turn it to another format to upload into an accounting system, then reconcile the accounting system to get back to the forces source system again.
What could possibly go wrong? So that's what we typically see is QuickBooks, a lot of, a lot of Excel spreadsheets. So what we typically do is say, well, let's get out of QuickBooks, especially if you're a multi entity because usually in the mid market pe back, you're going to become multi entity, multi currency. Let's get into a proper ERP now.
Let's actually create a finance data warehouse. And you know, it can live anywhere. It can live in aws, Azure, Snowflake, databricks. You know, they all work for us.
And let's automate this. This is magical things called APIs that a lot of accountants don't know about. And kudos to them like Campfire. Campfire is like an AI first erp.
They're actually helping automate the actual journal entry process within the whole entire system now. So we come in, we basically say you can basically double your work and double your revenues and have no additional people need to be hired because you're just inefficient. Yeah, I want to come back to a point you made in the middle of that there about Databricks, S3, Snowflake, etc. I often find when I start with clients, they think that folks like you and I have a magical secret menu of bi tools that we can stitch together that will solve Everybody's problems.
And 99% of the time my answer to that question is like, it really doesn't matter. Actually, what matters more is getting a clear definition of what you're trying to do with the warehouse, what reports you want out of it, what decisions you're trying to make out of it. Once we've got that sorted, unless there's some really complex ML that's going to take advantage of a cutting edge databricks feature, or unless there's some real Some huge amount of storage that's going to take advantage of snow flexibility to transition stuff across for cheaper compute.
It doesn't matter. Right. And usually none of those things are the case. None of the things are the case in the mid market.
It's not a technology tool issue, it's a process issue is what most people don't realize. Technology is not the silver bullet, it's the process that derives everything behind there and how you deliver it. Yeah, no, absolutely the case. A lot of things being held together with tape or as you say, maybe QuickBooks can be one of the better examples rather than some of the stuff's on QuickBooks, some of the stuff's over here on this one, some of the stuff's over here.
All things that are normal to find. And you know, I think we should stop to say from our position of privilege here, like having worked at big companies, this is a normal part of evolutionary growth in mid market. Right. This isn't mid market people don't know what they're doing or mid market people don't have the right idea.
It's just the company starts from nothing often and has kind of built up and scaled because it's so damn successful. And it's more often a case I've found of like, okay, this used to work three years ago or this used to work five years ago. This was perfectly acceptable. Right.
My business runs on QuickBooks. That's okay. I'm not dealing with 100 million transactions a month. So it's okay for me to be on QuickBooks in the future when I've got the world's biggest data consultancy in two years, Greg, then yeah, sure, maybe QuickBooks isn't going to cut it anymore.
But it's very often nobody's at fault here. It's just that time and scale have overtaken data architecture and it no longer meets the needs is typically what I'm seeing. No, absolutely. We quite often see that the finance and data departments are one of the last ones to get addressed.
Sales is the first one that gets addressed because they're the ones that are scaling up. Then you have the product that's going along with sales because the product. Sales needs a product and all that happens up there. Well, the back office, behind the scenes, they're the last ones to get addressed and it's not.
And it's just an evolution. It's just like you have, it's a small pie, limited pie, like how much money you're going to spend eventually it catches up though, right? Because a lot of People say, oh, we're just going to fake it till you make it. Right?
Well I always say that it's fake it till you break it. And if you break it, that's a lot more expensive, right? Like you can, you can drop a netsuite or you can drop Snowflake into a broken process and all you get is more expensive broken process. Totally agree, totally agree.
And I've seen that a lot of times. So yeah, that's the key, deciding the right time to make the decision. And again, the decisions to prioritize sales, marketing, product are completely rational and understandable and correct. Right.
Because those things go straight to the bottom line and that's what gets companies on the radar PE in the first place, that they've grown and grown and scaled so successfully. It was just that data and as you say, finance are more kind of horizontal enterprise investments that are the kind of rising tide that silently lift all the boats behind the scenes without saying spend whatever 5,000 on Snowflake this year and your revenue is going to go up 10%. It doesn't necessarily work like that.
There's a lot of risk mitigation that it offers and a lot of defense that comes from having a well governed data as well as all the opportunities that it unlocks. No, absolutely. As you mentioned, it's no one's fault. And quite often we're talking portfolio companies of private firms, quite often series A or series B companies, usually the VP of finance, or sometimes it's a head of analytics or a head of finance.
You know, they've been in the company for two, three, four years, even the company go from like four to five to $10 million. They've never been for that process of going from 10 to $100 million. It's a case of you don't know what you don't know. That's why like a scale path makes sense.
Because though as a company we've been through about 15 to 20 different scaling ups so far. Right. But we bring in the expertise that your head of finance just may not know or your head of data may not know. They were a business analyst five years ago, now they're the head of data three times the size, right?
Yes. It's just, it's just bringing in the experts who have been there to help you, not to replace you, but just to help. Zero. Yeah.
And in terms of actual impacts and again we can only ever guess because there's so many variables that go into the equation. But how much difference might a really good finance data stack versus A really kind of manual and lower quality one actually make to an exit multiple. So we have something we call the dirty data discount. It's usually worth about 10% of the valuation.
We did a survey about 120, 122 different advisors and kind of basically said, if we gave you this data, we gave you this data, what would your valuation be? And you know, we got anywhere between a 1% discount and a 22% discount. And actually about 12 or 13% actually said, we're just walking away, we won't even touch that deal. It's just a no.
Yeah, yeah. So like, you know, if you think about it, you know, in a mid, lower, mid market, that's 40 to 50 million dollars deals, 10% discount, that's 4 odd million dollars you're leaving on the table because you have this dirty data. Now if you give me a minute, I'll explain what we mean by dirty data. Yeah, go ahead.
A prime example was a SaaS company we're working with. No, they were on QuickBooks, $21.6 million of revenue. Great, great success story, going from three and a half to seven to 14 to 21.
6. But what's happening is during the exit process, they needed a product profitability report. Well, the problem is like most firms, they hadn't actually grown the back office, their product set in three databases. When they pulled that together, it's a 21.
15 million revenue. When they went into their CRM, they pulled $21.3 million of sales, 21.15, 21.
3, 21.6. Well, when you're looking at exit, you're telling a story, you're selling your story and the numbers. Well, now the buyer side's like, what's going on here?
We don't believe your numbers. And soon as you start losing that belief in the trust, down goes your valuation. Right? So they, they came out and said, okay, we're going to actually give you 21.
1. There's a half million dollar discount, but it's not really half a million dollars because it's a 5.8x on multiple. So you're talking about $2.
8 million of lost revenue of lost valuation because you didn't have the correct reports in place and you couldn't prove your 21.6 million of revenue. And even that's, I would say that's generous. They were lucky to get the 21.
1 because as you say, trust is more important than any of the three numbers. And it Captures perfectly what I describe as the difference between management grade reporting and investor grid reporting. So for management grade, like okay, year to date, we're on 21.3, 21.
1, 21.6. Either way, that's 7 million higher than it was last year. So I feel really good about it.
And we can all go about our jobs leading this company and feel good about ourselves. But with investors it raises, as you say, that important question of trust. And that's why I think the 21.1 is generous.
Because the other approach they could take is like, well, rather than believe the lowest number, we could just choose to not believe any of the numbers. Right. And rather than giving you a pass and giving you three minutes to work, three months to work out what's going on, as you say, the outcome might actually be more than the multi million discount. It might just be, okay, well we'll go find another company.
Because people are getting more and more selective with the whole period's lengthening. There's competition for deals around as well, of course, but people are getting more and more selective with the whole periods lengthening and it's easier to to walk away and find a company that has got numbers you can trust. Now the good news story about all that was that we came in, we actually wrote the finance data warehouse and applied our scale path methodology and the actual number was 21.
52 million. So we went back and we restated QuickBooks to be 21.52 and we gave the buy side six different reports. Slice and dice by any different views of the world that all tied back to 21.
52. So we actually ended up adding back over $2 million at valuation. No, this company was real lucky that the buy side was given the couple of months to come back in here and fix everything. It could have either been a dead deal in that moment.
Yeah, that's rare. Patience. Well, I hope they bought you a nice dinner. That they did.
Yeah, that was. They're very nice. I love it. My guest last week was talking a lot about how the old playbooks are not working anymore in private equity and how data is an increasingly important layer because, you know, I think some of the cliched old patterns of like I'm going to buy 10 chiropractors and the operational efficiency will be inevitable because I'm rolling them up and then because me and my buddies are smart and we'll dispense some executive wisdom every quarter of the board meeting, our success is all but assured.
Right. And playbooks like that aren't working anymore like we talked about, hold times are lengthening. And I'm certainly seeing data become a more and more important factor in not only monitoring value creation, but diligence in the first place. They want that extra layer of knowledge to make the right choices when they're buying.
Absolutely. We're seeing the same thing. The data rooms have gone from 25 to 30 documents now. Like, last data room I worked on was 306 items.
So they're becoming more and more diligent and they're wanting more. Now you need your cohort analysis, but not just know by time period. You need it by, by region, by client type, by client size. Like you need four or five different chords to now analyze by segment, by country like it's acquisition channel.
There's another endless. It is absolutely endless. Like, no, the. And then the, the.
If there's gaps in there, the inconsistencies of data, it becomes a really big thing now because it's not like, hey, we don't and don't get a week to produce report anymore. The buyers are expecting to see that in the data room. Not like they don't want to have to ask for it two months later and then wait a week or two to get into there. You need to have that ready.
So data room prep before you actually even get into the deal. Jones. And now we're actually even seeing some really interesting things where you can actually have start to query against the data room itself. Oh, right, okay.
And how do you set that up? Do you, do you push that into a model and then, and then run NLP queries on the model or are you creating some sort of structure and structuring it? Yeah, so there's, there's some data rooms out there that are. It's not our data room.
It's like it's someone else's data room. But you know, they're basically creating like MLP models that sit on top of their data room. So when you upload the documents, it says, oh, okay, I understand what this document is. I understand this document is.
And so if you ask them what is, can you show me what the revenue is across C3 documents? And it would actually tell you whether or not it's consistent or not. And the buy side is using that against the sell side. Oh, I love that.
I love that. So AI is hitting the data rooms now as well? Well, no, it's no surprise to me to see it continue its conquest of the world. Used in the right places and in the right way.
Finance is an interesting place, right? It's great for assessing things or comparing things. I think the risk that people need to watch out for, of course, is like if an AI model feels like it's got permission to fill in the gaps, right? For example, if it sees the 21.
1, 21.3215, is it going to say, oh, I'll just take an average and we'll call it 21.3 and we'll talk authoritatively on it in my response, because I haven't been told specifically to call it out, so there's a bit of model management needs to happen, but no surprise to me to see AI coming in. And I talked to a few people about this week as, as with anything in the news, right, Everything goes to the ends of the spectrum, right?
So you hear people say, I don't know if you saw the Matt Schumer essay last week called Something Big is Happening. I'll link it in the comments, but look it up. But it's, it's a very. It did well because places fear uncertainty and doubt in people.
And basically he was saying we are at a moment with AI that's basically like we were when we first heard about COVID of like, oh, it's no big deal whatever, a new virus, just like bird flu that we had before, it'll blow over. And then three weeks later our entire lives had changed and turned upside down in a way that we as human beings are incapable of fathoming that amount of change on that amount of time. So he was arguing that we're at that tipping point now. So that's on one end and then on the other end you get this AI is useless.
I have to spell strawberry or cut the oz and strawberry. And it got one thing wrong. I set a trap for it and failed. Therefore the entire technology is useless, right?
And you hear these two camps at the opposite end of the spectrum arguing with each other, but the truth is somewhere in the middle. And you've seen AI in high stakes scenario. Like a data room on an acquisition is a high stakes scenario I would say from a business perspective. And of course there's been loads in the news this week about the anthropic stopping the Pentagon from using Claude for some things that they'd like to use it for.
Definitely a high stakes scenarios going on there. Where do you think the truth lies, Greg, in terms of where are we at now? If it's zero, AI is useless. 10 the robots are going to take over the world and nobody has to work tomorrow.
Where do you think we're at. We are at Skysat. We firmly believe we're in the 7 range, like we're actually calling AI, just like our business partner now. So Claude released their Excel module a week or two ago and it's changing the FPA world.
I played around with it and built three different models out of it so far. Another one, my colleague Paul Bernhurst, FP, and a guy, LinkedIn, really great guy to follow. He's done a bit full on tests on it as well. And it's gone from a year and a half ago, its models were crap.
Now it's intermediate to senior analyst level. It still needs review, it still makes a couple of small mistakes, but it can actually explain to you what its logic is and why it's building it that way. So what used to take two days now is two hours in Excel using Claude. It's definitely what we call the business partner now.
And if anyone is still sitting there coding line by line by line by line, they're getting behind the times. Yeah, they're getting left behind for sure. I know a couple examples. A friend of mine, his son just left X.
He'd been working on Grok as a software engineer and he basically said that he told his dad that he was bored because he was no longer writing any code. He was mainly just reviewing code because the recursive loops had got so good that the model is basically writing its next version by itself, largely. So I think a lot of effort has gone in the software engineering space as well, for sure. Another example I've got, I've had this battle with like creating a personal CRM.
Like Facebook used to be amazing for reminding me of when my friends and family's birthdays were. But they put so many layers in the way and then like they want to turn on notifications and send adverts to me and it just became a nightmare. Personal CRM products were running like 50 bucks a month as a SaaS platform. I was like, ugh.
So I sat with cloud coding in an hour. It took my Excel spreadsheet that I exported from one of the SAS products and made like a personal CRM for me that I just kick up once a. Once a month, a week, check for birthdays. It reminds me I've got personalized settings where I can like snooze reminders for a custom number of days and whatever else.
It's like having seen it in action at its best. It's scary. And as my friend reminded me whose son worked on Grok, what we've got is six to 12 months behind what those guys have got. So yeah, the progress is exponential for sure.
Having said that, here's a negative side of it. Right? So just we were working on a scale path client and we, out of curiosity, we, you know, we automize the data and put no, ask Claude, ask ChatGPT, ask Gemini, what, what would your recommendation on this scenario be? And so we fed something.
We fed a P and L balance sheet in and they came back and said, oh, you know, this client needs to go to like a NetSuite. They have 17 legal entities. No, they didn't. The actual client, it was actually a single legal entity, single currency.
They had 17 related party transactions, which was the shareholders. It misread the chart of accounts as each payable as a separate legal entity. So I thought it was a 17 legal entity. So they was recommending spending 80, $90,000 on NetSuite when the problem wasn't QuickBooks.
The problem was all the information going into QuickBooks, not QuickBooks itself. The $225 per month subscription to NetSuite, to QuickBooks was all they needed, not a $90,000. But if you listen to Claude, he would have said, hey, you need to spend $90,000 on NetSuite. Yeah, that's interesting.
And it helps catch it when it explains its logic for sure. And I've seen, I've been playing a lot with Cortex code on Snowflake and that's very good and I think notably better than Claude and GPT at explaining what it's doing along the way. So Cortex basically gives you an NLP front end interface. I think it's powered by anthropic, actually.
I think it's got Opus 4. 6 sitting under it. But it actually says, oh, the user asked an ambiguous question. So sometimes it will come back and ask you, but other times it will say, I'm making this assumption because XYZ and I looked at the data and I think they mean this.
And it's line by line will explain to you every assumption that's made along the way. And I think increasingly that's especially for higher stakes things like we've talked about. That's an important feature for a model to have. The one thing that we're seeing as you kind of get into some of these M and A discussions now, is that we're seeing boards ask, is our data clean enough to use.
There's not a board meeting I've been to as an observer or as a member. No one's asked that Question. Right. Well that's kind of.
You're in now to get a budget for conversation that you would never had before. Yeah, it's critical. And that's where the competitive advantage lies because in the tool itself, like as it ships or as it's available on the web, it's no competitive end because everyone else has got it. Right.
If you don't have first party data or high quality data about your business or your customers, then there's no competitive advantage there for you to eke out because everyone else can use it to write marketing copy and all the other things that you might do as a business. There's SaaS and FinTech are under extreme pressure right now. And then you have to be very careful that the Vibe code you do, that's not a production ready product. Right.
That's just Vibe coding that gets you, it's an mvp, it's a great thing. Works for my personal CRM, doesn't work for a bank. You're not going to run a $100 million company off of Vibe coding and pass regulatory examinations. No, absolutely not.
And ultimately and correctly as well. I would say with reg exams you could vicode an AI substitute as much as you like, but there still needs to be a person whose name is on the block for a reg submission. Right. And if that person's smart, they're not going to outsource that to a model, especially one that they don't know they can't track and have full history of what it's doing.
Like at the bottom of the balance sheet, the audit financial statements, there's a signature on the bottom of the audited financial statement saying this was reviewed by person X. Right. And if that, if that's wrong, there's fiduciary duties and criminal charges can be brought against you against all this stuff now. So you can't just go blindly.
Yeah, peeps gotta be careful. Speaking of opportunities, you offer data monetization as a service, and this is no stranger to me, this idea that first party data becomes valuable and potentially even more so actually given what we just talked about with, with AI, that proprietary data could be more important. So you help companies figure out whether the data actually could be an asset that they could sell, which I'm sure is a novel thing to most mid market companies. What sort of situations does that make sense in?
So what it comes down to is the two biggest buyers of data are academia, npe, they're the ones who buy the data. How many universities out there, how much research are universities always doing? There's data libraries, data monetization libraries, just on marketplaces, just for academia. What makes your data more.
The more unique it is, the more it's worth. But if it's more unique, it's less tested. So less tested makes it worth less. So there's kind of the dynamic pull that everybody wants to have.
The, the different data points, especially in pe and like hedge funds, the analysts, they want this unique data point, but they need to be able to trust it. Right. So it comes back down to what's your internal data quality? Like, what if you don't have a good data quality to go publish it?
Don't even bother, because you are the hook for the quality of this data. There are ramifications having wrong data. But you know, some things that we had one company come to us, like, oh, I think we can make, we make probably a million dollars a year off our data. We looked at it, I'm like, so you are 13th in your marketplace.
You are one tenth the size of your largest competitor, and your data is no different than your competitors. What would make yours worth a million dollars and not theirs? Good questions, no answer. So that's the case.
Like, no, you don't want to, like you're going to spend 50, 70, 80, $100,000 doing all this. Your return is probably going to be 50 to 80, $100,000. Another one was like we had from pharma data. So we able to mask from pharma data that pharma data was worth a lot of money.
It was an $800,000 data set. Is this like who's using pharmaceutical products or how they're performing in testing or it was who's using pharma products. So it was basically like a lot of drugstore data across Canada. Right.
So we could tell you then, like by very, very high level that, you know, in this, you know, in this, in this city, here are the top five drugs in the 20 to 50 being sold at these drugstores between the ages of 20 and 30. Between 30 and 40 does. My mind jumps immediately to GLP1s and the huge demand there must be for that data right now, because there's a big market out there and I think they've got pills now, haven't they? As well as the having to inject yourself in the stomach things.
Yeah. So it's all about like, nose, like how, how unique is your data set and how trustworthy your data set. Those are two things that will drive your value. Yeah, no, absolutely.
And I, to clarify, I'm Not a user of ozempic. The. I'm just aware of what a hot market is now and what a big margins are available in there. So if, you know, if I think about those new to market folks who've got the GLP1 pills, for example, like to them knowing the regions of Canada and the type of people who are buying GLP1s is incredibly valuable.
Yeah. So there, so there's of course when you're also going through the whole entire data monetization, you have to look at hashing and anonymizing your data because there's some data you cannot sell. Right. There's all sorts of pii, gdpr, hipaa, like where.
But depending which region you're in, you have to really follow all these, you know, privacy regulations. You have to make sure that you nothing that you sell can come back to the user. Yeah. There was an example last year of location data.
I think people who were able to through not always 100% ethical means, like get apps installed on people's phones or whatever so they could track their location and they had some demographic information. So they had basically geo information down to the level of like a block or a building, even like where people were going, which includes of course people who are going to jails, people who are going to doctor surgeries, people who are going to abortion clinics and other such like sensitive subjects.
And it was. The government always interests me because the government like clamped down on this obviously, like you can't sell this stuff. Like this is not okay. Like you can't, you can't sell it.
And then they actually gave out a huge fine, but tempo'd the fine and said this is only allowed if you're giving the data to us. But this also leads to another interesting discussion about, interesting. Depending if you're a CPA or not, but right now, from a balance sheet finance perspective, Data is worth $0 on your balance sheet. Data is not considered technically an asset under accounting rules.
Right. You know, I am the personal opinion, I've been our regulate, I've been trying to argue with the regulators on this, the accounting standards that it falls under an intangible asset. Like if you have, if you have such things as goodwill, which is the amount you pay over for a company, if that's an intangible asset, why can data not be considered an actual. How do you value it?
There's like a lot of people have done some work on like how do you value it? Is it if you sell it for $800,000, isn't it worth $800,000 on your balance sheet. Now, if you have a data breach, what happens to the data breach? You get sued for millions and millions of dollars.
Yeah, yeah. How do you sue on a worthless asset? Yeah, that's interesting. But you know, and people say you're getting sued because of damages, but damages mean that it was valuable in the first place.
How can something go otherwise it wouldn't have been damaging to get it out? Yeah, yeah. So the. So you know, by trade accountants are very slow moving, hence why they're the last ones to adopt the AI, one of the last ones to get off Excel spreadsheets.
But it's just an interesting dynamic, technically speaking, from a corporate perspective, Data is worth $0 on a balance sheet. So pulling from your CPA knowledge, which far outstrips mine, what other sorts of assets are intangible assets, like intellectual property, for example? Is that one? Yeah, so it's all under the IP rules.
So you have a patent. A patent can actually be considered a valuable asset. Right. People sell patents for millions and millions of dollars.
People sell data. So like, no, it's anything that it's an intangible asset. Like there are, there are many things out there, but data's just not one of them. So it's an interesting conversation that's been going on for I'd say about eight years now that we have made absolutely naturally say about 10% gain on the accounting society on this argument so far.
But you know, because now you can't run AI without your data. So like, once again, how is data not valuable? We just spoke about how data is the thing which turns AI from a tool that everyone has into a competitive advantage. Right.
So how can we say that's not valuable? So I'm really, I would say excited, but I don't want to use the word accounting and exciting in the same sentence here, but really interested to see how the accounting societies start to value data in the future. Well, I'm down for it. And if you ever want an objective witness statement, Greg, I'm here for you.
I'll take the stand for you, don't worry. Say the word. Say the word. So, one last thing, I just want to get into patterns, particularly patterns that you might see repeat across PE backed companies and PE portfolios.
We're looking at largely the same companies from slightly different angles. What are some of the patterns that keep showing up, regardless of vertical, regardless of the size of the fund or the company? What do you keep seeing? Well, I'll see something like, well, we quite refer to the 100 day tech mistake where operating partners scope systems work before they actually understand the process underneath.
It's the hey, oh, we need a snowflake. We need a databricks to fix this. Once again, we talked about this earlier. You can drop netsuite into a broken process.
You need to get a more expensive broken process. Right? Yeah. Why not get them Palantir, let's go all the way.
They're not paying attention to the data layer. Like hey, 21.6, 21.1.
Everyone's like, oh, it's close enough. But close enough is good enough for now until it's not. Right. Like you don't want to be in a position where you're starting to do this, like trying to cleanse all this stuff three months before an exit.
This is a 12, 18, 24 month time frame. Like you want to get in front of it that way when you go into diligence and go into it because you know, as pe, we don't want to have seven and ten year old periods. You want to get out for it faster. Doing it at the last 11th hour doesn't do anyone any favors.
So these are the kind of mistakes that we keep seeing is that people put it off to the end. People start just replacing tech with understanding the process, not realizing it's a process that's broken, not necessarily the tech. Yeah. And it can be even more dangerous, I found because if you fix SWORD like that, then even if you put a great model in or great traceability and governance in going forward, you still have to fix all the historicals.
You can't just say oh, our data is wonderful since the day that we put snowflake or databricks in. These are not the droids you're looking for. For everything up until that point that doesn't work, does it? Say no.
And you do some of the same work that we do. We basically come in as data consultants. Right. But it becomes what we call the adversarial consultant problem.
It's how you frame the engagement. It's like we're not here to point fingers or auditors. We're actually trying to make your life easier and better and make us all more money here. Yes.
The way the engagement model and the way the operating partners present a data consultancy firm matters as well. Yeah. Is that part of the old playbook not working anymore? Yeah.
You can't just say we're bringing this person in and they're going to fix it. That hard approach isn't working anymore because what's going to happen is people get their backs up against a wall and finance teams like otters are going to try to hide things. If they try to hide it, then it's never going to get fixed. Yeah, we're playing a largely similar game.
And for me, it's so important that that psychological safety is in there going in for the existing employees of the Portco, because if it's not, you don't get the quality information that you need to be able to make good recommendations. No, what we're seeing there is a real gap in the market between financial diligence rigor and data diligence rigor. And we're starting to feel it. We're really starting to feel it.
Yep. Well, we're both going to be busy. That's the actual conclusion to draw from that. And with that, like time has flown by so fast.
So fast. Greg, like, I've already enjoyed our previous conversation so much and we'll continue to enjoy more in the future, I am sure. But for now, thanks for coming on. For those listening who want to find you want to find skysight and dig a bit more into the type of things that you've been talking about.
Where are the right places on the Internet to go and find that stuff? Greghood CPACMA on LinkedIn or www.skysiteanalytics.com Perfect.
Great domain name. Great podcast episode. Thanks so much for coming on, Greg. I appreciate you.
Thank you so much. It's great to be here. All right, take care. Bye.
Thanks for listening to the PE Data guy. The place where private equity meets data. Please forward this episode to your favorite private equity friend. Thanks for listening.
See you next time.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.