The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The Data Crunch
The Data Crunch artwork

Season 2: Episode #23 | Version Control in Analytics and Why It Matters

The Data Crunch · 2025-06-12 · 11 min

0:00--:--

Version control in analytics solves a pervasive problem: data teams often develop SQL queries, dashboards, and metrics without any documentation, versioning, or change tracking, leading to overwritten work, mysterious formula changes, and data integrity crises. Helen from Ox walks through the pain points - tribal knowledge, accidental metric inflation, and impossible-to-trace bugs - and explains how central Git repositories, pull requests, and clear commit messages create accountability and auditability. The episode covers specific implementation tactics: starting with core KPIs, using structured folder hierarchies (organized by department or product, not analyst names), maintaining README files, archiving deprecated SQL instead of deleting it, and adopting naming conventions that reflect what the code does rather than its editing history (no "final_final.sql"). Tools like dbt, GitHub, and BitBucket provide the foundation, though even basic setups with Notion or Fiberry plus disciplined Git workflows substantially improve visibility. The takeaway is that version control doesn't require perfection - just intentional structure that helps teams understand who changed what, when, and why.

Key takeaways

  • →Start version control with your core KPIs and revenue models, which have the highest impact and stakeholder visibility, rather than trying to version everything at once.
  • →Use clear, explanatory commit messages that capture why and how changes were made (not just what), because Git already shows the diff - the context is what saves time during reviews and debugging months later.
  • →Organize repository folders by business domain or product area, not by analyst name, and keep folder depth shallow (aim for one or two clicks to find a model) so team members can locate relevant logic without knowing who wrote it.
  • →Archive deprecated SQL with a note explaining why it's no longer in use instead of deleting it, preserving historical context for future metric reconciliation and stakeholder inquiries about how calculations have changed.
  • →Adopt dbt, GitHub, or similar tools for built-in versioning, testing, and pull request workflows that enable peer review before changes reach production, catching errors like accidental filter duplication that would otherwise inflate key metrics.

Topics in this episode

GitHubdbtKPI trackingBitbucketPull requestsVersion control systemsCommit messagesSQL versioningAnalytics documentationGit repositories

Questions this episode answers

What should I do when I find old SQL queries I'm not using anymore?

Archive them in a dedicated folder (e.g., 'archive' or 'deprecated') with a brief note explaining why they're no longer in use, rather than deleting them, because that old code may help explain how metrics were historically calculated or reveal context needed to debug future issues.

How do I write a good commit message for analytics changes?

Explain why and how you made the change, not just what changed - Git already shows the diff. Good context means when a similar issue emerges months later, you can find and understand the original fix without hunting down the analyst who wrote it.

What's the best folder structure for organizing analytics code?

Organize by business domain, product area, or vendor rather than by analyst name, keep folder depth shallow (aim for one or two clicks to find anything), and include README files with context so team members can find relevant logic without knowing who wrote it.

Can we use version control if we're not ready for Git yet?

Yes - you can start with tools like Notion or Fiberry combined with basic documentation and naming conventions to create structure and visibility, though Git-based workflows with pull requests provide stronger change tracking and review capabilities.

What happens when you don't use version control for analytics?

Teams overwrite each other's work silently, copy logic into multiple places creating conflicting sources of truth, and struggle to trace metric changes - leading to errors like accidental filter duplication that can inflate key KPIs undetected for days.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B67%
  • Speaker A33%

Most-used words

version16control12helen8analytics7changed7names6final6data5logic5context5folders5back4today4file4analyst4start4

Episode notes

Tired of sticky-note SQL and fragile dashboards? In this episode, Vadym and Helen dig into version control in analytics - why treating your SQL like code matters, how to avoid "final_FINAL" disasters, and what real-world teams are doing to keep their data trustworthy. What you’ll learn: Why version control is critical for analytics teams Real-world mistakes caused by missing change history How to implement lightweight versioning with tools like Git and dbt Folder structure, naming tips, and archiving best practices A quick tour of helpful tools - from GitHub to Notion ️ Start making analytics more reliable with OWOX BI ️ Want to help build and improve our open‑source connectors? Join the OWOX Data Marts community on GitHub and contribute Get trusted reports with your context in 1 minute OWOX Website Analytics with OWOX BI YouTube Channel

Full transcript

11 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Hey, everyone. Welcome back to the Data Crunch podcast. I'm Vadim, your host, and today we're tackling a question I think most data teams have silently screamed at some point. Why do we treat analytics logic like temporary sticky notes? You know, dashboards built on a fragile SQL versioned only by file names like all lowercase, final, all caps, final caps, this one use SQL, you know. And joining me today, you know her, you love her. It's Helen, head of customer success here at ox. Helen, welcome back.

Speaker B: Thanks, Vladimir. Yeah, it's always great to be here. And about the, uh, question you asked. Yes, the nightmare starts when someone opens the. The file six months later and says, wait, where did this number come from? I've been there. So, a SQL, um, that no one remembers. Written, no context, no backups, just, you know, vibes.

Speaker A: Yeah, exactly. And that brings us straight to today's topic, Version control in analytics. And if you're enjoying our podcast, hit that subscribe button. We drop episodes every Thursday with stories, real life lessons, and just enough drama to keep it fun. All right, Helen, let's jump in. Version control. Really? Not the sexiest term, but wow, it is important.

Speaker B: Totally, totally. Uh, think of version control as giving your analyst logic, a brain and a memory. So without it, you are, uh, one broken query away from a crisis.

Speaker A: Yeah, let's talk about the pain points. What happens when analytics teams don't use version control?

Speaker B: Okay, first, um, people overwrite each other's work without knowing. Or even worse, they copy past logic into five other places, and suddenly your single source of truth is five sources with different numbers.

Speaker A: Yeah. And when things go wrong, you hear, oh, yeah, Anna changed that formula last week, I think. No changelog, no context, just a shrug.

Speaker B: Yeah, exactly, exactly. And, uh, you know, let's be honest, tribal knowledge is not a long term data strategy.

Speaker A: Yeah, agree. Um, all right, so. So what does good version control look like in the analytics world?

Speaker B: Um, think central git repos. So for DBT models, SQL scripts, even your BI dashboards, um, use pull requests, add comments, review changes. And, um, it's about. It's not about slowing things down. It's about knowing who changed what, when and why.

Speaker A: Exactly. And structured folders. Not finalfinal or Dashboard. New July, Edited.

Speaker B: Yeah, real names, please. And the ability to roll back when something breaks. Yeah, that's absolute lifesaver.

Speaker A: Now, let's say you're new to this. What is the best way to dip your toes into version control?

Speaker B: Start with your core metrics, the KPIs. Everyone depends on version. Those Queries first. And um, even if you're just using GitHub with basic folders and simple commit messages, it's a huge step forward. Assign code owners, track who's responsible for what. Yeah, and hey, m, it doesn't have to be perfect. Uh, it just has to be better than invisible logic like floating around in people's head.

Speaker A: Okay, let's shift gears a little bit. Um, Helen, hit us with a real world story. One where version control saved the day or where the lack of it almost sank the ship.

Speaker B: Um, sure. At one company in a faraway galaxy, a junior analyst updated the revenue model on Friday. Of course, uh, On Monday the CFO saw 30% jump in MRR and almost celebrated it with the champagne. Uh, but turns out the analyst accidentally duplicated a filter or something and inflated the numbers. Like no versioning, no pull request, just, you know, silent update and it took days to untangle. If we'd had version control, one glance at the commit would have caught it in seconds.

Speaker A: Yeah, champagne postponed. I guess that's the perfect case for why this matters. Alright, um, let's get to a rapid fire time. Helen, give us some quick version control tips every analyst should know.

Speaker B: Yeah, let's go. Um, first is always explain your commit messages. Fix stuff. It will not help anyone. A good message tells not just what changed because git already shows that, but why it changed and how you fixed it. Um, that context is also good. Um, especially when you hit a similar issue months later and need to retrace your steps. So you just should be able to find that version. Um, the second thing I'd say don't push it to main, open a pull request and get a second pair of eyes. Uh, of course, if your workflow allows it, uh, and this ties directly into the previous point. A clear, precise commit message makes the review process way faster so the person reviewing your PR will instantly understand what was changed and why. Like no guesswork, no back and forth, all these things. Next I say create the readme, uh, for your folders with context. Actually the folder structure itself matters a lot. I've seen the different setups. Uh, they all worked well, like organized by department, by product area, by vendor if relevant or by project. Basically, uh, one thing I do not recommend, uh, naming folders after analysts if you are looking for a change related to a country grouping names, uh, you won't think to check the Julie or Thomas folder even if their readmes are perfectly written. So keep folders shallow when possible. Um, if you need five clicks to find a model that's the red flag. Another one is archive and SQL. Do not delete them, please just archive. And even if it's no longer in use, old SQL can hold valuable context. You might need to reference it later to understand how metric was calculated in the past or debug something or answer what changed question like from stakeholders. So you can just create archive or deprecated folder, whatever and uh, add a quick note about why it's no longer in use and um, move it there. Clean structure, no lost history. Yeah, makes sense. Um, and the last one, I'd say the naming conventions use it like consistent report, version 2 final. Final is a lie. Yeah, I know it sounds obvious, um, but this naming chaos is still everywhere when you're using tools like Git. There is no need to cram metadata, uh, into file names like no dates, no initials. Git already tracks who changed what and when. So your file names should focus on what the uh, SQL actually does, like monthly revenue by region instead of final. Final. Use this one SQL, um, the clear names, faster onboarding, easier review and you know, just fewer headaches. That's it.

Speaker A: Yes, thank you so much, Helen. That's a great breakdown. And uh, now before we wrap up, Helen, give us a quick tool roundup. What should people be checking out?

Speaker B: Um, dbt, um, they have built in versioning, testing and documentation. It's just the game changer, um, GitHub or BitBucket. Don't overthink it. Uh, start with some basics and um, any internal tool you use for knowledge sharing like Notion or Fiberry plus SQL. Actually, uh, if your team is not ready for git, you can use this combo and it will still give you some structure.

Speaker A: Yeah, great advice.

Speaker B: So.

Speaker A: So there you have it. Version control isn't just a dev thing. It's the key to consistent, reliable analytics in fast moving teams.

Speaker B: Yeah, and it doesn't have to be complicated. Just pick a spot, start small and grow from there.

Speaker A: And if you want to trust your data and bring order to your analytics logic, check out oxbi. We'll make it easy to version your transformations, track changes and collaborate on data across your team. Head to ox.com and start your journey to reliable, transparent analytics today.

Speaker B: Yeah, thanks everyone for listening. And remember, you don't need perfection, you need visibility. Version control helps you get there.

Speaker A: Yes, thank you, Helen. Thanks to everyone listening. Subscribe, Leave us a comment and tell us your best version control. Horror story. We'll catch you guys in the next episode of the DataCrunch podcast.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • How do you turn AI coding chaos into a repeatable playbook?The Stack Overflow Podcast · on GitHub86 / 100
  • Is Your AI Actually Worth What You're Spending? with Parker ConradStrictlyVC Download · on dbt86 / 100
  • Orchestrating data across 30 companies at ittiThe Data Flowcast · on dbt81 / 100
  • Code Review Is a Taste Problem | David Poll ⁨@GitHub⁩Hangar DX Podcast · on GitHub81 / 100
  • How Finance Teams Are Actually Using AI | Opendoor, Datadog, PwCRun the Numbers · on GitHub78 / 100
  • The context engineering playbook (Claire Gouze)The Analytics Engineering Podcast · on GitHub75 / 100

More from The Data Crunch

All episodes →
  • Season 2: Episode #30 | Top 5 Reasons Data Analysts Hate SaaS Tools (And What They Really Want)42 / 100
  • Season 2: Episode #29 | From Firefighting to Strategic Analysis: Elevating the Analyst Role
  • Season 2: Episode #28 | Training Business Users: Setting Boundaries for Data Exploration
  • Season 2: Episode #27 | Self-Service Reporting: Empowering Users Without Losing Sleep
  • Season 2: Episode #26 | The Hidden Costs of Copy-Paste Reporting
Explore the best B2B AI & Data podcasts →
All The Data Crunch episodes →