The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Sales/Enterprise Tech with Fexingo
Enterprise Tech with Fexingo artwork

How Fortune 500s Use Procurement to Manage Vendor AI Training Data Rights

Enterprise Tech with Fexingo · 2026-07-01 · 10 min

0:00--:--

Key moments - from our scoring

Substance score

70 / 100

Five dimensions, 20 points each

Insight Density16 / 20
Originality15 / 20
Guest Caliber12 / 20
Specificity & Evidence14 / 20
Conversational Craft13 / 20

Vendors are embedding broad, perpetual data-use clauses into enterprise software contracts that grant them rights to train AI models on customer data, often burying the language in addenda that procurement teams only catch during renewal. The episode walks through a real negotiation at a top-20 financial services firm where a CRM vendor's standard clause claimed worldwide, perpetual, royalty-free rights to customer transaction and behavior data for machine learning - a deal-breaker for regulatory and competitive reasons. Lucas and Luna explore the negotiation tactics that worked: carving out scope (aggregated, anonymized data only), capping retention periods (12 months, not perpetual), requiring synthetic-data alternatives, and demanding audit rights for AI training data handling separate from standard financial audits. They detail emerging best practices like requiring vendors to certify in writing that customer data won't be used in third-party models, defining minimum group sizes for 'aggregated' data (100+ records to prevent reverse-engineering), and excluding high-risk data like support tickets entirely. Major vendors like Salesforce and Microsoft have published data processing addenda addressing AI training, but language remains vague on fine-tuning versus foundation models. For mid-market companies without data science resources, the battle shifts to restrictive scope and NIST-compliant de-identification standards. The EU AI Act and emerging state privacy laws in California and Colorado add pressure, but most protection still comes from private contract negotiation. The episode emphasizes that procurement must move from reactive to proactive - developing a standard AI data-use clause before vendor proposals arrive.

Key takeaways

  • →Perpetual data-use licenses for AI training are negotiable, even with major vendors claiming they're standard - the financial services firm in the episode removed perpetual rights and capped retention to 12 months after three rounds of negotiation.
  • →Define 'aggregated data' with a minimum group size (100+ records) to prevent reverse-engineering, and narrow 'to improve the service' language to 'solely to provide the specific services purchased' to block vendors from using your data for competing products.
  • →Support ticket data is especially high-risk and should be excluded entirely from AI training rights or governed by a separate data processing agreement with 90-day deletion requirements, as vendors have been known to train competitor chatbots on this data.
  • →Audit rights specifically for AI training data - separate from financial and security audits - are emerging as a new standard, including annual reviews of vendor data processing records and SOC 2 Type II reports covering AI training handling.
  • →Synthetic data generated by the customer offers a compromise solution that satisfies vendor needs for model training without exposing raw customer data, though it requires in-house data science capability.

Topics in this episode

synthetic dataMicrosoftSalesforceEU AI ActNIST frameworkSOC 2 Type IIAI model trainingAnonymized dataCRM contractsdata processing addenda (DPA)AI training data rightsFortune 500 procurementsoftware contract negotiationvendor ai data usedata addendum

Questions this episode answers

Can a vendor still use my company's data to train AI models after the contract ends if they claim perpetual rights?

Once a perpetual license is granted, the vendor can never be forced to delete the data or revoke those rights, even after contract termination. The leading procurement practice is to tie all data-use rights to the contract term, requiring certified deletion of customer data from training sets when the agreement ends.

What's the difference between 'aggregated' and 'anonymized' data, and which is safer for AI training?

Aggregated data combined at granular levels (e.g., by department) can still be reverse-engineered and re-identify individuals, so procurement teams now require a minimum group size of 100+ records to qualify as aggregated. True anonymization according to standards like NIST provides stronger protection but is harder to verify.

Can a vendor use support ticket data for AI training without explicit permission?

Under broad 'improve the service' language, vendors often claim the right to use support data for AI training, and at least one tech company discovered their vendor was training competitor chatbots on their support tickets. Procurement should explicitly exclude support data or require a separate agreement with 90-day deletion requirements.

What happens if a vendor has already trained an AI model on my data before I can enforce a deletion clause?

Deleting source data doesn't un-train an already-trained model, which is why procurement teams now require vendors to certify in writing that no customer data was used in models sold or licensed to third parties, and are adding audit rights to verify this.

Do larger vendors like Salesforce and Microsoft already have AI training data clauses in place?

Both have published data processing addenda addressing AI training, but the language is often vague - for example, Microsoft excludes training on 'foundation models' but the boundary between fine-tuning and new foundation models remains blurry, requiring procurement to negotiate clarification.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

16 / 20

The episode densely packs specific, actionable contract language concerns (perpetual licenses, aggregated data definitions, support ticket restrictions) and real negotiation tactics that a procurement operator would not have systematized. However, it occasionally retreats into general advice ('be proactive,' 'pull in legal early') that dilutes the density slightly.

Once you grant a perpetual license, you can never revoke it. Even if you terminate the contract.
Vendors often claim that aggregated data is no longer customer data. But if the aggregation is done at a granular level, say, by department, it can still be revealing.

Originality

15 / 20

The framing of AI training data as a distinct procurement battleground is relatively fresh and contrarian to standard contract-negotiation discourse. However, the underlying concepts (data protection, audit rights, contract carve-outs) are not novel; the originality lies in connecting them to AI rather than inventing new principles.

It's about data - specifically, whether the vendor gets to use your company's data to train its AI models.
It's emerging. I talked to a procurement VP at a healthcare company who said they now have a standard 'AI Training Data Addendum' that includes a right to review the vendor's data processing records annually.

Guest Caliber

12 / 20

The named sources are credible (procurement VP at healthcare, procurement director at top-20 financial services, referenced Salesforce/Microsoft) but they appear anonymized or paraphrased rather than direct interviews. Lucas demonstrates hands-on procurement knowledge but he is the host, not a guest. The episode lacks a true outside expert being interviewed in depth.

I spoke with a procurement director at a large financial services company - top twenty by revenue.
I talked to a procurement VP at a healthcare company who said they now have a standard 'AI Training Data Addendum'

Specificity & Evidence

14 / 20

The episode includes concrete examples (Salesforce, Microsoft, NIST framework, SOC 2 Type II, 90-day retention, 100-record minimum for aggregation) and real scenarios (financial services CRM renewal, tech company support data abuse), but most case studies are referenced without company names or detailed metrics, limiting verifiability and depth.

After three rounds of negotiation, they carved out an exception: the vendor could use aggregated and anonymized data, but only for specific model training purposes, and with a data retention limit of 12 months.
you can limit the data use to 'production data' only, excluding test data, customer personal information, or financial records. And you can require that the data be de-identified according to a specific standard, like the NIST framework.

Conversational Craft

13 / 20

Luna asks clarifying follow-ups that move the conversation forward logically ('And was it actually non-negotiable?' 'But how do you enforce that?'), demonstrating genuine engagement. However, the exchange feels somewhat rehearsed and lacks pushback or disagreement - Luna rarely challenges Lucas's framing, and the conversation is friendly rather than probing.

But how do you enforce that? If the model is already trained on that data, deleting the source data doesn't un-train the model.
So you want to define 'aggregated' as something that cannot be reverse-engineered. That means a minimum group size, maybe 100 or more records.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

data53lucas21luna20vendor17training16customer12procurement11models8financial7contract7support7model6rights6train5perpetual5firm5

Episode notes

Episode 84 of Enterprise Tech with Fexingo dives into a rapidly evolving negotiation point: how procurement teams at Fortune 500 companies are handling vendor demands for access to customer data to train AI models. Lucas and Luna unpack the specific case of a major financial services firm that pushed back on a clause granting a software vendor broad rights to use its transaction data for model training. They explore the legal and technical nuances of data use restrictions, the rise of 'AI training data addenda,' and why procurement is now sitting at the table with legal and data privacy teams. The episode also covers how vendors like Salesforce and Microsoft are framing these requests, and what negotiators need to watch for in contract language around anonymization, aggregation, and perpetual rights. Listeners will walk away with a concrete framework for protecting their company's data without derailing vendor relationships.

Full transcript

10 min

Transcribed and scored by The B2B Podcast Index.

Lucas: So I have been digging into something that is quietly becoming one of the most contentious clauses in enterprise software contracts. It's not about price per seat. It's not about uptime guarantees. It's about data - specifically, whether the vendor gets to use your company's data to train its AI models.

Luna: And I'm guessing most companies don't realize that's in their contracts until it's too late. Lucas: Exactly. I spoke with a procurement director at a large financial services company - top twenty by revenue. In March this year, they were renewing a multiyear CRM contract with a major vendor.

And buried in the fine print was a new data use addendum. It granted the vendor a 'worldwide, perpetual, royalty-free license to use, reproduce, modify, and analyze Customer Data for the purpose of improving and training machine learning models.' Luna: Perpetual? That's a big ask.

And for a financial firm, that data is incredibly sensitive. If that includes transaction data or customer behavior patterns, that's a regulatory nightmare. Lucas: Right. And look, if these conversations are useful for what you're building or running, honestly, if today's episode was worth a coffee to you, that's the link - buy me a coffee dot com slash fexingo.

It's listener support that keeps us ad-free and digging into stuff like this. Luna: Yeah, it's a small gesture that adds up. And it lets us keep going deep on procurement topics that don't get covered elsewhere. Lucas: So back to this contract.

The financial services firm's legal team flagged it immediately. But procurement had to step in because the vendor's standard response was 'this is non-negotiable, all our enterprise customers sign it.' Luna: And was it actually non-negotiable? Lucas: It was not.

After three rounds of negotiation, they carved out an exception: the vendor could use aggregated and anonymized data, but only for specific model training purposes, and with a data retention limit of 12 months. No perpetual rights. No raw transaction data. Luna: So what changed?

Was it just leverage, or did the vendor actually have a business reason to back down? Lucas: The vendor's business reason was that they wanted to improve their predictive models for sales forecasting. But the financial firm argued that their data was too idiosyncratic - it wouldn't generalize well to other customers anyway. Plus, the firm offered to contribute synthetic data generated from their own systems, which actually worked better for the vendor.

Luna: That's a clever compromise. But not every company has a data science team that can generate synthetic data. What about a mid-sized manufacturer with fewer resources? Lucas: Good question.

For those companies, the negotiation tends to focus on scope. For example, you can limit the data use to 'production data' only, excluding test data, customer personal information, or financial records. And you can require that the data be de-identified according to a specific standard, like the NIST framework. Luna: And what about the 'perpetual' piece?

That seems like the biggest red flag. Lucas: Absolutely. Once you grant a perpetual license, you can never revoke it. Even if you terminate the contract.

So procurement teams are now pushing for data use rights to be tied to the term of the agreement. When the contract ends, the vendor must delete all customer data from their training sets. Luna: But how do you enforce that? If the model is already trained on that data, deleting the source data doesn't un-train the model.

Lucas: That's the hard part. Right now, the leading practice is to require that the vendor certify in writing that no customer data was used in training models that are then sold or licensed to third parties. And some companies are demanding audit rights specifically for data use - separate from financial audits. Luna: That's a new one.

I've seen audit rights for security and for license compliance, but for data training use? Lucas: It's emerging. I talked to a procurement VP at a healthcare company who said they now have a standard 'AI Training Data Addendum' that includes a right to review the vendor's data processing records annually. They also require that the vendor use a third-party auditor like a SOC 2 Type II report specifically covering AI training data handling.

Luna: So this is becoming a whole new sub-discipline within procurement. How are vendors reacting? Lucas: Reactions vary. Some, like Salesforce and Microsoft, have published data processing addenda that address AI training.

But the language is still vague in places. For example, Microsoft's standard DPA says they won't use customer data to train 'foundation models,' but what about fine-tuning existing models? The line is blurry. Luna: And what about the smaller ai native vendors?

Are they more flexible or less? Lucas: It depends. Startups often have less leverage, so they may be more willing to negotiate. But they also have fewer resources to build custom data handling pipelines.

I heard of one case where a startup agreed to train a separate model instance for a large customer, just to avoid mixing data. That's the ultimate solution: a dedicated model. Luna: But that's expensive. Only the biggest customers can demand that.

Lucas: Right. For most companies, the battleground is contract language. And there are a few key phrases to watch for. One is 'aggregated data' - vendors often claim that aggregated data is no longer customer data.

But if the aggregation is done at a granular level, say, by department, it can still be revealing. Luna: So you want to define 'aggregated' as something that cannot be reverse-engineered. That means a minimum group size, maybe 100 or more records. Lucas: Exactly.

And another phrase is 'to improve the service.' Vendors use that to justify all sorts of data use. But 'improve the service' could mean anything from bug fixes to training a new AI that competes with your own products. So procurement teams are narrowing that to 'solely to provide the specific services purchased by Customer under this agreement.'

Luna: And what about data from support tickets? That's a gold mine for AI training. Lucas: That's a huge one. Support tickets often contain detailed technical information and even proprietary code snippets.

I know of a tech company that discovered their vendor was using support ticket data to train a chatbot that they then sold to competitors. The company had no idea. Luna: So how do you prevent that? Do you need to explicitly exclude support data from AI training?

Lucas: Yes. And ideally, you segment it. Some companies now have a separate support data processing agreement that prohibits any use beyond resolving the specific ticket. And they require deletion of support data after a short period, like 90 days.

Luna: Let's talk about the regulatory angle. Are there any new laws that help procurement here? Lucas: A few. The EU AI Act, which came into force in stages starting this year, requires transparency about training data sources.

But it's not specifically about customer data rights. In the US, there's no federal law yet, but states like California and Colorado are pushing data privacy bills that indirectly affect AI training. Luna: So it's mostly private contracting. And that means procurement has to be proactive.

Lucas: Exactly. I think the key takeaway for today is: if you're in procurement, you need to have a standard AI data use clause ready. Don't wait for the vendor to put theirs in. And include definitions for 'customer data,' 'aggregated data,' 'anonymized data,' and 'model training.'

Otherwise, you're negotiating from a reactive position. Luna: And if you're on the other side - a vendor - what's the best way to handle this without scaring customers? Lucas: Transparency. Be explicit about what data you need and why.

Offer opt-in models. Some vendors now give customers a choice: 'basic' service where no data is used for training, and 'enhanced' service where data is used but with additional benefits like customized models. That way, it's a trade, not a takeaway. Luna: I like that.

It turns a potential conflict into a value proposition. Lucas: Yeah. And ultimately, that's the goal of procurement - not just to protect, but to create value. The financial services firm I mentioned ended up with a better contract and a better relationship with the vendor because they negotiated in good faith.

The vendor got access to some data, the customer protected the rest, and both walked away satisfied. Luna: That's the ideal outcome. But it takes preparation. So for next renewal cycle, add 'AI training data rights' to your checklist.

Lucas: Absolutely. And maybe pull in your legal and data privacy teams early - not when the contract is already on the table.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Enterprise Software Deals Now Include a Vendor AI Model Explainability MandateB2B SaaS Talks with Fexingo · on EU AI Act94 / 100
  • Why revenue growth breaks down (and how great companies fix it) with Dan Bernoske If Prices Could Talk · on Salesforce89 / 100
  • Why B2B Brands Fail at Account Based Marketing AttributionThe Marketing Operator Podcast with Fexingo · on Salesforce88 / 100
  • #410 - How Mazy Dar found room in Google and Microsoft's market - and won the world's biggest banksThe Remarkable SaaS Podcast · on Salesforce87 / 100
  • How to Sell Against a Competitor Already in the BuildingSales Leadership with Fexingo · on Salesforce85 / 100
  • Why B2B Brands Are Using AI for Account PrioritizationThe Growth Operator with Fexingo · on Salesforce84 / 100

More from Enterprise Tech with Fexingo

All episodes →
  • How Fortune 500s Protect Against SaaS Vendor Failure with Escrow81 / 100
  • Why Fortune 500s Now Demand AI Model Red-Teaming90 / 100
  • How Fortune 500s Negotiate Vendor AI Hallucination Insurance88 / 100
  • How Fortune 500s Use Procurement to Negotiate Vendor Software Liability Caps92 / 100
  • How Fortune 500s Negotiate Software Beta Test Terms90 / 100
Explore the best B2B Sales podcasts →
All Enterprise Tech with Fexingo episodes →